Skip to content

test(watch): cap test's cadence precondition fails CI on loaded runners #397

Description

@dean0x

Problem

watch_debounce_cap_rebuilds_while_writes_never_stop (crates/mds-cli/tests/cli_watch.rs :1304) has a harness precondition: the writer thread must sustain a sub-200ms write cadence (largest inter-write gap < WINDOW) so that a rebuild observed mid-burst can be attributed to the debounce cap rather than a quiet period that ended on its own. The test retries this measurement up to 4 attempts, gated only on that precondition. When every attempt is inconclusive (the writer thread gets descheduled past the window on a loaded runner), the final assert! (:1379-1384) fails — which fails the required Rust — fmt, clippy, test check even though no product behaviour was exercised.

This is NOT one of the watch data-loss bugs (#317/#319/#321). Those were real silent-data-loss defects in the watch implementation. This is a harness/runner-load problem: the writer thread (a std::thread::sleep(5ms) loop) is not a real-time guarantee, and on a sufficiently loaded CI or local runner it can be descheduled long enough to blow the 200ms window, making the sample inconclusive rather than proving anything about the cap.

Occurrences (last week)

An earlier 747ms occurrence was noted in a prior comment — specifically the harness-precondition doc comment on this same test (introduced in #384: "a 747ms inter-write gap was observed on CI"), which is the source of the MAX_ATTEMPTS/window design in the first place.

Interim mitigation (this issue's companion PR)

  • Raises MAX_ATTEMPTS from 4 to 6.
  • On the last inconclusive attempt, the test now skips with a summary warning (drops the child process, prints a SKIPPED (inconclusive harness) line to stderr, and best-effort appends a warning line to $GITHUB_STEP_SUMMARY when set) instead of failing the required check.
  • Every behaviour assertion is unchanged and still fails hard on the first conclusive attempt — a real product regression is never retried or skipped away.

Follow-up (stays open)

This issue tracks root-causing the cadence problem on loaded runners, e.g.:

  • A dedicated/higher-priority writer thread or timing mechanism that isn't subject to ordinary scheduler jitter.
  • cargo-nextest test serialization/isolation so the writer thread doesn't compete with other test processes for CPU.
  • Identifying whether GitHub-hosted runner class/load correlates with the gap size, and whether a larger WINDOW or an entirely different cap-detection strategy (not wall-clock-gap-based) would remove the flake at the root.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions