Have Agents Reproduce Only Non-Flake Failures
When CI goes red, it is easy to tell an agent “reproduce the failure and fix it.” If that failure passes on retry of the same commit, or wobbles with isolation, time, remotes, or resources, the agent chases a phantom regression—or “fixes” a flake by raising timeouts. Agent repro of non-flake failures is not “run more tests.” It is a process that filters retryable signals, then reproduces and reports only deterministic product failures.
This post covers only how to filter flake signals · the minimal repro command set · evidence in the report. agent-test-gate owns the post-turn gate; agent-tdd-loop owns fix the failing test first. Here we only cover stripping flake before the repro ask. No pricing, affiliates, C++ samples, or invented pass rates.
Grounding: Martin Fowler — Eradicating Non-Determinism in Tests (quarantine and cause axes: isolation, async, remote, time, resource leaks), plus process/log heuristics used in CI: retries, history windows, co-failures, environment deltas, quarantine lists. Pass-rate numbers vary by team and tool—do not invent them; record what to compare.
How do you filter flake signals?
One-line answer: Treat as flake candidates first when you see same-SHA retry pass · quarantine list match · infra/env reason · co-failures with no change-set proximity. A retry-pass alone does not prove “the test is at fault”—infrastructure failures also pass on retry.
Filter order (process):
- Use retries as a detection signal, not a silent pass. Fail → re-run the same commit/job → pass ⇒ candidate for flake or infra. Do not leave only green; keep attempt count, per-attempt outcome, commit SHA, job URL.
- Check quarantine status. Fowler: keep non-deterministic tests out of the healthy suite. If the test is already on a quarantine/skip/known-flake list and the failure mode matches, route to the existing quarantine process—do not open a fresh “reproduce and fix product” task.
- Subtract environment/infra first. Runner image, container tag, cache, network, disk, permission errors that break before product asserts are not product repros. If CI summaries expose a reason/tag, follow that route.
- Look at a history window. Intermittent over recent runs vs a step drop from a date vs a red streak near one PR—read logs/dashboard history. Do not invent a percentage from “it fails a lot lately.”
- Combine co-failure with change-set proximity. Many fails in the same shard/suite ⇒ shared state/fixture candidate. Files near the call graph in this change ⇒ regression candidate. A lone intermittent fail with no diff proximity leans flake.
Prompt line for the agent:
Do not chase failures that: (a) pass on retry of the same SHA,
(b) match quarantine/known-flake list, (c) show infra/env reason
before product asserts, (d) have no change-set proximity and a
history of intermittent fails. Only reproduce deterministic product
failures. Record why each red was filtered.
What is the minimal repro command set?
One-line answer: Minimal set = only the failing test (or job) · fixed seed/env · one run, no retry · fixed log path. “Whole suite, many times” amplifies noise; it is not a repro.
Skeleton (swap runner names for your stack):
| Step | Purpose | Pass when |
|---|---|---|
| 1. Pin target | Failed test name / file / job from CI log | One-line id |
| 2. Pin SHA | git rev-parse HEAD equals failing CI commit | No other branch / dirty tree |
| 3. Single run | That test only, retries off | Non-zero exit recorded as-is |
| 4. Seed / time | Fix random seed, clock stub, TZ when possible | Fowler time/async axes |
| 5. Isolate | Minimize local cache and parallel workers | No shared DB/port fights |
| 6. Artifacts | stdout/stderr, JUnit/XML, artifact path | Files exist for the report |
Prompt-ready minimal set:
Repro (deterministic only):
1) cd <repo> && git checkout <failing-sha> && git status --porcelain # must be empty
2) export TZ=UTC CI=1 <SEED_ENV>=<fixed> # no live clock drift if wrap exists
3) <test-runner> --no-retry path/to/test::CaseName
4) On non-zero: save full log to artifacts/repro-<sha>-<case>.log
5) Do NOT: bump timeouts, add sleeps, re-run until green, edit quarantine list
6) If second identical command passes without code change → stop; label flake-candidate; do not "fix"
Do not: bare sleep for async (Fowler), hitch to live remotes that add non-determinism, quarantine and forget, or write invented pass rates like “passes M of N” without dashboard data.
What evidence belongs in the report?
One-line answer: The report needs verdict (flake-candidate / infra / deterministic-defect) · filter reasons · full repro commands · one-shot failure log link · SHA and job URL · attempt history. “Sometimes fails” is not evidence.
Evidence checklist:
- One-line verdict. One of
deterministic-defect|flake-candidate|infra/env|quarantined-known. If unclear, useunclassifiedand name the missing signal (history, co-failure, env). - Filter rationale. Retry outcomes, quarantine match, CI reason/tag, history pattern (intermittent vs step drop), diff proximity.
- Repro commands verbatim. Paste the minimal set. Do not let the agent “approximate.”
- One-shot failure artifacts. Log path, assertion/stack summary, JUnit fail node if any. Keep a passing retry log beside a failing one when labeling flake.
- Identifiers. Commit SHA, CI job/run URL, full test name, runner image/tag (for env delta).
- Next action. Deterministic ⇒ scoped fix; flake-candidate ⇒ quarantine + ticket + deadline (Fowler: quarantine, then fix soon); infra ⇒ platform issue. No timeout bumps pretending to be fixes.
Report sketch:
Verdict: deterministic-defect
Filtered out: none (same SHA retry still failed; not on quarantine; assert in product code)
SHA: abcdef1
Job: https://ci.example/jobs/12345
Repro:
git checkout abcdef1
<test-runner> --no-retry tests/foo::test_bar
Evidence: artifacts/repro-abcdef1-test_bar.log (assertion at foo.py:88)
Next: fix product path X; do not quarantine
Bottom line: Agent repro = filter flake/infra signals → one deterministic minimal command → report verdict, commands, logs, SHA. Aligns with agent-test-gate and agent-tdd-loop; this post is the front door that strips flake before the repro ask.
FAQ
Does a retry-pass always mean flake?
No. Infra and environment failures also pass on retry. Check CI reason/tag, runner/image delta, and co-failures first; only then label flake-candidate when product asserts wobble.
Is quarantine the end of the story?
No. Fowler: quarantine protects the healthy suite, but quarantined tests are work items to fix soon. Do not default the agent prompt to “quarantine if blocked.” Hand off only when list, owner, and deadline exist.
How is this different from agent-test-gate?
agent-test-gate is a post-turn pass/fail gate on agent output. This post is before you hand a red CI to an agent: filter flake/infra and pin repro scope and evidence.
Must the report include a pass rate?
Cite numbers that already exist on your team dashboard only. Do not invent pass-rate percentages. If none exist, record attempt counts, per-attempt outcomes, and the date window.
Sources
- Martin Fowler — Eradicating Non-Determinism in Tests — quarantine; isolation / async / remote / time / resource leaks
- CI process heuristics: treat retries as a detection signal, judge via history window, co-failure, env delta, quarantine list (from your CI logs and summary reasons)
- Adjacent: agent-test-gate (post-turn gate), agent-tdd-loop (failing test first)