A Flake-Only Retry Loop for Agents
When an agent sees the same test go red only intermittently in CI or locally, it tends to re-run the whole suite or patch product code at random. A flake test agent loop is not “fail → always patch.” It is a procedure to re-run only the suspected flake, within a hard limit, in an isolated retry, then decide whether to file an issue.
Adjacent posts stay on other axes. agent-test-gate is a pass/stop barrier on exit codes. agent-tdd-loop is an intentional red→green writing loop. This post covers only retry limits · isolation · when to file. No pricing, plans, or affiliate links. Grounded in pytest selection / rerun practice, Jest retryTimes, and GitHub Actions re-run patterns.
What retry limit?
One-line answer: Cap flake retries at the same command, same seed when possible, max N (team default 2–3). If it still fails inside N, leave the patch loop and take the issue / quarantine path. “Until green” is not a limit.
| Item | Prefer | Forbidden |
|---|---|---|
| Max retries | Local agent 2; CI quarantine job 2–3 | Infinite / “until it passes” |
| Target | Only the failed test id/file | Full-suite re-run |
| Backoff | Short fixed backoff (or none) | Random sleeps that hide races |
| Seed / env | Same seed, same worker count | Changing env each attempt for a lucky pass |
| On pass | Log flake-pass-on-retry k/N | Silent Done |
| After N fails | No product patch → issue draft | Speculative body edits |
Instruction snippet:
Flake retry policy (this task):
- Only re-run the failed test id(s) listed below — not the full suite
- Max retries: 2 (total runs = 1 initial + 2 retries)
- Same command flags / same cwd / do not change seed unless asked
- On pass after retry: report flake-pass-on-retry and STOP editing product code
- On fail after max retries: do NOT patch; draft an issue (see quarantine rules)
- Never treat “retry until green” as success criteria
Practice tips:
- Put the number in a Skill/Rule. “Try a bit more” lets agents exceed N.
- If CI already has a retry plugin, do not stack agent retries on top silently—log who ran how many times.
- One line vs agent-test-gate: the gate still must pass to proceed; the flake loop is a separate policy for intermittent failures, not a looser gate.
How to isolate?
One-line answer: Flake retries run only the failing case in an environment that reduces shared state and parallel interference. Isolation means narrowing scope, parallelism, and external deps—not moving files to another folder.
| Slot | Do | Don’t |
|---|---|---|
| Selective run | -k / --testPathPattern / single node id | Whole-package test |
| Parallelism | workers=1 or serialize that file | Max workers on a known flake |
| Shared resources | Temp schema, unique port, unique tmpdir | Global /tmp/fixed or shared accounts |
| Time | Fake clock / explicit timeouts only | Longer wall-clock sleeps |
| Network | Recorded fixtures / mocks | Hit real servers “one more time” |
| Agent scope | Logs, repro command, issue draft | Drive-by refactors “to make retry pass” |
Command patterns:
# Examples — pick the runner you already use
pytest path/to/test_foo.py::test_bar -vv --maxfail=1
npm test -- --runTestsByPath path/to/foo.test.ts --runInBand
cargo test test_bar -- --exact --nocapture
go test ./pkg -run '^TestBar$' -count=1
# Agent must print before each retry:
# attempt=k/N · cwd=… · command=…
Instruction snippet:
Isolation rules for flake retry:
1. Re-run ONLY: <test ids>
2. Prefer serial / single-worker for this retry loop
3. Do not start unrelated services; use existing mocks/fixtures
4. Do not edit product source to “make the retry pass”
5. Capture full stdout/stderr per attempt into one log blob
One line vs agent-tdd-loop: TDD intends failure and patches minimally to green; flake isolation reproduces and confirms without product patches. A green retry without a product fix is not “fix complete.”
When to file an issue?
One-line answer: After N failed retries, or when green appears only via retry, stop product patches and open an issue (or quarantine label). Agent Done means repro + issue draft, not “I fixed it.”
| Observation | Action | Done means |
|---|---|---|
| Fail once → all retries fail | Treat as hard fail; end flake loop | Failure log + minimal repro (hand off to TDD/gate) |
| Fail then pass only on retry | Open flake-candidate issue | Title, repro, attempt log, suspected axis |
| Same test flake-passes k×/week | Quarantine list / separate CI job | flake / quarantine + owning team |
| Other tests wobble during retry | Do not widen scope; separate ticket | Link “additional flake found” only |
| Agent wants a speculative patch | Refuse; hypothesis in the issue only | Exit with no product diff |
Issue draft template:
Title: [flake] <test id> intermittent failure
Body:
- Repro command: …
- Attempts: fail, fail, pass (or fail×N)
- Env: cwd, runner, workers, seed (if any)
- Suspected axis: parallel | timing | network | shared fs | order-dependent
- Product patch: none in this loop (by policy)
- Next: quarantine / owner / harden test isolation
Labels: flake, testing
Ops rules:
- flake-pass-on-retry is not a merge pardon. Note it on the PR; do not disable required gates (agent-test-gate).
- Quarantine is temporary. Skip without issues erodes suite trust.
- “Just fix it, no issue” leaks into the TDD loop and invites unrelated refactors—state no product patch in the brief.
Bottom line: Flake loops are failed case only · capped N · isolated re-run, ending in issue/quarantine—not a softer gate and not a TDD patch cycle.
FAQ
Is a CI retry plugin enough?
Auto-retry can hide signal. The agent loop records who / how many / which command, and repeated flakes become issues. Plugin ≠ policy.
How is this different from agent-test-gate / agent-tdd-loop?
agent-test-gate is the pass/stop barrier. agent-tdd-loop is intentional red→green writing. This post is the flake-only retry loop: limits, isolation, and filing criteria.
Isn’t N=1 (no retry) better?
For new, deterministic failures, N=0–1 is right. Flake policy applies only to cases already classed as intermittent. Retrying every red collapses the gate.
If a retry goes green, can we skip the issue?
Default no—open one (or enqueue quarantine). “Passed once” still signals unreproducible risk. Silent Done is forbidden.
Sources
- pytest documentation — selective runs, failure output, rerun practice
- Jest — retryTimes — per-test retry API
- GitHub Actions Docs — job re-runs and failure signals in CI
- Adjacent: agent-test-gate (pass/stop barrier), agent-tdd-loop (write red→green), agent-eval-checklist (pre-merge eval)