A Flake-Only Retry Loop for Agents

When an agent sees the same test go red only intermittently in CI or locally, it tends to re-run the whole suite or patch product code at random. A flake test agent loop is not “fail → always patch.” It is a procedure to re-run only the suspected flake, within a hard limit, in an isolated retry, then decide whether to file an issue.

Adjacent posts stay on other axes. agent-test-gate is a pass/stop barrier on exit codes. agent-tdd-loop is an intentional red→green writing loop. This post covers only retry limits · isolation · when to file. No pricing, plans, or affiliate links. Grounded in pytest selection / rerun practice, Jest retryTimes, and GitHub Actions re-run patterns.

What retry limit?

One-line answer: Cap flake retries at the same command, same seed when possible, max N (team default 2–3). If it still fails inside N, leave the patch loop and take the issue / quarantine path. “Until green” is not a limit.

ItemPreferForbidden
Max retriesLocal agent 2; CI quarantine job 2–3Infinite / “until it passes”
TargetOnly the failed test id/fileFull-suite re-run
BackoffShort fixed backoff (or none)Random sleeps that hide races
Seed / envSame seed, same worker countChanging env each attempt for a lucky pass
On passLog flake-pass-on-retry k/NSilent Done
After N failsNo product patch → issue draftSpeculative body edits

Instruction snippet:

Flake retry policy (this task):
- Only re-run the failed test id(s) listed below — not the full suite
- Max retries: 2 (total runs = 1 initial + 2 retries)
- Same command flags / same cwd / do not change seed unless asked
- On pass after retry: report flake-pass-on-retry and STOP editing product code
- On fail after max retries: do NOT patch; draft an issue (see quarantine rules)
- Never treat “retry until green” as success criteria

Practice tips:

  1. Put the number in a Skill/Rule. “Try a bit more” lets agents exceed N.
  2. If CI already has a retry plugin, do not stack agent retries on top silently—log who ran how many times.
  3. One line vs agent-test-gate: the gate still must pass to proceed; the flake loop is a separate policy for intermittent failures, not a looser gate.

How to isolate?

One-line answer: Flake retries run only the failing case in an environment that reduces shared state and parallel interference. Isolation means narrowing scope, parallelism, and external deps—not moving files to another folder.

SlotDoDon’t
Selective run-k / --testPathPattern / single node idWhole-package test
Parallelismworkers=1 or serialize that fileMax workers on a known flake
Shared resourcesTemp schema, unique port, unique tmpdirGlobal /tmp/fixed or shared accounts
TimeFake clock / explicit timeouts onlyLonger wall-clock sleeps
NetworkRecorded fixtures / mocksHit real servers “one more time”
Agent scopeLogs, repro command, issue draftDrive-by refactors “to make retry pass”

Command patterns:

# Examples — pick the runner you already use
pytest path/to/test_foo.py::test_bar -vv --maxfail=1
npm test -- --runTestsByPath path/to/foo.test.ts --runInBand
cargo test test_bar -- --exact --nocapture
go test ./pkg -run '^TestBar$' -count=1

# Agent must print before each retry:
# attempt=k/N · cwd=… · command=…

Instruction snippet:

Isolation rules for flake retry:
1. Re-run ONLY: <test ids>
2. Prefer serial / single-worker for this retry loop
3. Do not start unrelated services; use existing mocks/fixtures
4. Do not edit product source to “make the retry pass”
5. Capture full stdout/stderr per attempt into one log blob

One line vs agent-tdd-loop: TDD intends failure and patches minimally to green; flake isolation reproduces and confirms without product patches. A green retry without a product fix is not “fix complete.”

When to file an issue?

One-line answer: After N failed retries, or when green appears only via retry, stop product patches and open an issue (or quarantine label). Agent Done means repro + issue draft, not “I fixed it.”

ObservationActionDone means
Fail once → all retries failTreat as hard fail; end flake loopFailure log + minimal repro (hand off to TDD/gate)
Fail then pass only on retryOpen flake-candidate issueTitle, repro, attempt log, suspected axis
Same test flake-passes k×/weekQuarantine list / separate CI jobflake / quarantine + owning team
Other tests wobble during retryDo not widen scope; separate ticketLink “additional flake found” only
Agent wants a speculative patchRefuse; hypothesis in the issue onlyExit with no product diff

Issue draft template:

Title: [flake] <test id> intermittent failure
Body:
- Repro command: …
- Attempts: fail, fail, pass (or fail×N)
- Env: cwd, runner, workers, seed (if any)
- Suspected axis: parallel | timing | network | shared fs | order-dependent
- Product patch: none in this loop (by policy)
- Next: quarantine / owner / harden test isolation
Labels: flake, testing

Ops rules:

  1. flake-pass-on-retry is not a merge pardon. Note it on the PR; do not disable required gates (agent-test-gate).
  2. Quarantine is temporary. Skip without issues erodes suite trust.
  3. “Just fix it, no issue” leaks into the TDD loop and invites unrelated refactors—state no product patch in the brief.

Bottom line: Flake loops are failed case only · capped N · isolated re-run, ending in issue/quarantine—not a softer gate and not a TDD patch cycle.

FAQ

Is a CI retry plugin enough?

Auto-retry can hide signal. The agent loop records who / how many / which command, and repeated flakes become issues. Plugin ≠ policy.

How is this different from agent-test-gate / agent-tdd-loop?

agent-test-gate is the pass/stop barrier. agent-tdd-loop is intentional red→green writing. This post is the flake-only retry loop: limits, isolation, and filing criteria.

Isn’t N=1 (no retry) better?

For new, deterministic failures, N=0–1 is right. Flake policy applies only to cases already classed as intermittent. Retrying every red collapses the gate.

If a retry goes green, can we skip the issue?

Default no—open one (or enqueue quarantine). “Passed once” still signals unreproducible risk. Silent Done is forbidden.

Sources

  • pytest documentation — selective runs, failure output, rerun practice
  • Jest — retryTimes — per-test retry API
  • GitHub Actions Docs — job re-runs and failure signals in CI
  • Adjacent: agent-test-gate (pass/stop barrier), agent-tdd-loop (write red→green), agent-eval-checklist (pre-merge eval)