AI Agent TDD Loop: Fix the Failing Test First
An AI agent TDD loop means you do not let the agent “patch code first, tests later.” You lock failing test (red) → minimal fix to pass (green) → optional cleanup (refactor) into prompts, Skills, and Rules. No pricing or personal reviews.
This is not a product-button walkthrough. It covers only how to feed failing logs, when one cycle is done, and how to stop infinite fix loops.
How do you feed failing test logs?
One-line answer: Prefer a single failing command’s stdout/stderr, failing test name, and scoped files over “run the whole suite and fix whatever.” Make the agent fix that failure only.
Minimum payload:
| Piece | Why |
|---|---|
| Repro command | Something re-runnable as-is: npm test -- path/to/spec, pytest -k name, cargo test foo |
| Failure output | Assertions, stack, exit code—not just “it failed” |
| Scope | Paths/tests allowed to change; not the whole repo |
| Non-goals | “Do not touch other flaky tests,” “no new dependencies” |
Practical handoff:
- Reproduce locally first, then paste the terminal output or
@-attach the log file. - Put the target test name and success re-run command at the top of the prompt.
- Say explicitly: report unrelated failures; do not fix them in this cycle.
- Put “always: failing log → minimal patch → re-run same command” in a Skill/
SKILL.md. Keep short norms (“never claim done without the test”) in a Rule.
Minimal prompt skeleton:
Failing command (re-run exactly):
<paste command>
Output:
<paste failing stdout/stderr>
Scope: only make <test-name> pass.
Do not change unrelated tests or add dependencies.
Done when: the same command exits 0.
If the log is huge, keep the first assertion, last stack frames, and file:line; drop noise. Never drop the repro command and exit condition.
What counts as success for one cycle?
One-line answer: One cycle ends when the named failing test(s) pass under the same command, the diff stays in agreed scope, and only then (optionally) a small refactor. “I changed something” is not success.
Cycle definition for agents:
| Phase | Agent job | Done signal |
|---|---|---|
| red | Confirm the log/command; 1–2 cause hypotheses | One sentence: why this test fails |
| green | Minimal patch so that test passes | Same repro command exits 0 |
| refactor | Optional cleanup without behavior change | Same command still green; no out-of-scope files |
Done checklist:
- Repro command greened — the user-given (or Skill-fixed) command succeeds.
- Scope held — diff limited to agreed paths around that failure.
- Explainable — short note: what changed and why the failure went away.
- Next cycle separate — new failures/features become a new red, not the same turn.
Example Done-when block for a Skill:
## Done when
- Re-run the exact failing command; exit code 0.
- Diff touches only files needed for that failure.
- Report: root cause (1–2 sentences) + files changed.
- If still failing after N attempts, stop and ask (see stop rules).
Also lock: no refactor before green, and no “pass” via skip or deleted assertions.
How do you stop infinite fix loops?
One-line answer: Hard-stop on attempt count, repeated identical errors, and scope drift; use stop hooks / loop_limit and prompt “call a human after N” so auto-followups cannot run forever.
Stop if any of these fire:
| Signal | Action |
|---|---|
| Same failure signature / same file churn N times (e.g. 3) | Stop; hand log + hypotheses to a human |
| Patches spawn new failures while the original remains | Rollback / shrink scope; no new cycle |
| Skip tests, delete assertions, only bump timeouts | Immediate stop (deny list) |
| Repro command changes (different suite/flags) | Not “done”; re-verify |
Stop-rules snippet:
Stop rules:
- Max 3 patch→re-test attempts for this failure.
- If the error signature is unchanged after 2 attempts, stop and summarize.
- Never mark done by skipping tests or deleting assertions.
- Do not start a new feature while this red is open.
Tool assists:
- Cursor hooks
stop/subagentStopwithloop_limitcap automatic follow-ups (official Hooks docs). - Feed agents a single failing command; let humans/CI own the full suite.
Most infinite loops come from guessing without a log and fuzzy done criteria. Fix the log bundle and Done-when first, then add attempt caps.
Wrap-up
An AI agent TDD loop is red→green→refactor locked as inputs (failing logs), done criteria, and stop rules—not slogans. Feed the exact failing command, treat same-command exit 0 as one-cycle success, and cut loops with N attempts, identical signatures, and no-skip rules. Put procedures in Skills; short bans in Rules.
Sources
- Cursor Agent Skills — lock multi-step procedure in
SKILL.md - Cursor Rules — short norms and bans
- Cursor Hooks —
stop/loop_limitfor auto-loop caps