Put a Test Gate on Agent Output
Agents can stack patches quickly, but moving on while tests are still red only burns more cost. An agent test gate is not “a human scores quality items”—it is a mechanical barrier that passes or stops on command exit codes.
This post covers only where to check? · stop on fail? · overlap with CI?. agent-tdd-loop is the red→green write-loop axis fixed in prompts; agent-eval-checklist is the pre-merge human (or CODEOWNERS) eval items axis. Here we only cover the gate (stop/pass) axis. No pricing, plans, tokens, or affiliate links.
Where to check?
One-line answer: Place gates right after an agent turn (local/hooks) and before/after a PR opens (CI required checks). “The agent said it passed in chat” is not a gate.
Candidate check locations:
| Location | What to run | Best when |
|---|---|---|
| Editor / agent hooks | Smoke · unit subset, lint, typecheck | Cutting a turn quickly |
| Local pre-commit / pre-push | Fast suite + format | Final block before a human push |
| Agent Done definition | “Named command exits 0” fixed in Skill/Rule | Letting the loop retry itself |
| CI required check | Agreed suite · build · contract tests | Official gate in front of merge |
| Merge queue / protected branch | Merge only when required status is green | Blocking force-around pushes |
Practical placement order:
- Put the cheapest checks on agent hooks/Done (path-scoped unit tests,
tsc --noEmit, lint). - Put medium cost on pre-push or required PR CI jobs.
- Keep expensive E2E/integration in CI only—not on every agent turn.
- Write the gate command so it is reproducible in one line. “Just run the tests” is not a gate.
Copy-paste Done gate:
Gate (must exit 0 before claiming done):
<exact command, e.g. npm test -- path/to/spec>
Do not open PR / do not start next task while gate is red.
Boundary: agent-tdd-loop is how to use that command in a red→green cycle; this section is which layer gets the barrier.
Stop on fail?
One-line answer: Stop when the gate command is non-zero. Do not allow the next feature, next file, PR open, or “commit for now.” Retries use the same gate command only; for flakes, do not raise timeouts—isolate, reproduce, ticket.
Stop policy table:
| Signal | Default action | Exception |
|---|---|---|
| Unit/smoke red | End agent turn · hold the patch | None (inside gate scope) |
| Lint/type red | Stop, then minimal fix | Warn-only for generated code only via a documented exception list |
| Unrelated suite red | Report only; do not fix | Out-of-gate → separate ticket |
| Suspected flake | Retry same command up to team limit N | Over limit → stop + flake issue |
| Secret/auth scan red | Immediate stop · no merge | No exceptions |
Stop rules for the agent:
On gate failure:
1. Stop. Do not start unrelated edits.
2. Re-run the exact gate command.
3. If still red: minimal patch OR report blocker (log + command).
4. Never mark done, open PR, or push while gate is red.
5. Flaky: max <N> retries; then open flake ticket — do not weaken asserts.
Stop ≠ full eval: agent-eval-checklist items for docs/secrets may still need a human pass. The test gate only owns mechanical exit codes. Gate green with eval red still means do not merge.
Overlap with CI?
One-line answer: Overlap is a problem when roles collide, not merely when the same command runs twice. Local/hooks give fast feedback and agent stop; CI is the shared truth and merge authority. You need both, but split suite size and required status.
Role split:
| Layer | Purpose | How to reduce waste |
|---|---|---|
| Agent gate | Cut the turn early | Impact-path subset, cacheable commands |
| Local pre-* | Block mistakes before push | Call the same entrypoint script as CI |
| CI required | Official merge block record | Same scripts/test-gate.sh as local |
| CI optional | Slow or flaky jobs | Not required; do not block merge |
Overlap design checklist:
[ ] Gate command pinned to one repo script?
[ ] Agent Done / hooks / CI required call the same entrypoint?
[ ] Full E2E not on every agent turn?
[ ] Required job list matches CODEOWNERS / branch protection?
[ ] “Local green, CI red” → is the local gate weaker? (env, seed, deps)
[ ] Not skipping eval/secret tables just because CI is green?
Common failure: agent runs an npm test subset while CI uses another package manager or Node version → the gate lies. Align entrypoint and runtime.
One-liner: Test gate = where to place it / stop on red / same script as CI for truth. Different axis from the TDD loop and merge eval.
FAQ
How is this different from agent-tdd-loop?
That post is the write loop: hand failing logs and finish one red→green cycle. This post is where to bolt the success condition (command exit 0) onto hooks, Done, and CI, plus stop policy.
How is this different from agent-eval-checklist?
The eval checklist is a human pass on feature/security/docs items. The test gate is a mechanical pass on automated tests/lint/types. Gate green ≠ eval pass.
What if the agent turns the gate off?
State in Rule/Skill: no changing, skipping, or weakening the gate command/asserts. Exceptions need a ticket ID and owner approval.
Same gate for every package?
In a monorepo, gate changed packages + contract tests; lift the full workspace to CI required. Full workspace every turn is usually too much.
Sources
- Team practice: bind agent Done, hooks, and CI required to one gate script — tables above
- Adjacent axes: agent-tdd-loop (red→green while writing), agent-eval-checklist (pre-merge human eval), multi-agent-pr-workflow (branch/merge gates)