Put a Test Gate on Agent Output

Agents can stack patches quickly, but moving on while tests are still red only burns more cost. An agent test gate is not “a human scores quality items”—it is a mechanical barrier that passes or stops on command exit codes.

This post covers only where to check? · stop on fail? · overlap with CI?. agent-tdd-loop is the red→green write-loop axis fixed in prompts; agent-eval-checklist is the pre-merge human (or CODEOWNERS) eval items axis. Here we only cover the gate (stop/pass) axis. No pricing, plans, tokens, or affiliate links.

Where to check?

One-line answer: Place gates right after an agent turn (local/hooks) and before/after a PR opens (CI required checks). “The agent said it passed in chat” is not a gate.

Candidate check locations:

LocationWhat to runBest when
Editor / agent hooksSmoke · unit subset, lint, typecheckCutting a turn quickly
Local pre-commit / pre-pushFast suite + formatFinal block before a human push
Agent Done definition“Named command exits 0” fixed in Skill/RuleLetting the loop retry itself
CI required checkAgreed suite · build · contract testsOfficial gate in front of merge
Merge queue / protected branchMerge only when required status is greenBlocking force-around pushes

Practical placement order:

  1. Put the cheapest checks on agent hooks/Done (path-scoped unit tests, tsc --noEmit, lint).
  2. Put medium cost on pre-push or required PR CI jobs.
  3. Keep expensive E2E/integration in CI only—not on every agent turn.
  4. Write the gate command so it is reproducible in one line. “Just run the tests” is not a gate.

Copy-paste Done gate:

Gate (must exit 0 before claiming done):
  <exact command, e.g. npm test -- path/to/spec>
Do not open PR / do not start next task while gate is red.

Boundary: agent-tdd-loop is how to use that command in a red→green cycle; this section is which layer gets the barrier.

Stop on fail?

One-line answer: Stop when the gate command is non-zero. Do not allow the next feature, next file, PR open, or “commit for now.” Retries use the same gate command only; for flakes, do not raise timeouts—isolate, reproduce, ticket.

Stop policy table:

SignalDefault actionException
Unit/smoke redEnd agent turn · hold the patchNone (inside gate scope)
Lint/type redStop, then minimal fixWarn-only for generated code only via a documented exception list
Unrelated suite redReport only; do not fixOut-of-gate → separate ticket
Suspected flakeRetry same command up to team limit NOver limit → stop + flake issue
Secret/auth scan redImmediate stop · no mergeNo exceptions

Stop rules for the agent:

On gate failure:
1. Stop. Do not start unrelated edits.
2. Re-run the exact gate command.
3. If still red: minimal patch OR report blocker (log + command).
4. Never mark done, open PR, or push while gate is red.
5. Flaky: max <N> retries; then open flake ticket — do not weaken asserts.

Stop ≠ full eval: agent-eval-checklist items for docs/secrets may still need a human pass. The test gate only owns mechanical exit codes. Gate green with eval red still means do not merge.

Overlap with CI?

One-line answer: Overlap is a problem when roles collide, not merely when the same command runs twice. Local/hooks give fast feedback and agent stop; CI is the shared truth and merge authority. You need both, but split suite size and required status.

Role split:

LayerPurposeHow to reduce waste
Agent gateCut the turn earlyImpact-path subset, cacheable commands
Local pre-*Block mistakes before pushCall the same entrypoint script as CI
CI requiredOfficial merge block recordSame scripts/test-gate.sh as local
CI optionalSlow or flaky jobsNot required; do not block merge

Overlap design checklist:

[ ] Gate command pinned to one repo script?
[ ] Agent Done / hooks / CI required call the same entrypoint?
[ ] Full E2E not on every agent turn?
[ ] Required job list matches CODEOWNERS / branch protection?
[ ] “Local green, CI red” → is the local gate weaker? (env, seed, deps)
[ ] Not skipping eval/secret tables just because CI is green?

Common failure: agent runs an npm test subset while CI uses another package manager or Node version → the gate lies. Align entrypoint and runtime.

One-liner: Test gate = where to place it / stop on red / same script as CI for truth. Different axis from the TDD loop and merge eval.

FAQ

How is this different from agent-tdd-loop?

That post is the write loop: hand failing logs and finish one red→green cycle. This post is where to bolt the success condition (command exit 0) onto hooks, Done, and CI, plus stop policy.

How is this different from agent-eval-checklist?

The eval checklist is a human pass on feature/security/docs items. The test gate is a mechanical pass on automated tests/lint/types. Gate green ≠ eval pass.

What if the agent turns the gate off?

State in Rule/Skill: no changing, skipping, or weakening the gate command/asserts. Exceptions need a ticket ID and owner approval.

Same gate for every package?

In a monorepo, gate changed packages + contract tests; lift the full workspace to CI required. Full workspace every turn is usually too much.

Sources

  • Team practice: bind agent Done, hooks, and CI required to one gate script — tables above
  • Adjacent axes: agent-tdd-loop (red→green while writing), agent-eval-checklist (pre-merge human eval), multi-agent-pr-workflow (branch/merge gates)