Agent Output Eval Checklist

Agents can fill a PR, patch, and even commit messages quickly, but pre-merge eval stays a human (or CODEOWNERS) gate. Skimming only in “read order” still lets feature regressions, secret-adjacent edits, and thin docs/commit lines slip through. You need a per-axis output quality checklist.

This post covers only feature/regression? · security/secrets? · docs/commit messages?. agent-diff-review is the diff read-order axis (stat → intent → risk paths → test pairing). Here the focus is pass/fail eval items before merge. No pricing, plans, tokens, or affiliate links.

Feature and regression?

One-line answer: Against the ticket’s Done and out-of-scope, both new/changed behavior must be reproducible and core existing paths must stay green. “The agent said it passed” is not an eval result.

Pre-merge feature/regression eval:

ItemPass whenOn fail
Done mappingEvery Done line has evidence (command, UI, log path)Empty Done / claims only → hold merge
Out-of-scopeDeclared out-of-scope work is not in the diffDrive-by refactors / dependency bumps → revert or split PR
New pathsHappy path + at least one failure/edge scenarioHappy-only → hold or require follow-up ticket
RegressionTeam smoke / related suite greenRed or skip-to-fake-green → no merge
ContractsPublic API, schema, auth flag changes are named in DoneSilent breaking change → hold

Copy-paste:

[ ] Each Done item has an evidence path (command/CI/manual)?
[ ] Out-of-scope not mixed into the diff?
[ ] New behavior: happy + at least one failure/edge?
[ ] Related regression suite/smoke green? no skip / weakened asserts?
[ ] API/schema/auth changes named in Done?

Often missing in agent output: “fixed” with no repro command, flaky timeouts only raised, “behavior unchanged” with no suite. agent-tdd-loop is red→green while writing; this section is pass/fail on feature/regression at merge time.

Security and secrets?

One-line answer: Confirm diff, commits, logs, and issue bodies have no secrets, credentials, internal URLs, or PII, and do not merge privilege expansion, auth bypass, or loosened defaults without an explicit reviewer plus a rollback unit.

Security/secrets eval table:

SurfaceLook forPass
Secret mix-in.env*, keys/tokens/PEM, hardcoded passwords, pasted agent chatNone. If present: scrub history + rotation plan
Logs / screenshotsSample output, captures, CI logs with tokens/cookiesMasked or removed
AuthZDeleted checks, * allow, wider CORS/public ACLNamed in Done + owner approval
Supply chainLockfile, new deps, postinstall scriptsInstall/build/core tests + provenance
DataProd dumps, real-user PII fixturesSynthetic/anonymized only

Copy-paste:

[ ] No secrets/keys/tokens in diff, commits, or PR body?
[ ] Examples, logs, screenshots masked?
[ ] AuthZ / public ACL changes in Done with owner approval?
[ ] New deps / lockfile: reason + verification commands?
[ ] No real data / PII fixtures?
[ ] Suspected leak: no merge + key-rotation runbook

Note: This overlaps agent-diff-review’s “put risk paths in read order,” but here the point is pass/fail items and leak response. “We looked during review” alone is not an eval pass.

Docs and commit messages?

One-line answer: A human must still recover intent, out-of-scope, and how to verify six months later. Titles state what/why, bodies cover verification · risk · follow-ups, and docs (README, ADR, runbook) pair with behavior changes.

Docs/commit eval:

ItemPass whenCommon fail
Commit/PR titleScope readable in one line (module · behavior)fix, update, agent changes
BodyDone / Out-of-scope / verify commands + results“tests passed” with no evidence path
Risk one-linerSecrets, migration, flag, rollback unitBlank while risk paths changed
Doc pairingUser/ops change → README, runbook, API notes updatedCode only, docs stale
Issue linksTicket / follow-up IDs“later” with no ID

Copy-paste:

[ ] Title names module/behavior (not just fix/update)?
[ ] Body has Done / Out-of-scope / verification evidence?
[ ] Risk/rollback one-liner if the change is dangerous?
[ ] Doc pairing for behavior/ops changes?
[ ] Follow-ups have ticket IDs?
[ ] Do not replace the body with agent “self-review” prose alone

Boundary: agent-diff-review is in what order to read files; this section is whether merge history and onboarding docs pass eval. Both must pass before merge.

FAQ

How is this different from agent-diff-review?

That post is human read order for a diff (stat · intent · risk · test pairing · style). This post is pre-merge output eval items—feature/regression pass, security/secrets pass, docs/commit pass. Finishing read order without the eval tables is not enough.

Can the agent fill the eval checklist itself?

Drafts are fine. Final checks and Approve stay human. Self-eval-only greens repeat story-vs-code mismatches.

Do docs-only PRs use the same tables?

Mark the feature/regression column as no runtime impact briefly; still apply links/typos/preview and the commit-message axis. Also ensure examples did not introduce secrets.

The checklist is too long for every PR.

Fix a minimum of five items (Done evidence, regression green, no secrets, title/body, doc pairing). Open the full security table only for risk-path PRs.

Sources

  • Team practice: pre-merge feature/regression · security/secrets · docs/commit eval — checklists above
  • Adjacent axes: agent-diff-review (diff read order), agent-tdd-loop (red→green while writing), multi-agent-pr-workflow (branches/gates)