Agent Output Eval Checklist
Agents can fill a PR, patch, and even commit messages quickly, but pre-merge eval stays a human (or CODEOWNERS) gate. Skimming only in “read order” still lets feature regressions, secret-adjacent edits, and thin docs/commit lines slip through. You need a per-axis output quality checklist.
This post covers only feature/regression? · security/secrets? · docs/commit messages?. agent-diff-review is the diff read-order axis (stat → intent → risk paths → test pairing). Here the focus is pass/fail eval items before merge. No pricing, plans, tokens, or affiliate links.
Feature and regression?
One-line answer: Against the ticket’s Done and out-of-scope, both new/changed behavior must be reproducible and core existing paths must stay green. “The agent said it passed” is not an eval result.
Pre-merge feature/regression eval:
| Item | Pass when | On fail |
|---|---|---|
| Done mapping | Every Done line has evidence (command, UI, log path) | Empty Done / claims only → hold merge |
| Out-of-scope | Declared out-of-scope work is not in the diff | Drive-by refactors / dependency bumps → revert or split PR |
| New paths | Happy path + at least one failure/edge scenario | Happy-only → hold or require follow-up ticket |
| Regression | Team smoke / related suite green | Red or skip-to-fake-green → no merge |
| Contracts | Public API, schema, auth flag changes are named in Done | Silent breaking change → hold |
Copy-paste:
[ ] Each Done item has an evidence path (command/CI/manual)?
[ ] Out-of-scope not mixed into the diff?
[ ] New behavior: happy + at least one failure/edge?
[ ] Related regression suite/smoke green? no skip / weakened asserts?
[ ] API/schema/auth changes named in Done?
Often missing in agent output: “fixed” with no repro command, flaky timeouts only raised, “behavior unchanged” with no suite. agent-tdd-loop is red→green while writing; this section is pass/fail on feature/regression at merge time.
Security and secrets?
One-line answer: Confirm diff, commits, logs, and issue bodies have no secrets, credentials, internal URLs, or PII, and do not merge privilege expansion, auth bypass, or loosened defaults without an explicit reviewer plus a rollback unit.
Security/secrets eval table:
| Surface | Look for | Pass |
|---|---|---|
| Secret mix-in | .env*, keys/tokens/PEM, hardcoded passwords, pasted agent chat | None. If present: scrub history + rotation plan |
| Logs / screenshots | Sample output, captures, CI logs with tokens/cookies | Masked or removed |
| AuthZ | Deleted checks, * allow, wider CORS/public ACL | Named in Done + owner approval |
| Supply chain | Lockfile, new deps, postinstall scripts | Install/build/core tests + provenance |
| Data | Prod dumps, real-user PII fixtures | Synthetic/anonymized only |
Copy-paste:
[ ] No secrets/keys/tokens in diff, commits, or PR body?
[ ] Examples, logs, screenshots masked?
[ ] AuthZ / public ACL changes in Done with owner approval?
[ ] New deps / lockfile: reason + verification commands?
[ ] No real data / PII fixtures?
[ ] Suspected leak: no merge + key-rotation runbook
Note: This overlaps agent-diff-review’s “put risk paths in read order,” but here the point is pass/fail items and leak response. “We looked during review” alone is not an eval pass.
Docs and commit messages?
One-line answer: A human must still recover intent, out-of-scope, and how to verify six months later. Titles state what/why, bodies cover verification · risk · follow-ups, and docs (README, ADR, runbook) pair with behavior changes.
Docs/commit eval:
| Item | Pass when | Common fail |
|---|---|---|
| Commit/PR title | Scope readable in one line (module · behavior) | fix, update, agent changes |
| Body | Done / Out-of-scope / verify commands + results | “tests passed” with no evidence path |
| Risk one-liner | Secrets, migration, flag, rollback unit | Blank while risk paths changed |
| Doc pairing | User/ops change → README, runbook, API notes updated | Code only, docs stale |
| Issue links | Ticket / follow-up IDs | “later” with no ID |
Copy-paste:
[ ] Title names module/behavior (not just fix/update)?
[ ] Body has Done / Out-of-scope / verification evidence?
[ ] Risk/rollback one-liner if the change is dangerous?
[ ] Doc pairing for behavior/ops changes?
[ ] Follow-ups have ticket IDs?
[ ] Do not replace the body with agent “self-review” prose alone
Boundary: agent-diff-review is in what order to read files; this section is whether merge history and onboarding docs pass eval. Both must pass before merge.
FAQ
How is this different from agent-diff-review?
That post is human read order for a diff (stat · intent · risk · test pairing · style). This post is pre-merge output eval items—feature/regression pass, security/secrets pass, docs/commit pass. Finishing read order without the eval tables is not enough.
Can the agent fill the eval checklist itself?
Drafts are fine. Final checks and Approve stay human. Self-eval-only greens repeat story-vs-code mismatches.
Do docs-only PRs use the same tables?
Mark the feature/regression column as no runtime impact briefly; still apply links/typos/preview and the commit-message axis. Also ensure examples did not introduce secrets.
The checklist is too long for every PR.
Fix a minimum of five items (Done evidence, regression green, no secrets, title/body, doc pairing). Open the full security table only for risk-path PRs.
Sources
- Team practice: pre-merge feature/regression · security/secrets · docs/commit eval — checklists above
- Adjacent axes: agent-diff-review (diff read order), agent-tdd-loop (red→green while writing), multi-agent-pr-workflow (branches/gates)