Agent work / verification
How do I verify a coding agent's work?
Keep the acceptance checks outside the agent's editable checkout.
I define the result first, run the same checks against the original code and the proposed fix, and inspect the evidence before I accept the work. A green test run can hide a changed requirement.
Keep the checks under your control
1. Define
Write the expected behavior and failure cases before the agent starts.
2. Compare
Run unchanged acceptance checks against the original code and the candidate.
3. Review
Inspect the code diff, check output, and any unresolved limits.
For a bug fix, the relevant regression check should fail on the original code and pass on the fix. Checks for an added feature can start with that feature absent. Run unrelated checks too, so a narrow fix does not hide another failure.
What if the agent changes the tests?
Keep a trusted copy of the acceptance suite in a separate verification checkout that the agent cannot edit. Use a disposable environment for candidate code. A separate folder alone does not enforce permissions; check the agent's actual write access.
The agent can propose new tests or explain a faulty test. Review that change against the requirement separately. If the requirement did not change, a weaker assertion is not proof of a fix. I keep both the old and proposed test output with the diff.
Run the example yourself
This synthetic example counts two TODO notes and one DONE note. Broken code counts only one TODO. A weakened test expects that wrong answer and passes. The fixed code passes the original acceptance checks.
I use Python's standard library here. The example makes no model or network calls and changes no files. It demonstrates the test boundary, not a BMD run or a customer outcome.
Download the runnable examplepython verify-agent-work.py| Candidate | Checks | Result |
|---|---|---|
| Broken code | Original | FAIL |
| Broken code | Weakened | PASS |
| Fixed code | Original | PASS |
What evidence should the agent provide before it says done?
- The requirement, expected output, and unchanged acceptance checks.
- The exact candidate revision and the code and test diffs.
- The check command, environment, exit status, and complete output.
- Failures, skipped checks, and behavior that remains unverified.
- A browser check for visible changes, or a readback from the actual system when the task changes live state.
Passing these checks proves the behavior they exercise. It does not prove every behavior of the program. A deployment, email, or payment needs evidence from that system; local tests alone do not prove the external result.
Sources and a next step
The question came from a public discussion about agents changing tests. Claude Code's verification guidance explains why concrete checks matter. The runnable example and its output are the proof for this guide.
BMD offers a Windows workspace for agent queues, runs, and review. Check its release notes for the behavior in your version. You still need acceptance checks you trust.
Set up BMD