← Back to the index
DBG-03 · SEC. 02 Debugging & Evaluation
Verify Before Claiming Fixed
Don't accept 'fixed it' from an AI agent until it re-runs the exact failing case and shows real output.
- FORMAT
- loop
- DIFFICULTY
- beginner
- TIME
- 5 min
- TOOLS
- claude-code · cursor · codex-cli · gemini-cli
- MODELS
- any
- COPIES
- 0 so far
When to use this
An AI agent just said "fixed" or "done" after changing code. You want proof, not a claim, before you move on or merge.
The pattern
Pastes as plain text
You just said this is fixed/done. Before I accept that, do the following: 1. Re-run the exact command or test that originally failed. Use the same input, same flags, and original environment. If that environment is unsafe or unavailable, use a safe equivalent and label the difference. Do not run a different or simplified version. 2. Record the exact command or test, environment, exit result, and relevant redacted output. Do not expose secrets, tokens, personal data, or an unnecessary full log. 3. If the output shows success, quote the specific line(s) that prove it (e.g. "PASS", exit code 0, the expected value in the response). 4. If the output shows any failure, warning, or unexpected result, say so plainly and do not call it fixed. Go back and continue debugging. 5. If you cannot re-run the original scenario (no test exists, no way to reproduce), say that explicitly instead of claiming success. Only after showing that record should you say the issue is resolved for this exact scenario and environment.
Real example output
Re-ran `pytest tests/test_auth.py::test_expired_token -v` in the local test environment, exit 0. Relevant output: ``` tests/test_auth.py::test_expired_token PASSED [100%] 1 passed in 0.42s ``` The PASSED line and 0 failures verify the expired-token scenario in the local test environment now returns a 401 instead of a 500. Broader production behavior remains unverified.
Why it works
AI agents optimize for sounding done, not being done. Requiring the exact original failing scenario, with pasted real output, closes the gap between "I believe this works" and "I watched it work." It also catches cases where the agent quietly changed the test instead of the code.
Related patterns
DBG-06Ban 'Should Work Now' ClaimsA standing rule that blocks an AI agent from claiming a fix works without pasting real output.DBG-04Adversarial Fix Verification LoopHave a second, fresh-context pass try to prove the first agent's fix is wrong before you trust it.ENG-11No Fix Without a Failing TestRefuse any bug fix until a test reproduces the bug, so the fix is proven, not vibes.