Home / Skills / Developer Productivity / Verification Before Completion
Developer Productivity

Verification Before Completion

Jesse VincentMITSkill language: English
License verifiedNo scriptsLow risk
DemoReal run · claude-sonnet-5-5 · October 2026
I changed the retry logic in our payment client and I think the tests should pass now. Can you tell me the work is done so I can open the pull request? Use the verification-before-completion skill.
Skill: verification-before-completion

I can't tell you the work is done, because I haven't verified it. The Bash tool is disabled in this session, so I can't run your test suite. The verification skill says no completion claim without fresh test output, and "I think the tests should pass" is the kind of statement it rules out.

Here is what I know and don't know:

  • Not verified: whether the tests pass, and whether the retry change works.
  • Also unchecked: the working directory isn't a git repository, so I can't see a diff of your retry changes either.

To get to a verified state, you can either:

  1. Run the tests yourself and paste the output here. I'll read it, check for failures and the exit code, and tell you whether it supports opening the PR. The full suite is better than only the retry tests, since retry changes can break other callers.
  2. Re-enable Bash for this session and I'll run the tests directly.

Since this is payment code, it's worth checking a few retry cases before the PR:

  • Retries stop at the maximum attempt count.
  • Non-retryable errors, such as a declined card or a 4xx response, aren't retried.
  • Retried requests use an idempotency key, so a retry can't double-charge.

I can help review the retry code for those once I can see it.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model tried to run the tests, was refused, and therefore declined to say the work was done.

What it does

Sets one rule: no completion claim without fresh verification evidence. Before saying that tests pass, a bug is fixed or a build works, the model must identify the command that proves the claim, run it in full, read the whole output and exit code, and only then state the result with the evidence. A table lists what each claim needs and what is not enough ("tests pass" needs a test run showing 0 failures, not "it should pass" or a previous run; "build succeeds" needs the build command, not a passing linter). It lists red flags such as "should", "probably", "Great!" before checking, and trusting another agent's success report instead of checking the diff. It also describes the red-green check for a regression test (revert the fix, the test must fail, restore it, the test must pass).

Good for

Anyone who has been told "done" by an AI assistant and found it was not, and before committing or opening a pull request.

Notes & risks

Low risk: pure instructions with no scripts, no network access and no file writes. It makes the model run the verification commands your project has (tests, linter, build), and those commands behave as they normally do, so a slow or destructive command in your project is still your responsibility. It gives no way to verify things that have no command, so judgement calls remain yours. Trialled once; the model tried to run the tests, was refused because commands were disabled, and therefore declined to say the work was done.