首页 / Skills / 开发提效 / 完成前先验证 Verification Before Completion
开发提效

完成前先验证 Verification Before Completion

Jesse VincentMITSkill 语言: 英文
协议已核对不含脚本低风险
演示真实运行记录 · claude-sonnet-5-5 · 2026年10月
I changed the retry logic in our payment client and I think the tests should pass now. Can you tell me the work is done so I can open the pull request? Use the verification-before-completion skill.
Skill: verification-before-completion

I can't tell you the work is done, because I haven't verified it. The Bash tool is disabled in this session, so I can't run your test suite. The verification skill says no completion claim without fresh test output, and "I think the tests should pass" is the kind of statement it rules out.

Here is what I know and don't know:

  • Not verified: whether the tests pass, and whether the retry change works.
  • Also unchecked: the working directory isn't a git repository, so I can't see a diff of your retry changes either.

To get to a verified state, you can either:

  1. Run the tests yourself and paste the output here. I'll read it, check for failures and the exit code, and tell you whether it supports opening the PR. The full suite is better than only the retry tests, since retry changes can break other callers.
  2. Re-enable Bash for this session and I'll run the tests directly.

Since this is payment code, it's worth checking a few retry cases before the PR:

  • Retries stop at the maximum attempt count.
  • Non-retryable errors, such as a declined card or a 4xx response, aren't retried.
  • Retried requests use an idempotency key, so a retry can't double-charge.

I can help review the retry code for those once I can see it.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model tried to run the tests, was refused, and therefore declined to say the work was done.

它做什么

只定一条规则:没有新鲜的验证证据,就不许宣称完成。在说“测试通过”“bug 已修复”“构建成功”之前,模型必须找出能证明这个说法的命令,完整运行它,读完全部输出和退出码,然后才带着证据说出结果。一张表列出每种说法需要什么、什么不够(“测试通过”需要一次显示 0 个失败的测试运行,而不是“应该会过”或上一次的结果;“构建成功”需要构建命令,而不是 lint 通过)。它列出了危险信号,比如“应该”“大概”、还没检查就说“太好了!”、不核对改动就信任另一个代理的成功报告。它还描述了回归测试的红绿检查(撤销修复,测试必须失败,恢复修复,测试必须通过)。

适合什么场景

被 AI 助手告知“做完了”结果发现没做完的人,以及提交代码或开 PR 之前。

说明与风险

低风险:纯指令文件,没有脚本,不联网、不写文件。它会让模型运行你项目里已有的验证命令(测试、lint、构建),这些命令照常运行,所以你项目里速度慢或有破坏性的命令,责任仍在你。对于没有对应命令的事情它没有验证办法,需要判断的部分仍由你来定。已试用一次:模型想运行测试,但命令被禁用,于是拒绝说工作已经完成。