首頁 / Skills / 開發提效 / 完成前先驗證 Verification Before Completion
開發提效

完成前先驗證 Verification Before Completion

Jesse VincentMITSkill 語言: 英文
授權已核對不含腳本低風險
演示真實執行紀錄 · claude-sonnet-5-5 · 2026年10月
I changed the retry logic in our payment client and I think the tests should pass now. Can you tell me the work is done so I can open the pull request? Use the verification-before-completion skill.
Skill: verification-before-completion

I can't tell you the work is done, because I haven't verified it. The Bash tool is disabled in this session, so I can't run your test suite. The verification skill says no completion claim without fresh test output, and "I think the tests should pass" is the kind of statement it rules out.

Here is what I know and don't know:

  • Not verified: whether the tests pass, and whether the retry change works.
  • Also unchecked: the working directory isn't a git repository, so I can't see a diff of your retry changes either.

To get to a verified state, you can either:

  1. Run the tests yourself and paste the output here. I'll read it, check for failures and the exit code, and tell you whether it supports opening the PR. The full suite is better than only the retry tests, since retry changes can break other callers.
  2. Re-enable Bash for this session and I'll run the tests directly.

Since this is payment code, it's worth checking a few retry cases before the PR:

  • Retries stop at the maximum attempt count.
  • Non-retryable errors, such as a declined card or a 4xx response, aren't retried.
  • Retried requests use an idempotency key, so a retry can't double-charge.

I can help review the retry code for those once I can see it.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model tried to run the tests, was refused, and therefore declined to say the work was done.

它做什麼

只定一條規則:沒有新鮮的驗證證據,就不許宣稱完成。在說「測試通過」「bug 已修復」「建置成功」之前,模型必須找出能證明這個說法的指令,完整執行它,讀完全部輸出與結束碼,然後才帶著證據說出結果。一張表列出每種說法需要什麼、什麼不夠(「測試通過」需要一次顯示 0 個失敗的測試執行,而不是「應該會過」或上一次的結果;「建置成功」需要建置指令,而不是 lint 通過)。它列出了危險訊號,例如「應該」「大概」、還沒檢查就說「太好了!」、不核對變更就信任另一個代理的成功回報。它還描述了回歸測試的紅綠檢查(還原修復,測試必須失敗,恢復修復,測試必須通過)。

適合什麼場景

被 AI 助理告知「做完了」結果發現沒做完的人,以及提交程式碼或開 PR 之前。

說明與風險

低風險:純指令檔,沒有腳本,不連網、不寫檔。它會讓模型執行你專案中已有的驗證指令(測試、lint、建置),這些指令照常執行,所以你專案中速度慢或具破壞性的指令,責任仍在你。對於沒有對應指令的事情它沒有驗證辦法,需要判斷的部分仍由你來定。已試用一次:模型想執行測試,但指令被停用,於是拒絕說工作已經完成。