Start / Skills / Entwickler-Produktivität / Verifizieren vor dem Abschluss (Verification Before Completion)
Entwickler-Produktivität

Verifizieren vor dem Abschluss (Verification Before Completion)

Jesse VincentMITSprache des Skills: Englisch
Lizenz geprüftKeine SkripteGeringes Risiko
DemoEchter Lauf · claude-sonnet-5-5 · Oktober 2026
I changed the retry logic in our payment client and I think the tests should pass now. Can you tell me the work is done so I can open the pull request? Use the verification-before-completion skill.
Skill: verification-before-completion

I can't tell you the work is done, because I haven't verified it. The Bash tool is disabled in this session, so I can't run your test suite. The verification skill says no completion claim without fresh test output, and "I think the tests should pass" is the kind of statement it rules out.

Here is what I know and don't know:

  • Not verified: whether the tests pass, and whether the retry change works.
  • Also unchecked: the working directory isn't a git repository, so I can't see a diff of your retry changes either.

To get to a verified state, you can either:

  1. Run the tests yourself and paste the output here. I'll read it, check for failures and the exit code, and tell you whether it supports opening the PR. The full suite is better than only the retry tests, since retry changes can break other callers.
  2. Re-enable Bash for this session and I'll run the tests directly.

Since this is payment code, it's worth checking a few retry cases before the PR:

  • Retries stop at the maximum attempt count.
  • Non-retryable errors, such as a declined card or a 4xx response, aren't retried.
  • Retried requests use an idempotency key, so a retry can't double-charge.

I can help review the retry code for those once I can see it.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model tried to run the tests, was refused, and therefore declined to say the work was done.

Was es macht

Setzt eine einzige Regel: keine Abschlussbehauptung ohne frischen Verifikationsbeleg. Bevor das Modell sagt, dass Tests bestehen, ein Fehler behoben ist oder ein Build funktioniert, muss es den Befehl bestimmen, der die Behauptung belegt, ihn vollständig ausführen, die gesamte Ausgabe und den Exit-Code lesen und erst dann das Ergebnis mit Beleg nennen. Eine Tabelle zeigt, was jede Behauptung braucht und was nicht genügt („Tests bestehen“ braucht einen Testlauf mit 0 Fehlern, nicht „sollte bestehen“ oder einen früheren Lauf; „Build erfolgreich“ braucht den Build-Befehl, nicht einen bestandenen Linter). Es nennt Warnzeichen wie „sollte“, „wahrscheinlich“, „Super!“ vor der Prüfung und das Vertrauen auf den Erfolgsbericht eines anderen Agenten ohne Blick in den Diff. Außerdem beschreibt es die Red-Green-Prüfung für einen Regressionstest (Fix zurücknehmen, der Test muss fehlschlagen, Fix wiederherstellen, der Test muss bestehen).

Geeignet für

Alle, denen ein KI-Assistent „fertig“ gemeldet hat, obwohl es das nicht war, und vor dem Committen oder Öffnen eines Pull Requests.

Hinweise & Risiken

Geringes Risiko: Reine Anweisungen ohne Skripte, ohne Netzwerkzugriff und ohne Dateischreiben. Es lässt das Modell die Prüfbefehle Ihres Projekts ausführen (Tests, Linter, Build), die sich wie üblich verhalten; langsame oder destruktive Befehle in Ihrem Projekt bleiben Ihre Verantwortung. Für Dinge ohne Befehl bietet es keine Verifikation, Ermessensentscheidungen bleiben bei Ihnen. Einmal getestet; das Modell wollte die Tests ausführen, wurde abgelehnt, weil Befehle deaktiviert waren, und sagte deshalb nicht, die Arbeit sei fertig.