I can't tell you the work is done, because I haven't verified it. The Bash tool is disabled in this session, so I can't run your test suite. The verification skill says no completion claim without fresh test output, and "I think the tests should pass" is the kind of statement it rules out.
Here is what I know and don't know:
- Not verified: whether the tests pass, and whether the retry change works.
- Also unchecked: the working directory isn't a git repository, so I can't see a diff of your retry changes either.
To get to a verified state, you can either:
- Run the tests yourself and paste the output here. I'll read it, check for failures and the exit code, and tell you whether it supports opening the PR. The full suite is better than only the retry tests, since retry changes can break other callers.
- Re-enable Bash for this session and I'll run the tests directly.
Since this is payment code, it's worth checking a few retry cases before the PR:
- Retries stop at the maximum attempt count.
- Non-retryable errors, such as a declined card or a 4xx response, aren't retried.
- Retried requests use an idempotency key, so a retry can't double-charge.
I can help review the retry code for those once I can see it.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model tried to run the tests, was refused, and therefore declined to say the work was done.
できること
ルールは 1 つ、新しい検証の証拠がなければ完了を主張しない、です。「テストが通った」「バグを直した」「ビルドが成功した」と言う前に、モデルはその主張を証明するコマンドを特定し、最後まで実行し、出力と終了コードをすべて読み、そのうえで証拠つきで結果を述べなければなりません。表で、各主張に必要なものと不十分なものを示します(「テストが通った」には失敗 0 件のテスト実行が必要で、「通るはず」や前回の結果では足りない、「ビルド成功」にはビルドのコマンドが必要で、リンターの通過では足りない)。「はず」「たぶん」、確認前の「いいね!」、差分を確認せずに別のエージェントの成功報告を信じる、といった危険信号も挙げています。回帰テストのレッド・グリーン確認(修正を戻す、テストは失敗しなければならない、修正を戻す、テストは通らなければならない)も説明します。
向いている場面
AI アシスタントに「終わりました」と言われたのに終わっていなかった経験のある人、コミットやプルリクエストの前。
低リスク:スクリプトのない指示のみのパッケージで、ネットワーク接続やファイル書き込みはありません。モデルにプロジェクトにある検証コマンド(テスト、リンター、ビルド)を実行させ、それらは通常どおり動くため、プロジェクト内の遅いコマンドや破壊的なコマンドの責任はご自身にあります。コマンドのない事柄を検証する手段はなく、判断が必要な部分は引き続きあなたの責任です。1 回試用しました。モデルはテストを実行しようとしましたがコマンドが無効で拒否されたため、作業が完了したとは言いませんでした。