Don't add the try/catch yet. It hides the symptom without fixing anything. Checkout is a write path, so swallowing the error there is risky. If order is undefined at line 88, the code either carries on with a half-built order or returns a success to the user for an order that wasn't created. That can mean silent data loss or a charge with no order, which is worse than a visible 500.
Phase 1: find the root cause
- Read the full stack trace. Check what
undefinedis at line 88. It's the object whose.idis being read, for exampleorder.id,user.idorcart.id. Then check who called this function and what they passed. - Check what changed in yesterday's deploy. Diff the release and look for anything touching
orderService.js, its callers, the shape of the data it reads (DB queries, migrations, upstream API responses), dependency bumps, and config or env changes. The timing points to the deploy, so that diff is the highest-value evidence. - Work out why it's intermittent. Intermittent usually means the failure depends on the input or the timing. Compare failing and succeeding requests for these differences:
- Particular users, carts or payment methods, such as guest vs. logged-in, empty cart, or a coupon.
- A race condition, such as a read before a write commits, or a missing
await. - A cache miss, or one instance or region running a different version or config.
- A lookup that returns
nullorundefinedfor some records, such as a.find()with no match or a missing relation.
- Add logging if the logs don't answer this. Log the inputs entering and leaving the function that contains line 88. Log the request ID, user, cart contents, and the result of whatever lookup feeds line 88. Then collect several failing and passing cases.
- Trace the bad value backward to where the
undefinedfirst appears. Fix it there, not at line 88.
If users are hurting right now
- Roll back yesterday's deploy if the diff points to it or you can't find the cause quickly. It stops the errors without hiding anything, and you can debug in peace.
- If you can't roll back, a guard is acceptable as a temporary mitigation. It should fail explicitly, with an early check that logs the context and returns a clear error or 4xx, rather than a blanket try/catch. It must not let checkout continue with bad state. Treat it as a stopgap and keep investigating.
Once you find the cause
Write a failing test that reproduces it, fix it at the source, and verify the test passes. Then consider validating at the boundary where the bad data enters, so this class of bug fails loudly and early.
The directory I'm in has no git repo, and I haven't seen orderService.js. If you point me at the repo, the deploy diff, or a few failing-request logs, I can start on steps 1–3.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The scenario is invented and the model had no access to the code, so it laid out the investigation steps. The bundled shell script was not run.
できること
鉄則は 1 つ、根本原因の調査が済むまで修正してはいけない、です。フェーズ 1 では、エラーを最後まで読み、再現し、最近の変更を確認し、各コンポーネントの境界で証拠を集めます。フェーズ 2 では、動いている例を探してすべての違いを挙げます。フェーズ 3 では、仮説を 1 つ立て、最小の変更で検証します。フェーズ 4 では、まず失敗するテストを書き、修正は 1 か所だけにして検証します。3 回続けて修正に失敗したら、4 つ目のパッチを当てずに止まり、アーキテクチャを見直すよう求めます。推測に走らせる言い訳(「緊急」「1 行だけ」)や、手順が飛ばされているときにあなたが言う言葉も挙げています。補助ファイルは、逆方向の原因追跡、多層防御、任意のタイムアウトを条件の待機に置き換える方法(TypeScript の例つき)です。
向いている場面
テストの失敗、本番のバグ、不安定なテスト、ビルドや結合の問題。特に時間がないとき。
中リスク:調査の手順では、モデルにログを追加させ、診断コマンドを実行させます。例には `env | grep` や `security find-identity` があるため、環境変数や認証情報ストアの内容が会話に出る可能性があります。出力を共有する前に秘密情報を伏せてください。同梱の `find-polluter.sh` は、ファイルを残すテストを探すため、該当するテストファイルごとに `npm test <ファイル>` を 1 回ずつ実行します。つまりプロジェクトのテストとその副作用が実行されます。何も削除せず、ネットワークも使いません。このサイトの掲載版には、作者自身のテストシナリオと作成記録を含めていません。試用ではバグを言葉で説明しただけでコードは見せず、スクリプトも実行していません。