Don't add the try/catch yet. It hides the symptom without fixing anything. Checkout is a write path, so swallowing the error there is risky. If order is undefined at line 88, the code either carries on with a half-built order or returns a success to the user for an order that wasn't created. That can mean silent data loss or a charge with no order, which is worse than a visible 500.
Phase 1: find the root cause
- Read the full stack trace. Check what
undefinedis at line 88. It's the object whose.idis being read, for exampleorder.id,user.idorcart.id. Then check who called this function and what they passed. - Check what changed in yesterday's deploy. Diff the release and look for anything touching
orderService.js, its callers, the shape of the data it reads (DB queries, migrations, upstream API responses), dependency bumps, and config or env changes. The timing points to the deploy, so that diff is the highest-value evidence. - Work out why it's intermittent. Intermittent usually means the failure depends on the input or the timing. Compare failing and succeeding requests for these differences:
- Particular users, carts or payment methods, such as guest vs. logged-in, empty cart, or a coupon.
- A race condition, such as a read before a write commits, or a missing
await. - A cache miss, or one instance or region running a different version or config.
- A lookup that returns
nullorundefinedfor some records, such as a.find()with no match or a missing relation.
- Add logging if the logs don't answer this. Log the inputs entering and leaving the function that contains line 88. Log the request ID, user, cart contents, and the result of whatever lookup feeds line 88. Then collect several failing and passing cases.
- Trace the bad value backward to where the
undefinedfirst appears. Fix it there, not at line 88.
If users are hurting right now
- Roll back yesterday's deploy if the diff points to it or you can't find the cause quickly. It stops the errors without hiding anything, and you can debug in peace.
- If you can't roll back, a guard is acceptable as a temporary mitigation. It should fail explicitly, with an early check that logs the context and returns a clear error or 4xx, rather than a blanket try/catch. It must not let checkout continue with bad state. Treat it as a stopgap and keep investigating.
Once you find the cause
Write a failing test that reproduces it, fix it at the source, and verify the test passes. Then consider validating at the boundary where the bad data enters, so this class of bug fails loudly and early.
The directory I'm in has no git repo, and I haven't seen orderService.js. If you point me at the repo, the deploy diff, or a few failing-request logs, I can start on steps 1–3.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The scenario is invented and the model had no access to the code, so it laid out the investigation steps. The bundled shell script was not run.
它做什麼
定下一條鐵律:沒做完根因調查,不許修。第 1 階段:完整讀懂錯誤訊息、穩定重現、檢查最近的變更,並在每個元件邊界蒐集證據。第 2 階段:找到能正常運作的範例,列出所有差異。第 3 階段:提出一個假設,用最小的改動去驗證。第 4 階段:先寫一個會失敗的測試,只做一處修復並驗證。連續三次修復失敗後,它會停下來請你重新檢視架構,而不是再打第四個補丁。它還列出了導致亂猜的各種藉口(「緊急」「只改一行」)與你察覺流程被跳過時會說的話。配套檔案講解倒著追蹤根因、縱深防禦,以及用條件輪詢取代隨意的逾時(附 TypeScript 範例)。
適合什麼場景
測試失敗、正式環境 bug、不穩定的測試、建置與整合問題,尤其是時間緊迫的時候。
中風險:排查步驟會讓模型加日誌、執行診斷指令,範例中有 `env | grep` 與 `security find-identity`,所以環境變數或憑證儲存資訊可能出現在對話中,分享輸出前請先遮蔽金鑰。隨附的 `find-polluter.sh` 會對每個符合的測試檔各執行一次 `npm test <檔案>`,用來找出會留下檔案的測試,這等於執行你專案的測試及其副作用。不會刪除任何東西,也不連網。本站收錄的版本去掉了作者自用的測試情境與製作紀錄。試用時只描述了一個 bug、沒有程式碼,也沒有執行腳本。