Don't add the try/catch yet. It hides the symptom without fixing anything. Checkout is a write path, so swallowing the error there is risky. If order is undefined at line 88, the code either carries on with a half-built order or returns a success to the user for an order that wasn't created. That can mean silent data loss or a charge with no order, which is worse than a visible 500.
Phase 1: find the root cause
- Read the full stack trace. Check what
undefinedis at line 88. It's the object whose.idis being read, for exampleorder.id,user.idorcart.id. Then check who called this function and what they passed. - Check what changed in yesterday's deploy. Diff the release and look for anything touching
orderService.js, its callers, the shape of the data it reads (DB queries, migrations, upstream API responses), dependency bumps, and config or env changes. The timing points to the deploy, so that diff is the highest-value evidence. - Work out why it's intermittent. Intermittent usually means the failure depends on the input or the timing. Compare failing and succeeding requests for these differences:
- Particular users, carts or payment methods, such as guest vs. logged-in, empty cart, or a coupon.
- A race condition, such as a read before a write commits, or a missing
await. - A cache miss, or one instance or region running a different version or config.
- A lookup that returns
nullorundefinedfor some records, such as a.find()with no match or a missing relation.
- Add logging if the logs don't answer this. Log the inputs entering and leaving the function that contains line 88. Log the request ID, user, cart contents, and the result of whatever lookup feeds line 88. Then collect several failing and passing cases.
- Trace the bad value backward to where the
undefinedfirst appears. Fix it there, not at line 88.
If users are hurting right now
- Roll back yesterday's deploy if the diff points to it or you can't find the cause quickly. It stops the errors without hiding anything, and you can debug in peace.
- If you can't roll back, a guard is acceptable as a temporary mitigation. It should fail explicitly, with an early check that logs the context and returns a clear error or 4xx, rather than a blanket try/catch. It must not let checkout continue with bad state. Treat it as a stopgap and keep investigating.
Once you find the cause
Write a failing test that reproduces it, fix it at the source, and verify the test passes. Then consider validating at the boundary where the bad data enters, so this class of bug fails loudly and early.
The directory I'm in has no git repo, and I haven't seen orderService.js. If you point me at the repo, the deploy diff, or a few failing-request logs, I can start on steps 1–3.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The scenario is invented and the model had no access to the code, so it laid out the investigation steps. The bundled shell script was not run.
它做什么
定下一条铁律:没做完根因调查,不许修。第 1 阶段:完整读懂报错、稳定复现、检查最近的改动,并在每个组件边界收集证据。第 2 阶段:找到能正常工作的例子,列出所有差异。第 3 阶段:提出一个假设,用最小的改动去验证。第 4 阶段:先写一个会失败的测试,只做一处修复并验证。连续三次修复失败后,它会停下来请你重新审视架构,而不是再打第四个补丁。它还列出了导致乱猜的各种借口(“紧急”“只改一行”)和你察觉流程被跳过时会说的话。配套文件讲解倒着追踪根因、纵深防御,以及用条件轮询代替随意的超时(附 TypeScript 示例)。
适合什么场景
测试失败、线上 bug、不稳定的测试、构建和集成问题,尤其是时间紧迫的时候。
中风险:排查步骤会让模型加日志、运行诊断命令,示例里有 `env | grep` 和 `security find-identity`,所以环境变量或凭据存储信息可能出现在对话里,分享输出前请先遮蔽密钥。随包的 `find-polluter.sh` 会对每个匹配的测试文件各运行一次 `npm test <文件>`,用来找出会留下文件的测试,这等于执行你项目的测试及其副作用。不会删除任何东西,也不联网。本站收录的版本去掉了作者自用的测试场景和制作记录。试用时只描述了一个 bug、没有代码,也没有运行脚本。