Don't add the try/catch yet. It hides the symptom without fixing anything. Checkout is a write path, so swallowing the error there is risky. If order is undefined at line 88, the code either carries on with a half-built order or returns a success to the user for an order that wasn't created. That can mean silent data loss or a charge with no order, which is worse than a visible 500.
Phase 1: find the root cause
- Read the full stack trace. Check what
undefinedis at line 88. It's the object whose.idis being read, for exampleorder.id,user.idorcart.id. Then check who called this function and what they passed. - Check what changed in yesterday's deploy. Diff the release and look for anything touching
orderService.js, its callers, the shape of the data it reads (DB queries, migrations, upstream API responses), dependency bumps, and config or env changes. The timing points to the deploy, so that diff is the highest-value evidence. - Work out why it's intermittent. Intermittent usually means the failure depends on the input or the timing. Compare failing and succeeding requests for these differences:
- Particular users, carts or payment methods, such as guest vs. logged-in, empty cart, or a coupon.
- A race condition, such as a read before a write commits, or a missing
await. - A cache miss, or one instance or region running a different version or config.
- A lookup that returns
nullorundefinedfor some records, such as a.find()with no match or a missing relation.
- Add logging if the logs don't answer this. Log the inputs entering and leaving the function that contains line 88. Log the request ID, user, cart contents, and the result of whatever lookup feeds line 88. Then collect several failing and passing cases.
- Trace the bad value backward to where the
undefinedfirst appears. Fix it there, not at line 88.
If users are hurting right now
- Roll back yesterday's deploy if the diff points to it or you can't find the cause quickly. It stops the errors without hiding anything, and you can debug in peace.
- If you can't roll back, a guard is acceptable as a temporary mitigation. It should fail explicitly, with an early check that logs the context and returns a clear error or 4xx, rather than a blanket try/catch. It must not let checkout continue with bad state. Treat it as a stopgap and keep investigating.
Once you find the cause
Write a failing test that reproduces it, fix it at the source, and verify the test passes. Then consider validating at the boundary where the bad data enters, so this class of bug fails loudly and early.
The directory I'm in has no git repo, and I haven't seen orderService.js. If you point me at the repo, the deploy diff, or a few failing-request logs, I can start on steps 1–3.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The scenario is invented and the model had no access to the code, so it laid out the investigation steps. The bundled shell script was not run.
What it does
Sets one iron rule: no fixes without a root-cause investigation first. Phase 1 reads errors fully, reproduces the problem, checks recent changes and gathers evidence at each component boundary. Phase 2 finds working examples and lists every difference. Phase 3 states one hypothesis and tests it with the smallest possible change. Phase 4 writes a failing test, makes a single fix and verifies it. After three failed fixes it stops and asks you to question the architecture rather than trying a fourth patch. It also lists the excuses that lead to guessing ("emergency", "just one quick fix") and your own signals that the process is being skipped. Supporting files cover backward root-cause tracing, defence in depth, and replacing arbitrary timeouts with condition polling (with a TypeScript example).
Good for
Test failures, production bugs, flaky tests, build and integration problems, especially under time pressure.
Medium risk: the investigation steps tell the model to add logging and run diagnostic commands, and the examples include `env | grep` and `security find-identity`, so environment variables or credential-store details can end up in the conversation; mask secrets before sharing output. The bundled `find-polluter.sh` runs `npm test <file>` once per matching test file to find the test that leaves files behind, which executes your project's tests and any side effects they have. Nothing is deleted and nothing uses the network. In this listing the author's own test scenarios and creation log were left out. Trialled on a described bug without code access; the scripts were not run.