Don't add the try/catch yet. It hides the symptom without fixing anything. Checkout is a write path, so swallowing the error there is risky. If order is undefined at line 88, the code either carries on with a half-built order or returns a success to the user for an order that wasn't created. That can mean silent data loss or a charge with no order, which is worse than a visible 500.
Phase 1: find the root cause
- Read the full stack trace. Check what
undefinedis at line 88. It's the object whose.idis being read, for exampleorder.id,user.idorcart.id. Then check who called this function and what they passed. - Check what changed in yesterday's deploy. Diff the release and look for anything touching
orderService.js, its callers, the shape of the data it reads (DB queries, migrations, upstream API responses), dependency bumps, and config or env changes. The timing points to the deploy, so that diff is the highest-value evidence. - Work out why it's intermittent. Intermittent usually means the failure depends on the input or the timing. Compare failing and succeeding requests for these differences:
- Particular users, carts or payment methods, such as guest vs. logged-in, empty cart, or a coupon.
- A race condition, such as a read before a write commits, or a missing
await. - A cache miss, or one instance or region running a different version or config.
- A lookup that returns
nullorundefinedfor some records, such as a.find()with no match or a missing relation.
- Add logging if the logs don't answer this. Log the inputs entering and leaving the function that contains line 88. Log the request ID, user, cart contents, and the result of whatever lookup feeds line 88. Then collect several failing and passing cases.
- Trace the bad value backward to where the
undefinedfirst appears. Fix it there, not at line 88.
If users are hurting right now
- Roll back yesterday's deploy if the diff points to it or you can't find the cause quickly. It stops the errors without hiding anything, and you can debug in peace.
- If you can't roll back, a guard is acceptable as a temporary mitigation. It should fail explicitly, with an early check that logs the context and returns a clear error or 4xx, rather than a blanket try/catch. It must not let checkout continue with bad state. Treat it as a stopgap and keep investigating.
Once you find the cause
Write a failing test that reproduces it, fix it at the source, and verify the test passes. Then consider validating at the boundary where the bad data enters, so this class of bug fails loudly and early.
The directory I'm in has no git repo, and I haven't seen orderService.js. If you point me at the repo, the deploy diff, or a few failing-request logs, I can start on steps 1–3.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The scenario is invented and the model had no access to the code, so it laid out the investigation steps. The bundled shell script was not run.
Was es macht
Setzt eine eiserne Regel: Kein Fix ohne vorherige Ursachenuntersuchung. Phase 1 liest Fehler vollständig, reproduziert das Problem, prüft jüngste Änderungen und sammelt Belege an jeder Komponentengrenze. Phase 2 sucht funktionierende Beispiele und listet jeden Unterschied auf. Phase 3 formuliert eine Hypothese und testet sie mit der kleinstmöglichen Änderung. Phase 4 schreibt einen fehlschlagenden Test, macht genau einen Fix und verifiziert ihn. Nach drei gescheiterten Fixes stoppt es und bittet Sie, die Architektur zu hinterfragen, statt einen vierten Patch zu versuchen. Es nennt auch die Ausreden, die zum Raten führen („Notfall“, „nur ein schneller Fix“), und Ihre eigenen Signale, dass der Prozess übersprungen wird. Begleitdateien behandeln rückwärtiges Ursachen-Tracing, Defense in Depth und das Ersetzen willkürlicher Timeouts durch Bedingungs-Polling (mit TypeScript-Beispiel).
Geeignet für
Testfehler, Produktionsfehler, instabile Tests, Build- und Integrationsprobleme, besonders unter Zeitdruck.
Mittleres Risiko: Die Untersuchungsschritte lassen das Modell Logging einbauen und Diagnosebefehle ausführen, und die Beispiele enthalten `env | grep` und `security find-identity`; Umgebungsvariablen oder Details des Credential-Stores können so in der Unterhaltung landen, maskieren Sie Geheimnisse, bevor Sie Ausgaben teilen. Das mitgelieferte `find-polluter.sh` führt für jede passende Testdatei einmal `npm test <Datei>` aus, um den Test zu finden, der Dateien hinterlässt; das führt die Tests Ihres Projekts samt Nebenwirkungen aus. Es wird nichts gelöscht, und es gibt keinen Netzwerkzugriff. In dieser Listung fehlen die eigenen Testszenarien und das Erstellungsprotokoll des Autors. Getestet an einem beschriebenen Fehler ohne Codezugriff; die Skripte wurden nicht ausgeführt.