I couldn't save this to docs/postmortems/2026-02-10-checkout-api-500-errors.md because I have no file-write tool in this session. The full document is below, ready to paste into that path.
Owners and due dates are placeholders. The skill requires a named person for each action, and I only know Maya's name. Due dates are proposed, so please confirm them.
Post-Mortem: Checkout API 500 Errors
Date: 2026-02-10
Severity: P2
Duration: 1 hour 15 minutes (14:05 UTC – 15:20 UTC, detection to resolution)
Incident Commander: @Maya
Status: Resolved
Summary
After the v2.41 deploy at 13:58 UTC, the checkout API began returning 500 errors. A new query in the release loaded full result sets into memory, which caused pods to run out of memory and crash-loop. At peak, about 12% of requests failed, against a 0.1% baseline. Rolling back to the previous version at 15:00 restored service, and recovery was confirmed at 15:20. No data was lost and no SLA was breached.
Impact
| Dimension | Value |
|---|---|
| Affected services | Checkout API |
| User-facing impact | Errors (HTTP 500) on a portion of requests |
| Users affected | Not yet quantified (about 12% of requests failed at peak) |
| Peak error rate | ~12% vs 0.1% baseline |
| Data loss | None |
| SLA breach | No |
Duration note: The impact window began at the 13:58 deploy. Detection to resolution was 1h15m. Deploy to resolution was about 1h22m.
Timeline
All times UTC, 2026-02-10.
| Time | Event |
|---|---|
| 13:58 | v2.41 deployed |
| 14:05 | Alert fired for elevated checkout API errors |
| 14:10 | Maya paged and began coordinating as incident commander |
| 14:35 | Root cause identified: a new query loaded full result sets into memory, causing pod OOM crash-loops |
| 15:00 | Rollback of v2.41 performed |
| 15:20 | Recovery confirmed |
Gaps to verify from logs, alerts and Slack:
- When the first 500s appeared. The alert fired 7 minutes after the deploy, but the first symptom may have been earlier.
- What happened between 14:35 and 15:00. Root cause was known, but the rollback took 25 minutes. This could have been decision time, the rollback process, or the deploy pipeline.
- Whether any customer or internal communication was sent, and when.
- Whether the incident was formally declared, and at what time.
Root Cause
The release shipped a query that loaded entire result sets into memory instead of bounding or streaming them. For accounts with large result sets, memory use exceeded pod limits, and the pods crashed and restarted repeatedly. This produced the 500s.
The underlying gap is that the pre-release process had no way to catch this. Load tests covered only median accounts, so memory behavior for large accounts was never exercised before production.
5 Whys
- Why did users see 500s? The checkout API pods were crash-looping.
- Why were they crash-looping? They exceeded their memory limits.
- Why was memory exceeded? The new query loaded full result sets into memory.
- Why wasn't this caught before deploy? Load tests covered only median accounts.
- Why? There were no test fixtures or scenarios for large accounts, so a memory-heavy query path looked fine under test.
Contributing Factors
- Memory alert threshold at 90%. Pods can reach OOM before a 90% alert gives useful lead time, so the memory alert did not give an early signal.
- No runbook for OOM pods. Responders had no documented path for diagnosing memory crash-loops or deciding on a rollback.
- Load tests covered only median accounts. This is the same gap as the root cause, and it also meant the release had no large-account validation.
What Went Well
- The alert fired 7 minutes after the deploy.
- The on-call was paged within 5 minutes of the alert.
- Root cause was identified about 30 minutes after the page.
- Rollback resolved the issue, and recovery was explicitly confirmed.
- No data was lost and no SLA was breached.
What Could Have Gone Better
- 25 minutes passed between identifying root cause and starting the rollback.
- The release had no staged rollout or canary that might have limited exposure to the failing path.
- Memory pressure was not visible early, because the alert threshold left little warning.
- Responders had no runbook to speed up diagnosis and the rollback decision.
Action Items
| # | Action | Owner | Due Date | Priority |
|---|---|---|---|---|
| 1 | Add load-test fixtures for high-volume accounts and include them in the pre-release load test for the checkout API | @TBD | 2026-03-10 | High |
| 2 | Fix the v2.41 query to paginate or stream results with a bounded memory footprint, and add a test asserting a memory ceiling | @TBD | 2026-02-20 | High |
| 3 | Lower the checkout API pod memory alert threshold from 90% to a lower value (proposed 75%), and add an alert on OOMKilled and restart count | @TBD | 2026-02-24 | High |
| 4 | Write a runbook for OOM and crash-looping pods, including diagnosis steps and rollback criteria | @TBD | 2026-03-03 | Medium |
| 5 | Review the 14:35–15:00 gap and document the fastest rollback path. Consider a rollback-first guideline for post-deploy regressions | @Maya | 2026-03-03 | Medium |
| 6 | Evaluate canary or staged rollout for checkout API deploys | @TBD | 2026-03-31 | Medium |
Lessons Learned
- Tests built around typical accounts can hide failures that only appear at the tail. Memory-bound code paths need large-data fixtures.
- An alert threshold that is only reached when a pod is about to die gives little warning. OOM kills and restart counts are better early signals.
- Once the cause is known, rollback speed matters as much as diagnosis speed. A runbook and clear rollback criteria would shorten the gap between the two.
Next steps:
- Fill in owners for the TBD items.
- Check the flagged timeline gaps against logs and Slack, especially 14:35–15:00 and any customer communication.
- Add a user-count estimate if one is available.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so nothing was fetched from the web and no file was written. The example is fictional. The incident is invented. The model could not save the document to its usual path and printed it inline instead.
できること
障害や深刻な性能低下の後のポストモーテム作成を支援します。インシデントの基本情報を集め、正確なタイムラインを組み立てて空白を指摘し、5 つのなぜで根本原因と寄与要因を分け、影響(検知から復旧までの時間、ピーク時のエラー率、影響を受けたトラフィックの割合、ビジネスへの影響)の数値化を手伝い、各原因を担当者と期限が決まった具体的な改善項目にします。トーンは責任追及なし。壊れるのは人ではなく仕組みだという考え方です。
動き方
- モデルが書き始める前に、タイトル、時刻、重大度、影響を受けたサービス、おおまかなタイムラインを集めます。
- タイムライン、根本原因、寄与要因、影響を一緒に整理します。
- 要約、影響の表、タイムライン、根本原因、寄与要因、うまくいったこと、改善できたこと、改善項目、学びを含む完全な文書を作ります。
向いている場面
本番障害、ユーザーに見えるエラー、データ損失、SLA 違反、ヒヤリハット。48〜72 時間以内に書くのが理想です。
指示のみのパッケージです。スクリプトはなく、ネットワーク接続やアカウントは不要です。最後の手順で、プロジェクト内の `docs/postmortems/YYYY-MM-DD-<slug>.md` に文書を保存するため、保存先を伝えるか、会話内への出力を依頼してください。ポストモーテムには内部システムの詳細、顧客への影響の数値、氏名が含まれがちなので、共有する内容とファイルを読める人に注意してください。モデルが入れた担当者と期限は、チームが確認するまで仮の値です。正式なインシデント対応やコンプライアンスの手続きの代わりにはなりません。