홈 / Skills / 개발 생산성 / 장애 포스트모템 (Incident Post-Mortem, 비난 없음)
개발 생산성

장애 포스트모템 (Incident Post-Mortem, 비난 없음)

awesome-copilot communityMITSkill 언어: 영어
라이선스 확인됨스크립트 없음중간 위험
데모실행 기록 · claude-sonnet-5-5 · 2026년 10월
Using the incident-postmortem skill, write a post-mortem for this fictional incident. Title: Checkout API 500 errors. Detected 2026-02-10 14:05 UTC, resolved 15:20 UTC, severity P2, incident commander Maya. Affected: checkout API; about 12 percent of requests failed at peak versus 0.1 percent baseline; no data loss; no SLA breach. Timeline: 13:58 deploy of v2.41; 14:05 alert fired; 14:10 Maya paged; 14:35 root cause found (new query loaded full result sets into memory so pods crash-looped); 15:00 rollback; 15:20 confirmed. Contributing: memory alert threshold at 90 percent, load tests only covered median accounts, no runbook for OOM pods.
Skill: incident-postmortem

I couldn't save this to docs/postmortems/2026-02-10-checkout-api-500-errors.md because I have no file-write tool in this session. The full document is below, ready to paste into that path.

Owners and due dates are placeholders. The skill requires a named person for each action, and I only know Maya's name. Due dates are proposed, so please confirm them.


Post-Mortem: Checkout API 500 Errors

Date: 2026-02-10
Severity: P2
Duration: 1 hour 15 minutes (14:05 UTC – 15:20 UTC, detection to resolution)
Incident Commander: @Maya
Status: Resolved


Summary

After the v2.41 deploy at 13:58 UTC, the checkout API began returning 500 errors. A new query in the release loaded full result sets into memory, which caused pods to run out of memory and crash-loop. At peak, about 12% of requests failed, against a 0.1% baseline. Rolling back to the previous version at 15:00 restored service, and recovery was confirmed at 15:20. No data was lost and no SLA was breached.

Impact

Dimension Value
Affected services Checkout API
User-facing impact Errors (HTTP 500) on a portion of requests
Users affected Not yet quantified (about 12% of requests failed at peak)
Peak error rate ~12% vs 0.1% baseline
Data loss None
SLA breach No

Duration note: The impact window began at the 13:58 deploy. Detection to resolution was 1h15m. Deploy to resolution was about 1h22m.

Timeline

All times UTC, 2026-02-10.

Time Event
13:58 v2.41 deployed
14:05 Alert fired for elevated checkout API errors
14:10 Maya paged and began coordinating as incident commander
14:35 Root cause identified: a new query loaded full result sets into memory, causing pod OOM crash-loops
15:00 Rollback of v2.41 performed
15:20 Recovery confirmed

Gaps to verify from logs, alerts and Slack:

  • When the first 500s appeared. The alert fired 7 minutes after the deploy, but the first symptom may have been earlier.
  • What happened between 14:35 and 15:00. Root cause was known, but the rollback took 25 minutes. This could have been decision time, the rollback process, or the deploy pipeline.
  • Whether any customer or internal communication was sent, and when.
  • Whether the incident was formally declared, and at what time.

Root Cause

The release shipped a query that loaded entire result sets into memory instead of bounding or streaming them. For accounts with large result sets, memory use exceeded pod limits, and the pods crashed and restarted repeatedly. This produced the 500s.

The underlying gap is that the pre-release process had no way to catch this. Load tests covered only median accounts, so memory behavior for large accounts was never exercised before production.

5 Whys

  • Why did users see 500s? The checkout API pods were crash-looping.
  • Why were they crash-looping? They exceeded their memory limits.
  • Why was memory exceeded? The new query loaded full result sets into memory.
  • Why wasn't this caught before deploy? Load tests covered only median accounts.
  • Why? There were no test fixtures or scenarios for large accounts, so a memory-heavy query path looked fine under test.

Contributing Factors

  • Memory alert threshold at 90%. Pods can reach OOM before a 90% alert gives useful lead time, so the memory alert did not give an early signal.
  • No runbook for OOM pods. Responders had no documented path for diagnosing memory crash-loops or deciding on a rollback.
  • Load tests covered only median accounts. This is the same gap as the root cause, and it also meant the release had no large-account validation.

What Went Well

  • The alert fired 7 minutes after the deploy.
  • The on-call was paged within 5 minutes of the alert.
  • Root cause was identified about 30 minutes after the page.
  • Rollback resolved the issue, and recovery was explicitly confirmed.
  • No data was lost and no SLA was breached.

What Could Have Gone Better

  • 25 minutes passed between identifying root cause and starting the rollback.
  • The release had no staged rollout or canary that might have limited exposure to the failing path.
  • Memory pressure was not visible early, because the alert threshold left little warning.
  • Responders had no runbook to speed up diagnosis and the rollback decision.

Action Items

# Action Owner Due Date Priority
1 Add load-test fixtures for high-volume accounts and include them in the pre-release load test for the checkout API @TBD 2026-03-10 High
2 Fix the v2.41 query to paginate or stream results with a bounded memory footprint, and add a test asserting a memory ceiling @TBD 2026-02-20 High
3 Lower the checkout API pod memory alert threshold from 90% to a lower value (proposed 75%), and add an alert on OOMKilled and restart count @TBD 2026-02-24 High
4 Write a runbook for OOM and crash-looping pods, including diagnosis steps and rollback criteria @TBD 2026-03-03 Medium
5 Review the 14:35–15:00 gap and document the fastest rollback path. Consider a rollback-first guideline for post-deploy regressions @Maya 2026-03-03 Medium
6 Evaluate canary or staged rollout for checkout API deploys @TBD 2026-03-31 Medium

Lessons Learned

  • Tests built around typical accounts can hide failures that only appear at the tail. Memory-bound code paths need large-data fixtures.
  • An alert threshold that is only reached when a pod is about to die gives little warning. OOM kills and restart counts are better early signals.
  • Once the cause is known, rollback speed matters as much as diagnosis speed. A runbook and clear rollback criteria would shorten the gap between the two.

Next steps:

  • Fill in owners for the TBD items.
  • Check the flagged timeline gaps against logs and Slack, especially 14:35–15:00 and any customer communication.
  • Add a user-count estimate if one is available.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so nothing was fetched from the web and no file was written. The example is fictional. The incident is invented. The model could not save the document to its usual path and printed it inline instead.

하는 일

장애나 심각한 성능 저하 뒤의 포스트모템 작성을 돕습니다. 사고 기본 정보를 모으고, 정확한 타임라인을 재구성해 빈틈을 짚고, 5 Whys 분석으로 근본 원인과 기여 요인을 구분하고, 영향(탐지부터 해결까지의 시간, 피크 오류율, 영향받은 트래픽 비율, 비즈니스 영향)을 수치로 정리하도록 돕고, 각 원인을 담당자와 날짜가 정해진 구체적인 개선 항목으로 바꿉니다. 톤은 비난하지 않는 것을 원칙으로 합니다. 문제는 사람이 아니라 시스템이라는 관점입니다.

동작 방식

  1. 모델이 쓰기 전에 제목, 시각, 심각도, 영향받은 서비스, 대략적인 타임라인을 모읍니다.
  2. 타임라인, 근본 원인, 기여 요인, 영향을 함께 정리합니다.
  3. 요약, 영향 표, 타임라인, 근본 원인, 기여 요인, 잘한 점, 더 나아질 수 있었던 점, 개선 항목, 교훈이 담긴 완성된 문서를 만듭니다.

이런 때 좋습니다

운영 장애, 사용자에게 보이는 오류, 데이터 손실, SLA 위반, 아슬아슬하게 넘긴 사고. 48~72시간 안에 쓰는 것이 이상적입니다.

참고 및 위험

지침만 담긴 패키지입니다. 스크립트가 없고 네트워크 접속이나 계정이 필요 없습니다. 마지막 단계에서 프로젝트의 `docs/postmortems/YYYY-MM-DD-<slug>.md`에 문서를 저장하므로, 저장 위치를 알려 주거나 대화 안에서 출력해 달라고 요청하세요. 포스트모템에는 내부 시스템 세부 정보, 고객 영향 수치, 이름이 들어가기 쉬우니, 무엇을 공유하는지와 누가 파일을 볼 수 있는지 확인하세요. 모델이 채운 담당자와 날짜는 팀이 확인하기 전까지는 임시 값입니다. 정식 장애 대응이나 규정 준수 절차를 대신하지 않습니다.