홈 / Skills / 개발 생산성 / 체계적 디버깅 (Systematic Debugging)
개발 생산성

체계적 디버깅 (Systematic Debugging)

Jesse VincentMITSkill 언어: 영어
라이선스 확인됨스크립트 포함중간 위험
데모실행 기록 · claude-sonnet-5-5 · 2026년 10월
Using the systematic-debugging skill: our checkout API started returning intermittent 500 errors after yesterday's deploy. The log says "TypeError: Cannot read properties of undefined (reading id)" at orderService.js line 88. A teammate says to just wrap that line in a try/catch so users stop seeing errors. What should we do?
Skill: systematic-debugging

Don't add the try/catch yet. It hides the symptom without fixing anything. Checkout is a write path, so swallowing the error there is risky. If order is undefined at line 88, the code either carries on with a half-built order or returns a success to the user for an order that wasn't created. That can mean silent data loss or a charge with no order, which is worse than a visible 500.

Phase 1: find the root cause

  1. Read the full stack trace. Check what undefined is at line 88. It's the object whose .id is being read, for example order.id, user.id or cart.id. Then check who called this function and what they passed.
  2. Check what changed in yesterday's deploy. Diff the release and look for anything touching orderService.js, its callers, the shape of the data it reads (DB queries, migrations, upstream API responses), dependency bumps, and config or env changes. The timing points to the deploy, so that diff is the highest-value evidence.
  3. Work out why it's intermittent. Intermittent usually means the failure depends on the input or the timing. Compare failing and succeeding requests for these differences:
    • Particular users, carts or payment methods, such as guest vs. logged-in, empty cart, or a coupon.
    • A race condition, such as a read before a write commits, or a missing await.
    • A cache miss, or one instance or region running a different version or config.
    • A lookup that returns null or undefined for some records, such as a .find() with no match or a missing relation.
  4. Add logging if the logs don't answer this. Log the inputs entering and leaving the function that contains line 88. Log the request ID, user, cart contents, and the result of whatever lookup feeds line 88. Then collect several failing and passing cases.
  5. Trace the bad value backward to where the undefined first appears. Fix it there, not at line 88.

If users are hurting right now

  • Roll back yesterday's deploy if the diff points to it or you can't find the cause quickly. It stops the errors without hiding anything, and you can debug in peace.
  • If you can't roll back, a guard is acceptable as a temporary mitigation. It should fail explicitly, with an early check that logs the context and returns a clear error or 4xx, rather than a blanket try/catch. It must not let checkout continue with bad state. Treat it as a stopgap and keep investigating.

Once you find the cause

Write a failing test that reproduces it, fix it at the source, and verify the test passes. Then consider validating at the boundary where the bad data enters, so this class of bug fails loudly and early.

The directory I'm in has no git repo, and I haven't seen orderService.js. If you point me at the repo, the deploy diff, or a few failing-request logs, I can start on steps 1–3.

Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The scenario is invented and the model had no access to the code, so it laid out the investigation steps. The bundled shell script was not run.

하는 일

철칙은 하나, 근본 원인 조사를 마치기 전에는 고치지 않는다입니다. 1단계에서는 오류를 끝까지 읽고, 재현하고, 최근 변경을 확인하고, 각 구성 요소 경계에서 증거를 모읍니다. 2단계에서는 잘 동작하는 예를 찾아 모든 차이를 나열합니다. 3단계에서는 가설 하나를 세우고 가장 작은 변경으로 검증합니다. 4단계에서는 먼저 실패하는 테스트를 쓰고, 수정은 한 곳만 하고, 검증합니다. 수정이 세 번 연속 실패하면 네 번째 패치를 하는 대신 멈추고 아키텍처를 다시 따져 보자고 요청합니다. 추측으로 이끄는 핑계("긴급해서", "한 줄만")와 절차가 건너뛰어질 때 사용자가 하는 말도 나열합니다. 보조 파일은 거꾸로 원인 추적, 다층 방어, 임의의 타임아웃을 조건 대기로 바꾸는 방법(TypeScript 예제 포함)을 다룹니다.

이런 때 좋습니다

테스트 실패, 운영 버그, 불안정한 테스트, 빌드·통합 문제, 특히 시간에 쫓길 때.

참고 및 위험

중간 위험:조사 단계에서 모델이 로그를 추가하고 진단 명령을 실행하며, 예시에 `env | grep`과 `security find-identity`가 있어 환경 변수나 자격 증명 저장소 정보가 대화에 나올 수 있으니 출력을 공유하기 전에 비밀 값을 가리세요. 동봉된 `find-polluter.sh`는 파일을 남기는 테스트를 찾으려고 일치하는 테스트 파일마다 `npm test <파일>`을 한 번씩 실행하므로, 프로젝트의 테스트와 그 부작용이 실행됩니다. 아무것도 삭제하지 않고 네트워크도 쓰지 않습니다. 이 사이트의 등록본에는 저자 자신의 테스트 시나리오와 제작 기록을 뺐습니다. 시험 실행은 코드 없이 설명한 버그로만 했고 스크립트는 실행하지 않았습니다.