Cause
It's a race on shared state between threads. current_user is a class attribute, so every thread in a gunicorn worker process reads and writes the same slot. With --threads 4, two requests can interleave like this:
| Thread A (user Alice) | Thread B (user Bob) |
|---|---|
CartView.current_user = Alice |
|
_cart_cache.get(alice.id) → miss |
|
CartView.current_user = Bob |
|
load_cart(CartView.current_user) → loads Bob's cart |
|
_cart_cache[alice.id] = Bob's cart |
Alice is shown Bob's cart. Two details make it worse:
- The bad entry sticks. The wrong cart is stored in
_cart_cacheunder Alice's id. Her later requests to that worker hit the cache and keep showing Bob's cart until the process restarts. So one short race turns into a lasting leak for that user on that worker. - The key and the value come from different sources. The cache key uses the thread-local
request.user.id, but the value is loaded from the sharedCartView.current_user. Nothing checks that the two refer to the same user.
Why only production, and why rarely
-
Staging (
--threads 1): each worker process handles one request at a time, so nothing else can overwritecurrent_userbetween the write and the read. The bug can't happen there. Separate gunicorn workers are separate processes, so they don't share the class attribute. -
Production (
--threads 4): requests share the attribute. It only goes wrong when all three of these line up:- a cache miss (first visit to that worker, or after a restart)
- a second request writes
current_userin the short gap beforeload_cartreads it - the GIL switches threads at exactly that point
That combination fits "once in a few hundred requests."
How to confirm it before changing anything
These checks are cheap. I haven't run any of them; this diagnosis comes from reading the code.
- Log the mismatch. Temporarily log a warning in production or in a load test whenever
CartView.current_user.id != request.user.idjust beforeload_cart. Any hit proves this mechanism. - Reproduce in staging. Run staging with
--threads 4and send concurrent requests from many different users to the cart page. A fresh worker gives the most cache misses. Then check which user's items each response contains. Running the same load with--threads 1should show zero leaks. - Force it in a test. Make
load_cartsleep briefly or wait on athreading.Barrier, then callCartView().get()from two threads with different users. This reproduces it every time. - Check that it persists. Once a user has seen a wrong cart, their repeat requests to the same worker should keep showing that same wrong cart. Restarting the worker should clear it. If that's what happens, the cache is what turns the race into a lasting leak.
If check 1 or 2 fires, the cause is confirmed. Note that any fix will also need to clear the poisoned entries in _cart_cache, either by restarting the workers or by discarding the cache. Holding off on the fix as you asked.
Real run recorded with claude-code / claude-opus-5-5. Output is shown verbatim, unmodified.
하는 일
원인이 분명하지 않은 장애를 위한 짧은 작업 규칙입니다. 증거를 모아 원인을 확정한 뒤에 제품 코드를 수정합니다.
작동 방식
- 관찰된 증상과 추정되는 원인을 구분합니다.
- 입력, 상태 변화, 책임 경계를 따라가며, 가설을 증거의 강도와 반증 비용 순으로 정리합니다.
- 모든 증거를 설명하는 메커니즘을 찾기 전에는 코드를 고치지 않고, 찾으면 원인과 근거를 보고합니다. 수정 요청이 없으면 고치지 않습니다.
이럴 때 좋습니다
간헐적인 버그, 성능 저하, '스테이징에서는 되는데 운영에서만 실패하는' 문제.
알아 둘 점
오픈소스 프로젝트 Caveman에 포함된 범용 작업 패턴 중 하나입니다. 지시문 자체에는 Caveman 브랜드 요소가 없어 어떤 프로젝트에서도 쓸 수 있습니다.
지시문만 담긴 파일입니다. 스크립트가 없고, 네트워크에 연결하지 않으며, 파일을 쓰지 않습니다. 패키지에는 LICENSE, NOTICE, agents/openai.yaml(Codex용 표시 이름과 기본 프롬프트)도 들어 있습니다.