Start / Skills / Entwickler-Produktivität / Investigate First
Entwickler-Produktivität

Investigate First

Julius BrusseeApache-2.0Sprache des Skills: Englisch
Lizenz geprüftKeine Skripte
DemoEchter Lauf · claude-opus-5-5 · Oktober 2026
Our checkout page sometimes shows another user's cart, maybe once in a few hundred requests, only in production. I don't know why yet and I don't want a fix until we know the cause. Here is the relevant code: ```python # cart/views.py (Django, gunicorn with 4 threads per worker) _cart_cache = {} class CartView(View): current_user = None def get(self, request): self.__class__.current_user = request.user cart = _cart_cache.get(request.user.id) if cart is None: cart = load_cart(CartView.current_user) _cart_cache[request.user.id] = cart return render(request, "cart.html", {"cart": cart}) ``` Staging runs gunicorn with `--threads 1`. Production uses `--threads 4`. What is causing this?
Skill: investigate-first

Cause

It's a race on shared state between threads. current_user is a class attribute, so every thread in a gunicorn worker process reads and writes the same slot. With --threads 4, two requests can interleave like this:

Thread A (user Alice) Thread B (user Bob)
CartView.current_user = Alice
_cart_cache.get(alice.id) → miss
CartView.current_user = Bob
load_cart(CartView.current_user) → loads Bob's cart
_cart_cache[alice.id] = Bob's cart

Alice is shown Bob's cart. Two details make it worse:

  1. The bad entry sticks. The wrong cart is stored in _cart_cache under Alice's id. Her later requests to that worker hit the cache and keep showing Bob's cart until the process restarts. So one short race turns into a lasting leak for that user on that worker.
  2. The key and the value come from different sources. The cache key uses the thread-local request.user.id, but the value is loaded from the shared CartView.current_user. Nothing checks that the two refer to the same user.

Why only production, and why rarely

  • Staging (--threads 1): each worker process handles one request at a time, so nothing else can overwrite current_user between the write and the read. The bug can't happen there. Separate gunicorn workers are separate processes, so they don't share the class attribute.

  • Production (--threads 4): requests share the attribute. It only goes wrong when all three of these line up:

    • a cache miss (first visit to that worker, or after a restart)
    • a second request writes current_user in the short gap before load_cart reads it
    • the GIL switches threads at exactly that point

    That combination fits "once in a few hundred requests."

How to confirm it before changing anything

These checks are cheap. I haven't run any of them; this diagnosis comes from reading the code.

  1. Log the mismatch. Temporarily log a warning in production or in a load test whenever CartView.current_user.id != request.user.id just before load_cart. Any hit proves this mechanism.
  2. Reproduce in staging. Run staging with --threads 4 and send concurrent requests from many different users to the cart page. A fresh worker gives the most cache misses. Then check which user's items each response contains. Running the same load with --threads 1 should show zero leaks.
  3. Force it in a test. Make load_cart sleep briefly or wait on a threading.Barrier, then call CartView().get() from two threads with different users. This reproduces it every time.
  4. Check that it persists. Once a user has seen a wrong cart, their repeat requests to the same worker should keep showing that same wrong cart. Restarting the worker should clear it. If that's what happens, the cache is what turns the race into a lasting leak.

If check 1 or 2 fires, the cause is confirmed. Note that any fix will also need to clear the poisoned entries in _cart_cache, either by restarting the workers or by discarding the cache. Holding off on the fix as you asked.

Real run recorded with claude-code / claude-opus-5-5. Output is shown verbatim, unmodified.

Was es macht

Eine kurze Arbeitsregel für unklare Fehler: erst Belege sammeln und die Ursache festnageln, dann Produktcode ändern.

So funktioniert es

  1. Trennt das beobachtete Symptom von der vermuteten Ursache.
  2. Verfolgt Eingaben, Zustandsänderungen und Zuständigkeitsgrenzen und ordnet Hypothesen nach Beweislage und danach, wie günstig sie sich ausschließen lassen.
  3. Ändert nichts, bis ein Mechanismus alle Belege erklärt, und berichtet dann Ursache und Nachweis. Repariert wird nur, wenn die Aufgabe das verlangt.

Geeignet für

Sporadische Bugs, Performance-Regressionen und Fälle à la "läuft auf Staging, fällt in Produktion aus".

Gut zu wissen

Eines der allgemeinen Arbeitsmuster aus dem Open-Source-Projekt Caveman. Die Anweisungen selbst tragen kein Caveman-Branding und funktionieren in jedem Projekt.

Hinweise & Risiken

Reine Anweisungsdatei: keine Skripte, kein Netzwerkzugriff, keine Dateischreibvorgänge. Das Paket enthält außerdem LICENSE, NOTICE und agents/openai.yaml (Anzeigename und Standard-Prompt für Codex).