I fixed it: an invoice issued Friday 2026-10-02 now gets a due date of Friday 2026-10-16, and all 4 tests pass.
Cause: invoice_due_date in billing/due_dates.py had an extra step, commented "Invoices issued late in the week reach customers after the weekend." It added one calendar day whenever the invoice was issued on a Friday, Saturday or Sunday (issued.weekday() >= 4). The business-day count by itself was already correct (2026-10-16). Adding one calendar day to that Friday moved it to Saturday.
How I confirmed it:
- A one-line check of the reported invoice returned
2026-10-17 Saturday, matching the customer reports. - Running only the business-day count (
add_business_days) gave 2026-10-16, so that part was fine. - I checked every issue weekday against terms of 1–30 days. 18 combinations gave a weekend due date, and every one was issued Fri/Sat/Sun, which are exactly the days the extra step applied to.
Fix: I removed that extra step, so invoice_due_date now just returns add_business_days(issued, terms_days). A Saturday or Sunday invoice is still handled sensibly, because the count already starts from Monday.
Regression tests, added to tests/test_due_dates.py. Both failed before the fix and pass now:
test_invoice_issued_on_friday_is_due_on_a_friday: the exact case Finance reported.test_invoice_due_date_never_falls_on_a_weekend: checks every issue weekday with terms of 1–30 days.
For Finance to confirm: this removes a deliberate rule. Besides the weekend problem, it also moved some weekday due dates one day later (e.g. a Friday invoice with 3-day terms was due Thursday, and is now due Wednesday). Finance's expected date implies no grace day at all, so I removed it entirely. If they do want extra time for late-week invoices, it should be added as a business day instead (which would make the example due Monday 2026-10-19); that conflicts with the date they gave.
I haven't committed anything; the changes are in your working tree.
Code changes recorded in the trial sandbox (git diff):
diff --git a/billing/due_dates.py b/billing/due_dates.py
index 35d3932..1fd87e9 100644
--- a/billing/due_dates.py
+++ b/billing/due_dates.py
@@ -14,8 +14,4 @@ def add_business_days(start: date, days: int) -> date:
def invoice_due_date(issued: date, terms_days: int = 10) -> date:
"""Due date for an invoice with `terms_days` business-day payment terms."""
- due = add_business_days(issued, terms_days)
- # Invoices issued late in the week reach customers after the weekend.
- if issued.weekday() >= 4:
- due += timedelta(days=1)
- return due
+ return add_business_days(issued, terms_days)
diff --git a/tests/test_due_dates.py b/tests/test_due_dates.py
index 497750a..4d1cbde 100644
--- a/tests/test_due_dates.py
+++ b/tests/test_due_dates.py
@@ -1,4 +1,4 @@
-from datetime import date
+from datetime import date, timedelta
from billing.due_dates import add_business_days, invoice_due_date
@@ -9,3 +9,14 @@ def test_add_business_days_skips_weekend():
def test_invoice_issued_on_monday():
assert invoice_due_date(date(2026, 9, 28)) == date(2026, 10, 12)
+
+
+def test_invoice_issued_on_friday_is_due_on_a_friday():
+ assert invoice_due_date(date(2026, 10, 2)) == date(2026, 10, 16)
+
+
+def test_invoice_due_date_never_falls_on_a_weekend():
+ for offset in range(7):
+ issued = date(2026, 9, 28) + timedelta(days=offset)
+ for terms in range(1, 31):
+ assert invoice_due_date(issued, terms).weekday() < 5, (issued, terms)
Real run recorded with claude-code / claude-opus-5-5. Run in a temporary directory containing a sample project, with Edit and Bash enabled. Tool-call lines show only the first line of each command. The last block is the git diff recorded when the trial ended. Both model outputs are unmodified.
Was es macht
Ein diszipliniertes Vorgehen für harte Bugs und Performance-Regressionen. Die Kernthese: Sobald du ein klares Bestanden/Fehlgeschlagen-Signal für genau diesen Bug hast, ist die Ursachensuche Routine, also fließt der meiste Aufwand zuerst in dieses Signal.
So funktioniert es
- Feedback-Schleife: Einen Befehl bauen, der bei genau diesem Bug rot wird, deterministisch und schnell ist und vom Agenten unbeaufsichtigt laufen kann. Vorher wird nicht spekuliert.
- Reproduzieren und verkleinern: Prüfen, dass er dein Symptom zeigt, dann Eingaben streichen, bis jedes verbleibende Teil unverzichtbar ist.
- Hypothesen: 3 bis 5 priorisierte, widerlegbare Hypothesen aufstellen und dir vor dem Testen zeigen.
- Messen: Immer nur eine Variable ändern, Debugger vor Logs, jedes Debug-Log mit eindeutigem Präfix, damit eine Suche alles entfernt.
- Beheben: Wo es eine passende Test-Naht gibt, zuerst den Regressionstest schreiben, dann beheben.
- Aufräumen: Die ursprüngliche Schleife erneut laufen lassen, Debug-Code entfernen und die bestätigte Ursache in die Commit-Nachricht schreiben.
Geeignet für
Sporadische Bugs, "früher ging es"-Regressionen und langsamen Code.
Gut zu wissen
Es liest GLOSSARY.md und ADRs, falls das Repo sie hat. In unserem Test entfernte es eine fehlerhafte Anpassung und ergänzte zwei Tests.
Enthält eine Bash-Vorlage (hitl-loop.template.sh), die nur Hinweise ausgibt und auf Enter oder eine Eingabe wartet; wir haben sie gelesen: keine Netzwerkaufrufe, keine Dateischreibvorgänge. Der Skill selbst lässt den Agenten Befehle ausführen, temporäre Debug-Logs einfügen, einen Regressionstest schreiben und deinen Code ändern, verändert also dein Repository. Er weist den Agenten an, Geheimnisse in allem Angezeigten zu schwärzen. Das Paket enthält außerdem die MIT-Lizenz (LICENSE) des Original-Repositorys und agents/openai.yaml (Anzeigename für Codex).