I fixed it: an invoice issued Friday 2026-10-02 now gets a due date of Friday 2026-10-16, and all 4 tests pass.
Cause: invoice_due_date in billing/due_dates.py had an extra step, commented "Invoices issued late in the week reach customers after the weekend." It added one calendar day whenever the invoice was issued on a Friday, Saturday or Sunday (issued.weekday() >= 4). The business-day count by itself was already correct (2026-10-16). Adding one calendar day to that Friday moved it to Saturday.
How I confirmed it:
- A one-line check of the reported invoice returned
2026-10-17 Saturday, matching the customer reports. - Running only the business-day count (
add_business_days) gave 2026-10-16, so that part was fine. - I checked every issue weekday against terms of 1–30 days. 18 combinations gave a weekend due date, and every one was issued Fri/Sat/Sun, which are exactly the days the extra step applied to.
Fix: I removed that extra step, so invoice_due_date now just returns add_business_days(issued, terms_days). A Saturday or Sunday invoice is still handled sensibly, because the count already starts from Monday.
Regression tests, added to tests/test_due_dates.py. Both failed before the fix and pass now:
test_invoice_issued_on_friday_is_due_on_a_friday: the exact case Finance reported.test_invoice_due_date_never_falls_on_a_weekend: checks every issue weekday with terms of 1–30 days.
For Finance to confirm: this removes a deliberate rule. Besides the weekend problem, it also moved some weekday due dates one day later (e.g. a Friday invoice with 3-day terms was due Thursday, and is now due Wednesday). Finance's expected date implies no grace day at all, so I removed it entirely. If they do want extra time for late-week invoices, it should be added as a business day instead (which would make the example due Monday 2026-10-19); that conflicts with the date they gave.
I haven't committed anything; the changes are in your working tree.
Code changes recorded in the trial sandbox (git diff):
diff --git a/billing/due_dates.py b/billing/due_dates.py
index 35d3932..1fd87e9 100644
--- a/billing/due_dates.py
+++ b/billing/due_dates.py
@@ -14,8 +14,4 @@ def add_business_days(start: date, days: int) -> date:
def invoice_due_date(issued: date, terms_days: int = 10) -> date:
"""Due date for an invoice with `terms_days` business-day payment terms."""
- due = add_business_days(issued, terms_days)
- # Invoices issued late in the week reach customers after the weekend.
- if issued.weekday() >= 4:
- due += timedelta(days=1)
- return due
+ return add_business_days(issued, terms_days)
diff --git a/tests/test_due_dates.py b/tests/test_due_dates.py
index 497750a..4d1cbde 100644
--- a/tests/test_due_dates.py
+++ b/tests/test_due_dates.py
@@ -1,4 +1,4 @@
-from datetime import date
+from datetime import date, timedelta
from billing.due_dates import add_business_days, invoice_due_date
@@ -9,3 +9,14 @@ def test_add_business_days_skips_weekend():
def test_invoice_issued_on_monday():
assert invoice_due_date(date(2026, 9, 28)) == date(2026, 10, 12)
+
+
+def test_invoice_issued_on_friday_is_due_on_a_friday():
+ assert invoice_due_date(date(2026, 10, 2)) == date(2026, 10, 16)
+
+
+def test_invoice_due_date_never_falls_on_a_weekend():
+ for offset in range(7):
+ issued = date(2026, 9, 28) + timedelta(days=offset)
+ for terms in range(1, 31):
+ assert invoice_due_date(issued, terms).weekday() < 5, (issued, terms)
Real run recorded with claude-code / claude-opus-5-5. Run in a temporary directory containing a sample project, with Edit and Bash enabled. Tool-call lines show only the first line of each command. The last block is the git diff recorded when the trial ended. Both model outputs are unmodified.
What it does
A disciplined routine for hard bugs and performance regressions. Its central claim: once you have a tight pass/fail signal for the bug, finding the cause is mechanical, so most of the effort goes into building that signal first.
How it works
- Feedback loop: build one command that goes red on this specific bug, is deterministic and fast, and that the agent can run unattended. No theorising before it exists.
- Reproduce and minimise: confirm it shows your symptom, then cut inputs until every remaining piece is load-bearing.
- Hypothesise: write 3 to 5 ranked, falsifiable hypotheses and show them to you before testing.
- Instrument: change one variable at a time, prefer a debugger over logs, and tag every debug log so it can be removed with one search.
- Fix: write the regression test first where a correct test seam exists, then fix.
- Clean up: re-run the original loop, remove debug code, and state the confirmed cause in the commit message.
Good for
Intermittent bugs, "it used to work" regressions and slow code.
Worth knowing
It reads GLOSSARY.md and ADRs if the repo has them. In our trial it removed a faulty adjustment and added two tests.
Contains one bash template (hitl-loop.template.sh) that only prints prompts and waits for you to press Enter or type an answer; we read it and it makes no network calls and writes no files. The skill itself tells the agent to run commands, add temporary debug logs, write a regression test and edit your code, so it changes your repository. It tells the agent to redact secrets in anything it shows. The package also contains the original MIT LICENSE and agents/openai.yaml (display name for Codex).