Discount codes are now in the cart, and all 5 tests pass (python3 -m pytest -q). I built it one test at a time, writing each test first and watching it fail before adding the code to pass it. Every test uses only apply_code and total().
Tests added (tests/test_cart.py); the cart in each is 2 × 12.50, so the subtotal is 25.00:
SAVE10gives a total of 27.50 (22.50 after the discount, plus 5.00 shipping).FREESHIPgives a total of 25.00.- An unknown code raises
InvalidCode, and the total stays at 30.00. - After one code is applied, a second
apply_coderaisesCodeAlreadyAppliedand the first code stays in effect (total stays 27.50).
Code changes (shop/cart.py): a list of known codes, the InvalidCode and CodeAlreadyApplied errors, a code field on Cart, and apply_code. total() now applies the 10% discount to the subtotal or drops the shipping fee, depending on the code.
Decisions for you:
- What a second code does: "only one code per cart" could mean the second code is rejected or that it replaces the first. I made it reject, using the new
CodeAlreadyAppliederror. If you'd rather the second code replace the first, that's a one-line change. - Rounding: nothing rounds the
SAVE10result yet. A subtotal like 12.35 would give a total of 16.115. I left it alone because rounding wasn't in the spec; tell me the rule (e.g. half-up to 0.01) and I'll add a test for it first.
I haven't committed anything.
Code changes recorded in the trial sandbox (git diff):
diff --git a/shop/cart.py b/shop/cart.py
index 937c5d1..dbc57bd 100644
--- a/shop/cart.py
+++ b/shop/cart.py
@@ -2,6 +2,15 @@ from dataclasses import dataclass, field
from decimal import Decimal
SHIPPING_FEE = Decimal("5.00")
+CODES = {"SAVE10", "FREESHIP"}
+
+
+class InvalidCode(ValueError):
+ pass
+
+
+class CodeAlreadyApplied(ValueError):
+ pass
@dataclass
@@ -14,6 +23,7 @@ class LineItem:
@dataclass
class Cart:
items: list[LineItem] = field(default_factory=list)
+ code: str | None = None
def add(self, sku: str, price: str, qty: int = 1) -> None:
self.items.append(LineItem(sku, Decimal(price), qty))
@@ -21,5 +31,16 @@ class Cart:
def subtotal(self) -> Decimal:
return sum((i.price * i.qty for i in self.items), Decimal("0"))
+ def apply_code(self, code: str) -> None:
+ if code not in CODES:
+ raise InvalidCode(code)
+ if self.code is not None:
+ raise CodeAlreadyApplied(self.code)
+ self.code = code
+
def total(self) -> Decimal:
- return self.subtotal() + SHIPPING_FEE
+ subtotal = self.subtotal()
+ if self.code == "SAVE10":
+ subtotal -= subtotal * Decimal("0.10")
+ shipping = Decimal("0") if self.code == "FREESHIP" else SHIPPING_FEE
+ return subtotal + shipping
diff --git a/tests/test_cart.py b/tests/test_cart.py
index e0c976c..199e238 100644
--- a/tests/test_cart.py
+++ b/tests/test_cart.py
@@ -1,9 +1,42 @@
from decimal import Decimal
-from shop.cart import Cart
+import pytest
+
+from shop.cart import Cart, CodeAlreadyApplied, InvalidCode
def test_total_adds_shipping():
cart = Cart()
cart.add("mug", "12.50", qty=2)
assert cart.total() == Decimal("30.00")
+
+
+def test_save10_takes_ten_percent_off_subtotal():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ cart.apply_code("SAVE10")
+ assert cart.total() == Decimal("27.50")
+
+
+def test_freeship_removes_shipping_fee():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ cart.apply_code("FREESHIP")
+ assert cart.total() == Decimal("25.00")
+
+
+def test_unknown_code_raises_invalid_code():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ with pytest.raises(InvalidCode):
+ cart.apply_code("BOGUS")
+ assert cart.total() == Decimal("30.00")
+
+
+def test_only_one_code_per_cart():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ cart.apply_code("SAVE10")
+ with pytest.raises(CodeAlreadyApplied):
+ cart.apply_code("FREESHIP")
+ assert cart.total() == Decimal("27.50")
Real run recorded with claude-code / claude-opus-5-5. Run in a temporary directory containing a sample project, with Edit and Bash enabled. Tool-call lines show only the first line of each command. The last block is the git diff recorded when the trial ended. Both model outputs are unmodified.
What it does
Guides the agent through a red-green loop and, more importantly, what makes the tests worth keeping. It ships with two reference notes, one on good and bad tests and one on when to mock.
How it works
- Agree the seams first: before any test is written, it names the public interfaces to test through and confirms them with you. No test is written at an unconfirmed seam.
- One slice at a time: one test, one minimal implementation, repeat. It avoids writing all tests up front.
- Red before green: the failing test comes first, then only enough code to pass it.
- Avoids three anti-patterns: tests coupled to implementation details, tautological tests whose expected value is computed the same way as the code, and horizontal slicing.
- Mocks only system boundaries such as external APIs, time and randomness, never your own modules.
Good for
New features and bug fixes where you want tests that survive refactoring.
Worth knowing
Refactoring is deliberately outside the loop and belongs to review. It mentions the author's codebase-design skill for interface design, but only as optional reading.
Pure instruction files: no scripts and no network access. The skill tells the agent to write tests and implementation code in your repository and to run your test command. The package also contains the original MIT LICENSE and agents/openai.yaml (display name for Codex).