Discount codes are now in the cart, and all 5 tests pass (python3 -m pytest -q). I built it one test at a time, writing each test first and watching it fail before adding the code to pass it. Every test uses only apply_code and total().
Tests added (tests/test_cart.py); the cart in each is 2 × 12.50, so the subtotal is 25.00:
SAVE10gives a total of 27.50 (22.50 after the discount, plus 5.00 shipping).FREESHIPgives a total of 25.00.- An unknown code raises
InvalidCode, and the total stays at 30.00. - After one code is applied, a second
apply_coderaisesCodeAlreadyAppliedand the first code stays in effect (total stays 27.50).
Code changes (shop/cart.py): a list of known codes, the InvalidCode and CodeAlreadyApplied errors, a code field on Cart, and apply_code. total() now applies the 10% discount to the subtotal or drops the shipping fee, depending on the code.
Decisions for you:
- What a second code does: "only one code per cart" could mean the second code is rejected or that it replaces the first. I made it reject, using the new
CodeAlreadyAppliederror. If you'd rather the second code replace the first, that's a one-line change. - Rounding: nothing rounds the
SAVE10result yet. A subtotal like 12.35 would give a total of 16.115. I left it alone because rounding wasn't in the spec; tell me the rule (e.g. half-up to 0.01) and I'll add a test for it first.
I haven't committed anything.
Code changes recorded in the trial sandbox (git diff):
diff --git a/shop/cart.py b/shop/cart.py
index 937c5d1..dbc57bd 100644
--- a/shop/cart.py
+++ b/shop/cart.py
@@ -2,6 +2,15 @@ from dataclasses import dataclass, field
from decimal import Decimal
SHIPPING_FEE = Decimal("5.00")
+CODES = {"SAVE10", "FREESHIP"}
+
+
+class InvalidCode(ValueError):
+ pass
+
+
+class CodeAlreadyApplied(ValueError):
+ pass
@dataclass
@@ -14,6 +23,7 @@ class LineItem:
@dataclass
class Cart:
items: list[LineItem] = field(default_factory=list)
+ code: str | None = None
def add(self, sku: str, price: str, qty: int = 1) -> None:
self.items.append(LineItem(sku, Decimal(price), qty))
@@ -21,5 +31,16 @@ class Cart:
def subtotal(self) -> Decimal:
return sum((i.price * i.qty for i in self.items), Decimal("0"))
+ def apply_code(self, code: str) -> None:
+ if code not in CODES:
+ raise InvalidCode(code)
+ if self.code is not None:
+ raise CodeAlreadyApplied(self.code)
+ self.code = code
+
def total(self) -> Decimal:
- return self.subtotal() + SHIPPING_FEE
+ subtotal = self.subtotal()
+ if self.code == "SAVE10":
+ subtotal -= subtotal * Decimal("0.10")
+ shipping = Decimal("0") if self.code == "FREESHIP" else SHIPPING_FEE
+ return subtotal + shipping
diff --git a/tests/test_cart.py b/tests/test_cart.py
index e0c976c..199e238 100644
--- a/tests/test_cart.py
+++ b/tests/test_cart.py
@@ -1,9 +1,42 @@
from decimal import Decimal
-from shop.cart import Cart
+import pytest
+
+from shop.cart import Cart, CodeAlreadyApplied, InvalidCode
def test_total_adds_shipping():
cart = Cart()
cart.add("mug", "12.50", qty=2)
assert cart.total() == Decimal("30.00")
+
+
+def test_save10_takes_ten_percent_off_subtotal():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ cart.apply_code("SAVE10")
+ assert cart.total() == Decimal("27.50")
+
+
+def test_freeship_removes_shipping_fee():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ cart.apply_code("FREESHIP")
+ assert cart.total() == Decimal("25.00")
+
+
+def test_unknown_code_raises_invalid_code():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ with pytest.raises(InvalidCode):
+ cart.apply_code("BOGUS")
+ assert cart.total() == Decimal("30.00")
+
+
+def test_only_one_code_per_cart():
+ cart = Cart()
+ cart.add("mug", "12.50", qty=2)
+ cart.apply_code("SAVE10")
+ with pytest.raises(CodeAlreadyApplied):
+ cart.apply_code("FREESHIP")
+ assert cart.total() == Decimal("27.50")
Real run recorded with claude-code / claude-opus-5-5. Run in a temporary directory containing a sample project, with Edit and Bash enabled. Tool-call lines show only the first line of each command. The last block is the git diff recorded when the trial ended. Both model outputs are unmodified.
它做什麼
引導代理走「紅到綠」的循環,更重要的是講清楚什麼樣的測試才值得保留。附帶兩份參考筆記:好測試與壞測試,以及什麼時候該 mock。
運作方式
- 先約定切入點: 撰寫任何測試之前,先列出要透過哪些公開介面來測,並與你確認。未確認的切入點不寫測試。
- 一次一小片: 一個測試、一段最小實作,反覆進行,避免一口氣先把所有測試寫完。
- 先紅後綠: 先有失敗的測試,再只寫剛好能讓它通過的程式碼。
- 避開三種反模式: 與實作細節耦合的測試;期望值與程式碼用同樣方法算出來、永遠不會不一致的「自證」測試;以及先寫完所有測試再寫實作的橫向切片。
- 只在系統邊界處 mock: 如外部 API、時間、隨機數,絕不 mock 你自己的模組。
適合情境
做新功能與修 bug,並希望測試在重構後仍然有效。
需要了解
重構被刻意排除在循環之外,屬於審查階段。它提到作者的 codebase-design skill 用於介面設計,但僅為選讀參考。
純指令檔:沒有腳本、不連網。 Skill 會讓代理在你的儲存庫中撰寫測試與實作程式碼,並執行你的測試指令。 壓縮檔中另附原始儲存庫的 MIT 授權檔 LICENSE 與 agents/openai.yaml(供 Codex 使用的顯示名稱)。