I can't run commands or create files here, so nothing below has been executed. The "expected output" blocks are predictions in vitest's format. I also haven't looked at the repo's package.json, so I'm assuming vitest and src/slugify.ts plus src/slugify.test.ts. With jest the commands and output would differ slightly.
Cycle 1: lowercasing
RED. The break this test catches is slugify returning the input unchanged. The expected value is a hand-written literal.
// src/slugify.test.ts
import { describe, expect, test } from 'vitest';
import { slugify } from './slugify';
describe('slugify', () => {
test('lowercases the title', () => {
expect(slugify('Hello')).toBe('hello');
});
});
Verify RED. Run npx vitest run src/slugify.test.ts. I expect an error rather than a failure, because the module doesn't exist yet:
FAIL src/slugify.test.ts
Error: Failed to resolve import "./slugify" from "src/slugify.test.ts"
The skill says an error doesn't count as a failure, so I'd add only a stub to get a real assertion failure:
// src/slugify.ts
export function slugify(title: string): string {
return '';
}
Re-run the same command. Expected:
FAIL src/slugify.test.ts > slugify > lowercases the title
AssertionError: expected '' to be 'hello'
Tests 1 failed (1)
The feature is missing and there are no typos, so this is the right failure.
GREEN. This is the minimal code:
export function slugify(title: string): string {
return title.toLowerCase();
}
Re-run. Expected: Tests 1 passed (1).
Cycle 2: trimming
RED. The break caught is leading or trailing whitespace surviving.
test('trims surrounding whitespace', () => {
expect(slugify(' Hello ')).toBe('hello');
});
Verify RED. Run the same command. Expected:
FAIL ... > trims surrounding whitespace
AssertionError: expected ' hello ' to be 'hello'
Tests 1 failed | 1 passed (2)
GREEN.
return title.toLowerCase().trim();
Verify GREEN. Expected: Tests 2 passed (2).
Cycle 3: runs of spaces become a single hyphen
RED. The break caught is spaces being left in place, or each space becoming its own hyphen. The input has a single space and a run of three, so a / /g replacement would fail it.
test('turns each run of spaces into a single hyphen', () => {
expect(slugify('hello world')).toBe('hello-world');
expect(slugify('hello big world')).toBe('hello-big-world');
});
Verify RED. Expected:
AssertionError: expected 'hello world' to be 'hello-world'
Tests 1 failed | 2 passed (3)
GREEN.
return title.toLowerCase().trim().replace(/ +/g, '-');
Verify GREEN. Expected: Tests 3 passed (3).
Cycle 4: stripping disallowed characters
RED. The breaks caught are punctuation surviving, and an over-eager strip that removes existing hyphens. Using well-known in the input covers the second one.
test('removes characters that are not letters, digits or hyphens', () => {
expect(slugify('Hello, World!')).toBe('hello-world');
expect(slugify('Well-Known Thing #1')).toBe('well-known-thing-1');
});
Verify RED. Expected:
AssertionError: expected 'hello,-world!' to be 'hello-world'
Tests 1 failed | 3 passed (4)
GREEN. The strip runs after lowercasing, so uppercase letters aren't removed by mistake.
return title
.toLowerCase()
.trim()
.replace(/ +/g, '-')
.replace(/[^a-z0-9-]/g, '');
Verify GREEN. Expected: Tests 4 passed (4).
Refactor
I'd split the pipeline into one named step per line with no behavior change. I don't think it's worth doing here, since the chain is already four short lines. Re-run to confirm 4 passed if you do.
Final verification
The skill requires the whole project suite, not just my file. Run npx vitest run (or npm test if the repo defines it). I expect all test files to pass with no warnings. Any failure in a pre-existing test should be reported by name, even if slugify didn't cause it.
Mutation check
| Mutation | Test that fails |
|---|---|
Remove .toLowerCase() |
Cycle 1 |
Remove .trim() |
Cycle 2 |
/ +/g → / /g |
Cycle 3, the three-space input |
| Remove the strip step | Cycle 4 |
Strip regex without - in the allowed set |
Cycle 4, well-known |
| Strip before lowercasing | Cycle 4, H and W would be removed |
Every realistic mutation is caught.
Decisions for you
I left these out because you didn't specify them, and each would need its own failing test first:
- Non-ASCII letters:
cafécurrently becomescaf, because "letters" is implemented asa-z. Unicode support would use\p{L}with theuflag. - Adjacent hyphens:
Rock & Rollbecomesrock--roll, because the space-to-hyphen step runs before the strip. Collapsing repeated hyphens, or trimming leading and trailing ones, is a separate behavior. - Tabs and newlines: only the space character is treated as a separator, as specified.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model could not run the tests, so it says what it would run and what output it expects at each step. The model opened one of the skill's reference files.
它做什麼
在功能開發、修 bug、重構與行為變更時強制執行測試驅動開發。鐵律是:沒有先失敗的測試,就不許寫正式環境程式碼。循環是:為一個行為寫一個最小的測試,執行並確認它因為正確的原因失敗,寫出最簡單能通過的程式碼,再次執行並確認整個專案的測試套件都是綠的(而不只是你那個檔案),然後在測試保持綠色的前提下重構。其中有好壞範例、常見藉口對照表(「太簡單不用測」「我之後再補測試」)、危險訊號清單、驗證清單與一個修 bug 的範例。另一份檔案 writing-good-tests.md 講怎樣寫出真正能抓住問題的測試:先說出測試要抓的那個破壞、測真實行為而不是測 mock、不把僅用於測試的程式碼放進正式類別,以及突變檢查。
適合什麼場景
任何你想讓測試真正能證明點什麼的改動,以及修了就不再復發的 bug 修復。
中風險:它會讓模型執行你的測試指令,而測試會執行你專案的程式碼。它最嚴格的規則是:先於測試寫出的程式碼要刪除並重新開始(「刪除就是刪除」);這針對的是模型剛寫的程式碼,但為求穩妥,請先提交或備份你的工作,並提前告訴它允許的例外(一次性原型、產生的程式碼、設定檔)。它要求執行整個專案的測試套件,所以速度慢或具破壞性的測試所帶來的副作用要由你承擔。語氣嚴格而絕對,這是設計如此。已試用一次,模型無法執行指令,只說明了它會執行什麼。