I can't run commands or create files here, so nothing below has been executed. The "expected output" blocks are predictions in vitest's format. I also haven't looked at the repo's package.json, so I'm assuming vitest and src/slugify.ts plus src/slugify.test.ts. With jest the commands and output would differ slightly.
Cycle 1: lowercasing
RED. The break this test catches is slugify returning the input unchanged. The expected value is a hand-written literal.
// src/slugify.test.ts
import { describe, expect, test } from 'vitest';
import { slugify } from './slugify';
describe('slugify', () => {
test('lowercases the title', () => {
expect(slugify('Hello')).toBe('hello');
});
});
Verify RED. Run npx vitest run src/slugify.test.ts. I expect an error rather than a failure, because the module doesn't exist yet:
FAIL src/slugify.test.ts
Error: Failed to resolve import "./slugify" from "src/slugify.test.ts"
The skill says an error doesn't count as a failure, so I'd add only a stub to get a real assertion failure:
// src/slugify.ts
export function slugify(title: string): string {
return '';
}
Re-run the same command. Expected:
FAIL src/slugify.test.ts > slugify > lowercases the title
AssertionError: expected '' to be 'hello'
Tests 1 failed (1)
The feature is missing and there are no typos, so this is the right failure.
GREEN. This is the minimal code:
export function slugify(title: string): string {
return title.toLowerCase();
}
Re-run. Expected: Tests 1 passed (1).
Cycle 2: trimming
RED. The break caught is leading or trailing whitespace surviving.
test('trims surrounding whitespace', () => {
expect(slugify(' Hello ')).toBe('hello');
});
Verify RED. Run the same command. Expected:
FAIL ... > trims surrounding whitespace
AssertionError: expected ' hello ' to be 'hello'
Tests 1 failed | 1 passed (2)
GREEN.
return title.toLowerCase().trim();
Verify GREEN. Expected: Tests 2 passed (2).
Cycle 3: runs of spaces become a single hyphen
RED. The break caught is spaces being left in place, or each space becoming its own hyphen. The input has a single space and a run of three, so a / /g replacement would fail it.
test('turns each run of spaces into a single hyphen', () => {
expect(slugify('hello world')).toBe('hello-world');
expect(slugify('hello big world')).toBe('hello-big-world');
});
Verify RED. Expected:
AssertionError: expected 'hello world' to be 'hello-world'
Tests 1 failed | 2 passed (3)
GREEN.
return title.toLowerCase().trim().replace(/ +/g, '-');
Verify GREEN. Expected: Tests 3 passed (3).
Cycle 4: stripping disallowed characters
RED. The breaks caught are punctuation surviving, and an over-eager strip that removes existing hyphens. Using well-known in the input covers the second one.
test('removes characters that are not letters, digits or hyphens', () => {
expect(slugify('Hello, World!')).toBe('hello-world');
expect(slugify('Well-Known Thing #1')).toBe('well-known-thing-1');
});
Verify RED. Expected:
AssertionError: expected 'hello,-world!' to be 'hello-world'
Tests 1 failed | 3 passed (4)
GREEN. The strip runs after lowercasing, so uppercase letters aren't removed by mistake.
return title
.toLowerCase()
.trim()
.replace(/ +/g, '-')
.replace(/[^a-z0-9-]/g, '');
Verify GREEN. Expected: Tests 4 passed (4).
Refactor
I'd split the pipeline into one named step per line with no behavior change. I don't think it's worth doing here, since the chain is already four short lines. Re-run to confirm 4 passed if you do.
Final verification
The skill requires the whole project suite, not just my file. Run npx vitest run (or npm test if the repo defines it). I expect all test files to pass with no warnings. Any failure in a pre-existing test should be reported by name, even if slugify didn't cause it.
Mutation check
| Mutation | Test that fails |
|---|---|
Remove .toLowerCase() |
Cycle 1 |
Remove .trim() |
Cycle 2 |
/ +/g → / /g |
Cycle 3, the three-space input |
| Remove the strip step | Cycle 4 |
Strip regex without - in the allowed set |
Cycle 4, well-known |
| Strip before lowercasing | Cycle 4, H and W would be removed |
Every realistic mutation is caught.
Decisions for you
I left these out because you didn't specify them, and each would need its own failing test first:
- Non-ASCII letters:
cafécurrently becomescaf, because "letters" is implemented asa-z. Unicode support would use\p{L}with theuflag. - Adjacent hyphens:
Rock & Rollbecomesrock--roll, because the space-to-hyphen step runs before the strip. Collapsing repeated hyphens, or trimming leading and trailing ones, is a separate behavior. - Tabs and newlines: only the space character is treated as a separator, as specified.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The example is fictional. The model could not run the tests, so it says what it would run and what output it expects at each step. The model opened one of the skill's reference files.
What it does
Enforces test-driven development on features, bug fixes, refactors and behaviour changes. The iron law is no production code without a failing test first. The loop is: write one minimal test for one behaviour, run it and confirm it fails for the right reason, write the simplest code that passes, run it again and confirm the whole project suite is green (not just your file), then refactor with the tests staying green. It includes good and bad examples, a table of common rationalisations ("too simple to test", "I'll test after"), a red-flag list, a verification checklist and a bug-fix example. A second file, writing-good-tests.md, covers naming the break a test should catch, testing real behaviour rather than mocks, keeping test-only code out of production classes, and a mutation check.
Good for
Any change where you want tests that actually prove something, and bug fixes that stay fixed.
Medium risk: it has the model run your test commands, which execute your project's code. Its strictest rule says that code written before its test must be deleted and started over ("delete means delete"); that is aimed at code the model just wrote, but to be safe commit or back up your work first and tell it up front about exceptions it allows (throwaway prototypes, generated code, config files). It asks for the full project suite to be run, so side effects of slow or destructive tests are on you. The tone is strict and absolute by design. Trialled once; the model could not run commands and said what it would run.