Skip to content

Commit ec00f3c

Browse files
committed
chore: add TDD skill for agents and claude code
1 parent 8c436c0 commit ec00f3c

14 files changed

Lines changed: 619 additions & 0 deletions

File tree

.agents/skills/tdd/SKILL.md

Lines changed: 107 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,107 @@
1+
---
2+
name: tdd
3+
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
4+
---
5+
6+
# Test-Driven Development
7+
8+
## Philosophy
9+
10+
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
11+
12+
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
13+
14+
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
15+
16+
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
17+
18+
## Anti-Pattern: Horizontal Slices
19+
20+
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
21+
22+
This produces **crap tests**:
23+
24+
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
25+
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
26+
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
27+
- You outrun your headlights, committing to test structure before understanding the implementation
28+
29+
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
30+
31+
```
32+
WRONG (horizontal):
33+
RED: test1, test2, test3, test4, test5
34+
GREEN: impl1, impl2, impl3, impl4, impl5
35+
36+
RIGHT (vertical):
37+
RED→GREEN: test1→impl1
38+
RED→GREEN: test2→impl2
39+
RED→GREEN: test3→impl3
40+
...
41+
```
42+
43+
## Workflow
44+
45+
### 1. Planning
46+
47+
Before writing any code:
48+
49+
- [ ] Confirm with user what interface changes are needed
50+
- [ ] Confirm with user which behaviors to test (prioritize)
51+
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
52+
- [ ] Design interfaces for [testability](interface-design.md)
53+
- [ ] List the behaviors to test (not implementation steps)
54+
- [ ] Get user approval on the plan
55+
56+
Ask: "What should the public interface look like? Which behaviors are most important to test?"
57+
58+
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
59+
60+
### 2. Tracer Bullet
61+
62+
Write ONE test that confirms ONE thing about the system:
63+
64+
```
65+
RED: Write test for first behavior → test fails
66+
GREEN: Write minimal code to pass → test passes
67+
```
68+
69+
This is your tracer bullet - proves the path works end-to-end.
70+
71+
### 3. Incremental Loop
72+
73+
For each remaining behavior:
74+
75+
```
76+
RED: Write next test → fails
77+
GREEN: Minimal code to pass → passes
78+
```
79+
80+
Rules:
81+
82+
- One test at a time
83+
- Only enough code to pass current test
84+
- Don't anticipate future tests
85+
- Keep tests focused on observable behavior
86+
87+
### 4. Refactor
88+
89+
After all tests pass, look for [refactor candidates](refactoring.md):
90+
91+
- [ ] Extract duplication
92+
- [ ] Deepen modules (move complexity behind simple interfaces)
93+
- [ ] Apply SOLID principles where natural
94+
- [ ] Consider what new code reveals about existing code
95+
- [ ] Run tests after each refactor step
96+
97+
**Never refactor while RED.** Get to GREEN first.
98+
99+
## Checklist Per Cycle
100+
101+
```
102+
[ ] Test describes behavior, not implementation
103+
[ ] Test uses public interface only
104+
[ ] Test would survive internal refactor
105+
[ ] Code is minimal for this test
106+
[ ] No speculative features added
107+
```

.agents/skills/tdd/deep-modules.md

Lines changed: 33 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,33 @@
1+
# Deep Modules
2+
3+
From "A Philosophy of Software Design":
4+
5+
**Deep module** = small interface + lots of implementation
6+
7+
```
8+
┌─────────────────────┐
9+
│ Small Interface │ ← Few methods, simple params
10+
├─────────────────────┤
11+
│ │
12+
│ │
13+
│ Deep Implementation│ ← Complex logic hidden
14+
│ │
15+
│ │
16+
└─────────────────────┘
17+
```
18+
19+
**Shallow module** = large interface + little implementation (avoid)
20+
21+
```
22+
┌─────────────────────────────────┐
23+
│ Large Interface │ ← Many methods, complex params
24+
├─────────────────────────────────┤
25+
│ Thin Implementation │ ← Just passes through
26+
└─────────────────────────────────┘
27+
```
28+
29+
When designing interfaces, ask:
30+
31+
- Can I reduce the number of methods?
32+
- Can I simplify the parameters?
33+
- Can I hide more complexity inside?
Lines changed: 31 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,31 @@
1+
# Interface Design for Testability
2+
3+
Good interfaces make testing natural:
4+
5+
1. **Accept dependencies, don't create them**
6+
7+
```typescript
8+
// Testable
9+
function processOrder(order, paymentGateway) {}
10+
11+
// Hard to test
12+
function processOrder(order) {
13+
const gateway = new StripeGateway();
14+
}
15+
```
16+
17+
2. **Return results, don't produce side effects**
18+
19+
```typescript
20+
// Testable
21+
function calculateDiscount(cart): Discount {}
22+
23+
// Hard to test
24+
function applyDiscount(cart): void {
25+
cart.total -= discount;
26+
}
27+
```
28+
29+
3. **Small surface area**
30+
- Fewer methods = fewer tests needed
31+
- Fewer params = simpler test setup

.agents/skills/tdd/mocking.md

Lines changed: 59 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,59 @@
1+
# When to Mock
2+
3+
Mock at **system boundaries** only:
4+
5+
- External APIs (payment, email, etc.)
6+
- Databases (sometimes - prefer test DB)
7+
- Time/randomness
8+
- File system (sometimes)
9+
10+
Don't mock:
11+
12+
- Your own classes/modules
13+
- Internal collaborators
14+
- Anything you control
15+
16+
## Designing for Mockability
17+
18+
At system boundaries, design interfaces that are easy to mock:
19+
20+
**1. Use dependency injection**
21+
22+
Pass external dependencies in rather than creating them internally:
23+
24+
```typescript
25+
// Easy to mock
26+
function processPayment(order, paymentClient) {
27+
return paymentClient.charge(order.total);
28+
}
29+
30+
// Hard to mock
31+
function processPayment(order) {
32+
const client = new StripeClient(process.env.STRIPE_KEY);
33+
return client.charge(order.total);
34+
}
35+
```
36+
37+
**2. Prefer SDK-style interfaces over generic fetchers**
38+
39+
Create specific functions for each external operation instead of one generic function with conditional logic:
40+
41+
```typescript
42+
// GOOD: Each function is independently mockable
43+
const api = {
44+
getUser: (id) => fetch(`/users/${id}`),
45+
getOrders: (userId) => fetch(`/users/${userId}/orders`),
46+
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
47+
};
48+
49+
// BAD: Mocking requires conditional logic inside the mock
50+
const api = {
51+
fetch: (endpoint, options) => fetch(endpoint, options),
52+
};
53+
```
54+
55+
The SDK approach means:
56+
- Each mock returns one specific shape
57+
- No conditional logic in test setup
58+
- Easier to see which endpoints a test exercises
59+
- Type safety per endpoint

.agents/skills/tdd/refactoring.md

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,10 @@
1+
# Refactor Candidates
2+
3+
After TDD cycle, look for:
4+
5+
- **Duplication** → Extract function/class
6+
- **Long methods** → Break into private helpers (keep tests on public interface)
7+
- **Shallow modules** → Combine or deepen
8+
- **Feature envy** → Move logic to where data lives
9+
- **Primitive obsession** → Introduce value objects
10+
- **Existing code** the new code reveals as problematic

.agents/skills/tdd/tests.md

Lines changed: 61 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
1+
# Good and Bad Tests
2+
3+
## Good Tests
4+
5+
**Integration-style**: Test through real interfaces, not mocks of internal parts.
6+
7+
```typescript
8+
// GOOD: Tests observable behavior
9+
test("user can checkout with valid cart", async () => {
10+
const cart = createCart();
11+
cart.add(product);
12+
const result = await checkout(cart, paymentMethod);
13+
expect(result.status).toBe("confirmed");
14+
});
15+
```
16+
17+
Characteristics:
18+
19+
- Tests behavior users/callers care about
20+
- Uses public API only
21+
- Survives internal refactors
22+
- Describes WHAT, not HOW
23+
- One logical assertion per test
24+
25+
## Bad Tests
26+
27+
**Implementation-detail tests**: Coupled to internal structure.
28+
29+
```typescript
30+
// BAD: Tests implementation details
31+
test("checkout calls paymentService.process", async () => {
32+
const mockPayment = jest.mock(paymentService);
33+
await checkout(cart, payment);
34+
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
35+
});
36+
```
37+
38+
Red flags:
39+
40+
- Mocking internal collaborators
41+
- Testing private methods
42+
- Asserting on call counts/order
43+
- Test breaks when refactoring without behavior change
44+
- Test name describes HOW not WHAT
45+
- Verifying through external means instead of interface
46+
47+
```typescript
48+
// BAD: Bypasses interface to verify
49+
test("createUser saves to database", async () => {
50+
await createUser({ name: "Alice" });
51+
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
52+
expect(row).toBeDefined();
53+
});
54+
55+
// GOOD: Verifies through interface
56+
test("createUser makes user retrievable", async () => {
57+
const user = await createUser({ name: "Alice" });
58+
const retrieved = await getUser(user.id);
59+
expect(retrieved.name).toBe("Alice");
60+
});
61+
```

0 commit comments

Comments
 (0)