Quality-First TDD skill

Guides quality-first TDD for new behavior, bug fixes, and test changes.

by HoangNguyen0403·MIT license·★ 569 Stars on the repo·GitHub ↗

Use now

Files of Quality-First TDD

HoangNguyen0403/develop1 file shown
SKILL.md
Show the full text75 lines

Quality-First TDD

Priority: P0 (CRITICAL)

A passing test is insufficient; the test must prove an owned behavior and a distinct plausible fault.

Choose the mode

  • New behavior: strict RED -> GREEN -> REFACTOR. Do not write production code before the expected RED.
  • Legacy or bug fix: characterize only when needed, then reproduce the intended change as a failing regression (RED). Preserve unrelated existing code; do not delete it merely because it predates the test.

Before writing a test

Compare against the existing suite before adding or removing cases. Ensure the scenario is not already covered upstream or in nearby layers.

Create one Test Intent Record per behavior/risk:

  • contract: observable contract — an application-owned result or outward side effect
  • fault: distinct plausible fault this test catches
  • layer: smallest honest unit, component, contract, integration, or E2E layer
  • cases: minimal distinct equivalence classes; use a parameterized test for equivalent inputs
  • command: exact focused single-run command

Reject tests that duplicate existing suite coverage, assert internal mock choreography over outward side effects, depend on time/network/order, or invent numeric four-pillar scores without measured evidence.

Bounded loop

  1. Run configured lint/type checks, inspect nearby tests, and derive the smallest command.
  2. RED: add one intent group and run it in the foreground, sequentially, single-run mode.
  3. Classify RED as expected_red, invalid_red, unexpected_green, or verification_infra_failed. If it is unexpected_green, inspect existing coverage and remove a redundant or weak case before implementing production code.
  4. GREEN: implement only enough to satisfy expected_red; rerun the same command.
  5. REFACTOR: improve structure without changing behavior; rerun the same command.
  6. Escalate only when evidence requires it: related unit target, integration/contract target, then explicit release/full-suite gate.

Execution safety

  • Honor project timeouts; otherwise use a 120-second fallback to bound a focused command.
  • On timeout, terminate only the agent-owned process group and verify child cleanup.
  • Never watch, blanket-kill, or retry an unchanged failure. Record the new hypothesis or corrective change first.
  • Coverage percentages are diagnostic tools, never an admission criterion. Without a configured threshold, report risk gaps and never add padding tests for an arbitrary percentage.

Red flags and rationalizations

  • Stop on: add tests after, too small, passed first run, run the full suite again, or mock every collaborator.
  • Urgency, manual testing, test count, or a coverage target never bypasses the intent record, expected RED, bounded command, or fault proof.

Test shape

  • Verify one logical contract per test; multiple assertions are allowed when verifying related aspects or side effects of that single contract.
  • Assert observable outcomes and outward side effects, not mock choreography or call sequences.
  • Mock external boundaries only when isolation requires it; prefer real pure/domain behavior and simple fakes.
  • Keep test names behavior-focused, without ticket IDs or TODO/FIXME markers.

See references/quality-contract.md for the intent record, failure taxonomy, layer routing, and runner examples.

1---
2name: common-tdd
3description: "Guides quality-first TDD for new behavior, bug fixes, and test changes. Selects the smallest test layer, proves a distinct regression risk, and runs bounded RED-GREEN-REFACTOR verification."
4metadata:
5 triggers:
6 files:
7 - "**/*.test.ts"
8 - "**/*.spec.ts"
9 - "**/*_test.go"
10 - "**/*Test.java"
11 - "**/*_test.dart"
12 - "**/*_spec.rb"
13 keywords:
14 - tdd
15 - unit test
16 - write test
17 - red green refactor
18 - failing test
19 - test coverage
20---
21 
22# Quality-First TDD
23 
24## **Priority: P0 (CRITICAL)**
25 
26A passing test is insufficient; the test must prove an owned behavior and a distinct plausible fault.
27 
28## Choose the mode
29 
30- **New behavior:** strict RED -> GREEN -> REFACTOR. Do not write production code before the expected RED.
31- **Legacy or bug fix:** characterize only when needed, then reproduce the intended change as a failing regression (RED). Preserve unrelated existing code; do not delete it merely because it predates the test.
32 
33## Before writing a test
34 
35Compare against the existing suite before adding or removing cases. Ensure the scenario is not already covered upstream or in nearby layers.
36 
37Create one Test Intent Record per behavior/risk:
38 
39- `contract`: observable contract — an application-owned result or outward side effect
40- `fault`: distinct plausible fault this test catches
41- `layer`: smallest honest unit, component, contract, integration, or E2E layer
42- `cases`: minimal distinct equivalence classes; use a parameterized test for equivalent inputs
43- `command`: exact focused single-run command
44 
45Reject tests that duplicate existing suite coverage, assert internal mock choreography over outward side effects, depend on time/network/order, or invent numeric four-pillar scores without measured evidence.
46## Bounded loop
47 
481. Run configured lint/type checks, inspect nearby tests, and derive the smallest command.
492. **RED:** add one intent group and run it in the foreground, sequentially, single-run mode.
503. Classify RED as `expected_red`, `invalid_red`, `unexpected_green`, or `verification_infra_failed`. If it is `unexpected_green`, inspect existing coverage and remove a redundant or weak case before implementing production code.
514. **GREEN:** implement only enough to satisfy `expected_red`; rerun the same command.
525. **REFACTOR:** improve structure without changing behavior; rerun the same command.
536. Escalate only when evidence requires it: related unit target, integration/contract target, then explicit release/full-suite gate.
54 
55## Execution safety
56 
57- Honor project timeouts; otherwise use a 120-second fallback to bound a focused command.
58- On timeout, terminate only the agent-owned process group and verify child cleanup.
59- Never watch, blanket-kill, or retry an unchanged failure. Record the new hypothesis or corrective change first.
60- Coverage percentages are diagnostic tools, never an admission criterion. Without a configured threshold, report risk gaps and never add padding tests for an arbitrary percentage.
61 
62## Red flags and rationalizations
63 
64- Stop on: `add tests after`, `too small`, `passed first run`, `run the full suite again`, or `mock every collaborator`.
65- Urgency, manual testing, test count, or a coverage target never bypasses the intent record, expected RED, bounded command, or fault proof.
66 
67## Test shape
68 
69- Verify one logical contract per test; multiple assertions are allowed when verifying related aspects or side effects of that single contract.
70- Assert observable outcomes and outward side effects, not mock choreography or call sequences.
71- Mock external boundaries only when isolation requires it; prefer real pure/domain behavior and simple fakes.
72- Keep test names behavior-focused, without ticket IDs or TODO/FIXME markers.
73 
74See `references/quality-contract.md` for the intent record, failure taxonomy, layer routing, and runner examples.
75 

Discussion

Alternatives