Quality-First TDD skill
Guides quality-first TDD for new behavior, bug fixes, and test changes.
by HoangNguyen0403·MIT license·★ 569 Stars on the repo·GitHub ↗
npx degit HoangNguyen0403/agent-skills-standard/skills/common/common-tdd#develop ~/.claude/skills/common-tddChecked ·commit develop
Files of Quality-First TDD
Show the full text75 lines
Quality-First TDD
Priority: P0 (CRITICAL)
A passing test is insufficient; the test must prove an owned behavior and a distinct plausible fault.
Choose the mode
- New behavior: strict RED -> GREEN -> REFACTOR. Do not write production code before the expected RED.
- Legacy or bug fix: characterize only when needed, then reproduce the intended change as a failing regression (RED). Preserve unrelated existing code; do not delete it merely because it predates the test.
Before writing a test
Compare against the existing suite before adding or removing cases. Ensure the scenario is not already covered upstream or in nearby layers.
Create one Test Intent Record per behavior/risk:
contract: observable contract — an application-owned result or outward side effectfault: distinct plausible fault this test catcheslayer: smallest honest unit, component, contract, integration, or E2E layercases: minimal distinct equivalence classes; use a parameterized test for equivalent inputscommand: exact focused single-run command
Reject tests that duplicate existing suite coverage, assert internal mock choreography over outward side effects, depend on time/network/order, or invent numeric four-pillar scores without measured evidence.
Bounded loop
- Run configured lint/type checks, inspect nearby tests, and derive the smallest command.
- RED: add one intent group and run it in the foreground, sequentially, single-run mode.
- Classify RED as
expected_red,invalid_red,unexpected_green, orverification_infra_failed. If it isunexpected_green, inspect existing coverage and remove a redundant or weak case before implementing production code. - GREEN: implement only enough to satisfy
expected_red; rerun the same command. - REFACTOR: improve structure without changing behavior; rerun the same command.
- Escalate only when evidence requires it: related unit target, integration/contract target, then explicit release/full-suite gate.
Execution safety
- Honor project timeouts; otherwise use a 120-second fallback to bound a focused command.
- On timeout, terminate only the agent-owned process group and verify child cleanup.
- Never watch, blanket-kill, or retry an unchanged failure. Record the new hypothesis or corrective change first.
- Coverage percentages are diagnostic tools, never an admission criterion. Without a configured threshold, report risk gaps and never add padding tests for an arbitrary percentage.
Red flags and rationalizations
- Stop on:
add tests after,too small,passed first run,run the full suite again, ormock every collaborator. - Urgency, manual testing, test count, or a coverage target never bypasses the intent record, expected RED, bounded command, or fault proof.
Test shape
- Verify one logical contract per test; multiple assertions are allowed when verifying related aspects or side effects of that single contract.
- Assert observable outcomes and outward side effects, not mock choreography or call sequences.
- Mock external boundaries only when isolation requires it; prefer real pure/domain behavior and simple fakes.
- Keep test names behavior-focused, without ticket IDs or TODO/FIXME markers.
See references/quality-contract.md for the intent record, failure taxonomy, layer routing, and runner examples.
| 1 | |
| 2 | name common-tdd |
| 3 | description "Guides quality-first TDD for new behavior, bug fixes, and test changes. Selects the smallest test layer, proves a distinct regression risk, and runs bounded RED-GREEN-REFACTOR verification." |
| 4 | metadata |
| 5 | triggers |
| 6 | files |
| 7 | - "**/*.test.ts" |
| 8 | - "**/*.spec.ts" |
| 9 | - "**/*_test.go" |
| 10 | - "**/*Test.java" |
| 11 | - "**/*_test.dart" |
| 12 | - "**/*_spec.rb" |
| 13 | keywords |
| 14 | - tdd |
| 15 | - unit test |
| 16 | - write test |
| 17 | - red green refactor |
| 18 | - failing test |
| 19 | - test coverage |
| 20 | |
| 21 | |
| 22 | # Quality-First TDD |
| 23 | |
| 24 | ## **Priority: P0 (CRITICAL)** |
| 25 | |
| 26 | A passing test is insufficient; the test must prove an owned behavior and a distinct plausible fault. |
| 27 | |
| 28 | ## Choose the mode |
| 29 | |
| 30 | **New behavior:** strict RED -> GREEN -> REFACTOR. Do not write production code before the expected RED. |
| 31 | **Legacy or bug fix:** characterize only when needed, then reproduce the intended change as a failing regression (RED). Preserve unrelated existing code; do not delete it merely because it predates the test. |
| 32 | |
| 33 | ## Before writing a test |
| 34 | |
| 35 | Compare against the existing suite before adding or removing cases. Ensure the scenario is not already covered upstream or in nearby layers. |
| 36 | |
| 37 | Create one Test Intent Record per behavior/risk: |
| 38 | |
| 39 | `contract`: observable contract — an application-owned result or outward side effect |
| 40 | `fault`: distinct plausible fault this test catches |
| 41 | `layer`: smallest honest unit, component, contract, integration, or E2E layer |
| 42 | `cases`: minimal distinct equivalence classes; use a parameterized test for equivalent inputs |
| 43 | `command`: exact focused single-run command |
| 44 | |
| 45 | Reject tests that duplicate existing suite coverage, assert internal mock choreography over outward side effects, depend on time/network/order, or invent numeric four-pillar scores without measured evidence. |
| 46 | ## Bounded loop |
| 47 | |
| 48 | Run configured lint/type checks, inspect nearby tests, and derive the smallest command. |
| 49 | **RED:** add one intent group and run it in the foreground, sequentially, single-run mode. |
| 50 | Classify RED as `expected_red`, `invalid_red`, `unexpected_green`, or `verification_infra_failed`. If it is `unexpected_green`, inspect existing coverage and remove a redundant or weak case before implementing production code. |
| 51 | **GREEN:** implement only enough to satisfy `expected_red`; rerun the same command. |
| 52 | **REFACTOR:** improve structure without changing behavior; rerun the same command. |
| 53 | Escalate only when evidence requires it: related unit target, integration/contract target, then explicit release/full-suite gate. |
| 54 | |
| 55 | ## Execution safety |
| 56 | |
| 57 | Honor project timeouts; otherwise use a 120-second fallback to bound a focused command. |
| 58 | On timeout, terminate only the agent-owned process group and verify child cleanup. |
| 59 | Never watch, blanket-kill, or retry an unchanged failure. Record the new hypothesis or corrective change first. |
| 60 | Coverage percentages are diagnostic tools, never an admission criterion. Without a configured threshold, report risk gaps and never add padding tests for an arbitrary percentage. |
| 61 | |
| 62 | ## Red flags and rationalizations |
| 63 | |
| 64 | Stop on: `add tests after`, `too small`, `passed first run`, `run the full suite again`, or `mock every collaborator`. |
| 65 | Urgency, manual testing, test count, or a coverage target never bypasses the intent record, expected RED, bounded command, or fault proof. |
| 66 | |
| 67 | ## Test shape |
| 68 | |
| 69 | Verify one logical contract per test; multiple assertions are allowed when verifying related aspects or side effects of that single contract. |
| 70 | Assert observable outcomes and outward side effects, not mock choreography or call sequences. |
| 71 | Mock external boundaries only when isolation requires it; prefer real pure/domain behavior and simple fakes. |
| 72 | Keep test names behavior-focused, without ticket IDs or TODO/FIXME markers. |
| 73 | |
| 74 | See `references/quality-contract.md` for the intent record, failure taxonomy, layer routing, and runner examples. |
| 75 |
Discussion
Alternatives
Browse more free Claude skills or everything in Development.