Testing principles skill

Comprehensive guide to writing clean, maintainable tests that serve as executable documentation.

by wondelai·MIT license·★ 2,235 Stars on the repo·GitHub ↗

Use now

Files of Testing principles

wondelai/main1 file
testing-principles.md
Show the full text354 lines

Testing Principles

Comprehensive guide to writing clean, maintainable tests that serve as executable documentation. Based on Robert C. Martin's Clean Code, Chapter 9.

Table of Contents

  1. Why Tests Matter
  2. The Three Laws of TDD
  3. Clean Tests
  4. One Concept Per Test
  5. F.I.R.S.T. Principles
  6. Test Naming
  7. Test Patterns and Practices
  8. Tests as Documentation

Why Tests Matter

Test code is just as important as production code. It is not a second-class citizen. It requires thought, design, and care. Dirty tests are equivalent to, if not worse than, having no tests. Tests that are hard to read, fragile, or slow become a liability that developers avoid and eventually delete.

The fundamental equation: Clean tests = confidence to refactor = clean production code. Without tests, every change is a potential bug. With dirty tests, every change requires fighting through incomprehensible test code. With clean tests, refactoring is fearless.


The Three Laws of TDD

Test-Driven Development follows three simple rules:

Law Rule What it means
First You may not write production code until you have written a failing unit test Tests drive the design, not the other way around
Second You may not write more of a unit test than is sufficient to fail (compilation failures count) Write the minimum test that fails
Third You may not write more production code than is sufficient to pass the currently failing test Write the minimum code that passes
The Red-Green-Refactor Cycle
  1. Red: Write a failing test (it should fail for the right reason)
  2. Green: Write the simplest code that makes the test pass (even if ugly)
  3. Refactor: Clean up both production code and test code while keeping all tests green

This cycle runs in seconds to minutes, not hours. Each cycle produces one small, tested increment.

Benefits of TDD
Benefit Why
Nearly 100% coverage Every line of production code was written to pass a test
Tests as documentation Tests show exactly how the code is intended to be used
Fearless refactoring You know immediately if a change breaks something
Better design Hard-to-test code is hard to use; TDD pushes toward clean design
Debugging reduction When a test fails, the bug is in the last few lines you wrote

Clean Tests

What Makes a Test Clean?

Readability. The same thing that makes production code clean makes test code clean: readability. What makes tests readable? The same thing that makes all code readable: clarity, simplicity, and density of expression. In a test, you want to say a lot with as few expressions as possible.

The Build-Operate-Check Pattern

Every clean test has three distinct phases:

def test_should_apply_bulk_discount_when_quantity_exceeds_threshold():
    # BUILD: Create the test data
    order = an_order()
        .with_item(product="Widget", quantity=25, unit_price=10.00)
        .build()

    # OPERATE: Execute the behavior under test
    invoice = billing_service.generate_invoice(order)

    # CHECK: Verify the expected outcome
    assert invoice.total == 225.00  # 250 - 10% bulk discount
    assert invoice.discount_applied == "BULK_10"

Also known as Arrange-Act-Assert or Given-When-Then.

Domain-Specific Testing Language

Build utility functions and helpers that read like a domain-specific language for your tests.

# BAD: Raw setup code obscures the test's intent
def test_expired_subscription():
    user = User(
        id=uuid4(),
        name="Alice",
        email="[email protected]",
        created_at=datetime(2024, 1, 1),
        subscription=Subscription(
            plan="pro",
            status="active",
            expires_at=datetime(2024, 1, 15),
        ),
    )
    user.subscription.check_expiry(current_date=datetime(2024, 2, 1))
    assert user.subscription.status == "expired"

# GOOD: Helpers create a readable narrative
def test_expired_subscription():
    user = a_user().with_pro_subscription(expires_on="2024-01-15").build()

    user.check_subscription_on("2024-02-01")

    assert_that(user).has_expired_subscription()

The test reads like a specification: "Given a user with a pro subscription expiring on Jan 15, when we check the subscription on Feb 1, then the subscription should be expired."


One Concept Per Test

Each test function should test one concept. This does not necessarily mean one assert per test -- it means one logical assertion, one behavioral expectation.

One Concept, Multiple Asserts (Acceptable)
def test_should_create_valid_invoice_from_order():
    order = an_order().with_two_items().build()

    invoice = billing_service.generate_invoice(order)

    assert invoice.customer == order.customer
    assert invoice.line_items_count == 2
    assert invoice.total == order.calculated_total
    assert invoice.status == "pending"

All four asserts verify one concept: "generating an invoice from an order produces a valid invoice."

Multiple Concepts (Split Into Separate Tests)
# BAD: Two concepts in one test
def test_invoice_generation():
    order = an_order().build()
    invoice = billing_service.generate_invoice(order)
    assert invoice.total == order.calculated_total  # Concept 1: correct total

    invoice.mark_as_paid()
    assert invoice.status == "paid"  # Concept 2: payment status transition

# GOOD: Each test covers one concept
def test_should_calculate_correct_invoice_total():
    order = an_order().with_total(150.00).build()
    invoice = billing_service.generate_invoice(order)
    assert invoice.total == 150.00

def test_should_transition_to_paid_when_marked_as_paid():
    invoice = an_invoice().with_status("pending").build()
    invoice.mark_as_paid()
    assert invoice.status == "paid"

F.I.R.S.T. Principles

Clean tests follow five principles that form the acronym F.I.R.S.T.:

Fast

Tests should be fast. When tests run slowly, you won't run them frequently. When you don't run them frequently, you won't find problems early. When you don't find problems early, you won't fix them easily.

Guideline Target How
Unit test suite Under 10 seconds Mock all external dependencies
Individual test Under 100ms No I/O, no network, no database
Integration tests Separate suite Run separately, not on every save
Independent

Tests should not depend on each other. One test should not set up conditions for the next. You should be able to run each test independently and in any order.

# BAD: Test B depends on Test A's side effects
def test_a_create_user():
    global test_user
    test_user = UserService.create("Alice")

def test_b_update_user():
    UserService.update(test_user.id, name="Bob")  # Fails if A doesn't run first

# GOOD: Each test is self-contained
def test_create_user():
    user = UserService.create("Alice")
    assert user.name == "Alice"

def test_update_user():
    user = UserService.create("Alice")  # Own setup
    updated = UserService.update(user.id, name="Bob")
    assert updated.name == "Bob"
Repeatable

Tests should produce the same result every time, in any environment -- development machine, CI server, production-like staging. Tests that depend on network availability, current time, or random data are flaky.

Flaky dependency Fix
Current time Inject a clock; mock datetime.now()
Random data Use seeded random or fixed test data
Network calls Mock HTTP clients
Database state Use transactions that roll back, or in-memory DB
File system Use temp directories; clean up in teardown
Environment variables Set explicitly in test setup
Self-Validating

Tests should have a boolean output: pass or fail. No manual interpretation required.

# BAD: Requires human to check output
def test_report_generation():
    report = generate_report()
    print(report)  # Developer must read and visually verify

# GOOD: Automated assertion
def test_report_generation():
    report = generate_report()
    assert report.title == "Q4 Revenue Report"
    assert report.total_revenue == 142_500.00
    assert len(report.line_items) == 12
Timely

Tests should be written just before the production code that makes them pass (TDD). Tests written after the fact are harder to write because the production code may not be designed for testability. You may decide that some production code is "too hard to test" -- which really means it's too coupled.


Test Naming

Test names should describe the scenario being tested and the expected behavior.

Naming Patterns
Pattern Example When to use
should_[expected]_when_[condition] should_reject_login_when_password_expired Most common; clear cause-effect
[method]_[scenario]_[expected] withdraw_insufficient_funds_throws_exception When testing a specific method
given_[state]_when_[action]_then_[result] given_empty_cart_when_checkout_then_error BDD-style
test_[behavior_description] test_expired_tokens_are_rejected Simple, readable
Bad Test Names
Bad name Problem Better name
test1 Meaningless test_empty_input_returns_empty_list
testProcess What about process? test_process_skips_inactive_users
testCalculate Too vague test_calculate_applies_weekend_surcharge
testBug1234 Won't make sense in 6 months test_duplicate_orders_are_rejected

Test Patterns and Practices

Parameterized Tests

When testing the same behavior with different inputs, use parameterized tests instead of copy-pasting.

@pytest.mark.parametrize("input_email,expected_valid", [
    ("[email protected]", True),
    ("[email protected]", True),
    ("user@example", False),
    ("@example.com", False),
    ("[email protected]", False),
    ("", False),
])
def test_email_validation(input_email, expected_valid):
    assert validate_email(input_email) == expected_valid
Test Fixtures and Builders

Use the Builder pattern for test data to make tests readable and maintainable.

class UserBuilder:
    def __init__(self):
        self._name = "Default User"
        self._email = "[email protected]"
        self._role = "viewer"
        self._active = True

    def with_name(self, name): self._name = name; return self
    def with_role(self, role): self._role = role; return self
    def inactive(self): self._active = False; return self
    def build(self):
        return User(
            name=self._name, email=self._email,
            role=self._role, active=self._active,
        )

def a_user():
    return UserBuilder()

# Usage in tests
admin = a_user().with_name("Alice").with_role("admin").build()
inactive_user = a_user().inactive().build()
Testing Error Paths

Every error path in production code should have a corresponding test.

def test_should_raise_on_negative_amount():
    account = an_account().with_balance(100).build()
    with pytest.raises(ValueError, match="Amount must be positive"):
        account.withdraw(-50)

def test_should_raise_on_insufficient_funds():
    account = an_account().with_balance(100).build()
    with pytest.raises(InsufficientFundsError):
        account.withdraw(150)
Boundary Condition Tests

Test the edges, not just the middle.

Boundary Tests needed
Empty input [], "", None, {}
Single element List with one item, string with one char
Maximum values MAX_INT, full capacity, max length
Off-by-one n-1, n, n+1 for any threshold
Transition points Just below and just above limits
Overflow/underflow Values that exceed type boundaries

Tests as Documentation

Clean tests serve as the most reliable documentation of how the system behaves. Unlike comments or wiki pages, tests are always up to date -- if they weren't, they'd be failing.

Documentation type Tests provide
API usage Test setup shows how to call the API
Expected behavior Assertions describe what should happen
Edge cases Boundary tests document special cases
Error behavior Error path tests document failure modes
Business rules Test names describe domain rules

When a new developer asks "how does this work?", point them to the tests. Clean tests answer the question better than any comment or README.

1# Testing Principles
2 
3Comprehensive guide to writing clean, maintainable tests that serve as executable documentation. Based on Robert C. Martin's *Clean Code*, Chapter 9.
4 
5 
6## Table of Contents
71. [Why Tests Matter](#why-tests-matter)
82. [The Three Laws of TDD](#the-three-laws-of-tdd)
93. [Clean Tests](#clean-tests)
104. [One Concept Per Test](#one-concept-per-test)
115. [F.I.R.S.T. Principles](#first-principles)
126. [Test Naming](#test-naming)
137. [Test Patterns and Practices](#test-patterns-and-practices)
148. [Tests as Documentation](#tests-as-documentation)
15 
16---
17 
18## Why Tests Matter
19 
20Test code is just as important as production code. It is not a second-class citizen. It requires thought, design, and care. Dirty tests are equivalent to, if not worse than, having no tests. Tests that are hard to read, fragile, or slow become a liability that developers avoid and eventually delete.
21 
22**The fundamental equation:** Clean tests = confidence to refactor = clean production code. Without tests, every change is a potential bug. With dirty tests, every change requires fighting through incomprehensible test code. With clean tests, refactoring is fearless.
23 
24---
25 
26## The Three Laws of TDD
27 
28Test-Driven Development follows three simple rules:
29 
30| Law | Rule | What it means |
31|-----|------|---------------|
32| **First** | You may not write production code until you have written a failing unit test | Tests drive the design, not the other way around |
33| **Second** | You may not write more of a unit test than is sufficient to fail (compilation failures count) | Write the minimum test that fails |
34| **Third** | You may not write more production code than is sufficient to pass the currently failing test | Write the minimum code that passes |
35 
36### The Red-Green-Refactor Cycle
37 
381. **Red:** Write a failing test (it should fail for the right reason)
392. **Green:** Write the simplest code that makes the test pass (even if ugly)
403. **Refactor:** Clean up both production code and test code while keeping all tests green
41 
42This cycle runs in seconds to minutes, not hours. Each cycle produces one small, tested increment.
43 
44### Benefits of TDD
45 
46| Benefit | Why |
47|---------|-----|
48| **Nearly 100% coverage** | Every line of production code was written to pass a test |
49| **Tests as documentation** | Tests show exactly how the code is intended to be used |
50| **Fearless refactoring** | You know immediately if a change breaks something |
51| **Better design** | Hard-to-test code is hard to use; TDD pushes toward clean design |
52| **Debugging reduction** | When a test fails, the bug is in the last few lines you wrote |
53 
54---
55 
56## Clean Tests
57 
58### What Makes a Test Clean?
59 
60**Readability.** The same thing that makes production code clean makes test code clean: readability. What makes tests readable? The same thing that makes all code readable: clarity, simplicity, and density of expression. In a test, you want to say a lot with as few expressions as possible.
61 
62### The Build-Operate-Check Pattern
63 
64Every clean test has three distinct phases:
65 
66```python
67def test_should_apply_bulk_discount_when_quantity_exceeds_threshold():
68 # BUILD: Create the test data
69 order = an_order()
70 .with_item(product="Widget", quantity=25, unit_price=10.00)
71 .build()
72 
73 # OPERATE: Execute the behavior under test
74 invoice = billing_service.generate_invoice(order)
75 
76 # CHECK: Verify the expected outcome
77 assert invoice.total == 225.00 # 250 - 10% bulk discount
78 assert invoice.discount_applied == "BULK_10"
79```
80 
81Also known as **Arrange-Act-Assert** or **Given-When-Then**.
82 
83### Domain-Specific Testing Language
84 
85Build utility functions and helpers that read like a domain-specific language for your tests.
86 
87```python
88# BAD: Raw setup code obscures the test's intent
89def test_expired_subscription():
90 user = User(
91 id=uuid4(),
92 name="Alice",
93 email="[email protected]",
94 created_at=datetime(2024, 1, 1),
95 subscription=Subscription(
96 plan="pro",
97 status="active",
98 expires_at=datetime(2024, 1, 15),
99 ),
100 )
101 user.subscription.check_expiry(current_date=datetime(2024, 2, 1))
102 assert user.subscription.status == "expired"
103 
104# GOOD: Helpers create a readable narrative
105def test_expired_subscription():
106 user = a_user().with_pro_subscription(expires_on="2024-01-15").build()
107 
108 user.check_subscription_on("2024-02-01")
109 
110 assert_that(user).has_expired_subscription()
111```
112 
113The test reads like a specification: "Given a user with a pro subscription expiring on Jan 15, when we check the subscription on Feb 1, then the subscription should be expired."
114 
115---
116 
117## One Concept Per Test
118 
119Each test function should test one concept. This does not necessarily mean one assert per test -- it means one logical assertion, one behavioral expectation.
120 
121### One Concept, Multiple Asserts (Acceptable)
122 
123```python
124def test_should_create_valid_invoice_from_order():
125 order = an_order().with_two_items().build()
126 
127 invoice = billing_service.generate_invoice(order)
128 
129 assert invoice.customer == order.customer
130 assert invoice.line_items_count == 2
131 assert invoice.total == order.calculated_total
132 assert invoice.status == "pending"
133```
134 
135All four asserts verify one concept: "generating an invoice from an order produces a valid invoice."
136 
137### Multiple Concepts (Split Into Separate Tests)
138 
139```python
140# BAD: Two concepts in one test
141def test_invoice_generation():
142 order = an_order().build()
143 invoice = billing_service.generate_invoice(order)
144 assert invoice.total == order.calculated_total # Concept 1: correct total
145 
146 invoice.mark_as_paid()
147 assert invoice.status == "paid" # Concept 2: payment status transition
148 
149# GOOD: Each test covers one concept
150def test_should_calculate_correct_invoice_total():
151 order = an_order().with_total(150.00).build()
152 invoice = billing_service.generate_invoice(order)
153 assert invoice.total == 150.00
154 
155def test_should_transition_to_paid_when_marked_as_paid():
156 invoice = an_invoice().with_status("pending").build()
157 invoice.mark_as_paid()
158 assert invoice.status == "paid"
159```
160 
161---
162 
163## F.I.R.S.T. Principles
164 
165Clean tests follow five principles that form the acronym F.I.R.S.T.:
166 
167### Fast
168 
169Tests should be fast. When tests run slowly, you won't run them frequently. When you don't run them frequently, you won't find problems early. When you don't find problems early, you won't fix them easily.
170 
171| Guideline | Target | How |
172|-----------|--------|-----|
173| Unit test suite | Under 10 seconds | Mock all external dependencies |
174| Individual test | Under 100ms | No I/O, no network, no database |
175| Integration tests | Separate suite | Run separately, not on every save |
176 
177### Independent
178 
179Tests should not depend on each other. One test should not set up conditions for the next. You should be able to run each test independently and in any order.
180 
181```python
182# BAD: Test B depends on Test A's side effects
183def test_a_create_user():
184 global test_user
185 test_user = UserService.create("Alice")
186 
187def test_b_update_user():
188 UserService.update(test_user.id, name="Bob") # Fails if A doesn't run first
189 
190# GOOD: Each test is self-contained
191def test_create_user():
192 user = UserService.create("Alice")
193 assert user.name == "Alice"
194 
195def test_update_user():
196 user = UserService.create("Alice") # Own setup
197 updated = UserService.update(user.id, name="Bob")
198 assert updated.name == "Bob"
199```
200 
201### Repeatable
202 
203Tests should produce the same result every time, in any environment -- development machine, CI server, production-like staging. Tests that depend on network availability, current time, or random data are flaky.
204 
205| Flaky dependency | Fix |
206|-----------------|-----|
207| Current time | Inject a clock; mock `datetime.now()` |
208| Random data | Use seeded random or fixed test data |
209| Network calls | Mock HTTP clients |
210| Database state | Use transactions that roll back, or in-memory DB |
211| File system | Use temp directories; clean up in teardown |
212| Environment variables | Set explicitly in test setup |
213 
214### Self-Validating
215 
216Tests should have a boolean output: pass or fail. No manual interpretation required.
217 
218```python
219# BAD: Requires human to check output
220def test_report_generation():
221 report = generate_report()
222 print(report) # Developer must read and visually verify
223 
224# GOOD: Automated assertion
225def test_report_generation():
226 report = generate_report()
227 assert report.title == "Q4 Revenue Report"
228 assert report.total_revenue == 142_500.00
229 assert len(report.line_items) == 12
230```
231 
232### Timely
233 
234Tests should be written just before the production code that makes them pass (TDD). Tests written after the fact are harder to write because the production code may not be designed for testability. You may decide that some production code is "too hard to test" -- which really means it's too coupled.
235 
236---
237 
238## Test Naming
239 
240Test names should describe the scenario being tested and the expected behavior.
241 
242### Naming Patterns
243 
244| Pattern | Example | When to use |
245|---------|---------|-------------|
246| `should_[expected]_when_[condition]` | `should_reject_login_when_password_expired` | Most common; clear cause-effect |
247| `[method]_[scenario]_[expected]` | `withdraw_insufficient_funds_throws_exception` | When testing a specific method |
248| `given_[state]_when_[action]_then_[result]` | `given_empty_cart_when_checkout_then_error` | BDD-style |
249| `test_[behavior_description]` | `test_expired_tokens_are_rejected` | Simple, readable |
250 
251### Bad Test Names
252 
253| Bad name | Problem | Better name |
254|----------|---------|-------------|
255| `test1` | Meaningless | `test_empty_input_returns_empty_list` |
256| `testProcess` | What about process? | `test_process_skips_inactive_users` |
257| `testCalculate` | Too vague | `test_calculate_applies_weekend_surcharge` |
258| `testBug1234` | Won't make sense in 6 months | `test_duplicate_orders_are_rejected` |
259 
260---
261 
262## Test Patterns and Practices
263 
264### Parameterized Tests
265 
266When testing the same behavior with different inputs, use parameterized tests instead of copy-pasting.
267 
268```python
269@pytest.mark.parametrize("input_email,expected_valid", [
270 ("[email protected]", True),
271 ("[email protected]", True),
272 ("user@example", False),
273 ("@example.com", False),
274 ("[email protected]", False),
275 ("", False),
276])
277def test_email_validation(input_email, expected_valid):
278 assert validate_email(input_email) == expected_valid
279```
280 
281### Test Fixtures and Builders
282 
283Use the Builder pattern for test data to make tests readable and maintainable.
284 
285```python
286class UserBuilder:
287 def __init__(self):
288 self._name = "Default User"
289 self._email = "[email protected]"
290 self._role = "viewer"
291 self._active = True
292 
293 def with_name(self, name): self._name = name; return self
294 def with_role(self, role): self._role = role; return self
295 def inactive(self): self._active = False; return self
296 def build(self):
297 return User(
298 name=self._name, email=self._email,
299 role=self._role, active=self._active,
300 )
301 
302def a_user():
303 return UserBuilder()
304 
305# Usage in tests
306admin = a_user().with_name("Alice").with_role("admin").build()
307inactive_user = a_user().inactive().build()
308```
309 
310### Testing Error Paths
311 
312Every error path in production code should have a corresponding test.
313 
314```python
315def test_should_raise_on_negative_amount():
316 account = an_account().with_balance(100).build()
317 with pytest.raises(ValueError, match="Amount must be positive"):
318 account.withdraw(-50)
319 
320def test_should_raise_on_insufficient_funds():
321 account = an_account().with_balance(100).build()
322 with pytest.raises(InsufficientFundsError):
323 account.withdraw(150)
324```
325 
326### Boundary Condition Tests
327 
328Test the edges, not just the middle.
329 
330| Boundary | Tests needed |
331|----------|-------------|
332| Empty input | `[]`, `""`, `None`, `{}` |
333| Single element | List with one item, string with one char |
334| Maximum values | `MAX_INT`, full capacity, max length |
335| Off-by-one | `n-1`, `n`, `n+1` for any threshold |
336| Transition points | Just below and just above limits |
337| Overflow/underflow | Values that exceed type boundaries |
338 
339---
340 
341## Tests as Documentation
342 
343Clean tests serve as the most reliable documentation of how the system behaves. Unlike comments or wiki pages, tests are always up to date -- if they weren't, they'd be failing.
344 
345| Documentation type | Tests provide |
346|-------------------|---------------|
347| **API usage** | Test setup shows how to call the API |
348| **Expected behavior** | Assertions describe what should happen |
349| **Edge cases** | Boundary tests document special cases |
350| **Error behavior** | Error path tests document failure modes |
351| **Business rules** | Test names describe domain rules |
352 
353When a new developer asks "how does this work?", point them to the tests. Clean tests answer the question better than any comment or README.
354 

Discussion

Alternatives