What a characterisation test is for
When a module has no tests and you must change it, the first question is not whether it is correct but what it does. Michael Feathers calls tests that document a system's actual behaviour characterisation tests. They are not a statement of what the code should do. They are a safety net: if a refactor, an upgrade or a rewrite changes a result, one of them fails straight away. That is the whole point. The tests describe today, so that tomorrow's difference is visible. They are different from a regression test for a known bug, which starts from a wrong result and a decision about the right one.
- Use them before a refactor, a dependency or runtime upgrade, or a handover.
- Do not use them to decide what the module should do. Keep that decision separate.
The method, step by step
Feathers' method starts with the hardest part: getting the code into a test harness, which usually means breaking the dependencies around it, such as the clock, the network and the database, so a test can run it. Then write a test that calls the code with an input and asserts a placeholder you know is wrong. Run it and read the failure, which tells you what the code returns. Paste that real value into the assertion and rename the test to say what you learned. Repeat with other inputs, including boundary values such as empty, zero and very large, and with inputs that should fail. The names of the tests will change as you understand the module better.
- Cover each public entry point with a normal input, a boundary input and an invalid one.
- Isolate time, randomness, network and file access so repeated runs give the same answer.
- Keep each test small enough that a failure points at one behaviour.
Surprising results are not automatically bugs
While characterising, you will find results that look wrong. Feathers' own example strips an empty tag to nothing and calls the question of whether that is a bug a trick question, because context decides. His advice is to leave the test in place unless you have concluded it is a bug, since users may rely on odd behaviour and once software is in production it becomes its own specification. What matters is knowing when you change existing behaviour. Keep a short behaviour note of the surprises, so a person with the context can decide which to correct. Correcting one later is a deliberate test change with a reason attached.
- Record surprising results as they are and mark them as unjudged.
- Turn a surprising result into a bug-fix job only once someone has decided what the correct result is.
What a coverage figure tells you, and how the paid job is accepted
Coverage tools help find code no test reaches. coverage.py distinguishes statement coverage, which records whether each line ran, from branch coverage, which also records the jumps between lines, so a line can run while one of its branches never does. Martin Fowler's view is that coverage is useful for finding untested code and of little use as a number for how good the tests are, because high numbers are easy to reach with low-quality tests. So a suite can reach a high figure while asserting nothing useful. The regression suite job therefore reports coverage for information and accepts the work on a different test: ten deliberate behaviour changes on a throwaway copy, each visible through the public interface, must make at least one test fail, with any survivor listed and explained. The suite must also pass 20 consecutive runs, in random order where the runner supports it.
Sources and limits
- Michael Feathers: Characterization Testing Checked 2026-10-11.
- Characterization tests document a system's actual behaviour and act as a safety net for refactoring, and getting the code into a test harness by breaking dependencies is the hardest part.
- The method writes a test with a placeholder expectation, reads the failure to learn what the code returns, then records that value and renames the test.
- Surprising behaviour is not automatically a defect; leave the test in unless you conclude it is a bug, because users may depend on it.
- coverage.py: branch coverage Checked 2026-10-11.
- Statement coverage records whether each line ran; branch coverage also records which jumps between lines happened, so a line can run while a branch from it never does.
- Martin Fowler: Test Coverage Checked 2026-10-11.
- Coverage is useful for finding untested parts of a codebase but of little use as a numeric statement of how good the tests are, and high numbers are easy to reach with low-quality tests.