Prove it is order, not luck
A test that passes alone and fails in the full run usually depends on something an earlier test left behind. Before changing anything, make the failure repeatable. Run the failing test by itself and note the result. Run its whole file, then the whole suite, and note each result. If the suite can be run in a different order, run it that way too. Then name the one or two tests you suspect, run them together and swap their order. If the result changes when the order changes, you have shown an order dependence and you have a short list of suspects, which is more useful than a hunch.
- Write down the command, the order and the first failing line for each run, so that you are comparing like with like.
- A test that fails every time, even alone, is not this problem: it is reporting something repeatable.
Where the shared state hides
pytest's fixtures have five scopes and the default is function, which means each test gets its own fresh object and the fixture is destroyed when the test ends. A module-scoped fixture is different: the documentation's own example hands the same connection object to two test functions. If one test changes that object, the next test sees the change. The same applies to class and session scopes, and to anything outside fixtures that lives longer than a test: module-level variables, class attributes, caches, singletons, rows left in a database and files in a shared folder.
- List every fixture with a scope wider than function that the failing test uses, directly or through another fixture.
- Search for module-level state that tests assign to, such as lists, dictionaries and cached lookups.
- Check for environment variables or the working directory being changed by hand instead of through monkeypatch.
- Check whether tests write to a fixed path rather than a per-test temporary directory.
Fix the leak at its source
The durable fix is isolation, not ordering. pytest's monkeypatch undoes its changes when the requesting test or fixture finishes, so use it for environment variables, attributes and the working directory. tmp_path gives every test function a unique directory, so shared fixed paths are unnecessary. Narrow a fixture's scope where its object is mutated, or return a fresh copy to each test. Martin Fowler's advice on non-deterministic tests is to rebuild a known state at the start of each test rather than tidy up afterwards, and to make teardown errors fail loudly. Putting the two tests in a fixed order, or adding a pause, hides the dependency without removing it.
- Prefer creating state at the start of a test over cleaning up at the end, because a failed cleanup then cannot poison the next test.
- Give each fixture one state-changing action and its own teardown, as the pytest documentation recommends.
- Re-run the original order and a changed order after the change, not just the one that failed.
What does not fit, and how a paid fix is accepted
If the test fails every time on its own, treat it as a bug with a regression test. If it fails only on the build server, compare the environments, which is a different investigation. If it fails at random even when run alone, look at time and randomness instead. When you need a named set of such tests stabilised, the one-off job reproduces each failure by repeated runs, alone, in its file, in the full suite and in a different order, classifies the cause with evidence, then fixes it or quarantines it openly. It counts as fixed only when the test that failed in the baseline passes the agreed number of consecutive runs under the same conditions and no assertion was weakened. A clean count lowers the odds that a flake remains; it cannot prove that there is none. Send test names and redacted failure messages, never code or credentials, to start an enquiry.
Sources and limits
- pytest: how to use fixtures Checked 2026-10-11.
- Fixtures have five scopes, function, class, module, package and session, and function is the default, with the fixture destroyed at the end of the test.
- A module-scoped fixture gives the same object to every test in the module, so state changed by one test is visible to the next.
- Teardown code after a yield runs once the test finishes, and a fixture should make one state-changing action together with its own teardown.
- pytest: how to monkeypatch and mock modules and environments Checked 2026-10-11.
- Modifications made with monkeypatch are undone after the requesting test function or fixture has finished.
- pytest: how to use temporary directories and files Checked 2026-10-11.
- tmp_path gives each test function its own temporary directory, and tmp_path_factory is session-scoped for directories that are meant to be shared.
- Martin Fowler: Eradicating Non-Determinism in Tests Checked 2026-10-11.
- Lack of isolation is a named cause of non-deterministic tests; each test should start from a known state, and teardown errors should fail loudly.