Most pytest tutorials stop right after assert 1 + 1 == 2, which is roughly where the
interesting part begins. This one keeps going — through fixtures, parametrization and the
configuration that makes a suite pleasant to live with, and out the other side into hooks and
plugins.
There is a moment in most Python projects where the test suite stops helping. It takes four minutes
to run, three tests fail intermittently, and nobody remembers what conftest.py is for. The
tool is rarely the problem — pytest has an answer for each of those, and the answers are small.
So this is written as a path rather than a reference. Part 1 gets you productive with nothing but
functions and assert. Parts 2 and 3 cover fixtures and parametrization, which is where
pytest stops being "unittest with less typing" and starts being a different tool. Part 4 is the
boring, high-leverage material: layout, configuration, mocking, CI. Part 5 opens the hood.
Every runnable snippet here was executed against pytest 9.1.1 before publishing, and the terminal output is copied from real runs. Anything version-sensitive is flagged inline.
Contents
- Basics — why pytest, then the mechanics you use every day
- Fixtures — the feature that makes pytest a different tool
- Parametrizing — one function, many independently reported cases
- Practical — the unglamorous, high-leverage material
- Advanced — under the hood, and how to extend it
- Reference — what to avoid, and a one-screen summary
Part 1 · Basics
Why pytest
Python ships with unittest, a port of Java's JUnit. It works, but it makes you inherit from a
base class, remember a different assert method for every comparison, and wrap setup logic in
setUp methods that run for every test whether you need them or not.
pytest replaces all of that with plain functions and the plain assert keyword.
import unittest
class TestMath(unittest.TestCase):
def setUp(self):
self.values = [1, 2, 3]
def test_sum(self):
self.assertEqual(sum(self.values), 6)
def test_membership(self):
self.assertIn(2, self.values)
self.assertNotIn(9, self.values)
import pytest
@pytest.fixture
def values():
return [1, 2, 3]
def test_sum(values):
assert sum(values) == 6
def test_membership(values):
assert 2 in values
assert 9 not in values
The pytest version has no class, no inheritance, and no assertion vocabulary to memorise. It is also
strictly more capable: values is a fixture, and unlike setUp it only runs for tests
that actually ask for it.
pytest also runs unittest test cases as-is, so adopting it never requires a big-bang migration.
Your first test
python -m pip install pytest
Create two files side by side.
import re
def slugify(title: str) -> str:
"""Turn a human title into a URL-safe slug."""
cleaned = re.sub(r"[^\w\s-]", "", title.lower())
return re.sub(r"[\s_]+", "-", cleaned).strip("-")
from slugify import slugify
def test_lowercases_and_joins_words():
assert slugify("Hello World") == "hello-world"
def test_strips_punctuation():
assert slugify("What's new?!") == "whats-new"
def test_collapses_whitespace():
assert slugify("too many spaces") == "too-many-spaces"
... [100%]
3 passed in 0.01s
Three dots, three passing tests. There is no registration step, no test suite object, no
if __name__ == "__main__". pytest found the file, found the functions, and ran them.
pytest inserts the test file's directory (technically its rootdir for that file) into
sys.path, which is why from slugify import slugify works here. For real projects use
the layout in Project layout instead of relying on this.
How tests are found
Discovery is convention-driven. pytest walks the directories you give it (or the current one) and collects:
| Level | Default rule | Config key |
|---|---|---|
| Files | test_*.py or *_test.py | python_files |
| Classes | Test*, and it must not define __init__ | python_classes |
| Functions | test*, at module level or inside a collected class | python_functions |
Ask pytest what it would run without running anything. This is the single most useful command when a test "isn't running":
pytest --collect-only -q
A test class with an __init__ method is silently skipped — pytest cannot instantiate it, so
it warns and moves on. If a whole class mysteriously vanishes, that is almost always why.
The assert statement
pytest rewrites the bytecode of your test modules at import time so that a failing assert
reports the value of every subexpression. You get JUnit-grade diagnostics from the plain keyword.
def build_report():
return {"rows": 41, "status": "ok", "tags": ["a", "b"]}
def test_report_shape():
report = build_report()
assert report == {"rows": 42, "status": "ok", "tags": ["a", "c"]}
def test_report_shape():
report = build_report()
> assert report == {"rows": 42, "status": "ok", "tags": ["a", "c"]}
E AssertionError: assert {'rows': 41, ...': ['a', 'b']} == {'rows': 42, ...': ['a', 'c']}
E
E Omitting 1 identical items, use -vv to show
E Differing items:
E {'rows': 41} != {'rows': 42}
E {'tags': ['a', 'b']} != {'tags': ['a', 'c']}
E Use -v to get more diff
test_report.py:6: AssertionError
1 failed in 0.02s
pytest knows how to diff dicts, lists, sets, strings and dataclasses. Add -vv when the diff is
truncated — it disables the shortening.
assert x == y, "x should equal y" is noise; the rewriting already showed you that.
A good message supplies context the values cannot: assert resp.ok, f"upstream said {resp.text}".
Exceptions and warnings
Use pytest.raises as a context manager. The test fails if the block does not raise.
import pytest
def withdraw(balance, amount):
if amount <= 0:
raise ValueError(f"amount must be positive, got {amount}")
if amount > balance:
raise ValueError("insufficient funds")
return balance - amount
def test_rejects_negative_amount():
with pytest.raises(ValueError):
withdraw(100, -5)
def test_error_message_names_the_amount():
# match= is a regex applied with re.search against str(exception)
with pytest.raises(ValueError, match=r"must be positive, got -5"):
withdraw(100, -5)
def test_inspect_the_exception_object():
with pytest.raises(ValueError) as excinfo:
withdraw(100, 500)
assert "insufficient" in str(excinfo.value)
assert excinfo.type is ValueError
Code placed after the raising line inside a with pytest.raises(...) block never executes.
Anything you want to check about the exception goes after the block using
excinfo.value.
Warnings have a mirror-image API:
import warnings, pytest
def legacy_api():
warnings.warn("use new_api() instead", DeprecationWarning, stacklevel=2)
return 42
def test_warns():
with pytest.warns(DeprecationWarning, match="new_api"):
assert legacy_api() == 42
Comparing floats
0.1 + 0.2 == 0.3 is False in binary floating point. pytest.approx wraps a value
with a tolerance and works on scalars, sequences, dicts and numpy arrays.
from pytest import approx
def test_scalar():
assert 0.1 + 0.2 == approx(0.3)
def test_explicit_tolerance():
assert 4.7 == approx(4.7005, abs=1e-3) # absolute
assert 1_000_100 == approx(1_000_000, rel=1e-3) # relative
def test_containers():
assert [0.1 + 0.2, 1 / 3] == approx([0.3, 0.33333333])
assert {"loss": 0.1 + 0.2} == approx({"loss": 0.3})
The default is a relative tolerance of 1e-6 with a small absolute floor, which is the right default for most numeric code. Say what you mean when precision actually matters.
Running and selecting tests
| Command | What it does |
|---|---|
pytest | Everything under the current directory |
pytest tests/unit | One directory |
pytest tests/test_api.py | One file |
pytest tests/test_api.py::test_login | One test — this is a node id |
pytest "tests/test_api.py::test_login[admin]" | One parametrized case |
pytest -k "login and not slow" | Substring/boolean match on names |
pytest -m integration | Match a marker |
pytest -x | Stop at the first failure |
pytest --maxfail=3 | Stop after three failures |
pytest -q / -v / -vv | Quieter / one line per test / no truncated diffs |
pytest -rA | Summary lines for every outcome, including passes |
pytest -s | Don't capture stdout (see your print calls live) |
pytest --lf | Only the tests that failed last run |
pytest --ff | Everything, but last run's failures first |
pytest --sw | Stepwise: stop at a failure, resume there next time |
Node ids are copy-pasteable. When a test fails, the id in the output is exactly the argument you need to re-run just that case.
Part 2 · Fixtures
Fixture basics
A fixture is a function that produces something a test needs. Tests request fixtures by naming them as parameters, and pytest supplies them. That is the entire mental model — it is dependency injection with a decorator.
import pytest
@pytest.fixture
def inventory():
return {"apple": 3, "pear": 0}
def test_has_apples(inventory): # ← the parameter name is the fixture name
assert inventory["apple"] == 3
def test_pears_are_out_of_stock(inventory):
assert inventory["pear"] == 0
Each test gets a fresh inventory by default, so one test mutating the dict cannot
affect another. That isolation is the reason fixtures beat module-level constants.
Fixtures can request other fixtures, forming a graph that pytest resolves for you:
@pytest.fixture
def config():
return {"currency": "EUR", "vat": 0.21}
@pytest.fixture
def cart(config): # fixtures compose
return {"config": config, "items": []}
def test_cart_uses_config(cart):
assert cart["config"]["vat"] == 0.21
Setup and teardown
Replace return with yield and everything after the yield becomes teardown. It runs
even if the test fails, because pytest wraps it in a finalizer.
import sqlite3
import pytest
@pytest.fixture
def db():
conn = sqlite3.connect(":memory:")
conn.execute("CREATE TABLE users (id INTEGER PRIMARY KEY, name TEXT)")
yield conn # ← the test runs here
conn.close() # ← always runs afterwards
def test_insert_and_read_back(db):
db.execute("INSERT INTO users (name) VALUES (?)", ("ada",))
assert db.execute("SELECT name FROM users").fetchone() == ("ada",)
If the code before yield raises, pytest reports the test as an error rather
than a failure, and the test body never runs. That distinction in the summary line tells you
instantly whether your production code or your scaffolding broke.
Scopes
Expensive setup should not repeat per test. scope controls how long one instance is cached and
shared.
| Scope | Created once per | Typical use |
|---|---|---|
function default | Test | Anything mutable |
class | Test class | Shared object for a group of methods |
module | File | A parsed file, a compiled artifact |
package | Directory | Rare; a shared subsystem |
session | Whole run | Docker container, server process, big download |
@pytest.fixture(scope="session")
def http_server():
proc = start_server() # slow: do it once for the whole run
yield proc
proc.terminate()
@pytest.fixture
def client(http_server): # cheap per-test wrapper over the shared server
return Client(http_server.url)
A fixture may only depend on fixtures of equal or wider scope. A
session fixture requesting a function fixture is an error, because the narrow one
would be destroyed while the wide one still holds it.
And never let a widely scoped fixture hand out something mutable. One test appending to a
session-scoped list poisons every later test, producing failures that depend on run order.
conftest.py
Fixtures defined in conftest.py are visible to every test in that directory and below, with no
import needed. Nested conftest.py files stack, and the closest definition wins.
tests/
├── conftest.py # fixtures for everything below
├── unit/
│ └── test_parser.py
└── integration/
├── conftest.py # extra fixtures, only for integration tests
└── test_api.py
import pytest
@pytest.fixture(scope="session")
def sample_data():
return {"users": ["ada", "grace"], "version": 3}
@pytest.fixture
def frozen_clock(monkeypatch):
"""Pin time.time() so timestamp assertions are deterministic."""
import time
monkeypatch.setattr(time, "time", lambda: 1_700_000_000.0)
return 1_700_000_000.0
It is also where hooks, custom CLI options and plugin registration live. pytest imports it
automatically — it must never be imported by hand, and it needs no __init__.py.
Where does a fixture come from? Ask:
pytest --fixtures # every fixture visible here, with docstrings
pytest --fixtures-per-test tests/test_api.py::test_login
Built-in fixtures
These ship with pytest. Reach for them before writing your own.
tmp_path — a real, empty directory per test
def test_writes_a_csv(tmp_path):
target = tmp_path / "out.csv" # tmp_path is a pathlib.Path
target.write_text("id,name\n1,ada\n")
assert target.exists()
assert target.read_text().splitlines()[0] == "id,name"
pytest keeps the last few runs' directories on disk for post-mortem inspection and cleans up older
ones. Use tmp_path_factory for a directory shared across a session.
monkeypatch — scoped, auto-undone patching
def test_reads_token_from_environment(monkeypatch):
monkeypatch.setenv("API_TOKEN", "secret")
monkeypatch.delenv("HTTP_PROXY", raising=False)
assert load_config().token == "secret"
def test_patch_an_attribute(monkeypatch):
monkeypatch.setattr("myapp.net.fetch", lambda url: {"ok": True})
assert myapp.summarize("http://x") == "ok"
def test_run_from_another_directory(monkeypatch, tmp_path):
monkeypatch.chdir(tmp_path) # undone automatically
assert Path.cwd() == tmp_path
Every change is reverted when the test ends, including failures. That is the whole point: hand-rolled
os.environ[...] = ... leaks into subsequent tests.
capsys and caplog — captured output
import logging
def test_prints_a_banner(capsys):
print("READY")
captured = capsys.readouterr() # consumes the buffer
assert captured.out == "READY\n"
assert captured.err == ""
def test_logs_a_warning(caplog):
with caplog.at_level(logging.WARNING):
logging.getLogger("app").warning("disk almost full")
assert "disk almost full" in caplog.text
assert caplog.records[0].levelno == logging.WARNING
Use capfd instead of capsys when the output comes from a subprocess or a C extension.
request — introspection
@pytest.fixture
def workspace(request, tmp_path):
"""Name the directory after the test that asked for it."""
d = tmp_path / request.node.name
d.mkdir()
return d
Factory fixtures
When a test needs several of a thing, or needs to choose parameters, return a function instead of a value. Track what you create so teardown stays correct.
import pytest
@pytest.fixture
def make_user(db):
created = []
def _make(name, admin=False):
cur = db.execute(
"INSERT INTO users (name, admin) VALUES (?, ?)", (name, admin)
)
created.append(cur.lastrowid)
return {"id": cur.lastrowid, "name": name, "admin": admin}
yield _make
for user_id in created: # clean up exactly what this test made
db.execute("DELETE FROM users WHERE id = ?", (user_id,))
def test_admins_can_see_everyone(make_user):
admin = make_user("ada", admin=True)
make_user("grace")
make_user("alan")
assert len(visible_users(admin)) == 3
Part 3 · Parametrizing
parametrize
One test function, many cases. Each case is a separate test: it gets its own node id, its own pass/fail, and one failing case does not hide the others.
import pytest
@pytest.mark.parametrize(
("title", "expected"),
[
("Hello World", "hello-world"),
("What's new?!", "whats-new"),
(" padded ", "padded"),
("", ""),
],
)
def test_slugify(title, expected):
assert slugify(title) == expected
test_slug.py::test_slugify[Hello World-hello-world] PASSED
test_slug.py::test_slugify[What's new?!-whats-new] PASSED
test_slug.py::test_slugify[ padded -padded] PASSED
test_slug.py::test_slugify[-] PASSED
Stacking decorators produces the cartesian product:
@pytest.mark.parametrize("codec", ["gzip", "zstd", "none"])
@pytest.mark.parametrize("size", [0, 1_000_000])
def test_roundtrip(codec, size):
payload = b"x" * size
assert decompress(compress(payload, codec), codec) == payload
Stacking is convenient and dangerous. Three stacked decorators of five values each is 125 tests. Prefer an explicit list of meaningful tuples once the product stops being all-interesting.
Readable IDs
Auto-generated ids come from the values, which is unreadable for objects and awkward for long strings. Name your cases — the id is what you will read in CI output and paste on the command line.
@pytest.mark.parametrize(
("payload", "expected"),
[
({"v": 1}, True),
({"v": 1, "extra": None}, True),
({}, False),
],
ids=["minimal", "with-nulls", "empty"],
)
def test_validate(payload, expected):
assert validate(payload) is expected
A dict keeps the name and the data in one place, which stops them drifting apart:
CASES = {
"empty-input": ("", []),
"single-token": ("a", ["a"]),
"trailing-comma": ("a,b,", ["a", "b"]),
"quoted-comma": ('"a,b",c', ["a,b", "c"]),
}
@pytest.mark.parametrize(
("raw", "expected"), list(CASES.values()), ids=list(CASES.keys())
)
def test_parse_csv_line(raw, expected):
assert parse_csv_line(raw) == expected
Per-case marks
pytest.param attaches a mark or an id to a single case, so one known-broken input does not force
you to skip the whole function.
@pytest.mark.parametrize(
("raw", "expected"),
[
("2024-01-01", date(2024, 1, 1)),
pytest.param("01/01/2024", date(2024, 1, 1), id="us-format"),
pytest.param(
"2024-W01-1", date(2024, 1, 1),
marks=pytest.mark.xfail(reason="ISO week dates not supported yet"),
),
pytest.param(
"yesterday", date(2024, 1, 1),
marks=pytest.mark.skipif(sys.platform == "win32", reason="needs GNU date"),
),
],
)
def test_parse_date(raw, expected):
assert parse_date(raw) == expected
Parametrized fixtures
Put params on the fixture and every test using it runs once per value. This is how you sweep a
whole suite across backends without touching a single test body.
@pytest.fixture(params=["sqlite", "postgres"], ids=["sqlite", "pg"])
def store(request):
backend = request.param # ← the current value
s = open_store(backend)
yield s
s.close()
# both of these now run twice, once per backend
def test_put_then_get(store):
store.put("k", "v")
assert store.get("k") == "v"
def test_missing_key_returns_none(store):
assert store.get("nope") is None
Indirect parametrization
indirect=True routes the parameter through a fixture rather than into the test. Use it when
the value needs setup before the test can see it.
@pytest.fixture
def user(request):
role = request.param # comes from parametrize, not from the test
return create_user(role=role)
@pytest.mark.parametrize("user", ["admin", "viewer"], indirect=True)
def test_dashboard_access(user):
assert user.can_open_dashboard() is (user.role == "admin")
Mix direct and indirect by naming which arguments are indirect:
@pytest.mark.parametrize(
("user", "expected"),
[("admin", True), ("viewer", False)],
indirect=["user"], # only `user` goes through the fixture
)
def test_dashboard_access(user, expected):
assert user.can_open_dashboard() is expected
Part 4 · Practical
Project layout
The src layout is the one that avoids import surprises: your package is not importable from the
repository root, so tests are forced to import the installed package — the same thing your users get.
myproject/
├── pyproject.toml
├── src/
│ └── myproject/
│ ├── __init__.py
│ └── core.py
└── tests/
├── conftest.py
├── unit/
│ └── test_core.py
└── integration/
└── test_end_to_end.py
python -m pip install -e ".[dev]"
pytest
__init__.py in tests/
Without it, test modules are imported as top-level modules — which means two files named
test_utils.py in different directories will collide. Either give them distinct names, or add
__init__.py files and accept package semantics. The distinct-names route is simpler.
Configuration
One config block removes a lot of repeated typing and makes local runs match CI. pyproject.toml
is the modern home; pytest.ini and setup.cfg still work.
[tool.pytest.ini_options]
minversion = "8.0"
testpaths = ["tests"]
addopts = [
"-ra", # summary for all non-passing outcomes
"--strict-markers", # unknown @pytest.mark.* becomes an error
"--strict-config", # typos in this very section become an error
"--import-mode=importlib",
]
markers = [
"slow: takes more than a second",
"integration: needs external services",
]
filterwarnings = [
"error", # warnings fail the suite ...
"ignore::DeprecationWarning:third_party.*" # ... except from code you don't own
]
| Option | Why you want it |
|---|---|
--strict-markers | A typo like @pytest.mark.slwo silently does nothing otherwise |
filterwarnings = ["error"] | Catches deprecations while they are cheap to fix |
testpaths | Bare pytest stops walking .venv, build, docs |
--import-mode=importlib | Modern import semantics; avoids sys.path surgery |
Markers, skip and xfail
Markers are labels. You attach them to tests and select on them later.
import pytest
@pytest.mark.slow
def test_full_reindex():
...
@pytest.mark.integration
class TestPaymentGateway: # applies to every method in the class
def test_charge(self): ...
def test_refund(self): ...
pytestmark = pytest.mark.integration # module-level: applies to the whole file
pytest -m slow # only slow
pytest -m "not slow" # everything else
pytest -m "integration and not slow"
skip vs xfail
| Construct | Meaning | When |
|---|---|---|
@pytest.mark.skip(reason=…) | Never run it | Temporarily irrelevant |
@pytest.mark.skipif(cond, reason=…) | Run only if the condition is false | Platform / version / optional dependency |
pytest.skip(reason) | Skip from inside the test | Decision needs runtime information |
pytest.importorskip("numpy") | Skip if the import fails | Optional dependency |
@pytest.mark.xfail(reason=…) | Run it; a failure is expected | Known bug, with a test that proves it |
import sys, pytest
@pytest.mark.skipif(sys.version_info < (3, 12), reason="uses PEP 695 syntax")
def test_new_generics(): ...
def test_needs_a_gpu():
if not gpu_present():
pytest.skip("no GPU on this runner")
...
pandas = pytest.importorskip("pandas", minversion="2.0")
@pytest.mark.xfail(strict=True, reason="bug #482: rounds half-down")
def test_rounds_half_up():
assert round_half_up(0.5) == 1
xfail(strict=True) over skip for known bugs
A strict xfail that starts passing becomes a failure. That is the feature: the day someone fixes the bug, the suite tells you to delete the marker. A skipped test just rots.
Mocking
Two tools, two jobs. monkeypatch replaces an attribute for the duration of a test.
unittest.mock additionally records how the replacement was used.
def test_retries_on_timeout(monkeypatch):
calls = []
def flaky_get(url):
calls.append(url)
if len(calls) < 3:
raise TimeoutError
return {"status": "ok"}
monkeypatch.setattr("myapp.client.get", flaky_get)
assert fetch_with_retry("http://x")["status"] == "ok"
assert len(calls) == 3
from unittest.mock import patch, MagicMock
def test_sends_exactly_one_email():
with patch("myapp.mailer.send") as send:
send.return_value = MagicMock(status=202)
register_user("ada@example.com")
send.assert_called_once_with(
to="ada@example.com", template="welcome"
)
If myapp/service.py does from myapp.mailer import send, then patching
"myapp.mailer.send" has no effect — service already holds its own reference.
Patch "myapp.service.send", the name in the module using it.
Mocking a network boundary is good. Mocking five internal collaborators means the test now asserts your implementation's shape rather than its behaviour, and it will break on every refactor while catching no bugs. Consider a fake object or a real in-memory implementation instead.
Output and logging
pytest captures stdout, stderr and logging, then prints them only for failing tests. That is why
your print statements seem to vanish on success.
| Flag | Effect |
|---|---|
-s | Disable capture entirely — output streams live |
--capture=no | Same as -s |
--log-cli-level=DEBUG | Stream log records to the terminal as they happen |
--show-capture=no | Hide captured output even on failures |
Coverage
python -m pip install pytest-cov
pytest --cov=myproject --cov-report=term-missing
pytest --cov=myproject --cov-report=html # then open htmlcov/index.html
pytest --cov=myproject --cov-fail-under=85 # gate in CI
A test that calls a function and asserts nothing still counts as full coverage of it. Treat the
number as a way to find untested code, never as evidence that tested code is correct.
Branch coverage (--cov-branch) is a strictly better signal than line coverage.
Speed
pytest --durations=10 # the 10 slowest tests, with setup/teardown split out
python -m pip install pytest-xdist
pytest -n auto # one worker per CPU
pytest -n 4 --dist loadfile # keep each file on a single worker
Parallelism exposes hidden coupling: tests that only pass in a particular order will start failing. That is a bug being surfaced, not a bug being introduced.
Debugging a failure
pytest -x -q # 1. stop at the first failure
pytest --lf -x # 2. iterate on just that one
pytest --lf -x --pdb # 3. drop into a debugger at the failure point
pytest --lf -x -vv --tb=long # 4. full diffs and full traceback
pytest -q # 5. confirm the whole suite is green again
| Flag | Traceback style |
|---|---|
--tb=long | Every frame, fully expanded default |
--tb=short | One line per frame — best for CI logs |
--tb=line | One line per failure |
--tb=no | Names only |
breakpoint() works inside tests as long as capture is off (-s), and --pdb opens a
debugger at the moment of failure with the whole frame still alive — usually faster than
guessing where to put the breakpoint.
CI recipe
name: tests
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12", "3.13"]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
cache: pip
- run: python -m pip install -e ".[dev]"
- run: pytest -ra --tb=short --durations=10 --cov --cov-report=xml -n auto
fail-fast: false matters: without it, one failing Python version cancels the others and hides
whether the problem is version-specific.
Part 5 · Advanced
Hooks
pytest is a plugin system with a test runner attached. Every phase — collection, setup, running,
reporting — publishes hooks, and conftest.py is a plugin that is loaded automatically. Implement
a hook by defining a function with the right name.
def pytest_collection_modifyitems(config, items):
"""Everything under tests/integration/ gets @pytest.mark.integration."""
import pytest
for item in items:
if "/integration/" in str(item.path):
item.add_marker(pytest.mark.integration)
import pytest
@pytest.hookimpl(hookwrapper=True)
def pytest_runtest_makereport(item, call):
outcome = yield
report = outcome.get_result()
# stash "rep_setup" / "rep_call" / "rep_teardown" on the test item
setattr(item, f"rep_{report.when}", report)
@pytest.fixture
def browser(request):
driver = start_browser()
yield driver
if getattr(request.node, "rep_call", None) and request.node.rep_call.failed:
driver.save_screenshot(f"/tmp/{request.node.name}.png")
driver.quit()
That pattern — capture an artifact only when the test failed — is the standard answer for browser screenshots, server logs and database dumps.
| Hook | Fires |
|---|---|
pytest_addoption | Once at startup, to register CLI flags |
pytest_configure | After config is read; register markers here |
pytest_collection_modifyitems | After collection — reorder, filter, mark |
pytest_generate_tests | Per test function, to generate parameters |
pytest_runtest_setup | Before each test |
pytest_runtest_makereport | After each phase, with the result |
pytest_sessionfinish | Once, at the very end |
Custom CLI options
import pytest
def pytest_addoption(parser):
parser.addoption(
"--runslow", action="store_true", default=False,
help="run tests marked slow",
)
parser.addoption(
"--env", action="store", default="staging",
choices=("staging", "prod"), help="target environment",
)
def pytest_configure(config):
config.addinivalue_line("markers", "slow: takes more than a second")
def pytest_collection_modifyitems(config, items):
if config.getoption("--runslow"):
return
skip = pytest.mark.skip(reason="needs --runslow")
for item in items:
if "slow" in item.keywords:
item.add_marker(skip)
@pytest.fixture(scope="session")
def env(request):
return request.config.getoption("--env")
pytest # slow tests are skipped
pytest --runslow # everything
pytest --env prod
pytest_generate_tests
When the set of cases is only known at runtime — read from a directory, a manifest, a database — generate them in this hook. Each generated case is still a first-class test.
# conftest.py
from pathlib import Path
CASES = Path(__file__).parent / "cases"
def pytest_generate_tests(metafunc):
if "case_file" in metafunc.fixturenames:
files = sorted(CASES.glob("*.input"))
metafunc.parametrize(
"case_file", files, ids=[f.stem for f in files]
)
# test_render.py
def test_render_matches_golden(case_file):
expected = case_file.with_suffix(".expected").read_text()
assert render(case_file.read_text()) == expected
Dropping a new pair of files into cases/ adds a test. No code change, and the id is the file
name so failures point straight at the data.
Writing a plugin
A plugin is a module of hooks and fixtures. Once it lives in a package with the
pytest11 entry point, installing the package is enough — no imports, no conftest.py.
import time
import pytest
@pytest.fixture
def timer():
"""Yield an object whose .elapsed is the time spent inside the block."""
class _Timer:
start = time.perf_counter()
@property
def elapsed(self):
return time.perf_counter() - self.start
return _Timer()
def pytest_addoption(parser):
parser.addoption("--warn-slower-than", type=float, default=None)
@pytest.hookimpl(hookwrapper=True)
def pytest_runtest_call(item):
started = time.perf_counter()
yield
limit = item.config.getoption("--warn-slower-than")
took = time.perf_counter() - started
if limit is not None and took > limit:
item.warn(pytest.PytestWarning(f"{item.name} took {took:.2f}s"))
[project.entry-points.pytest11]
timer = "pytest_timer.plugin"
Test your plugin with pytest's own pytester fixture, which runs a nested pytest session in a
temporary directory:
pytest_plugins = ["pytester"]
def test_warns_about_slow_tests(pytester):
pytester.makepyfile("""
import time
def test_slow():
time.sleep(0.05)
""")
result = pytester.runpytest("--warn-slower-than=0.01")
result.assert_outcomes(passed=1, warnings=1)
Custom assertion helpers
Shared assertion functions clutter tracebacks with frames nobody cares about. Set
__tracebackhide__ and the failure points at the caller instead.
import pytest
def assert_valid_invoice(invoice):
__tracebackhide__ = True # hide this frame in the report
if invoice.total < 0:
pytest.fail(f"negative total: {invoice.total}")
if invoice.lines and invoice.total != sum(l.amount for l in invoice.lines):
pytest.fail(
f"total {invoice.total} != sum of lines "
f"{sum(l.amount for l in invoice.lines)}"
)
def test_generated_invoice(order):
assert_valid_invoice(build_invoice(order)) # ← failure is reported here
To keep pytest's value-introspection inside a helper module, register it for rewriting before it is imported:
import pytest
pytest.register_assert_rewrite("mypackage.testing.helpers")
Why plain assert works
Python's assert throws away everything but the boolean. pytest installs an import hook that
rewrites the AST of test modules before they are compiled, replacing each assert with code that
stores intermediate values and builds an explanation on failure.
# what you write
assert compute(x) == expected
# roughly what gets compiled
@py_left = compute(x)
@py_right = expected
if not (@py_left == @py_right):
raise AssertionError(explain("==", @py_left, @py_right))
Three consequences worth knowing:
- Rewriting applies to test modules,
conftest.py, and modules you explicitly register. A plain library module keeps bareAssertionErrors with no detail. - Running Python with
-Ostripsassertstatements entirely, so never run a suite optimised. - Stale
__pycache__from a non-pytest run can shadow rewritten bytecode;pytest --cache-clearor deleting the caches fixes the "my assertions lost their detail" mystery.
Async tests
pytest cannot await a coroutine on its own — an async def test without a plugin is
skipped with a warning, which looks like a pass in a hurry. Install one of the two plugins.
# pyproject.toml
# [tool.pytest.ini_options]
# asyncio_mode = "auto" # no per-test marker needed
import asyncio, pytest
@pytest.fixture
async def client():
c = await open_client()
yield c
await c.aclose()
async def test_fetches_concurrently(client):
a, b = await asyncio.gather(client.get("/a"), client.get("/b"))
assert a.status == b.status == 200
anyio is the alternative when you need the same tests to run on both asyncio and trio.
Property-based testing
Instead of listing examples, describe the shape of valid input and assert a property that must hold for all of them. Hypothesis generates cases, and on failure shrinks them to the smallest reproducer.
from hypothesis import given, strategies as st
@given(st.text())
def test_slugify_is_idempotent(s):
once = slugify(s)
assert slugify(once) == once
@given(st.lists(st.integers()))
def test_sort_is_a_permutation(xs):
out = my_sort(xs)
assert len(out) == len(xs)
assert sorted(out) == sorted(xs)
assert all(a <= b for a, b in zip(out, out[1:]))
It composes with pytest normally — @given functions are collected like any other test. Good
properties are round-trips (decode(encode(x)) == x), invariants (length, ordering) and
equivalence to a slow reference implementation.
@given with function-scoped fixtures
Hypothesis calls the test body many times, but a function-scoped fixture is created once for the whole thing. Shared mutable state leaks between generated examples. Build what you need inside the test, or use a factory fixture.
Flaky tests and isolation
A flaky test is worse than no test: it trains people to re-run CI instead of reading it. The usual causes are few, and each has a direct fix.
| Cause | Fix |
|---|---|
| Shared mutable state across tests | Narrow the fixture scope; return copies |
| Order dependence | pytest -p no:randomly to confirm, then fix the coupling |
| Real clock / real timezone | Freeze time (monkeypatch, freezegun, time-machine) |
| Unseeded randomness | Seed it, and log the seed |
sleep()-based waiting | Poll for the condition with a timeout |
| Dict/set iteration order assumptions | Compare sets, or sort before comparing |
| Leftover files or env vars | tmp_path and monkeypatch, never manual mutation |
python -m pip install pytest-randomly
pytest # shuffles order every run, prints the seed
pytest -p no:randomly # reproduce the deterministic order
pytest -p randomly --randomly-seed=12345 # reproduce a specific shuffle
pytest-rerunfailures can keep a pipeline moving, but a test that passes on the second attempt is
telling you something real about your system. Retry only at genuinely non-deterministic boundaries
(a live network), and never on unit tests.
Reference
Anti-patterns
Logic in the test
A test with branching needs its own tests. Split it into parametrized cases so each path is visible in the report.
def test_prices():
for currency in ("EUR", "USD"):
if currency == "EUR":
assert fmt(1, currency) == "1,00 €"
else:
assert fmt(1, currency) == "$1.00"
@pytest.mark.parametrize(
("currency", "expected"),
[("EUR", "1,00 €"), ("USD", "$1.00")],
)
def test_prices(currency, expected):
assert fmt(1, currency) == expected
Asserting the implementation
mock.assert_called_once_with(...) on an internal helper locks in today's call graph. Rename the
helper and the test fails while the behaviour is unchanged. Assert on the observable result instead.
One test, many concerns
When a 40-line test fails you learn "something in registration is broken". Several focused tests tell you which part, and the names double as documentation.
Depending on execution order
A test that only passes after another one ran is not a test, it is half of one. Each test must set up everything it needs.
Sleeping
start_worker()
time.sleep(2) # slow AND flaky
assert job_done()
def wait_until(pred, timeout=5, interval=0.05):
__tracebackhide__ = True
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
if pred():
return
time.sleep(interval)
pytest.fail(f"condition not met in {timeout}s")
start_worker()
wait_until(job_done)
Cheat sheet
Command line
pytest path::test_name | Run one test |
-k "expr" / -m "expr" | Select by name / by marker |
-x / --maxfail=N | Stop early |
--lf / --ff / --sw | Last-failed / failures-first / stepwise |
-q / -v / -vv / -ra | Verbosity and outcome summary |
-s | Show output live |
--pdb | Debugger at the failure |
--tb=short | Compact tracebacks |
--collect-only | List tests without running |
--fixtures | List available fixtures |
--durations=10 | Slowest tests |
-n auto | Parallel (xdist) |
-W error | Turn warnings into failures |
API
pytest.fixture(scope=, params=, autouse=, ids=) | Define a fixture |
pytest.mark.parametrize(argnames, argvalues, ids=, indirect=) | Generate cases |
pytest.param(*values, id=, marks=) | One case with metadata |
pytest.raises(Exc, match=) | Expect an exception |
pytest.warns(W, match=) | Expect a warning |
pytest.approx(value, rel=, abs=) | Float comparison |
pytest.skip(reason) / pytest.fail(msg) | Skip / fail imperatively |
pytest.importorskip(name, minversion=) | Optional dependency |
pytest.register_assert_rewrite(module) | Rewrite asserts in a helper module |
Built-in fixtures
tmp_path, tmp_path_factory | Temporary directories |
monkeypatch | Patch attributes, env vars, cwd — auto-undone |
capsys, capfd | Captured stdout/stderr (Python level / fd level) |
caplog | Captured log records |
request | Test metadata, request.param, config access |
pytester | Run a nested pytest — for testing plugins |
recwarn | Record all warnings raised |
Wrapping up
What actually moves the needle
Reading a tool's full feature list rarely changes how anyone works. If you adopt three things from this post, make them these.
The short list
-
Turn on
--strict-markersandfilterwarnings = ["error"]. Two lines of config that convert a class of silent failures into loud ones, today, in an existing project. - Replace your ad-hoc setup with fixtures, and keep them function-scoped until proven slow. Almost every order-dependent flake traces back to state shared more widely than it needed to be.
-
Reach for
parametrizethe moment a test grows a loop or anif. Each case gets its own name and its own verdict, which is the difference between "billing is broken" and "billing is broken for zero-quantity line items".
The advanced material matters less often, but it is worth knowing it exists. The day you find
yourself copying the same conftest.py into a fourth repository, that is the signal to turn it
into a plugin — and by then the entry point is a two-line change.
Further reading
- docs.pytest.org — the reference; the "How-to guides" section is better than its name suggests
- Plugin list — around 1,500 plugins, and the one you need probably exists
- Hypothesis — property-based testing, covered briefly above
- coverage.py — what
pytest-covwraps, worth reading directly for branch coverage
No comments:
Post a Comment