15 - pytest¶
Previous: 14 - Async Python | Index: All guides | Next: 16 - Jupyter
Quick reference for testing Python code with pytest: writing tests, fixtures, parametrising, mocking APIs and LLM calls, testing FastAPI, and coverage.
Last verified: 2026-09-27. For newer changes, check the Official docs links in the Introduction.
Introduction¶
Before you start¶
You should know: Python functions and assert (10 - Python Basics), and how to run commands in your project environment (12 - uv).
The problem it solves: every change to your code can silently break something that used to work. Checking everything by hand after each change is slow, boring and gets skipped. Automated tests check your code in seconds, every time, and tell you exactly what broke. They also let you change code with confidence, because the tests will catch mistakes.
Before pytest: Python's built-in unittest module (copied from Java's JUnit) requires test classes and special methods like self.assertEqual(a, b). pytest lets you write plain functions with plain assert, finds them automatically, and shows readable explanations when they fail, which is why it became the standard.
Think of it like: a checklist a pilot runs before every flight. It is fast, it is the same every time, and it catches the problem on the ground instead of in the air (in production).
What is testing and what is pytest?¶
A test is a small piece of code that runs your code and checks the result is what you expect. pytest is the most popular Python testing tool: you write plain functions starting with test_ that use assert, and pytest finds them, runs them and reports which passed or failed with a clear explanation.
Mental model¶
Tests are a safety net under your code. Every test follows three steps, often called Arrange - Act - Assert:
Arrange -> prepare inputs and fake dependencies df = make_sample_df()
Act -> call the thing you are testing result = clean(df)
Assert -> check the outcome assert result["age"].isna().sum() == 0
The test pyramid tells you how many of each kind to write:
/\ end-to-end tests (few, slow): full app + real services
/ \
/----\ integration tests (some): API endpoint + database
/ \
/--------\ unit tests (many, fast): one function, fake dependencies
For AI apps, deterministic tests check your code (parsing, tool routing, prompts built correctly) with the LLM mocked; evals check the model's behaviour (see 35 - Evals and Observability).
Why use it?¶
- Change code without fear: tests tell you instantly if you broke something.
- Faster debugging: a failing test points to the exact function and input.
- Living documentation: tests show how code is meant to be used.
- Required in teams and CI: pull requests run tests automatically (44 - GitHub Actions).
Key terms¶
| Term | Meaning |
|---|---|
| Test function | def test_something(): with assert statements |
| Assertion | A check that must be true |
| Fixture | Reusable setup (data, client, temp folder) injected by name |
| Parametrize | Run the same test with many inputs |
| Mock / fake / stub | A stand-in for a real dependency (API, database, LLM) |
| Monkeypatch | Temporarily replace an attribute or env var during a test |
| Coverage | Share of code lines executed by tests |
| Regression test | Test that reproduces a fixed bug so it never returns |
| Flaky test | Sometimes passes, sometimes fails (timing, randomness, network) |
Where it fits: tests code from 10 - Python Basics to 40 - FastAPI; runs in CI via 44 - GitHub Actions; LLM quality is measured with 35 - Evals.
Official docs¶
Where to read the latest, authoritative documentation:
| Resource | Link |
|---|---|
| pytest documentation | https://docs.pytest.org/ |
| unittest.mock | https://docs.python.org/3/library/unittest.mock.html |
| pytest-asyncio | https://pytest-asyncio.readthedocs.io/ |
| pytest-cov | https://pytest-cov.readthedocs.io/ |
Contents¶
- Flags and Parameters
- Install and First Test
- Project Layout and Discovery
- Assertions
- Testing Exceptions and Warnings
- Fixtures
- Fixture Scope and conftest.py
- Built-in Fixtures
- Parametrize
- Markers (skip, xfail, custom)
- Mocking with unittest.mock
- Mocking LLM Calls
- Mocking HTTP Requests
- Testing pandas Code
- Testing FastAPI
- Async Tests
- Coverage
- Configuration (pyproject.toml)
- Good Testing Habits
- Troubleshooting
- Try It
0. Flags and Parameters¶
The pytest command-line options you use every day.
pytest [options] [paths or node ids].Use this when you see
pytest -x -k "login and not slow" -vv tests/and want to know what each part does.
pytest -x -k "login and not slow" -vv tests/test_auth.py::test_login_ok
| | | | |
| | | | +-- node id: one file, one test
| | | +------- very verbose: full diffs
| | +-------------------------------- only tests whose name matches the expression
| +------------------------------------ stop at the first failure
+-------------------------------------------- test runner
| Flag | Meaning |
|---|---|
-q / -v / -vv |
Quiet / verbose / very verbose |
-x |
Stop after the first failure |
--maxfail=3 |
Stop after 3 failures |
-k "expr" |
Run tests whose names match (and, or, not) |
-m slow |
Run tests with a marker (-m "not slow" to skip them) |
--lf |
Re-run only the tests that failed last time |
--ff |
Run failed tests first, then the rest |
-s |
Show print() output (disable capture) |
--pdb |
Open the debugger at the failure |
-l |
Show local variables in tracebacks |
--durations=10 |
Show the 10 slowest tests |
-n auto |
Run in parallel on all cores (plugin pytest-xdist) |
--cov=src --cov-report=term-missing |
Coverage report (plugin pytest-cov) |
file.py::TestClass::test_name |
Run one specific test |
1. Install and First Test¶
Writing and running the simplest test. A file named
test_*.pywith functions namedtest_*that useassert.Use it at the start of every project, even with one function.
# src/pricing.py
def add_tax(price: float, rate: float = 0.19) -> float:
"""Return price including tax, rounded to cents."""
return round(price * (1 + rate), 2)
# tests/test_pricing.py
from pricing import add_tax
def test_add_tax_default_rate():
assert add_tax(100) == 119.0
def test_add_tax_custom_rate():
assert add_tax(100, rate=0.07) == 107.0
2. Project Layout and Discovery¶
Where tests live and how pytest finds them. pytest collects files
test_*.py/*_test.py, functionstest_*, classesTest*(no__init__).Use it for setting up a new project.
project/
pyproject.toml
src/
myapp/
__init__.py
pricing.py
tests/
conftest.py shared fixtures
test_pricing.py
test_api.py
With a src/ layout install your package in editable mode (pip install -e . / uv sync) or set pythonpath = ["src"] in the pytest config so imports work.
3. Assertions¶
Checking results with plain
assert. pytest rewrites asserts to show both sides of a failed comparison.Use it in every test.
assert result == 42
assert "error" not in message
assert items == ["a", "b"]
assert user.is_active
assert len(rows) > 0
assert 0.1 + 0.2 == pytest.approx(0.3) # floats: never use == directly
assert {"a": 1}.items() <= result.items() # dict contains these keys / values
assert result == 42, f"unexpected result for input {x}" # custom message
4. Testing Exceptions and Warnings¶
Checking that bad input raises the right error.
pytest.raisesas a context manager;match=checks the message with a regex.Use it for validation logic, error paths.
import pytest
def test_negative_price_raises():
with pytest.raises(ValueError, match="must be positive"):
add_tax(-5)
def test_deprecation():
with pytest.warns(DeprecationWarning):
old_function()
5. Fixtures¶
Reusable setup code that tests receive as arguments. Decorate a function with
@pytest.fixture; any test with a parameter of the same name gets its return value. Code afteryieldruns as cleanup.Use it for sample data, clients, temp files, database connections.
import pandas as pd
import pytest
@pytest.fixture
def sample_df():
return pd.DataFrame({"age": [25, None, 40], "city": ["Berlin", "Paris", None]})
@pytest.fixture
def db():
conn = connect_test_db()
yield conn # the test runs here
conn.close() # cleanup, even if the test failed
def test_clean_fills_missing(sample_df):
cleaned = clean(sample_df)
assert cleaned["age"].isna().sum() == 0
Fixtures can use other fixtures by listing them as parameters.
6. Fixture Scope and conftest.py¶
How often a fixture is created, and where shared fixtures live.
scope="function"(default, fresh per test),"module","session"(once per run); fixtures inconftest.pyare available to all tests in that folder.Use it for expensive setup (load a model, start a DB) shared by many tests.
# tests/conftest.py
import pytest
@pytest.fixture(scope="session")
def embedding_model():
return load_small_model() # loaded once for the whole test run
@pytest.fixture(autouse=True)
def no_real_api_keys(monkeypatch):
monkeypatch.setenv("ANTHROPIC_API_KEY", "test-key") # applied to every test automatically
7. Built-in Fixtures¶
Fixtures pytest provides out of the box. Just add their name as a test parameter.
Use it for temp files, env vars, captured output, logs.
| Fixture | Gives you |
|---|---|
tmp_path |
A fresh temporary folder (pathlib.Path) |
monkeypatch |
Set / delete env vars, attributes, dict items temporarily |
capsys |
Captured print output (capsys.readouterr().out) |
caplog |
Captured log records |
request |
Info about the running test (for advanced fixtures) |
def test_writes_report(tmp_path):
out = tmp_path / "report.csv"
write_report(out)
assert out.read_text().startswith("date,total")
def test_reads_model_from_env(monkeypatch):
monkeypatch.setenv("LLM_MODEL", "test-model")
assert get_settings().llm_model == "test-model"
8. Parametrize¶
Running one test function with many input / expected pairs.
@pytest.mark.parametrize("args", [cases]); each case is reported as its own test.Use it for edge cases, tables of inputs, regex / parser tests.
@pytest.mark.parametrize(
("price", "rate", "expected"),
[
(100, 0.19, 119.0),
(0, 0.19, 0.0),
(9.99, 0.07, 10.69),
],
ids=["normal", "zero", "reduced-rate"],
)
def test_add_tax(price, rate, expected):
assert add_tax(price, rate) == expected
9. Markers (skip, xfail, custom)¶
Labels that change how tests run.
@pytest.mark.<name>; select or skip with-m.Use it for slow tests, tests needing real API keys, known bugs.
import os
import sys
import pytest
@pytest.mark.skip(reason="feature not ready")
def test_future(): ...
@pytest.mark.skipif(sys.platform == "win32", reason="Linux only")
def test_permissions(): ...
@pytest.mark.skipif(not os.getenv("ANTHROPIC_API_KEY"), reason="needs a real key")
@pytest.mark.integration
def test_real_llm_call(): ...
@pytest.mark.xfail(reason="bug #123, fix pending")
def test_known_bug(): ...
Register custom markers in pyproject.toml (section 17), then run pytest -m "not integration" locally and in fast CI.
10. Mocking with unittest.mock¶
Replacing a real dependency with a fake object you control.
unittest.mock.patchswaps an attribute during the test;MagicMockrecords calls and returns what you tell it.Use it for external APIs, time, randomness, anything slow or non-deterministic.
from unittest.mock import MagicMock, patch
def test_sends_email_on_signup():
with patch("myapp.users.send_email") as fake_send: # patch WHERE IT IS USED
signup("ana@example.com")
fake_send.assert_called_once_with("ana@example.com", subject="Welcome")
fake = MagicMock(return_value=42)
fake() # 42
fake.side_effect = ValueError("boom") # raise instead
fake.side_effect = [1, 2, 3] # return different values on each call
fake.call_count ; fake.call_args
Patch the name in the module that uses it (myapp.users.send_email), not where it is defined.
11. Mocking LLM Calls¶
Testing your LLM app code without calling a real model. Inject a fake client (dependency injection) or patch the SDK method to return a canned response.
Use it for unit tests for prompt building, output parsing, tool routing, error handling. Fast, free and deterministic.
# app.py - the client is a parameter, so tests can pass a fake one
def summarize(text: str, client) -> str:
msg = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[{"role": "user", "content": f"Summarize:\n\n{text}"}],
)
return next(b.text for b in msg.content if b.type == "text")
# tests/test_app.py
from types import SimpleNamespace
from unittest.mock import MagicMock
from app import summarize
def fake_message(text: str):
return SimpleNamespace(content=[SimpleNamespace(type="text", text=text)], stop_reason="end_turn")
def test_summarize_returns_text_and_sends_prompt():
client = MagicMock()
client.messages.create.return_value = fake_message("Short summary.")
result = summarize("Long article...", client)
assert result == "Short summary."
sent = client.messages.create.call_args.kwargs
assert "Long article..." in sent["messages"][0]["content"]
Also test the unhappy paths: empty response, stop_reason == "max_tokens", API errors (side_effect=...), invalid JSON.
12. Mocking HTTP Requests¶
Faking HTTP responses for code that uses requests or httpx. Plugins intercept outgoing calls:
responses(requests),respx(httpx).Use it for testing API clients and error handling (404, 429, timeouts).
import responses
@responses.activate
def test_get_user():
responses.get("https://api.example.com/users/1", json={"id": 1, "name": "Ana"}, status=200)
assert get_user(1)["name"] == "Ana"
@responses.activate
def test_retries_on_429():
responses.get("https://api.example.com/x", status=429)
responses.get("https://api.example.com/x", json={"ok": True})
assert fetch_with_retry("https://api.example.com/x") == {"ok": True}
13. Testing pandas Code¶
Comparing DataFrames and Series in tests.
pandas.testinghelpers give readable diffs and handle NaN and dtypes.Use it for data cleaning and feature engineering functions.
import pandas as pd
from pandas.testing import assert_frame_equal, assert_series_equal
def test_clean_names():
df = pd.DataFrame({"name": [" ana ", "BO"]})
expected = pd.DataFrame({"name": ["Ana", "Bo"]})
assert_frame_equal(clean_names(df), expected)
assert_frame_equal(a, b, check_dtype=False, check_like=True) # ignore dtypes / column order
14. Testing FastAPI¶
Calling your API endpoints in tests without starting a server.
TestClient(app)sends requests directly to the app;dependency_overridesswaps dependencies (DB, settings, LLM client).Use it in every endpoint: status codes, validation errors, response shape.
import pytest
from fastapi.testclient import TestClient
from app.main import app, get_llm_client
@pytest.fixture
def client():
fake_llm = MagicMock()
fake_llm.messages.create.return_value = fake_message("positive")
app.dependency_overrides[get_llm_client] = lambda: fake_llm
yield TestClient(app)
app.dependency_overrides.clear()
def test_classify_ok(client):
r = client.post("/classify", json={"text": "I love it"})
assert r.status_code == 200
assert r.json() == {"label": "positive"}
def test_classify_validation_error(client):
assert client.post("/classify", json={}).status_code == 422
15. Async Tests¶
Testing
async deffunctions. Thepytest-asyncioplugin runs async tests in an event loop.Use it for async clients, agents, async FastAPI dependencies.
import pytest
from unittest.mock import AsyncMock
@pytest.mark.asyncio
async def test_classify_async():
client = MagicMock()
client.messages.create = AsyncMock(return_value=fake_message("neutral"))
assert await classify("ok", client) == "neutral"
Set asyncio_mode = "auto" in config to skip the decorator.
16. Coverage¶
Measuring which lines your tests execute.
pytest-covruns coverage.py during the test run.Use it for finding untested code; CI quality gates. High coverage does not guarantee good tests.
pip install pytest-cov
pytest --cov=src --cov-report=term-missing # shows missing line numbers
pytest --cov=src --cov-report=html # open htmlcov/index.html
pytest --cov=src --cov-fail-under=80 # fail if below 80%
17. Configuration (pyproject.toml)¶
Project-wide pytest settings.
[tool.pytest.ini_options]inpyproject.tomlis read automatically.Use it for set test paths, default flags and markers once.
[tool.pytest.ini_options]
testpaths = ["tests"]
pythonpath = ["src"]
addopts = "-q --strict-markers"
asyncio_mode = "auto"
markers = [
"integration: calls real external services (deselect with -m 'not integration')",
"slow: takes more than a few seconds",
]
filterwarnings = ["error::DeprecationWarning"]
18. Good Testing Habits¶
Practices that keep tests useful. Small, independent, fast, deterministic tests with clear names.
Use it always.
- Name tests after behaviour:
test_refund_fails_when_order_already_refunded. - One behaviour per test; several asserts about that behaviour are fine.
- Tests must not depend on each other or on execution order.
- No real network / LLM calls in unit tests; mark real ones
integration. - Fix randomness with seeds (
random.seed(0),np.random.default_rng(0)). - When you fix a bug, first write a test that reproduces it.
- Run tests before every commit and in CI.
19. Troubleshooting¶
| Problem | Fix |
|---|---|
ModuleNotFoundError for your package |
pythonpath = ["src"] in config or pip install -e . |
collected 0 items |
File / function names must start with test_; check testpaths |
Fixture 'x' not found |
Typo, or fixture defined in a conftest.py outside the test's folder |
| Patch has no effect | Patch where the name is used (myapp.module.func), not where it is defined |
Float comparison fails (0.30000000000000004) |
pytest.approx |
PytestUnknownMarkWarning |
Register the marker in pyproject.toml |
| Async test "passes" without running / coroutine never awaited | Install pytest-asyncio, add marker or asyncio_mode = "auto" |
| Tests pass alone, fail together | Shared state between tests; use fresh fixtures, reset globals, dependency_overrides.clear() |
| Flaky test | Remove time / network / randomness dependencies; mock them |
print output not shown |
pytest -s |
20. Try It¶
Short exercises to practise this guide. Try each task yourself first, then open the solution.
Use it right after reading the guide, or later as a quick self-test.
Exercise 1: Parametrize¶
Test add_tax(price, rate) for (100, 0.19) -> 119.0, (0, 0.19) -> 0.0 and (9.99, 0.07) -> 10.69 in one test.
Solution
Exercise 2: Temp file fixture¶
Write a fixture that creates a small CSV in a temporary folder.
Solution
Exercise 3: Mock the LLM¶
Run the tests of examples/tool_agent and read how fake_tool_use_message drives the agent loop without an API key.
Solution
client.messages.create.side_effect = [tool_use_message, text_message] returns a different fake response on each call, so the test walks through the loop step by step.
Previous: 14 - Async Python | Index: All guides | Next: 16 - Jupyter