13 - Pydantic¶
Previous: 12 - uv | Index: All guides | Next: 14 - Async Python
Quick reference for Pydantic v2: data validation with Python type hints, used by FastAPI, LLM structured outputs, agent frameworks and settings management.
Last verified: 2026-09-27. For newer changes, check the Official docs links in the Introduction.
Introduction¶
Before you start¶
You should know: Python classes, type hints and dicts (10 - Python Basics), and what JSON looks like (08).
The problem it solves: data from outside your program cannot be trusted: a form sends "42" instead of 42, an API leaves out a field, an LLM returns JSON with a missing key, a config file has a typo. Without checks, the bad value travels deep into your code and crashes somewhere far from the cause, with a confusing error.
Before Pydantic: you wrote checks by hand for every field (if "age" not in data, int(data["age"]), if age < 0), which is long, repetitive and easy to forget. Python type hints alone do not help at runtime: they describe types but never check them. Pydantic uses those same type hints to actually check and convert data.
Think of it like: customs at a border (the mental model below): everything entering the country is checked once, and inside you can trust what passed.
What is Pydantic?¶
Pydantic is a Python library that checks and converts data using type hints. You describe the shape of your data as a class (a model), and Pydantic makes sure any incoming data (JSON from an API, output from an LLM, values from a config file) matches that shape. If it does, you get a clean Python object with the right types. If not, you get a clear error saying exactly which field is wrong and why.
Mental model¶
Think of a Pydantic model as a customs checkpoint at the border of your program:
Outside world (untrusted) Border (Pydantic model) Inside your code (trusted)
------------------------- ----------------------- --------------------------
'{"age": "42", "email": "a@b.c"}' -> check types, convert "42" -> 42, -> user.age == 42 (int)
LLM output, API request body, check rules (age >= 0), user.email is a valid email
.env values, CSV rows reject with a clear error if bad safe to use, autocomplete works
Validate once at the boundary, then trust the object everywhere inside. The same model can also export a JSON Schema, which is exactly what LLM APIs and FastAPI use to describe inputs and outputs.
Why use it?¶
- Catch bad data early with precise error messages instead of crashes deep in your code.
- Automatic conversion:
"42"->42, ISO strings ->datetime, dicts -> nested models. - One definition, many uses: validation, JSON export, JSON Schema, editor autocomplete, docs.
- Core of the AI stack: FastAPI bodies, Claude / OpenAI structured outputs, tool definitions, LangChain, agent SDKs and
pydantic-settingsall use it. - Fast: the core is written in Rust.
Key terms¶
| Term | Meaning |
|---|---|
| Model | A class inheriting BaseModel that defines fields and types |
| Field | One attribute with a type (and optional rules / default) |
| Validation | Checking and converting input data |
| Serialization | Turning a model back into a dict / JSON |
| Validator | Your own function that checks or transforms a field |
| JSON Schema | Machine-readable description of the model, generated automatically |
| Coercion | Converting compatible values ("1" -> 1) |
| Strict mode | No coercion: types must match exactly |
Where it fits: used in 40 - FastAPI, 27 - LLM APIs (structured outputs), 29 - Tool Use, 33 - Agent Frameworks; builds on 10 - Python Basics type hints and 08 - JSON. In a web API the flow is: 41 - Uvicorn receives the request -> 40 - FastAPI routes it -> Pydantic validates it.
Official docs¶
Where to read the latest, authoritative documentation:
| Resource | Link |
|---|---|
| Pydantic documentation | https://docs.pydantic.dev/latest/ |
| pydantic-settings | https://docs.pydantic.dev/latest/concepts/pydantic_settings/ |
| Migration guide (v1 to v2) | https://docs.pydantic.dev/latest/migration/ |
Contents¶
- Install and First Model
- Field Types
- Optional Fields and Defaults
- Field Constraints
- Nested Models and Lists
- Validation Errors
- Custom Validators
- Serialization (dump)
- Parsing JSON and Dicts
- JSON Schema
- Model Configuration
- Enums and Literals
- Aliases (Different Names in JSON)
- Computed Fields
- TypeAdapter (Validate Without a Model)
- Settings from Environment (pydantic-settings)
- Pydantic for LLM Structured Output
- Pydantic vs dataclass vs TypedDict
- v1 to v2 Cheat Sheet
- Troubleshooting
- Try It
1. Install and First Model¶
Defining a model and creating validated objects. Subclass
BaseModel, declare fields with type hints, create instances with keyword arguments.Use it any time data enters your program from outside.
from pydantic import BaseModel
class User(BaseModel):
"""A registered user."""
id: int
name: str
active: bool = True
u = User(id="42", name="Ana") # "42" is converted to 42
u.id # 42
u # User(id=42, name='Ana', active=True)
u.name = "Bo" # models are mutable by default
2. Field Types¶
The types Pydantic understands and validates. Standard Python types plus special ones for common formats.
Use it for choosing the most precise type gives the best validation for free.
| Type | Accepts / validates |
|---|---|
int, float, str, bool |
Basic types (with sensible coercion) |
list[str], dict[str, int], set[int], tuple[int, int] |
Collections with element types |
datetime, date, time, timedelta |
ISO strings like "2026-09-27T10:00:00" |
UUID |
UUID strings |
Decimal |
Exact decimals (money) |
EmailStr |
Valid email address (needs pydantic[email]) |
HttpUrl, AnyUrl |
Valid URLs |
SecretStr |
Hidden when printed (**********) |
Path |
File system paths |
Literal["a", "b"] |
Only these exact values |
X \| None |
X or None |
Any |
Anything (no validation) |
3. Optional Fields and Defaults¶
Fields that may be missing or empty. A default value makes a field optional;
X | None = Noneallows missing and null.Use it for partial updates, optional metadata, LLM outputs where a value may not exist.
from pydantic import BaseModel, Field
class Item(BaseModel):
name: str # required
description: str | None = None # optional, may be null
tags: list[str] = [] # safe in Pydantic (copied per instance)
quantity: int = 1
created: datetime = Field(default_factory=datetime.now) # computed at creation
Required vs optional is decided by the default, not by | None: x: int | None without a default is still required (but may be null).
4. Field Constraints¶
Extra rules on values: ranges, lengths, patterns.
Field(...)with keyword arguments, orAnnotated[type, Field(...)]for reusable types.Use it for prices >= 0, names not empty, codes matching a pattern.
from typing import Annotated
from pydantic import BaseModel, Field
Percent = Annotated[float, Field(ge=0, le=100)] # reusable constrained type
class Product(BaseModel):
name: str = Field(min_length=1, max_length=100, description="Product name")
price: float = Field(gt=0)
discount: Percent = 0
sku: str = Field(pattern=r"^[A-Z]{3}-\d{4}$")
tags: list[str] = Field(default_factory=list, max_length=10)
| Keyword | Meaning |
|---|---|
gt, ge, lt, le |
>, >=, <, <= |
min_length, max_length |
Length of strings / lists |
pattern |
Regex the string must match |
multiple_of |
Number must be a multiple |
default, default_factory |
Default value / function producing it |
description, examples, title |
Documentation (shown in JSON Schema, FastAPI docs, LLM tool schemas) |
5. Nested Models and Lists¶
Models inside models, for tree-shaped data. Use a model as a field type; dicts are converted into model instances automatically.
Use it for orders with line items, API responses with nested objects, complex LLM outputs.
class Address(BaseModel):
street: str
city: str
class LineItem(BaseModel):
sku: str
qty: int = Field(ge=1)
class Order(BaseModel):
id: int
shipping: Address
items: list[LineItem]
order = Order.model_validate({
"id": 1,
"shipping": {"street": "Main 1", "city": "Berlin"},
"items": [{"sku": "ABC-0001", "qty": 2}],
})
order.items[0].qty # 2
6. Validation Errors¶
What happens when data does not match. Pydantic raises
ValidationErrorlisting every problem with its location and message.Use it for returning helpful errors to users, re-asking an LLM to fix its output.
from pydantic import ValidationError
try:
Product(name="", price=-5, sku="bad")
except ValidationError as e:
print(e.error_count()) # 3
for err in e.errors():
print(err["loc"], err["msg"])
# ('name',) String should have at least 1 character
# ('price',) Input should be greater than 0
# ('sku',) String should match pattern '^[A-Z]{3}-\d{4}$'
FastAPI turns this automatically into a 422 response.
7. Custom Validators¶
Your own checking / cleaning logic.
@field_validatorfor one field;@model_validatorfor rules that involve several fields.Use it for normalising text, cross-field rules (end date after start date).
from pydantic import BaseModel, field_validator, model_validator
class Booking(BaseModel):
guest: str
start: date
end: date
@field_validator("guest")
@classmethod
def clean_name(cls, v: str) -> str:
v = v.strip()
if not v:
raise ValueError("guest must not be empty")
return v.title()
@model_validator(mode="after")
def check_dates(self) -> "Booking":
if self.end <= self.start:
raise ValueError("end must be after start")
return self
mode="before" runs on raw input before type conversion (for example to split a comma-separated string into a list).
8. Serialization (dump)¶
Turning a model back into a dict or JSON string.
model_dump()-> dict,model_dump_json()-> str; options to include / exclude fields.Use it for returning API responses, saving to files / databases, sending to LLMs.
order.model_dump() # dict (nested models become dicts)
order.model_dump_json(indent=2) # JSON string
order.model_dump(exclude={"id"}) # leave out fields
order.model_dump(include={"items"}) # only some fields
order.model_dump(exclude_none=True) # drop None values
order.model_dump(exclude_unset=True) # only fields that were explicitly set (PATCH)
order.model_dump(by_alias=True) # use JSON aliases
order.model_dump(mode="json") # dict with JSON-safe values (datetime -> str)
order.model_copy(update={"id": 2}) # copy with changes
9. Parsing JSON and Dicts¶
Creating models from dicts or raw JSON text.
model_validate(dict)ormodel_validate_json(str)(faster thanjson.loads+ validate).Use it for API responses, LLM JSON, files.
user = User.model_validate({"id": 1, "name": "Ana"})
user = User.model_validate_json('{"id": 1, "name": "Ana"}')
users = [User.model_validate(row) for row in rows]
with open("user.json", encoding="utf-8") as f:
user = User.model_validate_json(f.read())
10. JSON Schema¶
An automatically generated description of the model.
model_json_schema()returns a dict following the JSON Schema standard, including descriptions and constraints.Use it for LLM tool definitions, structured output schemas, API documentation.
class WeatherQuery(BaseModel):
"""Get the weather for a city."""
city: str = Field(description="City name, e.g. Berlin")
days: int = Field(1, ge=1, le=7, description="Forecast length")
WeatherQuery.model_json_schema()
# {'description': 'Get the weather for a city.',
# 'properties': {'city': {'description': 'City name, e.g. Berlin', 'title': 'City', 'type': 'string'},
# 'days': {'default': 1, 'maximum': 7, 'minimum': 1, ...}},
# 'required': ['city'], 'title': 'WeatherQuery', 'type': 'object'}
Good descriptions matter: the LLM reads them to decide how to fill the fields.
11. Model Configuration¶
Settings that change how a model behaves.
model_config = ConfigDict(...)inside the class.Use it for rejecting unknown fields, immutability, reading from ORM objects.
from pydantic import BaseModel, ConfigDict
class Strict(BaseModel):
model_config = ConfigDict(
extra="forbid", # unknown fields -> error ("ignore" is default, "allow" keeps them)
frozen=True, # immutable (and hashable)
str_strip_whitespace=True, # trim strings
strict=False, # True = no coercion ("42" stays invalid for int)
from_attributes=True, # build from objects (SQLAlchemy rows): Model.model_validate(row)
populate_by_name=True, # accept field name as well as alias
)
12. Enums and Literals¶
Restricting a field to a fixed set of values.
Literal[...]for simple cases;Enumwhen you want a named, reusable type.Use it for status fields, categories, LLM classification labels.
from enum import Enum
from typing import Literal
class Sentiment(str, Enum):
positive = "positive"
negative = "negative"
neutral = "neutral"
class Review(BaseModel):
sentiment: Sentiment
priority: Literal["low", "medium", "high"]
Review(sentiment="positive", priority="high").sentiment # <Sentiment.positive: 'positive'>
13. Aliases (Different Names in JSON)¶
Using one name in Python and another in JSON.
Field(alias=...), or analias_generatorfor all fields (e.g. camelCase).Use it for external APIs that use
camelCaseor names that are Python keywords.
from pydantic import ConfigDict
from pydantic.alias_generators import to_camel
class ApiUser(BaseModel):
model_config = ConfigDict(alias_generator=to_camel, populate_by_name=True)
user_id: int
first_name: str
u = ApiUser.model_validate({"userId": 1, "firstName": "Ana"})
u.model_dump(by_alias=True) # {'userId': 1, 'firstName': 'Ana'}
14. Computed Fields¶
Read-only values derived from other fields, included in output.
@computed_fieldon a@property.Use it for totals, full names, derived flags you want in the JSON output.
from pydantic import computed_field
class Cart(BaseModel):
prices: list[float]
@computed_field
@property
def total(self) -> float:
return round(sum(self.prices), 2)
Cart(prices=[1.5, 2.25]).model_dump() # {'prices': [1.5, 2.25], 'total': 3.75}
15. TypeAdapter (Validate Without a Model)¶
Validating plain types like
list[int]ordict[str, User]without writing a model class. Wrap the type inTypeAdapter, then callvalidate_python/validate_json.Use it for validating a list of models from an API, quick checks.
from pydantic import TypeAdapter
users = TypeAdapter(list[User]).validate_json('[{"id": 1, "name": "Ana"}]')
TypeAdapter(list[User]).json_schema()
16. Settings from Environment (pydantic-settings)¶
A typed config object loaded from environment variables and
.env. SubclassBaseSettings; field names map to env var names (case-insensitive).Use it in every app with API keys, URLs or feature flags.
from pydantic import SecretStr
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
"""Application configuration."""
model_config = SettingsConfigDict(env_file=".env", env_prefix="APP_", extra="ignore")
anthropic_api_key: SecretStr # reads APP_ANTHROPIC_API_KEY
llm_model: str = "claude-opus-5"
max_tokens: int = 4096
debug: bool = False # "true", "1", "yes" -> True
settings = Settings() # fails fast at startup if a required value is missing
settings.anthropic_api_key.get_secret_value()
17. Pydantic for LLM Structured Output¶
Getting LLM answers as validated Python objects instead of free text. Define a model; the SDK sends its JSON Schema and parses the reply into the model.
Use it for extraction, classification, any time code (not a human) reads the LLM output.
import anthropic
from pydantic import BaseModel, Field
class Invoice(BaseModel):
vendor: str
total: float = Field(description="Total amount in EUR")
due_date: date | None = None
line_items: list[str]
client = anthropic.Anthropic()
response = client.messages.parse(
model="claude-opus-5",
max_tokens=16000,
messages=[{"role": "user", "content": f"Extract the invoice data:\n\n{invoice_text}"}],
output_format=Invoice,
)
invoice = response.parsed_output # validated Invoice instance
More in 27 - LLM APIs and 28 - Prompt Engineering.
18. Pydantic vs dataclass vs TypedDict¶
Three ways to describe structured data in Python. They differ in whether data is validated at runtime.
Use it for Pydantic at boundaries (untrusted input); dataclasses for internal data; TypedDict to type plain dicts.
BaseModel |
@dataclass |
TypedDict |
|
|---|---|---|---|
| Runtime validation | Yes | No | No |
| Type conversion | Yes | No | No |
| JSON / Schema export | Built in | Manual | Manual |
| Speed / overhead | Small cost | Fastest | Plain dict |
| Best for | API input, LLM output, config | Internal objects | Typing existing dicts |
19. v1 to v2 Cheat Sheet¶
Renamed methods between Pydantic v1 and v2. Old tutorials use v1 names; v2 (current) uses the
model_prefix.Use it for reading older code or Stack Overflow answers.
| v1 | v2 |
|---|---|
.dict() |
.model_dump() |
.json() |
.model_dump_json() |
.parse_obj(d) |
.model_validate(d) |
.parse_raw(s) |
.model_validate_json(s) |
.schema() |
.model_json_schema() |
.copy() |
.model_copy() |
@validator |
@field_validator |
@root_validator |
@model_validator |
class Config: |
model_config = ConfigDict(...) |
orm_mode = True |
from_attributes=True |
BaseSettings in pydantic |
pip install pydantic-settings |
20. Troubleshooting¶
| Error | Fix |
|---|---|
ValidationError: Field required |
Missing key in input, or give the field a default |
Input should be a valid integer |
Wrong type; check the data or use a looser type |
Extra inputs are not permitted |
extra="forbid" and an unknown key was sent |
AttributeError: 'X' object has no attribute 'dict' |
Pydantic v2: use model_dump() |
PydanticImportError: BaseSettings has been moved |
pip install pydantic-settings; import from pydantic_settings |
email-validator is not installed |
pip install "pydantic[email]" |
| Validator not running | Check the field name string in @field_validator("...") and @classmethod |
datetime not JSON serializable |
Use model_dump_json() or model_dump(mode="json") |
| Secret printed in logs | Use SecretStr and .get_secret_value() only where needed |
| Mutable default shared between instances (dataclass) | Pydantic copies defaults; for dataclasses use field(default_factory=list) |
21. Try It¶
Short exercises to practise this guide. Try each task yourself first, then open the solution.
Use it right after reading the guide, or later as a quick self-test.
Exercise 1: Validate an order¶
Create Order with id (int > 0), items (list of str, at least 1) and total (float >= 0). Validate {"id": 0, "items": [], "total": -1} and print each error.
Solution
from pydantic import BaseModel, Field, ValidationError
class Order(BaseModel):
id: int = Field(gt=0)
items: list[str] = Field(min_length=1)
total: float = Field(ge=0)
try:
Order.model_validate({"id": 0, "items": [], "total": -1})
except ValidationError as e:
for err in e.errors():
print(err["loc"], err["msg"])
Exercise 2: Normalise a field¶
Add a validator that strips and title-cases a name field.
Solution
Exercise 3: Settings from env¶
Read APP_DEBUG (bool) and APP_MODEL (default claude-opus-5) from the environment.
Solution
Previous: 12 - uv | Index: All guides | Next: 14 - Async Python