Python gives you several ways to model data classes. The right choice depends on what you need.
Option 1: @dataclass
from dataclasses import dataclass
@dataclass
class User:
id: int
name: str
email: str
is_active: bool = True
u = User(id=1, name="Alice", email="alice@example.com")
print(u) # User(id=1, name='Alice', email='alice@example.com', is_active=True)
What you get:
__init__with all fields.__repr__(nice printing).__eq__(value equality).- Type hints are recorded.
Mutable by default. Pass frozen=True for immutable:
@dataclass(frozen=True)
class Point:
x: float
y: float
p = Point(1.0, 2.0)
p.x = 3.0 # raises FrozenInstanceError
When to use dataclass
- Internal data containers.
- Configuration objects.
- DTOs (data transfer objects) within your codebase.
- When you don't need runtime validation.
What dataclass DOESN'T do
- No runtime type checking.
User(id="not_an_int", name=1, email=None)happily runs. - No JSON deserialization (need to write it).
- No validation logic (need to add in
__post_init__).
Option 2: Pydantic
from pydantic import BaseModel, EmailStr, field_validator
class User(BaseModel):
id: int
name: str
email: EmailStr
is_active: bool = True
@field_validator("name")
@classmethod
def name_must_be_non_empty(cls, v):
if not v.strip():
raise ValueError("name cannot be empty")
return v
u = User(id=1, name="Alice", email="alice@example.com")
What Pydantic adds:
- Runtime validation:
User(id="not_int", ...)raisesValidationError. - Coercion:
User(id="42", ...)works; "42" coerced to int. - Custom validators: arbitrary logic on fields.
- JSON serialization/deserialization:
User.model_dump_json(),User.model_validate_json(json_str). - Email/URL/IP types: built-in validators.
When to use Pydantic
- API request/response models (FastAPI uses Pydantic).
- Configuration with validation.
- External data ingestion (parsing config files, API responses).
- Anywhere data crosses a trust boundary.
Pydantic v1 vs v2
Pydantic v2 (released 2023) is the current standard. Significant API changes from v1:
class Config:→model_config = ConfigDict(...).dict()→model_dump().parse_obj()→model_validate().Field(default=..., env=...)→Field(default=..., validation_alias=...).- Faster (Rust-based core).
If you see v1 syntax in old tutorials, look up the v2 equivalent.
Option 3: NamedTuple
from typing import NamedTuple
class Point(NamedTuple):
x: float
y: float
p = Point(1.0, 2.0)
print(p.x, p[0]) # both work: 1.0
What you get:
- Immutable (tuple-based).
- Indexable (
p[0]) AND attribute access (p.x). - Hashable.
- Very lightweight (no dict; tuple-backed).
When to use NamedTuple
- Immutable records (coordinates, RGB, timestamps with metadata).
- Returns from functions where you want named results without a full class.
- Where you want tuple semantics (unpacking, sorting).
def get_dimensions() -> Point:
return Point(x=10.0, y=20.0)
x, y = get_dimensions()
Lightweight and unpacks naturally. Compared to dataclass: faster construction, smaller memory, immutable always.
The decision matrix
| Need | Choice |
|---|---|
| Internal data, mutable | dataclass |
| Internal data, immutable | dataclass(frozen=True) or NamedTuple |
| Tuple-like with attribute access | NamedTuple |
| External data (API, config) requiring validation | Pydantic |
| API request/response (FastAPI) | Pydantic |
| Simple struct, max performance | NamedTuple |
Performance differences
For 1M instances:
- NamedTuple: fastest (tuple subclass).
- dataclass: fast.
- Pydantic: slower (validation overhead).
- regular class with
__init__: middle.
For most apps, performance is irrelevant. Choose by features. Pydantic's overhead is worth it at API boundaries; not worth it for internal hot loops.
Pydantic settings (for configuration)
from pydantic_settings import BaseSettings
class Settings(BaseSettings):
database_url: str
debug: bool = False
api_key: str
class Config:
env_file = ".env"
settings = Settings() # loads from env vars / .env file
The 12-factor config pattern: environment variables; validated on startup.
attrs (the third option, less common)
import attrs
@attrs.define
class User:
id: int
name: str
Predates dataclass; very similar. Use dataclass for new code (stdlib); attrs only if codebase already uses it.
Common mistakes
- Using Pydantic for internal data. Performance hit for no benefit if data doesn't cross trust boundary.
- Using dataclass for API boundaries. No validation; bad data crashes deep in logic.
- Mutable default in dataclass.
field: list = []is a footgun (shared across instances). Usefield: list = field(default_factory=list). - NamedTuple with too many fields. Becomes hard to use; dataclass is clearer at 5+ fields.
- Forgetting
frozen=Truefor hashable dataclass. Mutable dataclass isn't hashable by default.
Takeaway
dataclass for internal data (mutable by default). NamedTuple for immutable tuples with names. Pydantic for trust-boundary data (APIs, config, external sources). Choose by validation needs and use case, not by hype. Pydantic v2 is the current version; check tutorials are not v1.
Production Note: Dataclass Equality Gotcha (eq and eq=False)
By default, @dataclass generates an __eq__ method that compares field tuples, but only if both objects are of the exact same class:
from dataclasses import dataclass
@dataclass
class PointA:
x: int
y: int
@dataclass
class PointB:
x: int
y: int
a = PointA(1, 2)
b = PointB(1, 2)
print(a == b) # False! Dataclasses check type(self) is type(other)
Even though PointA and PointB share identical attribute names and values, a == b returns False because the generated __eq__ checks isinstance(other, self.__class__). To disable generated equality comparisons or support cross-class structural equality, use @dataclass(eq=False) and implement a custom __eq__.