Generators are one of Python's best features. They let you process data of any size with O(1) memory and compose pipelines elegantly.
Generator basics
def squares(n):
for i in range(n):
yield i * i
for sq in squares(10):
print(sq)
yield pauses the function; next call resumes where it left off. Function returns a generator OBJECT, not a list.
Generator expressions
Compact syntax:
squares = (i * i for i in range(10))
Like list comprehension but with parens — produces a generator, not a list.
Why generators
For large datasets:
# BAD: loads all 1B lines into memory
lines = open("huge.log").readlines()
for line in lines:
process(line)
# GOOD: streams line by line, O(1) memory
for line in open("huge.log"):
process(line)
open() is a generator (iterates lines lazily). Same pattern for any streaming source.
Composing generators
def read_lines(path):
with open(path) as f:
for line in f:
yield line.strip()
def parse_json(lines):
for line in lines:
yield json.loads(line)
def filter_errors(records):
for r in records:
if r.get("level") == "ERROR":
yield r
def get_timestamps(records):
for r in records:
yield r["timestamp"]
# Pipeline: each stage is lazy
pipeline = get_timestamps(filter_errors(parse_json(read_lines("logs.txt"))))
for ts in pipeline:
print(ts)
Each generator stage processes one item at a time; no intermediate list created. Memory: O(1). Reads 1TB of logs without OOM.
itertools — the standard toolkit
Built-in functions for working with iterators:
chain — concatenate iterables
from itertools import chain
result = list(chain([1, 2], [3, 4], [5])) # [1, 2, 3, 4, 5]
islice — slice an iterator
from itertools import islice
first_10 = list(islice(open("huge.log"), 10)) # first 10 lines
groupby — group consecutive equal items
from itertools import groupby
for key, group in groupby([1, 1, 2, 2, 2, 3]):
print(key, list(group))
# 1 [1, 1]
# 2 [2, 2, 2]
# 3 [3]
Note: groups CONSECUTIVE items; sort first if you want to group all equals.
product, permutations, combinations
from itertools import product
list(product([1,2], [3,4])) # [(1,3), (1,4), (2,3), (2,4)]
from itertools import combinations
list(combinations([1,2,3,4], 2)) # all pairs: (1,2), (1,3), (1,4), (2,3), (2,4), (3,4)
count, cycle, repeat — infinite iterators
from itertools import count, cycle, repeat
counter = count(1) # 1, 2, 3, ...
cycler = cycle(['A', 'B', 'C']) # A, B, C, A, B, C, ...
Use with islice to bound:
list(islice(count(), 5)) # [0, 1, 2, 3, 4]
tee — split one iterator into multiple
from itertools import tee
a, b = tee(iter(range(5)))
list(a) # [0, 1, 2, 3, 4]
list(b) # [0, 1, 2, 3, 4]
Useful when you need to iterate same data twice; works on generators that you can't restart.
accumulate — running totals/maximums
from itertools import accumulate
list(accumulate([1, 2, 3, 4])) # [1, 3, 6, 10] (running sum)
list(accumulate([3, 1, 4, 1, 5], max)) # [3, 3, 4, 4, 5] (running max)
When generators win
- Streaming: file processing, log analysis.
- Pipelines: each stage transforms one item at a time.
- Infinite sequences: don't materialize what you don't need.
- Large datasets: process anything that fits on disk.
When generators DON'T win
- Random access needed (can't index a generator).
- Multiple iterations over same data (generators are exhausted after one pass).
- Small datasets (overhead of generator vs list is negligible; list is simpler).
Reading a generator twice
Generators are one-shot:
gen = (x*x for x in range(5))
list(gen) # [0, 1, 4, 9, 16]
list(gen) # [] — exhausted
To iterate twice: re-create the generator, OR convert to list (loses memory benefit).
itertools.tee splits an iterator into multiple copies (caches internally).
Generator vs list comprehension
squares_list = [x*x for x in range(10**6)] # 1M items in memory
squares_gen = (x*x for x in range(10**6)) # generator object
sum(squares_list) # works; uses memory
sum(squares_gen) # works; O(1) memory
For sum, max, min, any True, all: pass generator directly.
Generator-based coroutines (legacy)
Before async/await, Python had generator-based coroutines using yield. Still works but use async def for new code.
Common generator mistakes
- Using lists where generators would do. Memory waste on large datasets.
- Iterating a generator twice (it's exhausted). Convert to list if you need multi-pass.
- Forgetting that generators are stateful. Side effects of iteration.
- Premature materialization.
list(gen)defeats the purpose. - Confusing
yieldwithreturn. Yielding suspends; returning ends generator.
Practical example: log processor
def process_logs(log_path):
return (
record
for line in open(log_path)
for record in [json.loads(line)]
if record.get("level") == "ERROR"
if record["timestamp"] > "2026-01-01"
)
for record in process_logs("huge.log"):
send_to_alerting(record)
Streams a 100GB log file with constant memory. Each stage is lazy.
Takeaway
Generators: lazy iteration, O(1) memory, perfect for streaming and pipelines. yield instead of return. Generator expressions: (x for x in ...). itertools toolkit: chain, islice, groupby, product, combinations, accumulate. Compose generators into pipelines — each stage lazy. One-shot; convert to list if multi-pass needed. Foundational for memory-efficient Python.