"Premature optimization is the root of all evil." — Knuth.
But once you've identified a bottleneck, knowing how to fix it matters. Measure first, then optimize.
The golden rule
Don't guess where the bottleneck is. You'll be wrong 80% of the time. Profile.
Profiling tools
cProfile (stdlib)
python -m cProfile -o profile.stats your_script.py
Or in code:
import cProfile
cProfile.run("expensive_function()")
Output: per-function call count and time. Use pstats to analyze:
import pstats
p = pstats.Stats("profile.stats")
p.sort_stats("cumulative").print_stats(20)
Or visualize: snakeviz profile.stats (browser-based flame graph).
py-spy (sampling profiler, no code change)
pip install py-spy
sudo py-spy top --pid 1234 # like Linux `top` but for Python functions
sudo py-spy record -o flame.svg --pid 1234 # flame graph
Works on running processes; doesn't slow them down (samples). Great for production debugging.
memray (memory profiler)
pip install memray
memray run script.py
memray flamegraph memray-*.bin
Tracks allocations. Find memory leaks and allocation hotspots.
line_profiler (per-line timing)
@profile # added by line_profiler
def slow_function():
...
kernprof -l -v script.py
Per-line execution time. Useful for narrowing within a hot function.
Time it quickly
from timeit import timeit
timeit("'.'.join(['a', 'b', 'c'])", number=1000000)
Or in shell:
python -m timeit "'.'.join(['a', 'b', 'c'])"
For micro-benchmarks.
Common Python performance traps
Trap 1: List membership check
# Slow: O(N)
if x in big_list: ...
# Fast: O(1)
big_set = set(big_list)
if x in big_set: ...
in list is linear; in set is constant. Common 100x speedup.
Trap 2: String concatenation in loop
# Slow: O(N²) — each concat copies the whole string
result = ""
for s in many_strings:
result += s
# Fast: O(N)
result = "".join(many_strings)
"".join() is the idiom.
Trap 3: Repeated computation
# Slow: recomputes every iteration
for item in items:
if expensive_check(item):
...
# Fast: compute once
threshold = compute_threshold()
for item in items:
if item.value > threshold:
...
Hoist invariants out of the loop.
Trap 4: Function call overhead
# Hot loop in Python; each call has overhead
def small_op(x):
return x * 2
result = [small_op(x) for x in big_list]
# Faster: inline or use builtins
result = [x * 2 for x in big_list]
For micro-operations in hot loops, function calls add up.
Trap 5: Wrong data structure
# Slow: linear search per check
for item in items:
if any(item.id == x.id for x in lookup_list):
...
# Fast: hash lookup
lookup_set = {x.id for x in lookup_list}
for item in items:
if item.id in lookup_set:
...
Set/dict for lookups; list only for ordered iteration.
Trap 6: Reading huge files into memory
# Bad: loads 100GB
lines = open("huge.log").readlines()
for line in lines:
process(line)
# Good: stream
for line in open("huge.log"):
process(line)
Trap 7: Pickling cost
Multiprocessing serializes arguments. Big args = slow.
# Slow: 1GB dataframe pickled to each worker
process_pool.map(work, [big_df] * 100)
# Faster: shared memory or chunking
When Python is the wrong choice
If you've profiled and optimized but it's still too slow:
- NumPy / Pandas / Polars: vectorize numeric work. 10-1000x speedup.
- Cython: compile hot loops to C.
- PyPy: alternative interpreter; JIT compiles. 2-10x typical.
- Rewrite hot paths in Rust with PyO3 bindings.
- Use a C extension: many libraries (lxml, msgpack, ujson) are C-backed.
Most "slow Python" is solvable with better algorithms or better tooling; rewriting in Rust is the last resort.
The 80/20 of optimization
Top wins, in order:
- Better algorithm: O(N²) → O(N log N) is the biggest single win possible.
- Right data structure: set instead of list for lookups.
- NumPy/Pandas for numeric work.
- Cache repeated computations (
@cache). - Reduce work: filter early, lazy evaluation.
Things that rarely matter:
- Local var vs global var lookup (negligible).
- Tuple vs list for fixed data (negligible).
- f-string vs %-format (negligible).
- Micro-optimizations in cold code.
When to optimize
Don't optimize:
- Before profiling.
- Code that runs once.
- Code that's "slow" but takes 100ms (humans don't notice).
Do optimize:
- Code in a hot loop measured to be the bottleneck.
- Code on the critical path of a user-facing operation.
- Code consuming significant production resources.
Common profiling/optimization mistakes
- Optimizing without profiling. Wrong code; wasted time.
- Premature optimization. Code unreadable for negligible speedup.
- Micro-benchmarks for code in real context.
timeitlies about real-world performance. - Not measuring after. Did your optimization actually help?
- Refusing to use NumPy because "Python should be fast enough". Use the right tool.
Takeaway
Profile first (cProfile, py-spy for live, memray for memory). Identify the bottleneck. Fix it with: better algorithm > better data structure > NumPy/Pandas > caching > filtering early. Common traps: in list, string concat in loop, repeated computation, wrong data structure. Rewriting in another language is the LAST resort. Most Python perf issues are fixable in Python.
Bridge to Week 9: this profile-first habit scales with you. When the same pipeline runs on Spark, you will profile stages in the Spark UI instead of
cProfile— but the rule is identical: measure the bottleneck before touching the code.