Python has three concurrency tools: threading, multiprocessing, asyncio. The right choice depends on the workload. Knowing the GIL helps.
The GIL (Global Interpreter Lock)
CPython has a global lock: only one thread can execute Python bytecode at a time.
Implications:
- Multi-threaded Python doesn't actually run in parallel for CPU work.
- Threads share memory; switching is cheap.
- I/O operations release the GIL (so threads ARE useful for I/O concurrency).
The GIL exists because reference counting (Python's memory management) isn't thread-safe without it.
In 2026, Python 3.13 added "free-threaded mode" (no-GIL): experimental, opt-in. Major libraries are slowly becoming compatible. Mainstream by ~2027-2028.
Threading
Multiple threads in one process. Share memory. GIL limits CPU parallelism but not I/O.
from concurrent.futures import ThreadPoolExecutor
def fetch_url(url):
return requests.get(url).text
urls = [...]
with ThreadPoolExecutor(max_workers=10) as executor:
results = list(executor.map(fetch_url, urls))
Use when:
- I/O-bound + you're stuck with sync libraries (can't use asyncio).
- Few hundred concurrent operations.
- Need shared mutable state across workers (use locks carefully).
Don't use when:
- CPU-bound (GIL prevents parallelism).
- Massive concurrency (1000+ threads → expensive context switching).
- You can use asyncio instead (better for I/O).
Multiprocessing
Multiple processes, each with its own Python interpreter. No shared GIL. True parallelism.
from concurrent.futures import ProcessPoolExecutor
def heavy_compute(x):
return sum(i * i for i in range(x))
with ProcessPoolExecutor(max_workers=8) as executor:
results = list(executor.map(heavy_compute, [10**6] * 10))
Use when:
- CPU-bound work (number crunching, image processing).
- You have multiple cores.
- Workers don't need shared state (or can share via queues).
Don't use when:
- I/O-bound (overhead exceeds benefit).
- Small per-task work (process startup cost dominates).
- Sharing large data is needed (serialization overhead).
Multiprocessing gotchas
- Each process has its own memory. No shared state by default. Use queues, pipes, or shared memory.
- Functions must be picklable. Lambdas, closures fail.
- Process startup is expensive (~50ms each). For many tiny tasks, threading wins.
asyncio
Single-threaded concurrency for I/O. Already covered.
The decision matrix
| Workload | Primitive |
|---|---|
| CPU-bound, multi-core | multiprocessing |
| I/O-bound, async-friendly | asyncio |
| I/O-bound, only sync libs available | threading |
| Mix CPU + I/O | combine (asyncio + asyncio.to_thread for sync I/O; multiprocessing for CPU) |
| GUI / interactive | asyncio (often) or threading |
| Tiny number of tasks (<10) | sync; concurrency overhead not worth it |
Combining primitives
Real systems often mix:
# I/O-bound (async): fetch many URLs concurrently
# CPU-bound (process pool): heavy parsing of results
import asyncio
from concurrent.futures import ProcessPoolExecutor
async def fetch_and_parse(url, pool, loop):
text = await fetch(url)
parsed = await loop.run_in_executor(pool, heavy_parse, text)
return parsed
async def main():
loop = asyncio.get_event_loop()
with ProcessPoolExecutor() as pool:
results = await asyncio.gather(*[
fetch_and_parse(url, pool, loop) for url in urls
])
asyncio for the I/O concurrency; process pool for the CPU work.
Sharing state
Threading (shared memory + locks)
import threading
counter = 0
lock = threading.Lock()
def increment():
global counter
with lock:
counter += 1
Easy to share; dangerous if you forget locks (race conditions).
Multiprocessing (queues / shared types)
from multiprocessing import Queue, Manager
queue = Queue()
def worker(q):
q.put("result")
# Or via Manager for dict/list:
with Manager() as mgr:
shared_dict = mgr.dict()
Higher overhead but no race conditions on shared memory (each process has its own).
asyncio (no real sharing needed)
Single thread; no race conditions for in-process state. For external coordination: locks, queues are async-aware versions.
Python 3.13 free-threaded mode
Python 3.13 (released Oct 2024) shipped experimental no-GIL build. In 3.13t and 3.14t:
- Threads run truly in parallel (CPU work scales).
- ~10-15% single-thread slowdown (cost of disabling GIL).
- Some C extensions need updates.
By 2027-2028: likely default. Today: experimental for adventurous teams.
Common concurrency mistakes
- Threading for CPU work. GIL prevents speedup; complexity for nothing.
- Multiprocessing for I/O. Process startup overhead exceeds benefit.
- Async without async libraries. Sync requests inside async = bottleneck.
- Race conditions in threading. Shared mutable state without locks.
- Pickling non-picklable. Lambdas in multiprocessing.
Concurrency overhead
For tiny tasks (~ms each), concurrency overhead dominates. Run sync.
For longer tasks:
- Threading: ~µs per task switch.
- asyncio: ~µs per task switch (lighter than threading).
- Multiprocessing: ~50ms process spawn (use a pool to amortize).
Takeaway
Three concurrency primitives. asyncio for I/O-bound when async libraries exist. Threading for I/O-bound with sync libraries (or shared state needs). Multiprocessing for CPU-bound on multiple cores. Mix when needed: asyncio + ProcessPoolExecutor for I/O + CPU. The GIL prevents threading from helping CPU work; Python 3.13 free-threaded mode is changing this experimentally.