Every data pipeline operates in one of three modes. The choice has profound implications for cost, complexity, and reliability.
Mode 1 — Batch
Data processed in scheduled chunks. Common schedules:
- Daily (midnight or pre-business-hours).
- Hourly (more responsive but more frequent infra).
- Every few minutes (approaching near-real-time).
[Source] → [Snapshot] → [Storage] → [Transform on schedule]
Pros:
- Simple to reason about. State = "what was true at last batch".
- Cheap. One compute job, one DB load.
- Easy debugging. Replay a single batch.
- Standard tools (Airflow, dbt) optimized for it.
Cons:
- Latency. Even an hourly batch is 30-60 min stale.
- Re-processing required for corrections.
90% of production data pipelines are batch. They should be.
Mode 2 — Streaming
Data processed event-by-event as it arrives.
[Event Source] → [Streaming Platform: Kafka/Pub-Sub] → [Stream Processor] → [Storage]
Pros:
- Low latency (sub-second possible).
- Reactive to events.
- Necessary for real-time features (fraud, recommendations, alerts).
Cons:
- Operational complexity 5-10x batch.
- Harder to debug (no "batch boundary" to look at).
- More expensive infrastructure.
- State management is tricky (exactly-once semantics, late events, out-of-order events).
- Reprocessing is harder.
10% of pipelines truly need streaming. Most don't.
Mode 3 — Hybrid (Lambda / Kappa architectures)
Run both batch and streaming. Streaming for real-time hot path; batch for historical re-processing.
Lambda:
[Source] → [Streaming] → [Speed layer] ↘
[Serving layer]
[Source] → [Batch] → [Batch layer] ↗
Used when:
- Real-time matters for current data.
- Historical accuracy matters more than streaming can guarantee.
- Different consumers have different needs.
Modern simplification: Kappa architecture treats batch as a special case of streaming (replay the stream from offset 0). Combined with frameworks like Materialize / RisingWave, you get a single pipeline that handles both.
Decision tree
Does anyone need to act on data within seconds?
├─ Yes → streaming
└─ No
├─ Within minutes? → frequent batch (5-15 min)
└─ Hourly/daily is fine → batch
Be honest. "We'd like real-time" usually means "we'd like fresh data, batch every 15 min is fine".
When streaming is genuinely required
- Fraud detection (block before transaction completes).
- Personalization on a live page (show user-specific content immediately).
- Alerting (system anomalies needing immediate response).
- Operational metrics on live production systems.
- Real-time bidding (ads).
- Stream processing of events too large to batch (millions/sec).
When batch is the right answer (usually)
- Dashboards (most stakeholders can tolerate 1-hour lag).
- Reports.
- Marketing analytics.
- LTV / retention analytics.
- ML model training data prep.
- Most reporting use cases.
If you're not sure, batch. You can add streaming later if needed.
Cost reality
For 1M events per day:
| Mode | Approximate cost | Latency |
|---|---|---|
| Daily batch | $5-20/day | 24h |
| Hourly batch | $30-100/day | 1h |
| Streaming (managed Kafka + processor) | $200-500/day | <1 min |
Streaming is 10-50x more expensive than daily batch for the same volume. Plus operational headcount cost (specialized skills).
Engineering complexity reality
For a typical pipeline:
- Batch: 1 engineer can build and maintain.
- Hourly batch: same.
- Streaming: usually 1-2 engineers full-time once at scale.
Streaming infra fails in interesting ways (out-of-order events, consumer lag, schema evolution mid-stream). The expertise required is real.
Hybrid pattern: batch + streaming SLI
A common middle ground:
- Batch processes data normally (cheap, simple).
- Streaming computes a small set of real-time signals (last-hour active users, error rates).
- Two pipelines, separate budgets.
The streaming pipeline is intentionally narrow — only metrics that genuinely need real-time. Everything else stays in batch.
Migration risk
Once batch, easy to stay batch. Once you've added streaming, hard to remove (consumers depend on the freshness).
Plan accordingly. Streaming is a one-way door — only add when sure.
Common mistakes
- "Real-time" requests that don't need real-time. Question every "we need it immediately" with "what action does it enable?"
- Lambda architecture for things batch alone could handle. Operational pain without benefit.
- Streaming for low-volume data. 1000 events/day doesn't need Kafka.
- Underestimating streaming operations cost. Includes monitoring, schema management, replay capabilities.
- Building streaming first. Batch first, prove value, then maybe add streaming.
Takeaway
Batch is the default. Streaming for genuine sub-minute needs. Hybrid when both matter. Start batch; add streaming reluctantly. The operational cost difference is large and persistent.