Data engineering in 2026 is shaped by the modern data stack — a set of layered, mostly-cloud-native tools that move data from sources to BI consumption. Knowing the layers is how you talk shop with other DEs and reason about where work fits.
The 5 layers
[ Source systems ] ← OLTP DBs, SaaS apps, files, events
↓ (extract)
[ Ingestion layer ] ← Fivetran, Airbyte, Meltano, custom
↓
[ Storage (raw layer) ] ← Snowflake / BigQuery / S3+Iceberg / Lakehouse
↓ (transform)
[ Transformation layer ] ← dbt, SQLMesh, Coalesce
↓
[ Storage (mart/serving) ] ← Same warehouse / Lakehouse
↓
[ Consumption layer ] ← BI tools, ML, APIs, reverse-ETL
You'll often hear "EL + T" stack: Extract & Load (ingestion) + Transform (dbt).
Layer 1 — Source systems
Where data is born:
- OLTP databases (Postgres, MySQL): app data.
- SaaS apps (Stripe, Shopify, HubSpot, Salesforce): business data.
- Files (S3, GCS): logs, exports, third-party feeds.
- Events (Kafka, Kinesis): clickstream, IoT.
Each has different access patterns. OLTP via JDBC/CDC. SaaS via REST APIs. Files via cloud storage. Events via streaming consumers.
Layer 2 — Ingestion
Moves data from sources into your storage. Three approaches:
Managed connectors (Fivetran, Airbyte Cloud)
Click-config tools. Cover 100s of common sources. You pay for ease.
Open-source connectors (Airbyte OSS, Meltano)
Same idea, self-hosted. Cheaper but you operate it.
Custom pipelines
For unusual sources or specific control. Python scripts + orchestration.
Module 2 covers ingestion deeply.
Layer 3 — Raw storage
Data lands here untransformed. Three options:
Cloud data warehouses
- Snowflake — most popular for enterprise.
- BigQuery (GCP) — serverless, scales massively.
- Redshift (AWS) — solid, AWS-native.
- Databricks — Spark + lakehouse.
Pros: easy SQL access, optimized for analytics, columnar. Cons: per-query/storage costs, vendor lock-in to a degree.
Lakehouse (open format on cloud storage)
- Iceberg / Delta / Hudi tables stored as Parquet files on S3/GCS/Azure.
- Query with Trino, Spark, Snowflake (Iceberg external), Databricks.
Pros: open formats, no lock-in, separates storage from compute. Cons: more moving parts, requires more engineering.
In 2026, the lakehouse pattern is dominant for new builds. Existing teams often stay on warehouse for inertia.
Pure data lake (S3 + Parquet, no table format)
Older approach. Limited transactional guarantees. Mostly being replaced by lakehouse.
Layer 4 — Transformation
Take raw, produce business-ready models. The dbt layer.
Industry standard: dbt (covered in dbt Fundamentals course). Models source data into staging, intermediate, marts.
Alternatives:
- SQLMesh — newer, slightly different philosophy.
- Coalesce — visual, enterprise-focused.
- Custom SQL/Python — for cases dbt doesn't fit.
Layer 5 — Consumption
How data leaves the warehouse:
BI tools
- Looker, Tableau, Power BI, Metabase, Hex.
- Query the mart layer directly.
ML training
- Pull training data from marts.
- Tools: Snowflake → Databricks, BigQuery → Vertex AI.
Reverse-ETL
- Push transformed data BACK to operational systems (Salesforce, HubSpot).
- Tools: Hightouch, Census.
Direct APIs
- App reads from a "serving" mart for personalization, dashboards, etc.
Where data engineering lives
Data engineering = layers 2-4 (ingestion, storage, transformation) + the reliability/observability layer that spans all of them.
Specifically:
- Designing ingestion pipelines.
- Operating orchestration (Airflow, Dagster, Prefect).
- Modeling the warehouse (covered in Data Modeling course).
- Writing dbt (covered in dbt Fundamentals).
- Streaming pipelines where needed.
- Ensuring data quality and reliability.
- Cost management.
Analytics engineers focus on layer 4 (dbt models). Pure data engineers focus on layers 2-3 + reliability. The role spectrum overlaps.
What's NEW in 2026
A few shifts from 2020:
- Lakehouse becoming default for new builds. Iceberg + Snowflake/Trino is mainstream.
- Streaming is more accessible. Kafka, Pub/Sub, Kinesis combined with managed stream processors (Materialize, RisingWave) make CDC + near-real-time easier.
- Better orchestrators. Dagster and Prefect have matured; Airflow remains dominant but no longer the only option.
- Lineage and observability standardized. OpenLineage + Marquez + commercial offerings (Monte Carlo, Datafold).
- dbt's dominance for transformation. It's the default; alternatives must justify themselves.
What this course covers
Module 1: the landscape (this lesson is part of it). Module 2: ingestion + orchestration (the EL of ELT). Module 3: reliability, streaming, Spark. Module 4: end-to-end case walkthroughs.
By the end, you'll be able to design and build a production data pipeline from source to serving.
Common landscape mistakes
- Skipping the lakehouse question. New builds in 2026 should consider Iceberg/Delta seriously.
- Picking tools before defining requirements. Stack is a means, not an end.
- Treating data engineering as just "running scripts". It's a discipline with reliability, quality, lineage concerns.
- Ignoring reverse-ETL. Data only delivers value when consumed; reverse-ETL puts it back where teams use it.
Takeaway
5 layers: source → ingestion → raw storage → transformation → consumption. Modern stack tools fill each. Data engineering owns layers 2-4 plus reliability across all. Lakehouse becoming default for new builds; warehouse still common for inherited stacks. The rest of this course goes deep into each layer.