Modern lakehouses are built on a stack of open formats. Knowing the layers — and why they exist — is essential for greenfield 2026 data engineering decisions.
The stack
[ Query Engine: Trino, Snowflake, Spark, Athena ]
↓
[ Table Format: Iceberg, Delta Lake, Hudi ] ← Adds ACID, schema evolution, time travel
↓
[ File Format: Parquet (mostly), ORC, Avro ] ← Columnar, compressed
↓
[ Storage: S3, GCS, Azure Blob, HDFS ] ← Cheap object storage
Each layer has a job.
File format — Parquet
A columnar file format with compression. Single file containing rows of data, stored column-by-column with per-column compression.
Why Parquet dominates:
- Columnar — efficient for analytics queries (read only needed columns).
- Compression — typically 5-10x smaller than CSV.
- Schema embedded — self-describing.
- Open format — works with every modern query engine.
Other file formats:
- ORC — similar to Parquet. Hive-era. Less common in modern stacks.
- Avro — row-oriented. Used for streaming/CDC where row-by-row writes matter.
For 99% of analytical workloads: Parquet.
What's missing from raw Parquet
A directory of Parquet files isn't a table:
- No ACID transactions (concurrent writes corrupt).
- No schema evolution (changing a column breaks readers).
- No time travel ("show me data as of yesterday").
- No DELETE/UPDATE (immutable files; you'd rewrite everything).
- No partition management.
For production, you need a table format layered on top.
Table format — Iceberg / Delta Lake / Hudi
Three competing open table formats:
Apache Iceberg
- Originated at Netflix.
- Now the de facto open standard (2026).
- Supported by Snowflake, BigQuery, Trino, Spark, Athena, Databricks (interop).
- Used at Apple, LinkedIn, Adobe.
- Maintained by Apache Foundation.
Delta Lake
- Originated at Databricks.
- Open-sourced 2019; widely adopted, especially in Databricks ecosystem.
- Snowflake added external Delta support in 2024.
- Strong choice if you're already on Databricks.
Apache Hudi
- Originated at Uber.
- Best-in-class for streaming and frequent updates.
- Less common than Iceberg/Delta in new builds.
In 2026: Iceberg has emerged as the cross-vendor standard. Snowflake, BigQuery, AWS, Cloudflare all support it as external table format. Delta is still strong in Databricks-centric stacks.
What table formats add
ACID transactions
Multiple writers can update simultaneously. No corrupted reads. Atomic schema changes.
Schema evolution
Add columns, rename columns, change types — without rewriting the table.
ALTER TABLE customers ADD COLUMN segment STRING;
-- New column appears; old data has NULL; readers handle gracefully.
Time travel
Query as of any past point.
SELECT * FROM customers FOR SYSTEM_VERSION AS OF '2026-04-01';
-- Or
SELECT * FROM customers VERSION AS OF 1234;
Useful for audits, debugging, recovery from bad transformations.
Partition management
Logical partitions handled by the table format, not by you. Re-partition without rewriting data.
Updates and deletes
Delete or update rows efficiently (using copy-on-write or merge-on-read strategies).
DELETE FROM customers WHERE created_at < '2024-01-01';
UPDATE customers SET status = 'inactive' WHERE last_login < '2025-01-01';
Without a table format, these operations require rewriting the entire dataset.
The lakehouse architecture
[ Apps generate data ]
↓
[ Iceberg/Delta tables on S3, written in Parquet ]
↓
[ Multiple engines query the same tables ]
- Snowflake (external Iceberg tables)
- Trino / Athena / Spark
- Databricks (Delta or Iceberg via Unity Catalog)
The same data, accessible to many engines without copying. One table, many readers. Storage and compute completely decoupled.
This is the lakehouse promise: open formats + cheap storage + compute flexibility.
When lakehouse over pure warehouse
Lakehouse advantages:
- Less vendor lock-in (data on S3, queryable from multiple engines).
- Cheaper storage (S3 vs warehouse-internal storage).
- Multiple compute engines for different use cases.
- Open formats = future-proof.
Lakehouse disadvantages:
- More moving parts to operate.
- Catalog management (Unity Catalog, Tabular, AWS Glue, project Nessie).
- Optimization tuning (compaction, partitioning).
- Less mature tooling vs Snowflake/BigQuery for some use cases.
For a startup picking a stack today: Iceberg-on-S3 + dbt + Snowflake (querying Iceberg) is increasingly the modern choice. Existing Snowflake-native installations stay native.
Common lakehouse mistakes
- Raw Parquet without a table format. Production needs ACID; raw Parquet doesn't have it.
- Mixing Iceberg, Delta, Hudi without standardization. Pick one per project.
- Underestimating catalog operations. Iceberg needs a catalog (Glue, Nessie, Tabular). Don't skip this.
- No compaction strategy. Lakehouses accumulate small files; need periodic compaction.
- Treating it like a warehouse. Lakehouse is more DIY; budget more engineering time.
Catalog services
A catalog tracks which Iceberg/Delta tables exist and where. Options:
- AWS Glue Data Catalog — if you're on AWS.
- Apache Hive Metastore — classic; still common.
- Project Nessie — Git-style versioning for tables.
- Tabular — managed Iceberg service.
- Snowflake Polaris — Snowflake's catalog, open-sourced.
Pick one early; switching is painful.
Takeaway
File format (Parquet) holds the bytes. Table format (Iceberg/Delta) adds ACID + schema evolution + time travel + UPDATE/DELETE. Together with cheap object storage = lakehouse. Iceberg is the 2026 cross-vendor standard. For new builds, lakehouse with Iceberg is a strong default; Snowflake/BigQuery still excellent for teams wanting more managed simplicity.