Every "Iceberg vs. Delta Lake" comparison eventually has to reckon with an inconvenient fact: the company behind Delta Lake now employs the people who created Iceberg.
In June 2024, Databricks paid over a billion dollars for Tabular, the company Ryan Blue and Daniel Weeks founded after building Iceberg at Netflix — and the stated goal wasn't to kill Iceberg, it was interoperability. That single acquisition tells you more about where this format war actually stands in 2026 than most feature-by-feature breakdowns do.
Both Apache Iceberg and Delta Lake solve the same underlying problem: give a data lake full of Parquet files ACID transactions, schema evolution, and time travel, without forcing every reader and writer through one vendor's proprietary engine. They solve it with genuinely different internal designs — Iceberg tracks table state through a metadata tree of snapshots and manifests, Delta writes a sequential, checkpointed transaction log — and that difference still shapes how each one partitions data, handles row-level deletes, and scales under heavy ingestion.
But the question that used to define this comparison, "which format is technically better," has been substantially overtaken by a different one: which catalog, and which engines, does your organization actually need first-class access from. This piece covers both — the real architectural differences that still matter operationally, and the governance and interoperability shifts that changed what "choosing a format" even means going into 2027.
Prefer to skip the debugging and have an expert handle this? Book a 30-min diagnostic →
Metadata Tree vs. Transaction Log: The Split That Explains Everything Else
Apache Iceberg's own spec describes a table as a chain of atomic swaps: every change writes a new metadata.json file and replaces the old one in a single atomic operation. That metadata file points to a snapshot, the snapshot points to a manifest list carrying partition stats and file counts, and the manifest list points to one or more manifests, each holding a row per data file with that file's partition data and column-level metrics. No single engine owns this structure; it's a spec any compliant reader implements the same way.
Delta Lake takes a different physical approach to the same goal. Per Delta's own protocol spec, every write becomes a single-line JSON commit appended to a `_delta_log` directory, sequentially numbered, periodically compacted into Parquet checkpoints for fast replay. Delta layers on two independent versioning dimensions: protocol versions that set minimum reader/writer capability, and, for newer tables, named table features that gate specific capabilities like deletion vectors or column mapping.
Both models deliver serializable-isolation optimistic concurrency and full time travel. The practical difference shows up at the edges: Iceberg's spec-first design means no engine is privileged over another by construction, while Delta's log-and-features model means an engine on an older protocol version can be locked out of a table the moment a newer table feature gets turned on — not a bug, just a direct consequence of a protocol built to be extended quickly by whoever controls it.
Partitioning: Iceberg Hides It, Delta Replaced It With Liquid Clustering
Iceberg's answer to Hive's oldest partitioning problem is hidden partitioning: per Iceberg's own docs, the engine "produces partition values by taking a column value and optionally transforming it" and keeps track of that relationship itself, so a query never needs to know or filter on a physical partition column the way Hive requires. Hive can't validate a partition value at write time, so a small formatting mistake produces silently wrong query results instead of an error; Iceberg removes that failure mode by construction. When a table's partitioning scheme needs to change, Iceberg's evolution docs are explicit that old data stays in its old layout and new data lands in the new one, with both layouts planned separately in the same table as a pure metadata operation, no rewrite required.
Delta's current answer isn't partitioning at all — it's Liquid Clustering, which Databricks' own docs describe as a technique that "replaces table partitioning and ZORDER," letting clustering keys be redefined without rewriting existing data. It's a genuinely different mechanism aimed at the same underlying pain point: query patterns change, and a data layout chosen a year ago shouldn't require an expensive migration to fix. Liquid Clustering has real, stated limits worth knowing before you commit to it — a maximum of four clustering keys, no complex types like structs, maps, or arrays, and it's explicitly incompatible with traditional partitioning or ZORDER on the same table.
Functionally, Iceberg's partition evolution and Delta's Liquid Clustering are solving the identical problem from opposite directions: one hides and evolves partition metadata, the other replaces partitioning with a clustering key that can move.
Row-Level Deletes: Both Sides Converged on Merge-on-Read
Update-heavy and delete-heavy workloads were the weak spot of early table formats, and both projects solved it the same conceptual way: mark a row as deleted without rewriting the entire file it lives in, then resolve the deletion at read time. Iceberg's v2 spec introduced delete files that encode which rows in an existing data file are gone; v3 — now a complete, adopted spec version — replaced that with more compact binary deletion vectors. Delta Lake's deletion vectors, per Delta's own 2023 announcement and Databricks' current docs, encode deleted-row positions in RoaringBitmap, a highly compressed bitmap format, applied at query time so `DELETE`, `UPDATE`, and `MERGE` operations skip rewriting the underlying Parquet file entirely.
Both formats treat this as a temporary state, not a permanent one: Databricks' own docs are explicit that ordinary file compaction "don't have strict guarantees for resolving changes recorded in deletion vectors" — you need an explicit `OPTIMIZE` or `REORG ... APPLY (PURGE)` to actually rewrite the base files and clear the deletion vectors out. Skip that maintenance step on a delete-heavy table on either format, and you accumulate write-amplification debt that eventually shows up as degraded read performance, quietly, with no error to flag it.
Still Piecing This Together Yourself?
A senior engineer looks at your actual setup, not a generic checklist, and tells you exactly what's wrong and how to fix it.
Book a Diagnostic CallThe Maintenance You Can't Skip — And What Happens If You Do
Both projects' own documentation carries surprisingly blunt warnings about maintenance, worth reading before assuming either format is low-touch. Iceberg has three named operations: `rewriteDataFiles` (compaction, because small files cause "an unnecessary amount of metadata and less efficient queries from file open costs"), `expireSnapshots` (recommended to keep metadata size bounded, since data files persist as long as any live snapshot references them), and `deleteOrphanFiles`, which carries the sharpest warning in Iceberg's own docs: removing orphan files with too short a retention window "might corrupt the table if in-progress files are considered orphaned and are deleted," which is why the default retention sits at three days. Delta's equivalent is `OPTIMIZE` for compaction and `VACUUM` for retention cleanup, defaulting to seven days, with an almost identically worded warning — a shorter retention risks concurrent readers failing or, worse, the table being corrupted if `VACUUM` deletes files that haven't been committed yet, and it permanently forfeits time travel to anything older than the new retention window.
The risk this maintenance actually guards against is real and Iceberg-specific in one important way: metadata itself can become the bottleneck. An illustrative case published by e6data — not a named customer incident, worth treating as an architecturally plausible scenario rather than a verified postmortem — describes a 50 TB table built from 45 million small files under streaming ingestion, where a 5 TB metadata footprint across 200,000 manifests caused Spark and Trino query planners to crash with out-of-memory errors while trying to read stats for 22 million files during planning. The root cause tracks directly back to Iceberg's manifest-based planning model: cost scales with file and manifest count, not raw data volume, so any streaming or CDC-heavy Iceberg table needs a compaction cadence tight enough to keep that count bounded before problems appear, not after.
The Catalog Layer Is Where the Real Decision Lives Now
This is the section where "Iceberg vs. Delta" stops being a pure format question. Both projects solved the same interoperability problem from opposite directions, and both solutions are now shipping, GA infrastructure rather than roadmap promises.
On the Iceberg side, the REST Catalog spec puts catalog logic behind a common, OpenAPI-defined HTTP API, so any engine implements one client and gets access to any compliant catalog — Apache Polaris, Project Nessie, AWS Glue, Unity Catalog, and others all speak it. Snowflake built Polaris and donated it to the Apache Software Foundation in August 2024; it graduated to a full top-level ASF project on February 18, 2026, putting Iceberg's most vendor-neutral catalog under the same governance model as Iceberg itself. On the Delta side, Delta Universal Format — UniForm — asynchronously generates Iceberg (and separately Hudi) metadata from ordinary Delta commits, using the same compute that ran the transaction, with Delta's own docs stating "negligible Delta write overhead." It requires column mapping enabled, minimum reader/writer protocol versions, and Delta Lake 3.1 or newer, but where those prerequisites are met, a Delta table becomes directly readable by an Iceberg client without a migration.
Meanwhile every major cloud has shipped native, first-party Iceberg support as of this writing: Snowflake's Iceberg Tables (Snowflake-managed or externally-managed via a REST catalog), AWS S3 Tables (a dedicated managed Iceberg storage layer AWS claims delivers up to 3x faster queries and 10x higher transactions per second than self-managed Iceberg on standard S3 — a vendor number with no independently published methodology found, worth treating as a claim rather than a verified benchmark), and Google BigQuery's managed Iceberg tables. Delta's equivalent first-party distribution outside Databricks itself remains comparatively thinner. That asymmetry, not any remaining spec-level feature gap, is the more decisive input to a real format decision today.
The Tabular Acquisition: What Actually Changed
On June 4, 2024, Databricks announced it was acquiring Tabular — the company founded by Iceberg's original creators — for a reported $1 billion-plus. Databricks' own blog post framed the rationale directly: "This acquisition brings the original creators of Apache Iceberg and those of Linux Foundation Delta Lake, the two leading open source lakehouse formats, together," with the stated long-term goal of "evolving toward a single, open, and common standard of interoperability." One day later, at the same event, Databricks open-sourced Unity Catalog under Apache 2.0 through the Linux Foundation's LF AI & Data, explicitly built to support "Delta Lake, Apache Iceberg via UniForm... and all the formats out there," implementing both the Iceberg REST Catalog API and the Hive Metastore API in the same catalog.
Read together, these two moves are the clearest signal available that the era of a genuine, permanent Iceberg-versus-Delta rivalry is largely over at the format layer. What it doesn't resolve is governance concentration: Databricks now controls Delta's primary commercial distribution and employs the engineers who built Iceberg's reference implementation, even though Iceberg itself remains governed independently by the Apache Software Foundation. Snowflake's parallel move — donating Polaris to the ASF rather than building a Databricks-equivalent acquisition — reads as a direct hedge against exactly that concentration.
Neither company's public statements change the underlying open-source licensing of either format; Apache 2.0 for Iceberg and the Linux Foundation's governance model for Delta are the durable guarantees to watch, not the acquisition headlines themselves.
Get a Straight Answer, Not a Sales Pitch
Tell us what you're running into. We'll tell you directly what's causing it and what it takes to fix, before you sign anything.
Contact UsSo Which One Should You Actually Choose?
Start from your actual engine mix, not an abstract feature checklist. If your organization is Databricks-centric and Spark is the dominant write path, Delta remains the simplest choice — it's the most mature, most deeply integrated option in that specific environment, and UniForm gives you a credible path to multi-engine read access later without a migration now. If your estate spans multiple engines and clouds — Snowflake plus Trino plus BigQuery plus Spark, none of which should have privileged access to the data — Iceberg via a REST catalog is the more defensible default, precisely because no single vendor's engine is structurally favored by the spec.
If you're already running Iceberg and ingesting from streaming or CDC sources, treat metadata scaling as a first-class operational concern from day one, not something to discover after a planner starts throwing out-of-memory errors. And if you're inheriting an existing deployment rather than choosing greenfield — which is the more common starting point for most engagements — the higher-value work usually isn't a format migration at all. It's auditing whether compaction cadence, snapshot expiration, and VACUUM retention were ever tuned past the defaults, because both projects' own documentation is unambiguous that skipping that maintenance is where real production incidents come from, regardless of which format sits underneath.
Related: Data Engineering
Iceberg and Delta Lake solve the same problem with genuinely different internal architectures, and that difference — a metadata tree of atomic snapshots versus a sequential, checkpointed transaction log — still shapes how each one partitions, deletes, and scales under load. But the industry-defining event of the past two years wasn't a spec update on either side; it was Databricks buying the company founded by Iceberg's own creators and open-sourcing a catalog built to speak both formats' languages in the same breath. That leaves most organizations choosing based on their actual engine and catalog footprint rather than a permanent technical verdict, and it leaves the operational discipline — compaction cadence, retention windows, metadata growth under streaming ingestion — as the thing that actually determines whether either format holds up in production.
Pick the format your engine mix already favors. Then treat the maintenance both projects' own docs insist on as non-optional, because that's where the real failures happen, not in the choice between the two.
Frequently Asked Questions
Is Apache Iceberg better than Delta Lake?
Neither is strictly better; they're built on different internal models solving the same problem. Iceberg's metadata-tree, spec-first design makes it the stronger fit for multi-engine, multi-cloud environments where no single vendor's engine should have privileged access. Delta Lake's transaction-log model is more deeply integrated for Databricks-centric, Spark-dominant environments. As of 2026, every major cloud ships native Iceberg support, while Delta's first-party distribution outside Databricks remains comparatively thinner.
Can Delta Lake tables be read by Iceberg clients?
Yes, via Delta Universal Format (UniForm), which asynchronously generates Iceberg (and separately Hudi) metadata from ordinary Delta commits with, per Delta's own docs, negligible write overhead. It requires column mapping enabled, minimum reader/writer protocol versions, and Delta Lake 3.1 or newer — worth verifying against an existing table's actual configuration before assuming it's ready.
Why did Databricks buy Tabular if Tabular built Iceberg?
Databricks' own stated rationale was interoperability, not elimination: bringing Iceberg's original creators together with Delta Lake's stewards to work toward, in Databricks' words, "a single, open, and common standard of interoperability." The concrete near-term product of that acquisition is Delta UniForm and Unity Catalog's native Iceberg REST Catalog API support, both shipping today, not roadmap promises.
What is the Iceberg REST Catalog and why does it matter?
It's an OpenAPI-defined HTTP API that puts catalog logic — table discovery, commits, credential vending — behind a common interface any engine can implement once and use against any compliant catalog (Apache Polaris, Project Nessie, AWS Glue, Unity Catalog). It's arguably more consequential for a buyer than the Iceberg table-format spec itself, since it determines whether a table written by one engine's client can actually be discovered and safely written to by a different engine.
How do Iceberg and Delta Lake handle row-level deletes?
Both use merge-on-read: a row is marked deleted in metadata without rewriting its data file, resolved at query time. Iceberg's v3 spec uses binary deletion vectors; Delta Lake uses deletion vectors encoded in RoaringBitmap, a compressed bitmap format. Both require an explicit compaction step (Iceberg's rewriteDataFiles, Delta's OPTIMIZE or REORG with PURGE) to eventually rewrite the base files — skipping it accumulates write-amplification debt on either format.
Why do Iceberg query planners run out of memory on large tables?
Planning cost in Iceberg scales with file and manifest count, not raw data volume, because the planner reads manifest entries to prune files before executing a query. A table with tens of millions of small files from high-frequency streaming or CDC ingestion can produce a metadata footprint large enough to exhaust planner memory. The fix is a compaction cadence sized to the ingestion rate, keeping manifest and file counts bounded before the problem appears.
Is it safe to lower the VACUUM retention period on a Delta table?
Not below the 7-day default without a specific reason. Delta's own docs warn that a shorter retention risks concurrent readers failing or the table being corrupted if VACUUM deletes files that haven't been committed yet, and it permanently forfeits the ability to time travel to any version older than the new window. Iceberg's equivalent operation, deleteOrphanFiles, carries the same category of warning with a 3-day default.
Who can help me choose between Iceberg and Delta Lake for my lakehouse?
DharmOps scopes this under Data Engineering by actual engine and catalog footprint, the same framework this guide uses, rather than treating it as a permanent technical verdict between the two formats.
Can someone help migrate an existing Delta Lake table to Iceberg, or vice versa?
Often the better first move isn't a migration at all — Delta UniForm and Iceberg REST catalogs, both covered in this guide, can give multi-engine read access without one. Worth confirming that's not sufficient before scoping a full migration.
How much does a table format migration cost?
It depends far more on compaction, snapshot-retention, and metadata-scaling maintenance debt than on data volume — the operational gaps this guide flags as where real production incidents actually come from. That maintenance audit is the right starting point before pricing a migration.
Get a Straight Answer on Your Setup
Tell us what you're running into. We read every message personally and reply within 24 hours with times for a free call.

