Iceberg gives every query engine the same transactional view of the same files on object storage. The format decision is settled; the catalog, compaction, and engine decisions that follow it decide whether an enterprise AI program stays open.
By

Billy Allocca

Table of Contents
Why Apache Iceberg Is Becoming the Default Storage Format for Enterprise AI, and What to Build Around It
Apache Iceberg is an open table format that adds ACID transactions, schema and partition evolution, and snapshot-based time travel to Parquet files on object storage, with the layout defined by a public specification. Any compliant engine (Trino, Spark, Flink, DuckDB) reads the same tables, so an enterprise can change compute without moving or rewriting its data.
The format debate closed within two years. Databricks acquired Tabular, the company founded by Iceberg's creators, on June 4, 2024, for a price it confirmed was above $1 billion [1][2]. Snowflake open-sourced its Polaris Catalog under the Apache 2.0 license on July 30, 2024 [3]. AWS shipped Amazon S3 Tables, Iceberg support built into the object store itself, on December 3, 2024 [4]. By May 2026, Databricks had declared Iceberg v3 and managed Iceberg tables generally available in Unity Catalog [5], and Google had previewed a serverless Iceberg REST catalog so Spark, Flink, and Trino could work on BigQuery data without copies [6].
AI workloads made the consolidation urgent. In Dremio's Q4 2024 survey of 563 IT decision-makers, 85% said they use a data lakehouse for AI model development [7], and training, feature computation, and agent retrieval each read the same tables through different engines. The same argument applied to compute is in the earlier post on why composability is required at every layer.
What Does Apache Iceberg Provide That a Parquet Data Lake Does Not?
Iceberg turns a directory of files into a table with database guarantees, entirely through metadata. An open table format is a specification for the metadata that tracks which data files belong to a table, what schema they carry, and how they are partitioned, so reads and writes behave like operations on a database table. The Iceberg specification states that reads always use a committed snapshot, the complete set of data files that made up the table at one point in time, and that writers coordinate through optimistic concurrency without readers taking locks [8]. Version 2 added row-level deletes; version 3 added binary deletion vectors, a variant type for semi-structured data, row lineage, and default column values [8], and AWS made the v3 deletion vectors and row lineage available across EMR, Glue, and S3 Tables on November 26, 2025 [9].
Capability | What Iceberg does | What it means for AI workloads |
|---|---|---|
ACID transactions on object storage | Atomic snapshot commits; readers never see a half-written table [8] | A training job and a streaming upsert can touch the same table at once without producing a corrupt feature set |
Schema evolution | Add, drop, rename, widen, and reorder columns as metadata-only changes with no file rewrite [10] | New features are added to a training table without a backfill; models that never asked for the column are unaffected |
Time travel | Every write creates a snapshot that queries can target by ID or timestamp [8] | A model can be reproduced against the exact table state it was trained on, which is the provenance evidence an audit asks for |
Hidden partitioning and pruning | Partition values are derived from column transforms; planners skip files that cannot match a filter [11] | Retrieval queries scan only the files they need; Hive-style tables return silently wrong results when a writer uses the wrong date format [11] |
Partition evolution | Change the partition spec; old files keep the old layout and queries plan across both [10] | A daily-partitioned table moves to hourly partitions as agent traffic grows, with no rewrite of history |
Hidden partitioning means the table itself owns the mapping between a column and its physical partition, so a filter on event_time prunes correctly without anyone knowing the table is partitioned by day(event_time).
Why Use Apache Iceberg as the Storage Layer for Enterprise AI?
Iceberg is the interface contract that makes the compute layer swappable, and for an AI program that property outweighs any single performance feature. Once tables are Iceberg, Spark writes them, Trino serves interactive queries, Flink streams into them, and DuckDB reads them on a laptop, all through the same specification. DuckDB added insert, update, and delete support for Iceberg tables through REST catalogs in November 2025 [12], and Trino's connector supports six catalog types, including REST and Nessie [13].
The origin explains the neutrality. Ryan Blue and Daniel Weeks built Iceberg at Netflix to replace Hive tables that failed under concurrent writes, then donated it to the Apache Software Foundation [1], where it graduated to a top-level project in May 2020 [14]. The specification was written for engines Netflix did not control, which is why competitors could adopt it.
Date | Event | What it settled |
|---|---|---|
June 4, 2024 | Databricks acquires Tabular for more than $1 billion [1][2] | The Delta Lake vendor commits to converging with Iceberg |
July 30, 2024 | Snowflake open-sources Polaris Catalog [3] | The catalog, previously the closed piece, gets an Apache-licensed implementation |
December 3, 2024 | AWS launches Amazon S3 Tables [4] | Iceberg becomes a primitive of the object store, with built-in compaction |
May 28, 2026 | Databricks: Iceberg v3 and managed Iceberg GA in Unity Catalog [5] | The largest cloud-era platform treats Iceberg as a first-class managed format |
For AI, the payoff is one governed copy of the data that every engine and every agent reads. One copy per platform in a proprietary format is the pattern the Hadoop modernization post identifies as the cause of multi-year timelines.
How Does Iceberg Compare to Delta Lake and Apache Hudi for the Enterprise?
All three formats deliver transactions on object storage, so the enterprise distinction is who governs the specification and whether the catalog has an open protocol, and there Iceberg has the widest neutral footprint. Delta Lake is hosted by the Linux Foundation and describes itself as independent of any single company [15], though Databricks leads its protocol, and its UniForm feature exposes Delta tables to Iceberg and Hudi readers [15]. Apache Hudi, which grew out of Uber's incremental processing work beginning in 2016, describes itself as a lakehouse platform with a table format built for high-performance writes on incremental pipelines [16][17].
Dimension | Apache Iceberg | Delta Lake | Apache Hudi |
|---|---|---|---|
Governance model | Apache Software Foundation top-level project since May 2020 [14] | Linux Foundation project; protocol led by Databricks [15] | Apache Software Foundation top-level project [16] |
Catalog openness | Public REST catalog specification implemented by Polaris, Nessie, Gravitino, Glue, Unity Catalog, and BigLake [18] | Catalog is engine-side; UniForm publishes Iceberg metadata for Iceberg catalogs to read [15] | Syncs metadata to Hive Metastore and AWS Glue; no independent catalog protocol [16] |
Engine support | Spark, Trino, Flink, DuckDB, Snowflake, BigQuery, Databricks, Dremio, Athena, EMR | Deepest on Databricks and Spark; read paths elsewhere | Deepest on Spark and Flink; read paths in Trino and Presto |
Adoption | Netflix, Apple, Airbnb, Adobe, Expedia; native in AWS, Google Cloud, Snowflake, Databricks [14] | Databricks reports 10,000+ production environments [15] | Uber, Amazon, ByteDance, Robinhood [16] |
Best fit | Multi-engine estates that want format and catalog independent of any compute vendor | Estates standardized on Databricks that want Iceberg readability | Streaming upsert-heavy pipelines |
Merge-on-read is a write strategy that stores deletes and updates as separate files that are reconciled at query time, which favors write speed; copy-on-write rewrites affected data files at write time, which favors reads. Both Hudi and Iceberg v2 support the merge-on-read pattern. The convergence Databricks described in June 2024, two formats that share Parquet yet became incompatible through independent development [1], is now moving toward Iceberg: v3 deletion vectors and row lineage use encodings compatible with Delta Lake's, and Databricks has proposed that Delta 5.0 adopt the adaptive metadata tree planned for Iceberg v4 [5].
What Are the Three Decisions That Follow an Iceberg Commitment?
Choosing Iceberg removes one decision and creates three, and the catalog decision determines whether the openness survives. A catalog is the service that maps a table name to its current metadata file; because every commit goes through it, whoever runs the catalog controls access, credential vending, and which engines may write. The Iceberg REST catalog is the specification's HTTP protocol for that service, with endpoints for namespaces, tables, views, and vended credentials, so a client written against it can point at Polaris, Nessie, Gravitino, Glue, or Unity Catalog without code changes [18].
Compaction is the rewriting of many small data files into fewer large ones. Iceberg tracks every file in manifests, so small files inflate metadata and slow queries through per-file open costs [19]. AWS measured 8.5 times more read operations on 1 MB files than on 512 MB objects, and automatic compaction in S3 Tables sped up eight benchmark queries by 1.12 to 3.2 times [20].
Decision | Options | Trade-offs |
|---|---|---|
Catalog | Apache Polaris (1.0 released July 10, 2025, with RBAC and credential vending) [21]; Project Nessie (Git-like branches and cross-table transactions) [22]; Apache Gravitino (federated metadata lake whose Iceberg REST service proxies Hive, JDBC, and other REST catalogs) [23]; AWS Glue; Unity Catalog; BigLake | Self-hosted Apache catalogs keep the control point inside your estate and run identically on-prem and in any cloud. Vendor catalogs remove operations work but tie access policy and credential vending to one provider's identity system |
Compaction strategy | Scheduled Spark | Self-run compaction gives control over file size, sort order, and cost windows but needs an owner and a schedule. Managed compaction only runs inside the vendor's storage or catalog, so tables outside it go unmaintained |
Query engine | Trino for interactive federation; Spark for training-set preparation; Flink for streaming ingestion; DuckDB for local reads; Apache Kyuubi as a multi-tenant SQL gateway over Spark | Several engines can run against the same tables. Identity and policy then have to be enforced above the engines, or each engine becomes its own governance island |
Snapshot expiration and time travel interact: the retention window set in expire_snapshots is also the reproducibility window for any model trained on that table [19]. The engine question continues in the guide on separating storage, compute, and serving layers, and catalog ownership in the CDO's guide to a vendor-neutral data estate.
Why Not Just Use a Cloud Vendor's Managed Iceberg Tables?
For a single-cloud estate with one dominant engine, managed Iceberg from AWS, Databricks, Google, or Snowflake is a reasonable start, because someone else runs compaction and the files stay readable by any compliant engine. What the arrangement costs is the control point. Databricks' managed Iceberg tables are governed by Unity Catalog and maintained by Predictive Optimization [5]; S3 Tables enforce table-level permissions through AWS IAM [4]. Both are sound designs, and both mean the access policy for a table lives in one vendor's identity system, an agent running in another cloud authenticates through that vendor, and a deployment in a data center you own cannot use the same catalog. The files stay open while the governance becomes vendor-specific.
The thesis holds because the REST protocol exists. You can adopt a managed service today and keep an exit by requiring that every reader and writer speak the Iceberg REST catalog API, so moving the catalog to Polaris, Nessie, or Gravitino later is a configuration change [18].
NexusOne Reads Iceberg Natively Across Every Engine and Environment
NexusOne treats Iceberg as the default table format and runs the three follow-on decisions as one configured layer. Tables created through the platform register automatically in an Apache Gravitino catalog that exposes the Iceberg REST protocol, and Gravitino federates external catalogs (Hive Metastore, Glue, other REST endpoints, plus Databricks and Teradata metadata) into one namespace, so Trino, Apache Spark, and Apache Kyuubi see the same tables whether they sit in an on-prem object store or a cloud bucket [23]. Identity arrives through one Keycloak layer federated from Active Directory, Okta, LDAP, or SAML. Policy is defined once in Apache Ranger and enforced across Trino, Spark, Kyuubi, and S3 object storage per object, and tagging a dataset in DataHub generates those policies at bucket, schema, and table level. Compaction, snapshot expiration, and orphan cleanup run as platform maintenance tasks. Agent and MCP requests pass through the AI & Data Control Plane, which answers deterministic questions from the query engines with zero tokens, applies per-role budgets and PII redaction, and logs every request. The layer deploys identically on-prem, in any cloud, hybrid, or air-gapped. To work through how your catalog and engine choices map onto this design, talk with a NexusOne architect.
Key Takeaways
Iceberg adds ACID commits, metadata-only schema and partition evolution, and snapshot time travel to Parquet on object storage, defined by a public specification any engine can implement.
The format consolidation ran from June 2024 (Databricks buys Tabular, Snowflake open-sources Polaris) to May 2026 (Iceberg v3 GA on Databricks, Google's serverless REST catalog).
Iceberg, Delta Lake, and Hudi all provide transactions; they differ on who governs the specification and whether the catalog has an open protocol, and only Iceberg has a widely implemented REST catalog standard.
Committing to Iceberg creates three decisions: which catalog holds the control point, who runs compaction and snapshot expiration, and which engines run against the same tables.
Managed Iceberg tables keep the files open but move governance into one vendor's identity system; requiring the REST catalog protocol from every client preserves the exit.
Snapshot retention is the reproducibility window for any model trained on the table, so set it against the audit requirement.
FAQ
What Is the Apache Iceberg Use Case for Enterprise AI?
Iceberg lets training, feature engineering, streaming ingestion, and agent retrieval run on the same tables with different engines and no copies. Snapshot time travel reproduces the exact table a model was trained on, schema evolution adds features without backfills, and partition pruning keeps retrieval queries from scanning full tables. Because the specification is public, the engine can change as the workload changes while the data stays in place.
Why Use Apache Iceberg as the Foundation of a Data Platform?
Iceberg gives you database guarantees on commodity object storage and makes the compute layer replaceable. Every major cloud and both large cloud-era platforms now read and write it natively, so a table written today remains readable by whichever engine wins the next decision. It also removes the silent correctness failures of Hive-style partitioning.
How Do Iceberg, Delta Lake, and Hudi Compare for the Enterprise?
All three provide ACID transactions on Parquet files. Iceberg is an Apache project with a public REST catalog specification implemented by Polaris, Nessie, Gravitino, Glue, Unity Catalog, and BigLake. Delta Lake is a Linux Foundation project led by Databricks whose UniForm feature publishes Iceberg-readable metadata. Hudi is an Apache project built for streaming upserts. For multi-engine, multi-environment estates, Iceberg's catalog openness is the deciding factor.
What Is an Open Table Format in an Enterprise Data Lakehouse?
An open table format is a published specification for the metadata that turns files on object storage into a transactional table: which files belong to it, their schema, their partitioning, and the history of snapshots. A data lakehouse is an architecture that runs warehouse-style SQL and AI workloads directly on those tables in the lake, with no load into a proprietary warehouse. The format is what lets several engines share one copy of the data safely.
Which Iceberg Catalog Should an Enterprise Choose?
Choose the catalog by where the control point should live. Apache Polaris, Project Nessie, and Apache Gravitino are Apache-licensed and run anywhere, including on-prem and air-gapped, and Gravitino can federate other catalogs into one namespace. AWS Glue, Unity Catalog, and BigLake remove operations work but keep authorization and credential vending inside one vendor's identity system. Whatever you pick, require the Iceberg REST protocol from every client so the catalog can be moved later.
References
Databricks + Tabular, Databricks Blog, June 4, 2024. https://www.databricks.com/blog/databricks-tabular
Databricks $1B-plus Tabular acquisition adds Iceberg support, TechTarget, June 6, 2024. https://www.techtarget.com/searchdatamanagement/news/366588032/Databricks-1B-plus-Tabular-acquisition-adds-Iceberg-support
Polaris Catalog Is Now Open Source, Snowflake, July 30, 2024. https://www.snowflake.com/en/blog/polaris-catalog-open-source/
Announcing Amazon S3 Tables: Fully managed Apache Iceberg tables optimized for analytics workloads, AWS, December 3, 2024. https://aws.amazon.com/about-aws/whats-new/2024/12/amazon-s3-tables-apache-iceberg-tables-analytics-workloads
Advancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance, Databricks Blog, May 28, 2026. https://www.databricks.com/blog/unity-catalog-and-next-era-apache-icebergtm
Google Cloud Introduces Cross-Engine Iceberg Support in BigQuery, InfoQ, May 23, 2026. https://www.infoq.com/news/2026/05/google-cross-engine-iceberg/
New "State of the Data Lakehouse in the AI Era" Report Shows Data Lakehouses Accelerating AI Readiness for 85% of Firms, Dremio, January 9, 2025. https://www.dremio.com/press-releases/new-state-of-the-data-lakehouse-in-the-ai-era-report-shows-data-lakehouses-accelerating-ai-readiness-for-85-of-firms/
Iceberg Table Spec, Apache Iceberg. https://iceberg.apache.org/spec/
AWS announces support for Apache Iceberg V3 deletion vectors and row lineage, AWS, November 26, 2025. https://aws.amazon.com/about-aws/whats-new/2025/11/aws-apache-iceberg-v3-deletion-vectors-row-lineage
Evolution, Apache Iceberg Documentation. https://iceberg.apache.org/docs/latest/evolution/
Partitioning, Apache Iceberg Documentation. https://iceberg.apache.org/docs/latest/partitioning/
Writes in DuckDB-Iceberg, DuckDB, November 28, 2025. https://duckdb.org/2025/11/28/iceberg-writes-in-duckdb
Iceberg connector, Trino Documentation. https://trino.io/docs/current/connector/iceberg.html
Apache Iceberg, Wikipedia. https://en.wikipedia.org/wiki/Apache_Iceberg
Delta Lake project site (Linux Foundation, UniForm, adoption). https://delta.io/
Overview, Apache Hudi Documentation. https://hudi.apache.org/docs/overview
Apache Hudi: The Streaming Data Lake Platform, Apache Hudi Blog, July 21, 2021. https://hudi.apache.org/blog/2021/07/21/streaming-data-lake-platform/
Apache Iceberg REST Catalog API (OpenAPI specification), Apache Iceberg. https://github.com/apache/iceberg/blob/main/open-api/rest-catalog-open-api.yaml
Spark Procedures (rewrite_data_files, expire_snapshots, remove_orphan_files), Apache Iceberg Documentation. https://iceberg.apache.org/docs/latest/spark-procedures/
How Amazon S3 Tables use compaction to improve query performance by up to 3 times, AWS Storage Blog, December 4, 2024. https://aws.amazon.com/blogs/storage/how-amazon-s3-tables-use-compaction-to-improve-query-performance-by-up-to-3-times/
Apache Polaris (Incubating) 1.0 Released, Snowflake Engineering Blog, July 10, 2025. https://www.snowflake.com/en/engineering-blog/apache-polaris-1-0-release-open-source-catalog/
Project Nessie: Transactional Catalog for Data Lakes with Git-like semantics. https://projectnessie.org/
Iceberg REST catalog service, Apache Gravitino Documentation. https://gravitino.apache.org/docs/latest/iceberg-rest-service
AWS Glue Data Catalog supports automatic compaction for Apache Iceberg tables, AWS, November 15, 2023. https://aws.amazon.com/about-aws/whats-new/2023/11/aws-glue-data-catalog-compaction-iceberg-tables

