Why Apache Iceberg Is Becoming the Default Storage Format for Enterprise AI, and What to Build Around It

Why Apache Iceberg Is Becoming the Default Storage Format for Enterprise AI, and What to Build Around It

Iceberg gives every query engine the same transactional view of the same files on object storage. The format decision is settled; the catalog, compaction, and engine decisions that follow it decide whether an enterprise AI program stays open.

By

Billy Allocca

Table of Contents

Why Apache Iceberg Is Becoming the Default Storage Format for Enterprise AI, and What to Build Around It

Apache Iceberg is an open table format that adds ACID transactions, schema and partition evolution, and snapshot-based time travel to Parquet files on object storage, with the layout defined by a public specification. Any compliant engine (Trino, Spark, Flink, DuckDB) reads the same tables, so an enterprise can change compute without moving or rewriting its data.

The format debate closed within two years. Databricks acquired Tabular, the company founded by Iceberg's creators, on June 4, 2024, for a price it confirmed was above $1 billion [1][2]. Snowflake open-sourced its Polaris Catalog under the Apache 2.0 license on July 30, 2024 [3]. AWS shipped Amazon S3 Tables, Iceberg support built into the object store itself, on December 3, 2024 [4]. By May 2026, Databricks had declared Iceberg v3 and managed Iceberg tables generally available in Unity Catalog [5], and Google had previewed a serverless Iceberg REST catalog so Spark, Flink, and Trino could work on BigQuery data without copies [6].

AI workloads made the consolidation urgent. In Dremio's Q4 2024 survey of 563 IT decision-makers, 85% said they use a data lakehouse for AI model development [7], and training, feature computation, and agent retrieval each read the same tables through different engines. The same argument applied to compute is in the earlier post on why composability is required at every layer.

What Does Apache Iceberg Provide That a Parquet Data Lake Does Not?

Iceberg turns a directory of files into a table with database guarantees, entirely through metadata. An open table format is a specification for the metadata that tracks which data files belong to a table, what schema they carry, and how they are partitioned, so reads and writes behave like operations on a database table. The Iceberg specification states that reads always use a committed snapshot, the complete set of data files that made up the table at one point in time, and that writers coordinate through optimistic concurrency without readers taking locks [8]. Version 2 added row-level deletes; version 3 added binary deletion vectors, a variant type for semi-structured data, row lineage, and default column values [8], and AWS made the v3 deletion vectors and row lineage available across EMR, Glue, and S3 Tables on November 26, 2025 [9].

Capability

What Iceberg does

What it means for AI workloads

ACID transactions on object storage

Atomic snapshot commits; readers never see a half-written table [8]

A training job and a streaming upsert can touch the same table at once without producing a corrupt feature set

Schema evolution

Add, drop, rename, widen, and reorder columns as metadata-only changes with no file rewrite [10]

New features are added to a training table without a backfill; models that never asked for the column are unaffected

Time travel

Every write creates a snapshot that queries can target by ID or timestamp [8]

A model can be reproduced against the exact table state it was trained on, which is the provenance evidence an audit asks for

Hidden partitioning and pruning

Partition values are derived from column transforms; planners skip files that cannot match a filter [11]

Retrieval queries scan only the files they need; Hive-style tables return silently wrong results when a writer uses the wrong date format [11]

Partition evolution

Change the partition spec; old files keep the old layout and queries plan across both [10]

A daily-partitioned table moves to hourly partitions as agent traffic grows, with no rewrite of history

Hidden partitioning means the table itself owns the mapping between a column and its physical partition, so a filter on event_time prunes correctly without anyone knowing the table is partitioned by day(event_time).

Why Use Apache Iceberg as the Storage Layer for Enterprise AI?

Iceberg is the interface contract that makes the compute layer swappable, and for an AI program that property outweighs any single performance feature. Once tables are Iceberg, Spark writes them, Trino serves interactive queries, Flink streams into them, and DuckDB reads them on a laptop, all through the same specification. DuckDB added insert, update, and delete support for Iceberg tables through REST catalogs in November 2025 [12], and Trino's connector supports six catalog types, including REST and Nessie [13].

The origin explains the neutrality. Ryan Blue and Daniel Weeks built Iceberg at Netflix to replace Hive tables that failed under concurrent writes, then donated it to the Apache Software Foundation [1], where it graduated to a top-level project in May 2020 [14]. The specification was written for engines Netflix did not control, which is why competitors could adopt it.

Date

Event

What it settled

June 4, 2024

Databricks acquires Tabular for more than $1 billion [1][2]

The Delta Lake vendor commits to converging with Iceberg

July 30, 2024

Snowflake open-sources Polaris Catalog [3]

The catalog, previously the closed piece, gets an Apache-licensed implementation

December 3, 2024

AWS launches Amazon S3 Tables [4]

Iceberg becomes a primitive of the object store, with built-in compaction

May 28, 2026

Databricks: Iceberg v3 and managed Iceberg GA in Unity Catalog [5]

The largest cloud-era platform treats Iceberg as a first-class managed format

For AI, the payoff is one governed copy of the data that every engine and every agent reads. One copy per platform in a proprietary format is the pattern the Hadoop modernization post identifies as the cause of multi-year timelines.

How Does Iceberg Compare to Delta Lake and Apache Hudi for the Enterprise?

All three formats deliver transactions on object storage, so the enterprise distinction is who governs the specification and whether the catalog has an open protocol, and there Iceberg has the widest neutral footprint. Delta Lake is hosted by the Linux Foundation and describes itself as independent of any single company [15], though Databricks leads its protocol, and its UniForm feature exposes Delta tables to Iceberg and Hudi readers [15]. Apache Hudi, which grew out of Uber's incremental processing work beginning in 2016, describes itself as a lakehouse platform with a table format built for high-performance writes on incremental pipelines [16][17].

Dimension

Apache Iceberg

Delta Lake

Apache Hudi

Governance model

Apache Software Foundation top-level project since May 2020 [14]

Linux Foundation project; protocol led by Databricks [15]

Apache Software Foundation top-level project [16]

Catalog openness

Public REST catalog specification implemented by Polaris, Nessie, Gravitino, Glue, Unity Catalog, and BigLake [18]

Catalog is engine-side; UniForm publishes Iceberg metadata for Iceberg catalogs to read [15]

Syncs metadata to Hive Metastore and AWS Glue; no independent catalog protocol [16]

Engine support

Spark, Trino, Flink, DuckDB, Snowflake, BigQuery, Databricks, Dremio, Athena, EMR

Deepest on Databricks and Spark; read paths elsewhere

Deepest on Spark and Flink; read paths in Trino and Presto

Adoption

Netflix, Apple, Airbnb, Adobe, Expedia; native in AWS, Google Cloud, Snowflake, Databricks [14]

Databricks reports 10,000+ production environments [15]

Uber, Amazon, ByteDance, Robinhood [16]

Best fit

Multi-engine estates that want format and catalog independent of any compute vendor

Estates standardized on Databricks that want Iceberg readability

Streaming upsert-heavy pipelines

Merge-on-read is a write strategy that stores deletes and updates as separate files that are reconciled at query time, which favors write speed; copy-on-write rewrites affected data files at write time, which favors reads. Both Hudi and Iceberg v2 support the merge-on-read pattern. The convergence Databricks described in June 2024, two formats that share Parquet yet became incompatible through independent development [1], is now moving toward Iceberg: v3 deletion vectors and row lineage use encodings compatible with Delta Lake's, and Databricks has proposed that Delta 5.0 adopt the adaptive metadata tree planned for Iceberg v4 [5].

What Are the Three Decisions That Follow an Iceberg Commitment?

Choosing Iceberg removes one decision and creates three, and the catalog decision determines whether the openness survives. A catalog is the service that maps a table name to its current metadata file; because every commit goes through it, whoever runs the catalog controls access, credential vending, and which engines may write. The Iceberg REST catalog is the specification's HTTP protocol for that service, with endpoints for namespaces, tables, views, and vended credentials, so a client written against it can point at Polaris, Nessie, Gravitino, Glue, or Unity Catalog without code changes [18].

Compaction is the rewriting of many small data files into fewer large ones. Iceberg tracks every file in manifests, so small files inflate metadata and slow queries through per-file open costs [19]. AWS measured 8.5 times more read operations on 1 MB files than on 512 MB objects, and automatic compaction in S3 Tables sped up eight benchmark queries by 1.12 to 3.2 times [20].

Decision

Options

Trade-offs

Catalog

Apache Polaris (1.0 released July 10, 2025, with RBAC and credential vending) [21]; Project Nessie (Git-like branches and cross-table transactions) [22]; Apache Gravitino (federated metadata lake whose Iceberg REST service proxies Hive, JDBC, and other REST catalogs) [23]; AWS Glue; Unity Catalog; BigLake

Self-hosted Apache catalogs keep the control point inside your estate and run identically on-prem and in any cloud. Vendor catalogs remove operations work but tie access policy and credential vending to one provider's identity system

Compaction strategy

Scheduled Spark rewrite_data_files plus expire_snapshots and remove_orphan_files [19]; Trino ALTER TABLE EXECUTE optimize [13]; managed compaction in S3 Tables (64 to 512 MB target) [20], Glue (since November 15, 2023) [24], or Unity Catalog Predictive Optimization [5]

Self-run compaction gives control over file size, sort order, and cost windows but needs an owner and a schedule. Managed compaction only runs inside the vendor's storage or catalog, so tables outside it go unmaintained

Query engine

Trino for interactive federation; Spark for training-set preparation; Flink for streaming ingestion; DuckDB for local reads; Apache Kyuubi as a multi-tenant SQL gateway over Spark

Several engines can run against the same tables. Identity and policy then have to be enforced above the engines, or each engine becomes its own governance island

Snapshot expiration and time travel interact: the retention window set in expire_snapshots is also the reproducibility window for any model trained on that table [19]. The engine question continues in the guide on separating storage, compute, and serving layers, and catalog ownership in the CDO's guide to a vendor-neutral data estate.

Why Not Just Use a Cloud Vendor's Managed Iceberg Tables?

For a single-cloud estate with one dominant engine, managed Iceberg from AWS, Databricks, Google, or Snowflake is a reasonable start, because someone else runs compaction and the files stay readable by any compliant engine. What the arrangement costs is the control point. Databricks' managed Iceberg tables are governed by Unity Catalog and maintained by Predictive Optimization [5]; S3 Tables enforce table-level permissions through AWS IAM [4]. Both are sound designs, and both mean the access policy for a table lives in one vendor's identity system, an agent running in another cloud authenticates through that vendor, and a deployment in a data center you own cannot use the same catalog. The files stay open while the governance becomes vendor-specific.

The thesis holds because the REST protocol exists. You can adopt a managed service today and keep an exit by requiring that every reader and writer speak the Iceberg REST catalog API, so moving the catalog to Polaris, Nessie, or Gravitino later is a configuration change [18].

NexusOne Reads Iceberg Natively Across Every Engine and Environment

NexusOne treats Iceberg as the default table format and runs the three follow-on decisions as one configured layer. Tables created through the platform register automatically in an Apache Gravitino catalog that exposes the Iceberg REST protocol, and Gravitino federates external catalogs (Hive Metastore, Glue, other REST endpoints, plus Databricks and Teradata metadata) into one namespace, so Trino, Apache Spark, and Apache Kyuubi see the same tables whether they sit in an on-prem object store or a cloud bucket [23]. Identity arrives through one Keycloak layer federated from Active Directory, Okta, LDAP, or SAML. Policy is defined once in Apache Ranger and enforced across Trino, Spark, Kyuubi, and S3 object storage per object, and tagging a dataset in DataHub generates those policies at bucket, schema, and table level. Compaction, snapshot expiration, and orphan cleanup run as platform maintenance tasks. Agent and MCP requests pass through the AI & Data Control Plane, which answers deterministic questions from the query engines with zero tokens, applies per-role budgets and PII redaction, and logs every request. The layer deploys identically on-prem, in any cloud, hybrid, or air-gapped. To work through how your catalog and engine choices map onto this design, talk with a NexusOne architect.

Key Takeaways

  • Iceberg adds ACID commits, metadata-only schema and partition evolution, and snapshot time travel to Parquet on object storage, defined by a public specification any engine can implement.

  • The format consolidation ran from June 2024 (Databricks buys Tabular, Snowflake open-sources Polaris) to May 2026 (Iceberg v3 GA on Databricks, Google's serverless REST catalog).

  • Iceberg, Delta Lake, and Hudi all provide transactions; they differ on who governs the specification and whether the catalog has an open protocol, and only Iceberg has a widely implemented REST catalog standard.

  • Committing to Iceberg creates three decisions: which catalog holds the control point, who runs compaction and snapshot expiration, and which engines run against the same tables.

  • Managed Iceberg tables keep the files open but move governance into one vendor's identity system; requiring the REST catalog protocol from every client preserves the exit.

  • Snapshot retention is the reproducibility window for any model trained on the table, so set it against the audit requirement.

FAQ

What Is the Apache Iceberg Use Case for Enterprise AI?

Iceberg lets training, feature engineering, streaming ingestion, and agent retrieval run on the same tables with different engines and no copies. Snapshot time travel reproduces the exact table a model was trained on, schema evolution adds features without backfills, and partition pruning keeps retrieval queries from scanning full tables. Because the specification is public, the engine can change as the workload changes while the data stays in place.

Why Use Apache Iceberg as the Foundation of a Data Platform?

Iceberg gives you database guarantees on commodity object storage and makes the compute layer replaceable. Every major cloud and both large cloud-era platforms now read and write it natively, so a table written today remains readable by whichever engine wins the next decision. It also removes the silent correctness failures of Hive-style partitioning.

How Do Iceberg, Delta Lake, and Hudi Compare for the Enterprise?

All three provide ACID transactions on Parquet files. Iceberg is an Apache project with a public REST catalog specification implemented by Polaris, Nessie, Gravitino, Glue, Unity Catalog, and BigLake. Delta Lake is a Linux Foundation project led by Databricks whose UniForm feature publishes Iceberg-readable metadata. Hudi is an Apache project built for streaming upserts. For multi-engine, multi-environment estates, Iceberg's catalog openness is the deciding factor.

What Is an Open Table Format in an Enterprise Data Lakehouse?

An open table format is a published specification for the metadata that turns files on object storage into a transactional table: which files belong to it, their schema, their partitioning, and the history of snapshots. A data lakehouse is an architecture that runs warehouse-style SQL and AI workloads directly on those tables in the lake, with no load into a proprietary warehouse. The format is what lets several engines share one copy of the data safely.

Which Iceberg Catalog Should an Enterprise Choose?

Choose the catalog by where the control point should live. Apache Polaris, Project Nessie, and Apache Gravitino are Apache-licensed and run anywhere, including on-prem and air-gapped, and Gravitino can federate other catalogs into one namespace. AWS Glue, Unity Catalog, and BigLake remove operations work but keep authorization and credential vending inside one vendor's identity system. Whatever you pick, require the Iceberg REST protocol from every client so the catalog can be moved later.

References

  1. Databricks + Tabular, Databricks Blog, June 4, 2024. https://www.databricks.com/blog/databricks-tabular

  2. Databricks $1B-plus Tabular acquisition adds Iceberg support, TechTarget, June 6, 2024. https://www.techtarget.com/searchdatamanagement/news/366588032/Databricks-1B-plus-Tabular-acquisition-adds-Iceberg-support

  3. Polaris Catalog Is Now Open Source, Snowflake, July 30, 2024. https://www.snowflake.com/en/blog/polaris-catalog-open-source/

  4. Announcing Amazon S3 Tables: Fully managed Apache Iceberg tables optimized for analytics workloads, AWS, December 3, 2024. https://aws.amazon.com/about-aws/whats-new/2024/12/amazon-s3-tables-apache-iceberg-tables-analytics-workloads

  5. Advancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance, Databricks Blog, May 28, 2026. https://www.databricks.com/blog/unity-catalog-and-next-era-apache-icebergtm

  6. Google Cloud Introduces Cross-Engine Iceberg Support in BigQuery, InfoQ, May 23, 2026. https://www.infoq.com/news/2026/05/google-cross-engine-iceberg/

  7. New "State of the Data Lakehouse in the AI Era" Report Shows Data Lakehouses Accelerating AI Readiness for 85% of Firms, Dremio, January 9, 2025. https://www.dremio.com/press-releases/new-state-of-the-data-lakehouse-in-the-ai-era-report-shows-data-lakehouses-accelerating-ai-readiness-for-85-of-firms/

  8. Iceberg Table Spec, Apache Iceberg. https://iceberg.apache.org/spec/

  9. AWS announces support for Apache Iceberg V3 deletion vectors and row lineage, AWS, November 26, 2025. https://aws.amazon.com/about-aws/whats-new/2025/11/aws-apache-iceberg-v3-deletion-vectors-row-lineage

  10. Evolution, Apache Iceberg Documentation. https://iceberg.apache.org/docs/latest/evolution/

  11. Partitioning, Apache Iceberg Documentation. https://iceberg.apache.org/docs/latest/partitioning/

  12. Writes in DuckDB-Iceberg, DuckDB, November 28, 2025. https://duckdb.org/2025/11/28/iceberg-writes-in-duckdb

  13. Iceberg connector, Trino Documentation. https://trino.io/docs/current/connector/iceberg.html

  14. Apache Iceberg, Wikipedia. https://en.wikipedia.org/wiki/Apache_Iceberg

  15. Delta Lake project site (Linux Foundation, UniForm, adoption). https://delta.io/

  16. Overview, Apache Hudi Documentation. https://hudi.apache.org/docs/overview

  17. Apache Hudi: The Streaming Data Lake Platform, Apache Hudi Blog, July 21, 2021. https://hudi.apache.org/blog/2021/07/21/streaming-data-lake-platform/

  18. Apache Iceberg REST Catalog API (OpenAPI specification), Apache Iceberg. https://github.com/apache/iceberg/blob/main/open-api/rest-catalog-open-api.yaml

  19. Spark Procedures (rewrite_data_files, expire_snapshots, remove_orphan_files), Apache Iceberg Documentation. https://iceberg.apache.org/docs/latest/spark-procedures/

  20. How Amazon S3 Tables use compaction to improve query performance by up to 3 times, AWS Storage Blog, December 4, 2024. https://aws.amazon.com/blogs/storage/how-amazon-s3-tables-use-compaction-to-improve-query-performance-by-up-to-3-times/

  21. Apache Polaris (Incubating) 1.0 Released, Snowflake Engineering Blog, July 10, 2025. https://www.snowflake.com/en/engineering-blog/apache-polaris-1-0-release-open-source-catalog/

  22. Project Nessie: Transactional Catalog for Data Lakes with Git-like semantics. https://projectnessie.org/

  23. Iceberg REST catalog service, Apache Gravitino Documentation. https://gravitino.apache.org/docs/latest/iceberg-rest-service

  24. AWS Glue Data Catalog supports automatic compaction for Apache Iceberg tables, AWS, November 15, 2023. https://aws.amazon.com/about-aws/whats-new/2023/11/aws-glue-data-catalog-compaction-iceberg-tables

ABOUT

1115 Howell Mill Rd
Suite 430,
Atlanta, GA 30318
An Insight Partners Company


Product Updates and News

@2026 NexusOne® - All rights reserved.

ABOUT

1115 Howell Mill Rd
Suite 430,
Atlanta, GA 30318
An Insight Partners Company


Product Updates and News

@2026 NexusOne® - All rights reserved.

ABOUT

1115 Howell Mill Rd
Suite 430,
Atlanta, GA 30318
An Insight Partners Company


Product Updates and News

@2026 NexusOne® - All rights reserved.