Zero Lock-in Compute: How to Separate Storage, Compute, and Serving Layers Without an Integration Nightmare

Zero Lock-in Compute: How to Separate Storage, Compute, and Serving Layers Without an Integration Nightmare

A composable data architecture holds together only when each layer exposes a stable abstraction: an open table format at storage, a federated query interface at compute, a semantic API at serving. What each abstraction guarantees, why most composable attempts collapse into point-to-point integration, and how to test whether a component is swappable before you buy it.

By

Billy Allocca

Table of Contents

Zero Lock-in Compute: How to Separate Storage, Compute, and Serving Layers Without an Integration Nightmare

A composable data architecture separates storage, compute, and serving into layers that communicate only through stable, open abstractions: an open table format (Apache Iceberg) at storage, a federated query interface at compute, and a semantic API at serving. Because each layer depends on the abstraction and never on a product, any component can be replaced without re-engineering the rest.

The cost of getting this wrong is measurable. MuleSoft's 2026 Connectivity Benchmark Report, published February 5, 2026, found the average organization manages 957 applications, only 27 percent of them connected, with IT teams spending 36 percent of their time building and testing custom integrations [1]. Okta's Businesses at Work 2025 report counted an average of 101 apps per company [2]; a data estate wired tool-to-tool inherits the same curve.

Snowflake's Iceberg tables, on object storage you manage, reached general availability on June 10, 2024 [3], a week after Snowflake announced the Polaris Catalog implementing Iceberg's REST API [4]. Databricks announced its acquisition of Tabular on June 4, 2024, with a stated goal of Delta Lake and Iceberg format interoperability [5], and on June 12, 2025 opened Unity Catalog to any REST-compatible Iceberg client for reads and writes [6]. What has to be true at compute and serving for that openness to be usable is the subject here; the post on why composability is required at every layer covers why query routing alone does not qualify.

What Does It Mean to Separate Storage, Compute, and Serving in a Data Architecture?

Separation means each layer exposes an abstraction the layer above depends on, and nothing above that abstraction knows which product implements it. The test is how many other systems need changes when you replace the product behind one abstraction; in a separated architecture that number is zero for every non-adjacent layer.

An open table format is a foundation-held specification for how a table's data files, schema, partitions, and version history are laid out on object storage so any conforming engine can read and write it. The Apache Iceberg spec commits to serializable isolation, multiple concurrent writers through atomic metadata swaps, and O(1) remote calls to plan a scan [7].

A federated query layer is a compute tier that resolves one SQL statement across several data sources, each reached through a connector. Trino's documentation states that many catalogs, each backed by a connector, can be queried within the same SQL statement [8]; Starburst's documentation shows an S3 catalog joined to a MySQL catalog by fully qualified catalog.schema.table names [9].

A semantic layer is a serving tier that maps physical tables to business-defined views (metrics, dimensions, entities) so consumers query the definition and never the physical location. Dremio's documentation states the property that matters: when data migrates to a new source, the view definition updates and the consumer's query is unchanged [10].

Layer

Stable abstraction

What "composable" means here

What breaks without it

Storage

Open table format (Apache Iceberg spec plus REST catalog API) on any S3-compatible object store

Any conforming engine reads and writes the same tables; object store and catalog swap independently

Tables are readable only by the engine that wrote them; a storage change means re-registering every table

Compute

Federated query interface (Trino, Apache Spark, Apache Kyuubi as the SQL gateway) under one identity and one policy model

Engines are added or retired behind one SQL endpoint; sources are joined in place without copying

Each engine carries its own connectors and access rules; every new engine is a new policy silo

Serving

Semantic API (metric and entity definitions exposed over SQL, REST, and MCP)

BI tools, applications, and agents query definitions; physical moves behind the definition are invisible to them

Dashboards and agents hard-code table paths; any storage or compute change becomes a consumer-by-consumer rewrite

What Are the Composable Data Architecture Patterns That Hold Up in Production?

One abstraction per boundary, which the layer above may never bypass, is the pattern that survives. Delta Lake UniForm generates Iceberg metadata alongside Delta metadata without rewriting Parquet files, so one copy of data serves both Delta and Iceberg clients, and the Databricks documentation is explicit that Iceberg client support is read-only [11]. That is a real interoperability step and a one-directional one: the write path still belongs to a single engine. The June 2025 Managed Iceberg tables announcement opened reads and writes to any client conforming to the Iceberg REST spec, naming Spark, Flink, and Trino [6]. The abstraction became a boundary, and that is what makes storage swappable.

The compute boundary follows the gateway pattern: Apache Kyuubi describes itself as a distributed, multi-tenant gateway providing serverless SQL on lakehouses over Spark, Flink, and Trino [12], so consumers connect to one JDBC endpoint and the engine behind it is a deployment decision.

Why Do Most Composable Architecture Attempts Fail?

They fail because "best of breed" gets built as direct connections between tools, and every connection is a contract renegotiated when either side changes; eight tools wired pairwise is 28 contracts. The Iceberg REST Catalog OpenAPI definition deletes one class of them by specifying namespaces, tables, views, and multi-table commits over HTTP, so a client can talk to any catalog implementation without a catalog-specific library [13].

The LY Corporation (Yahoo! JAPAN) engineering team documented an engine swap in a November 15, 2024 post: migrating more than 200 statistical workflows from Apache Hive to Trino and Spark on unchanged HDFS storage and an unchanged Hive Metastore, at 600 GB per partition and 15 partitions a day, with Spark delivering roughly 2.4 times Hive's CPU efficiency. The friction came from what was never abstracted: Trino inserts into partitioned Hive tables that did not reliably synchronize the metastore, type-casting differences, and memory-heavy tasks that had to be steered to Spark [14].

Symptom

Root cause

Fix

Swapping one engine forces changes in dashboards and notebooks

Consumers query physical table paths, so compute and serving are fused

Put a semantic layer between consumers and engines; expose definitions only

Each new engine needs its own access policy configuration

Authorization lives inside each engine

Externalize policy to Apache Ranger and identity to Keycloak, enforced by every engine

Tables written by one tool are unreadable or read-only for others

Proprietary metadata, or a one-directional compatibility layer

Standardize on the Iceberg spec with a REST catalog; require read and write conformance

Query results differ between engines on the same data

SQL dialect and type-casting differences reach users

Route through one SQL gateway (Apache Kyuubi or Trino) with a tested dialect contract

Metadata drifts between engines after writes

Multiple catalogs, or engines updating a metastore inconsistently

One catalog of record over the REST catalog API; every engine commits through it

Integration backlog grows faster than the team

N squared point-to-point contracts

Normalize every tool to one integration surface per boundary (format, SQL, semantic API)

How Do You Avoid Vendor Lock-in in a Data Platform?

You avoid it by testing each component against one criterion before purchase: it can be removed without touching non-adjacent layers. The checklist turns the criterion into evidence you can request from a vendor.

  • It reads and writes tables through the published Apache Iceberg spec, with any REST catalog implementation, including one it does not sell.

  • Data files stay on object storage you control, in Parquet, with metadata you can list and copy while the vendor's software is off.

  • Identity is federated from your provider and authorization lives in a policy store you own, so revoking the component revokes nothing else.

  • Its SQL surface is the one your gateway already exposes, so retiring it changes no consumer connection string.

  • Metric definitions and lineage live in a catalog outside the component and export in an open format.

  • Another engine in your estate has already run the same workload against the same tables.

  • The vendor's documentation states what an external client can do with your tables when their service is off.

Why Not Just Standardize on One Vendor's Platform?

Standardizing on one platform is a legitimate choice when your estate fits inside it. The counterargument is scope. The MuleSoft figure of 957 applications with 27 percent connected [1] describes an estate no single platform contains, so under standardization the integration problem relocates to the platform's perimeter, where its connectors become the point-to-point layer.

There is also a direction-of-control problem. A platform can open the format of the tables it manages while holding the catalog, the identity model, and the policy engine inside its boundary, so you can walk away with your Parquet files and still lose every access rule and metric definition. That is why the post on the semantic layer that travels with the data argues semantics have to move with the data. Openness at storage is necessary and insufficient without open compute and serving.

NexusOne Makes the Storage Layer Swappable From the Compute and Serving Tiers

NexusOne occupies the compute and serving tiers and treats storage as an interface. At compute, it runs Trino, Apache Spark, and Apache Kyuubi as a federated query layer over every source in the estate, on Apache Iceberg tables on any S3-compatible object store. The engines share one identity model (Keycloak) and one policy engine (Apache Ranger, enforced simultaneously across Trino, Spark, Kyuubi, and per-object on S3), so adding or retiring an engine creates no policy silo. Tagging a dataset in DataHub and mapping a role to the tag generates the Ranger policies for every engine at once.

At serving, the semantic layer holds metric and entity definitions and exposes them over SQL and through the AI & Data Control Plane, which governs agent and MCP requests with per-role token budgets, PII redaction before egress, and a log of every request. Normalization is the principle underneath: each tool maps to one standard at each boundary, so the estate has one integration surface. It runs identically on-prem, in the cloud, or air-gapped, and if you leave, your tables remain Iceberg on storage you own. To score your components against the checklist, talk with a NexusOne architect.

Key Takeaways

  • A composable data architecture is defined by its abstractions: an open table format at storage, a federated query interface at compute, and a semantic API at serving.

  • Point-to-point integration scales as N squared; the 2026 MuleSoft benchmark puts the average enterprise at 957 applications with 27 percent connected.

  • Read-only interoperability (UniForm) and read-write interoperability (REST catalog conformance) differ in kind; only the second makes storage swappable.

  • Policy, identity, and semantic definitions must live outside any single engine, or the engine remains the lock-in.

FAQ

How Do You Separate Storage, Compute, and Serving in a Data Architecture?

Put a published abstraction at each boundary and forbid anything above it from bypassing that abstraction. Storage exposes Apache Iceberg tables through a REST catalog on object storage you control. Compute exposes one SQL surface through a federated engine layer (Trino, Spark, Kyuubi) carrying one identity and one policy model. Serving exposes business definitions through a semantic layer so consumers never reference physical tables. Dependencies then run downward only, and any layer can be replaced without changing the ones that are not adjacent to it.

What Are the Main Composable Data Architecture Patterns?

The patterns that hold up are the open-table-format pattern at storage (Iceberg spec plus REST catalog), the gateway pattern at compute (one SQL endpoint such as Apache Kyuubi or Trino fronting several engines), and the semantic-view pattern at serving (views or a metrics layer mapping definitions to physical tables). Each removes a class of pairwise contract. Compatibility shims such as Delta Lake UniForm are transitional: they widen reads without opening writes, so they help without completing the boundary.

How Do You Avoid Vendor Lock-in in a Data Platform?

Evaluate each component on whether it can be removed without touching non-adjacent layers. Require read and write conformance to the Iceberg spec through any REST catalog, keep data files on storage you control, federate identity from your own provider, hold authorization in a policy store you own such as Apache Ranger, and keep metrics and lineage in a catalog outside the component. Then prove it by running the same workload on a second engine against the same tables in a non-production tenant.

What Does an Open Data Architecture Look Like for an Enterprise CDO?

It is a three-tier estate where the choices that are expensive to reverse (table format, catalog protocol, identity model, policy engine, semantic definitions) are open standards or software you operate, and the choices that are cheap to reverse (which engine, which object store, which BI tool) are treated as replaceable. Snowflake and Databricks both shipped Iceberg support with external REST catalog access between June 2024 and June 2025, so the storage tier is now open across the cloud-era platforms. The remaining work is making compute and serving equally open, which the sibling guides on Apache Iceberg for enterprise AI and the vendor-neutral data estate take up in detail.

Does Query Federation Replace a Data Lakehouse?

No. Federation reads data where it lives and joins across sources in one statement, while a lakehouse stores data in an open table format on object storage. In a composable estate they are adjacent layers: the federated engine is the compute tier and the Iceberg lakehouse is the storage tier. Federation lets you query operational systems, warehouses, and the lakehouse together without copying, and the lakehouse gives every engine a shared, versioned, concurrently writable table.

References

  1. Key Findings from MuleSoft's 2026 Connectivity Benchmark Report, MuleSoft Blog, https://blogs.mulesoft.com/news/connectivity-benchmark-report/

  2. Businesses at Work 2025: 10 years of data show how critical security has become, Okta Newsroom, https://www.okta.com/newsroom/articles/businesses-at-work-2025/

  3. June 10, 2024, Apache Iceberg tables, General Availability, Snowflake Documentation, https://docs.snowflake.com/en/release-notes/2024/other/2024-06-10-iceberg-tables

  4. Polaris Catalog: An Open Source Catalog for Apache Iceberg, Snowflake Blog, https://www.snowflake.com/en/blog/introducing-polaris-catalog/

  5. Databricks + Tabular, Databricks Blog, https://www.databricks.com/blog/databricks-tabular

  6. Announcing full Apache Iceberg support in Databricks, Databricks Blog, https://www.databricks.com/blog/announcing-full-apache-iceberg-support-databricks

  7. Iceberg Table Spec, Apache Iceberg, https://iceberg.apache.org/spec/

  8. Trino concepts, Trino Documentation, https://trino.io/docs/current/overview/concepts.html

  9. Federate queries from different sources, Starburst Galaxy Documentation, https://docs.starburst.io/starburst-galaxy/working-with-data/query-data/federate-queries-from-different-sources.html

  10. Self-Service Semantic Layer, Dremio Documentation, https://docs.dremio.com/current/help-support/well-architected-framework/semantic/

  11. Use UniForm to read Delta tables with Iceberg clients, Databricks Documentation, https://docs.databricks.com/aws/en/delta/uniform

  12. Apache Kyuubi, Apache Software Foundation, https://kyuubi.apache.org/

  13. Iceberg REST Catalog Spec (OpenAPI), Apache Iceberg, https://iceberg.apache.org/rest-catalog-spec/

  14. Migration from Hive to Trino and Spark: Challenges and insights, LY Corporation Tech Blog, https://techblog.lycorp.co.jp/en/migration-from-hive-to-trino-and-spark

Other posts

Other posts

Trusted at every layer.

One security model across every system. Cell-level encryption. Row-level security. Agent permission impersonation. VPC isolation. 500+ audits/year passed at production customers.

Newsletter

Keep updated

1115 Howell Mill Rd, Suite 430,
Atlanta, GA 30318

Back to top

©2026 NexusOne® All rights reserved.

Trusted at every layer.

One security model across every system. Cell-level encryption. Row-level security. Agent permission impersonation. VPC isolation. 500+ audits/year passed at production customers.

GitHub

LinkedIn

Careers

About

Blog

Newsletter

Keep updated

1115 Howell Mill Rd, Suite 430,
Atlanta, GA 30318

@2026 NexusOne® -
All rights reserved.

Back to top