All glossary terms
D Data & reasoning Architecture

Data lakehouse

The lakehouse solved the storage problem and left the one agents care about. It already holds time travel, lineage and row-level permissions, and almost nobody wires any of it to their agents. The capability is sitting there unused.

Definition

A data lakehouse is an architecture that puts warehouse guarantees, meaning transactions, schema enforcement and governance, directly onto low-cost object storage using open table formats. Databricks introduced the category in 2020. For agentic systems what matters is less the storage unification than whether the governance metadata sitting above it can be read by a machine at query time.

What is a data lakehouse?

A data lakehouse stores all your data cheaply, like a lake, yet keeps it as tidy and reliable as a warehouse. A simple record book tracks every change, so nothing gets lost or overwritten by mistake. Databricks coined the term in 2020 and owns the category.

Attribution matters here. Lakehouse is a coined category with an identifiable origin, and treating it as a generic architecture term rather than crediting where it came from is the kind of thing careful buyers notice.

The mechanism is a transaction log sitting alongside the data files, which is what turns a directory of Parquet into a table with guarantees.

  • Transactions. Writes either complete or do not, so readers never see a half-written table. This is the property a plain data lake lacked and the reason lakes so often became unusable.
  • Schema enforcement and evolution. The table has a declared shape, and changing it is a recorded operation rather than an accident discovered downstream.
  • Versioning. Every write produces a new version, with the old ones still addressable. This is the property this page argues you are underusing.
  • Governance above the storage. A catalog holding permissions, lineage and descriptions, separate from the files themselves.

What problem has Data Lakehouse solved?

The data lakehouse removed the need to run two systems. Companies once stored raw data in a cheap lake and tidy data in a costly warehouse. Teams copied data between them and fixed mismatches. The lakehouse keeps everything in one place, with one copy.

The problem it solved is real and the solution has largely worked, which is worth saying plainly before arguing about what it did not solve.

Data lake, data warehouse and lakehouse compared on cost, guarantees and workload fit
Architecture What it was good at What it could not do
Data lake Cheap storage of anything, including unstructured and semi-structured data No transactions or schema enforcement, so quality degraded until nobody trusted it
Warehouse Reliable, governed, fast on structured analytics Expensive at volume, and poorly suited to unstructured data or machine learning workloads
Lakehouse Warehouse guarantees on lake storage, one copy, both workload types Solves storage and compute. Does not by itself make the data meaningful to a machine

That last cell is the whole argument of this page. Unifying storage was a genuine achievement and it addressed an infrastructure problem. An agent's difficulty is rarely that the data lives in two systems. It is that the data does not explain itself, which is the separate subject covered at data intelligence.

What does an agent need that a dashboard does not?

Governance metadata it can read and act on at query time, rather than metadata a person consults beforehand. A dashboard's permissions are applied by whoever built it. An agent retrieves what it can reach and then acts, so the constraint has to be readable in the query path.

Four properties, and the useful observation is that a lakehouse holds all four and typically exposes none of them to the agent.

Four properties compared across two columns. For each of permissions, history, lineage and freshness, the left column shows how a dashboard consumes it: applied by the person who built the report, rarely needed, checked manually, and known from context. The right column shows what an agent needs: enforced in the query path, queried as of a date, attached to the retrieved fact, and expressed as data it can reason about. A caption notes the lakehouse already holds all four.
Four governance properties, how a dashboard uses each and what an agent requires
Property For a dashboard For an agent
Permissions Applied once by whoever built the report Enforced in the query path, per requester, because the agent chooses what to read
History Rarely needed. The current state is the state Queried as of a date, so a fact can be checked for whether it still holds
Lineage Consulted manually when a number looks wrong Attached to the retrieved fact, so provenance travels into the answer
Freshness Known from context. The reader knows it runs nightly Expressed as data, because an agent has no intuition that a figure looks old

Reading the right-hand column back, it describes something close to a context graph: facts carrying the conditions under which they hold. The lakehouse is not that, and it holds most of the raw material for it.

Which capability do you already have and not use?

In a data lakehouse, time travel shows your data as it stood on any past date. Every change saves a new version and keeps the old ones. Few firms connect this to their AI agents. It answers the question every audit asks: what did the data say then?

This is the strongest practical point on the page and it costs nothing to adopt, because the capability is already paid for.

A table version history along a timeline with four versions, each written on a different date. A standard query reads only the latest version and returns the current value. An as-of query reads the version that was current on a chosen past date and returns what was true then. A panel below notes that this answers what an auditor asks, that it is native to open table formats, and that it is rarely exposed to agents.

What it enables, in order of how often it comes up.

  • Answering as-of questions. What did this record say on the date the decision was taken. Every dispute, audit and complaint starts here, and without versioning the honest answer is that you cannot reconstruct it.
  • Reproducing an agent run. If a trace records which table version the agent read, you can replay against exactly that data rather than against a table that has since changed. This turns an argument about what happened into a check.
  • Distinguishing the two clocks. When a fact became true and when your system learned it are different, and versioning gives you the second for free. The first still has to be modelled deliberately.
  • Detecting silent change. A column whose meaning shifted between versions is visible in the history and invisible in the current state.

One retention caution before you rely on this. Table history is not kept indefinitely. Maintenance operations that reclaim storage will remove old versions past a configured window, and the default window is usually far shorter than any audit retention obligation. If you intend to use versioning as evidence rather than as a convenience, the retention setting is a governance decision rather than a storage one. See AI audit trail for the retention periods different regimes actually require.

Where does the lakehouse fight the agentic loop?

On latency and on access pattern. A lakehouse is optimised for scanning large volumes efficiently, which is the opposite of what an agent does: many small, selective, interactive reads inside a loop where every round trip adds to a response time somebody is waiting for.

This is an architectural mismatch rather than a flaw, and it is manageable if you plan for it rather than discover it.

  • Cold start on object storage. First reads against infrequently touched data can take seconds. Acceptable in a nightly job, unacceptable inside a retrieval loop that may run several rounds per question.
  • The loop multiplies the cost. An agentic retrieval pattern issues multiple queries per question, so a query time that is fine once becomes the dominant latency when it happens four times. See agentic RAG.
  • Batch freshness meets interactive expectation. A table refreshed hourly is entirely reasonable, and an agent asked a question about the last twenty minutes has no way to know it is looking at a stale view unless freshness is recorded as data.
  • Compute economics invert. Elastic compute priced for large scheduled jobs behaves differently under many small concurrent interactive queries, and the bill is the place most teams discover this.

The workable pattern is not to abandon the lakehouse but to stop treating it as the agent's direct query surface. Serve agents from a layer built for selective low-latency reads, keep the lakehouse as the governed source of record behind it, and carry the governance metadata forward rather than leaving it behind at the boundary.

Do you need one before running agents?

No, and the belief that you do has delayed a lot of useful work. Agents read from whatever holds the data, and consolidating an estate is a multi-year programme. If your agents need three systems, connecting to three systems is a faster and more reversible path than centralising first.

Two situations, and they call for different things.

  1. If you do not have a lakehouse

    Do not build one in order to start. Connect agents to the sources that hold the data your live use cases actually read, which is usually a short list, and address meaning and permissions at the connector boundary. See enterprise connectors for what to demand of those integrations, particularly deletion propagation and permission capture.

  2. If you already have one

    The work is not migration, it is exposure. Make the catalog's permissions enforceable in the agent's query path, surface table versions so as-of questions are answerable, carry lineage through to the retrieved fact, and record freshness as data rather than as an assumption. All four already exist in the platform and typically stop at the analytics boundary.

  3. Either way, fix meaning before architecture

    A consolidated platform does not make a field mean the same thing in two systems. If your agents produce confident errors, the cause is more often ambiguous definitions than distributed storage, and no migration resolves that.

The honest summary: a lakehouse is a good answer to a question about analytics infrastructure and a partial answer to a question about agents. Where one already exists, the governance layer above it is worth considerably more to an agentic programme than the storage layer beneath it, and it is the part most organizations have never exposed.

Frequently asked questions about data lakehouses

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is a data lakehouse?

Warehouse behaviour running on lake economics. Open table formats add a transaction log over files sitting in object storage, which produces atomic writes, schema enforcement and versioning on storage costing a fraction of a warehouse. Databricks introduced the term in 2020 and the category is theirs. The mechanism is the transaction log alongside the data files, which is what turns a directory of Parquet into a table with guarantees.

How is a lakehouse different from a data lake or warehouse?

A lake stored anything cheaply and offered no transactions or schema enforcement, so quality degraded until people stopped trusting it. A warehouse was reliable and governed and expensive at volume, and poorly suited to unstructured data. A lakehouse puts the warehouse's guarantees onto the lake's storage, so you keep one copy and serve both workload types. It solves storage and compute; it does not by itself make data meaningful to a machine.

Do you need a data lakehouse to run AI agents?

No, and believing otherwise has delayed a lot of useful work. Agents read from whatever holds the data, while consolidating an estate is a multi-year programme. If your live use cases need three systems, connecting to three systems is faster and more reversible than centralising first. Where a lakehouse already exists the work is exposure rather than migration: make its permissions, versions, lineage and freshness reachable from the agent's query path.

What does an agent need from a lakehouse that a dashboard does not?

Governance metadata readable at query time rather than consulted beforehand. A dashboard's permissions are applied once by whoever built the report; an agent retrieves what it can reach and then acts, so the constraint has to sit in the query path. Similarly, history has to be queryable as of a date, lineage has to travel with the retrieved fact, and freshness has to be expressed as data because an agent has no intuition that a figure looks old.

What is time travel in a lakehouse?

The ability to query a table as it stood at a past point. Every write to an open table format produces a new version and leaves the previous ones addressable, so you can read the table as of a date or a version number rather than only as it is now. It is native to the format, already paid for, and rarely wired to agents despite answering the question that every audit and dispute begins with.

Why does time travel matter for AI agents?

Because it answers as-of questions and makes agent runs reproducible. What did this record say on the date the decision was taken is where every dispute starts, and without versioning the honest answer is that you cannot reconstruct it. If a trace also records which table version the agent read, you can replay against exactly that data rather than against a table that has since changed, which turns an argument about what happened into a check.

How long is lakehouse table history retained?

Not indefinitely, and the default is usually far shorter than any audit obligation. Maintenance operations that reclaim storage remove versions past a configured window. If you intend to rely on versioning as evidence rather than as a convenience, that retention setting becomes a governance decision rather than a storage one, and it should be set against the retention periods your regulatory regimes actually require rather than against the platform default.

Is a lakehouse fast enough for agentic retrieval?

Frequently not, at least not as the surface the agent queries directly, and the mismatch is architectural rather than a defect. Scanning wide swathes of data efficiently is what these platforms are tuned for, while an agent fires off many narrow interactive reads inside a loop with somebody waiting on the other end. Cold reads against infrequently touched object storage can take seconds, which is fine in a nightly job and not fine inside a retrieval loop running several rounds.

How should you serve agents if the lakehouse is too slow?

Stop treating it as the agent's direct query surface without abandoning it. Serve agents from a layer built for selective low-latency reads, keep the lakehouse as the governed source of record behind that layer, and carry the governance metadata forward rather than leaving it at the boundary. The failure mode to avoid is a fast serving layer that has dropped the permissions, lineage and version information that made the lakehouse worth having.

Does a lakehouse solve data quality for AI?

It solves a category of it and not the category that usually breaks agents. Transactions and schema enforcement stop the structural degradation that made plain data lakes unusable. They do not make a field mean the same thing in two systems, and they do not record which definition of a term a given table uses. If your agents produce confident errors, ambiguous definitions are a more likely cause than distributed storage.

What are open table formats?

The layer that makes a lakehouse possible. They maintain a transaction log alongside data files in object storage, which is what provides atomic writes, schema enforcement and version history over what would otherwise be a directory of files. Being open matters commercially: the table format determines how portable your data is between engines, so it is one of the more consequential choices in a platform decision and one often made by default.

Should you migrate to a lakehouse before starting with agents?

Almost never in that order. A migration is a multi-year programme with its own risk, and agents can read from the sources you already have while it proceeds. The sequencing that works is to connect to the short list of systems your live use cases actually read, address meaning and permissions at the connector boundary, and let any consolidation happen on its own timeline for its own reasons rather than as a prerequisite.

The governance layer, not the storage layer
You already paid for time travel. Is it wired to anything?

SERAA Axon is the reasoning layer above the lake rather than a replacement for it, supporting Databricks data-readiness patterns as a deployment pattern and reading in place across 100+ connectors, so a governed lakehouse becomes something an agent can query with permissions, history and lineage intact.