What is a data lakehouse?
A data lakehouse stores all your data cheaply, like a lake, yet keeps it as tidy and reliable as a warehouse. A simple record book tracks every change, so nothing gets lost or overwritten by mistake. Databricks coined the term in 2020 and owns the category.
Attribution matters here. Lakehouse is a coined category with an identifiable origin, and treating it as a generic architecture term rather than crediting where it came from is the kind of thing careful buyers notice.
The mechanism is a transaction log sitting alongside the data files, which is what turns a directory of Parquet into a table with guarantees.
- Transactions. Writes either complete or do not, so readers never see a half-written table. This is the property a plain data lake lacked and the reason lakes so often became unusable.
- Schema enforcement and evolution. The table has a declared shape, and changing it is a recorded operation rather than an accident discovered downstream.
- Versioning. Every write produces a new version, with the old ones still addressable. This is the property this page argues you are underusing.
- Governance above the storage. A catalog holding permissions, lineage and descriptions, separate from the files themselves.
What problem has Data Lakehouse solved?
The data lakehouse removed the need to run two systems. Companies once stored raw data in a cheap lake and tidy data in a costly warehouse. Teams copied data between them and fixed mismatches. The lakehouse keeps everything in one place, with one copy.
The problem it solved is real and the solution has largely worked, which is worth saying plainly before arguing about what it did not solve.
| Architecture | What it was good at | What it could not do |
|---|---|---|
| Data lake | Cheap storage of anything, including unstructured and semi-structured data | No transactions or schema enforcement, so quality degraded until nobody trusted it |
| Warehouse | Reliable, governed, fast on structured analytics | Expensive at volume, and poorly suited to unstructured data or machine learning workloads |
| Lakehouse | Warehouse guarantees on lake storage, one copy, both workload types | Solves storage and compute. Does not by itself make the data meaningful to a machine |
That last cell is the whole argument of this page. Unifying storage was a genuine achievement and it addressed an infrastructure problem. An agent's difficulty is rarely that the data lives in two systems. It is that the data does not explain itself, which is the separate subject covered at data intelligence.
What does an agent need that a dashboard does not?
Governance metadata it can read and act on at query time, rather than metadata a person consults beforehand. A dashboard's permissions are applied by whoever built it. An agent retrieves what it can reach and then acts, so the constraint has to be readable in the query path.
Four properties, and the useful observation is that a lakehouse holds all four and typically exposes none of them to the agent.
| Property | For a dashboard | For an agent |
|---|---|---|
| Permissions | Applied once by whoever built the report | Enforced in the query path, per requester, because the agent chooses what to read |
| History | Rarely needed. The current state is the state | Queried as of a date, so a fact can be checked for whether it still holds |
| Lineage | Consulted manually when a number looks wrong | Attached to the retrieved fact, so provenance travels into the answer |
| Freshness | Known from context. The reader knows it runs nightly | Expressed as data, because an agent has no intuition that a figure looks old |
Reading the right-hand column back, it describes something close to a context graph: facts carrying the conditions under which they hold. The lakehouse is not that, and it holds most of the raw material for it.
Which capability do you already have and not use?
In a data lakehouse, time travel shows your data as it stood on any past date. Every change saves a new version and keeps the old ones. Few firms connect this to their AI agents. It answers the question every audit asks: what did the data say then?
This is the strongest practical point on the page and it costs nothing to adopt, because the capability is already paid for.
What it enables, in order of how often it comes up.
- Answering as-of questions. What did this record say on the date the decision was taken. Every dispute, audit and complaint starts here, and without versioning the honest answer is that you cannot reconstruct it.
- Reproducing an agent run. If a trace records which table version the agent read, you can replay against exactly that data rather than against a table that has since changed. This turns an argument about what happened into a check.
- Distinguishing the two clocks. When a fact became true and when your system learned it are different, and versioning gives you the second for free. The first still has to be modelled deliberately.
- Detecting silent change. A column whose meaning shifted between versions is visible in the history and invisible in the current state.
One retention caution before you rely on this. Table history is not kept indefinitely. Maintenance operations that reclaim storage will remove old versions past a configured window, and the default window is usually far shorter than any audit retention obligation. If you intend to use versioning as evidence rather than as a convenience, the retention setting is a governance decision rather than a storage one. See AI audit trail for the retention periods different regimes actually require.
Where does the lakehouse fight the agentic loop?
On latency and on access pattern. A lakehouse is optimised for scanning large volumes efficiently, which is the opposite of what an agent does: many small, selective, interactive reads inside a loop where every round trip adds to a response time somebody is waiting for.
This is an architectural mismatch rather than a flaw, and it is manageable if you plan for it rather than discover it.
- Cold start on object storage. First reads against infrequently touched data can take seconds. Acceptable in a nightly job, unacceptable inside a retrieval loop that may run several rounds per question.
- The loop multiplies the cost. An agentic retrieval pattern issues multiple queries per question, so a query time that is fine once becomes the dominant latency when it happens four times. See agentic RAG.
- Batch freshness meets interactive expectation. A table refreshed hourly is entirely reasonable, and an agent asked a question about the last twenty minutes has no way to know it is looking at a stale view unless freshness is recorded as data.
- Compute economics invert. Elastic compute priced for large scheduled jobs behaves differently under many small concurrent interactive queries, and the bill is the place most teams discover this.
The workable pattern is not to abandon the lakehouse but to stop treating it as the agent's direct query surface. Serve agents from a layer built for selective low-latency reads, keep the lakehouse as the governed source of record behind it, and carry the governance metadata forward rather than leaving it behind at the boundary.
Do you need one before running agents?
No, and the belief that you do has delayed a lot of useful work. Agents read from whatever holds the data, and consolidating an estate is a multi-year programme. If your agents need three systems, connecting to three systems is a faster and more reversible path than centralising first.
Two situations, and they call for different things.
-
If you do not have a lakehouse
Do not build one in order to start. Connect agents to the sources that hold the data your live use cases actually read, which is usually a short list, and address meaning and permissions at the connector boundary. See enterprise connectors for what to demand of those integrations, particularly deletion propagation and permission capture.
-
If you already have one
The work is not migration, it is exposure. Make the catalog's permissions enforceable in the agent's query path, surface table versions so as-of questions are answerable, carry lineage through to the retrieved fact, and record freshness as data rather than as an assumption. All four already exist in the platform and typically stop at the analytics boundary.
-
Either way, fix meaning before architecture
A consolidated platform does not make a field mean the same thing in two systems. If your agents produce confident errors, the cause is more often ambiguous definitions than distributed storage, and no migration resolves that.
The honest summary: a lakehouse is a good answer to a question about analytics infrastructure and a partial answer to a question about agents. Where one already exists, the governance layer above it is worth considerably more to an agentic programme than the storage layer beneath it, and it is the part most organizations have never exposed.