What is a semantic layer?
The place where business logic is defined once rather than re-implemented per report. It holds metrics, the dimensions you slice them by, the joins between tables, and the hierarchies inside those dimensions, so every consumer asking for revenue receives the same number computed the same way.
The concept is old. Business intelligence tools have shipped versions of it for three decades under different names, which is why data platform owners recognise the idea immediately and often assume it is solved.
What changed is what a definition has to cover. A metrics layer states formulas. A semantic layer also models the relationships, the grain and the permitted join paths, which is the part that decides whether a query is even coherent. Modern implementations blur the two, and the distinction still matters when evaluating whether something will serve an agent or only a dashboard.
Four terms that get used interchangeably and should not be. An ontology is the schema of what exists. A semantic layer is a governed business abstraction over data, frequently with no graph beneath it. A knowledge graph is instantiated entities and relationships. A context graph adds the conditions under which each assertion holds. A vendor conversation that slides between them is worth slowing down.
Why does a Semantic layer matter more for agents than analysts?
Because an analyst who meets an ambiguous metric asks somebody. That single behaviour is what a semantic layer was historically a convenience for rather than a dependency on. Remove the person who can ask, and the layer becomes the thing deciding whether the answer is right.
A schema cannot tell a model whether refunded orders count toward revenue, whether last quarter means order date or settlement date, or which of five join paths avoids double counting. Those are institutional decisions that live in people's heads, in dashboard logic and in finance policy.
So an agent pointed at raw tables does what a new analyst would do without a colleague to ask: it picks an interpretation and computes confidently. The practitioner description of this is exact, and it is that the agent was free-handing SQL and inventing its own metric definition. The numbers look plausible. They do not match the dashboard, and nobody finds out until somebody compares.
What does the measured evidence show?
The evidence shows that the AI gives noticeably more correct answers when it uses a semantic layer, which is a shared dictionary spelling out what each business term means. But the more useful finding is how each approach goes wrong. Published tests report that when the semantic layer fails, it usually admits it cannot answer. When the AI writes its own database requests from scratch, it usually gives a wrong number with total confidence. An honest ‘I don't know’ is far safer than a confident mistake, and that difference matters more than any percentage.
Both halves are worth having in front of you, because the accuracy figures alone understate the case.
| Finding | What was measured |
|---|---|
| Failure mode differs | Semantic layer failures are typically refusals. Text-to-SQL failures are confident wrong numbers. The distinction is silent versus explicit |
| Accuracy rises | A 2026 paired benchmark reports models moving from around 90% on text-to-SQL to around 98%, and from 84% to 100%, on a well-modelled project |
| In-scope questions become deterministic | For questions falling inside what has been modelled, benchmark models returned correct results every time. The logic is codified, so the same question cannot drift between runs |
| Even a little context helps | Adding a small business-semantics document, on the order of a few kilobytes, lifted accuracy by roughly 17 to 23 percentage points across three frontier models |
| Coverage is the trade | A semantic layer can only answer what has been modelled, which is precisely why its failures are refusals rather than inventions |
The third row is the one to take to a risk committee. If the agent picks the right metric and dimensions, the query is guaranteed correct, because the layer generates the SQL rather than the model. That converts a probabilistic step into a deterministic one, which is a rare thing to be able to say about any part of an agentic system.
Read the accuracy figures with the usual caution. They come from specific benchmark suites on specific modelled projects, and the magnitude varies by dataset and implementation. The direction is consistent across every study; the number is not yours until you measure it.
How does an agent actually reach one?
Through an interface that is not a dashboard. A semantic layer locked inside a business intelligence tool is invisible to everything outside that tool, which is why the architecture matters more now than it did when the only consumer was a report.
The shift that makes this term newly relevant is headless delivery: definitions served over an API rather than embedded in a single application.
- Embedded. Definitions live inside the reporting tool. Excellent for that tool, unreachable by anything else, and the reason many organizations have a semantic layer that an agent cannot use.
- Headless. Definitions sit behind an API that any consumer queries, which is what makes one layer serve dashboards and agents from the same source rather than two definitions drifting apart.
- Warehouse-native. Semantic objects defined inside the data platform itself, a category that reached general availability during 2026. Attractive if your data is genuinely centralised, less so if it is not.
In practice an agent now commonly reaches governed metrics over the Model Context Protocol, listing what is available and requesting measures and dimensions by name rather than writing SQL against tables. That is a meaningful architectural point: the agent asks for a definition instead of reconstructing one, and access rules can be applied before the query runs rather than filtering results afterwards.
There is also movement toward portability. An open specification for expressing metric and semantic definitions in a vendor-neutral format entered the Apache Incubator after launching in late 2025, with a first specification version published in early 2026. It is early, there is no stable release yet, and it is worth tracking rather than depending on.
What can a semantic layer not do?
Rather more than vendor material suggests. It resolves what a metric means. It does not make an ambiguous question answerable, does not repair the underlying data, and does not stop a finance director maintaining a private definition of churn in a spreadsheet.
Being explicit about the boundary is what keeps the investment honest, because each limit maps to work somebody else has to do.
- It cannot fix source data. A governed definition computed over wrong values returns a wrong answer with excellent provenance. That work sits at data quality for AI.
- It does not carry validity or freshness. Whether a fact still holds, when it became true and who may rely on it are properties of the assertion, covered at context graph.
- It cannot cover everything, by design. Coverage is the trade for determinism. Questions outside the model get refused, which is the right behaviour and still a gap somebody has to fill.
- It does not end the definitional argument. It relocates it. Somebody still has to decide whose definition of active customer governs, and that decision has budget attached.
- It can become a second source of truth. If the layer drifts from the transformations beneath it, you now have two authoritative definitions and a harder problem than you started with.
The honest summary is that a semantic layer removes one large class of failure and leaves the rest. Regression tests on known-answer questions, provenance on every response, access control at query time and an escalation path for consequential answers are all still yours to build.
How do you build a Semantic layer without boiling the ocean?
Start small by defining only the numbers your AI is asked about, which is a much shorter list than every number your company tracks. The usual way this goes wrong is a team that sets out to define everything, takes eighteen months to finish, and ends up describing a business that has already changed by the time the work is done.
Five steps, ordered so the first one usually removes most of the scope.
-
Collect the questions, not the metrics
Take fifty questions people actually put to the agent or the analytics team. The metrics behind them are your scope. Everything else is optionality you are proposing to pay for.
-
Start with the contested ones
The terms your business argues about, such as active customer, recognised revenue or churn. Those carry the most risk of an agent picking the wrong interpretation, and defining them delivers value on the first day.
-
Model joins and grain, not only formulas
A formula without a declared join path leaves the model to choose one, which is where double counting arrives. This is the difference between a metrics list and a semantic layer.
-
Serve it headless from the start
Even if the only consumer today is a dashboard. Definitions reachable by an API can serve an agent later; definitions inside a reporting tool have to be rebuilt.
-
Write regression questions with known answers
A small set of questions with agreed correct answers, run whenever the layer changes. This is the cheapest guard against the layer silently drifting from the transformations beneath it.
Keep one idea in mind throughout the project. A semantic layer decides what counts as revenue, but it does not limit which questions people can ask about it. Good versions stay flexible when a question is asked. An AI agent can add its own filters and groupings on top of the agreed definitions, so it is not stuck with a fixed set of reports. If a proposal looks like it limits curiosity, then it has been designed the wrong way.