All glossary terms
E Data & reasoning Operating model

Enterprise data management

Every enterprise data platform was designed on one assumption: that you know who will consume the data and what they will ask. Agents are an unknown number of unknown consumers asking unanticipated questions, which is the assumption breaking rather than the platform failing.

Definition

Enterprise data management is the discipline and the systems that make an organisation's data available, correct, governed and usable across the business. Its design assumption is that data consumers are known in advance and can be modelled for, which is the assumption agentic systems break. Note that the acronym EDMS more commonly denotes electronic document management, so the spelled-out term is safer in writing.

What is enterprise data management?

The discipline of making an organisation's data available, correct, governed and usable across the business, together with the systems that deliver it. It is an umbrella rather than a product, which is why it is defined differently by every vendor who sells a piece of it.

One naming point first, because it causes real procurement errors. The acronym EDMS almost always means electronic document management in the wider market: document storage, version control, retrieval, the category served by long-established records management vendors. Enterprise data management is a different discipline entirely, and using the acronym invites the confusion. Spell it out.

What distinguishes the discipline from any single component is scope. It spans how data arrives, where it lives, what it means, who may see it, and whether it is trustworthy, across systems that were bought at different times for different reasons by people who no longer work there.

What does it actually cover?

There are six separate things you need to worry about here, and each one has its own set of tools and its own specialists who focus on just that piece. It matters that we name them separately, because most organizations end up good at two or three of them. Then it's easy to assume that being strong in a couple of areas means the whole thing is covered, when really the other pieces are just being ignored.

The map below doubles as a route into the specifics, since each concern has a page of its own in this glossary.

Six concerns of enterprise data management arranged as a stack: arrival covering ingestion and connectors, storage covering warehouse and lakehouse, meaning covering semantics and definitions, entity resolution covering master data, access covering permissions and lineage, and trust covering quality and freshness. Each is annotated with the question it answers, and a note states that most organizations are strong on arrival and storage and weak on meaning and trust, which is the opposite of what agents need.
Six concerns within enterprise data management, the question each answers, and where each is covered
Concern The question it answers Covered at
Arrival How does data get here, how often, and what happens when it is deleted at source Data ingestion, enterprise connectors
Storage Where does it live, with what guarantees, at what cost Data lakehouse
Meaning What does each field actually mean, and does it mean the same thing everywhere Data intelligence
Resolution Are these two records the same customer Master data management
Access Who may see which record, and where did this value come from Agent identity for the enforcement side
Trust Is it correct, is it current, and would anyone notice if it were not Context graph for validity as a property

The pattern across most estates is consistent and worth checking against your own. Arrival and storage tend to be well handled, because they are engineering problems with vendors attached. Meaning and trust tend not to be, because they require somebody to make decisions rather than buy something. That is the opposite of the order agents need them in.

Why did its core assumption break?

The whole discipline was built around knowing who your consumers were. You'd gather requirements, design the data around the reports and applications that would read it, and curate it with that in mind. Agents break that pattern completely, because they're an unknown number of unknown consumers, asking questions that nobody ever gathered requirements for.

This is a genuine break rather than a shortcoming, and it explains why a mature data platform can still fail an agentic programme.

Two models compared. On the left, a small fixed set of known consumers: four dashboards, two applications and a reporting team, each with gathered requirements and a curated view built for it. On the right, an unbounded set of agent consumers asking arbitrary questions, with no requirements gathering possible, shown reading the same underlying data directly. A caption states that you can curate for known queries and cannot curate for arbitrary ones, so data has to describe itself instead.
  • You cannot gather requirements from an agent. The traditional cycle asks the consumer what they need, then models for it. An agent's questions emerge from whatever a user asks it, which is not a stable set and cannot be enumerated in advance.
  • Curated views stop being the answer. A view built for a known report is exactly right for that report and arbitrary for anything else. An agent given a curated view inherits somebody else's assumptions about which columns matter.
  • Documentation written for people fails a machine. A data dictionary saying a field is "the standard customer status" works for an analyst who can ask a colleague. An agent reads it literally and acts.
  • The volume of distinct queries goes up sharply. A reporting estate issues a bounded set of queries repeatedly. An agentic estate issues a wide set once each, which changes both the performance profile and what has to be documented.

The consequence is a single shift in approach. You can no longer make data usable by anticipating the question. You have to make it self-describing, so that a consumer you never met can establish what a field means, how current it is, and whether it may rely on it. That is the whole subject of data intelligence, arrived at here from the systems side rather than the semantic one.

What do agents need that reports did not?

Four properties, and a reporting estate had no reason to provide any of them. Meaning attached to the data rather than held in a wiki, permissions enforceable at query time rather than at build time, freshness expressed as a value, and provenance travelling alongside the record itself.

Each has the same character: something a person supplied from context now has to be present in the data.

  1. 01
    Meaning, attached rather than documentedAn analyst reads a column name and knows what it means from experience. An agent has only what is recorded. The test: could a competent outsider, given only your metadata, use this field correctly on their first day.
  2. 02
    Permissions enforceable in the query pathA dashboard filters what it displays. An agent retrieves what it can reach, then acts. The test: query as two users with different entitlements and see whether the results differ.
  3. 03
    Freshness as a value, not an assumptionA reader knows the report runs nightly. An agent has no such intuition and treats a stale figure as current. The test: can a consumer read how old this record is without asking anybody.
  4. 04
    Provenance travelling with the recordLineage in a catalog is consulted when a number looks wrong. An agent needs the source attached to the fact it retrieved. The test: does an answer carry where each element came from, or only the answer.

Notice that all four are things a person supplied silently. The reporting estate was never incomplete; it was complete for a consumer who brought context with them. Agents bring none.

Do you need one before running agents?

No, and treating it as a prerequisite has delayed a great deal of useful work. A data management programme is a multi-year undertaking. Agents read from whatever holds the data, so the sequencing that works is narrow and specific rather than total and deferred.

Three positions, and the third is the one worth arguing for.

  • Fix everything first. Defensible on paper and rarely survives contact with a budget cycle. By the time the programme lands, the use cases that justified it have moved on.
  • Fix nothing and hope. Produces agents that give confident answers from ambiguous data, which is worse than no agent because the output is plausible enough to act on.
  • Fix the sources your live use cases actually read. Usually a handful of tables in two or three systems. Establish what each field means, whether permissions survive the read, and how fresh it is. Narrow, quick, and it delivers the property agents need without the programme.

The third option also produces a better business case for the wider work, because you arrive at the funding conversation with a specific list of what broke and what it cost, rather than with a maturity model.

How do you assess what you already have?

By testing it as an agent would rather than by scoring it against a framework. Take one real question an agent will be asked, follow it to the data, and see whether a consumer with no institutional knowledge could answer it correctly from what is recorded.

Four checks, each of which takes an afternoon and tells you more than a maturity assessment.

  1. The outsider test on one table

    Take a table your live use case reads. Hand its metadata to somebody competent who has never worked with it, and ask them what three specific fields mean and when they were last correct. What they cannot answer is what an agent cannot answer either.

  2. The two-user query

    Run the same retrieval as two people with different entitlements. If the results are identical, your access control is applied after retrieval rather than during it, which means the agent's reach is bounded by a filter somebody can misconfigure.

  3. The deletion trace

    Delete a test record at source and time how long it takes to disappear from everywhere downstream. If the answer is never, you are holding content the source system removed, which is covered in more detail at enterprise connectors.

  4. The two-system definition check

    Pick a term your business argues about, such as active customer or recognised revenue, and find how it is defined in two systems that both claim to hold it. A discrepancy here is not a data quality issue; it is the reason an agent will produce two defensible and contradictory answers.

One closing point about the discipline itself. Enterprise data management was never a technology problem and agents have not made it one. The four checks above mostly surface decisions nobody has made rather than systems nobody has bought, and that is why the work resists being solved by a purchase.

Frequently asked questions about enterprise data management

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is enterprise data management?

An umbrella covering both the practice and the machinery: getting data where it needs to be, keeping it right, controlling who touches it, and making it fit to use. Because it is not one product, each vendor with a stake in one part tends to describe the whole in terms of that part. What distinguishes it from any single component is scope: how data arrives, where it lives, what it means, who may see it, and whether it can be trusted, across systems bought at different times for different reasons.

Does EDMS mean enterprise data management?

Usually not, and this causes real procurement errors. Across most of the market EDMS denotes electronic document management: storage, version control and retrieval of documents, a category served by long-established records management vendors. Enterprise data management is a different discipline concerned with data rather than documents. If you mean the data discipline, spell it out, because the acronym will be read the other way by most people and most search engines.

What does enterprise data management cover?

Six concerns, each with its own tooling market: how data arrives, where it is stored, what each field means, whether two records refer to the same entity, who may access what, and whether the data can be trusted. Most estates handle arrival and storage well, because those are engineering problems with vendors attached, and handle meaning and trust poorly, because those require somebody to make decisions rather than buy something.

Why do agents break enterprise data management?

Because it presumed you knew your audience. The working method was to interview the people who would read the data, build to what they described, and shape the tables around those answers. That method has no equivalent for an agent: there is nobody to interview and no stable list of questions, so a purpose-built view is guesswork and a data dictionary written for an analyst gets read literally by something that then goes and acts.

What does self-describing data mean?

Data that carries what a consumer needs in order to use it correctly, rather than relying on the consumer already knowing. Since you cannot anticipate what an agent will ask, you cannot curate a view for it, so the alternative is making each field explain what it means, how current it is, and who may rely on it. That is the shift from anticipating questions to describing the data, and it is the practical core of the change.

What do agents need that reports did not?

Four properties, all of them things a person previously supplied silently. Meaning attached to the data rather than held in a wiki or somebody's head. Permissions enforceable at query time rather than applied when a dashboard was built. Freshness expressed as a value, because an agent has no intuition that a figure looks old. And provenance travelling with the record rather than sitting in a catalog somebody consults when a number looks wrong.

Do you need to fix your data before deploying agents?

Not all of it, and treating that as a prerequisite has delayed a great deal of useful work. Fixing everything first rarely survives a budget cycle, and fixing nothing produces agents that give confident answers from ambiguous data. The approach that works is narrow: establish meaning, permissions and freshness for the handful of tables your live use cases actually read, which is usually two or three systems rather than the estate.

How do you test whether your data is agent-ready?

Four checks, each taking about an afternoon. Hand one table's metadata to a competent outsider and ask what three fields mean and when they were last correct. Run the same retrieval as two users with different entitlements and see whether results differ. Delete a test record at source and time its disappearance downstream. And find how a contested term such as active customer is defined in two systems that both claim to hold it.

What is the outsider test?

Giving your metadata to somebody competent who has never worked with the data, and asking them to explain specific fields and their currency. Whatever they cannot answer is precisely what an agent cannot answer either, because an agent has no colleague to ask and no institutional memory to draw on. It is the cheapest diagnostic available and it usually surfaces more than a formal maturity assessment.

Why does a definition mismatch matter more than data quality?

Because it produces two defensible and contradictory answers rather than one wrong one. If active customer means something different in two systems that both claim to hold it, an agent querying either will be correct according to that system and inconsistent with the other. No amount of cleansing resolves it, because nothing is dirty. Somebody has to decide which definition governs, which is a governance act rather than an engineering one.

Is enterprise data management a technology problem?

It never was, and agents have not changed that. The diagnostic checks that matter mostly surface decisions nobody has made rather than systems nobody has bought: which definition governs, who is entitled to which records, how current is current enough. That is why the discipline resists being solved by a purchase, and why a programme framed as a platform selection tends to deliver infrastructure without resolving the underlying ambiguity.

Where does master data management fit?

It answers one of the six concerns: whether two records refer to the same entity. That question becomes sharper with agents, because an agent acting on an unresolved entity may act twice, or act on the wrong record, without anything in its trace looking wrong. Where a reporting estate tolerated duplicate customers as a known annoyance in a dashboard, an agent turns the same ambiguity into a duplicated action.

Start with the tables you actually read
Could an outsider use your data correctly on day one?

SERAA Axon connects across ERP, CRM, HRMS and unstructured sources and builds the relationships as data arrives, without moving everything into a lake first. Bring the handful of systems your live use cases read and we will run the four checks with you.