What is enterprise data management?
The discipline of making an organisation's data available, correct, governed and usable across the business, together with the systems that deliver it. It is an umbrella rather than a product, which is why it is defined differently by every vendor who sells a piece of it.
One naming point first, because it causes real procurement errors. The acronym EDMS almost always means electronic document management in the wider market: document storage, version control, retrieval, the category served by long-established records management vendors. Enterprise data management is a different discipline entirely, and using the acronym invites the confusion. Spell it out.
What distinguishes the discipline from any single component is scope. It spans how data arrives, where it lives, what it means, who may see it, and whether it is trustworthy, across systems that were bought at different times for different reasons by people who no longer work there.
What does it actually cover?
There are six separate things you need to worry about here, and each one has its own set of tools and its own specialists who focus on just that piece. It matters that we name them separately, because most organizations end up good at two or three of them. Then it's easy to assume that being strong in a couple of areas means the whole thing is covered, when really the other pieces are just being ignored.
The map below doubles as a route into the specifics, since each concern has a page of its own in this glossary.
| Concern | The question it answers | Covered at |
|---|---|---|
| Arrival | How does data get here, how often, and what happens when it is deleted at source | Data ingestion, enterprise connectors |
| Storage | Where does it live, with what guarantees, at what cost | Data lakehouse |
| Meaning | What does each field actually mean, and does it mean the same thing everywhere | Data intelligence |
| Resolution | Are these two records the same customer | Master data management |
| Access | Who may see which record, and where did this value come from | Agent identity for the enforcement side |
| Trust | Is it correct, is it current, and would anyone notice if it were not | Context graph for validity as a property |
The pattern across most estates is consistent and worth checking against your own. Arrival and storage tend to be well handled, because they are engineering problems with vendors attached. Meaning and trust tend not to be, because they require somebody to make decisions rather than buy something. That is the opposite of the order agents need them in.
Why did its core assumption break?
The whole discipline was built around knowing who your consumers were. You'd gather requirements, design the data around the reports and applications that would read it, and curate it with that in mind. Agents break that pattern completely, because they're an unknown number of unknown consumers, asking questions that nobody ever gathered requirements for.
This is a genuine break rather than a shortcoming, and it explains why a mature data platform can still fail an agentic programme.
- You cannot gather requirements from an agent. The traditional cycle asks the consumer what they need, then models for it. An agent's questions emerge from whatever a user asks it, which is not a stable set and cannot be enumerated in advance.
- Curated views stop being the answer. A view built for a known report is exactly right for that report and arbitrary for anything else. An agent given a curated view inherits somebody else's assumptions about which columns matter.
- Documentation written for people fails a machine. A data dictionary saying a field is "the standard customer status" works for an analyst who can ask a colleague. An agent reads it literally and acts.
- The volume of distinct queries goes up sharply. A reporting estate issues a bounded set of queries repeatedly. An agentic estate issues a wide set once each, which changes both the performance profile and what has to be documented.
The consequence is a single shift in approach. You can no longer make data usable by anticipating the question. You have to make it self-describing, so that a consumer you never met can establish what a field means, how current it is, and whether it may rely on it. That is the whole subject of data intelligence, arrived at here from the systems side rather than the semantic one.
What do agents need that reports did not?
Four properties, and a reporting estate had no reason to provide any of them. Meaning attached to the data rather than held in a wiki, permissions enforceable at query time rather than at build time, freshness expressed as a value, and provenance travelling alongside the record itself.
Each has the same character: something a person supplied from context now has to be present in the data.
- 01
Meaning, attached rather than documentedAn analyst reads a column name and knows what it means from experience. An agent has only what is recorded. The test: could a competent outsider, given only your metadata, use this field correctly on their first day.
- 02
Permissions enforceable in the query pathA dashboard filters what it displays. An agent retrieves what it can reach, then acts. The test: query as two users with different entitlements and see whether the results differ.
- 03
Freshness as a value, not an assumptionA reader knows the report runs nightly. An agent has no such intuition and treats a stale figure as current. The test: can a consumer read how old this record is without asking anybody.
- 04
Provenance travelling with the recordLineage in a catalog is consulted when a number looks wrong. An agent needs the source attached to the fact it retrieved. The test: does an answer carry where each element came from, or only the answer.
Notice that all four are things a person supplied silently. The reporting estate was never incomplete; it was complete for a consumer who brought context with them. Agents bring none.
Do you need one before running agents?
No, and treating it as a prerequisite has delayed a great deal of useful work. A data management programme is a multi-year undertaking. Agents read from whatever holds the data, so the sequencing that works is narrow and specific rather than total and deferred.
Three positions, and the third is the one worth arguing for.
- Fix everything first. Defensible on paper and rarely survives contact with a budget cycle. By the time the programme lands, the use cases that justified it have moved on.
- Fix nothing and hope. Produces agents that give confident answers from ambiguous data, which is worse than no agent because the output is plausible enough to act on.
- Fix the sources your live use cases actually read. Usually a handful of tables in two or three systems. Establish what each field means, whether permissions survive the read, and how fresh it is. Narrow, quick, and it delivers the property agents need without the programme.
The third option also produces a better business case for the wider work, because you arrive at the funding conversation with a specific list of what broke and what it cost, rather than with a maturity model.
How do you assess what you already have?
By testing it as an agent would rather than by scoring it against a framework. Take one real question an agent will be asked, follow it to the data, and see whether a consumer with no institutional knowledge could answer it correctly from what is recorded.
Four checks, each of which takes an afternoon and tells you more than a maturity assessment.
-
The outsider test on one table
Take a table your live use case reads. Hand its metadata to somebody competent who has never worked with it, and ask them what three specific fields mean and when they were last correct. What they cannot answer is what an agent cannot answer either.
-
The two-user query
Run the same retrieval as two people with different entitlements. If the results are identical, your access control is applied after retrieval rather than during it, which means the agent's reach is bounded by a filter somebody can misconfigure.
-
The deletion trace
Delete a test record at source and time how long it takes to disappear from everywhere downstream. If the answer is never, you are holding content the source system removed, which is covered in more detail at enterprise connectors.
-
The two-system definition check
Pick a term your business argues about, such as active customer or recognised revenue, and find how it is defined in two systems that both claim to hold it. A discrepancy here is not a data quality issue; it is the reason an agent will produce two defensible and contradictory answers.
One closing point about the discipline itself. Enterprise data management was never a technology problem and agents have not made it one. The four checks above mostly surface decisions nobody has made rather than systems nobody has bought, and that is why the work resists being solved by a purchase.