All glossary terms
D Data & reasoning Measurement

Data quality for AI

A 2% error rate is fine on a dashboard, because the person reading it spots the outlier. An agent acting on the same data acts wrongly 2% of the time and nothing flags it. The dimensions did not change. The tolerance collapsed.

Definition

Data quality for AI is fitness for machine consumption rather than human interpretation. The traditional dimensions still apply and their thresholds do not, because analytics quality is statistical while agentic quality is per instance: a dashboard is judged on the average row and an agent acts on one particular row.

What is data quality for AI?

Data quality for AI is fitness for a consumer that cannot sanity-check. Quality programmes assumed a person would notice a wrong number and question it. An AI agent does not. It treats every value as true and acts on it, so the same data becomes riskier.

This is not a new set of dimensions. It is the same set with the tolerances rewritten, and the reason is a shift in how errors get consumed.

The same dataset consumed two ways. On the left, an analytics consumer aggregates a thousand rows into one chart, where twenty incorrect rows change the total by a fraction of a percent and a human notices any visible outlier, marked as statistically acceptable. On the right, an agentic consumer reads one row at a time and acts on it, where the twenty incorrect rows become twenty wrong actions with no error raised and nothing to notice, marked as per instance.

The distinction is worth stating precisely because it changes what you measure. Analytics consumes data in aggregate, so quality is a property of the distribution. An agent consumes one record at a time, so quality is a property of that record. A dataset that is 98% correct is excellent for reporting and is a system that takes a wrong action every fiftieth time it runs.

Why do the classic dimensions need redefining?

Because every one of them was written with a human reader in mind. Completeness meant the field had something in it, consistency meant the formats lined up, timeliness meant the refresh ran on schedule. All three quietly assume a person who will look at the result and interpret it, and an agent interprets nothing.

The same six dimensions, re-read for a machine consumer.

The classic data quality dimensions and what each has to mean for an agentic consumer
Dimension What it meant for reporting What it has to mean now
Completeness The field is populated Absence is distinguishable from zero. A null rendered as 0 makes an agent compute confidently and wrongly
Consistency The formats match across systems The meanings match. Two systems agreeing on the format of active customer while disagreeing on the definition is the harder failure
Timeliness Refreshed on the agreed schedule The age is readable as a value, because an agent has no sense that a figure looks old
Validity Conforms to the expected type and range Unchanged, and now worth enforcing at read rather than trusting at write
Uniqueness Duplicates are deduplicated for counting An unresolved duplicate becomes a duplicated action, not an inflated total
Accuracy Close enough on average Correct on this record, because this record is the one being acted on

The two marked rows cause the most damage in practice and neither is usually monitored. A null that arrives as a zero and a field whose meaning differs between systems both pass every conventional quality check, because nothing about either is malformed. The meaning problem is covered at length at data intelligence.

Does automated data processing improve data quality?

It improves throughput and leaves quality where it was. Automated data processing moves, transforms and loads data faster and more reliably than people did, which is genuinely valuable. What it does not do is decide what correct means, and that decision is where quality actually lives.

The distinction matters because pipeline investment is often presented as quality investment, and they are different budgets solving different problems.

  • What automation does fix. Transcription errors, missed runs, inconsistent manual steps, and the delay between something changing and the change arriving. Real problems, genuinely solved.
  • What it cannot fix. Whether the source value was right, whether two systems mean the same thing by a field, whether a rule encoded in 2019 still reflects how the business works. No pipeline resolves a definitional disagreement.
  • What it makes worse. Scale. Processing ambiguous data faster produces ambiguity faster, and a well-run pipeline will propagate a wrong definition to every downstream consumer with perfect reliability.

The trap in the phrase itself. Automated data processing describes the movement of data, and quality is a property of its meaning. A programme that buys the first and reports against the second will deliver a faster pipeline, an unchanged error rate, and a governance conversation nobody expected. If the business case says quality, the measurements need to be quality measurements rather than throughput and uptime.

Why is sampling the wrong instrument?

Because it answers a question about the population when the risk lives in the tail. Profiling a sample tells you roughly what proportion of records are sound, which is exactly what a reporting consumer needs and exactly what an agent's exposure does not depend on.

An agent does not encounter your average record. It encounters the specific record attached to the specific request, and the requests that reach an agent are frequently the awkward ones a person could not resolve.

  1. 01
    The tail is the workload, not the exceptionStraightforward cases were automated years ago. What arrives at an agent skews toward the unusual. The consequence: a quality score computed on a representative sample systematically understates the error rate in production.
  2. 02
    Errors are correlated, not scatteredBad records cluster by source, by period, by the migration that went sideways in March. The consequence: a random sample either misses a cluster entirely or over-weights it, and neither result is useful.
  3. 03
    The measure has to follow the queryWhat matters is the quality of the records your live use cases actually read, which is a narrow slice. The consequence: profiling the estate is expensive and answers a question nobody asked, while profiling the slice is cheap and predicts production.

The practical replacement is to measure quality where retrieval happens rather than where data rests. Sample the records your agents actually touched last week, not the table they came from, and the number you get will bear some relationship to the failures you are seeing.

How do you deal with unstructured data?

Unstructured data is the gap most quality programmes cannot close. Conventional tools profile columns for null rates and data types. Those checks do not apply to contracts or email threads. Yet unstructured content is a large share of what agents read. A semantic backbone helps by linking each document to the business entity it describes.

There is no null rate for a PDF, so the measurement apparatus most organizations bought does not reach the data creating most of the risk.

  • Is it the current version. The single most consequential question, and the one least often answerable. A superseded policy document is indistinguishable from the current one once its text is in an index.
  • Is it authoritative or a draft. Shared drives are full of working copies, and nothing in the text says which is which.
  • Was it extracted correctly. A table misread during parsing produces plausible numbers in the wrong columns, which is worse than failing to parse at all.
  • Should this reader see it. Permission is a quality property once retrieval is automated, because nothing downstream will catch what should not have been returned.

Three of those four are metadata questions rather than content questions, which is the useful insight. Most unstructured quality is answerable at the boundary, when the document arrives, and unanswerable afterwards from the text alone. That is the argument set out at data ingestion, and it is why the connector properties at enterprise connectors turn out to be quality controls.

How do you measure Data Quality for agents?

By the outcome rather than the dataset. The question that matters is not what proportion of records are sound, it is how often an agent acted on something wrong, and that is measurable directly once traces record which records were read.

Five measures, and the first one replaces most of a conventional quality dashboard.

  1. Failures traced back to data, as a share of all failures

    When an agent gets something wrong, establish whether the cause was the model, the instructions or the data it read. The proportion attributable to data is the only quality metric that connects to something the business cares about.

  2. Quality of the slice your agents actually read

    Not the estate. Identify the tables and sources your live use cases touch, which is usually a short list, and profile those properly rather than everything thinly.

  3. Null-versus-zero conformance on numeric fields

    A specific and cheap check that catches a specific and expensive failure. Anywhere an absent value is being rendered as a zero, an agent will compute with confidence and be wrong.

  4. Definition conflicts between systems, counted

    Take the terms your business argues about and establish how many systems define each differently. That count is a governance metric rather than a data metric, and it predicts agent errors better than any completeness score.

  5. Freshness, expressed as data rather than assumed

    Record the age of what was retrieved on every retrieval. This makes staleness visible in the trace and lets an agent reason about it, instead of treating everything as current.

One closing point about where this work belongs. Most of what fails here is a decision nobody made rather than a system nobody bought. Which definition governs, how current is current enough, who may see what. Framing a quality programme as a tooling selection tends to deliver dashboards that measure the dimensions which were never the problem.

Frequently asked questions about data quality for AI

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is data quality for AI?

Fitness for a consumer that cannot sanity-check its inputs. Every quality programme ever built assumed a person at the end who would notice a number looked wrong and query it. Remove that person and the same data carries a different risk profile. The dimensions did not change; the tolerances did, because analytics judges the average row while an agent acts on one particular row.

Does automated data processing improve data quality?

Throughput improves. Quality stays broadly as it was. What automated data processing gives you is movement and transformation running faster and more dependably than any manual equivalent, which does genuinely eliminate mistyped values, skipped runs and lag. What it cannot fix is whether the source value was right, whether two systems mean the same thing by a field, or whether a rule encoded years ago still reflects the business.

Can a data pipeline make quality worse?

At scale, yes. Processing ambiguous data faster produces ambiguity faster, and a well-run pipeline propagates a wrong definition to every downstream consumer with perfect reliability. That is the trap in treating pipeline investment as quality investment: you get a faster pipeline, an unchanged error rate, and a governance conversation nobody planned for. If the business case says quality, the measurements have to be quality measurements rather than throughput and uptime.

Why is a 98% accurate dataset a problem for agents?

Because it means a wrong action every fiftieth run. Analytics consumes data in aggregate, so twenty bad rows in a thousand shift a total by a fraction of a percent and any visible outlier gets noticed. An agent consumes one record at a time and acts on it, so the same twenty rows become twenty wrong actions, with no error raised and nothing for anybody to spot.

How do the classic data quality dimensions change for AI?

Each was defined against a human reader. Completeness has to mean that absence is distinguishable from zero, since a null rendered as 0 makes an agent compute confidently and wrongly. Consistency has to mean matching meanings rather than matching formats. Timeliness has to mean the age is readable as a value. Uniqueness matters because an unresolved duplicate becomes a duplicated action rather than an inflated total.

Why does a null rendered as zero matter so much?

Because nothing about it is malformed, so it passes every conventional check. A missing value arriving as 0 is structurally valid, correctly typed and within range. A person reading a report would question a zero that looked odd; an agent takes it as a measurement and computes with it. It is the cheapest check on this page to implement and one of the most expensive failures to leave in place.

Why is sampling the wrong way to measure quality for agents?

Because it describes the population when the risk lives in the tail. Straightforward cases were automated years ago, so what reaches an agent skews toward the unusual, and a score computed on a representative sample systematically understates the production error rate. Bad records also cluster by source and by period rather than scattering, so a random sample either misses a cluster or over-weights it.

How do you measure the quality of unstructured data?

Mostly through metadata rather than content, which is the useful insight. There is no null rate for a PDF, so column profiling does not reach it. The questions that matter are whether it is the current version, whether it is authoritative or a draft, whether extraction read it correctly, and whether this reader should see it. Three of those four are answerable at the boundary and unanswerable afterwards from the text alone.

What is the most consequential unstructured quality question?

Whether the document is the current version, and it is the one least often answerable. A superseded policy is indistinguishable from the current one once its text sits in an index, so an agent retrieves and acts on whichever scores higher for relevance. Nothing in the content announces that it was replaced, which is why version and validity have to be captured when the document arrives.

What should you actually measure?

The outcome rather than the dataset. When an agent gets something wrong, establish whether the cause was the model, the instructions or the data, and track the proportion attributable to data. Then profile only the slice your live use cases read, check null-versus-zero conformance on numeric fields, count how many systems define each contested term differently, and record the age of what was retrieved on every retrieval.

Is data quality a technology problem?

Rarely, and framing it as one is why quality programmes produce dashboards measuring the dimensions that were never the problem. Most of what fails is a decision nobody made rather than a system nobody bought: which definition governs, how current is current enough, who may see what. Tooling reports the symptoms accurately and cannot resolve any of the three, because all three need somebody with authority.

Do you need to fix data quality before deploying agents?

Not across the estate, and treating that as a prerequisite delays useful work indefinitely. Fix the slice your live use cases read, which is usually a handful of tables in two or three systems. That is narrow enough to finish, it produces a measurable reduction in agent errors, and it gives you a specific list of what broke and what it cost when you come to fund the wider programme.

Measure the slice, not the estate
When an agent gets it wrong, can you tell whether the data caused it?

SERAA Axon reads in place across 100+ connectors with no migration, composes SQL, vector search and graph traversal into one chain, and leaves a reasoned audit trail at every decision node, so an answer carries which source each element came from.