All glossary terms
A Responsible AI Risk

AI bias

AI bias is a repeatable pattern of unfair outcomes, not a one-off error. Fairness testing assumes one model making one decision, and an agent makes dozens. Every individual step can pass while the path as a whole does not.

Definition

AI bias is a systematic, repeatable pattern of unfair outcomes from an AI system, rather than a one-off error or evidence of intent. The term covers three different things in practice: statistical skew, discriminatory impact, and evaluator preference. Only the second is what regulators mean, and conflating them wastes most of the effort spent on the problem.

What is AI bias?

AI bias is when an AI system produces unfair outcomes in a consistent, repeatable way, not just occasionally, but repeatedly, in the same direction, against the same group.

The key word here is ‘systematic’. There is a real difference between two things that might look similar on the surface. A model that is occasionally wrong just has an error rate, mistakes happen randomly, spread across different people and situations. But a model that is consistently wrong in one direction about one group has a bias. It is not random anymore; it is a pattern.

Before anything else, three distinct things get called AI bias, and separating them is the single most useful move available. Teams routinely spend months on one while the risk sits in another.

Three columns distinguishing the senses of the word bias. Statistical bias is systematic deviation from a true value, which can be intentional and useful. Discriminatory bias is disparate outcomes for protected groups, which is what regulators and courts mean. Evaluator bias is a model judge preferring longer or earlier answers, which distorts your measurements rather than your decisions. A note reads that only the middle one carries legal exposure.
Three distinct meanings of AI bias, what each refers to, and who is concerned with it
Sense What it means Whose problem it is
Statistical Systematic deviation between a model's estimates and true values. Sometimes deliberate, since trading a little bias for less variance is standard practice Data science. Not inherently a problem at all
Discriminatory Outcomes that differ by protected characteristic without justification. Independent of intent and of whether the attribute was used directly Legal, risk, and the board. This is what regulators mean
Evaluator A model judge systematically preferring longer answers, earlier options, or output from its own family Whoever owns measurement. It distorts your metrics, not your customers

Bias and discrimination are not the same thing, and the academic literature is increasingly firm about it. A statistically biased model can produce entirely fair outcomes, and an unbiased one can discriminate. Treating the words as interchangeable makes the conversation harder, because an engineer hears a technical property and a lawyer hears a liability. When somebody says "the model is biased," the useful next question is which of the three they mean. Evaluator bias is covered at agent evaluation; the rest of this page is about the second column.

Where does bias enter a system?

At every stage, and rarely at the one people inspect. Problem framing, data collection, labelling, model design, deployment context, and post-launch feedback loops each introduce it independently and by different mechanisms, which is why auditing the model alone reliably finds only a fraction of it.

The mechanism worth understanding in detail is the one that defeats the most common mitigation.

  • Problem framing. What you chose to predict encodes a judgement. Optimising for a proxy that correlates with the outcome you care about imports whatever else that proxy correlates with.
  • Historical data. A model trained on past decisions learns the past, including the parts nobody would defend. This is the best-known source and the least tractable, because the data is what you have.
  • Labelling. Annotators bring their own judgement, and inconsistent labelling across groups produces disparity that looks like signal.
  • Optimising for aggregate accuracy. A model tuned to maximise overall accuracy will trade minority-group performance for majority-group performance if that raises the average, without anyone choosing that trade.
  • Feedback loops. A system whose outputs shape the data it later trains on compounds whatever skew it started with, which makes a small initial disparity grow rather than wash out.

Why removing the sensitive attribute does not work

The instinct is to drop race, gender or age from the inputs and declare the problem solved. It does not work, and understanding why matters. Models find proxies: postcode, vocabulary patterns, device type, education history, employment gaps. Any variable correlated with a protected attribute carries its signal, and a sufficiently capable model will use it because it improves prediction.

So "we do not collect that data" is not a fairness control. It is often a reason you cannot measure fairness at all, which the last section returns to.

Why can you not simply make a system fair?

Because fairness is not one property. Several precise and entirely reasonable definitions exist, and they conflict mathematically, so satisfying one of them can mean violating another. Choosing which to satisfy is a documented business and legal decision rather than a technical default.

This is the part most treatments skip, and it is the part that determines whether your fairness work survives scrutiny.

Common fairness definitions and what each requires
Definition What it requires, and what it costs
Demographic parity The same rate of positive outcomes across groups. Straightforward to explain, and it ignores whether the groups differ on the thing being predicted
Equal opportunity The same true positive rate across groups, so qualified people are treated alike. Says nothing about the false positive rate
Equalised odds The same true and false positive rates across groups. Stricter, and generally impossible to hold alongside calibration when base rates differ
Disparate impact Outcome ratios within a threshold. The framing courts and regulators tend to use, so it often wins regardless of what your data scientists prefer

The consequence is uncomfortable and worth stating plainly. There is no configuration that is fair in every sense at once, so any system in production embodies a choice among competing definitions. If nobody made that choice deliberately, it was made by whichever metric happened to be in the training objective.

The practical requirement that follows is documentation rather than perfection: name the fairness definition you are optimising for, record why it fits this decision and this population, and note what it trades away. That record is what makes the decision defensible later, and producing it is a governance activity rather than an engineering one. See AI governance.

What changes when an agent acts rather than advises?

Two things, and the second breaks the testing methodology. A biased recommendation becomes a biased action that executes without review. And fairness metrics assume one model producing one decision, while an agent produces a trajectory of small decisions, none of which is individually the decision being tested.

The first change is the obvious escalation. The second is the one nobody has a method for.

On the left, the classical fairness testing setup: one model, one decision, one measurable outcome, with disparate impact computable across groups. On the right, an agent trajectory of many small decisions including which document to retrieve, how to phrase a query, which tool to call, whether to escalate, and how much effort to spend, each marked as passing individually while the aggregate outcome differs by group. A caption reads that no single step is the decision, so there is nothing for the standard test to measure.

Classical fairness testing needs three things: a defined decision, a measurable outcome, and group membership. Given those, disparate impact is computable. An agent handling a case makes dozens of choices instead: which records to retrieve, how to phrase a lookup, which tool to call, how much effort to spend before giving up, whether to escalate to a human or resolve it alone.

None of those is the decision. Each looks defensible in isolation, each would pass a fairness test if you could even define one for it, and the disparity appears only in the aggregate: cases from one group escalated more often, resolved with fewer steps, or abandoned earlier. The trajectory discriminates while every step is clean.

So measure outcomes at the case level, not the step level. This is the one practical adaptation that works today. Stop trying to certify the model and start comparing what actually happened to comparable cases across groups: escalation rate, resolution time, number of steps, how often the agent gave up, and how often a human overrode it. Those are computable from your traces without any new methodology, and they catch the aggregate disparity that step-level testing cannot see. It requires per-step tracing to exist in the first place, which is covered at AI agent observability.

The volume point applies as well and needs no elaboration. A person applying a prejudice affects the cases in front of them. A system applying one repeats it at the speed and scale of the business, consistently, without the variation that makes human bias partially self-limiting.

Can bias hide in retrieval rather than in the model?

Yes, and almost nobody audits for it. If the retrieval layer surfaces different evidence depending on how a query is phrased, the model can be entirely fair and the system still produces disparate outcomes, because the two cases were never decided on the same information.

This vector is specific to retrieval-based and agentic systems, and it sits outside every fairness toolkit, all of which examine models.

  • Query phrasing varies with the person. Vocabulary, dialect, formality and detail differ across populations, and retrieval is sensitive to phrasing. Two materially identical requests can return different documents because they were worded differently.
  • The corpus is unevenly distributed. If your knowledge base documents some products, regions or customer segments more thoroughly than others, retrieval returns better evidence for the well-documented ones. The model then answers well for some cases and poorly for others, faithfully.
  • Relevance ranking is not neutral. Ranking optimises similarity, and whatever correlates with similarity in your corpus becomes a de facto weighting nobody chose.
  • The audit trail looks clean. Every step is recorded, every retrieval was relevant, every answer was grounded in a source. Nothing in the trace signals that a different case received better evidence.

The test is straightforward and worth running: take matched cases that differ only in phrasing or in attributes correlated with a protected characteristic, and compare what the retriever returned rather than what the model concluded. If the evidence differed, the disparity was created before the model ran. The upstream causes sit in data ingestion and context engineering.

How do you test and control for it?

Five practices, and one genuine obstacle that has no clean resolution. You cannot measure disparate impact without knowing group membership, and privacy obligations often mean you should not be holding that data. The tension is real rather than a matter of insufficient effort.

Take the obstacle first, because it shapes everything else.

The measurement paradox. Demonstrating that outcomes do not differ by protected characteristic requires knowing each person's protected characteristics. Data protection principles push toward not collecting them, and in several jurisdictions collecting them requires a lawful basis you may not have. So the obligation to demonstrate fairness and the obligation to minimise sensitive data collection point in opposite directions. Established approaches exist, including collecting the attributes separately under a narrow legal basis used only for aggregate testing, or using statistical inference at population rather than individual level. Both need legal sign-off. What does not work is assuming the problem away.

  1. Decide and record the fairness definition first

    Before any testing, name which definition applies to this decision and why, and note what it trades away. Doing this after a result arrives makes it look like the metric was chosen to suit the answer, which is the worst position to be in during scrutiny.

  2. Test outcomes at the case level

    Compare what happened to comparable cases across groups rather than certifying the model: escalation rate, resolution time, step count, abandonment, override frequency. This is the adaptation that works for agents, and it needs no new methodology.

  3. Audit the retriever, not only the model

    Run matched cases differing only in phrasing and compare the evidence returned. Fairness toolkits examine models and will not find this.

  4. Test with proxies present

    Removing sensitive attributes does not remove their signal, so test with the proxies your model actually sees. A model that behaves fairly on stripped data and unfairly on real data has been tested on a system you do not run.

  5. Monitor continuously, not at release

    Populations shift, data shifts, and feedback loops compound. A fairness assessment describes the day it was run, which for an agent whose behaviour changes without a deployment is a considerably shorter window than for a static model.

  6. Keep a human path for contested outcomes

    Not a fairness metric, and the most reliable control. Somebody affected by a decision needs a route to challenge it that does not depend on the system recognising its own error, and the existence of that route is increasingly a regulatory expectation rather than a courtesy.

Regulators around the world do not always agree on the details, but they all point in the same direction. The EU AI Act sorts AI systems by risk level and requires the highest-risk ones to follow strict data governance rules and undergo bias checks. Several countries now require companies to run bias audits for specific uses, like automated hiring decisions. And even where rules are not mandatory, many organizations follow voluntary frameworks like the NIST AI Risk Management Framework to show they took the problem seriously.

Frequently asked questions about AI bias

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is AI bias?

A repeatable, structural pattern of unfair outcomes, rather than a one-off error or evidence that anyone intended harm. The word structural carries the meaning: a model that is occasionally wrong has an error rate, while one that is consistently wrong in the same direction about the same group has a bias. Three different things get called AI bias, and separating them is the most useful first step, because teams often spend months on one while the exposure sits in another.

Is bias the same as discrimination?

No, and the distinction matters more than it sounds. Statistical bias is systematic deviation from a true value, and it can be deliberate and beneficial, since trading a little bias for less variance is standard practice. Discrimination is unjustified difference in outcome by protected characteristic. A statistically biased model can produce fair outcomes and an unbiased one can discriminate. Using the words interchangeably means an engineer hears a technical property while a lawyer hears a liability.

Why does removing sensitive attributes not fix bias?

Because models find proxies. Postcode, vocabulary patterns, device type, education history and employment gaps all correlate with protected attributes, and a capable model will use any of them because doing so improves prediction. Dropping the attribute removes the label and leaves the signal. Worse, it often removes your ability to measure whether the system discriminates, so a control that does not work has also disabled the check that would have told you.

Which fairness metric should you use?

Whichever fits the decision, documented before you test. Demographic parity, equal opportunity, equalised odds and disparate impact are all precise and reasonable, and they conflict mathematically, so satisfying one can mean violating another. There is no configuration that is fair in every sense at once. What matters is that somebody chose deliberately, recorded why it fits this decision and population, and noted what it trades away. Courts and regulators tend to reason in disparate impact terms regardless of internal preference.

Why is bias harder to test in an agent than in a model?

Because fairness metrics assume one model producing one decision with a measurable outcome. An agent produces a trajectory of dozens of small choices: which records to retrieve, how to phrase a lookup, which tool to call, how much effort to spend, whether to escalate. None of those is the decision being tested, each looks defensible alone, and the disparity appears only in aggregate. The trajectory can discriminate while every individual step is clean.

How do you test an agent for bias?

Measure outcomes at the case level rather than certifying the model. Compare what actually happened to comparable cases across groups: escalation rate, resolution time, number of steps taken, how often the agent gave up, and how often a human overrode it. Those are computable from execution traces with no new methodology, and they catch the aggregate disparity that step-level testing cannot see. It does require per-step tracing to exist, which many deployments lack.

Can bias come from retrieval rather than the model?

Yes, and almost nobody audits for it because every fairness toolkit examines models. If retrieval surfaces different evidence depending on how a query is phrased, two materially identical cases were never decided on the same information, so the model can be entirely fair while the system produces disparate outcomes. An unevenly documented corpus does the same thing. The trace looks clean throughout: every retrieval was relevant and every answer was grounded in a source.

How do you measure fairness without collecting sensitive data?

With difficulty, and the tension is genuine rather than a failure of effort. Demonstrating that outcomes do not differ by protected characteristic requires knowing those characteristics, while data protection principles push toward not collecting them. Established approaches include collecting the attributes separately under a narrow lawful basis used only for aggregate testing, or inferring at population rather than individual level. Both need legal sign-off. What does not work is treating non-collection as a fairness control.

Where does bias enter an AI system?

At every stage, and rarely at the one people inspect. Problem framing encodes a judgement about what to predict. Historical data teaches the past including the parts nobody would defend. Labelling imports annotator judgement. Optimising for aggregate accuracy will trade minority-group performance for majority-group performance if that raises the average. And feedback loops compound whatever skew existed at the start, so a small initial disparity grows rather than washes out. Auditing the model alone finds a fraction of it.

Does AI bias require intent?

No, and this is the point most often misunderstood at senior level. Discriminatory impact is assessed on outcomes rather than on intentions, so a system can create unjustified disparity with nobody having chosen it and no protected attribute in the inputs. That also means "we never intended that" is not a defence, and it is not a useful thing to say during an inquiry. The practical implication is that fairness has to be measured rather than asserted from process.

Is AI bias worse than human bias?

Different rather than simply worse, and the differences cut both ways. A system is consistent, auditable, and correctable in a way an individual is not, which are genuine advantages. It also applies whatever bias it has at the speed and volume of the business, without the variation across individuals that makes human bias partially self-limiting. So the same properties that make a system testable make its errors uniform. The honest comparison is against your actual current process, not against an idealised one.

What do regulators require on AI bias?

The detail varies and the direction is consistent. Risk-tiered regimes such as the EU AI Act impose data governance and bias examination obligations on high-risk systems. Several jurisdictions require bias audits for specific applications, most established for automated employment decision tools. Voluntary frameworks such as the NIST AI Risk Management Framework are widely used as the structure for demonstrating diligence. None accepts non-collection of an attribute as an answer, and all of them expect documentation produced before rather than after scrutiny.

Compare cases, not models
Did comparable cases get comparable treatment?

Answering that requires the trace. SERAA Cortex records every model call and tool call per agent with actor attribution, so escalation rates, step counts and override frequency are things you can compare across cohorts rather than estimate.