What is agent washing?
Agent washing is what happens when a product gets relabelled rather than rebuilt. A chatbot, an assistant, or a scripted automation acquires agentic marketing language without acquiring the ability to decide its own steps. Gartner named the practice in June 2025.
This pattern isn't new. 'Greenwashing' was coined when companies made environmental claims bigger than what they did. 'AI washing' was when products claimed to use smart technology but were just following simple, fixed rules. 'Agent washing' is the same trick, just with a newer buzzword, and it matters more this time. People are paying extra for AI 'agents' because they're supposed to make decisions and act on their own, something a basic chatbot simply cannot do, no matter the price.
A Gartner study from June 2025 put a number on this problem, and it's still widely referenced today. Out of thousands of vendors claiming to offer agentic AI, only about 130 were found to genuinely deliver agentic capabilities. Even if you question how exact that number is, almost no one selling in this space disagrees with the bigger picture it shows.
The uncomfortable part: much of it is not deception
It is tempting to read agent washing as vendors lying, and some of it is. The more useful reading is that there is no standards body for what qualifies as an agent, no certification, and no agreed threshold. In that vacuum, every vendor sets the bar where their product already stands, and most of them believe their own framing.
That changes the remedy. If the problem were dishonesty, you would solve it by finding honest vendors. Because the problem is definitional, you cannot rely on the label at all, from anyone, including from vendors acting in complete good faith. You need a test you apply yourself. The test that does most of the work is on the AI agent page: who decided the order of the steps?
Why does agent washing happen?
Four forces, only one of which is dishonesty. The word has no agreed boundary, buyers ask for agents by name, budget follows the label rather than the capability, and a genuine agent is considerably harder to build than a chatbot with a few functions attached.
Knowing how this works matters because it helps you spot the signs that mean something. A vendor simply responding to what the market wants to hear acts differently from one deliberately trying to trick you, but both can end up with a slide that says 'agentic.'
- There's no line to cross. Nobody can be caught lying about a definition that doesn't exist. A product that plans one step and uses one tool sits somewhere on the scale of 'agentic,' and the vendor isn't technically lying by calling it that.
- Buyers ask for it by name. When RFPs specifically ask for agentic AI, a vendor with a great product that isn't quite agentic has to choose between losing the deal or changing the wording. Demand shapes the label.
- Money follows the word. Calling something 'agentic AI' unlocks budget that 'workflow automation improvements' simply doesn't, for both vendors and buyers. That same pull affects your own teams too, which we'll cover below.
- The real thing is hard to build. An agent that figures out its own path needs planning, memory, the ability to use tools, a way to know when to stop, safety checks, testing, and clear oversight. Slapping the 'agentic' label on a basic chatbot just takes a marketing push.
The picture above is the whole problem in one image. The same two words are applied across a capability range that spans an order of magnitude, which means the label carries no information. It is not that the label is often wrong. It is that it cannot be right or wrong, because it excludes nothing.
How do you detect agent washing?
Ask five questions that a simple demo can't answer. Each one checks for something the product must be built with, not just promised on a slide. A product that's faking it usually stumbles on the same two questions: how does it know when to stop, and can it show you exactly what went wrong when it failed.
Demonstrations are rehearsed and prove almost nothing about autonomy, because a scripted path and a chosen path look identical when the script was written for that input. These questions are designed to be answerable only with evidence.
- 01
Who decided the order of the steps?If a developer or analyst wrote the sequence, it is automation with AI inside it, whatever the label. What you are listening for: a straight answer. Deflection to "it depends on configuration" usually means a person configured it.
- 02
Show me a run where it did something you did not anticipateGenuine agency produces surprises, including good ones. What you are listening for: a specific example with a trace. A vendor who has never seen their product do anything unexpected is describing a workflow engine.
- 03
What decides when it stops?The stopping condition is the component most often absent, and the one no demonstration reveals, because demonstrations end when the presenter stops them. What you are listening for: named limits on steps, time, and spend, and a definition of done.
- 04
Show me the trace of a run that failedThis is the single most reliable tell. A real agent platform has per-step traces because it cannot be operated without them. What you are listening for: tool calls, arguments, and the decision points. Screenshots of a dashboard are not a trace.
- 05
What happens if I remove one of its tools?An agent adapts and finds another route or reports that it cannot proceed. A scripted product breaks at that step. What you are listening for: a willingness to try it in front of you.
Questions three and four do most of the work. They are the two capabilities that cannot be retrofitted into a demo, because both require the product to have been built for production rather than for a pitch. Ask them early. If the answers are vague, the remaining questions will be too, and you have saved yourself an evaluation cycle.
What does agent washing actually cost?
More than the licence fee. A washed purchase underdelivers, the organization concludes that agentic AI does not work, and the next proposal is harder to fund. The damage is to the category inside your own company, and it outlasts the contract.
The direct cost is easy to see and rarely the largest one. Four consequences follow a washed purchase, in roughly this order.
- You pay an agentic price for automation. The immediate loss, and the most recoverable. Contracts end.
- The pilot underdelivers and the diagnosis is wrong. When the results disappoint, the failure is attributed to agentic AI rather than to the product that was not agentic. The wrong lesson gets learned and written down.
- The category is poisoned internally. The next team proposing genuine agentic work now argues against an internal precedent. This is the expensive part, it does not appear in any budget line, and it can set a programme back a year.
- The skepticism becomes self-reinforcing. Washed products fail, buyers lose confidence in the category, and legitimate vendors face heightened suspicion, which raises the cost of every subsequent evaluation for everyone.
There is a variant worth naming separately, because it is not washing and it produces the same outcome. A vendor can describe their product accurately and still leave you disappointed if you assumed more autonomy than the description implied. Honest description and accurate expectation are two different things, and the five questions above close the gap in both cases.
Does agent washing happen inside enterprises too?
Yes, and it is discussed far less than the vendor version. Internal teams relabel existing automation as agentic to secure funding, because budget follows the word. The organization then counts projects as agentic that were never agentic, and draws conclusions from the wrong denominator.
The incentive is identical to the vendor's and the mechanism is the same. A team with a working rules engine and a funding round to survive describes it in the language that gets funded. Nobody involved feels dishonest, and in many cases the relabelling is a fair description of an ambition rather than of a system.
Two consequences follow, and both are worse than they first appear.
- Your agent inventory becomes fiction. If a third of what is registered as an agent is scripted automation, then the count, the risk assessment, and the governance requirements attached to it are all wrong in ways nobody can see. This is a different failure from agent sprawl and it compounds with it: you cannot say how many agents you run, and some of what you have counted is not an agent.
- Your cancellation rate misleads you. When relabelled automation is cancelled because it did not deliver agentic outcomes, that cancellation is recorded against agentic AI. Some proportion of every reported failure statistic, internal and industry-wide, is measuring agent washing being discovered rather than agentic AI failing.
How should you read agent washing statistics?
Check the date before you quote the number. The figures in circulation come from a single Gartner release in June 2025, and much of the 2026 coverage repeats them without the date, which makes a year-old prediction read as fresh research.
This is worth getting right, especially on a page that's warning others about overclaiming, because even the statistics used to sound that warning is often being repeated in ways that stretch the truth. Here's where every figure, you are likely to come across comes from.
| The figure | What it actually is |
|---|---|
| Over 40% of agentic AI projects cancelled by end of 2027 | A Gartner prediction, published 25 June 2025. Not an observed outcome, and not new. Causes given: escalating costs, unclear business value, inadequate risk controls |
| Only ~130 of thousands of vendors are real | A Gartner estimate from the same June 2025 release. An analyst judgement rather than a measured count, and the vendor population has moved since |
| 19% significantly invested, 42% cautious | A Gartner poll of 3,412 webinar attendees, January 2025. A self-selecting audience, so treat it as directional |
| 15% of day-to-day work decisions autonomous by 2028 | From the same release, and rarely quoted alongside the 40% figure because it points the other way |
| 33% of enterprise software will include agentic AI by 2028 | Also from June 2025, up from under 1% in 2024. Note that this measures inclusion of a feature, not agentic outcomes |
Two things follow from this. The same source that predicts the cancellations also predict substantial adoption, so coverage that quotes only the pessimistic half is cherry-picking. Both numbers appear in one press release, and when read together, they describe an ordinary adoption curve, not a collapse.
And a statistic quoted without its date is doing rhetorical work. When a mid-2026 article presents a June 2025 prediction as a current finding, that is the same move as agent washing performed on a number instead of a product: an existing thing relabelled to seem newer than it is. The habit that protects you is small. Ask when the figure was published, what it measured, and whether it was a prediction or an observation, before it reaches a board slide with your name on it.