All glossary terms
R Automation Comparison

Robotic Process Automation (RPA)

RPA's great advantage was that it did not need any integration. That is also exactly why it tends to break. Agents do not remove that dependency; they just change what you are depending on. Instead of relying on a screen staying the same, you are now relying on the AI's inference being correct, and when that fails, it fails much more quietly.

Definition

Robotic process automation is software that completes a task by operating the same interfaces a person would: clicking buttons, typing into fields, reading what appears on screen. Its defining property is that it requires no integration with the systems it drives, which is both why it was adopted at enormous scale and why it breaks when somebody else changes a screen.

What is robotic process automation?

Software that does a job the way a person would: opening the application, clicking the button, typing into the field, reading the screen. The bot follows a recorded or scripted sequence exactly, every time, and does nothing that was not written into it.

That literalness is the point. An RPA bot is deterministic in the strict sense: given the same inputs and the same screens, it does the same thing, which makes it testable, auditable and cheap to run once built.

Two operating modes are worth distinguishing because they carry different risks. Attended bots run alongside a person, triggered by them, on their machine. Unattended bots run on a schedule with no one watching, which is where the estate-level problems accumulate.

A note on how this page is written. A great deal of current material on RPA is published by vendors selling replacements for it, which colours both the framing and the figures. The position taken here is the engineering one rather than the displacement one: organizations that built on RPA were not wrong to automate, they were using the best tooling available at the time, and a meaningful part of what they built should stay exactly where it is.

Why was it adopted so widely?

Because it needed no cooperation from anybody. No API, no vendor engagement, no integration project, no change to the system being driven. If a person could do it on a screen, a bot could be recorded doing it, and that removed every dependency that normally makes automation slow.

That single property explains the adoption curve better than any capability claim, and it is worth appreciating rather than dismissing.

One property shown branching into two consequences. The property is that RPA needs no integration with the systems it drives. On the left branch, the upside: no API required, no vendor cooperation, no change to the target system, so automation could reach legacy and third-party systems and be delivered by a business team in weeks. On the right branch, the cost: the bot depends on a screen layout controlled by somebody else, so a single change upstream can break every bot that touches it.
  • It reached systems nothing else could. A mainframe with no API, a third-party portal with no integration option, a supplier's web form. RPA automated the unautomatable, which is a genuine engineering achievement.
  • It did not need the IT queue. An operations team could deliver working automation in weeks without waiting for a platform roadmap, which is why adoption happened from the business side outward.
  • The business case was unusually easy. Count the hours a task takes, multiply by volume, compare against a licence. That arithmetic closed deals fast, and it is the same arithmetic that agentic AI ROI argues is harder than it looks.

Now look at that same trait from the other direction. A bot that needs no integration is also a bot that depends entirely on an interface someone else controls, and that person can change it anytime without telling you. There is no contract or agreement holding that interface in place, so you have no warning and no say in it. That means the dependency stays completely invisible right up until the moment it fails on you.

What actually breaks in an RPA estate?

Not usually one bot. A change to a shared portal breaks every bot that touches it at the same moment, so somebody has to find the failures, identify the affected bots, rewrite, test and redeploy. In an estate of hundreds, that stops being an incident and becomes a standing workload.

The failure mechanism matters more than the failure rate, because it explains why maintenance cost grows faster than the estate does.

  • Failures correlate. Bots share the systems they drive, so one upstream change produces a cluster of simultaneous breakages rather than an isolated fault. Capacity planning based on average incident rates underestimates this badly.
  • Exceptions have nowhere to go. A bot meeting a case it was not written for throws an error and stops. That is the correct behaviour and it means every unanticipated variation becomes a human queue.
  • The estate outlives its documentation. Bots built by people who have since moved on, driving screens that have since changed, for processes nobody has reviewed. This is the same accumulation problem described at agent sprawl, arriving a decade earlier.
  • Maintenance competes with new work. Once a team spends most of its capacity keeping existing bots running, the programme stops delivering and starts defending.

What do agents actually change?

What matters is the nature of the dependency, not the fact that one exists at all. An RPA bot depends on a screen or a button staying exactly where it was, so it breaks the moment that layout shifts. An agent depends on something different, meaning its judgment about what to do next. Both are things you cannot fully control, but they go wrong in very different ways, and that difference is what you need to plan for.

The distinction that matters operationally is how each one fails, and this is where the honest comparison favours RPA more than displacement material admits.

Two failure modes compared. An RPA bot meeting a changed screen stops and raises an error, described as loud, immediate, and obvious that nothing happened. An agent meeting the same changed screen adapts and continues, which is usually correct and occasionally means it does something plausible and wrong, described as quiet, delayed, and requiring somebody to notice. A caption notes that a loud failure is a better failure and that agents therefore need the apparatus RPA never did.
RPA bots and agents compared by what each depends on and how each fails
Dimension RPA bot Agent
Depends on The interface staying as it was Its inference about the task being correct
Handles variation No. Unanticipated cases stop it Yes. That is the capability being bought
How it fails Loudly. It stops, and nothing happened Quietly. It proceeds, and something plausible happened
Cost per run Near zero once built Per token, every run, varying with difficulty
Testability Unit-testable. Same input, same output Not unit-testable. Requires evaluation on a distribution

Read the third row carefully, because it is the one that gets skipped. A loud failure is a better failure. A stopped bot is a visible queue somebody addresses that morning. An agent that adapted its way to a wrong outcome produces no error, no alert and no queue, and is found later by whoever is affected. Which is why an agentic estate needs tracing, evaluation and guardrails that an RPA estate genuinely did not, and why that apparatus is a cost of the change rather than an optional extra.

The other change worth naming is who can build. RPA development generally needed a specialist in a proprietary tool. Agentic work needs somebody who can specify a goal and a boundary, which is a different and more widely held skill. That widens the pool of builders, and widens the governance problem correspondingly.

Where is RPA still the right answer?

Wherever the sequence is knowable and the interface is stable. If you can write down every step in advance and nothing upstream is going to move, a deterministic bot is cheaper per run, testable, auditable and incapable of inventing anything. Those are real advantages and no agent supplies them.

This is the section RPA teams should be handed first, because it is the one that is true and rarely said.

  • Stable legacy interfaces. A mainframe screen that has not changed in fifteen years is not going to change next quarter. The brittleness argument does not apply, and RPA remains the cost-effective choice.
  • High volume, low variation. The same operation thousands of times on identically structured input. Per-run cost dominates at that volume, and near-zero beats per-token.
  • Where determinism is the requirement. Regulatory processes where the same input must produce the same output and be shown to have done so. A deterministic bot is explainable by construction; an agent needs a faithfulness argument, as set out at agent explainability.
  • Where a wrong action is expensive and a stopped process is not. The loud-failure property becomes a feature rather than a limitation.

The framing to resist is that this is a choice between two technologies. The estates that work in practice run both: deterministic automation where precision and predictability matter, agents where judgment and variation matter, and often a deterministic workflow with a single agentic step inside it. That hybrid is the actual architecture, and it is set out from the coordination side at AI orchestration.

How do you modernise an estate?

By triaging what you have into three groups rather than migrating on principle. Most published guidance offers two, replace and keep. The third is the one that saves the most money: a share of any mature estate should simply be switched off.

Work through the estate in this order, because the third bucket removes work from the first two.

Three buckets for an existing RPA estate. Delete covers bots automating a process that no longer needs doing, duplicates, and bots whose output nobody consumes. Keep covers stable interfaces, high volume with low variation, and processes where determinism is a requirement. Replace covers the bots with the highest maintenance burden and failure rates, and processes that were rejected for automation because they needed judgment. A note states that the delete bucket should be worked first because it removes work from the other two.
  1. Delete first, and expect this bucket to be larger than anyone thinks

    Bots automating a step in a process that has since been retired. Duplicates built by two teams who did not know about each other. Bots whose output nobody has consumed in a year. Migrating any of these to an agent is paying to make an old decision more expensive.

  2. Rank the rest by maintenance burden

    How many times has each bot been rewritten, and how many hours went into it. The heaviest consumers are your migration candidates, because they are already failing on their own terms and replacing them is the fastest measurable saving.

  3. Leave the stable ones alone

    A bot that has run for three years without intervention against an interface nobody controls any more is doing its job. Replacing it introduces non-determinism, per-run cost and an evaluation requirement in exchange for nothing.

  4. Check whether the answer is an API rather than an agent

    A significant share of RPA exists because an integration was unavailable at the time. If the target system has since published an API, the correct modernisation is an integration, which is cheaper and more reliable than either a bot or an agent.

  5. Point new development at the work RPA could never take

    Processes rejected for automation because they involved unstructured documents or judgment calls. This is where agents add capability rather than substitute for it, and where the return is clearest.

  6. Budget the apparatus into the migration, not after it

    Anything moved from a deterministic bot to an agent needs tracing, evaluation and bounded permissions, because a quiet failure is now possible where a loud one used to be. Costing the migration without that is the mistake described at agentic AI ROI.

One important point for anyone leading this with an existing RPA team already in place. The people who built that automation estate carry a lot of knowledge in their heads that never made it into any document. They know which bots quietly break every few weeks, which processes have stopped working without anyone officially noticing, and which screens are about to change and take a bot down with them. That knowledge is incredibly valuable when you are deciding what to fix first and what to leave behind. But if your programme treats the whole estate as a technical debt, rather than as a body of work that needs sorting through carefully, you lose access to exactly the people who know the best.

Frequently asked questions about RPA

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is robotic process automation?

A bot driving your applications from the outside, the way an employee does. It was either recorded from somebody performing the task or scripted step by step, and it will not deviate from that script under any circumstances. Rigidity is the feature rather than the flaw: hand the same bot the same screens and the same data on any given day and the behaviour is identical, which is what makes it testable and auditable.

Why was RPA adopted so widely?

Because it required no cooperation from anyone. No API, no vendor engagement, no integration project, no change to the system being driven. That removed every dependency that normally makes automation slow, so it reached mainframes with no interface, third-party portals with no integration option, and suppliers' web forms. An operations team could deliver working automation in weeks without waiting for a platform roadmap, which is why adoption spread from the business outward.

Why do RPA bots break?

Because the property that made them easy to build is a dependency on an interface somebody else controls. A bot needing no integration is a bot tied to a screen layout that can change without notice, and nothing about the arrangement is contractual, so the dependency stays invisible until it fails. The change is usually upstream and shared, which is why failures arrive as clusters rather than as isolated faults.

Is it true that 30 to 50% of RPA projects fail?

That range recurs widely and it is worth checking before quoting. Across current material it is attached to three different claims: bots requiring rework within twelve months, projects failing outright, and projects abandoned within two years. Those are not the same measurement, and most citations trace back to vendors selling RPA replacements. The direction matches what practitioners report; the figure is not one to put in a board paper without finding the primary source.

Do AI agents replace RPA?

Not wholesale, and framing it as a choice between technologies is the error. Agents change the dependency rather than removing it: an RPA bot depends on a screen staying put, an agent depends on its inference being correct. The estates that work run both, using deterministic automation where precision and predictability matter and agents where judgment and variation matter, often with a single agentic step inside an otherwise deterministic workflow.

How do RPA and agent failures differ?

RPA fails loudly and agents fail quietly, which is the comparison most often skipped. A bot meeting a changed screen stops, raises an error, and nothing has happened, so the result is a visible queue somebody addresses that morning. An agent adapts and continues, which is usually right and occasionally means it did something plausible and wrong, producing no error and no alert. A loud failure is a better failure.

Where should you keep using RPA?

Wherever the sequence is knowable and the interface is stable. Specifically: legacy screens that have not changed in years and are not about to, high volume with low variation where near-zero cost per run beats per-token pricing, regulatory processes where the same input must demonstrably produce the same output, and any case where a wrong action is expensive while a stopped process is merely inconvenient. Those are real advantages and no agent supplies them.

What is the biggest hidden cost of RPA?

Correlated maintenance. Bots share the systems they drive, so a single change to a shared portal breaks every bot touching it simultaneously, and somebody has to locate the failures, identify the affected bots, rewrite, test and redeploy. In an estate of hundreds that stops being an incident and becomes a standing workload, which is why maintenance cost grows faster than the estate does and eventually competes with new development.

How should you triage an existing RPA estate?

Into three groups rather than two. Delete covers bots automating a step in a process since retired, duplicates built by teams unaware of each other, and bots whose output nobody has consumed in a year. Keep covers stable interfaces, high volume with low variation, and anywhere determinism is the requirement. Replace covers the heaviest maintenance consumers. Work the delete bucket first, because it removes work from the other two.

Should you migrate every high-maintenance bot to an agent?

Check for an API first. A significant share of RPA exists because an integration was unavailable when the bot was built, and if the target system has since published one, the correct modernisation is an integration rather than either a bot or an agent. It will be cheaper, faster and more reliable than both. Reaching for an agent where a supported interface now exists is solving a solved problem expensively.

Where do agents add the most value over RPA?

On the work RPA was never able to take. Processes rejected for automation because they involved unstructured documents, inconsistent formats or judgment calls are where agents add capability rather than substituting for something that already worked. That is also where the return is clearest, because the comparison is against a manual process rather than against a functioning bot whose replacement has to justify itself.

What should you not lose when replacing RPA?

The team's knowledge of the estate. The people who built it know which bots break, which processes have quietly died, and which screens are about to change, and none of that is documented anywhere. That knowledge is the most valuable input to a triage and it is not available to a programme that treats the estate as technical debt to be cleared rather than as work to be sorted properly.

Sort the estate, do not clear it
Which of your bots should become agents, and which should just be switched off?

Where a bot becomes an agent, a quiet failure becomes possible where a loud one used to be. SERAA Cortex records every model and tool call per agent with cost and actor attribution, scopes what each can reach, and gates promotion to production.