What is an agentic workflow?
A process where at least one step decides for itself what to do. The rest of the process can still be a perfectly ordinary sequence with fixed order, retries and error handling. What makes it agentic is that somewhere inside, the path is chosen at runtime rather than written down in advance.
Defining it as a spectrum rather than a category is the useful move, because it turns an all-or-nothing architecture argument into a series of small decisions you can make one at a time.
| Degree | What is decided at runtime | Where it fits |
|---|---|---|
| None | Nothing. Every branch and step written in advance | Stable, high-volume processes. See RPA |
| One step | The content or classification produced at a single point | The common production shape. Fixed process, one judgment call |
| Bounded path | Which of a defined set of steps runs, and in what order | Where cases vary but the possible actions do not |
| Open | The goal is given and the whole approach is chosen | Genuinely unpredictable work. Rare, and expensive to govern |
The two middle rows account for most of what actually ships. Fully open agents get the attention and bounded ones get the results, which is worth knowing before a design conversation starts at the wrong end of the table.
Which steps actually need to be agentic?
The ones where you cannot write down the rule, not the ones where writing it down is tedious. That distinction does most of the work. If a rule exists but is long, code it. If no rule exists because the input varies in ways nobody can enumerate, that is where discretion earns its cost.
Walk the process step by step and ask one question at each: could somebody write this step down completely, given a week and access to the people who do it today?
- Genuinely agentic: unstructured input. A document, email or conversation whose shape varies without limit. No amount of rule-writing covers the next format somebody invents.
- Genuinely agentic: judgment against context. Whether this exception is reasonable, whether these two records describe the same thing, whether this answer addresses the question asked.
- Genuinely agentic: an unbounded action space. Where the right next step depends on what the previous step returned, and the possibilities were not known when the process was designed.
- Not agentic, despite appearances: complicated rules. Forty branches is not judgment, it is forty branches. Coding them is cheaper, faster, testable and will not surprise you at three in the morning.
- Not agentic: variation you have simply never documented. Sometimes the rule exists in somebody's head and nobody wrote it down. Writing it down is the fix, and it is cheaper than inference.
That last one is worth sitting with. A process that looks like it needs judgment needs a discussion with the person who has done it for nine years. If you reach for an agent because no one wrote the rules down, you turn a documentation problem into a per-token cost and an ongoing evaluation obligation. And you still will not know what the rule was.
What do you give up at each agentic step?
Four properties, and you give them up per step rather than per system. Determinism, unit-testability, near-zero cost per run, and a failure that announces itself. Each agentic step should be justified against that list rather than adopted because the technology is available.
Being explicit about the trade is what keeps the architecture honest, and it is the same trade discussed from the incumbent side at RPA.
| What you had | What you get instead, and what it costs |
|---|---|
| Same input, same output, always | Variation between runs. Reproducing an incident now needs the model version and the inputs recorded together |
| A unit test that passes or fails | Evaluation across a distribution of cases, maintained over time. See agent evaluation |
| Cost that is effectively zero once built | A per-token charge on every execution, varying with how hard the case is |
| A failure that stops the process | A step that produced something plausible and wrong, and carried on. Nothing alerts |
The fourth row is the one that changes what you have to build. A deterministic workflow that breaks tells you. An agentic step that is wrong does not, so the apparatus of tracing, evaluation and bounded permissions is part of the cost of that step rather than a separate programme. Cost the step with it included, as argued at agentic AI ROI.
Plan first, or decide as you go?
Two architectures travel under the same name and they part company on one point that matters for governance: whether the intended sequence ever exists as an object a person can look at before anything happens. Plan-first produces exactly that. Decide-as-you-go never produces one at all.
This distinction gets very little attention and it determines whether human approval is even possible.
| Question | Plan first | Decide as you go |
|---|---|---|
| Approval | Yes. The plan is an object. Approve once, then execute | No. Only per action, which does not scale past a handful of steps |
| Surprises | Poorly. A plan built on a wrong assumption executes anyway unless it replans | Well. That is the entire point of the pattern |
| Cost | One planning call, then cheap execution | A model call at every step, so cost scales with path length. See model routing |
| The trace | Intended sequence and actual sequence, comparable against each other | Actual sequence only. Intent has to be inferred |
The practical recommendation follows from the first and last rows. Where an action is irreversible or expensive, use plan-first and put the gate on the plan. Approving a plan once is a workable human control; approving twenty individual actions is a control nobody will operate by the third week. Where the work is exploratory and cheap to undo, decide-as-you-go is the better pattern and the weaker governance story, which is a trade to make deliberately rather than by default.
Hybrids exist and tend to be the best answer: plan first, execute, and replan only when a step returns something the plan did not anticipate. That keeps the inspectable object and recovers most of the adaptability.
What does the deterministic shell provide?
Everything that makes a process operable: ordering, retries, idempotency, timeouts, error handling and a record of what ran. Those are solved problems in workflow engineering, and an agentic system that abandons them is re-solving them badly at the same time as it is solving a new problem.
This is the argument for keeping the outer process boring, and it is the part most often lost in enthusiasm for autonomy.
- Idempotency. Running a step twice must not charge the customer twice. A workflow engine handles this with keys and deduplication. An agent deciding to retry has no such guarantee unless something around it provides one.
- Ordering and dependency. Some steps genuinely cannot run before others. Expressing that as a graph is clearer, cheaper and more reliable than hoping the model infers it each time.
- Timeouts and compensation. What happens when step seven hangs, and what has to be undone. Mature workflow engines have well-understood answers; a loop does not.
- A record of what ran. The shell knows the intended steps, so it can report which completed and which did not, which is the backbone an agentic trace attaches to. See AI agent observability.
- A stopping point. A bounded process ends. An agent without a stopping condition is the failure described at AI agent, where the component most often left out is the rule for when to stop.
Put plainly: the shell supplies guarantees and the agentic step supplies judgment. Asking a model to provide both means asking a probabilistic component to deliver deterministic properties, which is the wrong tool for half the job and expensive for the other half.
How do you migrate a workflow without rewriting it?
By changing one step at a time inside the process you already have. The failure mode is a rewrite: replacing a working process with an autonomous system, then discovering that the guarantees you had were doing more work than anyone documented.
Six steps, and the first two are about what not to touch.
-
Write the process down as it actually runs
Not as documented. Include the exceptions people handle by hand, because those are usually where the agentic value is and they are almost never in the diagram.
-
Mark the steps where a rule exists
Even a long or tedious rule. Those stay deterministic, and confirming this early prevents a migration from expanding to cover work that was never the problem.
-
Pick the single step causing the most manual handling
Usually the one generating a queue: an unstructured document somebody rekeys, an exception somebody judges. One step, inside the existing process, with the shell unchanged around it.
-
Bound it before you tune it
Decide what that step may read, what it may do, what it costs at most, and what happens when it is unsure. An unsure step should hand back to a person rather than guess, and that path has to exist before launch.
-
Measure against the process, not the step
Cycle time and manual touches for the whole workflow. A step that performs beautifully while shifting work into review has made things worse, and step-level metrics hide that.
-
Only then consider a second step
Two agentic steps in sequence compound both the cost and the uncertainty, and the second one inherits whatever the first got wrong. If the first has not been stable for a while, the answer to the second is not yet.
A closing note for anyone doing this alongside an existing automation team. The steps worth making agentic are the ones that team already knows are painful, because they are the ones generating rework and exceptions. That makes this an extension of what they built rather than a replacement for it, which is both the accurate framing and the one that gets the migration done.