What is agentic RAG?
Retrieval wrapped in a loop instead of run as a pipeline. The agent plans, retrieves, judges whether what came back is enough, and searches again if it is not. One conditional step after the evaluation is what makes the system agentic: it can loop.
Classic retrieval-augmented generation is a straight line. Embed the query, pull the top matches, put them in the prompt, generate. It works, and it fails the moment a question needs more than one lookup: multi-part queries, comparisons across sources, or follow-ups that depend on what was just found.
Agentic RAG is the architectural response to that specific failure. And the useful way to identify it is the same test that identifies any agent, applied to retrieval.
Who decided the retrieval strategy? In classic RAG a developer did: how many chunks, from which index, reranked how, once. That is automation with retrieval inside it, and it is often the right build. In agentic RAG the agent decides, per query, at runtime. Which means agentic RAG is not a better retrieval technique. It is an agent whose principal tool happens to be a retriever, and everything true of agents becomes true of your retrieval layer.
How is it different from classic RAG?
Five differences that matter operationally. Who decides the retrieval, whether the cost per query is knowable in advance, how long a response takes, what a failure looks like, and what you have to evaluate. Only the first is architectural. The other four are consequences.
The comparison usually presented is accuracy, which is the least useful axis because it depends entirely on query difficulty.
| Dimension | Classic RAG | Agentic RAG |
|---|---|---|
| Who decides | A developer, once, at build time | The agent, per query, at runtime. This is the only architectural difference |
| Cost per query | Fixed and knowable. One retrieval, one generation | Unbounded by default. Reported at roughly ten times classic for the same job |
| Latency | One round. Predictable | Each round adds a vector search plus a rerank. Three rounds lands in the five to fifteen second range |
| Failure mode | One bad retrieval, one wrong premise, visible | A bad first retrieval steers the next query. Errors compound rather than add |
| What you evaluate | Was the retrieved chunk relevant | Was the retrieval path reasonable, across searches the agent chose |
The honest summary across published comparisons is consistent: the agentic branch wins on hard questions and loses on easy ones. Reported gains are largest on multi-hop and complex queries and negligible on simple factual lookups, where you have paid ten times the cost and added seconds of latency for a result the straight line would have produced.
Worth noting an inversion that has caught teams out: retrieval, not generation, is now the bottleneck in these systems. Each round costs a few hundred milliseconds for the vector search and a few hundred more for reranking, so the loop is where the latency lives rather than the model call.
What are the agentic RAG patterns?
Five patterns cover almost every enterprise case: router, ReAct, plan-and-execute, multi-agent retrieval, and self-RAG. Most teams need only two of them. Start with routing and a basic reason-act loop, then add self-evaluation gates once that first loop is stable in production.
The patterns differ in where the agent's discretion sits, which is a more useful way to hold them than by name.
- 01
RouterThe agent chooses which source or index to query for a given question. Discretion sits in: source selection. The cheapest pattern to add and usually the highest return, because most retrieval failures are querying the wrong corpus.
- 02
ReActReason, act, observe, repeat. The agent searches, reads the result, and decides its next search from what it learned. Discretion sits in: the sequence of queries. This is the pattern most people mean by agentic RAG.
- 03
Plan-and-executeThe agent decomposes a complex question into sub-questions up front, retrieves for each, then synthesises. Discretion sits in: the decomposition. Better than ReAct where sub-questions are independent and can run in parallel.
- 04
Multi-agentSeparate agents handle separate sources or domains, with a coordinator combining results. Discretion sits in: delegation. Inherits every coordination problem covered at multi-agent orchestration, including the compounding reliability arithmetic.
- 05
Self-RAGThe agent explicitly judges four things: whether retrieval is needed at all, whether what came back is relevant, whether its answer is actually grounded in that evidence, and whether the answer is useful. Discretion sits in: self-assessment, which is also this pattern's weakness.
The sequencing advice in the field is unanimous and worth following: router and ReAct first, self-evaluation gates afterwards. A self-evaluating loop built before the basic loop is stable produces a system that confidently iterates in the wrong direction.
Why is the stopping condition the whole design?
Because the loop has no natural end. The agent itself decides if its results are good enough, so nothing tells it to stop.
That is why a hard limit, like a max number of tries, is not optional. Without one, the agent can just keep searching forever on a hard question, burning time until the system grinds to a halt. Skipping this is not a small gap. It is the one thing standing between a system that works and one that quietly breaks.
This is the component most often absent, and it is the same gap identified on the AI agent page: a stopping condition is the piece nobody lists among an agent's parts and the piece whose absence is most expensive.
- Set a hard iteration cap. Production systems typically bound the loop at two or three rounds. Beyond that the marginal retrieval rarely changes the answer and reliably changes the latency.
- Add early exit on self-assessment. If the agent's sufficiency scores are consistently high after two rounds, stop. The cap catches the pathological case; the early exit saves the ordinary one.
- Cap spend, not only steps. A step limit and a token ceiling fail differently. One long retrieval can cost more than three short ones, so bound both.
- Decide what happens at the cap. The design decision people skip. Does the agent answer from what it has, say it could not determine the answer, or escalate? Answering from insufficient evidence at the iteration limit is how a bounded loop produces a confident wrong answer.
The latency consequence is worth planning for rather than discovering. Three retrieval rounds plus the accompanying model calls commonly lands between five and fifteen seconds, which most interfaces cannot present as a synchronous response. The practical patterns are streaming intermediate progress, running independent retrievals in parallel, and handing genuinely multi-step queries to an asynchronous job.
What gets worse, and not just better?
Three things, and none of them appears in a demonstration. Errors compound instead of adding, the attack surface moves from your index to anything the agent can reach at runtime, and evaluation stops being about chunks and becomes about paths.
Every published comparison covers the accuracy gain. These are the costs on the other side of it, and they are the reason agentic RAG needs the apparatus rather than just the loop.
Errors steer, they do not merely accumulate
In classic RAG a poor retrieval produces one wrong premise, and it is at least visible in the retrieved context. In an agentic loop the first retrieval shapes the second query, which shapes the third. A confidently wrong first search produces a confidently wrong assessment that it was sufficient, because the agent judging sufficiency is the same agent that chose the search. Self-RAG's weakness is exactly here: it grades its own homework.
The attack surface moves from the corpus to the runtime
Classic RAG has a bounded exposure: an attacker has to get content into the index you built. An agent that decides what to fetch may search the web, follow a link, or call a tool, so the surface becomes anything reachable at execution time. That is a materially larger problem and it is the delivery route described at prompt injection. Relevance ranking cannot separate hostile content from useful content, because an attacker who wants their content retrieved writes content that is highly relevant.
Evaluation becomes trajectory evaluation
You cannot score whether the retrieved chunk was relevant when there were six retrievals the agent selected. You have to score the retrieval path: were the queries sensible, did the agent stop at the right point, did it recognise insufficient evidence. That is the same problem covered at agent evaluation, and it means per-iteration tracing is a prerequisite rather than an optimisation.
Should you use agentic RAG?
Not as a replacement, and not everywhere. Route between the two rather than choosing, because a classifier at the front door sending simple queries down the straight line and hard ones into the loop captures most of the benefit at a fraction of the cost.
Framing this as a choice is the mistake, and it is the framing most vendor material encourages.
- Classic retrieval is the right default. Self-contained questions, clean documents, cost and latency as constraints. If precision is short, add hybrid search and a reranker before adding an agent. That is the best value available and it is not agentic.
- Agentic RAG earns its cost on genuinely hard queries. Multi-hop reasoning, comparisons requiring several sources, questions where the second search depends on the first, and cases where every claim needs grounding in a citable source.
- Route, do not choose. A cheap classifier at the entrance deciding which path a query takes is reported to cut both cost and latency substantially against running everything through the loop. Most production systems that work are hybrids.
- Treat it as an engineering discipline. Per-iteration observability, prompt caching, hybrid routing, and explicit stop conditions. A published assessment puts it bluntly: everything beyond those four is marketing.
One upstream point that decides more than the architecture. An agentic loop searching an ambiguous corpus retrieves more of the ambiguity rather than resolving it, and iteration does not fix data that means different things in different systems. If your retrieval quality problem is actually a data problem, the loop will make it more expensive rather than better. See data intelligence.