All glossary terms
A Data & reasoning Architecture

Agentic RAG

Agentic RAG turns retrieval from a pipeline into a loop the agent controls. The change is not better retrieval. It is that a developer no longer decides the retrieval strategy. Which fixes the real limitation and inherits every consequence of agency.

Definition

Agentic RAG turns retrieval from a fixed pipeline into a loop the agent itself controls. The real change is not better retrieval, it is that a developer no longer decides the retrieval strategy, the agent does. That fixes a genuine limitation, but it also means retrieval now inherits every risk that comes with giving an agent control in the first place.

What is agentic RAG?

Retrieval wrapped in a loop instead of run as a pipeline. The agent plans, retrieves, judges whether what came back is enough, and searches again if it is not. One conditional step after the evaluation is what makes the system agentic: it can loop.

Classic retrieval-augmented generation is a straight line. Embed the query, pull the top matches, put them in the prompt, generate. It works, and it fails the moment a question needs more than one lookup: multi-part queries, comparisons across sources, or follow-ups that depend on what was just found.

Agentic RAG is the architectural response to that specific failure. And the useful way to identify it is the same test that identifies any agent, applied to retrieval.

Who decided the retrieval strategy? In classic RAG a developer did: how many chunks, from which index, reranked how, once. That is automation with retrieval inside it, and it is often the right build. In agentic RAG the agent decides, per query, at runtime. Which means agentic RAG is not a better retrieval technique. It is an agent whose principal tool happens to be a retriever, and everything true of agents becomes true of your retrieval layer.

How is it different from classic RAG?

Five differences that matter operationally. Who decides the retrieval, whether the cost per query is knowable in advance, how long a response takes, what a failure looks like, and what you have to evaluate. Only the first is architectural. The other four are consequences.

The comparison usually presented is accuracy, which is the least useful axis because it depends entirely on query difficulty.

Classic retrieval-augmented generation compared with agentic RAG across five operational dimensions
Dimension Classic RAG Agentic RAG
Who decides A developer, once, at build time The agent, per query, at runtime. This is the only architectural difference
Cost per query Fixed and knowable. One retrieval, one generation Unbounded by default. Reported at roughly ten times classic for the same job
Latency One round. Predictable Each round adds a vector search plus a rerank. Three rounds lands in the five to fifteen second range
Failure mode One bad retrieval, one wrong premise, visible A bad first retrieval steers the next query. Errors compound rather than add
What you evaluate Was the retrieved chunk relevant Was the retrieval path reasonable, across searches the agent chose

The honest summary across published comparisons is consistent: the agentic branch wins on hard questions and loses on easy ones. Reported gains are largest on multi-hop and complex queries and negligible on simple factual lookups, where you have paid ten times the cost and added seconds of latency for a result the straight line would have produced.

Worth noting an inversion that has caught teams out: retrieval, not generation, is now the bottleneck in these systems. Each round costs a few hundred milliseconds for the vector search and a few hundred more for reranking, so the loop is where the latency lives rather than the model call.

What are the agentic RAG patterns?

Five patterns cover almost every enterprise case: router, ReAct, plan-and-execute, multi-agent retrieval, and self-RAG. Most teams need only two of them. Start with routing and a basic reason-act loop, then add self-evaluation gates once that first loop is stable in production.

The patterns differ in where the agent's discretion sits, which is a more useful way to hold them than by name.

Two flows compared. On the left, classic RAG as a straight line from query through embedding, top-k retrieval, prompt assembly and generation, annotated that a developer decided every step at build time. On the right, agentic RAG as a loop: the agent plans, retrieves, evaluates sufficiency, and either loops back to retrieve again or proceeds to synthesis, annotated that the agent decides at runtime and that the conditional edge after evaluation is what makes it agentic. A stop condition gate is marked on the loop.
  1. 01
    RouterThe agent chooses which source or index to query for a given question. Discretion sits in: source selection. The cheapest pattern to add and usually the highest return, because most retrieval failures are querying the wrong corpus.
  2. 02
    ReActReason, act, observe, repeat. The agent searches, reads the result, and decides its next search from what it learned. Discretion sits in: the sequence of queries. This is the pattern most people mean by agentic RAG.
  3. 03
    Plan-and-executeThe agent decomposes a complex question into sub-questions up front, retrieves for each, then synthesises. Discretion sits in: the decomposition. Better than ReAct where sub-questions are independent and can run in parallel.
  4. 04
    Multi-agentSeparate agents handle separate sources or domains, with a coordinator combining results. Discretion sits in: delegation. Inherits every coordination problem covered at multi-agent orchestration, including the compounding reliability arithmetic.
  5. 05
    Self-RAGThe agent explicitly judges four things: whether retrieval is needed at all, whether what came back is relevant, whether its answer is actually grounded in that evidence, and whether the answer is useful. Discretion sits in: self-assessment, which is also this pattern's weakness.

The sequencing advice in the field is unanimous and worth following: router and ReAct first, self-evaluation gates afterwards. A self-evaluating loop built before the basic loop is stable produces a system that confidently iterates in the wrong direction.

Why is the stopping condition the whole design?

Because the loop has no natural end. The agent itself decides if its results are good enough, so nothing tells it to stop.

That is why a hard limit, like a max number of tries, is not optional. Without one, the agent can just keep searching forever on a hard question, burning time until the system grinds to a halt. Skipping this is not a small gap. It is the one thing standing between a system that works and one that quietly breaks.

This is the component most often absent, and it is the same gap identified on the AI agent page: a stopping condition is the piece nobody lists among an agent's parts and the piece whose absence is most expensive.

  • Set a hard iteration cap. Production systems typically bound the loop at two or three rounds. Beyond that the marginal retrieval rarely changes the answer and reliably changes the latency.
  • Add early exit on self-assessment. If the agent's sufficiency scores are consistently high after two rounds, stop. The cap catches the pathological case; the early exit saves the ordinary one.
  • Cap spend, not only steps. A step limit and a token ceiling fail differently. One long retrieval can cost more than three short ones, so bound both.
  • Decide what happens at the cap. The design decision people skip. Does the agent answer from what it has, say it could not determine the answer, or escalate? Answering from insufficient evidence at the iteration limit is how a bounded loop produces a confident wrong answer.

The latency consequence is worth planning for rather than discovering. Three retrieval rounds plus the accompanying model calls commonly lands between five and fifteen seconds, which most interfaces cannot present as a synchronous response. The practical patterns are streaming intermediate progress, running independent retrievals in parallel, and handing genuinely multi-step queries to an asynchronous job.

What gets worse, and not just better?

Three things, and none of them appears in a demonstration. Errors compound instead of adding, the attack surface moves from your index to anything the agent can reach at runtime, and evaluation stops being about chunks and becomes about paths.

Every published comparison covers the accuracy gain. These are the costs on the other side of it, and they are the reason agentic RAG needs the apparatus rather than just the loop.

Three panels. The first shows a bad first retrieval steering the second query and then the third, so errors steer rather than merely accumulate. The second contrasts the classic RAG attack surface, limited to the indexed corpus, with the agentic surface covering anything the agent can fetch at runtime including web pages, tools and followed links. The third shows evaluation shifting from scoring one retrieved chunk to scoring a path of searches the agent chose.

Errors steer, they do not merely accumulate

In classic RAG a poor retrieval produces one wrong premise, and it is at least visible in the retrieved context. In an agentic loop the first retrieval shapes the second query, which shapes the third. A confidently wrong first search produces a confidently wrong assessment that it was sufficient, because the agent judging sufficiency is the same agent that chose the search. Self-RAG's weakness is exactly here: it grades its own homework.

The attack surface moves from the corpus to the runtime

Classic RAG has a bounded exposure: an attacker has to get content into the index you built. An agent that decides what to fetch may search the web, follow a link, or call a tool, so the surface becomes anything reachable at execution time. That is a materially larger problem and it is the delivery route described at prompt injection. Relevance ranking cannot separate hostile content from useful content, because an attacker who wants their content retrieved writes content that is highly relevant.

Evaluation becomes trajectory evaluation

You cannot score whether the retrieved chunk was relevant when there were six retrievals the agent selected. You have to score the retrieval path: were the queries sensible, did the agent stop at the right point, did it recognise insufficient evidence. That is the same problem covered at agent evaluation, and it means per-iteration tracing is a prerequisite rather than an optimisation.

Should you use agentic RAG?

Not as a replacement, and not everywhere. Route between the two rather than choosing, because a classifier at the front door sending simple queries down the straight line and hard ones into the loop captures most of the benefit at a fraction of the cost.

Framing this as a choice is the mistake, and it is the framing most vendor material encourages.

  • Classic retrieval is the right default. Self-contained questions, clean documents, cost and latency as constraints. If precision is short, add hybrid search and a reranker before adding an agent. That is the best value available and it is not agentic.
  • Agentic RAG earns its cost on genuinely hard queries. Multi-hop reasoning, comparisons requiring several sources, questions where the second search depends on the first, and cases where every claim needs grounding in a citable source.
  • Route, do not choose. A cheap classifier at the entrance deciding which path a query takes is reported to cut both cost and latency substantially against running everything through the loop. Most production systems that work are hybrids.
  • Treat it as an engineering discipline. Per-iteration observability, prompt caching, hybrid routing, and explicit stop conditions. A published assessment puts it bluntly: everything beyond those four is marketing.

One upstream point that decides more than the architecture. An agentic loop searching an ambiguous corpus retrieves more of the ambiguity rather than resolving it, and iteration does not fix data that means different things in different systems. If your retrieval quality problem is actually a data problem, the loop will make it more expensive rather than better. See data intelligence.

Frequently asked questions about agentic RAG

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is agentic RAG in simple terms?

Retrieval that can try again. Classic retrieval-augmented generation searches once and answers from whatever came back. Agentic RAG lets the system look at what it found, decide whether that is enough, and search again with a better query if it is not. One conditional step after the evaluation is what makes it agentic: the ability to loop. Everything else follows from that single change.

How is agentic RAG different from RAG?

Who decides the retrieval strategy. In classic RAG a developer decided at build time how many chunks to pull, from which index, reranked how, once. In agentic RAG the agent decides per query at runtime. That is the only architectural difference; the rest are consequences of it, including unbounded cost, higher latency, compounding errors, and evaluation that has to score a path rather than a chunk.

What are the five agentic RAG patterns?

Router, where the agent picks which source to query. ReAct, where it reasons, searches, reads the result and decides the next search. Plan-and-execute, where it decomposes the question up front and retrieves for each part. Multi-agent retrieval, where separate agents handle separate domains. And self-RAG, where the agent judges whether retrieval is needed, whether results are relevant, whether its answer is grounded, and whether it is useful. Most teams need the first two.

How much more does agentic RAG cost?

Roughly ten times a classic pipeline for the same job, according to published comparisons, plus several seconds of added latency. The more important point is that the cost is unbounded by default rather than simply higher, because the agent decides how many times to search. The same question can cost once or twenty times depending on how hard the agent finds it, which is why an explicit iteration budget is part of the architecture rather than a tuning option.

Why does agentic RAG need a stopping condition?

Because nothing outside the agent ever says stop. It is the agent that judges whether the evidence it gathered is sufficient, so the loop has no natural end. Production systems typically cap iterations at two or three and add early exit when self-assessment scores are consistently high. The decision most often skipped is what happens at the cap: answering from insufficient evidence when the iteration limit is reached is how a bounded loop still produces a confident wrong answer.

Why is latency worse with agentic RAG?

Because retrieval rather than generation is the bottleneck in these systems. Each round costs a few hundred milliseconds for the vector search and a few hundred more for reranking, so three rounds plus the accompanying model calls commonly lands between five and fifteen seconds. Most interfaces cannot present that synchronously. The workable patterns are streaming intermediate progress, running independent retrievals in parallel, and handing genuinely multi-step queries to an asynchronous job.

Do errors get worse in an agentic retrieval loop?

Yes, and in a specific way: they steer rather than merely accumulate. A poor first retrieval shapes the second query, which shapes the third, so the loop can travel confidently in the wrong direction. The self-evaluation does not save you, because the agent judging whether the evidence is sufficient is the same agent that chose the search. That is self-RAG's structural weakness: it grades its own homework.

Is agentic RAG less secure than classic RAG?

The exposure is larger. Classic RAG is bounded: an attacker has to get content into the index you built. An agent that decides what to fetch may search the web, follow a link, or call a tool, so the surface becomes anything reachable at execution time. Relevance ranking cannot separate hostile content from useful content, because an attacker who wants their material retrieved writes material that is highly relevant.

How do you evaluate agentic RAG?

By scoring the retrieval path rather than the retrieved chunk, since there were several searches the agent selected. Ask whether the queries were sensible, whether the agent recognised insufficient evidence, and whether it stopped at the right point. That requires per-iteration tracing: record every search issued, what came back, and what the agent concluded about sufficiency, as separate spans. Without it you can see an answer was wrong and not which search sent it wrong.

When should you use classic RAG instead?

Whenever queries are self-contained, documents are reasonably clean, and cost or latency are real constraints. That covers considerably more enterprise use than vendor material implies. If precision is falling short, add hybrid search and a reranker before adding an agent, because that is the best value available and it is not agentic. Published comparisons are consistent: the agentic branch wins on hard questions and loses on easy ones.

Can you use both classic and agentic RAG together?

That is the answer most production systems arrive at. Put a cheap, fast classifier at the front door to decide which path a query takes: simple factual lookups go down the straight line, multi-hop and comparative queries go into the loop. Reported reductions in both cost and latency against running everything agentically are substantial. Framing agentic RAG as a replacement rather than a routing decision is the mistake most vendor material encourages.

Will agentic RAG fix poor retrieval quality?

Not if the problem is the data. An agentic loop searching an ambiguous corpus retrieves more of the ambiguity rather than resolving it, and iteration does not reconcile fields that mean different things in different systems. So if your retrieval quality problem is actually a data problem, adding the loop makes it more expensive without making it better. Check whether an outside reader could interpret your data correctly before deciding the architecture is at fault.

Which search sent it wrong?
A loop you cannot see inside is a loop you cannot fix

SERAA Axon composes SQL, vector search and graph traversal in a single reasoning chain across 100+ connectors, and logs every reasoning step. SERAA Cortex traces each model and tool call per agent, so a retrieval path is something you can read rather than infer.