What is agent memory?
It is the memory that the agent saves after finishing one run, so it can use it again later. That is different from the context window, which is just temporary working space. The context window clears out once a task ends, like a whiteboard getting wiped clean.
The memory has two parts. One, a place where information gets stored. Two, a decision about what to pull back out and use later. That second part, choosing what to bring back and when, is where most of the real difficulty lives.
Two problems that get conflated because both involve putting text in front of a model.
| Question | Context | Memory |
|---|---|---|
| Lifespan | One run. It empties | Indefinite, until somebody removes it |
| The hard problem | A budget. What fits, and in what order | Trust. Whether this should be believed, and by whom |
| Blast radius of a bad entry | The current session | Every future session that retrieves it, including other agents |
| Who else is affected | Nobody. It is per run | Anyone sharing the store, which is usually the point of having one |
The context side of this is covered at context engineering, which argues that context poisoning is injection that persisted. Memory is the mechanism by which it persists, and that is the reason this term deserves its own treatment rather than a paragraph there.
What kinds of memory are there?
Four, borrowed from cognitive science and reasonably stable across implementations. The taxonomy earns its place because each type carries a different risk and needs its own retention answer, and because most systems in the market implement one of the four while describing themselves as having all of them.
- 01
WorkingWhat the agent holds during one run: the task, intermediate results, what it just did. The risk: it runs out, and truncation drops the wrong thing. This is a context problem rather than a memory one.
- 02
EpisodicWhat happened. Past conversations, previous runs, what the user asked last Tuesday. The risk: it is the largest store and the least curated, so it is where retrieval quality degrades first.
- 03
SemanticFacts the agent has come to hold. This customer prefers email. That region requires a second approval. The risk: this is the one that becomes a belief, and the one poisoning targets.
- 04
ProceduralHow to do things. Learned sequences, successful approaches, workarounds. The risk: a learned shortcut that worked once becomes the default route, including around a control somebody put there deliberately.
The distinction worth holding onto is between the third and the second. Episodic memory records that something was said. Semantic memory records that something is true, and the step between them is where an agent stops reporting and starts asserting. That step is the subject of section four.
Why is agent memory an attack surface?
Because it is usually much easier to write something into an agent's memory than it is to carefully check what gets read back. If an agent learns from its conversations, then anyone who can chat with it can try to feed it something to ‘learn’. And if there is no checkpoint between what the agent simply observes and what it keeps, that is an open door. Bad information can slip in quietly today and cause problems much later.
The economics of the attack change completely once persistence enters the picture.
- One write, many reads. An injection into a context window pays off once. An injection into memory pays off on every future retrieval, which inverts the attacker's effort-to-return ratio.
- It reaches agents that were never attacked. A shared memory store is shared. An entry planted through one agent's conversation is retrievable by every other agent reading that store.
- Provenance disappears on the way in. Once a claim is a memory entry, it looks like every other entry. Whether it came from an authoritative document or from something a user typed is usually not recorded.
- Retrieval looks like success. Nothing errors. The trace shows a relevant memory retrieved and used, which is the same signature as the system working correctly. This is the 200 OK problem described at AI agent observability.
The delivery route matters less than you would think. An entry can be planted by a user talking to the agent, by a document the agent ingested, or by a tool description it read, which is the vector covered at MCP and prompt injection. What memory adds is not a new way in. It is the removal of an expiry date.
When does an observation become a belief?
This is the design question that decides whether a memory system is useful or dangerous, and most implementations answer it by default rather than deliberately. One conversation should not become a fact the organisation holds, and a thousand conversations probably should.
Three questions sit inside that, and answering them explicitly is what separates a memory system from a log with retrieval bolted on.
- What counts as evidence? A single conversation is an anecdote. Recurrence across genuinely distinct people and sessions is a pattern. Systems that promote on first observation learn whatever the last person said.
- Who approves it, and what do they see? A reviewer shown a summary of what will be stored is approving a paraphrase. A reviewer shown the exact text that will be written is approving the thing itself, and only the second is a real gate.
- How far does it travel? Something learned at one site is true at that site. Letting it climb one level at a time, accumulating evidence at each step, stops one branch's local workaround becoming organisational policy.
There is a structural point underneath all three. A flat memory store cannot express that a corporate policy, a regional playbook and a local procedure are different altitudes of the same truth. Flatten them and retrieval either returns noise or needs manual filtering forever, because nothing in the store records which one should win when they disagree.
Who can delete a memory?
In most implementations, nobody, and this is the gap that turns a memory feature into a regulatory problem. If an agent derived a fact from a record somebody later asked you to erase, deleting the source record does not delete the derived memory, and usually nothing links the two.
Four questions worth putting to any memory system before it holds anything consequential.
| The question | Why it matters, and what a weak answer costs you |
|---|---|
| Can you delete one entry? | If removal means rebuilding the store, nobody will do it. The erasure obligation then sits unmet against a system nobody wants to touch |
| Can you find what derived from a source? | Deleting a customer record while the memory derived from it remains is a deletion that did not happen. This is the propagation problem from enterprise connectors, one layer further downstream |
| Can you explain why the agent believes something? | Without provenance per entry, the honest answer is that it read it somewhere, which is not an answer anybody can act on |
| Can something be demoted rather than only deleted? | Much of what goes wrong is over-generalisation rather than falsehood. If the only remedy is deletion, correct-but-too-broad knowledge gets destroyed instead of narrowed |
The last row is the one most often missing and the cheapest to add. A memory that can only be created or destroyed forces a binary decision on something that is usually a scoping error, and recording the reason for a demotion gives you the one audit artefact that explains how the agent's beliefs changed over time.
How do you implement an agent's memory system responsibly?
Start with the smallest memory that solves your problem, because retention is easy to add and almost impossible to walk back. Most use cases described as needing memory need session continuity, which is a considerably narrower thing and carries none of the same exposure.
Six decisions, in the order that determines the ones after it.
-
Establish whether you need memory or just continuity
Remembering within a conversation is context handling. Remembering across conversations is memory, with everything on this page attached. A surprising share of requirements turn out to be the first.
-
Separate what the agent was told from what it concluded
Episodic and semantic stores should not be the same store. Conflating them is how a customer's offhand remark becomes a fact about that customer.
-
Record provenance on every entry, at write time
Where it came from, when, and via which agent. This is unrecoverable afterwards and it is what makes every later question answerable.
-
Put a gate between observation and retention
Whether that gate is recurrence, human review or both, decide it explicitly. A system that writes on first observation is a system that learns from whoever spoke to it last.
-
Scope retrieval before you scale the store
Which memories this agent, acting for this person, may read. Without it a shared store becomes a route around access control, and the constraint has to travel with the entry rather than be applied afterwards.
-
Build deletion and demotion on day one
Including the link from a source record to everything derived from it. Retrofitting this is the single most expensive thing on the list, and the obligation arrives whether or not you built for it.
One honest note on evaluation. Memory makes an agent non-reproducible by design: the same input with a different memory state produces a different output, which means a failure you cannot reproduce may be a memory difference rather than a flake. Recording which memory entries were retrieved on each run is what keeps evaluation meaningful once memory is in the loop.