What is a kill switch?
A single action that stops a running agent and prevents it being re-invoked until somebody lifts the suspension. Two properties make it a control rather than a feature: a defined role can trigger it without an engineering cycle, and it completes inside a defined window.
The regulatory basis is unusually specific for this field. Article 14(4)(e) of the EU AI Act requires that oversight persons be enabled to intervene in a high-risk system's operation, or interrupt it through a stop button or similar procedure that allows the system to come to a halt in a safe state.
Note what that sentence does and does not say. It does not require a button. It requires that the system can be brought to a halt in a safe state, which is a considerably higher bar and the reason this page exists separately from the surface you press it on.
Where this sits. The AI Agent Control Tower is the surface a kill switch is operated from, and it makes the point that graduated intervention such as narrowing a permission or capping spend gets used far more often than a full stop. This page is about the stop itself: what it has to interrupt, what state it leaves behind, and how you know it works. Two reasonable time targets from published guidance: under five minutes for a production agent, under one minute for an agent with transaction authority.
Why is stopping harder than not starting?
Because most implementations block new invocations and leave the current trajectory running. Preventing an agent from picking up new work is a configuration change. Interrupting the tool call it has already dispatched requires you to have designed for interruption, and almost nobody has.
The measured position bears this out. Research reported during 2026 found roughly 37 to 40% of enterprises holding genuine containment controls, against 58 to 59% reporting monitoring and oversight. Twice as many organizations can watch an agent as can stop one.
Three moments, and only one of them is clean.
| When it lands | What actually happens | What you are left with |
|---|---|---|
| Before dispatch | The agent has decided on an action and not yet sent it | Clean. Nothing external happened. This is the only genuinely safe stop |
| After dispatch | The call is in flight. The downstream system never hears about your stop | An effect in the world your trace may not record, because the agent stopped before logging the result |
| Mid-trajectory | Several steps completed, the rest abandoned | An incomplete multi-step operation. Frequently worse than either finishing or never starting |
The engineering consequence is that a kill switch is not something you add to a running system. It is a property the system has to be built with, because interruptibility means the agent checks whether it should continue at points you chose in advance.
What does a halt in a safe state require?
That the system is left in a condition somebody can reason about. A half-completed multi-step operation is not safe: money moved and not recorded, a record created and never linked, a notification sent about a change that was then abandoned. Stopping created that, not the original fault.
This is the part the phrase in the legislation is pointing at, and it is why "add a kill switch" is a bigger request than it sounds.
-
Define safe-stop points in the trajectory
Places where the operation is coherent whether or not it continues: after a read, after a write that stands alone, between independent sub-tasks. The agent checks for a stop signal at those points rather than being killed wherever it happens to be. This is the single design decision that makes the rest possible.
-
Make multi-step operations resumable or reversible
One of the two, chosen deliberately. Resumable means the state is recorded well enough to continue later. Reversible means each step can be undone. An operation that is neither has no safe stop in the middle, which is a finding worth surfacing before deployment rather than during an incident.
-
Record the stop as an event, with attribution
Who triggered it, when, against which agent, and why. This is both operationally necessary and part of demonstrating that oversight functioned. See AI audit trail, and note that a stop is exactly the moment a trace is most likely to be truncated.
-
Reconcile what was in flight
After the halt, establish what actually completed downstream versus what the trace recorded. The gap between those two is the specific damage a stop can cause, and finding it requires a deliberate check rather than an assumption.
-
Decide the default when you cannot stop safely
Sometimes the choice is between an unsafe halt and letting a bad operation finish. That is a business decision about a specific process, it should be made in advance, and it should be written down next to the process rather than improvised by whoever is on call.
One architectural rule that applies throughout. The stop has to be enforced outside the model's reasoning loop, where no instruction reaching the agent can bypass or argue with it. A halt the agent evaluates is a halt an injected instruction can suppress, which is the same principle set out at AI guardrails.
What are the three controls people conflate?
Pause, halt and revoke. Each interrupts something different, leaves the system in a different condition, and needs a different route back. Calling all three a kill switch is why teams find out mid-incident that they only ever built one of them.
Worth deciding which ones you actually have, because the answer is usually not all three.
| Control | What it stops | State it leaves | Recovery |
|---|---|---|---|
| Pause | New work only. In-flight work finishes | Clean, because nothing was interrupted | Resume. The safest control and the slowest to take effect |
| Halt | Everything, at the next safe-stop point or immediately | Possibly incomplete. Depends entirely on whether safe-stop points exist | Reconcile, then restart. This is what most people mean by kill switch |
| Revoke | The agent's ability to affect anything, by removing its credentials | The agent may still be running and can no longer do damage | Reissue credentials. The most reliable control, because it does not depend on the agent cooperating |
Revoke deserves more attention than it gets. A halt asks the agent's runtime to stop, which requires that runtime to be responsive. Revoking the credential asks nothing of the agent at all and works even when the process is wedged, looping, or ignoring signals. It is the control that holds when the others fail, and it depends on the agent having its own identity rather than sharing a service account, which is the argument at agent identity.
Why must a kill switch cascade?
Because stopping an orchestrator does not stop the agents it delegated to. Published analysis of the multi-agent case is explicit: a stop button at the orchestrator level does not satisfy the requirement if sub-agents continue executing independently. The stop has to propagate through the whole graph.
This is the failure most likely to be discovered in production, because a single-agent kill switch tested on a single agent appears to work perfectly.
- Sub-agents need their own stop path. If the orchestrator is the only thing listening for the signal, killing it orphans everything below rather than stopping it. Those agents keep working on instructions from a coordinator that no longer exists.
- Regulatory treatment reinforces this. The Digital Omnibus agreement of May 2026 clarified that multi-agent systems are treated as a single regulated system, so the oversight obligation applies to the whole graph rather than to each component separately.
- Cross-boundary delegation may be unreachable. Work handed to another organization's agent over the Agent-to-Agent protocol cannot be stopped by you, because the protocol has no stop primitive that reaches inside a remote agent. What you can do is stop sending, and revoke whatever access you granted.
- Revocation cascades better than halting. If sub-agents draw their access from credentials you issue, revoking at the source reaches every one of them at once, without any of them needing to be listening.
Why is an untested kill switch not a control?
Because it is a belief about your system rather than a property of it. Nobody tests these, for an understandable reason: testing means deliberately halting something in production. Which is exactly why it needs to be scheduled rather than left to the first real incident.
Treat it the way a serious organization treats disaster recovery, where an untested backup is assumed not to work until proven otherwise.
- 01
Run the drill on a scheduleHalt a real production agent, mid-trajectory, on purpose. What you are measuring: whether it stopped, how long it took, and what state it left. Quarterly is defensible; after any change to the agent's tools is better.
- 02
Time it against a targetPublished guidance suggests under five minutes for a production agent and under one minute where the agent has transaction authority. What you are measuring: the interval from decision to actual cessation, not from decision to somebody clicking.
- 03
Test with the person who would really do itNot the engineer who built the agent. What you are measuring: whether the business or risk owner can find the control, has the permission, and understands what it does, on a Sunday.
- 04
Check the state afterwards, deliberatelyReconcile what completed downstream against what the trace recorded. What you are measuring: the damage the stop itself caused, which is the number that tells you whether your safe-stop points are in the right places.
- 05
Rehearse the restartA stopped agent that nobody is authorised to restart stays stopped, and that is sometimes worse than the original problem. What you are measuring: who decides it is safe to resume, on what evidence, and how long that takes.
The restart question is the one most often left undefined and it belongs to governance rather than engineering. Somebody has to hold the authority to say the fault is understood and the agent may run again, and if that person is not named in advance the decision defaults to whoever is most senior and available, which is not the same thing as whoever is best placed to judge. See AI governance.