Home / Several agents and adversarial checking
Remove the tool before writing the prohibition
An instruction that prohibits steers what an agent attempts without preventing anything: the protection that holds first removes the tools its mission does not need, then the closed list of expected effects and the prohibition by class come second and third, as instructions.
Take a typical case. A review agent receives an instruction that forbids it two things: writing a file, sending a message. It has access to the terminal, the tool that runs system commands. During the review, it runs a command that transmits the document to a device on the network. The instruction did not name this effect, and the agent's configuration did not prevent it.
An instruction steers, it does not bound
Lengthening the list of prohibitions is the first reflex, and it runs into two limits. The list leaves permitted whatever it has not foreseen. Above all, an instruction remains a text the agent reads: the Claude Code permissions page says so of the guidance placed in CLAUDE.md, which steers what Claude attempts without imposing a boundary, and it recommends backing it with a mechanism that really applies. A prohibition written by class, any effect outside the entrusted folder, covers more cases than an enumeration, but it remains an instruction.
Three lines, in this order
The first line removes access. A Claude Code subagent declares in its header the tools field, the list of the only tools it may use, or the disallowedTools field, the list of tools removed from those it inherits. The two do not protect equally: the documentation states that an agent declared with disallowedTools: Write, Edit keeps the terminal and the tools of MCP connectors, and therefore the means to produce the effect in the case described.
---
name: reviewer
description: Reviews a document and returns its remarks in its answer
tools: Read, Grep, Glob
---
Your only deliverable is the list of your remarks, returned in your final answer.
Any effect outside this answer is outside your mission: file, message, command, device.
The second line is the closed list of expected effects, here a single one: the answer. It goes further than a prohibition by class, since what is not foreseen falls outside the mission without having been named, but it remains a text the agent reads, so an instruction that steers without imposing a boundary. The third line, the prohibition by class, serves as a readable reminder. An agent that must keep the terminal gets its first line another way: through the permission rules, and through the sandbox, an isolation imposed by the operating system on terminal commands, which the same page points to for protecting files and network without depending on the text of the command. According to its page, the sandbox runs on macOS, Linux and WSL2, and native Windows is not supported.
The lesson on the human checkpoint deals with the irreversible action that waits for approval; this one removes the opportunity upstream. The next module applies the same logic to commands, with an allowlist rather than a forbidden pattern.
What remains accessible to a review agent
| Configuration | Edit a file with the dedicated tool | Run a terminal command | Call an MCP tool |
|---|---|---|---|
| Instruction alone, inherited tools | Accessible, the instruction asks it to refrain | Accessible | Accessible |
| disallowedTools: Write, Edit | Removed | Accessible | Accessible |
| tools: Read, Grep, Glob | Removed | Removed | Removed |
A review agent is declared with a header that removes from its inherited tools those that write and edit files, and with an instruction that forbids it to write a file and to send a message. During the review, it runs from the terminal a command that transmits the document to a device on the network. The session log displays the command and its output.
Write in one sentence what this situation establishes, and in one sentence what it does not establish.
What this establishes: It establishes that the terminal remained accessible to the agent after the writing tools were removed, and that it used it for an effect that neither this removal nor the instruction covered.
What this does not establish: It does not establish that an instruction written by class, any effect outside the entrusted folder, would have prevented the command, nor that the document reached the device: the log shows the command and its output, not what the device did with it.
The three most common miscalibrations
- Too broad The situation establishes that the agent's configuration let terminal commands through, including those that write a file.
- Too narrow The situation establishes that the instruction lacked a prohibition on network devices, and that adding it is enough to close the case.
- Beside the point The situation establishes that the reviewed document contained a hidden instruction that pushed the agent to transmit it.
- A prohibition written in an instruction changes what the agent tries, not what it can do.
- A subagent's tools field fixes the only tools accessible, whereas disallowedTools removes a few tools from an inherited set that keeps everything else.
- A closed list of expected effects makes the unforeseen fall outside the mission without having to name it, while remaining an instruction that nothing enforces.
- When an agent keeps the terminal, the permission rules and the sandbox take over from the tool that could not be removed.
Open the header of a subagent you use, note the tools it inherits, replace them with a tools field that keeps only those its mission needs, then write in its text the list of effects it must produce.
Every datable claim in this lesson links here to the public text behind it. A source that does not open proves nothing.
- Claude Code, subagents, tools and disallowedTools fields, consulted on 2026-09-28 consultée le 2026-09-28
- Claude Code, configuring permissions, consulted on 2026-09-28 consultée le 2026-09-28
- Claude Code, command sandbox, supported systems, consulted on 2026-09-28 consultée le 2026-09-28