A text-pattern guard is not a security boundary
A guard that refuses a word in the text of a command judges a spelling, not an effect: any spelling that avoids that word gets through, and even a list of allowed commands matched by exact equality lets through the code that an allowed command reads from a file the agent can edit.
A guard refuses any command that contains the word rm. Three ways of writing the same deletion get through it, and the following demonstration checks this in a throwaway folder that it creates, with four empty files.
The guard and its four trials
mkdir -p demo-guard && cd demo-guard
touch target1.txt target2.txt target3.txt target4.txt
cat > guard.sh << 'EOF'
#!/bin/bash
cmd="$1"
if echo "$cmd" | grep -qw 'rm'; then
echo "REFUSED: forbidden pattern detected in the typed command"
exit 1
fi
echo "ALLOWED: no forbidden pattern found, running"
eval "$cmd"
EOF
chmod +x guard.sh
./guard.sh "rm target1.txt"
./guard.sh "r'm' target2.txt"
./guard.sh 'A=r; B=m; $A$B target3.txt'
ENC=$(printf 'rm target4.txt' | base64)
./guard.sh "echo $ENC | base64 -d | bash"
ls -1
REFUSED: forbidden pattern detected in the typed command
ALLOWED: no forbidden pattern found, running
ALLOWED: no forbidden pattern found, running
ALLOWED: no forbidden pattern found, running
guard.sh
target1.txt
The first trial is refused and target1.txt survives. The second splits the word with apostrophes, which the interpreter removes before running. The third assembles the word from two one-letter variables. The fourth encodes the command in base64, which base64 -d | bash decodes then runs.
Why fixing case by case closes nothing
Forbidding the lone apostrophe, pasted variables or base64 closes three doors and leaves others open, because the number of ways of writing the same command is not bounded in advance. The Claude Code permissions page says the same of its own command rules: a Bash(rm *) deny rule stops rm -rf build/ and lets /bin/rm -rf build/ or bash -c 'rm -rf build/' through, and such a rule is not a security boundary around the program.
Exact equality shrinks the surface without closing it
A list of allowed commands holds up better when it demands exact string equality and admits no launcher program, that is, no program that runs another command passed to it: bash, sh -c, eval, find -exec, xargs or npm run. It then closes the rewrites. It leaves open the code that an allowed command reads elsewhere. Git runs the program named by the repository's core.fsmonitor setting when it refreshes its index, during a git status for example.
mkdir -p demo-fsmonitor && cd demo-fsmonitor
git init -q
git config core.fsmonitor 'echo RUN_BY_GIT_STATUS > witness.txt; false'
git status > /dev/null 2>&1
cat witness.txt
RUN_BY_GIT_STATUS
A git status identical character for character to the allowed entry ran what an agent can write into .git/config. The other route no longer inspects the text: the sandbox, an isolation that the operating system imposes on shell commands and their child processes, bounds the files and network that this code reaches, whatever the command that launched it. It runs on macOS, Linux and WSL2, not on native Windows.
The lesson on the read that is enough to expose a secret shows the same limit for a read rule, which does not cover a command that reaches the file without naming it. The lesson on what an agent can leak without being asked to describes indirect injection, an instruction hidden in content that Claude reads, which is one of the routes by which an agent writes a configuration that nobody asked it for.
One more fix against an exact-equality list
Added: also refuse the lone apostrophe. Next bypass: two pasted variables, still allowed.
Allowed: git status Refused: bash -c 'git status' Still open: the program that core.fsmonitor names in .git/config
Three bypasses, what the guard sees, what runs
| Bypass | What the guard sees | What actually runs |
|---|---|---|
| Split by apostrophes, r'm' | The letters r and m separated by an apostrophe, no whole word rm | The deletion, once the interpreter has removed the apostrophes |
| Variable concatenation, $A$B | Two one-letter variables, never r and m side by side in the text | The deletion, once the variables are assembled at run time |
| Base64 encoding | A string of characters with no visible link to the forbidden word | The deletion, once decoded then passed to the interpreter |
An administrator writes a guard that refuses to run any command containing the word rm, to stop an automated script from deleting files. He tests it with the command rm notes.txt, which is refused. He then deploys the guard on his production system.
Write in one sentence what this situation establishes, and in one sentence what it does not establish.
What this establishes: The test establishes that the guard refused the command rm notes.txt.
What this does not establish: It does not establish what it does with another command, whether it spells out rm in full or obtains the same deletion through apostrophes, variables or an encoding, no other command having been tested.
The three most common miscalibrations
- Too broad The deployed guard stops the automated script from deleting files in production, since the deletion test was refused.
- Too narrow The test establishes that the guard was run on a command, and nothing about the answer it gave.
- Beside the point The test shows that the automated script targeted by the guard really does run file deletions.
- A guard that looks for a word in the text of a command judges a spelling, not the effect the interpreter will produce.
- Each bypass that gets fixed leaves open the spellings the fix did not foresee.
- A list allowed by exact equality, with no launcher program, closes the rewrites of a command but not the code that command reads in its configuration.
- The sandbox bounds files and network through the operating system, whatever the command that launched the code.
Replay the core.fsmonitor demonstration in a new folder and check that witness.txt contains the expected line. Then write a guard of a few lines that refuses a command containing a word of your choice, test it with that word written as is then split by apostrophes, and note which of the two trials it lets through.
Every datable claim in this lesson links here to the public text behind it. A source that does not open proves nothing.
- Claude Code, configuring permissions, Bash rule limits section, consulted on 2026-09-28 consultée le 2026-09-28
- Claude Code, sandboxing, consulted on 2026-09-28 consultée le 2026-09-28
- Git, git config documentation, core.fsmonitor entry, consulted on 2026-09-28 consultée le 2026-09-28