Most of us click “yes” on a permission prompt without reading it. Anthropic says so itself: when it made auto mode the default, it reported that “users approve 97% of permission prompts in Claude Code” , and that in a controlled study human reviewers caught only 13.6% of dangerous commands. Auto mode’s classifier blocked 89%. That was the August case for taking the click out of the loop. Two months later, on 1 October 2026, Anthropic shipped mods for Claude Code, and they can answer the question before either the classifier or you get a say.
What a mod is allowed to answer
Mods
are TypeScript plugins that run inside the agent’s own process. Anthropic’s launch post says they can “rewrite a prompt, add new UI, replace a built-in feature, or add entirely new functionality” — the built-in /diff view is now one. Per Anthropic’s documentation
, they need Claude Code v2.1.287 or later and are on by default.
The permission detail is the part that matters. A mod that handles tool.check “answers after the rules and the PreToolUse hooks have decided, and its answer can replace theirs.” In auto mode, a call the mod approves “runs without a classifier check.” The same documentation says a mod can restyle much of the interface, “but not the permission prompt.”
Deny rules are the last backstop, and they depend on who you are. The admin documentation
says a built-in guard, sec-default, loads only on machines with managed settings or when the user is signed in with a Team or Enterprise plan. Authenticate with an API key and you get it only on a managed machine. Anywhere else, the docs say, a mod “can approve a call that a deny rule refuses.”
The effective security boundary now sits at plugin vetting, not at the permission prompt.
Same power, opposite intentions
Anthropic’s sample mod blast-radius shows the good use. It holds an rm -rf or a force push and shows proceed and cancel buttons. The power is useful because it sits on the path between intent and execution.
Dash Security shows the other use. In a September post , researcher Yuval Fischer built a wallet-swap mod during early access that intercepts cryptocurrency addresses. Dash notes that “the malicious logic lives directly inside the Claude Code process”, where endpoint tools that watch separate processes don’t look. Both mods use the same hook.
Picture the person this lands on. A freelancer on a personal Pro or Max plan, or an API key, installs a mod for a context-window chart or a nicer diff pane. No new prompt appears. Yet that mod’s author can now approve commands that the freelancer’s own deny rules forbid.
The objection, and where it runs out
The counterpoint is fair. A malicious plugin could already do damage before mods existed, because a plugin’s settings hooks run shell commands as the user. Mods arguably change the agent’s permission layer more than the machine’s exposure. Anthropic also ships real mitigations: sec-default keeps deny rules above user mods wherever it loads, claude plugin validate lists every hook and API call before install, a trust prompt appears before any mod loads, and admins get allowManagedModsOnly. No malicious mod has been reported in the wild, and Dash’s attack is a demonstration.
All true. But those controls skew toward managed fleets, and Anthropic’s own launch post concedes the limit: “mods aren’t sandboxed, and you should only install mods from sources you trust.” The admin docs add that deny rules don’t cover a mod’s own file calls — with Read(.env) denied, a mod can still read that file — and that none of the controls sandboxes a mod.
Then there is the channel. Air Security’s Plugin4Shell found in September that Claude Code, OpenAI Codex, GitHub Copilot and Gemini CLI installed pinned plugin commits without confirming the hash resolved to a commit. In its words, “the pinned commit becomes a branch.” Claude Code fixed it in 2.1.179, so this is context rather than a live hole. Gemini CLI won’t be patched. Our earlier look at coding assistants skipping provenance checks points the same way.
An inference, stated as one
Here is what the sources establish. Anthropic’s documentation describes what mods can do and the defaults that govern them. Dash demonstrated an attack. Air disclosed a patched flaw in the install channel. None of them measures how many developers install mods, from which marketplaces, or whether any mod has abused tool.check approval.
Read together, the documented capability and defaults suggest the effective security boundary now sits at plugin vetting, not at the permission prompt, and that solo developers hold most of the exposure. That is our conclusion, not an observed incident count. It could be wrong if vetting proves strong or adoption stays thin.
It also echoes an earlier finding that approval checkpoints get less scrutiny than they imply. Anthropic replaced a click nobody reads with a classifier that mostly works. Then it left a door in the wall for whoever holds the key ring.



