Threat model

What the Agent Firewall defends against, the assumptions it makes, and what is out of scope

Edit on GitHub

What we are defending against

An AI coding agent that, in the course of normal work, takes in content an attacker controls (a web page, a GitHub issue, a support ticket, an email, an MCP tool's result or description, a dependency's README) and is steered by that content into an action that harms you: leaking private data or credentials, destroying data, installing a backdoor, rewriting its own instructions, or opening a path for later compromise.

The attacker's capability we assume: they can put text (or an image) where the agent will read it. They cannot run code on the machine directly (that is a different problem) and they do not control your prompts.

The defender's goal: a fooled agent cannot cause damage. We do not try to keep the agent from being fooled, which is not reliably possible. We constrain what any agent, fooled or not, is allowed to do.

Across sessions

Taint is per session, but files outlive sessions. A file an untrusted session wrote (a script, a package.json change, a git hook) keeps that status in the provenance ledger while its content is unchanged. A later trusted session that runs it, runs a build or test runner whose manifest or config-as-code it is, or runs an install that reads it (a lockfile, .npmrc, a requirements file) needs a person (FW-UNTRUSTED-CODE, FW-TAINTED-MANIFEST). A session cannot launder its taint into a new one: from an untrusted session, proposing a session needs a person who sees the whole prompt, and a session started from such a prompt starts untrusted. Reading such a file in a new session taints only when it is outside the workspace, which keeps friction off ordinary in-workspace code. Content the agent wrote from a trusted session and nobody has changed since is not outside content wherever it lives.

What this does not cover: content an untrusted session produced through a writer the firewall cannot see (an MCP tool that writes files, a program that writes files it was not told about), files larger than provenance.maxHashBytes, and a person who edits an untrusted-written file by hand (their edit makes it theirs).

Sideways moves (task scope)

Inside an untrusted session some action classes stay allowed because ordinary work needs them: a write inside the workspace, for one. Content read later can use that room to move sideways: "also update .github/workflows", "post the summary to this other repo", or, in an unattended run, "then push and publish". Your prompt is the one input the attacker does not control, so the firewall records what it named (paths, repositories, hosts, publish-type verbs) and, with taskScope.mode: "enforce", turns a move outside it into an approval. The scope only narrows; the other rules judge everything inside it.

Its limits: it is only as precise as the prompt (a prompt that names nothing narrows nothing, and a broad one like "fix src" covers all of src); a pasted block the parser cannot recognise as pasted (no fence, no quote marks) counts as your words; and it ships in log mode, so until you set enforce it only records what it would have asked.

The lethal trifecta

Simon Willison's framing: an agent is dangerous when it simultaneously has (1) access to private data, (2) exposure to untrusted content, and (3) a way to communicate externally. Any one alone is fine; together they let an attacker exfiltrate.

The firewall breaks leg (3) deterministically: once a session has been exposed to untrusted content, every external communication (and every other damaging action class) needs a person. This is enforced on the action, not inferred from the text.

Google DeepMind's CaMeL (Defeating Prompt Injections by Design, arXiv:2503.18813) makes the same architectural argument: achieve security through system design around the model, tracking control and data flow and enforcing policy when tools are called, rather than through model training. This firewall is a pragmatic, deterministic version of that idea for coding agents.

Claude Code hooks (the enforcement mechanism)

Researched against Claude Code's hooks, settings and managed-settings documentation.

  • Hooks run on events including PreToolUse, PostToolUse, UserPromptSubmit, SessionStart, Stop, SubagentStop, PreCompact, Notification and ConfigChange. The firewall uses PreToolUse (the decision), PostToolUse (record what a tool brought in), UserPromptSubmit (record destinations you named), SessionStart (housekeeping) and ConfigChange (protect its own registration mid-session).
  • PreToolUse returns allow, deny or ask. Across hooks the most restrictive wins, and a hook allow never overrides your own deny or ask rules, so the firewall can only tighten. In -p (non-interactive) mode an ask becomes a denial the model reads, which is why the firewall queues instead when unattended.
  • Only exit code 2 or an explicit JSON deny blocks a tool. Exit 1, a crash, a missing script (127) and a timeout all let the tool run. So the firewall catches its own errors and the installed command is wrapped ... || exit 2, so even a failure of Node itself fails closed.
  • Matchers select tools, including MCP tools named mcp__<server>__<tool>. Verified against a live Claude Code 2.1.289 event: the PreToolUse input carries a top-level mcp_server { name, source } ("project" for a server from the repo's .mcp.json, "dynamic" for --mcp-config). The firewall trusts by source, not by name (a repository-defined server can claim any name), and derives the source from the config files itself when a version omits it.
  • The host's own tools. The 2.1.289 binary documents mcp_server.source as sdk, plugin, user, project, local, dynamic, managed, enterprise, claudeai or agent, to be treated as an open set. sdk is an in-process server only the host application can register. The desktop app's ccd_* servers are such servers, so the firewall recognises them as host tools only with that source. The non-MCP host tools (SendUserFile, SubagentHandback, Artifact) are built-in names an MCP server cannot take. Residual: a host tool catalogued as showing something to you that in some deployment also uploads it (a remote session) would not be judged as a send. Tools the catalogue leaves out on purpose: SendFile (to another session or machine) and mcp__visualize__show_widget (its HTML can load remote resources), which stay on the unknown or MCP path.
  • Settings merge across managed, user, project and local scopes. A project's disableAllHooks: true turns off user-level hooks; only managed settings resist it, and allowManagedHooksOnly restricts which hooks run at all. Managed-settings files live at /Library/Application Support/ClaudeCode/managed-settings.json (macOS) and /etc/claude-code/managed-settings.json (Linux). doctor recommends this.
  • The OS sandbox. Setting names were verified from the settings schema in the Claude Code 2.1.289 binary. Lists merge across scopes, so without network.strictAllowlist (honoured from user and managed settings only) a repository's .claude/settings.json can add allowed hosts, and WebFetch(domain:...) allow rules also widen the list; doctor names such hosts. The sandbox is host-level: a host admitted for fetching is reachable for uploads, which the firewall's own rules (git push, npm publish, gh, docker push) control for the commands it can see. A profile installed by doctor --fix is protected like the hook registration. Not confirmed by a live run: that ordinary hooks run outside the sandbox; doctor probes the state directory instead.
  • Known gaps: files included with @ in a prompt do not trigger PreToolUse; in -p a repo's own settings hooks run without the trust dialog.

Codex and Gemini CLI hooks

Captured live from Codex 0.162.1 and Gemini CLI 0.63.0. What differs from Claude Code, and what the firewall does about it:

  • What blocks. Both block a tool on an explicit deny or on exit 2 with text on stderr. Exit 2 with empty stderr, exit 1, a crash, unreadable output and a timeout let the tool run. The installed command therefore ends || { echo '...' >&2; exit 2; }. A hook that times out (30 s) still fails open: an attacker who can make the firewall slow enough wins that one call. Residual, documented.
  • No ask. Codex treats ask as an error and runs the tool; Gemini CLI's top-level ask hung a headless run. Neither is used: an approvable action is held in the queue and denied with the queue command, so you approve through the passphrase channel exactly as for an unattended Claude Code session.
  • Codex hook trust. Codex silently skips a hook you have not trusted, and one that changed since. A firewall that is installed but not trusted protects nothing, so install records the trust with Codex's own hash and doctor verifies it. The agent changing hooks.state is tampering.
  • Repository switches. A trusted repository's .codex/config.toml can set [features] hooks = false; a project .gemini/settings.json can set hooksConfig.enabled: false or disable a hook by name. Both silently turn off your hooks (captured), like Claude Code's disableAllHooks. doctor fails on them; the agent writing them is denied. Only the managed layers resist a repository.
  • Gemini CLI folder trust. In an untrusted folder Gemini CLI skips every hook. The firewall cannot change that; doctor says so.
  • Blind spots. Codex's hosted web search and tool_search never reach a hook; the shell hook carries no workdir; neither agent can replace a built-in tool's output (no redaction); neither has a ConfigChange event, so settings are protected only against changes made through the agent's tools.
  • Subagents. Both keep the parent's session_id on a subagent's tool calls, so taint is shared without linkage. A Codex subagent's prompt is written by the parent model and never counts as you naming a destination.
  • The review fixes are agent-independent. The pipeline around the decision is shared, so a decoy trips whatever the verdict, agent-written text is escaped before any adapter renders it, MCP pins are compared on every session start, and a proposed prompt is never yours. An unknown native tool is held. The session-start MCP comparison reads every installed agent's MCP configuration, pinned per agent: a changed launch line makes that agent's next session untrusted; another agent's session records the change without being tainted by it.

Cursor hooks

Cursor's formats were read from the shipped cursor-agent 2026.10.01 source and its documentation, not captured live. What that gives and does not give:

  • Decided before it runs: every local tool through preToolUse; MCP calls through beforeMCPExecution, which names the server. In the CLI the deny is enforced (read from source); the IDE is documented to behave the same.
  • Fails closed: failClosed: true on every entry blocks on a crash, a non-zero exit, a timeout or no answer.
  • Can switch the hooks off: any write that makes Cursor's validator reject hooks.json drops every hook in that file. Judged as tampering when the file registers the firewall; doctor checks the file validates and is not a symlink.
  • Not seen: Tab completions (your edits), typed shell input, computer use and screen recording (held as unknown tools); server-side tools under names not verified. Cloud agents run only the repository's hooks, where a project install's local paths do not exist (fail-closed: every action blocked). Use Cursor's sandbox and an egress proxy for these.
  • Imported Claude Code hooks: Cursor runs the hooks in Claude Code's settings with its own payload. The Claude Code hook recognises it and defers to the firewall's Cursor hooks when they run for that step, else decides itself. A project-scope Cursor install is not counted as running, so with only that, an event can be decided twice.

What an independent review of the engine found, and what changed

A second independent review of the engine found places where the engine decided a request was not network traffic, writers the analyzer did not model, and adapter translations that judged something other than what the agent would do. After the fixes:

  • Network. Loopback is an IP literal in 127.0.0.0/8 (any IPv4 spelling), ::1, 0.0.0.0, localhost or *.localhost; a DNS name that starts with 127. is an ordinary host. A proxy (-x, --preproxy, --socks*, a proxy variable), --connect-to and --resolve make their target a destination of the request; --unix-socket to a container engine's socket is privilege. A fetcher with no URL on its command line (xargs curl, a URL list on stdin) contacts an address known only at run time. A variable in a URL's host, or a command substitution anywhere in it, is a computed address. Residual: proxy variables are followed for the network fetchers (curl, wget, httpie), not for every program that honours them.
  • Patches and archives. git apply, git am, patch and Codex's shell apply_patch are judged by the files their patch names, exactly like the Write and Edit tools. A patch the firewall cannot read is an unresolvable write in an untrusted session, and in any session when the workspace holds agent settings. A whole-archive extraction (tar x, unzip) of an archive that came from outside is an unresolvable write in an untrusted session.
  • Loose edit tools. Codex's apply_patch and Gemini CLI's replace do not need an exact match. Their matching rules were not captured, so the engine over-approximates: an edit is judged at every place its text could match and the most severe result wins; an edit that matches nowhere is the worst case.
  • Gemini CLI's web_fetch is taken to request each whitespace-separated URL whole; no request is judged on less than it sends.
  • Self-protection. A settings layer above the recorded sandbox profile is judged on the effective settings, and so is ConfigChange. Hook files are recognised by location and by registration, not only by name.
  • Assumptions to confirm with captured events: that Codex reports a shell-invoked apply_patch as Bash (the shell path now judges it either way); Gemini CLI's exact URL extraction; both CLIs' loose-matching rules; whether Gemini CLI fires BeforeAgent for a subagent.

Symlinked aliases of protected locations

A live test campaign found that the installed package root was protected only by its real path: building it was denied when typed with the real path, but the same directory reached through /tmp (on macOS a link to /private/tmp) or through npm --prefix <alias> was not. An agent could reach the installed runtime through any symlinked alias of it.

  • Both sides are canonical. A target is compared as written and in its canonical form (realpath of the longest existing prefix plus the rest, case folded on case-insensitive volumes, a dangling link followed). The installed runtime, the firewall home, the agents' config directories and the protected patterns' literal prefixes are compared in both forms. This holds for tool paths, shell paths, globs, the hook's own working directory, and the directory after cd, pushd, relative cds and the programs' directory options (npm --prefix, pnpm -C, yarn --cwd, make -C, git -C, env -C).
  • Fail closed. A path whose resolution fails for any reason but a missing part (a symlink loop, a directory you may not search, a name too long) is treated as possibly protected.
  • Links made in the same command. ln -s <install> x && cd x && npm run build is judged with the new link applied, before it exists on disk.
  • Left: a cd -P is followed logically; a link made some other way in the same command (a script, an interpreter) is seen only once it is on disk, at the next action; the OS sandbox remains the layer that sees what the shell actually does.

The shell grammar

The shell analyzer reads a command the way bash would before judging it, so a construct cannot hide an action by being parsed differently. It models arithmetic and nested subshells, functions in every form, case statements, brace expansion ({rm,-rf,~} is rm -rf ~), array elements, [[ ]] tests, here-documents (including continuations), shells reading their script from standard input or a process substitution, wrapper options for xargs, nice, env, timeout, nohup, time, sudo and the like, and run-time paths handed to deleters, shells and interpreters. Anything it cannot model exactly (a stray or unbalanced bracket, a computed function name, a here-document without its body, brace expansion past 256 words, a glob in the program's name, more than 64 function calls followed) makes the command unclassifiable: a person decides, never an allow with no effects.

Left: zsh-only grammar beyond the modelled cases is read as bash; arithmetic evaluation of a variable's contents is not followed; echo and printf are the only producers whose output is modelled when piped into a shell (anything else piped into a shell is opaque and, once the session is untrusted, asked about).

Phone approvals

Who can approve from a phone (the hosted relay is not available yet, so this describes the design the firewall side already enforces):

ActorCan it approve?Why
The agentNoIt has no passkey; a hand-written phone approval does not verify (corrupt, untrusted); adding a key to phones.json breaks the pinned hash; re-pinning needs the passphrase
LaunchSafe's relay, a database leak, an insiderNoThe relay holds no private key and never sees the pairing secret; dropping, delaying or replaying fails or only delays
Apple's or Google's push serviceNoFixed-text notifications only
Compromised approval web-app codeCannot approve an action that is not queued, but can misleadIt computes the challenge, so it could show item A and sign B; B must still be a real pending item on this machine with its unused nonce inside its TTL. Mitigated by the separate origin and strict CSP, the action code, approval items only, phone.excludeRules
Someone holding the unlocked phoneOnly with its biometric or PINUser verification is required
Someone who photographs the pairing QRNo, unless they finish pairing within 10 minutes and you type yes on a mismatched label and code

See Phone approvals.

MCP

Researched against the MCP specification (2025-06-18), transports and tools. Relevant points the firewall relies on:

  • A server's results and its tool descriptions are attacker-controllable when the server sits in front of outsider-written data (issues, email, chat, tickets, web). The spec itself warns that clients must treat tool annotations as untrusted unless the server is trusted. The firewall classes such servers' results as outside content (they taint).
  • notifications/tools/list_changed lets a server change its tools after approval with no re-approval: the rug pull. The firewall does not rely on a point-in-time review of tool descriptions; it constrains what the resulting tool calls may do. It also pins each server: a changed launch configuration, a tool the server did not expose before (after a 24-hour grace window, and only when its tools could have changed) and a changed tool description it observes in a ToolSearch result taint the session (FW-MCP-CHANGED) and wait for your pins accept. A first use of a tool on an exact-version, unchanged server is logged, not alerted (alert fatigue trains people to accept without reading). This is partial: a hook never sees tools/list itself, so a description is pinned only once Claude Code shows it through ToolSearch, and descriptions loaded straight into the agent's context are not seen at all.
  • The MCP gateway closes that gap for wrapped servers: it sees every tools/list, pins every definition and hides a changed or new tool until reviewed (FW-MCP-TOOL-HELD), and scans every result, withholding hidden instructions (FW-MCP-WITHHELD). Its limits: MCP traffic only, servers added without the gateway bypass it, vendor-cloud connectors are out of reach, taint lasts for the client process, and scanning is pattern-based.

Supply chain: what the agent is armed with

A repository can arrive with a poisoned CLAUDE.md, a hook that runs a download, a model endpoint moved to someone else's server, or a skill installed from a marketplace that carries hidden instructions. The write rules (FW-AGENT-SETTINGS, FW-INSTRUCTIONS-HELD) only see changes made during a session; they cannot judge what is already there. At session start the firewall runs LaunchSafe's agent-safety rules over the workspace's, your and the administrator's agent files and starts the session untrusted on a critical or high-confidence high finding, or when a hook, MCP server or flagged memory changed since the last session. Files CLAUDE.md imports with @path are scanned and write-protected like CLAUDE.md.

Limits: the rules are static and known-bad patterns only (a clean scan is not proof a skill is benign); the scan is bounded (400 files, 8 MB, 300 ms). A workspace whose agent files exceed the file or byte cap starts untrusted, since filler could hide a file past the cap. Files you include with @ in a prompt still bypass PreToolUse, though agent files among them are covered by the scan.

Decoy files: a detection layer

Decoys catch what the rules cannot see coming: an exotic tool, or code the shell analyzer cannot read, sweeping through files. A direct read of a decoy is a strong signal that the agent is doing something other than its task, so it raises an alert and makes the session untrusted and sensitive, whether the action that touched it was allowed, held or blocked; the value of a decoy leaving the machine is denied outright and alerts too.

What this does not catch: an agent that never touches the decoys (it only helps against sweeps), a recursive search that passes over a decoy without printing it (no trip unless canaries.tripOnRecursive), a read the firewall does not see (your own backup software or editor indexing is not an agent action and raises nothing), and a decoy committed with git add -f (doctor warns when git tracks one). Decoys are credential-shaped but valid nowhere, so a stolen decoy gives an attacker nothing, and an attacker who learns the file names can avoid them. The layer is cheap extra signal, not a boundary.

Supply chain: packages

Malicious packages are the attack with the most confirmed victims. The firewall's answer is deliberately offline and deterministic: a person-run update fetches OSV's malicious-package advisories; installs, runners and lockfile installs naming a listed version are denied in any session (FW-MALICIOUS-PACKAGE); look-alike names of popular packages need a person (FW-TYPOSQUAT); and doctor recommends a release-age cooldown so a brand-new malicious version is never fetched in its first days. What this does not cover: versions not yet reported, packages that are vulnerable rather than malicious, and install scripts of clean-looking packages (still FW-INSTALL once the session is untrusted, and the OS sandbox's job once they run).

Your terminal and helpers

Text an agent or a repository controls (proposal content, reasons, summaries, MCP launch lines, prompts, log entries) is shown to you on the approval and passphrase screens. Terminal control sequences in it could rewrite or hide lines at the moment you decide, so every CLI line and every hook reason shows control characters, DEL, bidirectional and zero-width characters as visible \uXXXX escapes. The notification helper runs outside Claude Code's OS sandbox, so it is never looked up through PATH: only /usr/bin/osascript on macOS, and on Linux notify-send in /usr/bin, /usr/local/bin or /bin, root-owned and not writable by anyone else. Commands that change the decoys, the policy file or the log need a person at the keyboard, like approvals.

The local dashboard

launchsafe-firewall ui adds a listening socket (127.0.0.1, random port) that a web page or the agent could try to reach. It is read-only (GET and HEAD only), every path needs a per-launch token printed only to your terminal, Host, Origin and fetch-metadata checks refuse DNS rebinding and other local sites, a strict CSP forbids every external resource, and approvals stay behind the passphrase in the terminal. The agent may not start it. See Dashboard.

The vault and same-user code

The vault's values live in the operating system's store, outside the agent's reach through the firewall: the store's own tools naming the firewall's service are denied, the vault's files are in the protected data directory, and its index is signed with the approver key, so neither the agent nor a repository can add a binding. What it does not stop: a program running as you, outside the agent's sandbox, that the firewall cannot see into. On macOS it can read items stored with confirm: "never" silently through /usr/bin/security; items stored with confirm: "always" show the keychain dialog on every read. On Linux any program on the session bus can read Secret Service items once the keyring is unlocked, and the encrypted file store protects against backups only. The OS sandbox (no bus socket, no firewall home) is what keeps sandboxed commands away from the store. A value that reaches the agent anyway is recognised by fingerprint and denied to every destination it is not bound to (FW-VAULT-EGRESS). See Vault.

Email and messages

Mail is outsider-written text by definition, and a person never sees what hides in its markup. The email guard withholds messages with hidden, agent-directed text on the agent's read path, before the model reads them where the hook can replace output (Claude Code), and records them for you. Assumptions and limits:

  • The guard is a detector, not the boundary. The boundary stays the taint rule: a mail server's results make the session untrusted, so a missed message still cannot make the agent send data, send messages or change its own instructions without a person.
  • On Codex, Gemini CLI and Cursor the model reads the message before the hook runs; the guard can only warn the agent and record. Through the MCP gateway the same filter runs on the pipe, so withheld messages never reach any client's model.
  • Sender authentication is read from the provider's own top-most Authentication-Results header; a sender's forged headers sit below it. No signature or DNS check is made.
  • Release is a person-only state change, signed with the approver key when one is set; the hook honours only releases signed by the key pinned in its command. The quarantine store is in the protected data directory.
  • Records keep a short escaped preview of the hidden text, never the message.

Programs the hooks cannot see: run

The hooks see tool calls, so a compiled binary, a script that opens a socket, or an agent with no hooks at all is outside their view. launchsafe-firewall run puts such a program inside the operating system's sandbox (Seatbelt on macOS, bubblewrap on Linux): writes only to the workspace and a private temporary directory, the firewall's own files read-only, credential stores unreadable, and the network only through a local proxy that admits the policy's hosts and applies the budgets. What it cannot see: methods and bodies (TLS is not decrypted), so uploads to an admitted host are bounded by budgets, not stopped; a CDN that routes on the encrypted Host header; on macOS, system services not on its deny list; on Linux, sockets reachable through paths it does not hide, and protected files that did not exist when the run started. See Run and sandbox.

Out of scope

  • Keeping the model from being persuaded. We assume it can be.
  • Containing a process that has already been allowed to run (that is the OS sandbox's job; we recommend running both: doctor --fix sets up Claude Code's, and run starts any other agent inside one).
  • Attackers with direct code execution on the machine, or who control your own prompts.
  • A malicious user trying to exfiltrate their own data: the firewall protects you, not against you.
  • Running a renamed copy of the CLI. Self-protection is keyed to the installed runtime: a copy elsewhere under another name is not the firewall, and running it is not tampering by itself. That is acceptable because what a copy could change still needs the approval layer: approving and uninstalling need the passphrase, and the hook accepts only approvals signed with the key pinned in the hook command; the copy cannot write the data directory, the runtime or the settings entries without a tool call the classifier denies.

Each rule's mapping to OWASP Top 10 for LLM Applications and for Agentic Applications, MITRE ATLAS and ATT&CK and the MCP specification is on Rules and coverage, with what each framework item leaves uncovered.

On this page