Limitations

The one list of what the Agent Firewall does not do. Read this before you rely on it.

Edit on GitHub

This is the one list of what the firewall does not do. Each item is stated once, plainly. The page for each part has the detail; this page has the summary. If a claim elsewhere in the docs reads stronger than an item here, this page is right and the other page has a bug: please report it.

The firewall was independently reviewed before each release, and the fixes are listed in its changelog. Nothing here is a finding that is still open without being named as open.

What kind of tool this is

  • It is a layer, not a sandbox and not a guarantee. The hooks decide whether a tool call runs. They do not contain a program that does run. Use the agent's own OS sandbox with it (doctor --fix for Claude Code, run for agents with no sandbox).
  • It cannot stop a model from being fooled. It limits what a fooled agent can do. It assumes the model can be persuaded.
  • It does not defend against you, or against code that already runs on your machine as you. A program running as your user, outside the agent's sandbox, that the firewall cannot see into, can read your files and your keychain.
  • Only approvals are bound to your passphrase. Session state, the provenance ledger, pins and counters are as strong as the analyser that guards the files they live in. The decision log is signed with a key in the firewall's own directory: anyone who can write that directory can rewrite the log. log verify proves integrity against everyone else. It catches an edited, reordered or truncated log, and a log file deleted or emptied while its recorded head (or the anchor of pruned files) remains; deleting the log and its head together leaves what a log that was never written leaves. Because a head left behind by a power cut is accepted, someone who can write the directory and kept an older signed head can cut the log back to that entry; keep an off-machine copy of the log if that matters to you.
  • A full disk does not stop the agent. Every decision is logged before its action runs; values an agent chooses are bounded, and an entry the log refuses is replaced by a short fallback entry. When the log refuses even that, an allowed action is asked about, or blocked with no one there (FW-LOG-UNRECORDED). An I/O failure is different and kept as it was: on a full disk or a failing drive the decision stands (a deny is still a deny, an allow still runs) and the error is printed on the hook's stderr, because blocking every tool, including the reads needed to free space, would leave you with an agent that cannot help. Such a decision has no log entry, and log verify cannot show the gap later.
  • Platforms: macOS and Linux, Node 22 or newer. Windows is not supported. The Linux suite has been run in a container; run on Linux has not been verified on a real machine with a real agent in this release.

What the firewall sees

  • Hooks only see what the agent reports. Files pulled into a prompt with @ never pass through PreToolUse, so they cannot taint a session. Treat an @-included file as input you chose.
  • A tool it has no model for is held. Coverage of exotic tools grows over time. Until then an unknown tool is sent to a person in any session.
  • In a trusted session it does not second-guess opaque code. An inline script that computes a path and writes there is held only once the session is untrusted. Inline code is judged read-only only by a list of known APIs; a write API missing from the list is the residual risk.
  • CDPATH is not modelled. A bare cd name is resolved against the current directory.
  • Symlinks are resolved as they are on disk when the hook runs. A link ln makes earlier in the same command is applied; one a script or an interpreter makes is seen at the next action. A path that cannot be resolved (a symlink loop, no permission) is judged as a protected one.
  • Shell writes to CI files, dev-container files and env-hook files in a trusted session are not checked. Only git hooks and editor tasks are promised. A custom core.hooksPath directory is not followed there.
  • Names only known at run time get no package verdict (npm i "$PKG").
  • Provenance hashes are compared when the hook runs. A process outside the agent could change a file between the hook and the tool's own read.
  • The first session in a workspace records receipts and taints for nothing. A freshly cloned repository's hooks and MCP servers are judged by the scan rules, not as drift.

Data leaving the machine

  • Secret fingerprints catch the documented encodings only. Verbatim, URL-encoded (once or twice), base64, base64url, base32, hex, and split across DNS labels are found. Reversed, case-changed, ROT13, compressed, encrypted, custom-alphabet, interleaved, or a few characters per request are not.
  • Slow exfiltration under a budget is not seen. Budgets bound how much can go to an allowed host; they do not detect it. Rotating subdomains get fresh per-host budgets (the total still holds). A web-search query has no host.
  • The taint rule, not pattern matching, is the main control on sending data out of an untrusted session. It depends on the firewall seeing the outside content come in.
  • The sandbox sees hosts, not methods. A host admitted for fetching (github.com, registry.npmjs.org) accepts uploads too. doctor labels such hosts. Uploads are controlled by the rules for commands the firewall can see (git push, npm publish, gh, docker push), not by the sandbox.
  • Hosts named in a prompt are allowed by the firewall for that session but not by the sandbox. A fetch to one fails inside the sandbox until you add the host to fetchAllowHosts and run doctor --fix again.
  • Package checks are only as good as their data. A malicious version published after your last update, or not yet in OSV, is not denied. A typosquat of a name outside the shipped popular lists is not flagged. The cooldown setting covers the first days of a release; npm and pip have no rolling cooldown.
  • Task scope is log-only by default (taskScope.mode: "log"). It records what it would have asked and changes no decision until you set "enforce". Enforce mode has known friction: tests next to a named file, a path that does not exist yet, paths with spaces. The gateway never sees a prompt, so scope does not apply there.
  • Loop detection can arrive fast for agents that cannot ask. Eight runs of the same command in ten minutes asks. Polling (gh run view, git fetch) can hit it. On Codex, Gemini CLI and Cursor every ask is held, so 25 holds in ten minutes make every non-read action a loop hit.

Agents

  • Each agent's hooks set the ceiling. Codex, Gemini CLI and Cursor show the firewall less than Claude Code does. None of the three can ask you: held actions wait in the queue. Codex and Gemini CLI let a tool run when a hook times out. Codex's hosted web search never reaches a hook. Codex cloud runs no local hooks. Cursor's cloud agents run only the repository's hooks, where a project install's paths do not exist (and, failing closed, block every action there). See the per-agent pages.
  • Codex runs its prompt hook only without a matcher. Codex 0.162.1 skips a UserPromptSubmit group that has any matcher, silently (captured). Installs from before the fix wrote "*", so on Codex the firewall never saw a prompt: no task scope and no destinations you named. doctor fails on such an install; run install --agent codex again. Until a prompt arrives, a Codex session has no task scope (today's rules apply) and no named destinations (a fetch after outside content is held).
  • Unattended Gemini CLI is detected from its command line only. Gemini CLI tells its hooks nothing about -p or yolo mode. The firewall reads the command line with ps, and with pgrep on macOS where a seatbelt sandbox forbids ps. If neither can run, the session is treated as unattended.
  • Gemini CLI's own sandbox (-s) needs the firewall's profile. Without it the hook cannot save its state there and blocks every action (one message says what to do). install --agent gemini --seatbelt permissive-open writes a profile that adds write access to the hook's state directories only, used with SEATBELT_PROFILE=launchsafe. The agent's commands share that sandbox, so the hook's state is protected there by the analyser alone, as without -s; the policy, keys, records and vault stay read-only. Re-run the command after updating Gemini CLI. mcp wrap inside -s and Gemini CLI's container sandboxes were not tested.
  • Cursor's formats were read from source, not captured live. So were Cursor's server-side tool names, its subagent events and its IDE hooks. They are marked in the fixtures.
  • Some behaviours of Codex and Gemini CLI were not captured. The firewall assumes the riskiest reading: whether Gemini CLI fires BeforeAgent for a subagent; whether Codex reports a shell-run apply_patch as Bash; the exact URL extraction of Gemini's web_fetch; the loose matching of Codex's apply_patch and Gemini's replace.
  • Tab completions, typed shell input, computer use and screen recording in Cursor carry nothing the firewall can judge. The last three are held as unknown tools.
  • A hook never sees an MCP server's full tool list. Without the gateway, description pinning covers only what Claude Code shows through its ToolSearch tool.
  • A repository can switch hooks off. Claude Code's disableAllHooks, Codex's [features] hooks = false and Gemini CLI's hooksConfig are set in files a repository can ship. Only managed settings resist this, and doctor fails when it finds one set. In a -p run a repository's hooks load without the trust dialog.

Approvals

  • No passphrase, no approvals. Installing without a terminal sets none. Held actions cannot be approved until approver init and a reinstall.
  • The commands that change state need a real terminal. They refuse to run from a script, from an agent, or as a descendant of a Claude Code process. A script or CI job of yours cannot approve either.
  • The phone signs what its app shows. A passkey prompt does not show the item. A compromised approval web app could show one pending item and have you sign another. The firewall accepts only a real pending item on this machine, with its unused nonce, inside its time window.
  • Synced passkeys report a counter of 0. Replay protection then rests on the single-use nonce and the firewall's own protection of the session file.
  • The phone sanitiser removes only known secret formats. A generic password or a -u user:pass argument can pass into a summary. Summaries are fixed sentences and counts for most rules; Bash summaries carry more.
  • Phone approvals need LaunchSafe's hosted relay, which is not available yet. Terminal approvals work now. Once the relay exists it is outside the firewall's repository, and with no network it will be unreachable while the terminal still works.

The vault

  • It is not a hardware boundary. A program running as you, outside the sandbox, that the firewall cannot see into can read the keychain: on Linux always, on macOS for secrets stored without per-use confirmation (--confirm never, the default for credentials).
  • Only curl and wget get values from a shell command. Other programs need vault allow <name> --process and run --env.
  • A bound host that echoes a request back, or redirects with curl -L (which resends -H headers), can return the value. What comes back is filtered in the documented encodings only, not after arbitrary transformation such as reversal.
  • A host binding compares the host name only. It ignores the port and a trailing dot.
  • Fields shorter than 12 characters cannot be fingerprinted (a CVC, an expiry date). vault add says which.
  • Not built: vault import --from-config and vault unlock or lock.

The MCP gateway

  • MCP only. Native tools (shell, file edits, built-in web fetch) never pass through it. Nothing covers Claude Desktop's or ChatGPT's own tools.
  • A server added directly bypasses it unless managed configuration forbids that. gateway status reports unwrapped entries; it cannot stop them.
  • Connectors called from a vendor's cloud (claude.ai connectors, ChatGPT apps) cannot be reached locally. .mcpb extensions and remote custom connectors in Claude Desktop cannot be wrapped.
  • Scanning is pattern-based. It finds hidden text and known agent-directed phrasings. It does not find paraphrased or semantic injections, or text in images, audio or blobs.
  • No conversation boundaries. Taint lasts for the client process (stdio) or until an idle gap (HTTP). That is stricter, never looser.
  • A new tool with no findings, inside the 24-hour grace window, is pinned silently.
  • Approval needs a retry. Elicitation asks are off: the list of verified clients is empty.
  • A Codex HTTP wrap drops bearer_token_env_var and http_headers. That breaks the server and fails closed.
  • gateway status --json prints the upstream URLs, including a secret in a query string that the client configuration already held.
  • Not built: HTTP for a legacy client's standalone GET stream, era translation of input requests, a login item for mcp serve, mcp pin --connect. Each wrapped stdio server costs one Node process (about 40 to 50 MB, not measured).

The email guard

  • Mail MCP tools only. Mail an agent reads another way (a shell command that prints a mailbox, a file, a vendor's cloud connector) is not split into messages. The taint rule still applies.
  • A server counts as mail by its name or its tools' names (mail, imap, gmail, outlook, smtp, and the words m365, o365, exchange, jmap, inbox, mailbox) or by emailGuard.servers. A mail server named otherwise, with tools named otherwise, is not guarded until you list it. exchange beside a market word (exchange-rates) is not counted.
  • On Codex, Gemini CLI and Cursor the model has read the message before the hook runs. The guard warns, records and untrusts the session. Only Claude Code's hook, and the gateway, keep it from the model.
  • Attachment contents are not scanned (PDF, DOCX, images). Names are.
  • The phrase lists are English only. Hidden-text rules do not depend on language.
  • Template-syntax exfiltration links and variation-selector smuggling are not covered.
  • Sender authentication is what the provider wrote. No signature or DNS check is made, and the rule looks at the top message, not forwarded ones.
  • Not built: mailbox-side quarantine (IMAP, Microsoft 365). When it is built it is a race after delivery and starts log-only.

Run

  • Hosts, not methods. An admitted host accepts uploads. Budgets bound how much; a fooled agent with an attacker's token can still push a small amount to an attacker's repository on an admitted host.
  • Domain fronting through a CDN. The SNI check stops a client that names another server in TLS. It cannot see the encrypted Host header, so a CDN that routes on it could serve another customer's site. ECH is not handled.
  • A name in run.privateHosts may resolve anywhere in the private ranges. The list is per name, not per address: whoever controls that name's DNS chooses which private address the proxy reaches.
  • No hooks means no taint tracking. No queue, no task scope, no payload fingerprints (TLS is not decrypted), no decoy tripwire on reads. Keep the hooks on an agent that has them.
  • Agents whose hooks run as child processes (Codex, Gemini CLI, Cursor) fail closed inside run, because the firewall's directory is read-only there. Use their own sandbox.
  • macOS uses a deny list. A system service not on it that does network or file work for its caller would not be stopped. sandbox-exec is deprecated by Apple. --allow-keychain and --allow-localhost widen what the process can reach.
  • Linux needs unprivileged user namespaces. Some distributions restrict them; doctor prints the error. Sockets outside the hidden directories stay reachable. A program that opens /dev/tty gets an error.
  • The read-only list for git and agent files is a list. A module that a hook script imports is not followed. A repository the agent creates inside the run is its own. Submodule and linked-worktree commands fail inside the run.
  • Caches in your home are read-only. npm, pip and many agents need --allow-write. /tmp is not writable.
  • No upstream proxy. An existing HTTPS_PROXY is replaced.
  • Other local processes can use the proxy on macOS (no credentials). It only reaches the policy's hosts, but they can use up the budgets.
  • If run is killed, an empty mount-point directory can stay behind on Linux (.mcp.json/, for example). Remove it.
  • Budgets are per run and counted on the wire. What stays under them is not seen.

Library

  • It does not run the checks that read a coding agent's own files: the session-start scan, MCP pins, desktop notifications, subagent linkage. It has no email guard. The provenance ledger and decoys work only with fileStore.
  • Approval is the host's job. The library never approves on its own, and a callback that is missing, throws or times out is a no.

Dashboard

  • It is read-only and as current as the page was when rendered.
  • A process running as you can read the browser history that holds the dashboard's private address. It gets the read-only view only. Approvals stay behind the passphrase. Run ui in a terminal the agent cannot read.

What is tested, and what is not

  • The decisions are tested: unit tests, the incident replays, the friction corpus and the mail corpus. A real model has never been talked into an attack against it in a test. The campaign harness drives real agent CLIs with a scripted model, so it tests the firewall's decisions on real hook traffic, not the model.
  • Latency is measured on a developer laptop. The budget is p95 under 50 ms and p99 under 150 ms for the decision and the warm hook path, and p95 under 120 ms and p99 under 200 ms for a whole hook process (node start included). It is not measured on slow or loaded machines.
  • The corpus counts (normal tasks, packages, mail messages) are the size of the corpora the LaunchSafe team wrote. A zero in them means none of those were interrupted, not that no ordinary task ever will be.

On this page