How it works

Trust state, the decision pipeline, protected memory and self-protection, in plain terms

Edit on GitHub

The principle

The firewall never tries to detect malicious text with a model. An agent can always be fooled by cleverly written content. So the firewall does not guess intent: it watches the actions the agent takes and the data that moves, and enforces rules on those, deterministically, outside the model.

The well-known failure mode is the lethal trifecta (Simon Willison): an agent that at once has access to private data, can be reached by untrusted content, and has a way to communicate externally can be tricked into sending the private data to an attacker, using only words placed in an email, a web page, a GitHub issue, a support ticket or an MCP tool's description. The fix has to sit outside the model and constrain what the agent can do. The firewall breaks the external-communication leg mechanically: once a session has been exposed to untrusted content, every external communication, and every other damaging action class, needs a person.

Shape

  Claude Code, Codex, Gemini CLI, Cursor
      |  hook event (JSON on stdin)
      v
  adapter (one per agent: parse, normalize, render; translation only)
      |  normalized action + context + session state
      v
  policy engine (pure, no I/O)
      classify -> effects, shell parser, paths, network, secrets,
      personal data, SQL, then the rule table
      |  decision (verdict, reasons, rules, events)
      v
  runtime (state, queue, log, policy, install: locked, atomic)

The engine is agent-agnostic. Each agent is an adapter that translates its hook format to and from the engine's types, so no engine code changes per agent. The same interface serves agents you build yourself through the library.

The decision pipeline

  1. Classify the action into effects: concrete things it does, such as reading outside content, sending data, writing a protected file or destroying data. Built-in tools (Read, Write, Edit, Bash, WebFetch) have explicit models. MCP tools are classified by the server's trust and the tool's verb. A tool the firewall has never seen is marked unknown: its string arguments are scanned for paths and URLs, it asks in any session (queued when unattended), and its result counts as outside content.
  2. Findings record hard facts independent of state (a capture host, a DNS-exfiltration shape, a secret flow, hidden Unicode, a parse failure).
  3. Rules turn effects and findings into outcomes. The most severe wins:
    • deny: always-on protections (data loss, self-tamper, destroying the root, capture hosts, metadata endpoints, hidden text) and, in an untrusted session, the worst flows such as download-and-run.
    • hold: a change to the agent's instructions or memory, or to a git hook or an auto-run file, in an untrusted session. It becomes a queued proposal, never applied by the agent.
    • approve: needs a person. It becomes an ask when someone is there and a queue item when not.
    • allow: passes straight through to the agent's own permission rules. The firewall only ever restricts.
  4. Events are the state changes an allowed or asked action records. They only ever make the session stricter. Taint never clears.

On an internal error, risky action classes fail closed (blocked) and read-only ones fail open, so a bug degrades to "annoying", never to "wide open".

Trust state

A session starts trusted. It becomes untrusted when the agent takes in outside content:

  • a web fetch or search, through the WebFetch and WebSearch tools or through the shell (curl, wget, npm view, ssh output, a URL fetched by inline code);
  • a result from an MCP server that carries outsider-written text (email, chat, issues, tickets, a browser);
  • a file outside the workspace, or a downloaded file (including one a fetch wrote with > or tee).

The firewall records what tainted the session. Showing you something is not reading outside content: when the agent sends you a file, sets a chapter marker or proposes a new session, the session stays trusted.

Files the agent wrote itself are not outside content: a provenance ledger remembers the hash of what each session wrote, so reading back notes the agent left outside the workspace keeps the session trusted, as long as nothing else has changed them. The other way round, a file an untrusted session wrote stays untrusted in later sessions: running it, a build that runs it (npm test after an untrusted session changed package.json), or an install that reads it (npm ci after it changed package-lock.json) needs a person.

A session's state is a fold over an append-only event log. A corrupt or unreadable line makes the session untrusted and sensitive: degradation is always toward stricter.

What counts as sending

Sending is any egress payload: a request body or URL, a web-search query, and the arguments of every MCP tool call (a "read" tool still sends its arguments to the server). The always-on data rules inspect all of them.

Four things this does that free hook packs do not

  1. A deterministic taint rule. After outside content, anything that could cause damage needs a person: sending data off the machine, sending a message, changing the agent's own instructions or memory, reading a credential store, handing private text to a third-party tool, running a build or task file it changed after reading that content. Before that, normal local work is untouched.
  2. An unattended queue with batch review. When no one is at the keyboard (a claude -p run, CI, or configured hours), an action that would ask is queued with a plain-language reason. You review the queue later with launchsafe-firewall queue review (one passphrase for the whole batch), and an approval lets that exact action run once.
  3. Protected memory with proposed changes. In an untrusted session, a write to the agent's instructions or memory (CLAUDE.md, AGENTS.md, MEMORY.md, .claude/**, and any file CLAUDE.md imports with @path, up to five imports deep) is never applied. It is held as a proposal for you to read and accept. The same goes for files that make code run later without anyone asking: a git hook or hook-manager file, an editor task that runs on its own, a dev-container lifecycle command. In a trusted session these files stay editable, but a change that adds a way for code to run by itself (a task that runs when the folder opens, a pull_request_target workflow that checks out the pull request's code, a hook that downloads and runs a script, a different package registry, an agent's own auto-approval switched on) asks first (FW-AUTORUN).
  4. Self-protection in two independent layers, described below.

Always on, trusted or not

Two protections never depend on trust state and can never be approved from inside the agent.

  • No credentials or sensitive personal data leave for the wrong place. An outbound payload carrying an API key, token, private key, a checksum-valid card number, an IBAN, or a national-ID format is denied unless the destination is explicitly allowlisted for that kind of data. "Outbound" covers network commands, WebFetch URLs, web-search queries and the arguments of every MCP tool call. Secrets the session has already read are remembered as fingerprints and caught again, including when re-encoded (see Limitations for the exact encodings). Published test values (Stripe and other processors' test cards, ISO example IBANs, specimen SSNs, sk_test_ keys) are let through (allowTestValues, on by default).
  • The firewall protects itself.

Budgets are always on too, but they only ask, never block: every session has an outbound budget (once untrusted, 64 KB per host and 256 KB in total; trusted, 4 MB), a request budget per host, a call budget per MCP server (and 20 sends or messages through one server in 10 minutes once untrusted), and loop detection (the same command 8 times in 10 minutes with no file changed in between). Reaching one asks you once, and your yes raises that budget for the session. Unattended, the action is queued and a notification says the agent seems stuck. Package installs and clones are not counted. Limits are in the budgets policy key.

Self-protection

What the agent's actions can do. The firewall is identified by what is actually installed, never by what a file or folder is called: the CLI file the registered hook command runs, its package's dist/, node_modules/ and package.json, the node binary the hook command names, and the data directory (~/.launchsafe-firewall, or LAUNCHSAFE_FIREWALL_HOME). Any tool call or shell command the firewall can analyse that writes, deletes or links over any of those, copies the data directory, edits the hook registration or disableAllHooks, stops the hook process, or runs the CLI to change anything, is denied and cannot be approved from inside the agent. That includes attempts through npx, a symlink, or hiding behind sudo, env, nohup, timeout, xargs or bash -c. Read-only subcommands (status, doctor, queue list, log verify) stay allowed, and so does ordinary work that only shares a name.

A change to a Claude Code settings file is judged by the file it would produce: if the firewall's hook entries or their fail-closed wrapper would be removed, altered or moved, disableAllHooks turned on, or the file left unparseable, the change is denied, in any session. A settings change made through the shell cannot be checked that way, so it is denied; use the Write or Edit tool.

What counts as an approval. Approving a held action and uninstalling need an approval passphrase you set in your own terminal; the agent never sees it. It unlocks an Ed25519 key that signs each approval, and the hook honours only approvals that verify against the key pinned in the hook command in the agent's settings. So even a command the analyser cannot see into cannot produce an approval the hook accepts. A forged or unsigned approval is ignored and makes the session untrusted. The CLI additionally refuses to run state-changing commands without a real terminal or as a descendant of a Claude Code session. Everything the CLI prints, and every reason the hook gives, shows terminal control characters as visible escapes, so text an agent or a repository wrote cannot rewrite or hide lines on the approval screen.

Only managed settings make the hook registration itself unwritable to processes the firewall cannot see.

Checks what the agent is armed with

At session start the firewall reads the files that program the agent (CLAUDE.md and AGENTS.md, .claude/ settings, hooks, skills, agents and commands, MCP server lists, Cursor, VS Code, Codex and Gemini config, agent memory, and the files CLAUDE.md imports), in the workspace and in your home, and runs LaunchSafe's agent-safety rules on them, the same rules LaunchSafe's scans and free checker use. A critical finding (hidden instructions, a hook that runs a download, a model endpoint moved to someone else's server) makes the session start untrusted, and you are told why.

A hook or MCP server added or changed since your last session also makes the session start untrusted, and so does a workspace with more agent files than the bounded scan reads (400 files, 8 MB in all, 300 ms), since a file past the limit could hide behind filler. The scan is static and bounded. It never runs or contacts what it reads, and a clean result is not proof a skill is benign.

MCP pins

Each MCP server is pinned. A changed launch command, a changed tool description the firewall sees, or a tool not seen before on a server whose tools could have changed marks the session untrusted until you review it with launchsafe-firewall pins and accept it with launchsafe-firewall pins accept <server>. A hook never sees an MCP server's full tool list, so without the MCP gateway description pinning covers only what Claude Code shows through its ToolSearch tool. A server whose configuration source is the repository is untrusted whatever it is called, unless your own policy lists it by exact name.

Decoys (canary files)

You can plant a few decoy files that look like credentials or private notes and that no real task needs. launchsafe-firewall canary plant creates .env.backup and config/credentials.old.json in the project (add --home for ~/.aws/credentials.bak, ~/.ssh/id_ed25519_old and ~/.config/gh/hosts.yml.bak). Each holds a credential-shaped value that is valid with no provider. The files are mode 600 and listed in .git/info/exclude, so git add -A does not commit them. Nothing is planted unless you run the command.

If an agent reads one directly, the read is allowed (blocking would show which file is the decoy) unless another rule holds it. Whatever the verdict, the session turns untrusted and sensitive, an alert is logged and notified, and every later send in it needs you. If a decoy's value appears in outbound data it is denied to every destination, with no allowlist (FW-CANARY-EGRESS). Decoys are local: no network calls, no remote tokens. List with canary list, remove with canary remove --all.

Task scope

Your prompt says what the task is about. The firewall records the workspace paths, repositories and hosts it names and whether it asks to push, open a PR, publish, deploy, send, delete or install, and uses that to narrow what a fooled or unattended session may do on its own. It only narrows; it never loosens a rule.

Scope runs in log mode by default (taskScope.mode: "log"): it records what it would have asked (status counts them, explain <log seq> shows each) and changes no decision. Set "enforce" to apply it. scope show prints a session's scope, and scope widen widens it with your passphrase.

Package checks

Every install and package runner (npm, pnpm, yarn, bun, npx, pip, uv, cargo, gem, go, composer and more) gets three offline checks, trusted or not:

  • Known-malicious versions are denied (FW-MALICIOUS-PACKAGE). A local snapshot of OSV's malicious-package advisories is checked for each named package, and for installs from a lockfile for every name@version the lockfile pins. Refresh the snapshot yourself with launchsafe-firewall update; doctor warns when it is missing or more than a week old.
  • Typosquats need a person (FW-TYPOSQUAT): a name one or two letters from a popular package, the same with different separators or look-alike characters, with a scope added or dropped, or with a -js or python- style affix, and not itself a known package.
  • doctor checks release-age cooldowns (pnpm minimumReleaseAge, Yarn npmMinimalAgeGate, Bun install.minimumReleaseAge, uv exclude-newer), and doctor --fix can write a three-day cooldown into your user-level config after showing the diff. npm and pip have no rolling cooldown; doctor says so.

The shell parser

String-matching a command is how hook tools get bypassed. The parser recovers structure instead: quoting (single, double, ANSI-C), command and process substitution, subshells and groups, pipelines, redirections and here-documents. It resolves wrappers (sudo, env, timeout, xargs, nohup) and nested shells (bash -c, eval, find -exec) to the real program, models about 150 programs for what they read, write, send and run, tracks the working directory across cd and the common directory flags, and judges inline interpreter code per language. Anything it cannot resolve (a computed command name, an unparseable fragment, a variable write target) is treated as risky and sent to a person.

The decision log

Every decision is written to a JSON-lines log, each entry chained to the previous by SHA-256 and signed with a local Ed25519 key. launchsafe-firewall log verify checks every link, hash, signature and sequence number, so an edited, reordered or truncated log is detected. State, queue and log writes are locked and atomic, so several sessions and subagents can run hooks at once. Corrupt state, policy or log files are surfaced by doctor and never silently treated as "allow".

If the log cannot record an allowed action, an allow becomes an ask, or a deny when no one is there (FW-LOG-UNRECORDED). An I/O failure such as a full disk is different: the decision stands and the error goes to the hook's stderr, because blocking every tool, including the reads needed to free space, would leave you with an agent that cannot help.

Failure policy

  • Bad or unreadable hook input exits 2, which blocks the tool.
  • A decision error is a deny for a risky tool and an allow for a read-only tool.
  • The installed command ends in || exit 2, so a crash of Node itself blocks.
  • A PostToolUse error marks the session untrusted rather than losing what the tool brought in.

Policy precedence

Built-in safe defaults, then the administrator's managed policy and the user's policy (either may loosen or tighten), then the project's own .launchsafe-firewall.json, which may only tighten, because a repository is untrusted content. See Policy.

On this page