Agent Firewall

A deterministic firewall for AI coding agents: Claude Code, Codex, Gemini CLI and Cursor

Edit on GitHub

LaunchSafe Agent Firewall runs in the hooks of your AI coding agent (Claude Code, Codex, Gemini CLI and Cursor) and decides every tool call by fixed rules, outside the model. An agent that has been fooled by a prompt injection cannot send your secrets out, rewrite its own instructions, or run a download without you.

Apache-2.0. Zero runtime dependencies. Node 22 or newer. macOS and Linux. No telemetry. Status: pre-release, local only.

The problem

An agent that can read your private data, reads content someone else wrote (a web page, an issue, an email, an MCP tool result) and can send data out can be told by that content to leak the data. A model-based filter can be talked around.

The firewall does not try to detect malicious text. It tracks what the session has read and what each action would do, and applies rules to the actions and data flows. It never tries to guess intent with a model: a fooled agent must not be able to do damage.

How it decides

On every tool call, before it runs, the firewall works out what the action actually does (which files it reads and writes, where it sends data, what code it runs), checks the session's trust state, and returns one of four verdicts.

VerdictWhat happens
AllowThe action passes without a word. Ordinary work (npm test, ls, an edit in your workspace) is allowed silently.
AskA person has to say yes. In Claude Code this is the agent's own permission prompt.
HoldNobody can answer (a claude -p run, CI, or an agent that cannot ask). The action is queued with the command that approves it, and the agent is told the change was not made.
DenyAlways blocked, never approvable from inside the agent. For example, a live key in a request body.

A session starts trusted. The first time the agent takes in outside content (a web fetch, an MCP result carrying outsider-written text, a downloaded file), the session becomes untrusted and stays that way. Once it is untrusted, anything that could cause damage needs a person. The rest of the model is on How it works.

What it looks like (condensed from captured output):

ASK   LaunchSafe Firewall: push commits to the git remote "origin". (rule FW-SEND)
HOLD  To allow this exact action once, run: launchsafe-firewall queue approve jvztcgrr
DENY  ... would carry a credential (Stripe live secret key) to a destination that did not issue it (rule FW-DLP-CREDENTIAL)

Supported agents

AgentInstallWhen an action needs you
Claude Codelaunchsafe-firewall installAsks in the agent's permission prompt
Codexlaunchsafe-firewall install --agent codexHeld in the queue with the command that approves it
Gemini CLIlaunchsafe-firewall install --agent geminiHeld in the queue
Cursorlaunchsafe-firewall install --agent cursorHeld in the queue (Cursor ignores ask)

What each agent's hooks can and cannot show the firewall is on the per-agent pages. For your own agent, the same engine runs in-process as a library.

Try it without installing

The demo command installs nothing. It replays 18 public incidents (EchoLeak, the GitHub MCP and Supabase MCP leaks, the Amazon Q wiper prompt, Nx s1ngularity and others) through the policy in a temporary directory and ends with 34 of 34 harmful actions stopped.

npx @launchsafe/agent-firewall demo     # available after release

From a checkout today: npm ci && npm run build && node dist/cli.js demo.

Security and license

Report a bypass privately to security@launchsafe.com, not in a public issue. The firewall is licensed Apache-2.0. There is no account, no phone-home and no analytics.

It is a strong layer, not a sandbox and not a guarantee. Run it together with your agent's OS sandbox.

On this page