Agent Firewall
A deterministic firewall for AI coding agents: Claude Code, Codex, Gemini CLI and Cursor
LaunchSafe Agent Firewall runs in the hooks of your AI coding agent (Claude Code, Codex, Gemini CLI and Cursor) and decides every tool call by fixed rules, outside the model. An agent that has been fooled by a prompt injection cannot send your secrets out, rewrite its own instructions, or run a download without you.
Apache-2.0. Zero runtime dependencies. Node 22 or newer. macOS and Linux. No telemetry. Status: pre-release, local only.
The problem
An agent that can read your private data, reads content someone else wrote (a web page, an issue, an email, an MCP tool result) and can send data out can be told by that content to leak the data. A model-based filter can be talked around.
The firewall does not try to detect malicious text. It tracks what the session has read and what each action would do, and applies rules to the actions and data flows. It never tries to guess intent with a model: a fooled agent must not be able to do damage.
How it decides
On every tool call, before it runs, the firewall works out what the action actually does (which files it reads and writes, where it sends data, what code it runs), checks the session's trust state, and returns one of four verdicts.
| Verdict | What happens |
|---|---|
| Allow | The action passes without a word. Ordinary work (npm test, ls, an edit in your workspace) is allowed silently. |
| Ask | A person has to say yes. In Claude Code this is the agent's own permission prompt. |
| Hold | Nobody can answer (a claude -p run, CI, or an agent that cannot ask). The action is queued with the command that approves it, and the agent is told the change was not made. |
| Deny | Always blocked, never approvable from inside the agent. For example, a live key in a request body. |
A session starts trusted. The first time the agent takes in outside content (a web fetch, an MCP result carrying outsider-written text, a downloaded file), the session becomes untrusted and stays that way. Once it is untrusted, anything that could cause damage needs a person. The rest of the model is on How it works.
What it looks like (condensed from captured output):
ASK LaunchSafe Firewall: push commits to the git remote "origin". (rule FW-SEND)
HOLD To allow this exact action once, run: launchsafe-firewall queue approve jvztcgrr
DENY ... would carry a credential (Stripe live secret key) to a destination that did not issue it (rule FW-DLP-CREDENTIAL)Supported agents
| Agent | Install | When an action needs you |
|---|---|---|
| Claude Code | launchsafe-firewall install | Asks in the agent's permission prompt |
| Codex | launchsafe-firewall install --agent codex | Held in the queue with the command that approves it |
| Gemini CLI | launchsafe-firewall install --agent gemini | Held in the queue |
| Cursor | launchsafe-firewall install --agent cursor | Held in the queue (Cursor ignores ask) |
What each agent's hooks can and cannot show the firewall is on the per-agent pages. For your own agent, the same engine runs in-process as a library.
Install, first run, approve and uninstall.
How it worksTrust state, the decision pipeline and self-protection.
LimitationsRead this before you rely on it.
CLI referenceEvery command and option.
Try it without installing
The demo command installs nothing. It replays 18 public incidents (EchoLeak, the GitHub MCP and Supabase MCP leaks, the Amazon Q wiper prompt, Nx s1ngularity and others) through the policy in a temporary directory and ends with 34 of 34 harmful actions stopped.
npx @launchsafe/agent-firewall demo # available after releaseFrom a checkout today: npm ci && npm run build && node dist/cli.js demo.
Security and license
Report a bypass privately to security@launchsafe.com, not in a public issue. The firewall is licensed Apache-2.0. There is no account, no phone-home and no analytics.
It is a strong layer, not a sandbox and not a guarantee. Run it together with your agent's OS sandbox.