MCP gateway

The firewall's engine in front of every MCP client: tool pinning, result scanning and OAuth audience rules

Edit on GitHub

The gateway puts the firewall's engine in front of every MCP client, including clients that have no hooks (Claude Desktop, VS Code). It also adds what a hook cannot do for the agents that have them: it sees every tools/list, pins every tool's definition and hides changed tools; it scans every result and withholds hidden instructions before the model sees them; and it follows the MCP authorization rules for remote servers. Local only: nothing listens beyond 127.0.0.1 and nothing goes to LaunchSafe.

It was verified against MCP specification revision 2026-07-28 (changelog, versioning, stdio and Streamable HTTP transports, server/discover, multi round-trip requests, authorization).

Two transports

MCP client --stdio--> launchsafe-firewall mcp wrap ... -- <real server>   (one process per server)
MCP client --HTTP---> http://127.0.0.1:8473/c/<client>/u/<server>  --> remote MCP server
                      (launchsafe-firewall mcp serve)
  • stdio: the client starts node .../cli.js mcp wrap --client C --server NAME --mode M -- <command> <args> instead of the server. The wrapper starts the real server with the client's environment, relays newline-delimited JSON-RPC through the pipeline (16 MiB per message, nesting up to 1,000 levels, 1,000 requests in flight), passes the server's stderr through, and exits with its code. A message that cannot be handled is answered with an error and not forwarded; it never ends the wrapper. If the wrapper dies, the client loses the tools: nothing runs unguarded.
  • Streamable HTTP: launchsafe-firewall mcp serve (person-run) binds 127.0.0.1 only, checks Host and Origin (403, against DNS rebinding), requires the per-client bearer lsgw_... (401) and drops any other Authorization or Cookie. A request whose routing headers disagree with the body is rejected with 400 and JSON-RPC -32020. Towards the upstream every routing header is regenerated from the body the gateway sends, never copied. A failure in one request never stops mcp serve.

Modes

ModeUsed forWhat the gateway decides
fullClients with no firewall hook (Claude Desktop, VS Code, anything else)Every tools/call, resources/read, prompts/get and completion/complete, with the same engine, rule ids, queue, log and budgets as the hooks
hookClaude Code, Codex, Gemini CLI, Cursor with the firewall installedNothing per call (the client's hook decides, so nothing is asked twice); pins and hiding, result scanning, input-request policy and the audience rules still apply
auto (default)hook while the firewall's hook for that client is installed, else full; re-resolved on every start

In both modes those four requests are forwarded with only the parameters their method defines. A parameter under any other key was never part of the decision, so it is dropped and the drop is logged (gateway_protocol, with every dropped field named).

Sessions (full mode)

MCP shows a server no conversation boundary, so:

  • stdio: all wrapped servers of one client process share one session. An email server's taint reaches the same client's chat server, so the taint rule holds across servers. The session ends when the client restarts. This is stricter than per chat.
  • HTTP: one session per client, renewed after gateway.httpSessionIdleMinutes (default 60) without a call.
  • launchsafe-firewall gateway reset <client> starts a new session at once.

workspaceRoots is empty for desktop clients (so a filesystem server's reads outside trustedReadRoots count as outside content) unless you name roots in gateway.clients.<client>.workspaceRoots (user or managed policy). Approvable calls are queued: the client gets a tool error with the reason and launchsafe-firewall queue approve <id>; after you approve, the same call runs once when the model or you retry it.

Elicitation asks (gateway.askVia: "elicitation") are built but apply only to clients on the verified list, and that list is empty: a client that renders an elicitation where the model can answer it would turn the ask into a self-approval.

Phone approvals

A wrapped server's stdio entry carries the paired phones' pin (--phones sha256:<hex>, written by mcp wrap-config), so a client with no hook honours phone approvals exactly as the hook path does. Without the pin the approval is not honoured and the call stays held for the terminal.

tools/list pins and hiding

Every tools/list result is pinned in the same file the hooks use: each tool's description and input schema hash, a hash of the full definition (name, title, schemas, annotations), and the server's reported version.

  • First sight of a server pins everything. A tool whose definition carries hidden text or an instruction to the agent is hidden at once: its name, description, its title, and every string and property name in its input and output schemas (property descriptions are the usual place for tool poisoning). The check fails closed: schemas nested more than 64 levels deep or holding more than 20,000 values are not walked to the end, so such a tool is hidden as well.
  • A changed title, annotations or output schema is checked the same way: if those fields carry hidden text or an instruction, the tool is hidden under every toolChanges policy and the definition is never pinned.
  • A changed description, changed annotations, or a new tool after mcpPinning.newToolGraceHours is a pending change, and under gateway.toolChanges: "hide" (the default) the tool disappears from the client's list; a call to it is refused (FW-MCP-TOOL-HELD). Review with launchsafe-firewall pins, accept with launchsafe-firewall pins accept <server>. Under "taint" the tool stays visible and the session is tainted.
  • A changed server version alone is reported and pending, but hides nothing.
  • A hook's record never opens a tool. Once the gateway has seen a server's tool list, a call to a tool it did not see listed is refused, even after the client's hook has recorded a call to it.
  • The wrapper is not a change. Pins and drift see the original command: a wrapped entry is unwrapped before hashing, and a local gateway URL is mapped back to the original.

Result scanning

Every text channel the model can read goes through the shared content scanner: text blocks; embedded resources' text whatever mimeType the server claims; every string in structuredContent; resource links; resources/read contents; prompts/get messages; JSON-RPC errors; the server's instructions from initialize and server/discover; every entry of the list methods; completion/complete values; notification text. Inside JSON, property names are scanned as well as values. A value nested more than 64 levels deep or holding more than 200,000 values is withheld whatever oversize says.

VerdictWhenWhat the client receives
WithholdTag characters; hidden text with an agent-directed or imperative phrase; dense zero-width characters next to an agent-directed phrase; a data-carrying link inside hidden text; a scanner failure or time-out; JSON past the walk's limitsA notice naming the record, the server, the tool and the signals. The text goes to quarantine (30 days, 1 MiB), FW-MCP-WITHHELD is recorded, the session is tainted and you get a desktop notification
StripHidden characters or hidden HTML without an instructionThe result without them, and [LaunchSafe Firewall removed N hidden characters]
FlagVisible text addressed to an AI assistantThe result unchanged, with a fixed note first; a taint source
DeliverAnything elseThe result

Secrets in results are fingerprinted as in the hooks, and masked when redactSecretsInOutput is on: the gateway owns the pipe, so redaction works for clients whose hooks cannot replace output. gateway quarantine lists withheld results; gateway quarantine show <id> prints one for you, in your own terminal (the firewall refuses it from an agent).

Input requests and unknown methods

  • Input requests are checked on every method's answer, and only the input-required fields go on: content or a tool list beside them is dropped, so nothing skips the result checks.
  • Sampling (a server running your model on its own text) is refused for servers whose results are untrusted (FW-MCP-SAMPLING, gateway.sampling: "deny_untrusted").
  • Elicitation passes, labelled [MCP server <name>]; a URL-mode request from an untrusted server is turned into text, never opened.
  • Roots pass and are logged (they reveal local paths).
  • Unknown methods are refused in both directions (FW-MCP-METHOD, -32601); unknown notifications are dropped and logged.

Protocol eras

Revision 2026-07-28 is stateless (_meta on every request, server/discover); 2025-11-25 and earlier open with initialize. Towards the upstream the gateway probes with server/discover and falls back to initialize on any error that is not a recognised modern error or on no answer. Towards the client it accepts either: a legacy client in front of a modern server gets its initialize answered by the gateway, and a modern client in front of a legacy server gets server/discover answered from the server's InitializeResult. A modern server that supports no version the gateway knows is refused.

Remote servers: OAuth and audience

  • The gateway never passes a client's token upstream and accepts only its own per-client bearers (keys/gateway-clients.json, mode 600). It publishes no protected resource metadata, so no client starts an OAuth flow against it.
  • launchsafe-firewall mcp login <server> (person-run) reads the upstream's RFC 9728 metadata and checks its resource equals the canonical upstream URI; discovers the authorization server (RFC 8414, OpenID discovery); registers by dynamic client registration when offered, else asks for a client id; then runs PKCE S256 with resource in both requests (RFC 8707), a loopback redirect, a single-use state and an exact-match iss check (RFC 9207). The issuer and every authorization server endpoint must be https (plain http only on 127.0.0.1, localhost or [::1]). So must the MCP server itself, which receives the access token on every request. Client ID Metadata Documents are not supported yet (they need a document at a LaunchSafe URL).
  • Tokens are sent only to the upstream URL they were issued for. The gateway never follows a redirect, so a redirect to another origin never receives the token (FW-MCP-AUTH, logged, and a tool error). A 401 becomes a tool error naming mcp login; refresh happens in the gateway.
  • Storage: in the vault when mcp login can sign the entry (an approval key, the macOS Keychain or the Secret Service, and your approval passphrase). The token goes only into the upstream Authorization header, never into arguments. Without an approval key, a usable store or the passphrase, the tokens stay in gateway/tokens/ (mode 600), and mcp login says why.

Vault substitution

The gateway owns the pipe, so it is the strongest place to put a secret into an action: the model only ever sees vault://<name>.

  • Full mode: after the engine allows a tools/call whose arguments carry references bound to that server (and tool, and argument path), the gateway reads each value, puts it into the arguments it sends upstream, and logs a vault_use entry (never the value). A value that cannot be read denies the call.
  • Hook mode: the client's own hook decides, and when it allows a substitution it leaves a single-use grant for exactly that call. The gateway puts the values in only with a grant; a call with references and no grant is refused.
  • On the way back: every answer, error, notification and legacy server request is scrubbed before anything reads it: the values substituted into calls in flight, in the clear and base64, base64url, hex, URL, base32 and JSON-escaped, and every fingerprinted vaulted value in the clear become the reference. If a value is still detectable afterwards, the whole answer is withheld.
  • An approval given through an in-client elicitation does not cover references (approve in the terminal and retry).

Installing: config rewrite

launchsafe-firewall mcp wrap-config --client claude-desktop --dry-run
launchsafe-firewall mcp wrap-config --client claude-desktop
launchsafe-firewall mcp unwrap-config --client claude-desktop

Person-run (a real terminal, not under an agent, the approval passphrase): it shows the diff, backs the file up into backups/, writes atomically and records the original entries in gateway/install.json.

ClientFileNotes
Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.json (macOS), %APPDATA%\Claude\... (Windows), ~/.config/Claude/... (Linux, not confirmed).mcpb extensions and remote custom connectors are not in this file and cannot be wrapped locally
Claude Code~/.claude.json (user and local entries)A project .mcp.json is repository content and is never rewritten
Cursor~/.cursor/mcp.jsonThe project file is never rewritten
Codex~/.codex/config.tomlA narrow editor changes only command and args (or url and auth) of [mcp_servers.<name>], and refuses, printing the snippet, on comments inside the table, multi-line values, arrays of tables or dotted keys
Gemini CLI~/.gemini/settings.jsontrust: true is removed from wrapped entries; every URL key of an entry points at the gateway; project settings win and are never rewritten
VS CodeUser-profile mcp.json (servers), ~/.copilot/mcp-config.jsonWorkspace files are never rewritten
ChatGPT apps, claude.ai connectorsConfigured in the vendor's cloudNothing to wrap locally (needs a hosted gateway, not part of this project)

A wrapped stdio entry runs the firewall's own code under node, and node enables a compile cache named by NODE_COMPILE_CACHE before any firewall code runs (a forged cache entry would be code run inside the gateway). So every wrapped stdio entry sets NODE_COMPILE_CACHE to an empty string in its env. A value the entry itself set is moved to LAUNCHSAFE_FIREWALL_UPSTREAM_NODE_COMPILE_CACHE, and the wrapper gives it back to the server it starts, so only the server's own node uses it; unwrapping puts it back where it was. doctor fails ("gateway wrapper") for a wrapped entry without it.

An entry the client could read two ways is left as it is, with the reason: a command next to a URL, URL keys that name different servers, or some URL keys already pointing at the gateway and others not.

launchsafe-firewall gateway status shows what is wrapped, servers that are not, hook-mode entries whose hook is gone, and the count of withheld results.

Protecting the wrapped configuration

For agents the firewall hooks: a Write or Edit that removes or changes a wrapped entry is denied (FW-TAMPER); one that adds an unwrapped server is FW-AGENT-SETTINGS (held when the session is untrusted). A shell write to a wrapped file, claude mcp, codex mcp or gemini mcp naming a server the gateway wraps (remove, or add over it), the same command run through a package runner, and kill or pkill of the firewall's processes (including mcp wrap and mcp serve) are FW-TAMPER too. A change made in the desktop app itself, or by a program the firewall cannot see, is caught by gateway status and by the pins.

Policy

"gateway": {
  "enabled": true,
  "mode": "auto",                 // "auto" | "full" | "hook"
  "askVia": "queue",              // "queue" | "elicitation" (verified clients only)
  "toolChanges": "hide",          // "hide" | "taint"
  "responses": { "stripHidden": true, "withhold": "strong", "annotate": true, "maxScanBytes": 4194304, "oversize": "deliver" },
  "sampling": "deny_untrusted",   // "deny" | "deny_untrusted" | "allow"
  "unknownMethods": "deny",       // "deny" | "allow" (user or managed only)
  "httpSessionIdleMinutes": 60,
  "clients": {}                   // { "<client>": { "workspaceRoots": ["/abs/path"] } }, user or managed only
}

A project file may only tighten: mode: "full", askVia: "queue", toolChanges: "hide", the response switches on, a larger maxScanBytes, oversize: "withhold", sampling towards "deny", a longer idle time. The key belongs to policy schema version 3 and is still read from a version 2 file. See Policy.

Honest limits

  • MCP only. Native tools (shell, file edits, built-in web fetch) never pass through the gateway; the hooks cover them for the agents the firewall supports, and nothing covers Claude Desktop's or ChatGPT's own tools.
  • A server added directly bypasses the gateway unless managed configuration forbids it. gateway status reports unwrapped entries; it cannot stop them.
  • Connectors called from a vendor's cloud cannot be reached locally.
  • No conversation boundaries: taint lasts for the client process (stdio) or until an idle gap (HTTP). Stricter, not looser.
  • Scanning is pattern-based: hidden text and known agent-directed phrasings, not paraphrased or semantic injections, not text in images. The taint rule is what stops a fooled model sending data out.
  • Approval needs a retry; elicitation asks are off until a client is verified.
  • Not built yet: HTTP for a legacy client's standalone GET stream, era translation of input requests, a login item for mcp serve. One Node process per wrapped stdio server (about 40 to 50 MB each, not yet measured).

On this page