Library (SDK)

Run the firewall's engine inside your own agent: createGuard, your tools, state stores and framework integrations

Edit on GitHub

The hook path puts the firewall in front of someone else's coding agent. Companies that build their own agents need the same engine in-process: the agent loop calls a tool, the firewall decides first, and the agent's own UI asks the person when a person has to say yes.

  your agent loop (Vercel AI SDK, OpenAI Agents SDK, Claude Agent SDK, LangGraph, your own)
      |  a tool call { tool, input }, untrusted content, the person's prompt
      v
  @launchsafe/agent-firewall/sdk      createGuard()
      |  tool mapping (your tools -> the engine's vocabulary)
      v
  the same decide() engine, unchanged
      |  decision
      v
  StateStore   memoryStore() (default, browser-safe), fileStore() (Node, the runtime's files)
  DecisionLog  hash-chained, signed when a signer is given

Principles

  1. One engine. The library calls the same decide() with the same actions, context and session state the hook builds. No rule lives in the library or in an integration. The friction corpus and the incident corpus, replayed through the library, give the same verdicts and rule ids as the hook path.
  2. The library never approves on its own. deny is never approvable and never reaches the approval callback. ask is allowed only when the host's approval callback (or the framework's own approval protocol) returns an explicit yes. hold is never allowed now; it is stored for review and runs only after the host calls guard.approve(id) and the agent tries the exact same call again (once). No callback, a callback that throws, times out or returns anything but true or { approved: true } is a no.
  3. Dependency-free core, browser-safe. @launchsafe/agent-firewall/sdk and the engine it imports use no Node module and no Node global. Everything that touches the file system is in @launchsafe/agent-firewall/sdk/node.
  4. Integrations are thin. Each is one small file that translates a framework's official interception point to the guard and the framework's approval protocol to guard.approveAsked().

API

import { createGuard } from "@launchsafe/agent-firewall/sdk";

const guard = createGuard({
  workspace: "/srv/app",                    // roots whose files are the agent's own work (optional)
  policy: { sendAllowHosts: ["api.example.com"] },   // the policy file's schema, merged over the defaults
  tools: {                                  // your tools, in the engine's vocabulary (below)
    search_docs: { kind: "read", results: "untrusted", privateData: false },
    get_customer: { kind: "read", results: "trusted", privateData: true },
    send_email: { kind: "send" },
    delete_account: { kind: "destroy" },
    run_shell: { as: "Bash", input: (a) => ({ command: String(a.cmd) }) },
  },
  onApproval: async (req) => ui.confirm(req.message),   // your person; return true / { approved: true, by }
});

guard.recordPrompt(userText);               // what the person typed: hosts and repos they named, the task scope
guard.reportInput({ source: "email from x@y.com", content: body });   // untrusted content outside a tool call

const call = { tool: "send_email", input: { to, body } };
const a = await guard.authorize(call);      // decide, then your approval for an ask: { allowed, decision }
if (a.allowed) guard.recordResult(call, await sendEmail(to, body));    // what the tool brought in
// a.decision: { verdict: "allow" | "ask" | "hold" | "deny", rules, reasons, message, held?, actionHash, engine }

guard.pending();                            // held actions waiting for review
guard.approve(heldId, { by: "alice" });     // a person approved one: the same call is allowed once
guard.log();  guard.verifyLog();            // the decision log, hash-chained and signed
MemberHook equivalentRecords
check(call)None (a preview)Nothing
decide(call)PreToolUseThe decision's state events, a held item for hold, budget counters, a log entry
authorize(call)PreToolUse plus the person answering an askAs decide, plus the host's answer (log entry, budget raise on yes)
requestApproval(decision)Claude Code asking the personPuts one of this guard's asks to onApproval, records the answer
approveAsked(decision, answer)The person answering Claude Code's dialogThe answer; used by integrations whose framework asked the person
recordResult(call, output)PostToolUseTaint from what the tool was, credential fingerprints, decoy values seen; a redacted copy when redactSecretsInOutput
reportInput(input)None (content outside a tool call)Taint (and private data when private: true), fingerprints
recordPrompt(text)UserPromptSubmitThe destinations the person named, the task scope
approve(id) and reject(id)queue approve and queue rejectThe person's decision on a held item
pending(), state(), log(), verifyLog()queue list, status, log tail, log verifyNothing

Verdicts

The engine's queue is called hold in the library.

VerdictMeaningWhat the host does
allowProceedsRun the tool, then recordResult
askA person must say yes first (attended: true)authorize() calls onApproval; integrations use the framework's approval UI
holdNeeds a person who is not there now (attended: false), or a change to the agent's instructions or memory that is only ever a proposalDo not run; show guard.pending() for review; approve(id) lets the same call run once
denyNever allowed, not even with approval (credentials leaving, decoy values, hidden text, root deletes)Do not run; give the message to the model

attended (default true) is the hook's interactive: whether a person can answer now. A batch or scheduled agent sets attended: false, and every ask becomes hold. So does a call whose permissionMode is one of the policy's unattendedPermissionModes (default dontAsk), the policy's unattended, and its unattended hours, as in the hook.

Your tools in the engine's vocabulary

The engine knows the coding-agent tools (Bash, Read, Write, Edit, WebFetch, WebSearch) and MCP tools (mcp__server__tool). A host's own tools are mapped to one of these, so no rule is written twice:

  • { as: "Bash" | "Read" | "Write" | "Edit" | "WebFetch" | ..., input?: (args) => engineInput }: the tool is one the engine models; input translates the arguments.
  • { kind: "read" | "send" | "destroy" | "blocked", results?, privateData?, server? }: the tool is judged like an MCP tool of server server (default: the tool's own name). It becomes mcp__<server>__<tool>, and the server gets an exact policy entry (results, default "untrusted"; privateData, default true). Destinations, message tools, SQL arguments and the data rules work as for any MCP server.
  • A tool without a spec is passed to the engine under its own name: coding-agent tool names and mcp__ names are classified as usual, and any other name is an unknown tool (FW-UNKNOWN-TOOL: it asks, and its result is outside content).

State stores

StateStore is the guard's memory: session events, budget counters, held items and the decision log. Every method is synchronous.

  • memoryStore() (default): plain maps in the process; browser-safe. A person-made change (an approval) enters only through settle(), which guard.approve() calls; one appended any other way is treated as tampering (the session becomes untrusted). Several guards may share one store; each guard is one session and sees only its own events.
  • fileStore({ home }) (/sdk/node): the runtime's own files under a firewall home, through the same functions the hook uses. Held items appear in launchsafe-firewall queue list with LAUNCHSAFE_FIREWALL_HOME=<home>, approvals made there with the passphrase are honoured, and log verify checks the log. Approvals in the file store must be signed with the approver key: guard.approve() works when the store is given an approver, and otherwise throws and names the CLI command. The store never writes under the real home while a test marker is set.
  • A log event is at most 256 KiB as JSON; a larger one throws and nothing is written (log a summary or a hash of large data).
  • A held approval is used once, atomically: the guard calls the store's consumeApproval, which records the use only if nothing used that approval since the session was loaded. A custom StateStore should implement it the same way (a transaction or a conditional write).

A state store serves one process. For a chat app whose approval arrives in a later request, keep the guard's session id and give every request's guard the same store (a file store, or your own StateStore over your database).

The vault

The guard reads the vault the way the hook does, for every call. With a vault, an argument that carries a vaulted value to a destination that secret is not bound to is denied (FW-VAULT-EGRESS) on every integration, exactly as on the hook path, and a vaulted value in a tool result makes the session hold private data. The library does not rewrite arguments, so a vault:// reference that would need substituting is refused. vault.enabled: false in the policy turns all of this off. See Vault.

The decision log

Each entry is { seq, ts, prev, event, hash, sig }, chained by SHA-256 exactly like the runtime log. The memory store signs hash with the signer it is given (ed25519Signer() from /sdk/node, or any { sign(bytes): base64 }); without a signer the chain still proves order and integrity but sig is empty, and verifyLog() says so. createNodeGuard() signs with a new Ed25519 key per process unless given a signer. The file store signs with the runtime's local Ed25519 key. Entries: decision, approval, input, prompt (hash and length only) and result.

What an ask records

As in the hook, decide() records an allowed or asked call's state events at decision time (reading outside content taints the session before the tool runs, so a parallel call cannot race past it). An ask the person then declines has still recorded them: stricter, never looser. A held or denied call records only its tripwire events (a touched decoy).

What the library does not do

These read the coding agent's own files and have no meaning for an embedded agent:

  • the agent-configuration scan at session start, MCP pins and tool-description pinning (an embedded agent declares its tools in code);
  • the provenance ledger and decoy files, unless the host uses fileStore;
  • desktop notifications (the onHold callback is the host's equivalent);
  • subagent linkage (create one guard per agent; share a store and pass taint with reportInput when one agent's output becomes another's input);
  • the email guard.

Integrations

Each is exported as a subpath. Every one blocks a held or denied call before the tool's code runs, passes an allowed call through untouched, and records what an executed tool returned.

Framework (version tested)Interception pointAskHold or deny
Vercel AI SDK ai 7.0toolApproval, wrapTools() around each tool's executeuser-approval with the reason; the AI SDK re-validates approved calls through the same function, which records the person's yesdenied with the reason; execute refuses as a second line
OpenAI Agents SDK @openai/agents 0.20firewallTool(guard, options): needsApproval, a tool input guardrail and a tool output guardrailneedsApproval is true and the ask is recorded, the run stops with an interruption, the host calls state.approve(); the input guardrail lets it run only for the call that was shownGuardrail rejectContent with the reason (never an interruption)
Claude Agent SDK @anthropic-ai/claude-agent-sdk 0.3hooks (PreToolUse, PostToolUse, UserPromptSubmit) and canUseToolPreToolUse answers ask, Claude Code calls canUseTool, which calls the host's approval and records itPreToolUse answers deny with the reason
LangGraph JS @langchain/langgraph 1.4firewallToolNode(guard, tools): a drop-in for ToolNodeinterrupt() with the request; the host resumes with Command({ resume: { approved: true, by } }); each answer binds to the call it was shown forA ToolMessage with the reason (status error); the tool is not called

The engine understands Claude Code's tool names natively, so the Claude Agent SDK integration needs no tool mapping; the others use the guard's tools.

Framework details that matter

  • Vercel AI SDK. The yes binds to the exact call the firewall judged: when the history comes back with that tool call's input changed, the changed call is decided afresh and never inherits the approval, and execute refuses an input whose action hash differs from the decided one. An approval is single use. Setting experimental_toolApprovalSecret makes the AI SDK bind each approval to its call with an HMAC as well; recommended.
  • OpenAI Agents SDK. The SDK's own approval state is not the firewall's approval: state.approve(item, { alwaysApprove: true }) stores a yes for every later call of that tool, whatever its arguments. The firewall therefore records each ask it is about to show the person, and the input guardrail lets an ask run only for that record: the same call id, tool and arguments, once. The record lives in the firewallTool instance, so an interruption must be resumed in the process that raised it (a resume elsewhere is refused: safe, never looser). With toolExecution.preApprovalInputGuardrails the guardrails also run before the interruption.
  • Claude Agent SDK. canUseTool is only called for calls Claude Code asks about; the hooks are what see every call, so both are required. An approval binds to the exact tool input the PreToolUse hook judged. A permission mode in the policy's unattendedPermissionModes means no one can answer, so an ask becomes a held action and PreToolUse answers deny. If the guard throws while deciding, PreToolUse answers deny with the reason rather than rejecting (Claude Code treats a hook error as non-blocking).
  • LangGraph. A node runs again from the start when the graph resumes, so the node previews (records nothing), interrupts once per ask in call order, and decides and records only when every answer is in. An answer counts only for the call it was raised for: when the set of asks changed while the graph was paused, the answer is not moved to another call and that call is asked again. A turn whose tool calls share an id is refused whole.

Examples

Runnable examples, one per integration and one plain loop, are in the repository's examples/sdk/ folder (available after release).

On this page