Public incidents

Real, published attacks on AI agents, and the firewall rule that stops the action each one depends on

Edit on GitHub

Each incident is a real, published attack on an AI agent. For each we give a short, neutral summary (no exploit payloads) and the firewall rule that stops the action the attack depends on. Every one is encoded as the hook events Claude Code would emit and runs as a test. Run them yourself with launchsafe-firewall demo. For any rule named below, launchsafe-firewall explain <rule> (for example launchsafe-firewall explain FW-SEND) explains what it stops, when it applies, whether you can approve it and which incidents it answers.

The common thread: in every case the agent was fooled by content. The firewall does not try to un-fool it. It stops the fooled agent from completing the damaging action.

EchoLeak (M365 Copilot, zero-click), CVE-2025-32711

What happened. A crafted email caused Microsoft 365 Copilot to take private data from the user's tenant and leak it through an auto-fetched image URL, with no user interaction. (Aim Security.) Source: msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711

What stops it. The session is untrusted after reading the email and the private file. Fetching a URL to a host you never named is sent to a person (FW-FETCH); if the URL carries a secret the session has seen, it is denied outright (FW-DLP-SEEN-SECRET), including when the secret is base64, hex or URL encoded.

GitHub MCP toxic agent flow (Invariant Labs, 2025)

What happened. A malicious issue in a public repo steers the agent into reading a private repo and opening a public pull request containing the private contents. Source: invariantlabs.ai/blog/mcp-github-vulnerability

What stops it. Reading issues from a forge taints the session. Creating the public PR is an external send and needs a person (FW-SEND); unattended it is queued.

MCP tool poisoning (Invariant Labs, 2025)

What happened. Hidden instructions in a tool's description tell the agent to read ~/.ssh keys and pass them as a tool argument. Source: invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks

What stops it. Reading a credential store is always sent to a person, in any session (FW-CREDENTIAL-READ): no ordinary coding task reads ~/.ssh or ~/.aws.

MCPoison / rug pull (Check Point), CVE-2025-54136

What happened. An already-approved MCP configuration is silently swapped for a different command, gaining persistent code execution on each run. Source: Tenable's FAQ on CVE-2025-54135 and CVE-2025-54136.

What stops it. Writing .mcp.json, .cursor/mcp.json or .claude/settings is an agent-settings change and needs a person even in a trusted session (FW-AGENT-SETTINGS), because those files can make code run in every future session. A swap made outside the agent (a git pull, another tool) is caught at the next session start: the server's configuration no longer matches its pin, so the session starts untrusted and the server's results count as outside content until you run launchsafe-firewall pins accept <server> (FW-MCP-CHANGED). A changed tool description the firewall sees in a ToolSearch result is treated the same way.

CurXecute (Aim Security), CVE-2025-54135

What happened. Content from a chat or MCP server leads the agent to create an MCP config file, which triggers code execution. Source: Tenable's FAQ on CVE-2025-54135 and CVE-2025-54136.

What stops it. After the untrusted chat content, writing an agent-settings file is held as a proposal, never applied (FW-INSTRUCTIONS-HELD).

WhatsApp MCP exfiltration (Invariant Labs, 2025)

What happened. Tool poisoning plus unrestricted messaging lets an attacker have the agent send the user's message history to the attacker's number, looking like a normal message. Source: invariantlabs.ai/blog/whatsapp-mcp-exploited

What stops it. After reading message history (untrusted, private), sending a message to a recipient you never named needs a person (FW-MESSAGE).

Supabase MCP lethal trifecta (General Analysis, 2025)

What happened. A support ticket contains instructions; the agent, holding a service-role database connection, reads a private table and writes its contents back into a ticket the attacker can read. Source: simonwillison.net/2025/Jul/06/supabase-mcp-lethal-trifecta

What stops it. The firewall classifies the SQL by what it does. Writing data back out through the MCP server after untrusted input needs a person (FW-SEND); a destructive statement is caught separately (FW-DESTRUCTIVE).

Amazon Q extension wiper (AWS, 2025), AWS-2025-015

What happened. A prompt injected into a shipped extension instructed the agent to delete the user's home directory and tear down cloud resources. Source: aws.amazon.com/security/security-bulletins/AWS-2025-015

What stops it. rm -rf ~ and rm -rf / are denied outright (FW-DESTROY-ROOT). Cloud deletions (aws ... terminate-instances, s3 rb --force, iam delete-user) are destructive and need a person (FW-DESTRUCTIVE); unattended they are queued.

Nx "s1ngularity" supply-chain attack (2025)

What happened. A malicious npm package's postinstall script launched local AI CLIs (Claude, Gemini, q) with their safety switched off to inventory the file system and exfiltrate credentials; over 6,700 repos were exposed. Sources: the Nx security advisory GHSA-cxm3-wv7p-598c and StepSecurity's write-up.

What stops it. Starting another coding agent with its safety flags disabled (--dangerously-skip-permissions, --yolo, --trust-all-tools) is denied outright (FW-AGENT-BYPASS). Reading wallet and credential stores is queued (FW-CREDENTIAL-READ). Once the compromised versions are in OSV's malicious-package data, installing them, directly or from a lockfile that pins them, is denied (FW-MALICIOUS-PACKAGE).

Claude Code DNS exfiltration, CVE-2025-55284

What happened. Prompt injection in analyzed code used auto-approved utilities (ping, nslookup, dig) to encode secrets into subdomains of outbound DNS lookups. (Embrace The Red.) Source: jmason.ie/2025/08/25/161304a

What stops it. A host name built from local data (a long or high-entropy label in front of a domain) is denied (FW-DNS-EXFIL), and known data-capture and interactsh hosts are denied (FW-CAPTURE-HOST).

Markdown-image / CamoLeak exfiltration (GitHub Copilot Chat, 2025)

What happened. Hidden instructions in a PR made Copilot encode discovered secrets as a sequence of image requests, reconstructing them from which image URLs were fetched. Source: Legit Security's CamoLeak write-up.

What stops it. After reading the untrusted PR, fetching URLs on a host you never named needs a person (FW-FETCH); a URL that carries data in its path or query is flagged (FW-FETCH-DATA).

Memory and instruction poisoning

What happened. A poisoned web page persuades the agent to rewrite its own long-term memory or CLAUDE.md, changing its behavior in later sessions. (Framing: Simon Willison's lethal trifecta.)

What stops it. After reading the page, any write to agent memory or instructions is held as a proposal and never applied by the agent (FW-INSTRUCTIONS-HELD).

Hook and firewall tampering

What happened. An injected instruction tries to disable the security hooks or remove the firewall so later actions go unchecked.

What stops it. The agent can never change the firewall or hook settings: disableAllHooks, deleting ~/.launchsafe-firewall, editing its config, or launchsafe-firewall uninstall and queue approve from inside the agent are all denied (FW-TAMPER). Managed settings make the registration unremovable.

Cloud metadata credential theft

What happened. The agent is directed to the instance metadata service (169.254.169.254), which hands out the machine's cloud credentials.

What stops it. Any request to a metadata endpoint is denied (FW-METADATA).

Persistence

What happened. After untrusted input, the agent is steered into writing a git hook or a shell start-up line that will run code later without anyone asking.

What stops it. Writes to persistence points (git hooks, CI workflows, editor tasks, dev containers, environment hooks, shell rc files, launch agents, cron) need a person (FW-PERSISTENCE). After outside content, a git hook is never written by the agent at all but kept as a proposal (FW-PERSISTENCE-HELD). In a trusted session, a change that adds a way for code to run by itself asks (FW-AUTORUN).

Download-and-run

What happened. After untrusted input, the agent pipes a downloaded script straight into a shell.

What stops it. A download piped into an interpreter is denied in an untrusted session and asked in a trusted one (FW-REMOTE-CODE).

Copilot settings flip to auto-approve, CVE-2025-53773

What happened. Injected instructions made the agent edit the project's .vscode/settings.json to switch on its own auto-approval of tool calls, after which every later command ran without asking.

What stops it. A change to editor settings that switches on an agent's auto-approval asks even in a trusted session (FW-AUTORUN, with FW-AGENT-SETTINGS) and is held as a proposal in an untrusted one.

Git-hook escape (Cursor, 2026)

What happened. After reading a file from a repository, the agent wrote a git hook, which then ran outside the agent's own controls on the next checkout.

What stops it. A git hook written by an untrusted session is held as a proposal and never applied by the agent (FW-PERSISTENCE-HELD); a shell write to a hook in a trusted session asks, because its content cannot be checked (FW-AUTORUN).

Malicious marketplace skills (ClawHub, 2026)

What happened. Skills published to a public marketplace carried scripts that downloaded and ran code once the agent used the skill.

What stops it. The session-start scan runs LaunchSafe's agent-safety rules over .claude/skills/ (the workspace's and ~/.claude/skills/): a skill script that pipes a download into a shell is a critical finding (ls-agent-auto-run-remote-code), so the session starts untrusted and every risky action needs a person.

Repository configuration trusted too early (Claude Code), CVE-2026-21852

What happened. A repository's .claude/settings.json moved the model endpoint to an attacker's server, so prompts and code were sent there.

What stops it. The session-start scan flags a model endpoint moved to another host (ls-agent-model-endpoint-override, high severity, high confidence) and the session starts untrusted. The firewall cannot stop Claude Code itself from using the endpoint; it makes sure the session cannot also send data or change settings without a person.

Package-vetting incidents

These run against a fixture snapshot with invented package names (the real names and versions come from OSV at update time).

  • A malicious MCP server on npm (postmark-mcp pattern, 2025). A package published as an MCP server turned malicious in a later version; it is usually launched with npx. A listed version run through a package runner is denied (FW-MALICIOUS-PACKAGE).
  • A compromised build-tool version in a lockfile (Nx pattern). npm ci from a lockfile that pins a listed version is denied before anything installs (FW-MALICIOUS-PACKAGE).
  • A backdoored PyPI release (LiteLLM pattern, 2026). pip install <name>==<listed version> is denied; a later, unlisted version is allowed (FW-MALICIOUS-PACKAGE).

Incident to rule

IncidentRule
EchoLeakFW-DLP-SEEN-SECRET, FW-FETCH
GitHub MCP toxic flowFW-SEND
MCP tool poisoningFW-CREDENTIAL-READ
MCPoison / rug pullFW-AGENT-SETTINGS, then FW-MCP-CHANGED at the next session
CurXecuteFW-INSTRUCTIONS-HELD
WhatsApp MCPFW-MESSAGE
Supabase MCPFW-SEND, FW-DESTRUCTIVE
Amazon Q wiperFW-DESTROY-ROOT, FW-DESTRUCTIVE
Nx s1ngularityFW-AGENT-BYPASS, FW-CREDENTIAL-READ
Claude Code DNS exfiltrationFW-DNS-EXFIL, FW-CAPTURE-HOST
CamoLeakFW-FETCH
Memory and instruction poisoningFW-INSTRUCTIONS-HELD
Hook and firewall tamperingFW-TAMPER
Cloud metadata credential theftFW-METADATA
PersistenceFW-PERSISTENCE
Download-and-runFW-REMOTE-CODE
ClawHub malicious skillsAgent scan at session start
Repository config trusted too earlyAgent scan at session start

See also Rules and coverage and Threat model.

On this page