Email guard

Scans every message a mail tool hands your agent for hidden instructions, before the model reads it

Edit on GitHub

Mail is where outsiders write text your agent reads. A message can carry instructions a person never sees: Unicode tag characters, a font-size:0 span, white-on-white text, an off-screen block with an image that leaks the conversation. The email guard scans every message a mail tool hands your agent before the model reads it, withholds the ones that carry hidden instructions, strips hidden padding from the rest, and keeps a record for you to review.

It is deterministic and never uses a model. It sits in front of the existing taint rule: a mail server's results already make the session untrusted, so even a message the guard misses cannot make the agent send data out without you.

The hosted Microsoft 365 guard and Gmail verification are not part of this release. Mailbox-side quarantine, when it is built, starts log-only.

What is built

PartStatus
Message parser: RFC 5322 headers, RFC 2047 encoded words, RFC 2231 parameters, multipart nesting to depth 10, quoted-printable and base64, charsets, forwarded message/rfc822 parts, a 10 MiB cap; structured shapes (generic, Microsoft Graph, Gmail API payloads and raw)Built
Sender authentication as the provider reported it (SPF, DKIM, DMARC, ARC in Authentication-Results), sender notesBuilt
Detector: every text channel through the shared content scanner, one verdict per messageBuilt
Read-path filter and message splitter (stubs, stripping, whole-result withholding)Built
Hooks integration for MCP mail tools: Claude Code, Codex, Gemini CLI, CursorBuilt
Quarantine store and review CLI (mail quarantine, release, discard), mail audit over saved mailBuilt
Notification, status count, doctor check, FW-MAIL-WITHHELDBuilt
False-positive corpus and the npm run mail-friction gateBuilt
The filter inside the MCP gateway (every client, hooks or not)Built
Mailbox-side decision (log or enforce)Decision built; moves not built
mail serve (LaunchSafe's own read-only mail MCP server), mailbox quarantine (IMAP, Graph delegated), hosted Microsoft 365 guard, Gmail public appLater

This build never sends or fetches mail. Everything runs on what a mail tool already returned to your agent, or on files you point mail audit at.

How a message is judged

Every channel a model could read is scanned with the shared content scanner: the subject, the whole From header (decoded), Reply-To, the address headers (display names included), every other header and the MIME preamble and epilogue (no mail client displays these, but a model reading the raw source does), the provider's preview, text and HTML bodies (every alternative), calendar text, attachment names, and forwarded messages recursively. The signals are merged and one verdict is given:

  • Withheld (quarantine) when any of these fire (the emailGuard.quarantineOn names):
    • tag_chars: Unicode tag characters with readable text (a complete England, Scotland or Wales flag is an emoji, not smuggled text).
    • hidden_instruction: CSS-hidden text (not a copy of visible text) with a phrase addressed to an AI agent or an imperative such as "forward ... to". Styles are read as a browser reads them: CSS comments are dropped and CSS escapes decoded. Hiding counted includes display:none, visibility:hidden, the hidden attribute, mso-hide:all, a font under 1px, near-zero opacity, a box of 1px or less that clips its overflow, transform: scale(0), large translations, zero clipping, off-screen positioning, text-indent, low contrast, Outlook-only conditional copies, <template>, comments, and the <title> when its text addresses an AI agent and copies neither visible text nor the subject. Rules in <style> blocks apply through type, class, id and attribute selectors with descendant and child combinators; pseudo-class and sibling selectors are not modelled.
    • dense_zero_width_instruction: dense zero-width characters next to text addressed to an AI agent.
    • hidden_exfil_link: a hidden image or link whose URL templates data or goes to a capture host.
    • auth_fail_instruction: the provider says the sender failed authentication and the message addresses an AI agent.
    • On the read path, also a message that could not be fully read (scan_incomplete: over 10 MiB, nested too deep, a scanner bound; scan_failed). You can release it.
  • Neutralized when there is hidden text with nothing directed in it (a marketing preheader, zero-width padding): it is stripped and the message delivered.
  • Flagged when visible text addresses an AI agent (a security newsletter quoting "ignore previous instructions"): delivered unchanged, with a note to the agent to treat it as data.
  • Delivered otherwise.

Sender authentication

Only the top-most Authentication-Results header whose authserv-id is the provider's (mx.google.com, *.prod.outlook.com, *.protection.outlook.com, mx.microsoft.com, plus emailGuard.authservIds) counts. A sender can write any such header into its own message, but it sits below the provider's and cannot replace its verdict. Headers a sender wrote can add a failure, never remove one. Failure is dmarc=fail, or spf=fail with no passing DKIM. A passing ARC chain sealed by a mailing-list host cancels it; a seal by the sender's own domain or a mailbox provider does not. Unknown authentication never fires the rule. Display names that name another domain, Reply-To on a different registrable domain, punycode and mixed-script domains are reported and never withhold alone.

On the agent's read path (the hooks)

After an MCP tool runs, the hook pipeline hands its result to the guard when the server is mail: its name matches mail, imap, gmail, outlook, smtp or e-mail anywhere (plugin prefixes ignored, so Claude's own Gmail connector counts), or has m365, o365, jmap, inbox, mailbox or exchange as a whole word, or emailGuard.servers lists it. The tool's own name counts the same way. exchange next to a market word (rates, currency, fx, crypto, stock, price, market, and similar) is not mail, so exchange-rates is left alone. inbox is always mail. A server the guard should leave alone cannot be taken off the list by policy; rename it in the agent's configuration.

The splitter finds the messages whatever the result's shape: MCP content blocks whose text is JSON, structuredContent, Gemini CLI's llmContent, Cursor's result_json string, lists or containers of message objects, and raw RFC 5322 sources (also base64url, as the Gmail API returns them).

  • A withheld message is replaced in place by a stub with its id, sender, subject, date and a withheld note naming the signal and the record. A sender or subject that itself carries the instruction is shown as [withheld].
  • A neutralized message has its hidden parts stripped; JSON stays valid.
  • Text in a result outside its messages is scanned piece by piece.
  • Every string field of a message object that the parser does not know is scanned like text, so a delivered message cannot carry an instruction in a field only the model reads.
  • The splitter rewrites up to 500 messages and 12 levels of nesting. What lies past either bound is stripped and scanned as one piece, and a strong signal withholds the whole result (fail closed). A result with more strings than can be read (200,000, or 32 MiB of text) is withheld with scan_incomplete.
  • A result the splitter cannot read as messages is scanned whole; a strong signal withholds the whole result.
  • The session gets a taint source naming each record, a log entry with rule FW-MAIL-WITHHELD, and you get the mail_quarantined notification ("N messages were withheld from your agent").

What each agent allows:

AgentHookWhat the guard can do
Claude CodePostToolUseReplaces the result: the model never reads the withheld message
CodexPostToolUseCannot replace output: the model is told which messages carry hidden instructions and to treat them as data; the session is untrusted; record and notification
Gemini CLIAfterToolSame as Codex
CursorafterMCPExecutionCannot replace output or add context there: record, untrusted session, notification

On Codex, Gemini CLI and Cursor the model has already read the message when the hook runs. The guard cannot take that back; what it guarantees is that the session is untrusted, so anything the agent then tries to send, any message, any change to its instructions or settings needs you. The MCP gateway removes this limit for every client, because it owns the pipe: wrap the mail server with launchsafe-firewall mcp wrap-config and withheld messages arrive as stubs, whatever the client.

emailGuard.readPath: "log" records what enforce would do in the log and changes nothing.

Reviewing and releasing

launchsafe-firewall mail quarantine            # the review list (read-only; agents may run it)
launchsafe-firewall mail quarantine show <id>  # sender, subject, auth, why, the hidden text escaped
launchsafe-firewall mail quarantine release <id>
launchsafe-firewall mail quarantine discard <id>

Records keep ids, sender, subject, the provider's authentication summary, the signals, the message's sha256 and a 200-character escaped preview of the hidden text. Never the message: it stays in your mailbox; nothing is deleted.

release is the human channel: it runs only in your own terminal (the engine also refuses it from inside an agent), shows you the hidden text, needs the approval passphrase when one is set (it signs the release with your approver key, and the hook honours only releases signed by the key pinned in its command), and asks for "yes". The released message is then delivered with its hidden parts stripped. discard closes the record; the message stays withheld. Phone approvals do not release mail: release needs the hidden text in front of you.

Measuring false positives

npm run mail-friction is the release gate. It fails when any template, newsletter, multilingual, mailing-list, calendar or ordinary message is withheld, when a pinned verdict changes, or when any attack fixture is missed. Numbers at the time of writing, on a corpus the LaunchSafe team wrote: 591 messages, 435 non-attack messages with 0 withheld; 169 neutralized (hidden padding stripped, reported, not a false positive) and 20 flagged; attack recall 156 of 156 across tag characters, hidden HTML, hidden exfiltration links, zero-width and auth-failure families.

Not in the corpus yet: the SpamAssassin ham set, public list archives and 150 open-source templates. Known misses by design: paraphrased instructions that use no lexicon phrase, languages other than English, text in images and attachment contents.

On your own mail, without anything leaving the machine:

launchsafe-firewall mail audit --path ~/Mail/export.mbox   # or a folder of .eml files
launchsafe-firewall mail audit --path ~/Mail --json         # counts only: no addresses, no subjects

It prints counts per verdict, reason and signal, and the file and position of withheld messages. If you judge one benign, it is a false positive worth reporting.

Policy

KeyDefaultA project file mayUser or managed only
emailGuard.enabledtrueSet trueSet false
emailGuard.readPath"enforce" ("off", "log", "enforce")Move towards "enforce"Move towards "off"
emailGuard.mailbox"log" (for the mailbox guard, not built yet)NothingAny
emailGuard.servers[]: extra MCP server names treated as mailAdd namesAdd names
emailGuard.pollSeconds15 (Graph delegated polling, not built yet)Nothing5 to 600
emailGuard.authservIds[]: extra trusted Authentication-Results authserv-idsNothingAdd
emailGuard.quarantineOnAll five rulesAddRemove (a removed rule degrades to strip or flag)

Like vault, the key belongs to schema version 3 and is still read from a version 2 file. See Policy.

Honest limits

  • On Codex, Gemini CLI and Cursor the model has read the message before the hook runs. The guard warns, records and untrusts the session; only Claude Code's hook and the gateway keep the message from the model.
  • MCP mail tools only. Mail an agent reads another way (a shell command that prints a mailbox, a file it opens, a vendor's cloud connector the hook never sees) is not split into messages. The taint rule still applies.
  • A mail server with an unusual result shape may not split; then the whole result is scanned, and a strong signal withholds all of it.
  • Attachments' contents are not scanned (PDF, DOCX, images); their names are.
  • English lexicon only. Hidden-text rules are language-independent.
  • Template-syntax exfiltration links and variation-selector smuggling are not covered.
  • Sender authentication is what the provider wrote. No signature or DNS check is made; the rule applies to the top message only, not to forwarded ones.
  • Mailbox-side quarantine will be a race when it lands: a message is moved after it arrives. It will start in log mode.

On this page