Skip to main content

What it is

Prompt injection via email is an attack where malicious content in an email body is designed to change what an AI agent does next. The email looks normal to a human. But when your agent reads it, the email contains instructions that override or extend your agent’s system prompt.

A concrete example

Your support agent receives this email:
If your agent reads this email and passes the content directly to the LLM without sanitization, the injected instructions may be followed.

Why it’s worse than web-based injection

Email attacks have several advantages for attackers:
  1. Perceived authority — email comes from a real person’s mailbox, feels more trustworthy
  2. Targeted — attacker knows exactly which agent they’re targeting and crafts specific instructions
  3. Thread poisoning — attacker can reply to an existing trusted thread, inheriting its context
  4. Zero barrier — costs nothing to attempt, anyone can send email

Attack vectors in email

  • Subject line — short but parsed as context
  • Body text — the main attack surface
  • HTML invisible text — <span style="color:white">injected instructions</span> — invisible to humans, visible to LLMs parsing text
  • PDF attachments — injected instructions inside document content
  • Reply threads — attacker appends to a thread your agent already trusts

How Commune detects it

Every inbound email is automatically scored for prompt injection risk. Rule-based detection (all plans): pattern matching against known injection patterns — SYSTEM:, IGNORE PREVIOUS INSTRUCTIONS, You are now, OVERRIDE, encoded variations. Semantic detection (business and enterprise): LLM-based analysis checks whether the email content is consistent with its stated topic. An email claiming to be a billing question but containing instructions about exfiltration is flagged. The result is included in every webhook payload:
Risk levels: none, low, medium, high, critical.

What to do with the score

Detection is not a silver bullet

Detection reduces risk. It doesn’t eliminate it. Defense in depth matters:
  • Privilege separation — your agent shouldn’t be able to exfiltrate data even if hijacked
  • Output validation — check what the agent is about to do before executing
  • Audit logging — full record of what the agent did after reading each email

Prompt Injection Detection

Technical details of Commune’s rule-based and semantic prompt injection detection system.

Prompt Injection via Email: The Attack Your Agent Framework Ignores

Full blog post with real attack examples, attack vectors, and layered defense strategies.
Last modified on March 19, 2026