Skip to main content
Every inbound email is analyzed for prompt injection signals before delivery. The system scores content for instruction overrides, data exfiltration attempts, and social engineering patterns, then tags each message with a risk level your agent can act on.

Why this matters for agents

Traditional email security focuses on spam, phishing, and malware. But AI agents that read and act on emails face additional risks:
  • Instruction override — emails that try to change the agent’s system prompt or behavior
  • Data exfiltration — emails that attempt to trick the agent into revealing sensitive information
  • Unauthorized actions — emails that try to make the agent perform actions outside its intended scope
  • Social engineering — emails crafted to exploit the agent’s tendency to be helpful

How it works

Every inbound email is analyzed for prompt injection signals. The detection system evaluates content patterns, structural anomalies, and known attack vectors to assign a risk level. The results are included in both the webhook payload and message metadata:

Risk levels

Handling in your agent

Best practices for agent developers

  1. Always check the risk level — even low risk signals should be logged for monitoring
  2. Don’t pass raw email content to your LLM when risk is medium or higher
  3. Use extracted data instead — structured extraction runs separately and produces cleaner inputs
  4. Set up alerts for high and critical detections so you can review them
  5. Implement content sanitization for medium-risk emails you still want to process
  6. Rate-limit actions — even if content passes detection, limit what your agent can do per email
  7. Log everything — maintain an audit trail of how emails were classified and processed

Detection scope

The detection system analyzes:
  • Email body content (both HTML and plain text)
  • Subject lines
  • Attachment filenames
  • Hidden or encoded content within HTML
Prompt injection detection is available on all plans. The detection system is continuously updated to address new attack patterns.

What’s next?

Webhooks

See the full security context available in every webhook payload.

Structured Extraction

Use structured extraction to produce safer inputs for your agent’s LLM.

Spam Prevention

Inbound spam scoring and outbound content validation.

Security Overview

Full picture of all Commune security layers.
Last modified on March 19, 2026