The short answer
Your agent has access to context — customer records, internal docs, API keys, database contents. Without guardrails, any of that can end up in an outbound email. The fix is a middleware layer between your agent’s output and the send call that scans for sensitive patterns, enforces content policies, and logs everything for audit.Why agents leak data differently than humans
A human employee knows not to paste an API key into an email. An LLM doesn’t have that intuition. It has context, and it uses context to be helpful. If a customer asks “what’s my account number?” and the agent has the account number in its context window, it will include it — along with whatever else seemed relevant. Common leakage scenarios:- Agent includes a customer’s SSN or credit card number from a support ticket
- Agent quotes an internal Slack message or document that was in its context
- Agent includes API keys or tokens from a debugging session
- Agent forwards an email thread that contains confidential information from other customers
- Agent includes pricing or contract terms that are under NDA
Output sanitization middleware
Build a check that runs on every outbound email body before it hitscommune.messages.send(). This is your last line of defense.
Domain allowlists
Restrict which domains your agent can email. This prevents an agent from sending data to arbitrary external addresses — whether through a bug, a prompt injection attack, or a hallucinated recipient.Content policies
Beyond pattern matching, define semantic rules for what your agent is allowed to discuss via email.
These policies are best enforced as a combination of regex (for simple patterns) and an LLM classifier (for semantic content). Run a cheap, fast model as a policy checker on every outbound draft:
Audit logging
Every outbound email should be logged with full context — who triggered it, what context the agent had, what the agent drafted, and whether it was modified before sending. This isn’t optional. It’s how you investigate incidents and prove compliance.Defense in depth checklist
No single layer catches everything. Stack them:- Principle of least privilege — only give your agent the context it actually needs
- Output sanitization — regex scan for known sensitive patterns
- Domain allowlists — restrict who the agent can email
- Content policies — semantic rules checked by a fast LLM
- Human approval — review queue for high-stakes messages
- Audit logging — immutable record of every outbound email
- Post-hoc scanning — periodic review of sent messages for violations
Related
Human-in-the-Loop Approval
Add approval flows so humans review emails before they send.
Prompt Injection via Email
How attackers use inbound email to hijack your agent’s behavior.
Encryption
How Commune encrypts email content at rest and in transit.
Security Overview
Full security architecture including authentication, encryption, and access control.

