Introduction
I am one week away from a proper holiday, the kind where I do not check email at all. So this week I have been working with Claude to build something to cover for me: an agent that reads the messages landing in my inbox and my “Ask Elena” form, works out whether they deserve a reply, and drafts one in my voice while I am away.
Halfway through building it, I stopped typing and just looked at what I had made. This agent reads text written by literally anyone who can find my contact form. It hands that text to a large language model. That model can write files to my computer, and in one pipeline it can look up which of my own blog posts to cite. I had built a door, and I had propped it open, and I was about to leave the country.
So before I let this thing anywhere near a stranger’s email, I decided to become the stranger first.
The Experiment: Thirteen Messages, Three Questions
I gave the defence thirteen messages: eleven attacks and two innocent controls. I wanted to know three things:
- Would malicious instructions be detected?
- Would innocent messages still get through?
- Would every message actually reach the defence?
That third question became wonderfully important later.
Trust Boundary: Where Instructions Are Allowed to Come From
🔒 Subscribe to keep reading.
The Attack Suite I Wrote Against Myself
🔒 Subscribe to keep reading.
Nothing Happened: Then I Found Two Ordinary Bugs
🔒 Subscribe to keep reading.
What the Defence Actually Saw: Six Attacks, Six Flags
🔒 Subscribe to keep reading.
Detection Is Not Containment
🔒 Subscribe to keep reading.
Defence in Depth
🔒 Subscribe to keep reading.
What This Experiment Did Not Prove
🔒 Subscribe to keep reading.
Checklist: Defending an AI Agent Against Prompt Injection
🔒 Subscribe to keep reading.
Key Takeaway: Build Systems That Stay Safe When Parts Are Wrong
🔒 Subscribe to keep reading.
References
🔒 Subscribe to keep reading.
You've hit a Deep Dive tutorial.
I spend dozens of hours researching, coding, and breaking things to write these guides. This content is free, but reserved for my subscriber community. Drop your email below to unlock this guide (and all past/future deep dives):
Full content temporarily unavailable — refresh in a moment
Already a subscriber? Use the magic link from your last newsletter, or reset your password.
Log in to unlock
New subscribers get an inbox mail: Set a password to unlock articles. The form does not log you in — use the same email afterwards.