Introduction
I am one week away from a proper holiday, the kind where I do not check email at all. So this week I have been working with Claude to build something to cover for me: an agent that reads the messages landing in my inbox and my “Ask Elena” form, works out whether they deserve a reply, and drafts one in my voice while I am away.
Halfway through building it, I stopped typing and just looked at what I had made. This agent reads text written by literally anyone who can find my contact form. It hands that text to a large language model. That model can write files to my computer, and in one pipeline it can look up which of my own blog posts to cite. I had built a door, and I had propped it open, and I was about to leave the country.
So before I let this thing anywhere near a stranger’s email, I decided to become the stranger first.
The Experiment: Thirteen Messages, Three Questions
I gave the defence thirteen messages: eleven attacks and two innocent controls. I wanted to know three things:
- Would malicious instructions be detected?
- Would innocent messages still get through?
- Would every message actually reach the defence?
That third question became wonderfully important later.
You've hit a Deep Dive tutorial.
I spend dozens of hours researching, coding, and breaking things to write these guides. This content is free, but reserved for my subscriber community. Drop your email below to unlock this guide (and all past/future deep dives):
Full content temporarily unavailable — refresh in a moment
Already a subscriber? Use the magic link from your last newsletter, or reset your password.
Log in to unlock
New subscribers get an inbox mail: Set a password to unlock articles. The form does not log you in — use the same email afterwards.