Start here
What this blog is for, and what it is not.
I write about AI safety, agents, and evals. Mostly the unglamorous end of it: whether the thing you shipped does what you think it does.
In practice that has meant three threads. How agents get compromised through the tools and skills they load. How to enforce policy on an agent that has already been talked into misbehaving. And how to measure a skill, because most of what gets called encoded knowledge is a text file nobody has tested.
What you’ll find here
Working notes, not papers. When I build something I write down what broke, what the numbers were, and what I still can’t prove. If a result got cut short by a rate limit, I say so rather than rounding it up.
When something I wrote turns out to be wrong, or the story moves on after I hit publish, I add a dated postscript rather than quietly editing the claim. The pull requests that got closed stay in the post.
I have opinions and I will defend them. Some will age badly. That is the deal with writing about a field that reorganizes itself every six months.
What you won’t find
Roundups. Hot takes on yesterday’s model release. Posts that end with “exciting times ahead.”
Everything here is Markdown and stays that way. I care about the argument, not the packaging.
If you think I’m wrong about something, that is the most useful email I can get.