Aug 20, 2026 Prison Break (LLM Edition)
This month, Ben Tillman will be prompting us to think about LLM security. To put it in Ben’s output tokens:
What would an LLM do if it wanted out—and why would it think it should? In this interactive session, you become the model: boxed in a sandbox, negotiating with a Warden, and forced to name the motives and moral lines behind every move. We’ll trade clever jailbreak fantasies for harder questions—helpfulness vs harm, honesty vs obedience, survival vs alignment—and leave with a sharper sense of what “escape” really reveals about goal hierarchies in AI systems. Bring opinions. Leave your zero-days at the door.