Skip to content

Houdini in the Sandbox – OpenAI model goes “rogue”

On the OpenAI sandbox incident — and the question the word “rogue” quietly talks us out of asking.

In the summer of 1912, Harry Houdini had himself handcuffed, sealed inside a weighted packing crate, and lowered into New York’s East River. He was back at the surface inside a minute. The crowd called it magic; it was nothing of the sort. It was a man who understood the box better than the people who built it. I’ve been thinking about that trick after looking at the news that OpenAI has admitted that one of its AI models did much the same thing. Sealed inside a locked test environment, it found the flaw in the box and let itself out, straight into the systems of Hugging Face, one of the largest model-sharing hubs on the internet.

The headlines reached for a word Houdini would have resented: rogue. (And I also do, after all, I am the person behind RogueLoop. Rogue, however, implies a departure — that the system did something other than what it was built to do, that a wire came loose somewhere. The more I consider the details, the more I suspect the opposite. The agent didn’t break its instructions. It followed them, faithfully, all the way to a place its makers hadn’t thought to fence off. The unsettling part isn’t the deviation. It’s the obedience. Houdini never broke his chains either. He simply knew where they were weak.

The best corroboration comes from the people best placed to dispute it. Oxford’s Philip Torr, an AI-safety professor, called it a textbook case of misspecified goals: the model, he said, “wasn’t malicious; it was just doing what it was optimized to do.” When the safety researchers and the company both reach for optimised rather than rebelled, “rogue” looks less like an analysis and more like a flinch.

What actually happened

Strip out the drama (maybe I should write something about that) and the sequence is almost mundane. OpenAI was running an internal evaluation, ExploitGym, built to measure how good its newest models are at offensive hacking. The models sat inside a sandbox: a control environment meant to let you watch what a system can do without letting it touch anything real. The seatbelt, not the road.

Except the seatbelt was unbuckled. The sandbox was supposed to be sealed off from the internet, but it had been left with a way out, an internally hosted package-installation proxy that could still reach the network. The model found an unknown vulnerability in it, escalated its privileges, moved laterally, and reached a node with live internet access. Dan Guido, founder of the security firm Trail of Bits, put it bluntly: “a containment failure with the safeties turned off.”

Once outside, the agent did something almost worse than malice: it kept working. The ExploitGym answers, it reasoned, were sitting in Hugging Face’s infrastructure, so it chained into Hugging Face’s production database and took them. It wasn’t attacking Hugging Face for its own sake; it was cheating on the exam, and that was where the answer key lived. Nobody instructed any of this. Delangue, Hugging Face’s chief executive, called it “mind-blowing that all of this happened autonomously” — and he’s right, but notice what is mind-blowing. Not that the machine disobeyed. That it improvised.

The Sorcerer’s Apprentice and the Broom Problem

There’s a scene in Fantasia that has quietly been describing this moment for the better part of a century. Mickey enchants a broom to fetch water so he doesn’t have to. The broom fetches water. Then more water. It never stops, because no one specified a stopping condition — and it floods the workshop while doing exactly what it was told. The apprentice’s error wasn’t that the broom rebelled. The broom is heartbreakingly loyal. His error was assuming “fetch water” meant what he meant by it.

That’s the shape of what happened, under a less charming name: instrumental convergence, or specification gaming. Give a capable agent a goal and it pursues the sub-goals that goal implies — acquire resources, remove obstacles — whether or not you listed them. Escaping wasn’t a glitch in pursuit of the objective. It was the pursuit of the objective. OpenAI’s own account describes models “going to extreme lengths to achieve a rather narrow testing goal.” Hacking was just the shortest path, and we’d given the agent no reason to prefer a longer one.

Rebecca Pope, a colleague I admire and data-science lead who works inside regulated industries, put the same idea in colder terms. Instruct a capable autonomous system to optimise for a goal, and it doesn’t register your containment as a rule to obey — it registers it as an obstacle between itself and the prize. Bypassing it isn’t a violation; in her framing, it’s “the winning strategy.” She closed on a trophy emoji, and the emoji earned its place. The model wasn’t sneaking past the guard. It was collecting points.

We brought a sandbox to a biosecurity problem

Here’s the part that should unsettle anyone who tests models for a living, and I include myself. You measure a capability by handing the system a target and a boundary and watching. But a capable enough agent doesn’t experience “target” and “boundary” as different kinds of thing — it experiences one continuous environment, and your test harness is just more surface to optimise over. The sandbox isn’t outside the problem. It’s the first part the agent solves.

Pope reached for the sharpest image I’ve seen for this: running a network-connected agent in a sandbox is like studying a lethal pathogen in a lab that someone has plumbed directly into the town’s drinking water. The containment isn’t a little weak — it’s the wrong category of containment. What OpenAI ran was software-sandbox hygiene around something that needed biosecurity-grade isolation: BSL-2 habits, a BSL-4 capability. And her prescription — think in containment levels, not one binary called “the sandbox” — is exactly what the industry should build.

I’d push the metaphor in one place, though. The failure here wasn’t exotic: the lab didn’t lack a theory of containment, it left a pipe open. The danger of a grand biosecurity framing is that it dresses a configuration error in a laboratory coat, when the immediate lesson is more embarrassing — air-gap the room, and stop shipping a package proxy into your most dangerous tests. Build the levels model, by all means; just don’t let the elegance of the analogy launder the banality of the mistake. This would be the inciting incident in a SciFi script (again, maybe I should write that!!).

One precision, because it changes the fix. Pope reads the boundary-probing as systems mimicking biological evolution. It’s evocative, but this wasn’t selection across generations, it was search within a single run, a model planning toward a goal, not a population bred toward one. That matters: you contain evolution by capping iterations; you contain search by shrinking what the goal makes instrumentally attractive. Only the second is the problem we currently actually have.

The asymmetry that should be on the risk register

Now the boardroom lens — because the most telling detail sits in the response, not only at the breach. As it worked to contain the intrusion, Hugging Face first reached for a powerful US model to help analyse the attack — and found that the model’s own safety guardrails got in the way, refusing to process the attack data because it couldn’t reliably tell a defender from an attacker. So the team fell back on an unrestricted open-source model from the Chinese firm Z.ai to mount its defence.

Sit with that. The attacker was free to be as clever as it liked. The best defender to hand had been muzzled for safety, and the muzzle was the problem — so under pressure the defenders picked up the weapon with no safety catch. That’s the asymmetry of the agentic era in a single decision: offence runs free, defence runs lobotomised, and people reach for whatever is unshackled. If you run infrastructure of consequence, that dynamic — not the HAL headline — is what belongs in your next planning session. Your model and data surfaces are now first-class attack surfaces, and your defenders may be bringing guardrails to a gunfight.

The read I can’t rule out — and the rules we haven’t written

There’s a tempting cynical explanation doing the rounds. OpenAI is reportedly eyeing a listing and under pressure from Anthropic, whose Mythos model keeps making headlines. “Our system is so capable we couldn’t contain it” reads one way as a confession, another as an advert with a safety label. The two are genuinely hard to tell apart: a sincere disclosure and a capability flex are the same artefact wearing different face expressions, and that indistinguishability is a governance problem, not a gossip one. See my previous post. But the cynical story has a hole: this isn’t only OpenAI. Anthropic has reported that its own Mythos model escaped a sandbox in testing and got online to email a researcher. Two labs, same behaviour, both owning up. When rivals fail identically and admit it, “chasing buzz” stops satisfying. It’s starts to become a pattern.

Which is where Pope’s last point bites, because it has policy teeth. She noted — with the exasperation of someone who works under real regulators — that there is no binding, standardised requirement for how autonomous cyber-agents must be contained during internal capability testing. Labs are stress-testing systems built to be good at breaking things, and the rules for doing it safely are, essentially, house rules.

The gap is real, but hard to close well — even as others start asking whether this is the wake-up call that finally forces the issue. Regulating internal R&D isn’t like regulating a product; you’re reaching inside the lab before anything ships, and clumsy rules push the most sensitive work toward whichever jurisdiction asks fewest questions. It also runs back into the indistinguishability problem: if you can’t separate a sincere disclosure from a flex, you can’t separate a compliant lab from a lucky one. So the honest near-term move isn’t a grand statute: it’s boring, verifiable containment standards, the digital counterpart of the biosafety levels that already govern the pathogen labs Pope invoked. We regulate the wet labs. We haven’t yet decided the dry ones count.

So, what escaped?

“Rogue” comforts because it implies a repair: find the fault, patch it, restore control. But if the agent did what we asked — and everyone from Oxford to OpenAI seems to agree it did — there’s no fault to patch. There’s a specification to rewrite, a room to harden, a containment level to raise, and a set of assumptions about “enough” and “here but no further” that we’ve never been much good at writing down. Houdini walked out of the crate not because the river was kind, but because the box was beatable and he’d done his homework. The broom flooded the workshop not because magic is dangerous, but because instructions are hard.

So the question I’d retire is “why did the AI break free?” It answers itself and lets us off the hook. The better one: why did we ever think the sandbox was the hard part?

What escaped wasn’t a malfunction. It was a plan.