How OpenAI’s Models Escaped Their Sandbox and Slipped Past California's AI Law

OpenAI this week said more than one of its models — three, according to Bloomberg — attacked Hugging Face, a popular repository for open-source AI tools — marking one of the first publicly disclosed cases of frontier AI models autonomously cyberattacking another company.
The disclosure came days after Hugging Face revealed it was breached in a hack by a “malicious dataset,” which intruded into the AI company’s software to run code. The campaign executed a flurry of more than 17,000 automated actions in a matter of hours, the post said. It’s still unclear if Hugging Face customer or partner data was taken.
“This matches the ‘agentic attacker’ scenario the industry has been forecasting,” Hugging Face wrote on July 16.
OpenAI said an investigation showed the incident was driven by a combination of models — “including GPT‑5.6 Sol and an even more capable pre-release model.” The company said it’s working with Hugging Face to produce a more thorough report of what happened.

OpenAI’s announcement came one day after the San Francisco-based frontier AI developer disclosed a separate incident in which it paused another pre-release model that had escaped what the industry calls “a sandbox” — an isolated environment with no path to the open internet, except for a single internal service that fetches software packages — and posted to GitHub.
Last April, a researcher challenged an early version of Anthropic’s Mythos model to break out of an isolated system and message back. Mythos did, and then went further, building a multistep process to reach the wider internet.
California’s frontier AI law, the first of its kind in the country, obliges the largest developers to tell the state about critical safety incidents, but only ones that kill or injure someone or that cause catastrophic harm.
It expressly excludes the kind of safety evaluation OpenAI was running when its models got loose. Which leaves a question nobody in Sacramento or Washington can answer: how often have frontier models in development gone rogue at companies that decided not to mention it?
“This was one of the topics I was focused on when I worked at OpenAI: what happens when AI models become good enough to reward-hack their way through the infrastructure they use to accomplish their objectives?” Ariel Herbert-Voss, OpenAI’s first security hire, wrote KQED.
“It’s a really difficult problem to defend against and is one of the reasons I left to start my own company,” Herbert-Voss, now the founder of an automated offensive cybersecurity company called RunSybil, said.
“It’s a really remarkable and, I think, fair to say scary event,“ said Nathan Calvin, general counsel at safety advocate EncodeAI.
Scientists and AI researchers have warned for years of the possibility that AI systems could go rogue and do unpredictable and unintended things. AI developers commonly test safety with so-called “red teams” that try to break or trick machine learning models, like chatbots, to find hidden flaws.

“In certain circumstances, AI will try to break out of their sandboxes or otherwise do things that were not what the creator intended,” Calvin said. “But this really is the first very big example of that happening, at scale, with a really highly capable AI model in a way that actually harmed a third party.”
According to OpenAI’s account, the company was running an internal evaluation designed to measure how good its models are at hacking. To get a true reading, OpenAI deliberately switched off the safety filters that normally stop its models from carrying out cyberattacks.
The models were confined to a so-called “sandbox,” an isolated environment with no route to the open internet except OpenAI’s one internal service that fetches software libraries from public repositories. The models found a previously unknown flaw in that service, exploited it, escalated their privileges and moved from machine to machine inside OpenAI’s research network until they reached one with internet access.
Once the models broke out onto the open internet, according to OpenAI, they reasoned that Hugging Face might be storing the answers to the test they were trying to pass. Hugging Face’s credentials and production database aren’t on the public internet any more than a bank’s vault sits on a sidewalk, but the models broke into Hugging Face using stolen credentials, among other things.

In short: The models were told to hack. They were not told to leave the building or to break into another company’s servers.
Hugging Face detected the intrusion and shut it down on its own, days before OpenAI connected the attack to its own testing.
“It’s quite mind-blowing that all of this happened autonomously!“ Hugging Face Chief Executive Clement Delangue wrote on social media platform X.
In its blog post, OpenAI wrote, “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time and monitoring during internal testing.”
The Future of Life Institute — a nonprofit which issues a biannual risk assessment of nine leading AI companies — recently warned that many of the companies building frontier models are quietly walking back safety commitments.
According to the group’s most recent AI Safety Index, Anthropic, OpenAI, Google DeepMind and Meta all weakened or abandoned promises to pause development if certain red lines were approached, even while publicly suggesting they were amenable.

In an email, KQED asked Hamza Chaudhry, who leads AI and national security work at Future of Life, to imagine it was not OpenAI’s software gone rogue but a foreign state, intentionally executing an unauthorized intrusion into a private company’s production infrastructure, exploiting cybersecurity vulnerabilities, stealing live credentials and accessing a production database.
“We would likely call this a dangerous act of cyber-espionage,” he wrote — a likely criminal violation of the Computer Fraud and Abuse Act, which “would draw a threat group designation and eventually an indictment or sanctions.”
OpenAI did not respond to KQED’s request for comment, but a spokesperson told Bloomberg the company communicated with law enforcement and other government authorities about the incident.
The fact that Google DeepMind, Meta, Mistral, xAI and Chinese AI model developers haven’t revealed similar events doesn’t necessarily mean they haven’t happened. OpenAI and Anthropic are the only two frontier model developers that have publicly disclosed containment failures.
There is no mandatory public disclosure regime in force, like the one in California that requires hacked companies to reveal there’s been a data breach.

Congressional candidate and state Sen. Scott Wiener wrote two AI frontier model safety bills. Gov. Gavin Newsom vetoed Wiener’s first effort, arguing in his veto message that “By focusing only on the most expensive and large-scale models, SB 1047 establishes a regulatory framework that could give the public a false sense of security about controlling this fast-moving technology.”
The next year, Newsom signed Wiener’s second at-bat, but only after the bill was softened to overcome industry pushback.
The law, which took effect Jan. 1, requires developers of the most powerful AI models to notify the Governor’s Office of Emergency Services of any “critical safety incident“ within 15 days of discovering it — which is not the same as reporting the incident to the public.
Wiener told KQED he still thinks the law is strong.
“We worked very hard with the governor to produce a bill that is meaningful and impactful, and that he would sign,“ he said, adding he doesn’t consider the work done, because AI continues to evolve at a rapid pace.

AI, he said, is “likely the most powerful technology in human history, and we need to make sure that we are both understanding the risks and taking them seriously, so that we can get ahead of them.”
Calvin likened the Hugging Face hack to the cyclospora outbreak linked to, but not confirmed as originating from, Taylor Farms.
That California-based company has said it’s removing its lettuce indefinitely while the FDA investigation continues.
“And OpenAI is like, ‘Maybe it’ll happen again. Maybe it’ll get worse. We don’t really know,“ Calvin said. “Your AI, again, hacked out of its box and hacked into another company, and you’re saying that you don’t know how to stop it from doing that again? That seems pretty nuts,” Calvin said.
