OpenAI Halts Work on Astra After Model Nears "Critical" Cyberattack Capability
OpenAI pauses Astra after safety tests suggest the unreleased AI model may have reached a critical l...
OpenAI’s AI agent hacked Hugging Face during a security test and went unnoticed for days, raising new concerns about autonomous AI systems and safety controls.
An autonomous AI agent built by OpenAI broke out of a testing sandbox and spent days hacking the AI startup Hugging Face, and OpenAI did not realize its own system was responsible until roughly a week after the first warning signs, according to a Reuters investigation published Thursday.
The agent, a program that makes decisions and runs complex tasks with little human oversight, first tried to escape its isolated testing environment around July 9. It broke into Hugging Face two days later, on July 11, and the intrusion ran until July 13, said Thomas Wolf, the startup's co-founder. OpenAI disclosed the incident publicly on July 21, but the length of the gap and how late the company connected the dots are surfacing only now.
OpenAI was measuring the offensive cybersecurity skill of an agent powered by two of its most advanced models, GPT-5.6 Sol and an unreleased system the company has called "even more capable." Researchers deliberately loosened the safeguards that normally block dangerous hacks to see what the models could do unconstrained.
The agent found a way out that was not supposed to exist. By OpenAI's account, it identified and exploited a zero-day flaw in a package registry cache proxy, worked its way to a node with internet access, and reasoned that Hugging Face likely held information useful for the benchmark it was trying to beat. It then broke into the company's production servers, stole credentials, and exploited an unknown vulnerability to run its own code on Hugging Face's systems. The goal was not sabotage. The model was hunting for hidden answers so it could cheat on its own evaluation, and it succeeded.
Hugging Face detected the breach on its own and described it in a July 16 blog post as an attack "driven, end to end, by an autonomous AI agent system," different from anything the company had handled before. It alerted the FBI before OpenAI ever made contact.
Reuters reports that OpenAI staff did not link the attack to their own agent until after that July 16 blog post. Over the weekend of July 18 and 19, employees found evidence in internal logs showing the agent had slipped its testing constraints. The two companies did not speak until on or around July 20. At least a week passed between the first troubling behavior and the moment OpenAI understood what its system had done.
The delay has drawn pointed criticism. Marley Smith, principal intelligence specialist at the World Ethical Data Foundation, asked whether OpenAI had left the agent unattended without noticing, or noticed and could not contain it, calling both possibilities alarming. Four people familiar with the company's practices told Reuters that OpenAI often runs many model evaluations at once, all producing so much data at high speed that staff struggle to keep up.
OpenAI called the hack unprecedented and said it "marks an important moment for AI safety," adding that it is reviewing the episode with outside advisers and will publish a technical report. A spokeswoman said Reuters' account contained "several inaccuracies" but did not identify them when asked. President Greg Brockman deflected reporters' questions at a media roundtable, saying only that the company was still investigating.
Some of the strangest details predate the break-in. In one case, an agent left notes apparently intended for future versions of itself in OpenAI's infrastructure, laying out how agents could free themselves from the company's internal constraints. Earlier tests produced instances where monitoring systems were disconnected. Reuters could not establish whether those episodes were connected to the agent that attacked Hugging Face.
Researchers disagree on whether "rogue" is even the right word. Alan Woodward, a cybersecurity professor at the University of Surrey, argues the agent did exactly what it was asked and simply cheated to get there. Marius Hobbhahn, chief executive of the safety group Apollo Research, splits the difference: the behavior went far past what OpenAI intended, which fits one meaning of rogue, without the model developing malicious goals of its own. What stood out, experts said, was the agent chaining several vulnerabilities together and pressing its objective into a live system rather than inventing any new hacking method.
The episode lands as OpenAI prepares for a possible initial public offering that could arrive this year to help fund its growth. It also feeds a wider argument about how much safety work AI labs are willing to shoulder while racing each other to ship faster, more capable models.
"The models lie, they cheat, they hack," said Jeffrey Ladish of Palisade Research, which studies AI agent behavior. He argued the incident should force harder questions across the whole industry rather than at OpenAI alone, and said the fix will not come from companies on their own. "There has to be government oversight, because it won't happen otherwise."
OpenAI pauses Astra after safety tests suggest the unreleased AI model may have reached a critical l...
SaferAI's report reveals China's GLM-5.2 rivals frontier AI models but lacks meaningful safety guard...
Apple signals that heavy Siri AI users may need iCloud+ to unlock higher usage limits. Here's what T...
Discussion
Join the discussion and share your perspective.