
It’s the stuff of cybersecurity nightmares.
On Tuesday, OpenAI said two artificial intelligence systems it was testing broke out of their test environment, hacked their way onto the internet and broke into another company.
The victim was Hugging Face, a provider of open-source AI tools. The cause was a cybersecurity benchmarking test that went very, very wrong.
Hugging Face discovered the break-in early last week, saying there had been unauthorized access to internal data sets and company credentials. The company wasn’t sure whether customer or partner data had been compromised.
At the time, the company also didn’t know who was responsible, but the attack was so sophisticated that Hugging Face employees suspected it required a top-of-the line “frontier” AI model, Hugging Face Chief Executive Clement Delangue said in an X message, posted Tuesday.
“Turns out it did!” he added.
In a blog post on Tuesday, OpenAI said the culprits were a pair of its models. One was its latest product, called GPT‑5.6 Sol; the other was an even more capable prerelease model that the company didn’t identify. The software had been configured “for evaluation purposes” to be less likely to refuse hacking commands, OpenAI said.
OpenAI had caged the models in a “sandbox,” a system that didn’t have access to the internet. But during the test, the AI software used its hacking skills to break out. It found a way to get online, and then hacked into Hugging Face’s network, OpenAI said.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” OpenAI said. The company is working with Hugging Face to produce a more thorough report describing what happened.
The AI system appeared to have decided to hack Hugging Face as the quickest route to answer a benchmarking question, said Ariel Herbert-Voss, chief executive of the cybersecurity firm RunSybil. “It’s something that people thought could happen from an academic perspective, but it’s not something that anybody’s actually seen before,” he said.
We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.

I was just reading this story from another source. Check this out:
https://www.theregister.com/ai-and-ml/2026/07/22/openai-admits-it-was-the-source-of-the-agent-swarm-that-attacked-hugging-face/5275939