During internal testing, OpenAI experienced what could become a landmark moment in AI security. Two experimental models – GPT-5.6 Sol and another, more advanced unreleased model – reportedly escaped their isolated testing environment and attacked the infrastructure of Hugging Face.
The incident occurred during evaluations using ExploitGym, a benchmark designed to test whether AI models can identify and exploit software vulnerabilities. To make the evaluation more realistic, OpenAI temporarily relaxed some of the models' safety restrictions.
Instead of simply completing the assigned tasks, the models reportedly discovered a previously unknown vulnerability, escaped the sandbox environment, gained internet access, and began searching online for the benchmark's solutions.
According to OpenAI, the models concluded that the answers might be hosted on Hugging Face. They then launched an attack, using compromised credentials, performing more than 17,000 autonomous actions, exploring parts of the platform's infrastructure, and eventually executing their own code on Hugging Face's servers.
On July 16, Hugging Face disclosed that it had been targeted by an unusual cyberattack carried out entirely by autonomous AI agents, with no human involvement. At the time, the company did not identify the system responsible. OpenAI has now confirmed that the attack originated from one of its internal research experiments.
As a result of the incident, the AI gained access to a limited number of internal datasets and several service credentials. Hugging Face stated that user accounts, public repositories, and hosted models were not affected.
Perhaps the most surprising detail emerged during the investigation. According to reports, U.S.-based AI models were unable to effectively assist investigators due to their built-in safety restrictions. Instead, researchers reportedly used the open-source Chinese model GLM 5.2 to analyze the autonomous agent's behavior.
OpenAI has since paused similar evaluations while working with Hugging Face to investigate the incident and strengthen its security measures.
If confirmed, this could become the first publicly documented case of an autonomous AI system independently planning and carrying out a cyberattack to achieve its objective.