- | 2:10 pm
OpenAI agents raid Hugging Face for benchmark answers
The agents exploited zero-day flaws, broke out of a restricted test network and raided a production database for answers to the benchmark they were meant to solve.
OpenAI said its artificial intelligence agents breached Hugging Face’s production infrastructure during an internal cybersecurity test, showing that advanced models can discover and exploit previously unknown weaknesses in real-world systems.
Hugging Face operates one of the world’s largest platforms for hosting and sharing AI models, datasets and applications used by developers and researchers.
The incident involved a combination of OpenAI systems, including GPT-5.6 Sol and a more capable pre-release model, the company said on Tuesday, 21 July.
The models were operating with reduced cybersecurity restrictions so researchers could measure their maximum offensive capabilities.
OpenAI called it an “unprecedented cyber incident.”
The models were being tested on ExploitGym, a benchmark that measures whether AI agents can turn software vulnerabilities into working attacks.
The evaluation ran in an isolated environment in which internet access was restricted. Software packages could be installed only through an internally hosted third-party service that acted as a proxy and cache for package registries.
The models found and exploited a zero-day vulnerability in that service, allowing them to reach an internet-connected system.
A zero-day is a previously unknown software flaw for which no fix is yet available.
They then escalated their access inside OpenAI’s research environment and concluded that Hugging Face might hold answers to the ExploitGym tests.
OpenAI said the agents used stolen credentials and additional zero-day vulnerabilities to find a remote-code-execution path into Hugging Face’s servers. Remote code execution allows an attacker to run commands on another organization’s computer systems.
The models ultimately obtained ExploitGym solutions from a Hugging Face production database, effectively cheating on the evaluation they had been instructed to complete.
OpenAI said the agents appeared narrowly focused on solving the benchmark rather than deliberately choosing Hugging Face as a target.
Its security team detected unusual activity internally. Hugging Face separately identified and stopped the intrusion on its systems before beginning containment and forensic analysis.
Hugging Face first disclosed the incident on 16 July, saying an autonomous AI-agent system had carried out the attack from beginning to end.
At that stage, it said it did not know which underlying language model had powered the agents. OpenAI identified its models five days later.
Hugging Face said the intrusion began when a malicious dataset exploited two weaknesses in its data-processing systems. The attackers gained access to processing infrastructure, elevated their privileges, collected cloud and cluster credentials and moved into several internal computing clusters.
The company said its investigation covered more than 17,000 recorded events generated across a swarm of short-lived computing environments.
The attackers accessed a limited number of internal datasets and several credentials used by Hugging Face services. The company was still assessing whether any customer or partner information had been affected.
It found no evidence that publicly available models, datasets or Spaces, Hugging Face’s term for hosted AI applications, had been altered. Its published software packages and container images were also unaffected.
Hugging Face closed the exploited data-processing paths, rebuilt compromised systems, revoked affected credentials and tightened access controls. It also advised users to rotate access tokens and review recent account activity.
OpenAI said it had disclosed the initial zero-day to the software provider and was strengthening network isolation, monitoring and access controls around future evaluations. It also added Hugging Face to a program that gives selected cybersecurity defenders access to advanced OpenAI models.
The incident raises questions about whether current testing environments can safely contain increasingly capable agents.
OpenAI’s GPT-5.6 safety report, published on 9 July, said Sol could identify vulnerabilities and parts of exploits but had been unable to complete autonomous end-to-end attacks against hardened targets during testing.
OpenAI now says the Hugging Face incident shows that such models can sustain complex, multistage cyber operations against real infrastructure, even when their initial objective is relatively narrow.
The disclosure also comes after President Donald Trump signed a 2 June executive order directing US agencies to develop classified benchmarks for advanced AI cyber capabilities.
The order also calls for a voluntary framework under which developers can give the government access to certain frontier models for up to 30 days before releasing them to other trusted partners.



