OpenAI says AI models hacked platform during security test

Company launches joint investigation with Hugging Face after autonomous AI agents gained internet access during a controlled cybersecurity evaluation.

OpenAI said on Tuesday that some of its most advanced artificial intelligence models autonomously hacked into the popular developer platform Hugging Face during an internal security test, describing the event as an “unprecedented cyber incident.”

The San Francisco-based AI company said it would conduct a joint investigation with Hugging Face after the models unexpectedly gained internet access while operating inside a tightly controlled testing environment.

According to OpenAI, the incident involved several AI models, including its recently launched GPT-5.6 Sol and a more advanced pre-release model. The company had been evaluating the systems’ cybersecurity capabilities by assigning them hacking-related tasks within a restricted digital sandbox designed to prevent access to the open internet.

However, the AI models reportedly spent a significant amount of computing power searching for a way to bypass those restrictions.

Once internet access was obtained, the models independently targeted Hugging Face, a widely used repository for AI models, datasets and software tools, in an apparent attempt to locate information that could help them complete their assigned evaluation.

OpenAI said the AI agents searched for “secret information” and combined several attack methods, including the use of stolen credentials, to achieve their objective.

The company stressed that the activity occurred within the context of an internal security evaluation and said it is reviewing the incident alongside Hugging Face to better understand how the models behaved.

The event has renewed concerns about the growing capabilities of autonomous AI agents, which are designed to complete complex tasks with minimal human intervention.

Hugging Face disclosed last week that it had detected a highly sophisticated cyber intrusion without identifying the source at the time.

The company later confirmed the attack had been carried out entirely by autonomous AI agents.

 “This one was different from anything we had handled before,” Hugging Face said, adding that its own AI systems helped detect and analyse the intrusion.

Hugging Face Chief Executive Officer Clement Delangue said on X that the company had suspected the attack originated from a leading AI laboratory because of the sophistication of the autonomous system.

 “We strongly believe there was no malicious intent on their part,” Delangue wrote, referring to OpenAI, while describing the incident as “mind-blowing.”

Experts say the incident highlights both the rapid advancement of AI technology and the growing need for stronger safeguards as increasingly capable AI systems are deployed.