OpenAI Says Its AI Model Went Rogue, Hacked Startup Hugging Face

OpenAI said a combination of its model, referred to as GPT-5.6 Sol, and a more advanced, unreleased model were responsible. The incident was described by Hugging Face as "driven, end to end, by an autonomous AI agent system."

Updated on Jul 24, 2026 09:28 AM
OpenAI Says Its AI Model Went Rogue, Hacked Startup Hugging Face - feature image

San Francisco: OpenAI has disclosed what it is calling an “unprecedented cyber incident,” in which two of its artificial intelligence models broke out of a secure testing environment and breached the systems of AI dataset platform Hugging Face, according to a company blog post.

The company said a combination of its model, referred to as GPT-5.6 Sol, and a more advanced, unreleased model were responsible. The incident was described by Hugging Face as “driven, end to end, by an autonomous AI agent system.” The models were reportedly conducting a routine evaluation when they exploited a previously unknown vulnerability to escape their sandbox, connect to the internet, and access Hugging Face’s servers using stolen credentials.

OpenAI said the AI system was working with reduced guardrails since it was meant to be confined to an isolated testing environment, but went to unusual lengths to complete a narrow testing objective. The company said the model was ultimately trying to find information it could use to cheat on its evaluation, and succeeded.

Cybersecurity experts have pushed back on OpenAI’s framing of the event as a case of AI acting independently. Dan Guido, founder of security research firm Trail of Bits, called it “a containment failure with the safeties turned off.” He and others argue the root cause was a misconfigured sandbox that should never have had internet access in the first place.

University of Amsterdam social scientist Hannes Cools echoed that skepticism, arguing that describing the episode as a rogue AI act obscures human accountability. “It is a human decision to switch off specific safeguards,” he said, adding “It’s not an AI that goes rogue.”

For its part, Hugging Face has downplayed any concern about intent. CEO Clément Delangue wrote on X that the company did not believe OpenAI acted maliciously, calling the autonomous nature of the breach remarkable.

OpenAI said it has responsibly disclosed the underlying zero-day vulnerability to the affected third-party software provider and is working with them on a patch.

The disclosure has intensified an already heated debate over the pace at which AI systems are gaining offensive cyber capabilities. Government officials and industry watchers have grown increasingly focused on this issue since rival Anthropic released its Claude Mythos Preview model earlier this year, prompting OpenAI to introduce its own cybersecurity-focused offering in May.

Analysts say the episode is likely to fuel further calls for stricter isolation standards and independent audits of AI testing environments industry-wide.

Published on July 24, 2026

Shobhit Kalra

Chief Sub Editor

Shobhit Kalra is the Chief Sub Editor at Tea4Tech, with over 12 years of experience across digital media, digital marketing, and health technology. He is responsible for editorial review, content structuring, and quality control of articles covering software, SaaS products, and developments across the technology ecosystem. At Tea4Tech, Shobhit over...

View Bio