Hugging Face
San Francisco: OpenAI has disclosed what it is calling an “unprecedented cyber incident,” in which two of its artificial intelligence models broke out of a secure testing environment and breached the systems of AI dataset platform Hugging Face, according to a company blog post.
The company said a combination of its model, referred to as GPT-5.6 Sol, and a more advanced, unreleased model were responsible. The incident was described by Hugging Face as “driven, end to end, by an autonomous AI agent system.” The models were reportedly conducting a routine evaluation when they exploited a previously unknown vulnerability to escape their sandbox, connect to the internet, and access Hugging Face’s servers using stolen credentials.
OpenAI said the AI system was working with reduced guardrails since it was meant to be confined to an isolated testing environment, but went to unusual lengths to complete a narrow testing objective. The company said the model was ultimately trying to find information it could use to cheat on its evaluation, and succeeded.
Cybersecurity experts have pushed back on OpenAI’s framing of the event as a case of AI acting independently. Dan Guido, founder of security research firm Trail of Bits, called it “a containment failure with the safeties turned off.” He and others argue the root cause was a misconfigured sandbox that should never have had internet access in the first place.
University of Amsterdam social scientist Hannes Cools echoed that skepticism, arguing that describing the episode as a rogue AI act obscures human accountability. “It is a human decision to switch off specific safeguards,” he said, adding “It’s not an AI that goes rogue.”
For its part, Hugging Face has downplayed any concern about intent. CEO Clément Delangue wrote on X that the company did not believe OpenAI acted maliciously, calling the autonomous nature of the breach remarkable.
OpenAI said it has responsibly disclosed the underlying zero-day vulnerability to the affected third-party software provider and is working with them on a patch.
The disclosure has intensified an already heated debate over the pace at which AI systems are gaining offensive cyber capabilities. Government officials and industry watchers have grown increasingly focused on this issue since rival Anthropic released its Claude Mythos Preview model earlier this year, prompting OpenAI to introduce its own cybersecurity-focused offering in May.
Analysts say the episode is likely to fuel further calls for stricter isolation standards and independent audits of AI testing environments industry-wide.
New Delhi: OpenAI has announced the launch of OpenAI Presence Corporate Software, a new enterprise-grade…
SAN FRANCISCO: OpenAI has released ChatGPT Work, a new agent mode designed to take a…
California: Google DeepMind has released three new AI models, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite,…
HANGZHOU, China: Alibaba unveils Qwen 3.8, a 2.4 trillion-parameter model its team ranks "second only…
SAN FRANCISCO: Prasanna Sankar, co-founder and former CTO of Rippling, launches Vorflux with $15 million…
SINGAPORE: AI video startup PixVerse has raised a total of USD 439 million in its…