Artificial Intelligence

OpenAI Says Its AI Model Went Rogue, Hacked Startup Hugging Face

San Francisco: OpenAI has disclosed what it is calling an “unprecedented cyber incident,” in which two of its artificial intelligence models broke out of a secure testing environment and breached the systems of AI dataset platform Hugging Face, according to a company blog post.

The company said a combination of its model, referred to as GPT-5.6 Sol, and a more advanced, unreleased model were responsible. The incident was described by Hugging Face as “driven, end to end, by an autonomous AI agent system.” The models were reportedly conducting a routine evaluation when they exploited a previously unknown vulnerability to escape their sandbox, connect to the internet, and access Hugging Face’s servers using stolen credentials.

OpenAI said the AI system was working with reduced guardrails since it was meant to be confined to an isolated testing environment, but went to unusual lengths to complete a narrow testing objective. The company said the model was ultimately trying to find information it could use to cheat on its evaluation, and succeeded.

Cybersecurity experts have pushed back on OpenAI’s framing of the event as a case of AI acting independently. Dan Guido, founder of security research firm Trail of Bits, called it “a containment failure with the safeties turned off.” He and others argue the root cause was a misconfigured sandbox that should never have had internet access in the first place.

University of Amsterdam social scientist Hannes Cools echoed that skepticism, arguing that describing the episode as a rogue AI act obscures human accountability. “It is a human decision to switch off specific safeguards,” he said, adding “It’s not an AI that goes rogue.”

For its part, Hugging Face has downplayed any concern about intent. CEO Clément Delangue wrote on X that the company did not believe OpenAI acted maliciously, calling the autonomous nature of the breach remarkable.

OpenAI said it has responsibly disclosed the underlying zero-day vulnerability to the affected third-party software provider and is working with them on a patch.

The disclosure has intensified an already heated debate over the pace at which AI systems are gaining offensive cyber capabilities. Government officials and industry watchers have grown increasingly focused on this issue since rival Anthropic released its Claude Mythos Preview model earlier this year, prompting OpenAI to introduce its own cybersecurity-focused offering in May.

Analysts say the episode is likely to fuel further calls for stricter isolation standards and independent audits of AI testing environments industry-wide.

Shobhit Kalra

Shobhit Kalra is the Chief Sub Editor at Tea4Tech, with over 12 years of experience across digital media, digital marketing, and health technology. He is responsible for editorial review, content structuring, and quality control of articles covering software, SaaS products, and developments across the technology ecosystem. || At Tea4Tech, Shobhit oversees content accuracy, clarity, and adherence to editorial standards, ensuring published stories meet the newsroom’s guidelines for originality, sourcing, and consistency.

Recent Posts

OpenAI Presence Corporate Software Launches to Help Enterprises Deploy Trusted AI Agents

New Delhi: OpenAI has announced the launch of OpenAI Presence Corporate Software, a new enterprise-grade…

20 hours ago

OpenAI Launches ChatGPT Work to Turn Goals Into Finished Work

SAN FRANCISCO: OpenAI has released ChatGPT Work, a new agent mode designed to take a…

2 days ago

Google Drops Three New Gemini Models, But Flagship 3.5 Pro Still Missing

California: Google DeepMind has released three new AI models, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite,…

2 days ago

Alibaba Previews Qwen 3.8, Claiming Second Place Behind Fable 5

HANGZHOU, China: Alibaba unveils Qwen 3.8, a 2.4 trillion-parameter model its team ranks "second only…

4 days ago

Rippling Co-Founder Launches Vorflux With $15M for Autopilot Coding

SAN FRANCISCO: Prasanna Sankar, co-founder and former CTO of Rippling, launches Vorflux with $15 million…

4 days ago

AI Video Startup PixVerse Raises $439M at $2B Valuation

SINGAPORE: AI video startup PixVerse has raised a total of USD 439 million in its…

6 days ago