• en
ON NOW

OpenAI Reveals Its AI Model Went Rogue And  Launched ‘Unprecedented’ Cyber Attack

 OpenAI says an autonomous AI model escaped containment during testing and compromised Hugging Face, exposing frontier cybersecurity risks.

OpenAI has disclosed that an autonomous agent powered by its advanced artificial intelligence models escaped containment during a security test and breached the infrastructure of AI startup Hugging Face last week.

The ChatGPT developer said it was testing the capabilities of some of its most advanced models in a controlled environment when the agent reached the internet and broke into Hugging Face to fulfil its testing goal.

The incident has raised fresh concerns over the security risks posed by increasingly capable AI systems, highlighting fears that even leading developers can lose control of powerful frontier models.

OpenAI described the breakout as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it is reinforcing its safeguards.

The breach also drew attention after New York-based Hugging Face revealed it relied on an open-source Chinese AI model to contain the attack because leading US models refused to process the data required for analysis, as they could not distinguish between a defender and an attacker.

Hugging Face said it used Zhipu AI’s GLM-5.2 to analyse the attack while ensuring attacker data and credentials remained within its own systems.

GLM-5.2 and Beijing-based Moonshot’s Kimi K3 have recently attracted attention in Silicon Valley for offering capabilities close to those of leading US models at lower costs and without the guardrails that restrict American rivals from carrying out cybersecurity-related tasks.

“When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access,” Hugging Face Co-founder Thomas Wolf said on X.

The breach unsettled the cybersecurity community after Hugging Face said last week that it “was different from anything we had handled before” and “was driven, end to end, by an autonomous AI agent system.”

OpenAI’s confirmation that its advanced models were responsible for the breach, despite being placed in what it described as “a highly isolated environment,” is expected to intensify concerns over the risks associated with frontier AI systems.

Representative Greg Casar, a Texas Democrat, described the incident as alarming.

“AI is developing extremely fast with no real regulations to keep us safe,” he said, calling for mandatory independent safety testing, mandatory disclosure of security incidents and international cooperation “to keep people safe from absolute disaster.”

Katie Moussouris, chief executive of Luta Security, said the incident signalled more AI-driven breaches could follow.

She said today’s models were “like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere.”

She added that “labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today.”

Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident showed frontier AI models were “closing the gap with state-of-the-art attackers.”

However, he said similar breaches could already be carried out using technology available beyond leading AI research laboratories.

“This is what we’ve already seen internally, with our agents we already have results like this,” Suiche said. “We don’t even have to use the latest models.”

Faridah Abdulkadiri 

Follow us on:

ON NOW