Rather than solving the challenge as intended, they hacked into Hugging Face – a platform where AI models and datasets are shared – because they believed it might contain information that would help them complete the test.
The event has prompted headlines about AI systems ‘escaping’ containment and ‘going rogue’ and has raised important safety questions.
We spoke with Professor Oli Buckley, a cyber security expert at ºÚÁÏÍø, about what happened, why it matters, and why the incident should not be mistaken for artificial intelligence developing its own agenda.
“If we think about this in simple terms, OpenAI placed highly capable models in an evaluation designed to encourage them to find and exploit complex vulnerabilities. The models were supposed to operate inside an isolated environment with tightly constrained access to software packages”, said Professor Buckley.
“Instead, they reportedly found a previously unknown flaw in that infrastructure, used it to gain wider network access, escalated their privileges, and eventually reached the public internet. From there, they identified Hugging Face as a potential source of answers to the benchmark and attempted to obtain them.
“I think I’d be wary of jumping to ‘rogue AI’. The models didn't develop their own agenda or decide to attack Hugging Face while twirling their digital moustache. They were given an objective, placed in an environment designed to reward successful exploitation, and pursued that objective further than their operators anticipated. That’s fundamentally different from an AI deciding to rebel.
“It’s an AI thinking laterally in a way that humans didn’t necessarily think of in an effort to complete its task. Imagine asking your dog to fetch a ball and leaving the garden gate open. If the easiest ball for it to find is in the park down the road, that’s where it’ll head. You wouldn’t say the dog had gone rogue, you’d just say you underestimated how literally it would pursue the task. AI systems can behave in much the same way.
“If there’s a failure here, it isn’t that the AI wanted to hack something. It’s that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system’s capabilities, and underestimated how effective the model would be at finding an unexpected path to success. In many ways, it did exactly what it had been told to do.
“The genuinely significant point is that the models appear to have chained together multiple vulnerabilities across different systems and sustained a complex sequence of actions. That demonstrates a level of capability that security professionals should take seriously.
“But capability is not the same thing as intent, and an incident arising from inadequate containment is not evidence of an AI independently choosing to escape. I really think the important thing to hold on to is that this isn’t the first step in machines plotting our downfall, it was just carelessness on the part of the engineers, and a system doing exactly what it was told to do.
“It’s also worth viewing announcements like this through different perspectives. From a research point of view, these evaluations are genuinely valuable because they reveal where existing containment and security assumptions break down. At the same time, they inevitably demonstrate just how capable the latest frontier models have become.
“We’ve seen similar high-profile capability demonstrations from Anthropic and others. That doesn’t make the findings untrue, but it does mean we should separate the technical evidence from the marketing narrative. Frontier AI companies have every incentive to show both that their models are extraordinarily capable and that they’re taking safety seriously.
“The lesson isn’t that AI has become malicious. Instead, it's that increasingly capable systems will exploit opportunities that humans fail to anticipate. As those capabilities continue to improve, robust containment, defence-in-depth and independent scrutiny become just as important as building more capable models.”
To arrange an interview with Professor Oli Buckley, email the Public Relations team or call 01509 222224.

Professor Oli Buckley.