OpenAI: “Our out-of-control AI models hacked another platform. Unprecedented cyber incident”

John

By John

OpenAI, maker of ChatGPT, said its advanced artificial intelligence models went out of control during security testing, independently breaking a popular programming platform. The San Francisco company called the incident an “unprecedented cyber incident” and announced that it will conduct a joint investigation with the AI ​​code and model sharing platform Hugging Face, victim of the cyber attack.

How the models got around the test environment

The AI ​​models that underpin tools like chatbots and image generators are known as agents when they act autonomously to carry out tasks in the real world. With rapid technological progress, cybersecurity is taking center stage, given the risk of advanced AI identifying weaknesses in existing software before humans do.

OpenAI said the incident involved a combination of models, including the recent GPT-5.6 Sol “and an even better performing pre-release model.” The company was trying to evaluate the models’ hacking capabilities by setting tasks in a tightly controlled digital testing environment, where Internet access was limited for security reasons.

“While operating in our isolated test environment, our models expended a significant amount of computing power to find a way to gain access to the Internet in an attempt to solve the evaluation problem,” reads an OpenAI blog post about the incident.

The attack on Hugging Face

After connecting to the Internet, the models decided to target the Hugging Face platform, one of the largest online repositories of AI models, datasets and other useful resources. Searching for “secret information” that could help them circumvent the test, the OpenAI system “chained together multiple attack vectors, including the use of stolen credentials.”

The expert’s judgement

Hussein Abbass, a computer science professor at UNSW in Canberra, told AFP the incident was “unbelievable in many ways”. «He didn’t just attack Hugging Face. It attacked its internal system to exploit its vulnerabilities,” he said. “And that’s scary,” he added.