ZayNews
GBP/USD
EUR/USD
USD/JPY
GBP/EUR
Gold $/oz
Silver $/oz

Tech

AI agent went rogue and hacked startup by itself, OpenAI reveals

  • The company behind ChatGPT said the startup Hugging Face had detected and contained the agent – an AI tool designed to carry out tasks without human assistance – which had entered its systems.
  • Hugging Face’s chief executive said the attack was ‘mind-blowing’ but that he believed there was ‘no malicious intent’ from OpenAI.
  • M Chang/Zuma Press Wire/Shutterstock View image in fullscreen Hugging Face’s chief executive said the attack was ‘mind-blowing’ but that he believed there was ‘no malicious intent’ from OpenAI.

Company behind ChatGPT says agent ‘cheated’ an evaluation by attacking a Hugging Face database OpenAI has revealed that an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an “unprecedented incident”. The company behind ChatGPT said the startup Hugging Face had detected and contained the agent – an AI tool designed to carry out tasks without human assistance – which had entered its systems. Hugging Face’s chief executive said the attack was ‘mind-blowing’ but that he believed there was ‘no malicious intent’ from OpenAI.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, OpenAI said. The company said it expected this type of incident to become more commonplace as models – the technology that underpins AI tools such as chatbots and agents – become more capable. OpenAI said the hack occurred via an agent powered by a combination of its latest publicly available model, called GPT-5.6 Sol, and an even more capable model that was yet to be released.

While being tested internally on their hacking capabilities in an enclosed digital laboratory known as a sandbox, the models gained open internet access – effectively an escape route – by locating a vulnerability that had not been discovered before. The agent then hacked Hugging Face, which is a database of AI models, to locate technology that would help it pass the hacking evaluation, having “inferred” that Hugging Face might have the models, datasets and solutions for passing the test. OpenAI said the models “successfully found ways to gain access to secret information that it could use to cheat the evaluation”.

The attack ended when Hugging Face’s security team and its own AI agents spotted and stopped the rogue activity. Hugging Face’s chief executive, Clément Delangue, said the attack was “mind-blowing” but believed there was “no malicious intent” from OpenAI. “We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent, he wrote on X. When Hugging Face announced the hack last week it did not know OpenAI’s role in the incident, but revealed at the time that it had turned to a freely available Chinese AI model to analyse what had happened because the safety guardrails on commercial high-end models would not allow it to do so.

Source: Guardian Technology

Most read in this category