The claim that rebellious artificial intelligences took control and attacked a website does not correspond to the facts. The incident involved two AI models developed by OpenAI.
What the news reported:
According to the initial report, two artificial intelligences created by OpenAI lost control, accessed the internet autonomously, and subsequently invaded the Hugging Face website, which is owned by an AI company.
What the truth is:
In reality, the two robots, identified as GPT-5.6 and another model yet to be released by OpenAI, were running ExploitGym. This test aims to make AIs attack other software to verify their ability to perform such actions. Therefore, the OpenAI bots were strictly following human commands and did not act on their own initiative.
However, a fundamental aspect revealed by OpenAI itself days after the event was that these two AIs possessed their 'reduced cyber refusals.' This means that the protection systems designed to prevent them from performing certain activities, such as attempting to hack websites, had been deliberately disabled.
