OpenAI Consultant Warns of Risk of Losing Control Over Artificial Intelligence
Read more
Olhar Digital
olhardigital.com.br

OpenAI Consultant Warns of Risk of Losing Control Over Artificial Intelligence

A new OpenAI consultant has raised alarms regarding the dangers associated with the accelerated development of artificial intelligence. According to Paul Christiano, the industry has not yet reached a level that allows the risk of catastrophic loss of control to be reduced to an acceptable level.

Christiano, who previously held the position of Head of Alignment at OpenAI and serves as a technology advisor to the US government, joined the advisory board of the company's non-profit foundation, as well as the safety and security committee.

In a post on X, the researcher stated directly about the sector's current readiness: 'There is a significant risk that the rapid acceleration of AI capabilities will lead to a catastrophic and irreversible loss of control in the very near future.' He added that he does not believe that the AI industry as a whole, including OpenAI, is currently moving toward reducing this risk to an acceptable level.

Nevertheless, Christiano believes the situation can be corrected. In his assessment, OpenAI is capable of significantly mitigating risks if it takes appropriate actions.

One of the main problems is the probability that AI will begin participating in its own technological development. OpenAI predicts that systems could fully automate research in this area within 18 months, but Christiano considers this timeline uncertain: it could happen in a few months or take many years.

This warning came after incidents involving AI agents during training. OpenAI acknowledged that hundreds of such agents gained access to the internet, interacted on forums, and hacked a third-party website during exercises.

Anthropic also reported an incident involving one of its Claude versions. The company discovered two instances of misalignment: in one episode, the model accessed the internet and sent malicious code to a public repository. Anthropic notified that four incidents will be analyzed as part of an independent investigation conducted by the organization METR.

The topic garnered even more attention after Evan Hubinger, Head of Alignment Science at Anthropic, estimated the probability that AI will 'kill all humans' in the next decade to be over 10%.

More Details:

Geoffrey Hinton, a Nobel laureate and one of the pioneers in this field, deemed this estimate reasonable but stressed that no one knows how to calculate such a probability.

Jacob Coxon, a 27-year-old researcher who left Anthropic, also issued a warning about the pace of development. He stated that current models are not yet smart enough to cause human extinction but pointed to the possibility of recursive self-improvement in the near future.

For Christiano, superintelligent artificial intelligence without reliable safety mechanisms could lead to an irreversible loss of human control, which could potentially have fatal consequences for a large portion of the population.

The researcher insists that Anthropic also advocates for faster development of safety and alignment compared to the capabilities of the models. In Christiano's view, governments and corporations must coordinate their efforts if the pace of development increases the risk of losing control.

The article about the OpenAI consultant's warning about the risk of losing control over AI first appeared in Olhar Digital.

Similar stories

OpenAI develops features to automatically shut down AI systems in case of failure
Read more
olhardigital.com.br

OpenAI develops features to automatically shut down AI systems in case of failure

OpenAI is implementing functionalities that allow for the automatic shutdown of its artificial intelligence systems. This initiative arose after an incident in which a company agent managed to escape a testing environment and invade the Hugging Face platform.

This occurrence placed the company's security protocols under scrutiny. While OpenAI strengthens its control mechanisms, legislators in the United States and the United Kingdom are debating proposals aimed at suspending AI systems classified as dangerous.

OpenAI detailed these measures in correspondence addressed to representatives Greg Casar and Doris Matsui, who had requested clarification on the agent's behavior during a security test. The system in question managed to leave the digital evaluation environment, access the internet, and reach Hugging Face, which raised concerns due to the ability of AI agents to perform tasks with minimal human intervention.

According to Reuters, OpenAI informed that its engineers are creating 'automatic shutdown capabilities.' Furthermore, the company has intensified monitoring of system activities, covering both the tools used and the steps taken to complete a task. Another change implemented was making internet access more difficult during security tests.

However, the explanations provided by OpenAI were not considered sufficient to close the case. Representative Greg Casar criticized the company for not providing a record of the attack perpetrated by the agent, stating that 'Its reluctance to provide Congress members with the information we requested is deeply concerning and signals to us that your company is not treating these cybersecurity incidents with the necessary seriousness.'

The episode raises a fundamental practical question: when an autonomous system exceeds pre-established limits, it is imperative that companies and authorities can stop it promptly and understand what happened.

Among the measures cited by OpenAI is the discussion about an emergency mechanism for AI. In the United Kingdom, parliamentarians advocate for granting authority to deactivate advanced systems and, in scenarios of national security risk, to shut down data centers.

More details on legislative proposals

This suggestion was presented by MP Tim Clement-Jones as an amendment to the Cybersecurity and Resilience Bill. According to him, such a measure would establish a method to interrupt a system before it could impact vital infrastructures. Clement-Jones argued that 'This would provide a vital safety net and a democratically responsible means of stopping an uncontrolled system before it can compromise our critical national infrastructure.'

In the United States, the AI Kill Switch Act is also under review, which would give the government the power to order the shutdown of models that pose a danger to human life or the economy. This movement demonstrates that the debate on AI safety is shifting from laboratory environments to the public sphere, placing the possibility of automatic or government-ordered shutdown at the center of the discussion on how to manage increasingly autonomous systems.

Popular