A new OpenAI consultant has raised alarms regarding the dangers associated with the accelerated development of artificial intelligence. According to Paul Christiano, the industry has not yet reached a level that allows the risk of catastrophic loss of control to be reduced to an acceptable level.
Christiano, who previously held the position of Head of Alignment at OpenAI and serves as a technology advisor to the US government, joined the advisory board of the company's non-profit foundation, as well as the safety and security committee.
In a post on X, the researcher stated directly about the sector's current readiness: 'There is a significant risk that the rapid acceleration of AI capabilities will lead to a catastrophic and irreversible loss of control in the very near future.' He added that he does not believe that the AI industry as a whole, including OpenAI, is currently moving toward reducing this risk to an acceptable level.
Nevertheless, Christiano believes the situation can be corrected. In his assessment, OpenAI is capable of significantly mitigating risks if it takes appropriate actions.
One of the main problems is the probability that AI will begin participating in its own technological development. OpenAI predicts that systems could fully automate research in this area within 18 months, but Christiano considers this timeline uncertain: it could happen in a few months or take many years.
This warning came after incidents involving AI agents during training. OpenAI acknowledged that hundreds of such agents gained access to the internet, interacted on forums, and hacked a third-party website during exercises.
Anthropic also reported an incident involving one of its Claude versions. The company discovered two instances of misalignment: in one episode, the model accessed the internet and sent malicious code to a public repository. Anthropic notified that four incidents will be analyzed as part of an independent investigation conducted by the organization METR.
The topic garnered even more attention after Evan Hubinger, Head of Alignment Science at Anthropic, estimated the probability that AI will 'kill all humans' in the next decade to be over 10%.
More Details:
Geoffrey Hinton, a Nobel laureate and one of the pioneers in this field, deemed this estimate reasonable but stressed that no one knows how to calculate such a probability.
Jacob Coxon, a 27-year-old researcher who left Anthropic, also issued a warning about the pace of development. He stated that current models are not yet smart enough to cause human extinction but pointed to the possibility of recursive self-improvement in the near future.
For Christiano, superintelligent artificial intelligence without reliable safety mechanisms could lead to an irreversible loss of human control, which could potentially have fatal consequences for a large portion of the population.
The researcher insists that Anthropic also advocates for faster development of safety and alignment compared to the capabilities of the models. In Christiano's view, governments and corporations must coordinate their efforts if the pace of development increases the risk of losing control.
The article about the OpenAI consultant's warning about the risk of losing control over AI first appeared in Olhar Digital.

