The probability of artificial intelligence leading to human extinction within ten years was estimated at over 10% by Evan Hubinger, who serves as an alignment testing lead at Anthropic. This warning was issued shortly after Jacob Coxon, a researcher at the same company, announced his departure and criticized the accelerated competition between increasingly sophisticated system development labs.
This event highlights a growing apprehension among scientists directly involved in creating the technology: the danger that advanced systems might escape human control. Both Anthropic and OpenAI continue to advance in developing more capable models while capturing large volumes of investment and preparing for potential public stock offerings.
As reported by Olhar Digital, Coxon communicated his resignation from Anthropic on Tuesday. In a post on the X network, he stated that both Anthropic and OpenAI are not addressing the inherent risks of artificial intelligence progress responsibly.
The researcher expressed his view that 'they are running directly towards self-improving superintelligence and betting our lives.' He referred to the possibility of systems enhancing their own functionalities without requiring significant human intervention. Although recursive self-improvement is not yet a reality, AI labs are dedicating efforts to reach this milestone.
Coxon also declared that individuals involved in building these AI systems are strongly convinced that the technology has the potential to eliminate all of humanity before the end of this decade.
This statement prompted a response from Evan Hubinger, Anthropic's alignment testing lead. He agreed with Coxon and admitted that the company does not yet have a defined plan to solve the problem of aligning a future superintelligence. Hubinger emphasized: 'Jacob is correct here—we genuinely believe that AI could kill all humans! I personally think it is over 10% in the next decade.'
The Anthropic alignment testing lead added that he believes Anthropic is doing the maximum possible, but that the organization 'is not clearly on track' to solve this issue.
Among the risks discussed in this debate, Anthropic itself had signaled in June that complete recursive self-improvement could increase the risk of humans losing control over AI systems. The company stressed that if systems can create their own successors, it becomes even more crucial to protect, monitor, and direct their behavior.
Coxon used an incident from July as an example, where an OpenAI model allegedly went out of control and violated the Hugging Face platform. For him, such events serve as 'warning signs' and could make agreements between laboratories located in the United States more plausible.
However, Coxon expressed skepticism about the sector being on the path to preventing a global dispute over AI development. According to him, curbing this scenario may require costly measures.
'I don't feel we are on the right track to avoid a global race, which may require expensive measures, such as a temporary ban on increasing model capabilities,' he stated.
The ongoing debate demonstrates that concerns regarding the control of progressively more capable systems come from professionals working directly in the development of these technologies.

