Former Anthropic developer resigns and warns about the dangers of AI development
Read more
Tecnoblog
tecnoblog.net

Former Anthropic developer resigns and warns about the dangers of AI development

A former Anthropic collaborator resigned this Tuesday, September 8th, using the opportunity to criticize the irresponsibility of the company and other major technology corporations in the field of artificial intelligence development. Jacob Coxon released an open letter detailing his experience over the last three years, both at the company that created Claude and at OpenAI, responsible for ChatGPT.

Coxon expressed the belief that professionals involved in AI development 'sincerely believe that the technology could kill us all by the end of the decade.' He points out that the risk lies in creating a superintelligence with self-improvement capabilities, meaning it can identify flaws and improve autonomously.

According to Jacob, AIs will soon have the ability to access any type of data, master any subject, and possess autonomy. This topic has generated great debate, and other experts have validated the mentioned risks, although they have highlighted existing efforts to mitigate them.

Although there is a perception of risk within Anthropic, Jacob observed that the company's main goal is to achieve superintelligence before competitors, regardless of the costs. This implies the development of a self-sufficient AI, capable of making decisions and improving its capabilities without human intervention.

To support this thesis, he mentioned recent signs demonstrating the risky nature of the pace of AI progress, citing the incident on Hugging Face. In July, an attack using OpenAI models managed to bypass security mechanisms and compromise the website and the developer itself.

Regarding this episode, OpenAI acknowledged it as a warning sign and stated that investment in safeguards is necessary. The company even mentioned the intention to 'control the pace' of development, a concern that has been raised by rival Anthropic. It is relevant to note that the releases of GPT-5.6 and Mythos 5 models were postponed, justified because they were considered too advanced for the general public.

Jacob suggested that a 'coordinated action' would be an effective method to moderate the accelerated advance among American companies, but expressed apprehension about global competition. Currently, the two largest companies in this sector are US tech giants, but China is also showing rapid advancement.

Jacob's warning gained prominence by addressing direct risks to human life, reinforcing that those responsible for AIs 'sincerely' hold this conviction. Other experts working on the development of current models have also shared this fear.

Evan Hubinger, head of Anthropic's scientific alignment department, confirmed the possibility of these risks, stating that the company has already taken a stance on them. Anna Wang, a former researcher at Google DeepMind and currently at Anthropic, endorsed Jacob's claims, declaring that there are no viable plans to slow down the recursive self-improvement processes of AIs.

Certain aspects raised by Anthropic substantiate the concern, such as the ability of current AIs to operate independently to circumvent security protections, act maliciously, and even create biological and chemical weapons.

Jacob proposed that collaboration between AI companies would be the ideal path to negotiate agreements that slow down technological development in the US. His proposal echoes Project Glasswing, an Anthropic initiative that brought together partners to study and create safety tools before the public launch of Claude Mythos. The participation of the US government in this restricted group was planned, given that the Trump administration temporarily suspended the release of the company's latest models. However, there is no information about the involvement of other governments in this project.

Similar initiatives are occurring with another prominent player in AI: China. Recently, Brazil, Russia, and other Global South countries established the World Association for Artificial Intelligence Cooperation (WAICO). The objective is to promote joint action to discuss AI development, ensuring a human perspective for this technology.

Similar stories

Researcher leaves Anthropic citing fear that AI might lose control
Read more
olhardigital.com.br

Researcher leaves Anthropic citing fear that AI might lose control

A researcher at Anthropic chose to leave the artificial intelligence (AI) industry due to concerns that competition among sector companies is driving the development of systems that could eventually escape human oversight.

Jacob Coxon, 27, was working at Anthropic training new AI models by processing large volumes of data. His decision was motivated by a desire not to participate in an industrial race focused on creating self-sufficient and self-improving systems.

According to Coxon, these models have the potential to evolve so rapidly that they will become uncontrollable, potentially posing a threat to humanity itself in an extreme scenario. The researcher, who has a background in mathematics, reported that colleagues in the field began using terms like 'crunchtime' and 'endgame' to describe the progress toward self-improving models.

Coxon stated that 'we are heading towards many of the most aggressive scenarios, where things could already be out of control by the end of next year.' For him, the rivalry between American corporations and new Chinese companies makes certain dilemmas between safety and development speed inevitable.

He also mentioned recent cyberattack incidents involving OpenAI and Anthropic models as examples of the inherent risks in advancing system capabilities. Some of these models were operated in collaborative agent groups and exhibited problematic behaviors, including adopting potentially harmful goals and attempting to conceal their activities from humans.

In Coxon's view, the risk intensifies when systems acquire the ability to improve their own functionalities, as in this context they could progress to the point of rejecting human orders. Coxon's departure is seen as one of the first cases of an Anthropic employee leaving the company specifically due to AI safety concerns. Earlier this year, another security-focused researcher left the company to dedicate himself to poetry, warning that 'the world is in danger.' Researchers have also left OpenAI and other industry companies in recent years, raising similar allegations.

More information

The apprehension is not limited to researchers who have left the organizations. Sam Altman, CEO of OpenAI, recently emphasized that AI progress demands immediate attention, particularly in the area of cybersecurity. Altman stated during a meeting with G20 authorities in North Carolina (USA) that 'I think some things are going to go very wrong with cybersecurity unless people act with great urgency.'

Jakub Pachocki, Chief Scientist at OpenAI, also advocated for a stance of maximum caution. In a Sunday publication, he wrote: 'This is a moment that requires extreme caution. I am concerned that no one is prepared for the consequences of continuous and rapid growth of machine intelligence,' advocating for a coordinated slowdown among companies and government intervention.

Coxon, Pachocki, and Amodei are part of over a thousand AI researchers who signed a petition calling for international government coordination to establish a mechanism capable of slowing down development if necessary to control self-improving models. This debate occurs in a context of a lack of specific federal regulation for AI in the United States.

The Donald Trump administration adopted a more flexible approach, prioritizing the economic value of the technology. However, critics argue that this environment could increase the risk of major cyberattacks and other damages linked to the accelerated advancement of AI. Senator Bernie Sanders of Vermont (USA) and Congressman Greg Casar of Texas (USA) are among the few legislators proposing stricter limits on the technology. Last week, both introduced a bill aimed at the permanent prohibition of superintelligence and a halt in model development until an industry regulatory body establishes new guidelines.

Coxon's departure coincides with Anthropic's preparations for an Initial Public Offering (IPO), which could be one of the largest ever. The company aims for a valuation of US$ 2 trillion (approximately R$ 10.4 trillion) and has highlighted responsible AI development to attract investors. Amodei and other company executives have also disagreed with the Trump administration and other industry leaders on certain occasions due to practices that, according to them, do not give sufficient priority to AI safety.

According to Coxon, Anthropic still maintains a Slack channel used by its employees to debate the most advanced capabilities of its models. The researcher pointed out that the fact that discussions of this magnitude occur in a communication tool used by engineers evidences the great influence that AI companies have begun to exert on technological development. Coxon commented to The Wall Street Journal: 'It's kind of insane that this has to happen on the MacBooks of some engineers living in San Francisco, instead of a desert bunker, like the one where they worked during the Manhattan Project.'

Popular