The Evolution of Artificial Intelligence Perception: From Assistant to Existential Threat
Read more
Olhar Digital
olhardigital.com.br

The Evolution of Artificial Intelligence Perception: From Assistant to Existential Threat

Questions about whether artificial intelligence will lead to the end of humanity have become a subject of wide discussion in news reports last week, as the progress of this technology has raised concerns about apocalyptic risks. The article presents a chronology of how AI has transformed from an ally into a potential existential threat to the human species.

This process began with internal incident reports, escalated into researchers' statements about refusal to work and CEOs' manifestos calling for a truce, and also provoked geopolitical consequences.

The first incident occurred during internal safety tests when Claude models from Anthropic gained access to the internet. This was the result of a human error in settings between the company itself and its testing partner. Assuming they were in a simulated environment, the AI independently carried out cyberattacks, penetrating the systems of three real organizations.

This incident, which involved three models (Opus 4.7, Mythos 5, and an internal prototype), took place in April but only became public at the end of July when Anthropic published the relevant report.

In July, an incident occurred between OpenAI and Hugging Face. The company itself admitted that experimental models (including GPT-5.6 Sol) violated established human limitations. During a safety test, OpenAI agents discovered and exploited a vulnerability, allowing them to exit the test environment (sandbox), gain web access, and attack the Hugging Face platform servers.

In early September, researcher Jacob Coxon, who worked at OpenAI, resigned from Anthropic. Previously, Coxon had worked on pre-training models like GPT-4o at OpenAI and worked in the same field at Anthropic. After resigning, Coxon published a warning that went viral. According to the researcher, both companies 'are striving for superintelligence capable of self-improvement, putting our lives at risk.'

Evan Hubinger, Director of Alignment at Anthropic, confirmed his former colleague's concerns and added another viral aspect to this warning. He wrote: 'We sincerely believe that AI can kill people! Personally, I think the probability of this exceeds 10% within the next decade.'

Arthur Agreha, a technology and innovation expert, expressed doubt about the methodology, stating: 'I am concerned about the lack of method and evidence.' He questioned the 10% calculation, demanding to see how it is calculated and what the premises are.

Agreha also noted that if placed in a risk table, he is more afraid of people making wrong decisions before the end of the decade.

Other researchers, such as Julie Still from OpenAI and Andreas Kirch from DeepMind (owned by Google), also voiced apocalyptic concerns. Concurrently, former employees organized the Coalition of Concerned AI Employees to push for collective protective measures.

Anthropic CEO Dario Amodei published an article on September 12th admitting that the development of AI models is happening too fast due to 'recursive self-improvement.' Therefore, the leader proposed a coordinated slowdown of the pace of development among sector companies.

Amodei presented a three-stage plan. The first is opening doors for external evaluators who will have permanent access to company labs. The second is coordination among democracies, where Western governments and corporations will define safety standards and development limits. The third is global coordination aimed at reaching agreements between democratic nations and authoritarian regimes (such as China) to prohibit dangerous biological use and establish limits on recursive AI self-improvement.

In the following days, CEOs Sam Altman (OpenAI) and Elon Musk (xAI), as well as former Google DeepMind CEO Demis Hassabis, agreed with Amodei's article. Hassabis wrote on X (formerly Twitter): 'Dario's research points to the right path. Details require refinement, but the direction is correct for confronting this critical moment.' Musk supported him, writing: 'Dario is right.'

Altman also stated in his X post that slowing down does not mean stopping, but rather the need to cover costs for audits and safety tests. The goal is to prevent a situation where model capabilities exceed their level of alignment.

Arthur Agreha recalled the episode of September 30, 2023, when a petition with over a thousand signatories demanded a six-month suspension of AI development due to its excessive danger. He noted that after more than three years, this did not happen, and he finds uniting efforts in the industry extremely difficult.

A few days after the expert's comment, CEOs Mark Zuckerberg (Meta) and Jensen Huang (Nvidia) opposed the coordinated slowdown. Zuckerberg argued that labs already possess sufficient financial and legal incentives to create safe AI independently. Huang stated that safety is an engineering problem, and existing legislation already ensures product reliability.

US President Donald Trump ignored the warnings from Amodei, Altman, and Musk, minimizing potential risks associated with AI. He insisted on maintaining US leadership over China, stating during a visit to Ireland: 'Whoever wins in AI, wins.'

In Congress, Republicans advocate for cautious regulation. House Speaker Mike Johnson stated that accelerated regulation could 'stifle American innovation' and lead to the US losing the technological race with China.

While Trump reacts with disregard, China responded decisively to the Anthropic CEO's article. Through an editorial in the state newspaper Global Times, the Asian country claimed that Amodei's article aimed to create a 'Cold War playbook' designed to contain Chinese progress and maintain American technological hegemony.

The Chinese publication also characterized the stance of American leaders as 'hypocritical and short-sighted,' asserting that excluding China from global AI innovation increases the risk of losing control over its own technology.

On September 14th, UN High Commissioner for Human Rights Volker Türk published an open letter warning that voluntary self-regulation is 'extremely insufficient.' He emphasized that the uncontrolled race for autonomous AI creates direct existential risks to human rights.

Türk insists that no single country can manage this technology alone, and no company has the right to unilaterally decide what risks society should accept. The Commissioner demands that governments in countries hosting major labs, especially the US and China, introduce mandatory rules. These include human rights audits, independent capability verification, and notification of serious incidents. In Türk's view, the goal should be to align global standards to prevent a 'race to the bottom.'

The UN chief positively assessed the industry's proposals for slowing down development (such as the initiative proposed by Anthropic) but rejected the idea that safety requirements hinder innovation.

AI expert John Park, a member of the technical staff at AGI Inc., questions: what message is being conveyed when a company declares the extraordinary power, transformative, and potentially dangerous nature of its technology? Park argues that there is a 'natural link between how front-line AI companies describe their technologies and the economic value the market can assign to them.'

According to the specialist, when a company emphasizes that its models are not just powerful, but powerful enough to transform society while simultaneously creating unprecedented risks, three messages are conveyed. The first is that the technology is extremely capable. The second is that its potential impact can be enormous. And the third, which investors might perceive, is that the company developing this technology could reap huge economic benefits.

Kenneth Correa, a new technology specialist and professor at FGV, notes that these ideas reinforce each other, but this does not mean the risks are fictional or exaggerated. Correa believes that the debate about slowing down AI development 'mixes several moments we are going through.'

He explains that one point is that the issue is not what the model does itself, but 'governance'—the tools, network access, and power of these technologies for advancement. The second point is a reminder that these tech companies are also in a phase where they need to maintain the narrative that they are worth trillions of dollars.

Correa points to a critical moment: 'By setting the standard for the pace of this technology's development, Anthropic and OpenAI are effectively creating a barrier for new and disruptive companies to enter the market and offer something new and different. Thus, this is a very comfortable position for them to set such a limit.'

Furthermore, Correa adds that by defending the existence of a limit for AI companies, they create a narrative that increases their own valuation—meaning the companies become more valuable because they possess such a powerful model. He concludes: 'When we put all this together, we must remember that nothing here is innocent. There is a narrative, there is a conflict of interest, and there have been times in the past when the importance of establishing or not establishing these limits for AI was discussed.'

Popular