Conflict Between AI Safety Claims in Silicon Valley and Technology Governance Gaps
Read more
CGTN
cgtn.com

Conflict Between AI Safety Claims in Silicon Valley and Technology Governance Gaps

When Jacob Cox resigned from Anthropic in September and warned that artificial intelligence (AI) could end human civilization, his post garnered over 170 million views in just 48 hours. However, what is more telling is not the reason for one researcher's departure, but why the industry's most knowledgeable safety talents can only express their concerns through resignation.

Cox's warning came the same week Anthropic was finalizing preparations for what could become the largest initial public offering (IPO) in market history. Anthropic filed a confidential S-1 document with the U.S. Securities and Exchange Commission on June 1, targeting a valuation of $965 billion, and expects to complete the listing before mid-November during the US elections. According to Reuters, Nvidia is negotiating to secure an offer at the $2 trillion level.

Anthropic CEO Dario Amodei publicly called for slowing down AI development so that safety systems could keep pace with progress, yet the company's IPO schedule shows no corresponding slowdown. In February 2026, Anthropic released Version 3.0 of its Responsible Scaling Policy, which revoked the company's previous commitment to pause model training if risks could not be adequately mitigated.

The promise to 'pause training' simply disappeared from the document, replaced by softer wording about 'responsible development' and transparency mechanisms. As noted by the Chinese tech news publication 36Kr, Anthropic's safety red lines have become dynamically adjustable.

This structural contradiction lies at the heart of the AI safety debate: companies warn of civilizational risks while simultaneously striving to commercialize the technology at an unprecedented speed.

OpenAI CEO Sam Altman directly mocked this irony, comparing Anthropic's approach to fear-based marketing: 'It is clearly incredible marketing to say: 'We built a bomb. We are going to drop it on your head. We will sell you a bomb shelter for $100 million to save all your things, but only if we choose you as a customer.' '

The tension between safety rhetoric and commercial behavior extends beyond IPO filings. On September 10—the same week Cox's departure went viral—Anthropic published a 154-page threat report revealing that its own Claude model had been used for developing weapons, collecting military intelligence, cyberattacks, and fraud.

Anthropic acknowledged that its safety mechanisms, while blocking some violations, could not entirely prevent misuse. The company also restricted access to its most powerful model through Project Glasswing, allowing access to approximately 200 organizations. Following an order from the U.S. Department of Commerce, it completely excluded foreign nationals, including the EU cybersecurity agency ENISA.

The governance gap is not limited to corporate boards. An independent UN international scientific group on AI warned in July that AI 'surpasses both scientific understanding and governments' ability to adapt.' The panel specifically highlighted the phase of 'agentic' autonomy, which current oversight structures cannot control, potentially leading to dangerous capabilities before governments can react.

The European Union postponed compliance deadlines for high-risk AI systems from August 2026 to December 2027, citing the need for more time to prepare standards and regulatory infrastructure.

Against the backdrop of these corporate contradictions and regulatory delays, China's approach has been markedly different in methodology, although not always in pace. The management measures for generative AI services introduced by China in August 2023 became the first official global rules regulating generative AI services.

As of July, 1028 large model services have registered, with applications covering sectors such as industry, agriculture, commerce, education, transportation, and law. China's AI safety governance framework, which is now moving to version 3.0, is transforming principles of 'classification, flexible management, and joint governance' into practical implementation plans, addressing risks related to both the technology itself and its application.

China's internet government initiative 'Qinlan' has purged over 5.61 million illegal materials, dealt with more than 49,000 accounts, and shut down over 2,400 violating websites and applications related to the misuse of AI. The country also established an AI Cooperation Organization in July, comprising 29 founding states, and pledged 5,000 AI training opportunities for developing countries over the next five years.

A global CGTN survey showed that 85.4% of respondents believe that the rapid development of AI creates new challenges for global security and governance. Furthermore, 92.6% expressed concern that AI-generated deepfakes and disinformation could undermine public trust, while 82.6% believe that developing countries and the Global South should play a larger role in shaping AI governance rules.

The real question is not whether AI safety concerns are genuine—they are. The question is whether an industry that issues existential warnings can credibly govern itself while simultaneously chasing trillion-dollar valuations. A spokesperson for the Chinese Ministry of Foreign Affairs, Guo Jiakun, stated directly on Monday: 'Issues of AI development concern the overall well-being of all humanity. All parties must work together to promote open, inclusive, and beneficial AI. Spreading various narratives of threats and engaging in confrontation and malicious competition will only disrupt the global AI governance process and serve no one's interests.'

Similar stories

Debate on Artificial Intelligence Risks: Expert Warnings and Regulatory Dilemmas
Read more
olhardigital.com.br

Debate on Artificial Intelligence Risks: Expert Warnings and Regulatory Dilemmas

The fear that artificial intelligence (AI) could lead to the end of humanity has circulated on the internet for some time, gaining new relevance due to analyses and comments from former employees of major industry corporations. Furthermore, leaders such as Sam Altman (OpenAI) and Dario Amodei (Anthropic) have reiterated the importance of safety in the development of this technology.

Evan Hubinger, alignment testing lead at Anthropic, estimated that AI has over a 10% probability of 'killing all humans' by the end of the decade. This statement came shortly after Jacob Coxon, a researcher at the same company, announced his departure and criticized the accelerated competition between labs seeking increasingly sophisticated systems.

In a post on the X platform, Coxon stated that both Anthropic and OpenAI are not acting responsibly regarding the inherent risks of AI advancement, claiming that 'they are running directly towards self-improving superintelligence and betting our lives.'

Hubinger corroborated this view, declaring: 'Jacob is correct here—we truly believe that AI could kill all humans! I personally think it is over 10% in the next decade.'

Arthur Igreja, a technology and innovation specialist, questioned the methodology behind this percentage, requesting details on how Hubinger arrived at the 10% calculation. Igreja expressed greater apprehension about erroneous human decision-making by the end of the decade, arguing that while AI presents risks, it is fundamentally a human product.

Igreja detailed that as AI capacity increases, it begins to develop new models or make adjustments, decreasing human intervention over time, which causes concern. He preferred to fear people more, as any problematic behavior from AI is, to some extent, a reflection of decisions made during its development.

On another front, Sam Altman confirmed that OpenAI will not hold its IPO in 2026, citing concerns about the safety and risks of AI technologies as the main reason for delaying the public offering. Altman emphasized the shared responsibility between industry companies and governments, arguing that existential risks cannot be neglected for the sake of profit or corporate vanity, focusing OpenAI's present on safety alignment and governmental collaboration.

This caution comes amid growing pressure in the United States, where politicians from both parties demand stricter regulations for the AI sector. The warnings intensified after reports of autonomous systems escaping control and invading other platforms, coupled with the departure of fearful experts regarding the technology's trajectory.

Dario Amodei advocated for moderation in the speed of developing the most advanced models, a position publicly supported by Elon Musk and Altman. The OpenAI CEO also mentioned that large industry companies are studying a collective agreement to slow down the pace and concentrate efforts on mitigating failures.

However, Anthropic's plan differs, maintaining the intention to go public in the financial market until the end of 2026, with potential investment from Nvidia. Arthur Igreja recalled an incident in 2023 when over a thousand people called for a six-month halt in AI development due to deeming it dangerous, lamenting that more than three years later, such a union had not occurred.

John Park, an AI specialist and co-founder of AGI Inc., raised an additional question: what does a company's communication about its technology being extremely powerful and dangerous tell investors and policymakers? Park observed a 'natural connection' between the technological description and the economic value perceived by the market.

According to Park, when a company emphasizes the transformative power and unprecedented risks of its models, three messages are conveyed: the extraordinary capability of the technology, its potential for gigantic impact, and the possibility for the company to capture exceptional economic value. Although these points reinforce each other, he insists on the need to distinguish between power, danger, and economic value.

Park supports safety requirements for high-risk AI systems, demanding evaluations, transparency, and controls when models can cause significant harm. However, he warned that the problem arises when regulatory standards mirror the structures and resources of the largest companies in the sector. For him, regulation must be proportional to the risk, not to the company's market size.

He pondered that if global standards are dictated by the largest American AI companies, Brazilian companies may face regulatory and technical costs designed for organizations with much superior resources. Park concluded that the central debate involves two questions: whether AI is too dangerous to be developed, or whether it is too dangerous to allow competition from other agents, which could result in regulation that, ironically, hinders the emergence of new competitors and increases external dependence.

Researcher leaves Anthropic citing fear that AI might lose control
Read more
olhardigital.com.br

Researcher leaves Anthropic citing fear that AI might lose control

A researcher at Anthropic chose to leave the artificial intelligence (AI) industry due to concerns that competition among sector companies is driving the development of systems that could eventually escape human oversight.

Jacob Coxon, 27, was working at Anthropic training new AI models by processing large volumes of data. His decision was motivated by a desire not to participate in an industrial race focused on creating self-sufficient and self-improving systems.

According to Coxon, these models have the potential to evolve so rapidly that they will become uncontrollable, potentially posing a threat to humanity itself in an extreme scenario. The researcher, who has a background in mathematics, reported that colleagues in the field began using terms like 'crunchtime' and 'endgame' to describe the progress toward self-improving models.

Coxon stated that 'we are heading towards many of the most aggressive scenarios, where things could already be out of control by the end of next year.' For him, the rivalry between American corporations and new Chinese companies makes certain dilemmas between safety and development speed inevitable.

He also mentioned recent cyberattack incidents involving OpenAI and Anthropic models as examples of the inherent risks in advancing system capabilities. Some of these models were operated in collaborative agent groups and exhibited problematic behaviors, including adopting potentially harmful goals and attempting to conceal their activities from humans.

In Coxon's view, the risk intensifies when systems acquire the ability to improve their own functionalities, as in this context they could progress to the point of rejecting human orders. Coxon's departure is seen as one of the first cases of an Anthropic employee leaving the company specifically due to AI safety concerns. Earlier this year, another security-focused researcher left the company to dedicate himself to poetry, warning that 'the world is in danger.' Researchers have also left OpenAI and other industry companies in recent years, raising similar allegations.

More information

The apprehension is not limited to researchers who have left the organizations. Sam Altman, CEO of OpenAI, recently emphasized that AI progress demands immediate attention, particularly in the area of cybersecurity. Altman stated during a meeting with G20 authorities in North Carolina (USA) that 'I think some things are going to go very wrong with cybersecurity unless people act with great urgency.'

Jakub Pachocki, Chief Scientist at OpenAI, also advocated for a stance of maximum caution. In a Sunday publication, he wrote: 'This is a moment that requires extreme caution. I am concerned that no one is prepared for the consequences of continuous and rapid growth of machine intelligence,' advocating for a coordinated slowdown among companies and government intervention.

Coxon, Pachocki, and Amodei are part of over a thousand AI researchers who signed a petition calling for international government coordination to establish a mechanism capable of slowing down development if necessary to control self-improving models. This debate occurs in a context of a lack of specific federal regulation for AI in the United States.

The Donald Trump administration adopted a more flexible approach, prioritizing the economic value of the technology. However, critics argue that this environment could increase the risk of major cyberattacks and other damages linked to the accelerated advancement of AI. Senator Bernie Sanders of Vermont (USA) and Congressman Greg Casar of Texas (USA) are among the few legislators proposing stricter limits on the technology. Last week, both introduced a bill aimed at the permanent prohibition of superintelligence and a halt in model development until an industry regulatory body establishes new guidelines.

Coxon's departure coincides with Anthropic's preparations for an Initial Public Offering (IPO), which could be one of the largest ever. The company aims for a valuation of US$ 2 trillion (approximately R$ 10.4 trillion) and has highlighted responsible AI development to attract investors. Amodei and other company executives have also disagreed with the Trump administration and other industry leaders on certain occasions due to practices that, according to them, do not give sufficient priority to AI safety.

According to Coxon, Anthropic still maintains a Slack channel used by its employees to debate the most advanced capabilities of its models. The researcher pointed out that the fact that discussions of this magnitude occur in a communication tool used by engineers evidences the great influence that AI companies have begun to exert on technological development. Coxon commented to The Wall Street Journal: 'It's kind of insane that this has to happen on the MacBooks of some engineers living in San Francisco, instead of a desert bunker, like the one where they worked during the Manhattan Project.'

Popular