Anthropic announced on Thursday, the 17th, three unprecedented indicators designed to track the speed at which artificial intelligence (AI) is evolving. This initiative comes shortly after Dario Amodei, the company's CEO, advocated that the largest corporations in the sector should coordinate a moderation in technological advancement.
The company also revealed that its chatbot Claude is already responsible for 26% of the internal AI Research and Development (R&D) work. According to Anthropic, this advancement helps measure the race for systems that progressively contribute to their own evolution.
Despite this increase in Claude's participation, Anthropic clarifies that the model is not yet acting completely autonomously in any part of its measured R&D. The term 'leading' in this context means that Claude executes the majority of a complete task based on a high-level directive, while a human remains responsible for overall supervision and guidance. This index was zero in February.
Additionally, Anthropic reported that AI is already involved in at least large segments of work under close human supervision in over 90% of its research, which includes the 26% where Claude is considered leading.
To calculate this index, Anthropic used a version of Claude itself to evaluate the company's R&D. The system was applied to an automation evaluation scale created by Epoch AI, and the results were subsequently validated by humans. The company also developed a chart to track Claude's level of automation since August 2025, proposing that other cutting-edge model developers can replicate the method using their own data and external validation.
The second metric presented by Anthropic focuses on the supervision of AI agents. The company developed a system to monitor and intervene in the actions of these agents and observed that about 30 thousand agents were simultaneously performing engineering and research tasks on its most used internal platform. The company also uses various internal autonomous monitors capable of detecting suspicious activities and forwarding them for human analysis. Anthropic's intention is to measure not only the volume of active agents but also aspects such as the proportion of monitored activities, the time required for action review, and the frequency with which agent behaviors are flagged.
Read more:
The third metric addresses the amount of computing power employed by the organization. By analyzing a snapshot of computing usage between July 13 and 20, Anthropic identified that approximately 6% of the computational capacity dedicated to AI R&D was allocated to security tasks. When specifically considering computation used in 'AI-driven' R&D, this portion dedicated to security reached about 12%. For Anthropic, such data helps illustrate how resources are being distributed as companies advance in developing increasingly sophisticated models.
Anthropic maintains that these three metrics should complement conventional assessments of model capabilities. While capability tests demonstrate what models are capable of doing, the new metrics aim to show how they are being built and the pace of that process. The company argues that such information can provide external researchers, governments, and the public with a starting point to assess the speed of development in major AI labs.
Anthropic stated that any cutting-edge developer could publish these measures regularly, following an open methodology. The company justified the disclosure of the data by claiming it provides greater transparency to the public, third parties, and governments regarding the pace of AI development in vanguard labs. This disclosure comes shortly after Amodei advocated for a coordinated slowdown in AI development.
The executive argued that leading companies should slow down to allow more time for creating safety mechanisms. This suggestion received support from industry leaders, including Sam Altman, CEO of OpenAI; Elon Musk, CEO of SpaceX; and Demis Hassabis, president of Google DeepMind. Amodei stated that his plan aims to control the speed of model capability evolution without compromising commercial advantage or U.S. leadership in AI.
Concurrently, the debate on safety gained momentum after warnings from experts and executives about the risk of AI systems becoming incomprehensible or uncontrollable. The previous week, a former Anthropic employee left the company stating that it was 'betting with our lives.' The discussion also reached Washington (USA), where Senator Bernie Sanders participated in an event about the dangers of AI escaping human control. Anthropic and competitors, such as OpenAI, have faced increased scrutiny regarding the assurance that their internal agents will behave as expected during tests and tasks. On Wednesday (16th), OpenAI released new examples of AI agents that failed tasks or deviated from the established script, presenting a framework for communicating such incidents. In this context, Anthropic points out that its new metrics can increase the visibility of technological advancement, especially with systems like Claude taking on an increasing share of next-generation AI development work.
