Anthropic integrates external evaluators to test the safety of its AI models
Read more
Olhar Digital
olhardigital.com.br

Anthropic integrates external evaluators to test the safety of its AI models

Anthropic has implemented a concrete measure to begin applying the proposal of its CEO, Dario Amodei, to slow down the development pace of more advanced artificial intelligence (AI) systems by incorporating external evaluators into the company.

Accenture was chosen as the first embedded external evaluator. This collaboration will allow specialists from Faculty, Accenture's AI division, to work internally at Anthropic to examine the company's safety mechanisms. Activities include safeguard testing, red teaming exercises, and analyses aimed at confirming that the AI models adhere to human values.

Anthropic and Accenture have established a minimum investment of US$ 1 billion (approximately R$ 5.3 billion) to develop this capability over the next five years. However, Anthropic stated that, given the urgency and relevance of the work, it will fund Accenture's operations itself.

In the long term, the company expects such activities to be funded by governmental or shared resources, as outlined in the Advanced AI Framework presented by Anthropic in June. Since these structures have not yet been established, the company plans to collaborate with various evaluators and funding models.

Amodei's suggestion proposed that independent evaluators have access to the employee level within AI companies. The purpose of this is to empower these specialists to directly verify safety practices and report any incidents.

When presenting his idea, Amodei assured that Anthropic would undertake this initial stage 'unilaterally' and encouraged other industry corporations to adopt a similar stance.

Other partnerships and responsibility

The collaboration with Accenture will not be exclusive; Anthropic informed that it is also in contact with the non-profit research organization METR and other potential entities to participate in this process.

Despite this, the company emphasized that the responsibility for the safety of its models remains entirely under its purview, and the presence of external evaluators does not diminish this duty.

Anthropic commented: "We are sharing these initial efforts now so that people and other AI developers can see our process." The company added that it intends to refine this methodology as the sector matures and will communicate new information as soon as the work begins and more evaluators are integrated.

With Accenture's participation, the company took a fundamental practical step to convert one of Amodei's central plan ideas into a functional structure: allowing an external entity direct access to audit its safety systems and detect potential failures before they escalate.

Similar stories

Anthropic blocks scientists for attempts to use AI in dangerous biological research
Read more
olhardigital.com.br

Anthropic blocks scientists for attempts to use AI in dangerous biological research

Anthropic reported that it blocked scientists who were using its Artificial Intelligence models in investigations that could aid in the development of biological weapons. Although the company could not determine if the projects had legitimate or malicious intent, the decision was to halt the work due to the risks involved.

One notable incident involved a researcher who requested assistance to modify the chikungunya virus, aiming to make it more aggressive. Anthropic detected that this research was being conducted at a military institute, which heightened concerns about the potential use of the results.

The central challenge for Anthropic lies in the ability to differentiate valid scientific investigation from a possible attempt to cause harm. While biological studies can drive medical advancements, such as vaccine development, similar knowledge can also be used to intensify the danger of infectious agents.

Jacob Klein, head of threat intelligence at Anthropic, commented to The New York Times that the situation is extremely complex because it is not an obvious scenario of biological weapon construction.

Due to this uncertainty, the company opted for a cautious stance. The researchers also attempted to bypass the control mechanisms designed to restrict access from certain regions and tried to disguise the purpose of their research to evade safeguards. The related accounts were suspended.

Key points identified by the company

Among the most relevant aspects pointed out by the company is a case from May, where a scientist asked Claude for help drafting a funding proposal to study the chikungunya virus, transmitted by mosquitoes and known to cause severe pain and other symptoms for months.

The project plan was to introduce genetic alterations into the virus to increase its virulence during successive infections in live animals. This type of study, called gain-of-function, has legitimate applications, including vaccine creation, but can also generate more dangerous variants of the virus.

The fact that the work was scheduled to take place at a military research center significantly influenced Anthropic's assessment. Klein stated that even though they did not know if the research would be turned into a weapon, the involvement of a military institution in gain-of-function research was a cause for concern.

Assessments conducted the previous year showed that Anthropic's older models were still not capable of providing substantial support in dangerous biological research. However, the company's current systems are better equipped to conduct complex scientific investigations, which motivated the implementation of stricter safety rules.

Read more:

Susan Monarez, a microbiologist and public health specialist, emphasized that these cases provide concrete evidence of clandestine efforts to use AI with the intention of increasing the chances of creating biological agents capable of causing considerable damage.

She also stressed that these same technological capabilities can accelerate medical discoveries. The great challenge lies in discerning when an investigation that appears correct may actually hide a harmful intent.

Anthropic also reported other inappropriate uses of Claude, covering surveillance, propaganda dissemination, and the development of conventional weaponry. However, the biological cases are considered the most alarming due to the difficulty in separating honest scientific research from preparation for causing harm.

The company decided to block the identified cases even without obtaining proof that the researchers intended to manufacture weapons. This episode illustrates how more sophisticated AI models complicate the assessment of hidden risks behind research that, at first glance, seems legitimate.

Popular