Anthropic is recruiting professionals with salaries that can reach up to $455k annually to help test the safety limits of its artificial intelligence (AI) systems. The company maintains a specialized team for security and abuse prevention, covering biological, chemical, explosive, cybernetic, fraud, scam, as well as radiological and nuclear risks.
The available positions involve specialists responsible for testing the systems, investigating threats, and analyzing risks in fields such as chemistry, biology, cybersecurity, and nuclear and explosive weaponry. The purpose of these roles goes beyond simply asking Claude to generate a prohibited response to check its obedience; Anthropic tests its models in various ways to identify if third parties could bypass its safeguards and use the AI to cause harm, aiming to detect flaws before they are exploited.
Anthropic operates a team called the Frontier Red Team, whose function is to examine AI systems to understand their capabilities and anticipate risks as the models advance. This name derives from the concept of red teaming, a security methodology where a team assumes the role of a potential attacker to locate vulnerabilities in a system. In the context of AI, this means seeking methods to circumvent security barriers and determining if the model can be used to generate harm.
Areas of focus include chemical, biological, radiological, and nuclear threats, in addition to cybersecurity and risks related to system autonomy. The company integrates tests conducted by researchers, specialists from various fields, and automated systems. Anthropic emphasizes that collaboration with external experts aids both in testing the models and in developing new evaluation methods. Depending on the risk analyzed, professionals may work with public versions of Claude or with non-commercial models subject to different security mechanisms.
Experts' View on the Role
Gabriel Cunha, an AI specialist at Oracle, commented to Olhar Digital that the role consists of ensuring that the institutional capacity to identify risks and implement controls keeps pace with the progress of model capabilities. Specialists attack the system to verify whether the model offers a real advantage compared to what a person could achieve without AI assistance.
Anthropic itself details a process that begins with tests conducted by specialists and can be expanded by automated systems. Initially, professionals define the type of threat to be investigated and seek ways to induce the system to exhibit risky behaviors. As the problem is better understood, tests can be repeated in hundreds or thousands of distinct scenarios.
The company also uses AI models to support this task. One system can generate attack attempts to test another model, and the results obtained aid in developing new protections. This cycle is repeated to confirm that the barriers remain effective even after new attack tactics are discovered.
A fundamental reason for Anthropic involving specialists from such diverse fields is that an inadequate response alone does not prove that the model has the capacity to cause real-world harm. The specialist must judge whether what the system produced is correct and whether it can be applied in practice.
Rodrigo Gava, CTO of Vultus and cybersecurity specialist, emphasizes: 'Getting an AI to respond to something prohibited does not demonstrate a real risk on its own.' He adds that the focus should be on evaluating whether the response is technically accurate, operationally viable, and whether it significantly reduces the knowledge, time, or cost required to cause harm. In cybersecurity, he exemplifies: 'A code might look sophisticated and not work, while a sequence of seemingly harmless information can complete a real chain of exploitation.'
Anthropic adopts this same logic in its security tests. Specialists first determine which threat they want to investigate, what kind of information could be used to cause harm, and what would be needed for the model to effectively assist someone in that action. In a biological risk assessment, for example, over 150 hours were dedicated to tests with biosafety specialists.
The same occurred in the nuclear risk assessment. In collaboration with the U.S. National Nuclear Security Administration (NNSA), agency specialists spent a year testing Claude models in a controlled environment. This work allowed Anthropic to create a tool capable of identifying conversations potentially related to nuclear weapons development. In preliminary tests, the tool identified 94.8% of queries related to nuclear weapons and registered no false positives in the evaluated set, achieving an overall accuracy of 96.2%. According to Anthropic, the system was subsequently partially integrated into Claude's traffic to aid in detecting misuse.
According to Gava, 'the process normally starts with a threat model: who is the potential attacker, what is their objective, what resources do they have, and what would be considered success.' From this, scenarios are developed to try to provoke behaviors that signal a vulnerability.
Tests are not limited to a direct command to the model. Professionals can examine multi-stage conversations, reformulations of the same request, changes in language and context, use of tools, autonomous agents, and combinations of techniques. The intention is to find situations where an isolated protection seems functional but can be bypassed when various resources are combined.
After a problem is identified by specialists, Anthropic can convert this situation into a test that will be replicated in hundreds or thousands of variations. Thus, the company verifies if the flaw persists and if the modifications applied to the model truly solved the issue. AI is also used in this process: one model can generate attack attempts to test another system, and the results are used to develop new forms of defense. The procedure is repeated to ensure that the barriers remain operational.
The need to test the models increases as they become capable of finding ways to exploit systems. In May 2026, Anthropic declared that Claude Mythos Preview was capable of discovering vulnerabilities, transforming them into exploitation vectors, and combining distinct steps to execute full attacks.
In another study, Anthropic evaluated Claude Opus 4 and Claude Sonnet 4 in cybersecurity challenges ranging from individual exercises to complex network simulations. The company reported that the models showed progress both in identifying flaws and in executing multi-phase attacks, although they still faced difficulties in sustaining extensive plans when faced with unforeseen events.
Advancement is also notable in the search for novel flaws, known as zero-days. In February, Anthropic reported that Claude Opus 4.6 had substantially improved its ability to find critical vulnerabilities on a large scale compared to previous versions. The company argues that such capability can be used to strengthen system defenses.
For Gabriel Cunha, such occurrences help justify why corporations must test not only singular responses but also the models' ability to integrate different resources to achieve a goal. The more tasks a system executes autonomously, the greater the need to evaluate what can happen when these competencies are used together.
Anthropic itself maintains the Frontier Red Team to track this evolution. In 2026, the group released analyses on vulnerability exploitation, cyberattacks, and threats associated with the use of AI by malicious actors. The stated goal is to understand the current capabilities of the models and predict what may emerge with their evolution.
Despite all these efforts, tests do not guarantee the absolute security of an AI. There is a virtually infinite number of possible combinations between commands, context, tools, and model behaviors. Additionally, new features or versions can introduce vulnerabilities nonexistent in previous evaluations.
The primary limitation lies in the fact that no red team can exhaustively cover all scenarios that would definitively attest to the AI's security. Models possess a probabilistic nature, the interaction space is almost infinite, and attackers continuously adapt their tactics.
Rodrigo Gava, CTO of Vultus, points out that Anthropic itself recognizes these limitations. In an analysis of AI security, the company states that testing for national security risks is still more 'art than science,' given that there is no totally standardized method for provoking concerning behaviors in the models.
There is also a balance between two types of errors: a false negative occurs when an attempted abuse is not detected and passes through protection mechanisms. A false positive can result in the blocking of legitimate activity because it was erroneously classified as dangerous.
For this reason, Anthropic implements multiple protection mechanisms. These include security testing, reward programs for those who discover flaws, monitoring, and automated systems designed to identify attempts to bypass model barriers.
In the Frontier Safety Roadmap, the company set a goal to develop, by January 2027, a system capable of finding new ways to circumvent protections against chemical and biological weapons on a large scale. Anthropic also plans to create tools to automatically investigate complex cyberattacks involving Claude.
The open positions at Anthropic illustrate how this structure operates in practice. The company seeks professionals for various roles, such as specialists in biological safety, chemical and explosive risks, cybersecurity, fraud and scams, as well as radiological and nuclear risks.
The roles are distinct. One is Red Team Engineer, responsible for trying to find flaws in the systems before they are exploited. This position requires experience in AI security and security testing, along with knowledge of how to bypass model barriers. The announced salary range is between $320k and $405k annually.
Another position is Threat Intelligence Manager related to chemical, biological, radiological, nuclear, and explosive risks. The professional will lead a team dedicated to detecting and investigating attempts to use Anthropic's systems for these purposes, with an annual salary of $455k.
There are also roles that require direct experience in the field of risk. For example, a position focused on biological safety requires a minimum of eight years of practical experience in life sciences, in addition to knowledge of molecular biology, drug discovery, or computational biology. This diversity of roles demonstrates that Anthropic's strategy transcends the simple refusal of Claude to certain requests; it aims to hire professionals to discover vulnerabilities, assess risks, investigate potential abuses, and transform these discoveries into new forms of protection, recognizing that the increase in model capabilities demands an expansion in the scale of this work.