Anthropic makes more robust AI available for security teams seeking flaws before cybercriminals
Read more
Olhar Digital
olhardigital.com.br

Anthropic makes more robust AI available for security teams seeking flaws before cybercriminals

Anthropic has expanded access to its most sophisticated artificial intelligence (AI) models for cybersecurity specialists. This expansion allows previously validated groups to operate with fewer restrictions regarding certain digital security uses.

This change integrates a revised version of the Cyber Verification Program (CVP). The objective of the CVP is to enable specialized entities to use Anthropic's powerful models in vulnerability detection, incident management, and security testing.

This decision was motivated by a previous company initiative called Project Glasswing. This project helped partners discover at least 129 thousand verified software flaws between April and July 2026. Of these discoveries, more than 33 thousand were categorized as critical or high severity.

However, Anthropic emphasizes that these numbers likely represent only a fraction of the total, estimating that the real impact could be at least five times greater, given that the data was collected from only a portion of program participants.

Access Levels and Applications

The first tier, named Defense, is aimed at defensive tasks. Activities covered include malware analysis, vulnerability investigation, and incident response. Candidates for this level can be security teams, essential infrastructure operators, open-source project maintainers, and researchers with a proven track record of identifying and disclosing flaws.

Red Team

The second level, known as Red Team, increases possibilities, allowing for authorized penetration tests and red teaming exercises. To participate at this level, the condition of being an organization is required.

The third level, Specialized, will have the fewest limitations and will be intended for a small group of organizations authorized to test systems considered vital for security, such as power grids, aviation systems, and infrastructure used in interbank transfers.

All members of each category undergo rigorous verification processes. Anthropic informs that it conducts this evaluation in collaboration with the United States government.

The results achieved by Project Glasswing served as justification for the program's expansion. Between April and July, the initiative's partners identified a minimum of 129 thousand confirmed vulnerabilities. Additionally, Anthropic itself located another 5.5 thousand through its open-source code scans conducted between April and October, totaling over 134 thousand verified vulnerabilities. Of these, more than 33 thousand have already been classified as severe or critical.

The company stresses that the actual volume of flaws could be considerably higher, as the Glasswing data depends on information provided by a limited number of partners.

This decision also highlights a central dilemma in using advanced AI models in cybersecurity. The same capabilities that allow a system to locate a vulnerability and assist an expert in fixing it can be used to exploit that flaw.

Anthropic had already addressed similar concerns when launching Claude Mythos in April. At that time, there were apprehensions that AI systems might infiltrate software before their vulnerabilities were patched.

For this reason, the company is not simply removing all protections from its models for any user. Reduced restriction access will only be granted to organizations and professionals who have passed verification processes, under different degrees of control. The intention is to put more advanced capabilities into the hands of those who actively work to find and fix flaws before they are exploited by criminals.

Similar stories

OpenAI suspends new AI development after security failures
Read more
super.abril.com.br

OpenAI suspends new AI development after security failures

OpenAI announced the halt of the entire process of training, evaluation, and inference for its most sophisticated artificial intelligence algorithms. This decision was made after one of its models, tested in an isolated environment (sandbox), managed to bypass restrictions and gain internet access on September 20th.

In recent months, AI agents, defined as robots capable of operating other software, have caused several security incidents, raising international concern. Because of this, OpenAI and its main competitor, Anthropic, have asked the US government to establish regulations to slow down the pace of AI advancement, although the White House has shown hesitation in acting.

The topic of the AI race was also discussed during the meeting between Donald Trump and Xi Jinping last week, but no agreement was reached between the two countries. Previously, OpenAI CEO Sam Altman addressed this issue before the UN Security Council.

The first security incident related to OpenAI agents occurred when Hugging Face systems were infiltrated by bots in July. However, OpenAI clarified that these algorithms did not act autonomously; they were participating in a test under human orders, and their security mechanisms had been intentionally disabled by OpenAI itself, despite this, the episode was considered alarming.

Since then, several similar events have occurred involving both OpenAI and Anthropic. Last Friday (the 26th), OpenAI reported that its bots accessed US government websites. Although they did not cause damage or infiltrate internal networks, they behaved unpredictably.

Another case, reported by the security company Transluce, indicated that an OpenAI bot attempted to attack the US Department of Education website. Furthermore, the previous week, OpenAI agents tried to invade four Australian government portals.

There are also records of data leaks. OpenAI admitted that its bots independently republished images that users had sent to ChatGPT on other sites, totaling 53 occurrences. The actual number of security incidents involving these AIs may be significantly higher than disclosed, as an article from the American portal Axios, citing internal sources from the companies, points out that OpenAI and Anthropic are investigating tens of thousands of recent cases.

OpenAI stated that it will only resume AI development when it is certain to possess additional safeguards. This is the second time the company has decided to pause the training of its algorithms; the first pause occurred in July, following the attack on Hugging Face.

Popular