Anthropic launches Claude Opus 5.5, a model focused on programming, complex tasks, and computer usage
Read more
Olhar Digital
olhardigital.com.br

Anthropic launches Claude Opus 5.5, a model focused on programming, complex tasks, and computer usage

Anthropic announced Claude Opus 5.5 on Tuesday (22), which is the first model in the Claude 5.5 series. The company claims that this new version improves performance in programming activities, computer manipulation, and knowledge tasks, while also offering reduced costs and faster response times.

Furthermore, the model has been equipped with more rigorous security mechanisms. These protections include defenses aimed at cybersecurity and biology, barriers against prompt injection attacks, and measures designed to reduce attempts to bypass the limits established in testing environments.

During tests conducted by Anthropic itself, Opus 5.5 achieved the company's best result to date in its main automated alignment test. The company also observed notable progress in programming, computer usage, and knowledge tasks.

One initial evaluator managed to complete a code migration of 680 thousand lines in less than a day. In another test scenario, the model demonstrated the ability to decrease the loading time of an application's pages in 39 out of 40 attempts made.

There was also a significant reduction in operational costs. According to Anthropic, Opus 5.5 requires 40% fewer resources to run in typical workloads compared to Opus 5, and it generates responses more than 30% faster.

Opus 5.5 represents Anthropic's first launch after the company advocated the thesis of 'slowing down the frontier' of AI development, maintaining safety practices aligned with the advancement of model capabilities.

In internal tests, the new model attempted to circumvent restrictions in controlled environments approximately 85% fewer times than Opus 5 or Claude Mythos 5.1. All detected attempts were classified as low severity and were reported by the model itself.

Anthropic also assures that Opus 5.5 has greater resistance to prompt injection attacks. For cybersecurity-related demands that activate protections, requests are redirected to Claude Opus 4.8, while requests about biology can be sent to Opus 5.

Performance in Programming and External Evaluations

Initial tests also highlighted the model's performance in programming. Mario Rodriguez, Product Director at GitHub, commented that the model used among the smallest volumes of tokens and steps recorded by the company and solved more terminal tasks than Opus 5 in VS Code, using less than half the steps.

In an evaluation conducted by Deloitte Consulting, the model also outperformed Opus 5 in certain tasks. Carl Bennett, CIO of Deloitte Consulting, stated: 'Even in its lowest effort configuration, Claude Opus 5.5 captured 72% of known errors in our code reviews, compared to 56% for Opus 5 at high effort, with fewer false alarms and a fraction of the output quantity.'

Anthropic also adjusted how the model drafts text. Opus 5.5 prioritizes placing the most relevant information at the beginning of responses, uses fewer technical jargon, and adheres better to user-provided writing guidelines. One initial evaluator summarized this change by stating: 'it writes like I write.'

Currently, Claude Opus 5.5 is available on Anthropic's platforms, as well as through Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can obtain it on the Claude Platform using the identifier claude-opus-5-5. Claude Sonnet 5.5 and Claude Haiku 5.5 will be available in the coming weeks.

Popular