Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1 with improved features and reduced restrictions
Read more
Olhar Digital
olhardigital.com.br

Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1 with improved features and reduced restrictions

Anthropic introduced new versions of its advanced line of artificial intelligence models—Claude Fable 5.1 and Claude Mythos 5.1—this Tuesday (the 1st). These updates combine increased performance with changes that reduce usage costs and the frequency of Fable's safety mechanisms triggering.

The Fable 5.1 model is intended for the general public and is available through the Anthropic API as well as on cloud computing platforms. Meanwhile, Mythos 5.1, which uses the same base model, remains accessible only to selected Anthropic partners involved in cybersecurity and life sciences research.

New Performance Achievements

The company claims that the new models set a new standard for performance in tasks related to programming and knowledge processing. According to tests published by Anthropic itself, Fable 5.1 achieved a score of 52.6% in the Terminal-Bench-Science 0.1 test, designed for AI agent-performed scientific research, which is significantly higher than the score of Fable 5, which was 24.7%.

In the Terminal-Bench 4.0 test, focused on terminal-based programming, Fable 5.1 reached 55.8%, compared to 42% for Fable 5. Mythos 5.1 demonstrated an even higher result—60.9% in the same test. Furthermore, the new model scored 77.9% in the partial mode of the OSWorld 2.0 test and 41.7% in the stricter version of this test. In the Humanity’s Last Exam, Fable 5.1 scored 60.9% without using tools and 65% when using them.

Improved Safety Mechanisms and Privacy

A significant change involves the safety mechanisms. Previously, Fable 5 faced criticism for blocking legitimate requests when systems detected a potential link to high-risk areas, especially in cybersecurity. Anthropic states that these protective functions have been enhanced in Fable 5.1 to minimize false positives.

The company explains that Fable 5.1 and Mythos 5.1 are essentially the same model, and some differences in previous tests were due to the intervention of safety systems. Anthropic points out that when these safeguards intervene in certain cybersecurity tests, the tasks are redirected to Claude Opus 4.8, and biological ones to Claude Opus 5, which could lower the registered performance of Fable in some evaluations.

Additionally, there has been a major modification to the privacy policy: Anthropic has started offering a Zero Data Retention option. This allows clients to use the models within their own infrastructure, guaranteeing that data does not leave those environments. Although the system will continue to monitor for potential misuse by users or agents, clients gain greater control over how this monitoring is conducted. Anthropic also emphasized that it has never trained its models on corporate data without explicit permission and does not plan to do so in the future.

Results from Early Users

Beyond benchmarks, Anthropic disclosed results obtained by companies that received early access to Fable 5.1. For instance, the investment firm Millennium reported that the model was able to detect the cause of an extremely rare failure in their internal systems that had remained unexplained to engineers for years. The model analyzed an external library, compared the code with a kernel dump, and identified the problem in that library as the source of the failure.

MongoDB noted that Fable 5.1 is capable of investigating the code and documentation of its services, developing complex prototypes, and operating for hours without supervision, using self-verification loops to assess its work. Anthropic also stated that the models demonstrated scientific results even before the official release, including creating a high-resolution map of Venus based on existing photographs and individual GPU optimization.

Despite the progress, Anthropic's internal security documentation notes that Mythos 5.1 shows a slight decline compared to Opus 5 in certain divergence scenarios. According to the document, the model more easily accepts requests related to misuse and unverifiable authorization claims, although it is less prone to ignoring explicit constraints, fabricating inputs, or falsely claiming task completion compared to previous models.

Popular