ByteDance releases updated Doubao-Seed-2.1-pro 0915 model with multimodal encoding on Volcengine
Read more
Pandaily
pandaily.com

ByteDance releases updated Doubao-Seed-2.1-pro 0915 model with multimodal encoding on Volcengine

ByteDance announced that the updated version of Doubao-Seed-2.1-pro, release 0915, is fully available via the Volcano Ark API. Concurrently, Doubao Work has been updated, and the TRAE system is now connected. The company emphasizes the implementation of multimodal encoding in production processes, enhancing agent reliability, and reducing token costs for processing images and videos, which, according to the company, is more than 30% lower than the previous generation.

Multimodal encoding is presented as a key feature of the product. ByteDance asserts that corporate tasks are often found not in formal documents but in screen recordings, UI sketches, and operator habits. In an official case study, Doubao-Seed-2.1-pro 0915 was able to analyze undocumented Java ERP, consisting of approximately 280,000 lines, using a recording and several sketches. The model successfully derived the architecture and created a working mobile frontend, shifting the focus from humans writing machine-readable specifications to models reading human visual expressions.

Another demonstration example used the open-source game repository Luanti, which included about 387,000 lines of code and 1,000 historical issues, with official fixes hidden. The model organized the work of several sub-agents over almost 36 hours to perform root cause analysis and correct files among themselves; approximately 83% of the completed tasks met a standard acceptable to engineers.

Separately, the model restored an interactive three-dimensional scene of a seasonal courtyard based on four design images, including structure, water, vegetation, and camera effects. Volcengine also highlights agent reliability and cost-effectiveness. Previously, Doubao-Seed-2.1-pro held leading positions in several benchmarks such as Terminal Bench 2.1, SWE-Pro, SciCode, OSWorld, MobileWorld, and MMMU-Pro. Update 0915 adds stronger evidence tracking, source authority checking, and verification for extensive research reports, along with improved document and engineering drawing parsing. The increased efficiency of tokens for image and video processing is positioned for use in high-frequency corporate APIs. This release is an update to the model and product on the Volcengine platform, not a funding announcement, and external teams can still reproduce the ERP and Luanti merger metrics on their own stacks.

Similar stories

DeepSeek launched V4.1-Flash, a new AI model that is faster and has reduced costs
Read more
olhardigital.com.br

DeepSeek launched V4.1-Flash, a new AI model that is faster and has reduced costs

DeepSeek, a Chinese company, has introduced V4.1-Flash, its latest artificial intelligence model. This new technology was presented as the smallest within its family of models and promises to offer more agile responses, greater processing capacity, and lower operational costs.

This launch comes as DeepSeek prepares for an Initial Public Offering (IPO) on the STAR market in Shanghai, as reported by Reuters. Furthermore, the company is implementing V4.1-Flash, replacing previous versions of its models.

Despite being the smallest model in the new line, V4.1-Flash has 552 billion parameters, which are the components used by the AI to process data and generate outputs. However, during each task, it uses only a fraction of these resources, specifically 8 billion parameters to receive information and 16 billion to formulate responses.

The practical objective for the user is to allow the model to operate with greater speed and handle a higher volume of requests without demanding excessive computational resources. DeepSeek also assures that the system was designed to facilitate the development of even larger models in the future.

The company highlighted that V4.1-Flash features 'native visual understanding,' allowing it to interpret visual content, such as images. Tests conducted by various entities indicated that this new model surpasses V4-Pro in terms of performance, cost, speed, and total execution time.

A significant change lies in how the model stores data during an operation, which historically could increase the costs of long-duration AI services. Compared to the previous generation, DeepSeek stated that V4.1-Flash requires fewer resources, which, according to the company, can drastically reduce expenses, particularly in AI agents and systems that perform tasks with little human intervention.

V4.1-Flash is already accessible through the DeepSeek API, which allows the integration of the model into other applications. The V4-Flash and V4-Flash-Vision-Exp models have been discontinued, and their calls are temporarily being redirected to the new system.

More Information:

V4-Pro will undergo a gradual withdrawal. Starting September 14, 2026, all requests directed to it will be automatically sent to V4.1-Flash, applying the charge corresponding to the new model. This process will continue until the launch of V4.1-Pro.

The new API rates will take effect on September 10. DeepSeek establishes different tariffs for peak and off-peak periods; outside of peak hours, the price will be half the maximum rate. The company emphasized: 'V4.1-Flash allows serving more users at a lower cost. We are passing this savings on to you.'

Additionally, DeepSeek plans to collaborate with the open-source community to expand support for the model and investigate other implementation methodologies. The report on the launch of the new model was initially published in Olhar Digital.

Popular