DeepSeek releases first native multimodal model V4-Flash-Vision-Exp with open source
Read more
Pandaily
pandaily.com

DeepSeek releases first native multimodal model V4-Flash-Vision-Exp with open source

DeepSeek has introduced its first native multimodal model. The DeepSeek-V4-Flash-Vision-Exp model was made available on the Hugging Face platform on August 31st under the MIT license, marking the initial support for image input in the V4 series alongside text.

The release includes model files, a tokenizer, a minimal implementation for prompt encoding, and a minimal inference implementation in PyTorch. This implementation covers the vision encoder, aligner, DFlash attention, mixture-of-experts routing layer, Hyper-Connections modules, and DSpark.

DeepSeek noted that the model weights are an experimental build, 'Exp'; previously, the model was available via the DeepSeek API starting August 21st. Thanks to its visual capabilities, the model can describe images, read text from screenshots, and analyze charts, accepting inputs in JPEG, PNG, GIF, and WebP formats.

In tasks based solely on text—such as reasoning, agents, and world knowledge—the company asserts that performance matches the V4-Flash version, retaining its previous strengths. However, in agent benchmarks requiring visual understanding, the model significantly surpasses V4-Flash, bringing the capabilities of a multimodal agent closer to the level of Claude Opus 4.8.

By opening up the inference stack and the weights themselves, DeepSeek positions this model for integration into various frameworks and agent tools, expanding its strategy of open weights from text tasks to vision-supported tasks. For developers and enterprises, this release lowers the barrier to entry for locally deploying a competitive model with vision capabilities or using standard infrastructure, aligning with the industry-wide trend toward cost-effective and customizable open models.

Models capable of handling visual data have become key elements in the AI agent market, as developers combine screenshots and real images with reasoning functions. DeepSeek's move follows the trend of increasing open multimodal releases from Chinese labs, which use permissive licenses and aggressive pricing to gain popularity on US developer platforms, cloud service catalogs, and corporate model menus.

Since the model is relatively lightweight for its weight class and comes with a minimal inference reference, teams can quickly integrate it without needing proprietary orchestration tools. The company stated that it will continue to improve the V4 lineup, using the vision-enabled experimental build as a means to test architectural decisions—aligner components, DSpark, and Hyper-Connections—before a broader multimodal release. Currently, the experimental flag indicates the need for caution: chart recognition and understanding are strong, but DeepSeek advises testing the model on specific workloads rather than assuming production readiness across all image tasks. Early users, including research groups and application developers, are using the open stack to assess how well the multimodal agent's Opus-level skills hold up in their own processes.

Similar stories

Local Model StartLux with 27 Billion Parameters Surpasses DeepSeek V4 Flash in Chinese AI Benchmark
Read more
pandaily.com

Local Model StartLux with 27 Billion Parameters Surpasses DeepSeek V4 Flash in Chinese AI Benchmark

According to a report prepared by the China Academy of Information and Communications Technology (CAICT), a new player has emerged on the map of local models. The company StartLux, based in Shanghai and having developed the StartLux-V1.0-27B-Preview model, secured second place overall in the specialized MCP test within the lineup of verified AI benchmarks, surpassing DeepSeek-V4-Flash using only 27 billion parameters.

The MCP test evaluates six specialized tasks: location navigation, internet search, browser automation, financial analysis, code repository management, and 3D design, and also includes a comprehensive assessment focusing on coordinating multiple tools, performing complex tasks, and interacting in real-world scenarios. Among the tested models were DeepSeek-V4-Pro (1.6 trillion parameters), DeepSeek-V4-Flash-0731 (284 billion), Step-3.7-Flash (198 billion), StartLux-27B-260715 (27 billion), Qwen-3.6-27B (27 billion), and AgentCPM-Explore (4 billion).

The StartLux-V1.0-27B-Preview model achieved a score of 39.25, earning it second place, overtaking both DeepSeek-V4-Flash with 284 billion parameters and Step-3.7-Flash with 198 billion. With the same parameter size, it outperformed Qwen-3.6-27B by 5.34 points. The model ranked first in location navigation and also achieved first place or tied for first in browser automation and financial analysis, sometimes reaching the level of the trillion-parameter DeepSeek-V4-Pro model.

This model is built upon Qwen3.6-27B with additional post-training enhancements. StartLux implemented an approach they termed 'AI trains AI' (Automatic Search), which allows for autonomous experimental training and strategy refinement through feedback. The company claims this is the first application of this method for a local agent model in China.

This result reflects a broader shift in the industry. As the performance of cutting-edge models reaches practical thresholds, the era of the parameter race is transforming into a phase of homogenization; user priorities are increasingly shifting towards solving real problems, ensuring data security, and controlling costs, rather than solely achieving high benchmark scores. Global players are moving in the same direction: Google's Gemma 4, Meta's open Muse Glimmer, and Nvidia's Nemotron 3.5 Lightning are targeting local deployment.

StartLux-V1.0-27B-Preview can run on consumer PCs, and the company plans to release its first generation of local intelligence solutions this year. The CAICT results signal to enterprises that compact, locally deployable models are now capable of competing with much larger cloud solutions in tasks that matter in real workflows, beginning to change procurement decisions.

DeepSeek launches AI with computer vision capable of performing tasks autonomously
Read more
olhardigital.com.br

DeepSeek launches AI with computer vision capable of performing tasks autonomously

DeepSeek unveiled DeepSeek-V4-Flash-Vision-Exp this Friday, the 21st. This new experimental artificial intelligence model is not limited to processing text; it also has the capability to read, interpret, and understand images.

According to the Chinese company, the performance of this new model approaches that of Opus 4.8, developed by the American competitor Anthropic. DeepSeek's AI can now analyze photographs, screenshots, and documents to perform activities without the user needing to provide detailed instructions for each step.

Developers can integrate this functionality into their own projects. The model retains all the textual functionalities of its previous version, DeepSeek-V4-Flash, including logical reasoning, general knowledge, and the ability to execute automated routines. The main distinction is that the system now processes visual and textual data simultaneously.

In performance evaluations focused on autonomous tasks, known as agents, the Chinese model achieved results comparable to Anthropic's Opus 4.8. This feature allows the tool to observe an image, understand the context presented on the screen, and make decisions on its own to complete the requested work.

Programmers have the option to test this novelty through the DeepSeek platform by entering the name of the new model in the access code. Images can be sent to the system via internet links, file codes, or using the company's new submission mechanism.

In addition to the visual model, DeepSeek made available a free file management tool called Files API. This feature allows the user to upload an image to the servers only once and reuse it in several subsequent API calls, optimizing internet consumption and speeding up responses.

This launch intensifies the competition between technology companies from China and the United States. The strategy adopted by Chinese companies focuses on offering solutions with lower operational costs while maintaining a competitive capacity close to higher-cost American models.

MiniMax releases MiniMax Design, integrating the H3 video model into the comprehensive workflow
Read more
pandaily.com

MiniMax releases MiniMax Design, integrating the H3 video model into the comprehensive workflow

Although launching the video model in one day with Seedance 2.5 might have led to its displacement, MiniMax's H3 model was an exception. After it was released in early August, users demonstrated enormous interest, particularly using the Chinese Wonderland series. Users generated high-quality reference images using Midjourney and then uploaded them to H3. Users noted that videos created by H3 feature advanced visual taste, smooth motion, and character stability.

MiniMax has now embodied this model functionality into a ready-made product—MiniMax Design. Similar to how GPT-5.6 Sol plus Codex or DeepSeek V4 plus DeepSeek Harness unlock the true agentic potential of a model, MiniMax Design integrates H3 into a workflow that supports continuous modification, collaboration, and final project delivery.

The interface looks professional, offering basic elements in the style of ComfyUI and a full canvas, but the main work is performed by a team of agents. Users do not need to understand the canvas, camera positions in the 3D director's cabin, or ComfyUI connection schemes; they simply describe the desired outcome, and MiniMax Design handles the entire production chain, utilizing skills, models, and agents to parse and execute the request. Its main task is to structure professional capabilities into executable nodes that activate H3, image, music, and voice models, bringing AI video creation closer to full production.

Instead of writing text prompts for AI video and copying them from ChatGPT into Hailuo or other web tools, users can switch to MiniMax Design: a simple requirement is enough, and the system itself creates the canvas, generates scripts, storyboards, images, music, dialogues, editing, subtitles, and the video itself. Authors reproduced a scene from the film Niu Coming by inputting a prompt to stage the scene in a director's cabin, where the heroine and hero walk in a meadow, meet by a wide river, she jumps across it, and he walks along the bank. Creating this scene was complex but only required uploading materials and communication.

The 3D director's cabin allows users or agents to place character models and control camera positions and poses in three-dimensional space, creating references for storyboarding and reducing trial-and-error costs. The complexity lay in the spatial relationships between characters; MiniMax Design first generated a camera movement video from the cabin and then created the precise final video using this reference. The value of the cabin lies in transforming vague requirements into specific spatial and cinematic benchmarks: if the action is incorrect, users adjust the characters and cameras, after which they regenerate the frame. In addition to such complex scenes, MiniMax Design easily handles various spectacular videos.

Popular