The Qwen team at Alibaba has released Qwen3.8-Omni-Flash—a native multimodal model capable of simultaneously processing text, images, audio, and video using a one-million token context window. This model is available on the Qwen AI platform starting September 18th and is positioned not so much as a demonstration of captioning capabilities, but as a set of agents for long-form audio and video workflows, including task planning, tool calling, and providing ready-made media materials.
When tested on a set of about 30 public and internal benchmarks, Qwen3.8-Omni-Flash showed an average performance increase of over 26% compared to the previous version, Qwen3.5-Omni-Plus. The most significant improvements were observed in tasks related to agentic understanding of audio and video, as well as in long-horizon tasks: WildClawBench-MM increased by 36.5 points, AgenticVBench by 22.3 points, and UniClawBench reached 69.6.
Furthermore, basic perception improved: LongAudioSpan increased by 8.3 points, and OmniVideoBench increased by 9.6 points. Meanwhile, the diarization error rate in AliMeeting and the word concatenation error rate decreased from 88.11/89.61 to 3.35/17.18, respectively.
Practical applications include understanding long videos in an agent mode. Instead of analyzing every frame, the model independently determines which parts need to be viewed and listened to, concentrating tokens on the most relevant segments. In agent mode on OmniVideoBench, accuracy increased from 63.4 to 67.8, while token consumption decreased by approximately 45.7%, amounting to about 79,117 instead of 145,736.
Meeting workflows support the input of audio and video materials up to one hour long for speaker separation, transcription, minutes generation, and subsequent actions such as sending emails or coding via tools. To support these pipelines, Alibaba has expanded Qwen-MM-Plugins and opened Qwen-Live Harness for continuous real-time multimodal interaction. A special Realtime SKU provides low-latency streaming, accent-aware speech practice, and spatial audio cues that assess sound direction and distance. The company reported that API prices for hourly audio input were reduced by over 98%, and for hourly audio-video input by over 93%, presenting the release as a leap in capabilities and a more economical path for production multimodal agents.


