Unisound has introduced the U2-Flash model, positioning it as a full-fledged Flash-class working tool rather than a simplified version of its previous U2 lineup. The model utilizes a sparse mixture-of-experts design with a total parameter count of approximately 266 billion, while only about 10 billion parameters are active during inference, which accounts for less than 4% of the total weight volume.
The goal of this architecture is to simultaneously perform coding tasks, agent operations, mathematical reasoning, and instruction following at a single verified point. Unisound presents this release as a density enhancement strategy: achieving core task quality with low latency and token cost characteristic of the Flash class, optimized for production workflows rather than just demonstration chats.
Furthermore, the company emphasizes that over ten years of data from vertical scenarios were used to improve post-training quality, especially when models need to complete multi-step tasks. Two key development areas underpin the presentation. The first is the post-training loop, which allows the model to assist in generating tasks, analyzing trajectories, and verifying the training stack. This is Unisound's stated first step toward recursive self-improvement within defined human sandboxes and verification gateways.
The second direction is latent reasoning in the hidden state space, presented as a controlled interface of thought intensity rather than an opaque internal mode. Additional elements include asynchronous reinforcement learning of agents with parallel workers, multi-teacher online policy distillation between math, code, and agent specialists, and adaptive task generation that maintains complexity at the current capability level of the model.
Unisound reports that this cycle has already generated a set of nearly 100,000 software development tasks covering major programming languages. Moreover, multi-teacher distillation reduced the number of steps to a skill-appropriate level by approximately 55%.
According to published metrics, U2-Flash demonstrates a score of 64.6 on DeepSWE v1.1, 24.3 on TerminalBench 3.0, and 61.6 on SWE-Bench Pro. The latter score increased by 10.5 points compared to the previous generation, according to company materials, which also claim leadership over named competitor benchmarks in the Flash and Pro classes.
Regarding performance metrics, an average time to first token of less than three seconds, peak throughput up to 300 tokens per second, a reduction in agent iterations by 20–30%, and a shortening of the end-to-end task execution cycle by approximately 35% have been declared, with token savings comparable to U2.
The model is also optimized to work with major domestic accelerator stacks in terms of throughput, latency, and cluster scaling for deployment in government, financial, manufacturing, and energy sectors.
U2-Flash is available on Unisound's MaaS platform during a promotional period extending until September 30, 2026. This period includes token discounts and the provision of a trial volume of 100 million tokens for new and existing users. However, since the launch accompanying documentation did not include independent reproduction of the full benchmark or confirmation of hardware parity, buyers are advised to conduct workload-level validation before considering U2-Flash as standard production intelligence.
