Initially starting as fan art, a meme featuring DeepSeek founder Liang Wenfen in the guise of a 'cool guy' in a blue suit quickly gained significant resonance on Chinese social media. This meme served as a catalyst for the release of DeepSeek V4 Flash.
The DeepSeek V4 Flash model achieved a score of 7.1 trillion tokens in the weekly global OpenRouter ranking and instantly became the most used model globally. As a result, DeepSeek secured two spots in the world's top five, with four positions held by Chinese models, while GPT-5.6 Luna dropped to sixth place.
Economic indicators show a significant advantage. One day before the release of V4 Flash, OpenAI reduced the price of GPT-5.6 Luna by 80 percent, setting the cost at $0.2 per million input tokens and $1.2 per million output tokens. DeepSeek V4 Flash was released at a price equivalent to approximately $0.14 for input and $0.28 for output, while offering a 98 percent caching discount compared to the industry standard of 90 percent. The independent evaluator Artificial Analysis determined the cost of V4 Flash for a typical agent task to be about 60 percent lower than that of Luna. OpenCode recorded a caching rate of 96 percent, which lowered the average session cost to $0.06.
The performance gap is no longer substantial. DeepSeek's internal tests show that V4 Flash surpasses V4 Pro Preview and GLM 5.2, only trailing Opus-4.8. Furthermore, Artificial Analysis assigned it an intelligence index score of 50, whereas Luna received 51. On the VulcanBench benchmark, Bold Metrics CTO Morgan Linton characterized V4 Flash as 'exceptionally stable under moderate load,' noting its superiority over Claude Fable and GPT-4.5. Approximately half of the traffic directed to DeepSeek on OpenRouter now comes from developers in the US and Europe, indicating that this phenomenon extends beyond the domestic market.
The focus on agent development is deliberate. V4 Flash utilizes the V4 Flash Preview architecture and focuses on agent capabilities after training. Team member Cui Tianyi publicly hired engineers for a new group, Harness—a system that wraps the model with context, tools, task state, and feedback, using open GitHub projects as a selection filter. In May, the DeepSeek-TUI agent, developed by independent American developer Hunter Boone, gathered over 11 thousand stars on GitHub in just a few days. Boone's open letter, titled 'whale bros, I'm the American who builds DeepSeek-TUI,' became a cultural event that swept across the Pacific. Since international developers support these models through their choices, China's open LLMs are no longer just a curiosity; they have become the baseline price-to-performance ratio that all others must now surpass.