The DeepSeek V4-Flash model became the leader of the weekly OpenRouter ranking, processing 7.1 trillion tokens. Additionally, nine Chinese models entered the top ten global leaders.
The DeepSeek V4-Flash model became the leader of the weekly OpenRouter ranking, processing 7.1 trillion tokens. Additionally, nine Chinese models entered the top ten global leaders.
According to OpenRouter data, DeepSeek V4-Flash took first place in model call volume just four days after the public API beta version launched on July 31. In one week, from July 28 to August 3, the model processed 7.1 trillion tokens. The second through fifth places were also occupied by Chinese products: Xiaomi MiMo-V2.5 with 5.4 trillion tokens, Tencent Hy3 with 4.89 trillion, DeepSeek V4-Pro with 3.1 trillion, and Zhipu GLM-5.2 with 2.87 trillion. The DeepSeek V4-Flash 0731 model independently rose to eighth place, processing 2.33 trillion tokens. On August 3 alone, version 0731 registered 1.02 trillion tokens, surpassing the previously dominant version 0423.
The total global weekly volume of AI model calls reached 56.8 trillion tokens. The weekly volume of calls for Chinese models was 28.13 trillion, exceeding US volumes for the fourteenth consecutive week. Chinese models hold nine out of the top ten spots. The official V4-Flash version uses the same architecture as the preliminary version: MoE, 284 billion shared parameters, of which 13 billion are activated, with a context window of 1 million tokens. Changes were only made to post-training: the official V4-Flash version achieved a score of 25.2 in the Agent Ultimate test, approaching the Claude Opus-4.8 score of 25.7, which is significantly higher than the preliminary V4-Pro version, which scored 15.8.
Pricing is almost unmatched: $0.14 per million input tokens and $0.28 per million output tokens. In weighted task testing, V4-Flash costs 3 cents per task, Kimi K3 costs 86 cents, GPT-5.6 Sol costs $1.86, and Claude Fable 5 costs $3.15. Operating costs are less than one percent of the flagship product Anthropic Claude Fable 5. On the social network X, developers wrote that Musk should personally test V4 Flash, comparing it to magic. Users then noticed that Elon Musk had indeed subscribed to the official DeepSeek account. Nearly half of DeepSeek's traffic comes from developers in the US and Europe, and many agent automation projects have switched their default models to DeepSeek. Founders of several foreign AI startups publicly noted that after fully switching to DeepSeek, monthly inference costs decreased by 60–85 percent.
Industry analysis indicates that DeepSeek has set a clear boundary in the global developer market. This boundary is not just a performance threshold but a balance between price and quality. Models offering a similar level of capabilities but at a significantly higher price will continuously lose mid-tier developers; products weaker than DeepSeek in performance but more expensive will face a rapid narrowing of their survival space. OpenAI also felt the pressure: GPT-5.6 Luna lowered its prices by 80 percent within three weeks of its release. From the public beta API on July 31 to entering the top global call volume on August 3, DeepSeek V4-Flash took less than a week.
The DeepSeek V4 Flash model demonstrated massive demand, consuming 8 trillion tokens in a single day on the OpenCode tool. This volume exceeded the average daily usage of the OpenRouter platform, which is around 6.6 trillion tokens. This surge in usage reflects a fundamental shift in AI interaction methods: where tasks used to be simple question-and-answer exchanges, agents now perform more complex operations such as reading files, modifying code, running programs, and retrying after errors, which significantly increases token consumption.
An example of such complex interaction was shown by a user who created a single-page 3D game project using Claude Opus 5. This process resulted in one HTML file that required 690 million tokens and cost nearly 3000 yuan.
With the advent of agent-driven tasks, the evaluation criterion for models has shifted from the quality of a single response to the cost of completing a long-running task. For instance, one million output tokens cost 2 yuan with DeepSeek, whereas the same operation with Claude Opus 4.8 cost $25 USD, equivalent to approximately 170 yuan, demonstrating an 85-fold price difference. Although the V4 Flash Max model scored 50 points compared to 56 for Opus Max in the Artificial Analysis Intelligence Index v4.1, the cost to pass this test was $72.02 for V4 Flash versus $3752.55 for Opus, with Opus gaining 6 extra points at roughly a 52-fold increase in call costs.
V4 Flash cannot yet prove absolute superiority, but if both models appear in the same candidate list, the 85-fold price difference could change the standard selection. When a model only slightly outperforms its competitor but costs dozens of times more, the mathematics of choice no longer favors the more expensive option.
The first to feel the impact of this trend are not the weakest models, as smaller models remain suitable for routing, and flagship models continue to handle the most complex tasks. Mid-tier models are the most vulnerable—they are stronger than V4 Flash but cost dozens of times more and are not powerful enough to solve the tasks that V4 Flash handles. This is DeepSeek's 'killer line': models do not necessarily need to be the strongest; they just need to be good enough to perform most tasks, leveraging a price gap of almost two orders of magnitude to displace more expensive options from the default list.
In terms of pricing, Opus costs $25 per million output tokens, Fable 5 costs $50, and Chinese models include Qwen 3.8 Max priced at 36 yuan, Kimi K3 at 100 yuan, and GLM 5.2 at 28 yuan.
The ability to sell Mutai at mineral water prices is due to the direct inference cost per token, which depends on activated parameters, accuracy, attention, KV cache, output length, hardware utilization, and parallelism. V4 Flash has 284 billion total parameters but activates about 13 billion per token by using hybrid attention, low-precision weights, and cache optimization. With a context of one million tokens, the inference computational load is about 10% of V3.2, and the KV cache size is about 7%. The V4 series overhauled the attention mechanism, implementing CSA compress-then-focus and HCA full-browse mechanisms, and utilizes DeepSeekMoE routing, directing each token to a small subgroup of experts. Commercial pricing and service costs follow these technical features. Thus, the 'killer line' redefines the boundary of development: the combination of performance and price becomes the decisive factor, where the cost of task execution is the determining metric.
DeepSeek, a Chinese artificial intelligence company based in Hangzhou, has announced plans to construct a large data center in Inner Mongolia. This mega-center is estimated to have a computational capacity of one gigawatt and aims to increase the company's infrastructure to compete in the global AI market.
The undertaking will be established in Ulanqab, a city located approximately 350 kilometers northwest of Beijing. The plan includes both the creation of its own structure and the acquisition of additional capacity from other corporations. DeepSeek aims to begin some operations by the end of 2027 or early 2028.
This initiative is part of the global competition for computing power dedicated to artificial intelligence, where both American and Chinese companies are increasing their investments in large-scale data centers. DeepSeek seeks to consolidate its position after gaining recognition for developing AI models that require fewer computational resources.
The construction of this complex in Inner Mongolia ranks among the largest computing infrastructure projects linked to a Chinese AI company. An installation with a 1 GW capacity would surpass all operational AI data centers in the country, although it remains below American projects reaching capacities of 3 GW and 5 GW.
DeepSeek's scheme allows for the combination of proprietary facilities with leased space in third-party managed structures. Sources familiar with the matter indicate that more than ten companies already have planned projects in Ulanqab, as compiled by an independent source based on local government approvals.
Although the specific chip to be used has not been specified, Nvidia semiconductors are considered the dominant standard for data centers in the AI sector, while Huawei stands out as the main Chinese manufacturer in this area.
The choice of Ulanqab was motivated by the region's climatic conditions. With an average annual temperature close to 4 degrees Celsius, the location reduces the energy demand required to cool high-consumption servers, an attractive factor for companies wishing to operate large computing structures.
DeepSeek's expansion comes after the company gained international visibility by presenting AI models that, according to the company itself, demonstrated superior performance using fewer computational resources compared to its competitors. The strategy also included the development of open-source and low-cost services.
This advancement has placed the startup at the center of the technological dispute between China and the United States. Giants like OpenAI and Anthropic have invested billions of dollars in computing capacity, while Chinese groups such as Moonshot, Z.AI, Alibaba, and MiniMax are also seeking to expand their share in the sector.
DeepSeek's move occurs in an environment of growing geopolitical scrutiny. US authorities have raised suspicions that the company may have circumvented export restrictions by using Nvidia's Blackwell processors in a facility in Inner Mongolia, according to sources close to the matter.
Additionally, OpenAI and Anthropic have accused the Chinese startup of practicing a method called distillation, where results generated by existing AI models would be used to enhance new technologies.
The company has also begun preparations for a potential Initial Public Offering (IPO), potentially filing the application this year, as previously reported by Bloomberg News. Previously, the company had raised $7 billion, reaching an approximate valuation of $50 billion.
The construction of a data center of this magnitude also demonstrates the rising financial capacity of Chinese companies to invest in AI infrastructure. A 1 GW complex equipped with advanced accelerators, according to an estimate cited by Nvidia CEO Jensen Huang, could require about $50 billion in investment.
Despite the size of the project, data center costs vary depending on technology and location. AI-focused structures in China generally present lower values than equivalents installed in the United States.