Zhipu AI's GLM-5.3-Flash model runs on 100,000 domestic chips and leads in usage on OpenRouter
Read more
Pandaily
pandaily.com

Zhipu AI's GLM-5.3-Flash model runs on 100,000 domestic chips and leads in usage on OpenRouter

Zhipu AI has introduced the GLM-5.3-Flash model, which, according to the company, operates exclusively on domestic artificial intelligence chips. One hundred thousand locally produced accelerators are used to process all of the model's online traffic.

The model ranked tenth in the AAII rating from the analytical firm Artificial Analysis, surpassing DeepSeek's V4 Pro Max. Zhipu AI emphasized that the entire volume of requests to GLM-5.3-Flash is served by these 100,000 domestically manufactured chips.

GLM-5.3-Flash was initially released under the codename 'Ox Alpha' on August 20th and quickly topped the popularity lists on the OpenRouter AI model routing platform within a week. The company positions this model as high-performance, sufficiently efficient for large-scale operation on China's internal computing infrastructure.

Zhipu AI has not publicly disclosed the specific chip supplier used. However, according to CNBC reports, analysts suggest that the hardware belongs to Huawei's Ascend series. Ivan Lin, an analyst at the research firm Counterpoint, noted that developers of Chinese AI models are increasingly directing investments into AI servers and computing infrastructure built on domestic chips.

This launch is the clearest signal yet that at least one leading model development laboratory in China believes that the domestic supply chain is ready to support flagship workloads. Zhipu AI is among the most active 'Chinese tigers' in expanding computing power, and the daily operation of a frontier-class model on its own accelerators is a statement about both performance and chip availability.

This move also aligns with a broader industry trend. As export controls restrict access to advanced foreign chips, creators of Chinese models are rushing to optimize their software for the hardware they can actually acquire. A model demonstrating high results in benchmarks while running on domestic chips serves as proof that software efficiency can compensate for hardware gaps.

For Zhipu AI, the Flash line represents a strategy combining model quality with practical scalability. With the support of 100,000 chips, GLM-5.3-Flash aims to demonstrate that the Chinese laboratory is capable of providing competitive frontier-level performance without reliance on imported accelerators.

Similar stories

Flash Models Change the Flagship Segment of LLMs in China: Availability Outpaces High Performance
Read more
pandaily.com

Flash Models Change the Flagship Segment of LLMs in China: Availability Outpaces High Performance

The flagship lineup of large language models (LLMs) in China has undergone significant changes almost overnight. Zhipu AI released and open-sourced the GLM-5.3-Flash model, which had been an anonymous participant in 'Ox Alpha,' holding first place on OpenRouter since August. Concurrently, Alibaba introduced Qwen3.8-Flash, making its weights available on Hugging Face and within the ModelScope community.

Both models belong to the 'Flash' series and feature aggressive pricing, challenging the established notion that higher performance always equates to higher cost. The GLM-5.3-Flash model supports a one-million-token context window and is the first natively multimodal model in the GLM-5 family capable of processing images, videos, files, and even computer usage.

During a limited-time discount, the input cost for GLM-5.3-Flash is 0.4 RMB per million tokens, and the output cost is 1.4 RMB, making it approximately twenty times cheaper than GLM-5.3. Qwen3.8-Flash, presented as an early preview of the Qwen4 architecture, also maintains a million-token context window. At a price of 1 RMB per million input tokens and 3 RMB for output, it is about twelve times cheaper than Qwen3.8-Max while reducing token consumption by 75% in office work scenarios.

These rapid and competitive releases are putting pressure on DeepSeek, which has long served as the benchmark for price-to-performance ratio in China. The current price for DeepSeek V4 Flash is 1.5 RMB for input and 4.5 RMB for output per million tokens, including peak and off-peak load multipliers. In benchmark comparisons, GLM-5.3-Flash showed results close to Opus-class models—a level that 'Flash' promises to achieve. The anonymous launch of this model ended DeepSeek's 56-day dominance on OpenCode. Zhipu reported that the traffic was served by over 100,000 domestic AI chips, whose efficiency, they claim, is comparable to major NVIDIA GPUs.

A deeper signal lies in the strategic shift: the term 'Flash' no longer simply means a trimmed-down version. Developers are now positioning these models as true flagships due to their large context, multimodal input, competitive reasoning, and a price point that makes them the standard choice for everyday production tasks in code agents and office suites.

As a result, a market is forming where price sensitivity is rapidly fragmenting. Model vendors are no longer competing solely on maximum performance but are focusing on providing practical workloads at a sufficiently low price so that most developers never need to resort to the 'Pro' tier. In the Chinese LLM market, availability and generosity have quietly become the new standard for the flagship level.

Zhipu AI unveils GLM-5.3-Flash, a model previously known as 'Ox Alpha'
Read more
pandaily.com

Zhipu AI unveils GLM-5.3-Flash, a model previously known as 'Ox Alpha'

For several days, AI developers worldwide attempted to identify the owner of Ox Alpha—an enigmatic product that suddenly appeared on the OpenRouter platform without specifying the developer, disclosing parameters, or providing a technical report. This model quickly rose in the platform's trending list due to its capabilities in code generation, complex reasoning, and executing long-term agent tasks.

Last week, the secret was revealed: Ox Alpha turned out to be the latest release from Zhipu AI—the GLM-5.3-Flash model. The model was first posted anonymously on OpenRouter on August 20th as part of a practical blind test. During this period, its provider offered approximately 100 trillion tokens of daily service capacity for free, attracting a large number of real users. Patrick Collison, co-founder of Stripe, called it 'very impressive.' Notably, Zhipu stated that all traffic generated during the anonymous launch was served by a large cluster deployed on domestic AI chips.

GLM-5.3-Flash is a 'mixture-of-experts' type model with a total of 320 billion parameters, although only about 18 billion are actively used at each inference step. This differs from the GLM-4.5 series, which has 320 billion total and 32 billion active parameters, and features a reduction in layers from 92 to 45. The 'Flash' prefix in the name indicates an architecture designed to maintain near-flagship performance while simultaneously reducing the computational load per token.

Furthermore, it is the first natively multimodal model in the GLM-5 family, capable of accepting text, images, videos, and files into a context window of approximately one million tokens. It was trained on a multimodal corpus of 30 trillion tokens. To make million-token inference accessible, Zhipu combined linear attention with sparse attention—an approach the company calls the first hybrid architecture of this kind in an open state-of-the-art model—and added IndexPool to reduce indexer overhead. The company estimates that the attention computational load is reduced by about a third, and the KV cache by about a quarter compared to GLM-5.3.

Tests show the model's strengths in coding and agent tasks: it scored 63.4 in DeepSWE v1.1 versus 46.2 for GLM-5.2, 48.8 in AutomationBench versus 26.2, and 84.3 in Terminal-Bench 2.1. The model weights are now available on Hugging Face under an MIT permissive license, supporting local deployment via SGLang, vLLM, and related frameworks. By withholding information during the launch, Zhipu allowed developers to form an opinion based solely on the model's actual performance.

OpenAI develops proprietary AI chip Jalapeño to compete with Nvidia
Read more
www.aajtak.in

OpenAI develops proprietary AI chip Jalapeño to compete with Nvidia

OpenAI, the creator of ChatGPT, is developing its own artificial intelligence chip named Jalapeño. The company claims that this first AI chip will provide higher speed for AI model operation while reducing power consumption.

This move is expected to be highly significant as it will help improve artificial intelligence support and reduce the company's dependence on Nvidia. Although ChatGPT is already popular as a chatbot, the company is now aiming to expand into hardware by designing devices on which AI models run.

Partnership with Broadcom

Jalapeño is a specialized custom chip designed specifically for running AI models. The company prepared it in partnership with Broadcom.

Design Focused on Energy Efficiency

OpenAI announced that the chip is designed to demonstrate high performance with low energy consumption and minimal response time. The company tested it on several selected AI models, including GPT-OSS 120B, Deepseek R1, and Kimi K2.5 1T.

According to OpenAI's statements, the Jalapeno chipset solved one of the serious problems in AI hardware—achieving high performance without increasing latency. The Jalapeno chipset is specifically created for the inference process. Inference can be understood as the process where a trained AI model processes a user request and generates a response, such as when a user asks a question to ChatGPT.

The company's goal is to ensure complete performance and efficiency by adapting the chip, software, memory, network infrastructure, and service system to its own models.

Popular