Zhipu AI officially announced that the anonymous model Ox Alpha, known in Chinese developer circles as 'Niu Lai,' is the newly released GLM-5.3-Flash model. Before its official launch, the model was provided for free anonymous testing on OpenRouter and OpenCode platforms. In five days, the testing attracted over 50 trillion tokens of traffic, setting growth records on both platforms. Furthermore, the usage of this model exceeded DeepSeek's figures by more than twofold.
Market attention was further drawn to an additional clarification from Zhipu: all this traffic was served using domestic chips. In its technical documentation, the company stated that its inference service operates on a cluster consisting of over 100,000 domestic chips. Zhipu claimed that 'the equipment efficiency and the cost per token have reached a level comparable to major Nvidia GPUs.'
The semiconductor research firm SemiAnalysis commented on X that serving every request using domestic silicon with efficiency comparable to Nvidia 'once again puts the sea of CUDA to the test,' especially after OpenAI announced its own inference chip, Jalapeño. LatePost reported that these chips could be produced by Huawei, Moore Threads, and Hygon, but Zhipu declined to confirm the suppliers or specific models.
To scale up to a context window of one million tokens, Zhipu developed a specialized inference engine based on SGLang. This engine utilizes W8A8 quantization, mixed INT8/FP8/BF16 cache quantization, and node-level tensor parallelism. Additionally, an architecture was implemented that separates encoding, prompt pre-processing, and token-by-token decoding, allowing independent scheduling of various pools.
The company asserts that the throughput performance of the end-to-end service has improved threefold compared to the initial baseline on the same hardware. The GLM-5.3-Flash model has a total of 320 billion parameters, with 18 billion activated, which is roughly equivalent to DeepSeek V4 Flash. It scores 57 points in the Artificial Intelligence Analysis Index, comparable to Claude Opus 4.8 and surpassing DeepSeek V4 Pro. The model costs about one-tenth the price of GLM-5.3 and one-hundredth the price of Claude Opus 4.8, making it an affordable option amid the general rise in prices for Chinese models.
