XPower Technology has unveiled the RPP AE7100E architecture chips at the WAIC event. These nail-sized microchips are capable of supporting models with parameters ranging from 400 billion to 1.6 trillion by utilizing distributed computing, while maintaining minimal token cost.
Challenging the Dominance of Large Chips
XPower Technology has established itself as a significant player in the domestic AI computing market at WAIC 2026. The company demonstrated its matrix of artificial intelligence accelerators based on its proprietary Reconfigurable Parallel Processor (RPP) architecture. This architecture covers cloud, edge, and end computing scenarios. The company's core idea directly challenges the industry's prevailing doctrine regarding the necessity of using giant chips.
The Technological Advantage of RPP
The AE7100E chips, which are about the size of a fingernail, can be integrated in quantities of 12 per accelerator card. This configuration allows for running large language models with 400 billion to 1.6 trillion parameters through distributed computing, achieving an industry-optimal cost efficiency per token. XPower CEO, Li Yuan, noted that the shift from training-oriented computing to inference-oriented computing is a key trend. He emphasized that AI programming has become a critical use case for inference, as the company's experience has shown that AI can now perform CPU compiler development, which previously required a professional team for years, and this capability can be scaled for rapid application across various tasks.
Economic Viability and Architecture
Since 55% of model requests in global data centers are directed towards open-source models, the inference market is accessible not only to a few giants but to any enterprise, creating huge demand for cost-effective inference infrastructure. A key observation is the inefficiency of large chips during inference: decoding requires a computation-to-data ratio of 1:1, whereas GPUs provide a 1000:1 ratio, leading to throughput loss. XPower's RPP architecture uses distributed small chips, enabling fine specialization within a single server. Furthermore, XPower's approach replaces expensive HBM technology with the more mature LPDDR, asserting that three elements are fundamentally necessary for AI inference: high capacity, sufficient bandwidth, and low cost. LPDDR technology meets all these requirements at a significantly lower cost than HBM. The configuration of 12 chips per card allows for detailed memory partitioning and parallel data processing between chips within a single server form factor. This distributed server architecture differentiates XPower from the NVIDIA Groq approach, which also uses distributed multi-chip designs, but on a much larger cluster scale. The inherent flexibility of the RPP architecture, balancing general and domain-specific computations, allows one chip to be used for deployment scenarios in the cloud, at the edge, and on the end device.
Strategic Market Vision
During WAIC, the company showcased its product line, ranging from the AE7100E chip to high-density server racks with 8 cards. The company's transition from edge computing to cloud is a strategic bet that the inference market—which XPower estimates to be ten times larger than the training market by total addressable value—will be captured not by the largest computational power of a single chip, but by the most cost-effective distributed systems. Although confirmation of this hypothesis will depend on real-world implementation results, the RPP architecture represents one of the most original domestic approaches to computing, challenging the consensus on large chips by asserting that the future of AI inference belongs not to single monolithic processors, but to clusters of coordinated small chips.