Huawei's X-Router system enables model routing for optimizing agent queries
Read more
Pandaily
pandaily.com

Huawei's X-Router system enables model routing for optimizing agent queries

Huawei has introduced X-Router, a self-evolving model router that is part of the openJiuwen platform. This platform, developed by Huawei Cloud teams and Huawei Labs since 2012, as well as universities and enterprise developers, is designed to work with AI agents.

Announced on October 2nd, X-Router functions between an agent and the models it calls. Its task is to determine which model should process each incoming request: simple tasks are directed to lightweight and inexpensive models, while complex ones go to the most powerful.

The team notes that the system is friendly to Ascend and, during testing, allowed for a reduction in token consumption by over 50%.

Operating Principle and Architecture

X-Router is one of the algorithms in the openJiuwen model routing mechanism. It has a core written in Rust and an extension in Python, and its code is available on GitHub and AtomGit under the Apache-2.0 license.

The Qwen3-0.6B classifier, running in the host process, analyzes the dialogue and assigns it one of five complexity levels: SIMPLE, MEDIUM, COMPLEX, RESEARCH, or REASONING. Requests matching the assigned level or lower are processed by a local model, while more complex ones are sent to a cloud model corresponding to that level.

Furthermore, if the target model begins to function incorrectly, the router automatically switches to using a local model instead of escalating the issue.

The mechanism supports capability profiles for each model, which are created offline based on historical data and updated within a minute when a model's version or performance changes. At each step, it studies the recent conversation and tool invocation progress, considering the latest user query, and also weighs systemic parameters such as proximity to KV-cache, current load, and whether the model recently timed out or hit rate limits.

Additionally, an optional contextual bandit learns from the results of similar past requests and can override the classifier's decision if another level shows a clearly better result, while maintaining the initial choice if the evidence base is weak.

Since the routing algorithms are pure functions and memory is stored in a separate state layer, the loss of this state only reverts the system to cold routing, preventing request failure.

Testing Results

The team integrated X-Router into WorkSwarm, the openJiuwen agent working environment, and ran all 147 PinchBench tasks across 11 categories. When sending each request to Kimi-K2-Thinking, a score of 71.36% was achieved at a model cost of $11.11. Static routing using X-Router showed 66.3% with costs of $8.25.

With bandit training enabled, the result improved to 70.70%, and the cost decreased to $6.16, which is approximately 44.6% less than the baseline.

On Terminal-Bench via LLMRouterBench, using the Opus 4.8, Qwen3.5-122B, and Qwen3.5-35B model pool, the cost-saving oriented configuration achieved 95.2% success at 51.4% lower cost compared to Opus, while the quality-oriented configuration showed 98.6% success with a 15.7% reduction in costs.

The router is also integrated with the Multi-Model Collaboration (MoA) feature in WorkSwarm, allowing it to decide when multiple models should answer a complex question before their results are aggregated. The team plans further development of self-evolving features, including training routing strategies using reinforcement learning.

Popular