Stable AI and Tsinghua University introduce LimiX-2 model for structured data
Read more
Pandaily
pandaily.com

Stable AI and Tsinghua University introduce LimiX-2 model for structured data

Stable AI, in collaboration with the group of Professor Peng Cui from Tsinghua University, has released LimiX-2—a foundational model with 400 million parameters designed to work with structured and tabular data. The weights and inference code are available on Hugging Face at stable-ai/LimiX-2, and the technical report is posted on arXiv under number 2609.17488.

One of the tested models is capable of performing classification, regression, and missing value imputation in a single forward pass, without requiring fine-tuning for a specific task.

Architecturally, LimiX-2 utilizes Contextual Mechanism Networks (CMNs), which were pre-trained using Contextually Conditional Masked Modeling. Instead of focusing on a traditional objective function tuned for tabular data that predicts a given goal based on context, CMNs are trained to discover the joint structure of dependencies between features and labels, which is closer to identifying how variables jointly generate data rather than tuning one column at a time.

Pre-training was conducted on synthetic datasets derived from structural causal models covering diverse graphs, mechanisms, and observation processes. The model was then scaled according to previously published LimiX scaling laws up to the 400M class.

On public leaderboards, the team demonstrated an Elo of 1935 on TabArena, 1432 on BCCO, and 1506 on TALENT. These results are claimed to be the first among comparable tabular foundation models and AutoGluon-style baseline models within the published protocols. The TabArena results also highlight strong performance in regression and classification splits; the model card additionally notes the recovery of the causal skeleton through feature attention, which encodes direct causal relationships.

Previous LimiX releases have covered hundreds of corporate structured data scenarios; LimiX-2 represents a scaled generation of CMN aimed at advancing this area of general tabular models further.

For data teams, the practical value lies in having a single open weight that can classify, regress, and impute data in heterogeneous tables without the need for retraining for each dataset. This is particularly useful where corporate tables do not resemble publicly available text corpora. However, it should be noted that the benchmarks are presented as Elo scores claimed by the authors under specific configurations; end-users in production environments will need to conduct checks on held-out samples and monitor for data leakage before replacing specialized AutoML pipelines.

Similar stories

Shengshu Technology introduces Motus2—a self-evolving world model for high-precision manipulation
Read more
pandaily.com

Shengshu Technology introduces Motus2—a self-evolving world model for high-precision manipulation

Shengshu Technology has released Motus2, which is a self-evolving general world model specifically designed for complex robotic manipulations. Co-founder and CEO Luo Yihang demonstrated this system at the Bund summit in 2026.

The authors of the Motus2 project include Shengshu's research identity, GensPI, as well as Tsinghua University; the technical paper was published on arXiv under the number 2608.30237. Unlike many systems that connect an action module to a separate world simulator, Motus2 integrates policy, simulation, and evaluation into a single video-action network with shared parameters.

This network provides three control interfaces. The world-action model offers executable action blocks. The action-conditioned world model predicts the visual outcomes of these blocks. And the value model evaluates the predicted outcomes for selection and learning. During testing, Best-of-N planning analyzes several options, imagines their future, and executes the branch with the highest value before the robot observes the real scene again and replans.

After training, model-based reinforcement learning transforms the same value signal into policy updates, while the prediction and evaluation weights remain frozen so that feedback does not erase the learned dynamics. Furthermore, an action-oriented information mask prevents the policy from looking into future video tokens before making an action decision—this prevents a failure that the team associates with earlier simplified methods combining vision and control.

When performing tasks such as phone placement and multi-finger operations on real robots, the baseline policy achieved about 65% success. Self-planning showed around 67.5%; model-based reinforcement learning showed about 72.5%; and the combination of both methods provided success at about 75%. A lightweight tactile expert, which reuses base features to refine short sub-blocks, increased the success rate of cup retrieval and paper tearing from approximately 60% to 72.5%.

Training data includes a human-centric pyramid consisting of approximately 130,000 hours of egocentric recordings from monocular to synchronized stereo vision, as well as over 100 hours of robot trajectories and human-robot co-occurrence. Scaling stereo imagery from 2,000 to 20,000 hours continued to reduce action prediction error in deferred tests.

Hardware demonstrations include high-degree-of-freedom platforms such as Sharpa Wave and Wuji Hand 2, which are used for screwing in light bulbs, turning pages, and other multi-finger contact tasks requiring verification of both vision and touch.

The team positions Motus2 as an early cycle of recursive self-improvement within constrained tasks—using simulated outcomes to adjust the policy, rather than as open autonomous learning in unbounded environments. For potential technology implementation buyers, the key signal is the closed loop of WAM, AC-WM, and the value model, as well as the measured improvement in phone and multi-finger contact-rich tasks. Shengshu views this loop as progress toward its general world model from L3 to L4, without claiming full autonomy in the open world at this time.

Review of Internal Solutions for Connection Scaling in Chinese AI Chip Developments
Read more
pandaily.com

Review of Internal Solutions for Connection Scaling in Chinese AI Chip Developments

An industry report from September, prepared by Moyuan Think Tank and distributed by Semiconductor Industry Watch, analyzes how five registered Chinese AI accelerator suppliers are building their own factories for scaling. These factories represent connections between chips that allow multiple cards to function as a single machine within a server or rack.

The analysis is technical in nature and focuses on published throughput, product matching, and levels of implementation evidence for the ecosystem, excluding narratives about fundraising or stock prices.

Scaling differs from Ethernet or InfiniBand fabrics, which combine nodes into larger clusters. This review only considers accelerator-to-accelerator buses developed by the chip vendor itself. These are then evaluated based on public data: ranging from a complete lack of information to research papers, published specifications, samples, customer confirmation, first shipment, redeployment, and redeployment with a verifiable unit of connection area.

The maximum number of cards listed in datasheets is not considered proof that customers are already using such an area size in production.

Cambricon's MLU-Link is the most obvious example of first shipment among the materials reviewed. The product pages for MLU370-X8 list a cumulative bidirectional throughput of 200 Gbps per card with four ports. Earlier MLU290 class modules demonstrate an MLU-Link cumulative throughput of up to 600 Gbps on six ports, supporting multi-chip boards and inter-system connections beyond simple PCIe.

Enflame presents a patented path for scaling GCU-LARE through generations of CloudBlazer. Public data shows around 200 Gbps bidirectional on DTU 1.0 class cards and approximately 300 Gbps on DTU 2.0 / T20–T21 class modules, providing a clearer trajectory from chip to generation, even if the latter customer area sizes are harder to verify through open documentation.

Biren reveals its BLink scaling protocol. BR166 class materials indicate a connection bidirectional throughput of about 576 Gbps, while newer BLink 2.0 announcements hint at memory semantic links and optical super-node roadmaps targeting up to 1024 GPUs.

Moore Threads uses the MTLink brand, citing figures up to 240 Gbps on MTT S4000 bridge topologies covering two-, four-, and eight-card connections, as well as 784 Gbps from card to card on MTT S5000 devices. KUAE cluster materials emphasize the linearity of multi-card training.

MetaX described its own scaling relative to competitor products in official documents and demonstrated Xijing S600 rack-level designs connecting around 64 GPUs, although the report still requires a closer link between shipped SKUs and the exact fabric provided by customers.

Despite the throughput figures becoming discussed among the five companies, parameters such as bidirectional versus unidirectional units, counting per card versus counting per chip, and verified deployments within a single area remain uneven. None of the vendors reviewed in the report have reached the highest level of redeployment with a independently verifiable area size—this is a gap in evidence that operators should monitor as thousands of card clusters multiply.

Popular