Meituan releases LongCat-2.5-Preview model: multimodal agent with 1.6 trillion parameters and 1 million token context
Read more
Pandaily
pandaily.com

Meituan releases LongCat-2.5-Preview model: multimodal agent with 1.6 trillion parameters and 1 million token context

Meituan introduced LongCat-2.5-Preview on its LongCat API platform on September 25, 2026. According to information from iFeng Tech, AI Tool Lab, and related industry publications, this release is positioned as a long-term multimodal agent model, rather than just an update to chat functions. The official English name for the product is Meituan / LongCat.

This review presents the technical specifications of the Preview product and the intended use of the agent software. Meituan has not published official benchmarks for this release, so claims regarding capabilities should be viewed with caution.

Architecturally, the model retains the Mixture-of-Experts approach used in LongCat-2.0. The total parameter count is approximately 1.6 trillion, with about 48 billion parameters activated at each inference step. The model features a built-in context window of one million tokens, making it suitable for processing long documents, repositories, logs, and multi-turn tool cycles.

The main difference from version 2.0 is the integration of native multimodal capability directly into the base model. This includes image parsing for cross-modal question answering, as well as visual reasoning and summarization. Furthermore, there has been a clear shift from purely agentic coding to autonomous workflow management across terminals, browsers, user interfaces, spreadsheets, and design tools.

Access to the model is provided in two formats: via API (with endpoints compatible with OpenAI and Anthropic, as described in reviews) and through a web interface at longcat.chat. The maximum output text length is set at 128 thousand tokens.

Integration notes mention tools such as Codex, OpenCode, OpenClaw, CatPaw, Claude Code, Hermes, and Kilo Code, allowing existing agent systems to directly access the Preview version. Secondary reports cite promotional pricing indicating limited-time input usage at around $0.30 and output usage at around $1.20 per million tokens, with an additional tiered input level and free tokens for current users; commercial terms should be considered as temporary Preview offers.

Since LongCat-2.5-Preview does not come with a public rating in this cycle, readers are advised to view the improvements in agent software as a positioning claim awaiting independent evaluation. For developers tracking multimodal agents in China, the key takeaway is that Meituan has presented a 1.6T / ~48B MoE model with native 1M context, focused on long sequences of software operations rather than financial history or ranking changes.

Similar stories

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15
Read more
pandaily.com

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15

StepFun introduced the preliminary version of Step 5 on September 20, 2026, as its next flagship foundational model designed for long-horizon agents, according to information from Tencent Tech, DataLearner model cards, and company materials available at stepfun.com. The official English name of the model is StepFun / Step 5.

It is important to note that although the model identifier step-5-preview is already available via the product API and the open StepFun platform, the open weights in BF16 format are scheduled only for October 15, 2026, and were not available at launch. Therefore, availability via API and open weights should be considered as separate points.

The model architecture features a Mixture of Experts (MoE) sparse design with an approximate total of 600 billion parameters. Approximately 27 billion parameters are activated per token within a 92-layer 'narrow and deep' Transformer. The choice of depth is driven by agent requirements, as longer information paths through layers contribute to improved multi-step implicit reasoning and handling of long tool outputs.

The model supports a one-million-token context window, accepting both text and image input (though video input is mentioned on third-party cards). It is oriented towards applications in AI coding, software development, financial analysis, and professional agent usage. To maintain practicality with the million-token attention, sparse GQA was applied in combination with token block merging, reducing the cost of indexer and top-k selection by approximately one eighth.

During training, special attention was paid to bit-level alignment between training and inference to ensure stable MoE routing, and load-aware scheduling, speculative decoding, and FP8 paths were utilized. StepFun claims that these methods accelerate the long-horizon Reinforcement Learning (RL) process by more than three times overall.

According to data from Artificial Analysis's composite AI index, it scores 44 points. The company places this rating among leading models focused on open access. Aggregator cards also list API prices: about $1.00 per input token and $2.70 per output token per million tokens, but comparisons of metrics and cost should be viewed as statements from third parties or providers.

Until October 15, Step 5 Preview should primarily be regarded as a functional API for long-horizon agents, having only a planned commitment for open weights, rather than a release of weights on Hugging Face or a drop-in replacement for Meituan LongCat, MiniMax Code Flash, or NaiveAI OSS MoE.

TaichuAI releases open multimodal model ZDTaichu5.0-9B for spatial understanding
Read more
pandaily.com

TaichuAI releases open multimodal model ZDTaichu5.0-9B for spatial understanding

TaichuAI has made the ZDTaichu5.0-9B model publicly available. This multimodal foundation model has approximately 9 billion parameters and is designed for general visual understanding, spatial reasoning, agent tool usage, and embodied AI research.

The model's architecture combines the Qwen3.5-9B language backbone with the C-RADIOv4-H vision encoder. The model accepts text input, as well as one or more images and videos of any resolution, supporting a context length of up to 128 thousand tokens. The release, summarized by TMTPost on September 15th, focused on real-world spatial perception, transformations between different views, and task planning for embodied systems.

According to the public model card, TaichuAI demonstrates state-of-the-art results among approximately 10-billion general VLMs across a wide range of visual tasks, while also expanding its capabilities in spatial, embodied, and agentic scenarios. Noteworthy metrics in space and embodiment include ViewSpatial 62.50, MMSI-Bench 47.20, MindCube-tiny 78.27, ERQA 48.00, and RoboSpatial 56.00.

Regarding agent and instruction sets listed in the evaluation card notes, TAU2-Bench achieves a score of 87.70, and the average Claw-Eval score is 71.40, while IFEval shows 93.70. The entropy-prioritized adaptive recursive reasoning mechanism is described as a mechanism that allocates additional refinement steps for latent variables for more complex tokens.

The model's capabilities cover Optical Character Recognition (OCR) and document understanding, math on images, fine-grained 2D relationships, multi-view association, 3D scene and perspective perception, as well as tracking multiple images and videos within a long context window, including multi-step tool use planning. Actual tool execution remains the responsibility of the host application.

The spatial learning topics listed on the card include relative relationships, dense counting and framing, camera motion and depth ordering, egocentric versus allocentric views, and high-level capability definition and action planning for VLA-style adaptation.

The model weights are hosted on Hugging Face at TaichuAI/ZDTaichu5.0-9B under the NVIDIA Open Model License, while retaining the Apache-2.0 notices from Qwen3.5. A custom branch of vLLM 0.26.0 and a Docker image from TaichuAI are used for serving to provide OpenAI-compatible endpoints, with recommended sampling settings for spatial grounding and general tasks.

As with other vendor-released models, external laboratories should view benchmark scores as reported with the specified prompts and judges until independent reproduction results emerge. However, the combination of a mid-sized open multimodal checkpoint, a focus on space and agents, and ready vLLM packaging provides researchers with a concrete artifact for evaluation.

Infinigence AI releases APXInf, an inference engine for embodied models on the Jetson Thor edge
Read more
pandaily.com

Infinigence AI releases APXInf, an inference engine for embodied models on the Jetson Thor edge

Infinigence AI, in collaboration with teams from Tsinghua University and Shanghai Jiao Tong University, has released APXInf to the public. This engine is designed for edge inference and is built for embodied models that must close the perception-decision-action loop directly on the robot, rather than in the cloud.

The public package is available on GitHub at RLinf/APXinf-robo and includes a Rust runtime with Python bindings, optimized for Jetson class devices and desktop GPUs. The company positions this release as the 'last mile' between capable vision-language-action models and stable control directly on the robot.

According to Infinigence data and coverage by QbitAI, using the FP8 format with PI 0.5 on the NVIDIA Jetson Thor platform reduces end-to-end inference time from approximately 278 ms to less than 26 ms. This achieves a frequency of about 38.46 Hz, which is sufficient for real-time multi-scenario control. This improvement represents an approximate tenfold reduction in latency compared to the unoptimized baseline they cite.

This breakthrough was made possible through work on the pipeline, graph, kernel, and quantization, as well as the use of Agent4Kernel, an agent-assisted CUDA kernel fusion.

Initial model support includes PI 0.5 and WALL-OSS. Target hardware platforms include Jetson Orin, Jetson Thor, and GeForce RTX 4090. Architectural decisions emphasize long-term stability alongside peak speed: a minimal Rust runtime is used for memory safety, complemented by Python ergonomics for developers. Furthermore, an agent workflow has been implemented to adapt new VLA/WAM checkpoints without requiring manual writing of every kernel path.

This engine integrates with Infinigence's earlier RLinf reinforcement learning training stack, extending the same open ecosystem from post-training to deployment on the robot with OpenPI-compatible servicing capabilities.

Future development plans include ports for a larger number of VLA/VLM and world models (work with Qwen and GR00T classes is already planned), optimization of nvfp4, development of backends for AMD, and support for domestic chips for inference and operating systems for domestic robots. Teams are encouraged to independently reproduce the latency metrics for Thor/Orin on their own checkpoints and within specified power consumption parameters. The stated FP8 value below 26 ms sets a standard for embodied system inference on the edge but is not a guarantee for every model or package form.

Popular