DeepCybo releases PhysBrain 1.5 — an open physical foundation model achieving 72.5 across 28 benchmarks
Read more
Pandaily
pandaily.com

DeepCybo releases PhysBrain 1.5 — an open physical foundation model achieving 72.5 across 28 benchmarks

DeepCybo has introduced PhysBrain 1.5, a physical foundation model that integrates embodied world understanding, action generation, and future state prediction into a single autoregressive core. This model is built upon Qwen3-VL and was enhanced with tokens for actions and visual states.

Specifically, the 8-billion parameter version demonstrates an overall score of 72.5 across 28 public benchmarks related to the embodied world. This is claimed to be the first result among open-source models and is approximately one point behind leading proprietary systems such as GPT-6 Astra (73.3) and Gemini 3.6 Flash (73.0), according to the project's published table.

According to DeepCybo's report, the PhysBrain 1.5-8B version ranks first in 14 out of 28 benchmarks and second in 10 among open-source models. These tasks cover visuospacial perception, multi-view reasoning, embodied world planning, capability grounding, and visual tracking tasks.

A lighter version, PhysBrain 1.5-2B, achieved a score of 66.6 on the same test set, serving as a compatible deployment tool rather than a primary contender in the rankings. The weights for both versions, the technical report, and the evaluation set are available on Hugging Face and GitHub under the DeepCybo / PhysBrain 1.5 project.

The training process is implemented through a closed physical loop: observation, reasoning and grounding, action execution, and then predicting world changes. Language responses, spatial structures, end-effector trajectories via ActionPiece tokens, and aligned future RGB, depth, and robot mask frames all utilize a single common objective—the next token—without using separate task heads.

Pre-training relied on videos of human interaction (in egocentric, ego-exocentric, and panoramic modes). Additional fine-tuning was conducted by mixing human demonstrations, real robot trajectories, and simulation using DeepCybo's Human-as-Humanoid action bridge.

Compared to other Chinese open lines such as Intern W0, Motus2, or Cog-WM, PhysBrain 1.5 represents the release of a foundational model with claimed leadership across multiple benchmarks, rather than just a policy demonstration for a single robot. Teams using this model are advised to re-run the evaluation suite, investigate ActionPiece transfer on their own manipulators, and view the aggregated score of 72.5 as a starting point for comparison, rather than equating proximity to proprietary systems with immediate applicability.

Similar stories

OpenBMB releases MiniCPM5-2B as an open SOTA model for on-device operation with less than 4 billion parameters
Read more
pandaily.com

OpenBMB releases MiniCPM5-2B as an open SOTA model for on-device operation with less than 4 billion parameters

OpenBMB, an open project associated with ModelBest, has introduced MiniCPM5-2B—a dense language model designed for direct on-device operation. This model features approximately 2.52 billion parameters and supports a native context length of 131,072 tokens, distributed under the Apache-2.0 license.

MiniCPM5-2B is the second release in the MiniCPM5 series, following MiniCPM5-1B. It utilizes the standard LlamaForCausalLM architecture, which includes 42 layers with grouped-query attention, employing 16 query heads and 2 key/value heads. This design allows core engines to load the model without requiring custom kernels or code forks.

The non-embedding parameters amount to about 1.98 billion, enabling the model to fit within the sub-4-billion parameter class suitable for hosting on phones, laptops, and edge devices.

When compared on the 34-benchmark dataset from OpenBMB, MiniCPM5-2B achieves an average score of 53.9, surpassing several larger open base models. Among these, Qwen3.5-4B scored 51.1, and granite-4.2-3B scored 42.7. Conversely, similarly sized models like LFM2.5-2.6B and Qwen3.5-2B lag behind.

The model's strengths are concentrated in areas such as coding agents, tool usage, and long-context information extraction. For instance, on LiveCodeBench v6, MiniCPM5-2B achieved 69.1 compared to 56.4 for Qwen3.5-4B, and on SWE-bench Verified, it reached 46.4 versus 33.6. High results were also obtained on τ²-Bench Telecom (97.1), BFCL v4 (66.6), and NoLiMa (68.1 versus 43.5).

In tasks requiring deep knowledge, performance remains closer to the upper limit of its size; for example, on MMLU-Pro, the model scored 70.8, while the benchmark 4B Qwen model achieved 78.0. OpenBMB specifically notes rows sourced from Artificial Analysis to allow users to distinguish between vendor data and third-party results.

The post-training process followed the UltraData level management: initially, supervised fine-tuning was conducted based on deep reasoning using approximately 400 billion tokens. Subsequently, specialized reinforcement-learning teachers for mathematics, code, agents, and writing were applied using the critique-based JustRL II algorithm. The final stage involved on-policy distillation, which merged 16 RL experts—five of which were agentic—into a single final student model.

OpenBMB attributes the average gain of about 10.96 points in reasoning and general benchmarks, along with 6.96 points in agentic benchmarks, to the RL and distillation stages. To allow for precise measurement of each component's contribution, the company releases intermediate checkpoints: Base, Midtrain, and SFT-only.

Along with the weights, UltraData publishes datasets covering web, code, mathematics, SFT, and RL samples for agents, enabling result reproducibility. Runtime support includes vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, MLX, LiteRT, and multi-chip FlagOS builds. Packages in GGUF, MLX, and GPTQ formats are provided for local assistants that need to perform tool calls and handle long context without relying on datacenter GPUs.

For developers focused on on-device deployment, MiniCPM5-2B is positioned not so much as a general knowledge giant, but rather as a compact Apache-licensed agent and coding companion that already surpasses some open 4B class base models based on published average scores.

UniPat raises $300 million with Alibaba support to expand AI testing and benchmarking
Read more
ventureburn.com

UniPat raises $300 million with Alibaba support to expand AI testing and benchmarking

As artificial intelligence develops, the problem of testing has emerged. While many focus on creating impressive large language models to attract investors, there are only a small number of verified platforms to check the functionality of these complex systems after training is complete.

AI models are prone to hallucinations and can drift from the norm over time. To solve this significant operational problem, which constantly deters large corporate users, UniPat has raised $300 million. UniPat actively conducts stress tests on these complex systems to ensure a full level of trust.

UniPat differs from other AI companies; they do not just create new algorithms to demonstrate at expensive tech conferences. Instead, the startup specifically tests models based on real-world scenarios, rather than sterile laboratory conditions.

This can be compared to a rigorous training camp for artificial intelligence. Before a model gains access to real consumer data or begins performing automated financial operations, UniPat subjects it to intensive trials. This allows for the generation of high-quality system performance data that developers need to eliminate critical flaws, as what hasn't been thoroughly broken first cannot be fixed.

The tech giant Alibaba led this major funding round, causing a stir in the industry. The financial details of this round are quite impressive: the new influx of capital boosted the startup's valuation to an impressive $2.5 billion post-investment. It is also interesting that this specialized center is headed by a former company employee.

It is clear that Alibaba is interested in retaining its top talent. Significant funds continue to flow despite the caution of the broader venture capital market. Moreover, this specific deal ranks among the top three percent of all registered late-stage venture capital rounds.

The high cost of the testing platform is due to basic corporate economics and the need for competitive survival. Early investors included representatives from Sequoia China. This demonstrates that leading financial players are looking for infrastructure-related enterprises amid the current AI race. We have moved past the phase of simply admiring smart chatbots.

Now, large international corporations demand flawless analytics. They need solid proof that implementing multi-million dollar AI will not lead to public embarrassment. UniPat provides this necessary insurance policy. This is not just another routine technology investment. It is a calculated move. As global powers fiercely compete for dominance in artificial intelligence, the basic infrastructure for evaluating these tools becomes infinitely valuable.

Alibaba's massive bet signals a clear shift in market priorities. The future belongs not only to those who can build the largest model but also to those who can prove their model is the safest, fastest, and most reliable in real-world conditions.

Popular