SenseTime releases SenseNova U1 Pro image generation model supporting up to 8K resolution via Raccoon and API
Read more
Pandaily
pandaily.com

SenseTime releases SenseNova U1 Pro image generation model supporting up to 8K resolution via Raccoon and API

SenseTime has introduced the production version of its image creation model, SenseNova U1 Pro. This model is now available to users through the company's Raccoon application and also via the SenseNova API service. According to a report by TechNode on September 22, 2026, the information was sourced from an announcement published on Sina Finance. The official English name for the product is SenseTime / SenseNova.

This release differs from previous coverage of the SenseNova U1.5 research and open-source training code. SenseTime asserts that U1 Pro utilizes a reasoning chain alternating with a text-to-image process to create visual materials containing structured information and readable text directly on the canvas, as well as production-ready layouts.

Instead of treating captions and pixels as separate stages, the model integrates reasoning steps with image formation. This ensures that the layout hierarchy, typography placement, and graphic elements remain consistent throughout the generation process. The product supports outputting images with resolutions up to 8K and allows for custom aspect ratios, targeting large formats and high detail rather than just square crops for social media.

The announcement lists target use cases, including infographics, posters, architectural visualizations, and other content that requires a uniform style and composition across the entire finished piece. These scenarios highlight the necessity of controlling composition and text rendering within the image itself, rather than using a separate layout tool after generation is complete. Thus, marketing and design teams can request production-ready frames from a single prompt.

Distribution occurs through two channels: consumers and creators gain access via Raccoon, while product teams can utilize programmatic access through the SenseNova API to integrate generation into their workflows. SenseTime has not published a full public model card detailing parameters, licensing terms, or independent benchmarking tables in the TechNode summary in English. Therefore, buyers are advised to view the claims regarding the 8K limit and layout capabilities as company assertions until third-party evaluations are available. For buyers of multimodal systems tracking Chinese image models, the nearest signal is the availability of the ready-made U1 Pro SKU for production in the Raccoon and API channels with an explicit ultra-high resolution limitation.

Similar stories

SenseTime releases SenseNova U1.5: an 8-billion parameter model for unified vision with open training code
Read more
pandaily.com

SenseTime releases SenseNova U1.5: an 8-billion parameter model for unified vision with open training code

SenseTime has introduced SenseNova U1.5, an eight-billion parameter model that unifies understanding, reasoning, and generation within the pixel space into a single native multimodal system.

Unlike existing pipelines that use a vision encoder for perception and a separate variational autoencoder for synthesis, U1.5 maps pixels and text into a common framework without these external modules. This aligns with the company's NEO-unify line, while improving spatial reconstruction to achieve higher accuracy at production resolutions.

The architectural change involves replacing independent MLP patch decoding with a lightweight spatial decoder. Visual tokens are transformed into a two-dimensional feature field and then reconstructed using Pixel Shuffle and local convolutions, allowing neighboring regions to exchange information before pixel finalization.

SenseTime reports that the interface still displays each area as 32x32 per visual token, but it supports native generation up to 4K with fewer seam breaks and textural inconsistencies compared to the previous U1 version. This was made possible by resolution-aware noise conditioning across wider aspect ratios.

The post-training process follows a recipe of specialization followed by unification. Reinforcement learning specialists focus on visual aesthetics, bilingual text rendering, infographic layout, and image editing. Subsequently, multi-expert policy distillation merges these skills into a single student along its own generation trajectories.

SenseTime asserts that U1.5 results surpass U1 in image generation and editing datasets, including competitive bilingual text rendering and multi-reference editing, while maintaining generally high multimodal understanding scores on standard STEM, VQA, and OCR benchmarks.

Alongside the technical report published on arXiv and Hugging Face Papers, SenseTime is making the training code open source, covering controlled fine-tuning, reinforcement learning, and policy distillation. Model assets are available in the SenseNova and Hugging Face collections under the OpenSenseNova organization. This release represents the full U1.5 product, not an early U1.5-Lite-Preview, and is intended for developers who wish to obtain a compact native unified vision model that can be fully explored, customized, and retrained.

Kinetix AI raises $75 million angel funding round to develop AI model hardware
Read more
ventureburn.com

Kinetix AI raises $75 million angel funding round to develop AI model hardware

Engineers have long faced a problem: artificial intelligence capable of writing novels in seconds struggles to grasp fragile glass without crushing it. This issue stems from a significant gap between digital intelligence and physical execution.

Bridging this gap requires flawless and continuous feedback between software systems and mechanical bodies. To solve this narrow problem, Kinetix AI has secured $75 million in a major and significant angel funding round, marking a serious step forward in embodied intelligence.

The historic funding round was led by Vertex Ventures, the powerful venture arm of Singapore's Temasek Holdings. F&G Venture and Wanshi Capital also joined the round to invest over 500 million RMB (approximately $75 million) in the early-stage startup, Kinetix AI.

Kinetix AI is not a small garage project. Founded in September 2025, the company has rapidly grown to nearly 200 employees. Driving Kinetix AI's rapid growth is CEO Yu Ze, who previously commercialized autonomous mining trucks for Huawei.

The team also includes Luo Ping, an outstanding deputy dean from HKU, and Zheng Qunyuan, former head of robotics at XPeng. This group possesses deep knowledge combining academic theory, autonomous vehicle logic, and manufacturing capabilities.

Many robotics companies cut corners on quality, but Kinetix does not accept this. The company firmly believes that since the world's infrastructure is built for humans, robots must look and move exactly like us to integrate seamlessly.

Their flagship creation, KAIBot, stands 1.73 meters tall and boasts an impressive 115 degrees of freedom. It is distinguished by its high quality and ultra-realism. Although creating such a robot requires significantly higher initial costs, and the supply chain is a nightmare, co-founder Zheng Qunyuan argues that starting with cheap, low-quality equipment is simply foolish. By locking down a true human form factor, they ensure their AI models will not fail as soon as any joint is updated.

Imagine a robot that doesn't just mimic the human silhouette but perfectly matches human motion data. When the hardware directly reflects our biology, algorithmic work transfers effortlessly.

The core magic happens thanks to the attracted capital, which is directed towards accelerating their closed ecosystem. This can be visualized as a three-headed monster. First, it involves collecting multimodal, egocentric data. Then, this raw data is fed directly into the native fundamental model embodied in the body. Finally, hardware such as KAIBot and the highly maneuverable KAI Hand physically executes the learned actions.

This continuous loop creates an 'intelligence flywheel.' As the robot collects sensory data from the real world, the AI model becomes smarter, instantly enhancing the capabilities of the physical hardware. There is no longer a need to rebuild the physical machine with every software update.

The system's dynamism was demonstrated at the 2026 World Games for Human Robotics, where their system played table tennis against world champion Ding Ning. The goal is to create a premium class of robots ready for complex real-world tasks.

Popular