MiniMax has introduced its first open multimodal generative model, H3, on July 31st, marking a significant event in the field of AI video generation. The model supports 2K audio-video generation with native two-channel sound and a duration of 15 seconds, and it secured the top spot in the global video editing ranking on the Artificial Analysis platform.
Competitive Pricing and Model Capabilities
The cost for generating videos using the H3 model is set at 0.8 yuan per second, which is only one-third of the prices for comparable flagship products. Following the announcement of the release, the company's shares rose by 14%, recovering previous declines, as investors reacted positively to the cost control strategy.
Technical Features of H3
H3 is capable of understanding contexts related to text, images, video, and audio, while generating audio-video with native two-channel sound. The model's technical advantage lies in its ability to follow instructions, correctly represent textual and brand information, and feature V2V motion transfer—the ability to transfer motion from one video frame to another. This function enables controlled editing of multimodal content.
Application and Economic Efficiency
The H3 model facilitates content generation for various commercial use cases, including advertising, branding, e-commerce, product development, UI/UX, and gaming. This significantly simplifies the process of creating high-quality video content for enterprises and developers. The cost advantage is achieved through systematic optimization: using a highly compressed tokenizer reduces the required number of tokens for video generation, while heterogeneous training configuration, load balancing, and improved GPU utilization substantially lower both training and inference costs while maintaining model quality.
Open Source Strategy
A central element of the release is the open-source strategy. MiniMax plans to make the model weights available within a few days, subject to regulatory compliance. This will allow companies to deploy the model locally and adapt it using their own data to ensure security and compliance. The company stated that chip manufacturers and developers can participate in adaptation and optimization to reduce operational expenses and expand application scope, following the industry trend of moving from closed services to open, collaborative ecosystems.
Comparison with Competitors and Market Trends
MiniMax's approach aligns with the example of Moonshot AI's Kimi K3, the largest open-weights model at 2.8 trillion parameters, which reached cluster computing limits within 48 hours of its release and suspended new consumer subscriptions on July 19th for priority use of computational resources by paying users. MiniMax's strategy deliberately avoids the parameter race characteristic of the Chinese AI industry. While leading players face computational limitations and supply shortages, MiniMax focuses on an extreme price-to-performance ratio combined with open-source distribution to build a commercial advantage based on unit economics rather than model scale. Market observers will monitor the community adoption progress of H3 after the open-source release, the pace of commercial orders, and whether the volume-based pricing strategy can return MiniMax to a leading position in the AI race. The model positions economically viable, commercially deployable multimodal generation as a new frontier of competition, and its video editing capability indicates that production quality for professional workflows may become as important as basic generative power.