On July 31st, MiniMax released a model that should make content creators recalculate their budgets: H3. It is not "just another video generator" — it is a model that simultaneously handles video, image, audio generation and editing, outputting 2K resolution + native stereo sound, with a maximum duration of 15 seconds. Priced at 0.8 yuan per second, it is less than one-third the cost of similar flagship models. Moreover, the weights(the model's core files — get them and run it yourself) will soon be open-sourced.
One Model Handles It All: Video, Image, Audio, Editing
H3's core selling point is not "generating video" — this field is already crowded with competitors. Its selling point is unification: text, images, videos, and audio are all understood, generated, edited, and referenced within the same model, eliminating the need to piece together multiple tools.
Specifically: you can give it a video and say "reference this camera movement," give it an image and say "have this character appear," give it an audio clip and say "add this sound" — H3 understands all inputs within the same context and outputs the result in one go. In the official demo, a single command "reference the Hitchcockian camera movement from Video 1, have the character from Image 2 sing the song from Audio 3" directly produces the final video.
Behind this is MiniMax's three-generation journey from Hailuo 01 → Hailuo 02 → H3: the first generation built the system, the second refined the components, and the third eliminated task boundaries. No longer distinguishing between T2I(text-to-image; V stands for video, A for audio), T2V, T2A, editing, referencing — all are unified into a single architecture, with tasks described using natural language.
0.8 Yuan/Second: A 15-Second Ad Clip Costs 12 Yuan
Pricing is the most direct language. H3's 2K video generation price is 0.8 yuan per second. The official claim is "less than one-third the cost of similar flagship video models" — multiple media outlets (Daily Economic News, East Money) have cross-confirmed this figure.
Convert it: a 15-second 2K ad clip costs 12 yuan. 768p is even cheaper, less than half the price of mainstream models' 720p. What does this price mean? It means the trial-and-error cost of "running a sample version to see the effect" is almost zero — previously, you might have only generated videos after finalizing the plan because "generating once is too expensive," but now you can validate ideas in the brainstorming stage.
The core technology is H3-VAE (next-generation video compressor): the compression ratio is greatly improved, and the effective sequence length has gained a 4x boost — the same video content requires only a quarter of the tokens(the units AI bill by; roughly 1–2 per word) as before. Fewer tokens mean lower inference costs(what it costs to actually run the AI). This is not a price cut for volume; it is a cost advantage at the architectural level.
Not Piling on Models, But Cutting Them Down
What is the industry norm? One expert model for each task: one for text-to-image, one for text-to-video, one for video editing, one for subject referencing, one for motion referencing, one for style transfer, one for voice acting, one for sound effects... adding up to more than a dozen models, each with its own API(the plug that lets your own app call it directly) and quirks.
H3 takes the opposite approach: cutting them all down, unifying into a single architecture. MiniMax's official statement is "task generalization is an irreversible trend, and architectural techniques should give way to how the model is defined." They even abandoned the Hailuo 02 architecture — despite its significant advantages — because it would introduce unnecessary complexity into "task generalization."
The result? One model handles all tasks, and far from being "a jack of all trades, master of none," it is actually the global number one in video editing. Why? Because cross-modal contextual understanding allows the model to "see the big picture": it is not generating a single frame in isolation but is planning the output after understanding the camera movement, character appearance, and audio rhythm.
Video Generation Transforms from "Toy" to "Productivity Tool"
Over the past two years, video generation models have been seen as "technology demos" — they look cool, but with low resolution, no sound, lack of control, and high cost, they cannot enter real production workflows. The release of H3 marks a turning point: video generation is beginning to meet the conditions for commercial-grade production.
Three signals: First, 2K + stereo sound means the output can be delivered directly to clients without the need for later sound enhancement or resolution boosting. Second, 0.8 yuan per second means the trial-and-error cost is nearly zero — you can validate ideas in the creative stage without waiting for the final plan. Third, the upcoming open-sourcing means companies can deploy the model privately, fine-tuning(retraining it on your own materials so it fits you) it with their own product images and brand materials without worrying about data leakage.
MiniMax's official target scenarios are clear: advertising, branding, e-commerce, product design, UI/UX, and gaming. Note, not "film" — the 15-second length and current image quality are not yet suitable for movies. But for creating an e-commerce product video, an app startup animation, or a set of social media content? It's enough.
Who Should Go Run a Sample Clip Now
H3's API is already online. If you belong to any of the following categories, you can go run a sample clip today to verify:
E-commerce/advertising materials: product photo → 15-second display video with background music and sound effects. Previously outsourced for 500 yuan and up, now 12 yuan for a version.
App/game UI demonstration: use V2V motion reference + brand color rendering to quickly produce product promotional animations.
Social media content: multi-shot native modeling, one command outputs multiple coherent scenes without stitching.
Enterprises requiring private deployment: wait for the weights to be released ("in the next few days"), deploy locally + fine-tune with your own materials, data does not leave the domain.
When a model can simultaneously handle video + image + audio, 2K at 0.8 yuan per second, and is about to be open-sourced, the idea of "piecing together multiple tools" should be upgraded. The first thing you can do today: take the most repetitive type of material you have and go run a sample clip on the H3 API. At a cost of 12 yuan, the data will make the decision for you.
Pricing and ranking data are from MiniMax's official blog and the Artificial Analysis third-party list, cross-confirmed by multiple media outlets such as Daily Economic News, East Money, and NetEase. Technical details (H3-VAE, Omni Transformer, In-Context Regeneration) are as described by the manufacturer, with the full technical report to be released soon. The open-source time is subject to MiniMax's official announcement.