Examples
Configure parameters and click Generate
Cinematic Motion
A cinematic shot with smooth camera movement and realistic lighting.
Dynamic Product Scene
A product-focused video with polished motion and detailed textures.
Creative Story Beat
A short visual story with expressive motion and atmospheric composition.
Pro Narrative Shot
A refined cinematic sequence with stable framing and rich scene detail.
Audio-Ready Scene
A polished video scene designed for expressive motion and audio support.
Studio Composition
A controlled scene with deliberate movement and professional lighting.
MiniMax H3 — The Multimodal Video Model, Explained
MiniMax H3 is a universal multimodal video generation model from MiniMax, released through the MiniMax API and the Hailuo App on July 31, 2026, where it is also called Hailuo 3.0. It takes text, images, video clips and audio as input and produces video with native stereo audio — up to 2K resolution, 24fps, and 4 to 15 seconds per clip.
Everything you need to know about MiniMax H3, in one place.
MiniMax H3 text-to-video is now live in the generator above — write a prompt, pick 768P or 2K, choose a length from 4 to 15 seconds and one of six aspect ratios. Image-to-video and R2V multimodal reference generation are not in Studio yet; the image-to-video tab currently runs on Hailuo 02.
What Is MiniMax H3?
MiniMax H3 is a universal multimodal video generation model developed by MiniMax. Unlike models that accept only a prompt or only a single image, MiniMax H3 unifies text, image, video and audio input inside one model, and outputs a finished clip that already carries its own stereo soundtrack. The consumer Hailuo App ships the same model under the names Hailuo 3.0, Hailuo H3 and Hailuo 03 — they all refer to MiniMax H3.
Unified multimodal input
Text, images, video clips and audio all go into the same model. That is the core difference from image-to-video-only tools: a single generation can be conditioned on a written prompt, reference images, reference footage and reference audio at once.
Native stereo audio output
MiniMax H3 generates video with stereo audio produced by the model itself, rather than requiring a separate sound pass afterwards. Dubbing is supported in 11 languages.
2K, 24fps, 4–15 seconds
Output goes up to 2K resolution at 24fps, with clip length set in whole seconds anywhere from 4 to 15 seconds. The model is roughly 33B parameters.
From Hailuo 2.3 to MiniMax H3
MiniMax H3 is the current generation of MiniMax's video model line, following Hailuo 2.3. The jump between them is not just quality — H3 changes what the model accepts as input and what it produces as output.
Hailuo 2.3
Previous VersionThe previous generation of the Hailuo video model line. It handled text-to-video and image-to-video generation, and remains the version referenced by most existing Hailuo tutorials and workflows published before mid-2026.
MiniMax H3
Current VersionThe current release, launched on the MiniMax API and the Hailuo App on July 31, 2026. H3 unifies text, image, video and audio input in one model, outputs native stereo audio, and reaches 2K at 24fps with 4–15 second clips. Open weights for H3-Base followed on August 3, 2026.
The Three MiniMax H3 Generation Modes
MiniMax H3 exposes three distinct ways to drive a generation. The third one, R2V, is where the multimodal design actually shows up — it is the mode that has no direct equivalent in single-input video models.
Text to Video
Describe the shot in words and MiniMax H3 generates the clip — subject, camera movement and scene, together with the native stereo audio track. No source image needed.
Generate with MiniMax H3Image to Video & First–Last Frame
Start from a single image and animate it, or supply both a first and a last frame and let MiniMax H3 generate the transition between them. Useful when you need the clip to land on a specific closing frame.
Coming soon to StudioReference to Video
Condition one generation on up to 9 reference images, up to 3 reference video clips (each 2–15 seconds, 15 seconds total across all clips) and up to 3 reference audio clips. This is how MiniMax H3 carries a character, a style, a motion pattern or a voice from your own material into a new shot.
Coming soon to StudioMiniMax H3 Pricing on Wowovid
Wowovid runs on a single credit balance that works across every video and image model on the platform — there is no separate MiniMax subscription to buy and no per-model plan to manage. A 4-second MiniMax H3 text-to-video clip costs 93 credits at 768P or 151 credits at 2K, with the exact total shown before generation. Annual plans start at $10/month billed yearly; one-time credit packs start at $39.
MiniMax H3 vs. Other Video Models
The two comparisons people search for most when evaluating MiniMax H3.
MiniMax H3 vs. Kling AI
Kling AI is the stronger known quantity for polished text-to-video and image-to-video with physics-aware motion. MiniMax H3's differentiators are its unified multimodal input — reference images, video and audio in one generation — and native stereo audio in the output.
MiniMax H3 vs. Seedance 2.5
Both are current-generation video models that creators evaluate side by side. MiniMax H3's distinguishing features are R2V multimodal reference conditioning, native stereo audio, and an open-weight base model — Seedance 2.5 is a hosted model with a different feature set.
MiniMax H3 — Frequently Asked Questions
Common questions about MiniMax H3, its capabilities, its open-source status, and how it compares to other video models.
MiniMax H3 is a universal multimodal video generation model from MiniMax. It accepts text, image, video and audio input in a unified way and outputs video with native stereo audio, at up to 2K resolution and 24fps, with clip lengths from 4 to 15 seconds in whole-second steps. It supports dubbing in 11 languages and is roughly 33B parameters.
Yes. "MiniMax H3" is the official model name. The consumer-facing Hailuo App and much of the community refer to the same model as Hailuo 3.0, Hailuo H3 or Hailuo 03. They are all the same model, launched on the MiniMax API and the Hailuo App on July 31, 2026.
Partly, and the distinction matters. MiniMax released open weights for H3-Base on August 3, 2026, including the FL2VA and Ref2VA checkpoints, under the MiniMax H3 Community License. However, two components are not part of that release: H3-Context-IR, the prompt preprocessing and enhancement system, and H3-Regenerate-2K, the 2K upscale module. Both remain hosted-only, so a local deployment of the open weights does not reproduce the full hosted pipeline. See our MiniMax H3 ComfyUI page for the full technical picture.
T2V (text-to-video) generates from a written prompt alone. I2V (image-to-video) animates a single image, or generates a transition when you supply both a first and a last frame. R2V (reference-to-video) is the multimodal mode: it accepts up to 9 reference images, up to 3 reference video clips of 2–15 seconds each totalling no more than 15 seconds, and up to 3 reference audio clips within a single generation.
Yes. MiniMax H3 text-to-video is live in the Wowovid Studio generator: write a prompt, choose 768P or 2K, set a length anywhere from 4 to 15 seconds and pick one of six aspect ratios, from 21:9 through 9:16. Image-to-video and R2V multimodal reference generation are not yet available in Studio — the image-to-video tab currently runs on the earlier Hailuo 02 model.
Kling AI is a well-established choice for realistic motion and physics-aware image-to-video, and is the model Wowovid routes most video generations through today. MiniMax H3 stands out where you need multimodal reference conditioning — carrying a character, style, motion or voice from your own images, clips and audio into a new shot — and where you want the output to arrive with stereo audio already generated rather than added later.
Both sit in the current generation of AI video models. The features that separate MiniMax H3 are its R2V multimodal reference mode, its native stereo audio output with dubbing in 11 languages, and the fact that its base model weights were released openly. Seedance 2.5 is a hosted model with its own feature set and no open-weight release.
Not on Wowovid. Because the open weights exist, MiniMax H3 can be run locally — but that means a capable GPU, tens of gigabytes of model downloads and a working ComfyUI environment. Wowovid is a hosted platform: generation runs on our infrastructure and you only need a browser. If you are specifically researching local deployment, our MiniMax H3 ComfyUI page covers hardware requirements and quantized builds.
Wowovid uses one credit balance shared across every model on the platform, so there is no separate MiniMax H3 subscription. A 4-second clip costs 93 credits at 768P or 151 credits at 2K, and the generator shows the exact total before you submit. Annual plans start at $10/month billed yearly, monthly plans at $29/month, and one-time packs at $39.
Free previews may carry a watermark. Videos exported on a paid credit plan come out in HD without a watermark, so they can be used directly in client work, ads or social content.
Related Pages on Wowovid
MiniMax H3 in ComfyUI
Open weights, VRAM requirements, GGUF and quantized builds, and what running H3 locally actually involves.
Kling AI
The Kling AI 3.0 video model on Wowovid — image-to-video and text-to-video, no watermark on paid plans.
Image to Video
Turn a single photo into a moving clip with the video models available on Wowovid.
Pricing
Annual plans, monthly plans and one-time credit packs — one balance across every model.
Independent Platform
"MiniMax H3", "MiniMax" and "Hailuo" are trademarks of MiniMax and are used here for descriptive purposes only. Wowovid is an independent platform offering dozens of image and video generation models, and is not affiliated with, endorsed by, or sponsored by MiniMax.
Generate Video on Wowovid
MiniMax H3 text-to-video is live — write a prompt and start generating in your browser, with no GPU, no install and no model downloads. HD export with no watermark on paid plans.