Published ratesView pricing →

Hailuo 3.0 · released July 31, 2026

MiniMax H3 AI video generator — paid Hailuo 3.0 T2V

Create 4–15-second MiniMax H3 clips at 768P or 2K, with native audio. Every request has a published price before submission, with no free allowance or subscription. See exact credit pricing before generating.

Resolution
2K
Max length
15s
Audio
Native stereo
Hosted mode
Paid T2V

Model releasedMiniMax H3 · Hailuo 3.0

Open weights, released July 31, 2026

  • 2K
  • 15s
  • 24fps
  • Native stereo audio
See pricing →

The published paid T2V rate starts at 20 credits per output second.

What MiniMax H3 output looks like

These are MiniMax’s own launch demos, shown here for reference — we did not generate them. H3Video will only publish its own clips after generation launches and each result can be verified.

Three generation modes in MiniMax H3

MiniMax H3 supports three official workflows. H3Video is staging text-to-video first; image-to-video and multimodal reference inputs remain locked until secure upload handling ships. In every mode, picture and stereo sound are produced in a single pass.

  • Text to video

    Describe a shot and MiniMax H3 renders it with synchronized stereo audio in one pass.

    You supply
    A prompt.

  • Image to video

    Give MiniMax H3 a still and it animates it, keeping the subject and framing you started from.

    You supply
    A first frame, an optional last frame, and a prompt describing the motion between them.

  • Reference to video

    Lock a character, a style or a motion from material you supply, then generate a new shot that keeps it.

    You supply
    Up to 9 reference images, 3 reference videos and 3 standalone audio references.

    Reference material is what the model matches against — it is not pasted into the output frame by frame.

H3Video generation in three steps

  1. Step 1

    Draft the prompt

    Write one shot with a subject, an observable action, camera movement and intended audio. Nothing is submitted until you press Generate.

  2. Step 2

    Price the exact request

    Choose 4–15 seconds, an aspect ratio and 768P or 2K. The form calculates the exact credit reservation before submission.

  3. Step 3

    Generate and download

    Sign in, purchase credits, submit an allowed prompt, follow the task status and download the completed clip. Failed tasks return their reserved credits automatically.

Prompt structure is where many first attempts go wrong. The MiniMax H3 prompt guide has the subject → action → camera → audio structure the model responds to, with copy-paste examples you can load into the generator.

What MiniMax H3 does differently

MiniMax H3 is an omni-modal model: text, images, video and audio share one context, and picture and sound come out of a single forward pass rather than being stitched together afterwards. That is why the audio lines up with the action — and why leaving the audio block out of your prompt doesn’t give you silence, it gives you a soundtrack the model picked.

It is also why MiniMax H3 is heavier to run than a video-only model of the same nominal size. The audio branch is not free, and neither is the 9-image reference path that makes the third generation mode work.

MiniMax H3 specifications

Published limits, not measurements — this site has no first-party benchmark hardware, so nothing here is a timing claim.

MiniMax H3 published output specifications
Maximum resolution2K
Frame rate24fps
Maximum duration15 seconds per generation
Resolution cap768px short edge, capped at 768×1344
AudioNative stereo, generated with the picture
Reference inputs9 images, 3 videos, 3 audio clips
ReleasedJuly 31, 2026

Source: docs.comfy.org / MiniMax release notes.

Hosted MiniMax H3 vs. running it locally

MiniMax H3 weights are downloadable, so running it yourself is a real option — it is just a much larger commitment than the download page suggests.

The H3Video hosted workflow compared with running MiniMax H3 on your own GPU
H3Video hosted T2VMiniMax H3 on your GPU
HardwareRemote MiniMax H3 API execution; no local GPU required.MiniMax's own configuration calls for 4 GPUs. Community GGUF builds start at 8.49GB of weights before overhead.
SetupGoogle sign-in and a paid credit balance; no model download.ComfyUI 0.30.0 or later, plus four model files placed across three folders.
OutputText-to-video at 768P or 2K, 4–15 seconds, with native audio.Same model, same ceiling — if your card can hold it at that resolution.
Cost20 credits/second at 768P or 30 credits/second at 2K. No subscription and no free allowance.No per-clip cost, once you own hardware that can run it.
ControlText-to-video is available; image and multimodal reference inputs remain guide-only until secure uploads ship.Full node graph: samplers, sigma shift, LoRAs, quantization.

Why most people can’t run MiniMax H3 locally

Community GGUF builds of MiniMax H3 go as small as 8.49GB, and every roundup on the internet quotes that number as the requirement. It is not one. A weight file is not a VRAM figure — the VAE, text encoder, latents and the audio branch all allocate on top of it, and peak usage climbs with resolution and duration. The MiniMax H3 VRAM breakdown has every quant level against the card it actually fits, with the overhead nobody lists. If you do have the hardware, the MiniMax H3 ComfyUI workflow guide walks through all three official templates and what breaks on older versions.

If you do run MiniMax H3 yourself

Two things decide whether a local install is pleasant or painful. The first is that the text encoder is a 32B model — usually larger than the generation model people sized their card against, which the MiniMax H3 local setup guide covers file by file. The second is sampling time: the MiniMax H3 turbo LoRA cuts 20 steps down to 4, roughly a 5× reduction — with one audio artifact that catches almost everyone the first time.

What the hosted workflow changes

H3Video sends an allowed text prompt to the official MiniMax H3 V2 API, so local VRAM and ComfyUI versions are not part of the user workflow. The service reserves credits before submission, stores completed clips in its own Cloudflare bucket and returns the reservation if the task fails. H3Video is independent and is not affiliated with MiniMax.

MiniMax H3 guides

Running MiniMax H3 on your own hardware is where most people get stuck. These cover the parts the release notes skip — what actually allocates VRAM, why 4-step audio comes out broken, and which of the 15 model files you really need.

Tools on the roadmap

The paid text-to-video generator is available above. These separate helper tools remain roadmap items until their own routes ship.

  • Coming soonImage to videoTurn a still into a moving clip with synced audio.
  • Coming soonVideo to promptUpload a clip, get the prompt that would recreate it.
  • Coming soonScript to videoBreak a script into shots and render them in sequence.

Frequently asked questions

What is MiniMax H3?

MiniMax H3 (also called Hailuo 3.0) is an omni-modal generation model released on July 31, 2026. It understands text, images, video and audio in a single context and generates video with native stereo sound at up to 2K resolution, 24fps, and about 15 seconds per clip.

Can I run MiniMax H3 without a GPU?

Yes. H3Video runs paid text-to-video generation on remote hardware, so you do not need a local GPU. Running H3 locally requires substantial GPU memory: MiniMax's own guidance calls for four GPUs, and even quantized community builds start around 8.5GB of weights before overhead.

Does MiniMax H3 support text to video and image to video?

Both, plus a third mode. MiniMax H3 ships three official workflows: text to video from a prompt alone, image to video with optional first and last frame control, and reference to video, which accepts up to 9 reference images, 3 reference videos and 3 standalone audio references.

Do MiniMax H3 videos have a watermark?

H3Video uses the official API with its optional AIGC watermark setting disabled. You should still inspect every delivered clip and review current provider and platform rules before publishing it.

Does MiniMax H3 generate audio with the video?

Yes, and it is not optional. MiniMax H3 produces picture and stereo sound in a single pass, so leaving the audio out of your prompt does not give you a silent clip — it gives you a soundtrack the model chose. Silence has to be requested explicitly.

What resolution does MiniMax H3 output?

Up to 2K at 24fps, with the short edge capped at 768px and an overall cap of 768×1344. Requests above that get clamped back down, so pick an aspect ratio inside the cap and upscale afterwards if you need more.

Is MiniMax H3 free to use?

The model weights are downloadable under the MiniMax H3 Community License, but hosted generation consumes paid compute. H3Video has no free allowance: packs start at $9.99 for 1,000 credits, and text-to-video costs 20 credits per second at 768P or 30 at 2K.

How long can a MiniMax H3 video be?

About 15 seconds per generation at up to 2K and 24fps. For longer pieces, generate consecutive clips and edit them together rather than trying to stretch a single generation.

Is this the official MiniMax H3 site?

No. H3Video is an independent guide and hosted workflow that uses the MiniMax H3 API; it is not affiliated with or endorsed by MiniMax. The official model announcement, API documentation and licensing terms live on MiniMax's own sites.

Can I use MiniMax H3 videos commercially?

Commercial use can depend on the MiniMax terms, your input rights, the output itself and the platform where you publish it. H3Video does not guarantee ownership or non-infringement; read the current source terms and review each clip before commercial use.