MiniMax H3 · Acceleration

MiniMax H3 LoRA: v4 turbo at 6–8 steps and custom training

By Published Updated

Two different things get called a MiniMax H3 LoRA for Hailuo 3.0. One is the turbo LoRA: current v4 guidance starts at 6–8 sampling steps instead of the usual ~20. The other is training your own LoRA on a character, product or style. This page covers both, starting with the two supported v4 setup paths.

The turbo LoRA’s reputation problem is a single symptom: the audio comes back distorted at the 4-step minimum and everyone assumes the MiniMax H3 LoRA is broken. Four steps is a legacy or constrained case, not the current v4 quality default.

This downloadable community LoRA is not fal’s similarly named hosted model. The H3 Max and H3 Max Turbo comparison keeps those product names, prices and limits separate.

Check this first: audio distortion at the documented 4-step minimum can make an otherwise correct installation look broken, while the current v4 quality starting range is 6–8 steps.

The 4-step audio artifact

Symptom
Video looks usable, but the audio comes back distorted or blown out on a low-step run.
Cause
H3 advances video and audio on related but distinct schedules. Distortion usually means an old shared-schedule graph was used, or the original Larry LoRA and the pruned ComfyUI conversion were loaded with each other's sampler recipe. Exactly four steps also leaves the least recovery room.
Fix
Keep the paths intact: Larry's original v4 uses the MiniMax-H3 Turbo Sampler with simple at 6–8 steps; drbaph's pruned conversion uses the supplied current-ComfyUI workflow with Euler, beta and 6–8 steps, usually 8.

If audio is still damaged after matching the file, loader, sampler and scheduler, compare the same seed without the Turbo LoRA before changing sigma shifts. Four is a minimum for the original path, not the current quality recommendation.

Sources: Larry v4 model card · drbaph compatibility conversion

Installing the MiniMax H3 LoRA into the official workflow

The original full-model file and the pruned compatibility conversion have different loaders and sampler recipes. Pick one path first, then follow only the steps for that path.

  1. 01

    Choose one path before downloading: Larry's original full-model LoRA and custom node, or drbaph's _pruned_comfyui conversion and native ComfyUI workflow.

  2. 02

    For the original path, install Larryvrh/ComfyUI-MiniMax-H3-Turbo, put minimax_h3_turbo_v4_step600_ema.safetensors in ComfyUI/models/loras/, and load it with the MiniMax-H3 Turbo LoRA node.

  3. 03

    For the original path, feed SamplerCustomAdvanced with the MiniMax-H3 Turbo Sampler, use the simple scheduler, and start at 6–8 steps. Four is only the supported minimum.

  4. 04

    For a pruned/curve-form base, put minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors in the LoRA folder and open drbaph's supplied workflow: 6–8 steps, usually 8, Euler, beta, strength 1.0.

The two files are not interchangeable. Loading Larry's original through the stock LoRA node skips the custom compatibility logic; treating drbaph's partial pruned conversion as a full-model LoRA misstates what was converted.

Node repo
Original: Larryvrh/ComfyUI-MiniMax-H3-Turbo · conversion: stock ComfyUI
In ComfyUI-Manager
Original path: search "MiniMax-H3 Turbo" · conversion path: no custom Turbo node
LoRA folder
ComfyUI/models/loras/
Ready-made workflow
original: minimax_h3_t2v_turbo.json · pruned: fl_minimax_h3_turbo_lora_example_workflow.json

Sources: original model card · Larry custom node · pruned compatibility conversion

MiniMax H3 LoRA settings that actually matter

Sampling steps
Original v4: 6–8 · pruned conversion: 6–8 (8 recommended)
Four is Larry's documented minimum, not the current default. The legacy v1 ckpt850 remains a path-specific fallback for heavy or fast motion at exactly 4 steps.
Sampler + scheduler
Original: Turbo Sampler + simple · conversion: Euler + beta
Do not cross the recipes. Larry's original full LoRA is paired with the MiniMax-H3 Turbo Sampler from his custom node and a simple scheduler. drbaph's pruned conversion is tested through the native ComfyUI path with Euler and beta.
LoRA strength
1.0, usable 0.8–1.2
This is a dial, not a constant — the publisher documents which direction to move it for which artifact. See the strength note below.
Load path
Original: Larry custom node · conversion: native LoRA loader
The original file depends on Larry's loader and sampler logic. The file ending in _pruned_comfyui is the separate compatibility conversion for a standard ComfyUI LoRA loader; converting away incompatible AdaLN pairs is what makes that path possible.
Base checkpoint
Original: any H3 through custom node · conversion: pruned/curve-form H3
Larry's custom node adapts the original LoRA for full and pruned variants. drbaph's converted file is deliberately scoped to the pruned/curve-form architecture and should not be presented as the original full-model LoRA.
low_vram switch
Larry custom-node path only; off by default
Off applies the original LoRA at run time for the sharpest result. On merges it into the weights to reduce peak VRAM, with a possible softness trade-off on quantized or pruned bases. It is not a setting on drbaph's native compatibility workflow.

Strength is a symptom dial, not a quality slider

Blurry ghosting or smear
Nudge strength up, around 1.05 to 1.2.
Over-sharp grain or crunchy artifacts
Nudge strength down, around 0.8 to 0.95.

Both directions are documented by the publisher, which makes this the one setting worth changing before you start swapping checkpoints.

Source: larryvrh/MiniMax-H3-Turbo-Lora README

Which MiniMax H3 LoRA checkpoint to download

Six MiniMax H3 LoRA files circulate and the names are nearly identical, which is why people end up on a superseded one. The training step count is the distinguishing detail, and it has moved twice already.

FileTrainedNotes
minimax_h3_turbo_v4_step600_ema.safetensorsstart herev4 step 600, EMALarry's current original full-model LoRA. Use it through ComfyUI-MiniMax-H3-Turbo at 6–8 steps with the MiniMax-H3 Turbo Sampler and simple scheduler; 4 steps is the documented minimum, not the current quality default.
minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensorsstart herev4 step 600, EMA conversiondrbaph's current compatibility conversion for ComfyUI's pruned/curve-form H3 model. Use the native LoRA path at 6–8 steps, usually 8, with Euler, beta and strength 1.0.
minimax_h3_turbo_v4_step600.safetensorsv4 step 600Larry's non-EMA v4 comparison file. The EMA file above is the current general recommendation.
minimax_h3_turbo_4step_ema_ckpt850.safetensorsv1 ~850, EMALegacy v1 checkpoint. v4 replaces it for general use, although Larry notes that v1 ckpt850 can still be friendlier for heavy or fast motion when you insist on exactly 4 steps.
minimax_h3_turbo_4step_ema_ckpt500.safetensorsv1 ~500, EMALegacy intermediate EMA checkpoint, retained for comparison and superseded by v4 for general use.
minimax_h3_turbo_4step_ema.safetensorsv1 initial EMAInitial public EMA release. Keep only for historical comparison; it is not a current starting point.

Larry's original v4 file is the complete full-model LoRA used through his custom node. drbaph's pruned ComfyUI file is a compatibility conversion that excludes AdaLN pairs which do not map to the pruned/curve-form architecture. It is intentionally loadable by the native ComfyUI LoRA path, but it is not numerically identical to the original.

Sources: Larry original v4 · drbaph pruned conversion

v4 replaces the plastic-looking v1 default, but 4 steps is still the floor

Larry's current card recommends the v4 step-600 EMA checkpoint for general use and says it resolves the over-sharpened, plastic look of the earlier v1 line. Start at 6–8 steps. The old v1 ckpt850 remains a narrow exception for heavy or fast motion at exactly 4 steps, where v4 can smear more readily.

Source: larryvrh/MiniMax-H3-Turbo-Lora README

Sigma shift, and why it is a pair

Every MiniMax H3 LoRA run inherits this pair, so it is worth knowing even if you never touch it. The default is 12 / 3. Two shift values, one for video and one for audio. The model derives both video and audio timesteps from the current video sigma plus these two shifts.

Because the two streams share a sigma but not a schedule, the shift pair is what keeps them aligned. Changing it without understanding that relationship is how people end up with drifting audio.

In the dual-clock sampler these two values are fixed, not tunable — they are taken from the original implementation. If a workflow exposes them as sliders, leaving them alone is the correct move.

Sources: AI Toolkit H3 extension · dual-clock sampler

Acceleration beyond the MiniMax H3 LoRA

A separate approach from step reduction: skip work rather than take fewer steps. These stack with the MiniMax H3 LoRA but come with sharper constraints.

minimax-h3-velocity-cache-v1

Caches target audio and target video rows as one compact tensor to skip repeated work.

Constraint · Locked to 1344×768, 124 frames and shift 12·3. Any other parameter combination is rejected outright — this is a fixed-configuration accelerator, not a general speedup.

ComfyUI-Spectrum-MiniMax-H3

Forecasts post-transformer features with Chebyshev ridge regression so selected transformer evaluations can be skipped entirely.

Constraint · Adaptive scheduling with sampler-aware safeguards and native fallbacks. More flexible than a fixed cache, but it is prediction — expect it to trade some fidelity for the skipped compute.

What running the MiniMax H3 LoRA actually asks of your machine

Fewer steps is not less memory. A MiniMax H3 LoRA cuts sampling time; the base model it attaches to is the same size it always was.

Base model size
The base is a ~33B model, and the publisher notes an 80 GB GPU is comfortable at the largest resolutions. The ComfyUI node streams the base and offers the low_vram switch, which is why it still runs on far smaller cards.
Without a ComfyUI graph
Larry's repo ships a single generate.py that loads the base, applies the original LoRA, runs the custom dual-schedule sampler and muxes an mp4. It still needs a ComfyUI checkout for the model, VAE and text-encoder definitions; with current v4 weights, start at 6–8 steps rather than treating four as the default.
Trading VRAM for system RAM
In the standalone script, --offload-adaln moves roughly 13 GB from VRAM to CPU RAM. Lowering resolution or frame count is the other lever.
Validated sizes
Width and height are multiples of 32 with a 768px short edge, and frame counts follow the 17k+5 grid at 24fps. The publisher reports validating roughly 124 to 362 frames, which is about 5 to 15 seconds.

Source: Larry v4 model card

Training your own MiniMax H3 LoRA

The turbo LoRA changes how fast H3 samples. A MiniMax H3 LoRA you train yourself changes what it renders — a character, a product, a look. Ostris’ AI Toolkit added an H3 training extension on August 3, 2026, and the details below come from that extension’s own source rather than from a write-up of it.

What can be trained
AI Toolkit's minimax_h3 extension trains text-to-video and first-frame image-to-video LoRAs, with joint audio when the dataset provides it. Image-only datasets train as single latent frames, and sampling with one frame renders a still.
Which base it loads
By default it pulls the Comfy-Org repack — the pruned int8-ConvRot transformer and the NVFP4 AWQ Qwen3-VL text encoder, both kept quantized through training. A config switch picks fl2va, fl2va_pruned (the default), ref2va or ref2va_pruned.
Frame counts snap down, not up
Training clips are aligned to the previous valid 17k+5 frame count rather than the next one. This is the opposite direction from the ComfyUI workflow, which rounds your requested duration up — worth knowing if you are matching a dataset to a target length.
No CFG at validation
H3 is guidance-distilled and the released sampler runs without classifier-free guidance, so previews during training should be generated at guidance 1. Validating at a higher scale tells you nothing about how the LoRA will behave in the real workflow.
The audio stream trains on its own clock
The two streams ride different flow shifts — 12 for video, 3 for audio — and the audio sigma is derived from the video sigma at every step, in training exactly as in sampling.

Source: AI Toolkit H3 extension

We have not trained a MiniMax H3 LoRA ourselves yet, so there are no timings or loss curves here — only what the toolkit does and expects. Reference-driven training was not in the extension at the time of writing; text-to-video and first-frame image-to-video were.

Speed you don’t have to configure

A MiniMax H3 LoRA is worth setting up if you generate constantly on your own hardware. Lightloom’s hosted text-to-video generator skips the checkpoint hunt, sampler swap and local audio debugging.

Compare hosted pricing →

Before downloading anything, confirm your card can hold the base model — the VRAM requirements breakdown explains why the 32B text encoder, not the diffusion weights, is what usually decides whether a MiniMax H3 LoRA run fits. The local setup guide lists every file and folder, and the ComfyUI workflow guide covers the templates a MiniMax H3 LoRA plugs into, including the shipped sampler settings this one replaces.

Frequently asked questions

It cuts the usual roughly 20-step H3 workflow to a current 6–8-step starting range. Larry documents 4 steps as the minimum for the original v4 LoRA, not the quality default. Fewer steps reduce sampling work, but total wall-clock is not a fixed 5× because model loading, VAE decoding and audio work remain.

Yes. Ostris' AI Toolkit added a minimax_h3 extension on August 3, 2026 that trains text-to-video and first-frame image-to-video LoRAs, with joint audio when the dataset provides it. It loads the Comfy-Org repack by default and keeps the transformer and text encoder quantized through training, which is what makes it viable on consumer hardware.

Start at 1.0 and treat it as a symptom dial. If the result shows blurry ghosting or smear, raise it towards 1.05–1.2; if it shows over-sharp grain or crunchy artifacts, lower it towards 0.8–0.95. Both directions are documented by the publisher, so it is worth trying before you swap checkpoints.

Four steps leaves the least room for H3's separate video and audio schedules, and broken audio often means the original and converted LoRA recipes were crossed. Use Larry's original v4 with the MiniMax-H3 Turbo Sampler plus simple at 6–8 steps, or use drbaph's pruned conversion with its current-ComfyUI workflow, Euler, beta and 6–8 steps.

For Larry's original custom-node path, use minimax_h3_turbo_v4_step600_ema.safetensors at 6–8 steps with the MiniMax-H3 Turbo Sampler and simple. For a native ComfyUI pruned/curve-form path, use minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors at 6–8 steps, usually 8, with Euler, beta and strength 1.0.

12 and 3 — one for video, one for audio. The model derives both video and audio timesteps from the current video sigma plus these two shifts, which is what keeps the two streams aligned. Change them only if you understand that relationship.

Yes, through the matching path. Larry's original full-model LoRA uses his custom node, which adapts it for full and pruned H3 bases. A stock ComfyUI LoRA loader on a pruned/curve-form base should use drbaph's _pruned_comfyui conversion instead; that compatibility file excludes incompatible AdaLN pairs and is not numerically identical to the original.