Independent field guide checked September 3, 2026
fal H3 Max: five seconds in under three
fal took MiniMax H3, trained it further, and paired it with an inference stack fast enough to render a five-second clip in under three. That changes video from a file you wait for into something a product can keep generating while you watch.
The speed figure is fal-reported. Lightloom has not reproduced it and does not currently host H3 Max.
The short answer
H3 Max changes the wait, not the clip.
What the name actually means
MiniMax built the base. fal built the fast variant.
MiniMax H3 — also branded Hailuo 3.0 — is the original model. H3 Max is fal’s further-trained, hosted version of that model, not H4 and not a separate MiniMax generation.
The open-weight base model. It keeps the widest feature range, including the higher-resolution and video-editing paths described by MiniMax.
Built by MiniMaxThe same model family, trained further by fal for instruction following and visual preference, then paired with fal's fast inference stack.
Built on H3 by falA public fal option with much lower listed rates. fal has not published enough technical detail to call it a new model generation.
Details still limitedText, image and reference inputs with synchronized sound.
Targets ordered prompt beats, visual preference and legible on-screen text.
The backend speed that creates the sub-real-time headline.
Method details: fal’s launch note · base model: MiniMax’s H3 release
Plain-language glossary
Six terms that make the launch sound harder than it is.
Read only the row you need. Every definition here is about the product, not API plumbing.
- Open weights
- MiniMax released the model files. H3 Max is not itself an open-weight download; it is a hosted fal variant built from H3.
- Post-training
- Extra training after the base model already works. fal says it added new data and reward-based training to make prompts land more reliably.
- Inference
- The model's render time on fal's backend. It does not include every second you spend uploading, queueing or downloading.
- Native audio
- Picture and sound are generated together, so dialogue, ambience and effects can be synchronized without a separate audio pass.
- Faster than real time
- A five-second clip can finish rendering in less than five seconds. fal reports under three seconds for a five-second 768P clip.
- Infinite video
- Not one endless model output. It is a product loop that renders the next short clip while the current one is still playing.
The benchmark reality
The speed case is stronger than the “best quality” case.
Two evaluations can both be real and still answer different questions. Here is the honest way to read the launch claim.
Won fal’s preference comparison
fal says H3 Max ranked first against 12 models across overall preference, prompt understanding and aesthetics. The test was designed and published by fal.
Read fal’s methodologyNear the top, not clearly alone
Artificial Analysis showed H3 Max at 1235, standard H3 at 1227, and the top three models inside overlapping uncertainty bands. An eight-point gap is not a quality landslide.
Open the live leaderboardThree ways in
Use the lightest input that preserves what matters.
More references can improve continuity, but they also add upload time, complexity and potentially a separate input charge.
Text to video
Use it when the scene can be described from scratch. This is the cleanest way to test prompt adherence.
One promptImage to video
Animate a start image and optionally supply an end frame when the destination matters as much as the opening.
Start frame + optional end frameReference to video
Guide the output with images, video or audio. fal accepts up to 12 combined reference files, but those inputs may add cost.
Up to 12 referencesA useful first test
Draft cheap. Prove the motion. Then raise resolution.
- Draft480P · 5 seconds
Check the subject, camera move and action order before spending more.
- RefineBalanced prompt expansion
fal says the balanced rewrite adds about one second; quality mode may add up to 30.
- Finish768P · 5–10 seconds
Use the higher setting only after the structure works.
Prompt anatomy
Write in the order the viewer should experience it.
Name the subject, action, camera, setting and sound. Put visible text in quotation marks and keep the beats chronological.
Learn the full H3 prompt grammar →Subject A tired night-shift baker closes the shop. Action She flips the sign to “SOLD OUT,” then smiles. Camera Slow handheld push-in. Sound Rain outside, soft bell, no music.
Modes and recommended settings: fal’s H3 Max guide · reference limits: reference endpoint
Pick the branch
H3 vs H3 Max vs H3 Max Turbo
The right answer changes with the constraint: resolution, iteration speed or cost.
Swipe to compare all three →
Open-weight base
fal post-trained variant
Newer hosted sibling
Up to 2K
480P or 768P
480P or 768P
5–15 seconds
5–15 seconds
5–15 seconds
Resolution or editing
Fast iteration + control
Lowest listed rate
Slower than Max
768P ceiling
Quality story is thin
fal price snapshot
Compare the route, not just the model name.
Checked Sep 3, 2026. Reference media may add input charges.
Turbo’s displayed launch rate runs through September 7; fal lists the later standard rate as $0.025/s at 480P and $0.04/s at 768P. Prices change, so check the endpoint before a large run.
Price and feature sources: fal’s H3 comparison · H3 Max endpoint · H3 Max Turbo endpoint
Why “Infinite Slop” appeared
Fast generation turns video into a live loop.
H3 Max does not emit one endless movie. It gives a product enough time to render clip B while clip A is still on screen.
The next clip can finish before the current one ends.
Audience input can influence a clip that has not been generated yet.
The product keeps selecting, rendering and appending short segments.
- 1Audience inputWhat should happen next?
- 2Carry contextPrompt, references or last frame
- 3Render aheadClip B generates while A plays
- 4Append and repeatThe stream stays alive
Product example: the creator’s Infinite Slop write-up · see how the public channel and queue work
The honest limits
What the launch headline leaves out.
Prompt expansion, queueing, reference uploads, network transfer and playback setup can make the click-to-result wait longer than three seconds.
If your final deliverable needs the 2K path or instruction-based video editing, standard H3 remains the more capable branch.
The first 4,096 reference tokens are free, then fal lists $0.02 per 1,000 tokens. A five-second reference video can cost more than the output itself.
Speed makes the loop possible. Your product still has to choose the next scene, carry context, recover from failed renders and join clips smoothly.
The one-line decision rule
Choose the constraint you cannot negotiate.
Best fit when 768P, synchronized audio and fast reference-driven iterations cover the job.
Keep the open-weight path, the higher-resolution route and instruction-based editing.
Run your own side-by-side quality check because its technical and benchmark story is still limited.
Make the first render count
Build the prompt before you spend the generation.
Drop a reference clip into Lightloom and turn its subject, camera, cut timing and soundscape into a structured prompt you can use on fal or anywhere else. Try it before signing in; create an account only when you want to save the analysis and revisit it later.
How this page earns trust
Facts, claims and unknowns stay separate.
- Primary facts
- Inputs, output limits, prices and model lineage come from dated fal and MiniMax pages linked at the point of use.
- Attributed claim
- The under-three-second timing and #1 launch ranking are fal’s claims. Lightloom has not reproduced either result.
- Independent check
- Artificial Analysis provides the external quality snapshot. Its live ranking can change as votes and models change.
- Open question
- fal has not published a detailed technical note explaining what changed inside H3 Max Turbo.
See how we date and correct claims in our editorial standards.