H3 video generator
with native audio
Create H3 video with an AI video maker that writes picture and sound in one pass — 4–15 second clips at 768P or 2K, with the exact credit cost shown before you submit.
Native audio — measured, not claimed
The audio is not a soundtrack laid on afterwards
MiniMax H3 writes picture and sound in one forward pass, so the two are aligned by construction rather than by editing. The clip below comes straight from this AI video generator — the waveform under it is measured from the delivered file, and the three prompt fields beside it are the exact text that produced it.
a medium shot frames a blacksmith in a dim forge
Coals breathe with a low roar under the constant hiss of a bellows.
A low sustained drone at a slow tempo
The prompt behind it
Forge settles- Pictureintegrated_multimodal_description
- [Shot 1] Live-action, cinematic, a medium shot frames a blacksmith in a dim forge, orange coal light across one side of her face and leather apron. The camera holds a static shot as she lifts a glowing bar from the fire with tongs, carries it to the anvil, and strikes it three times in an even rhythm, sparks scattering with each impact. [Shot 2] At 00:04.500, the shot cuts to a low close shot of the quench trough as she plunges the bar into the water; a violent burst of steam floods the frame and the glow dies from orange to grey, and the camera pulls out with small amplitude at slow speed as the steam thins and the forge settles back to low ember light.
- Soundoverall_soundscape
- Coals breathe with a low roar under the constant hiss of a bellows. Hammer strikes ring in a clear three-beat rhythm with a metallic decay after each blow, then the quench erupts into a sharp violent hiss that falls away into dripping water and settling steam.
- Scorenon_diegetic_music
- A low sustained drone at a slow tempo with a single struck metallic tone repeating at wide intervals, dropping out entirely at the quench.
The prompt asked for three hammer strikes in an even rhythm. Measured off the delivered file, they land 0.95 seconds apart at about 22 dB above the forge ambience, and the score drops out at the quench — which is what the third field asked for. The audio track on this page is the model’s own, copied without re-encoding.
Every field above is documented in the MiniMax H3 prompt guide, along with 28 copy-paste examples ready to drop into the generator, and the reason an empty audio field gives you a soundtrack the model picked rather than silence.
Generated here — not stock footage
Three kinds of shot. One native-audio model.
Paper craft, 3D characters, macro ASMR — every H3 video on this wall came from this AI video maker, with sound and picture created in one pass. These previews are stripped of their audio so the wall can play silently; the clip you generate keeps the sound it was made with.
- Stop-motion & craft
A paper seed unfolds into a cut-paper tulip, frame by frame
- 3D character shorts
An H3 video of a small repair robot tapping its reflection in a rooftop puddle
- ASMR & sound-led
Soda pours over ice in macro, carried entirely by its audio
Each card opens the generator with the exact prompt that produced the clip — the same three-field format documented in the prompt guide.
Who it’s for
Who is this AI video maker for
Every chip below is a real preset from the library — open one and it arrives in the generator already written, with its soundscape and score fields filled in. Nothing is submitted until you press Generate.

Short-form creators
The format lives or dies on sound, and a silent clip you score afterwards never lines up with what is on screen. Here the audio arrives with the picture.

Product and ecommerce teams
A product shot needs controlled light, a controlled camera move and foley you actually chose — not a stock whoosh dropped over the top.

Directors and previz
You are testing whether a shot reads before anyone books a crew, which means the cut, the camera amplitude and a line of dialogue all have to survive the same eight seconds.

Explainer and course makers
A narrated piece falls apart when the style drifts between shots or the voiceover fights the score, so both belong in the prompt rather than in the edit.
These are 11 of the 21 text-to-video presets in this generator. The prompt guide has the rest, grouped by the technique each one demonstrates — dialogue that survives a cut, a style contract that holds for a whole clip, and on-screen text that renders as letters rather than noise.
Three H3 video generation workflows
H3 video generation supports three official workflows. This AI video generator currently runs text-to-video; image-to-video and multimodal reference inputs will be added once secure upload handling ships. In every mode, the engine produces picture and stereo sound in a single pass.
Text to video
Describe a shot and the model renders it with synchronized stereo audio in one pass.
You supply
A prompt.Image to video
Give the generator a still and it animates it, keeping the subject and framing you started from.
You supply
A first frame, an optional last frame, and a prompt describing the motion between them.Reference to video
Lock a character, a style or a motion from material you supply, then generate a new shot that keeps it.
You supply
Up to 9 reference images, 3 reference videos and 3 standalone audio references.Reference material is what the model matches against — it is not pasted into the output frame by frame.
Make an H3 video in three steps
- Step 1
Draft the prompt
Write one shot with a subject, an observable action, camera movement and intended audio. Nothing is submitted until you press Generate.
- Step 2
Price the exact request
Choose 4–15 seconds, an aspect ratio and 768P or 2K. The form calculates the exact credit reservation before submission.
- Step 3
Generate and download
Sign in, purchase credits, submit an allowed prompt, follow the task status and download the completed clip. Failed tasks return their reserved credits automatically.
Prompt structure is where many first attempts go wrong. The prompt guide has the subject → action → camera → audio structure the model responds to, with copy-paste examples you can load into the generator.
Tools
The paid AI video maker has its own page — text-to-video today. Video to prompt works the other way round: hand it a reference clip and it writes the prompt that would recreate it, sound included. The rest stay roadmap items until their own routes ship.
Inside the H3 video model
The MiniMax H3 video model is omni-modal — picture and stereo sound come out of a single forward pass, which is why the audio lines up with the action instead of being stitched on afterwards. It is also why running the model locally takes far more than the quoted file sizes suggest.
The MiniMax H3 model page has the published specifications, an honest hosted-vs-local comparison, and all five technical guides — VRAM, ComfyUI, local setup, turbo LoRA and prompting.
Seeing clips finish before they can play through? Our plain-language H3 Max explainer separates fal’s faster hosted variant, the newer Turbo name and the “infinite video” demos built around them.
If the channel is what brought you here, the Infinite Slop viewer’s guide explains where to watch, how suggestions enter the public queue and why the scarce resource is airtime rather than another Generate button.
H3 video frequently asked questions
Both terms describe this H3 video tool. Lightloom's AI video maker is a paid hosted workflow built on MiniMax H3: you write one prompt and it renders picture and stereo audio together in a single pass, with the H3 video cost shown in credits before you submit.
No. H3 video generation runs on remote hardware, so you do not need a local GPU. Running MiniMax H3 locally is a different matter: the official guidance calls for four GPUs, and even quantized community builds start around 8.5GB of weights before overhead.
The underlying MiniMax H3 model ships three official workflows: text to video from a prompt alone, image to video with optional first and last frame control, and reference to video with up to 9 reference images. This H3 video generator currently runs text to video; the other two arrive once secure uploads ship.
The current international MiniMax H3 V2 create API does not list a watermark parameter. Lightloom therefore does not claim to disable a provider watermark and does not add its own; inspect each returned file and the current MiniMax terms before publishing.
Yes, and it is not optional. H3 video generation produces picture and stereo sound in a single pass, so leaving audio out of your prompt does not give you a silent clip — it gives you a soundtrack the model chose. Silence has to be requested explicitly.
The underlying model's weights are downloadable under the MiniMax H3 Community License, but hosted H3 video generation consumes paid compute. New accounts start with 150 welcome credits and a daily check-in adds 30 more; beyond these grants, packs start at $9.99 for 10,000 credits, and text-to-video costs 200 credits per second at 768P or 300 at 2K.
An H3 video can be about 15 seconds long at up to 2K and 24fps. For longer pieces, generate consecutive clips and edit them together rather than trying to stretch a single generation.
No. Lightloom is an independent guide and hosted workflow that uses the MiniMax H3 API; it is not affiliated with or endorsed by MiniMax. The official model announcement, API documentation and licensing terms live on MiniMax's own sites.
Commercial use can depend on the MiniMax terms, your input rights, the output itself and the platform where you publish it. Lightloom does not guarantee ownership or non-infringement; read the current source terms and review each clip before commercial use.