Powered by MiniMax H3

H3 video generator
with native audio

Create H3 video with an AI video maker that writes picture and sound in one pass — 4–15 second clips at 768P or 2K, with the exact credit cost shown before you submit.

  • MiniMax H3
  • 5s
  • 16:9
  • 768P
  • 1000 credits

Native audio — measured, not claimed

The audio is not a soundtrack laid on afterwards

MiniMax H3 writes picture and sound in one forward pass, so the two are aligned by construction rather than by editing. The clip below comes straight from this AI video generator — the waveform under it is measured from the delivered file, and the three prompt fields beside it are the exact text that produced it.

a medium shot frames a blacksmith in a dim forge

Coals breathe with a low roar under the constant hiss of a bellows.

A low sustained drone at a slow tempo

The prompt behind it

Forge settles
Pictureintegrated_multimodal_description
[Shot 1] Live-action, cinematic, a medium shot frames a blacksmith in a dim forge, orange coal light across one side of her face and leather apron. The camera holds a static shot as she lifts a glowing bar from the fire with tongs, carries it to the anvil, and strikes it three times in an even rhythm, sparks scattering with each impact. [Shot 2] At 00:04.500, the shot cuts to a low close shot of the quench trough as she plunges the bar into the water; a violent burst of steam floods the frame and the glow dies from orange to grey, and the camera pulls out with small amplitude at slow speed as the steam thins and the forge settles back to low ember light.
Soundoverall_soundscape
Coals breathe with a low roar under the constant hiss of a bellows. Hammer strikes ring in a clear three-beat rhythm with a metallic decay after each blow, then the quench erupts into a sharp violent hiss that falls away into dripping water and settling steam.
Scorenon_diegetic_music
A low sustained drone at a slow tempo with a single struck metallic tone repeating at wide intervals, dropping out entirely at the quench.
0s2s4s6s8s

The prompt asked for three hammer strikes in an even rhythm. Measured off the delivered file, they land 0.95 seconds apart at about 22 dB above the forge ambience, and the score drops out at the quench — which is what the third field asked for. The audio track on this page is the model’s own, copied without re-encoding.

Run it →

Every field above is documented in the MiniMax H3 prompt guide, along with 28 copy-paste examples ready to drop into the generator, and the reason an empty audio field gives you a soundtrack the model picked rather than silence.

Generated here — not stock footage

Three kinds of shot. One native-audio model.

Paper craft, 3D characters, macro ASMR — every H3 video on this wall came from this AI video maker, with sound and picture created in one pass. These previews are stripped of their audio so the wall can play silently; the clip you generate keeps the sound it was made with.

Get started

Each card opens the generator with the exact prompt that produced the clip — the same three-field format documented in the prompt guide.

Who it’s for

Who is this AI video maker for

Every chip below is a real preset from the library — open one and it arrives in the generator already written, with its soundscape and score fields filled in. Nothing is submitted until you press Generate.

These are 11 of the 21 text-to-video presets in this generator. The prompt guide has the rest, grouped by the technique each one demonstrates — dialogue that survives a cut, a style contract that holds for a whole clip, and on-screen text that renders as letters rather than noise.

Three H3 video generation workflows

H3 video generation supports three official workflows. This AI video generator currently runs text-to-video; image-to-video and multimodal reference inputs will be added once secure upload handling ships. In every mode, the engine produces picture and stereo sound in a single pass.

  • Text to video

    Describe a shot and the model renders it with synchronized stereo audio in one pass.

    You supply
    A prompt.

  • Image to video

    Give the generator a still and it animates it, keeping the subject and framing you started from.

    You supply
    A first frame, an optional last frame, and a prompt describing the motion between them.

  • Reference to video

    Lock a character, a style or a motion from material you supply, then generate a new shot that keeps it.

    You supply
    Up to 9 reference images, 3 reference videos and 3 standalone audio references.

    Reference material is what the model matches against — it is not pasted into the output frame by frame.

Make an H3 video in three steps

  1. Step 1

    Draft the prompt

    Write one shot with a subject, an observable action, camera movement and intended audio. Nothing is submitted until you press Generate.

  2. Step 2

    Price the exact request

    Choose 4–15 seconds, an aspect ratio and 768P or 2K. The form calculates the exact credit reservation before submission.

  3. Step 3

    Generate and download

    Sign in, purchase credits, submit an allowed prompt, follow the task status and download the completed clip. Failed tasks return their reserved credits automatically.

Prompt structure is where many first attempts go wrong. The prompt guide has the subject → action → camera → audio structure the model responds to, with copy-paste examples you can load into the generator.

Open the generator →

Tools

The paid AI video maker has its own page — text-to-video today. Video to prompt works the other way round: hand it a reference clip and it writes the prompt that would recreate it, sound included. The rest stay roadmap items until their own routes ship.

Inside the H3 video model

The MiniMax H3 video model is omni-modal — picture and stereo sound come out of a single forward pass, which is why the audio lines up with the action instead of being stitched on afterwards. It is also why running the model locally takes far more than the quoted file sizes suggest.

The MiniMax H3 model page has the published specifications, an honest hosted-vs-local comparison, and all five technical guides — VRAM, ComfyUI, local setup, turbo LoRA and prompting.

Seeing clips finish before they can play through? Our plain-language H3 Max explainer separates fal’s faster hosted variant, the newer Turbo name and the “infinite video” demos built around them.

If the channel is what brought you here, the Infinite Slop viewer’s guide explains where to watch, how suggestions enter the public queue and why the scarce resource is airtime rather than another Generate button.

H3 video frequently asked questions

Both terms describe this H3 video tool. Lightloom's AI video maker is a paid hosted workflow built on MiniMax H3: you write one prompt and it renders picture and stereo audio together in a single pass, with the H3 video cost shown in credits before you submit.

No. H3 video generation runs on remote hardware, so you do not need a local GPU. Running MiniMax H3 locally is a different matter: the official guidance calls for four GPUs, and even quantized community builds start around 8.5GB of weights before overhead.

The underlying MiniMax H3 model ships three official workflows: text to video from a prompt alone, image to video with optional first and last frame control, and reference to video with up to 9 reference images. This H3 video generator currently runs text to video; the other two arrive once secure uploads ship.

The current international MiniMax H3 V2 create API does not list a watermark parameter. Lightloom therefore does not claim to disable a provider watermark and does not add its own; inspect each returned file and the current MiniMax terms before publishing.

Yes, and it is not optional. H3 video generation produces picture and stereo sound in a single pass, so leaving audio out of your prompt does not give you a silent clip — it gives you a soundtrack the model chose. Silence has to be requested explicitly.

The underlying model's weights are downloadable under the MiniMax H3 Community License, but hosted H3 video generation consumes paid compute. New accounts start with 150 welcome credits and a daily check-in adds 30 more; beyond these grants, packs start at $9.99 for 10,000 credits, and text-to-video costs 200 credits per second at 768P or 300 at 2K.

An H3 video can be about 15 seconds long at up to 2K and 24fps. For longer pieces, generate consecutive clips and edit them together rather than trying to stretch a single generation.

No. Lightloom is an independent guide and hosted workflow that uses the MiniMax H3 API; it is not affiliated with or endorsed by MiniMax. The official model announcement, API documentation and licensing terms live on MiniMax's own sites.

Commercial use can depend on the MiniMax terms, your input rights, the output itself and the platform where you publish it. Lightloom does not guarantee ownership or non-infringement; read the current source terms and review each clip before commercial use.