Powered by MiniMax H3

Video to prompt,
with the sound included

Paste a link or drop a clip. You get back the structured prompt that would recreate it — every subject, the framing, the timestamp of each cut, and the soundscape as a structure rather than a one-line tag.

Free with an account — 3 reads a dayDirect MP4 / MOV links and uploadsRendering a video is the paid step
  1. 01

    Point it at a clip

    Paste a direct MP4 / MOV link, or upload the file — audio included.

  2. 02

    It reads for 3–4 minutes

    Picture and sound both get read. Close the tab if you like; it resumes.

  3. 03

    Leave with the prompt

    Copy the six sections, export the shot table, or generate from it here.

Sample readThe 8s clip behind this page, read back by MiniMax H3-Context-IR — unedited.
subject_definitions

<Subject 1> is the paper cutout Earth, featuring torn green paper continents over a blue paper ocean, set against a solid brown background in <Video 1>. <Subject 2> is the paper cutout landscape, featuring a brown sloped hill on the left and a darker blue body of water on the right, set against an orange-brown sky background in <Video 1>. <Audio 1> contains a calm female voiceover, a light acoustic musical track, and water sound effects, serving as the exact audio track for the target video.

overall_soundscape

The overall soundscape is defined by <Audio 1>, featuring a calm female voiceover in the first half and the crisp, natural sound effects of falling rain and rushing water that take over in the second half.

non_diegetic_music

<Audio 1> provides a light, cheerful acoustic background score featuring plucked strings and a steady, mid-tempo rhythm that plays continuously beneath the narration and sound effects.

See the full read below, compared line by line with the original prompt

A round trip we ran on ourselves

We wrote a prompt, generated this clip with MiniMax H3, then fed the clip back in and asked for the prompt. Because we still have the original, the two can be compared line by line — which is the one thing a video to prompt tool working with someone else’s footage can never show you.

8s, generated with MiniMax H3 from the prompt on the right. The narration, score and rain are the model’s own audio output.

What we wrote

integrated_multimodal_description: [Shot 1] Paper collage animation, torn-edge cut-paper shapes on a warm textured background, a centred flat composition frames a cut-paper Earth made of layered blue and green scraps. The camera holds a static shot as a paper cloud slides in from the left and a torn yellow sun rotates into the top corner in stepped jumps. A warm, measured female narrator (S1) says in an off-screen voiceover: <d>[English] Every drop of rain has already been somewhere else.</d> while no character appears on screen. [Shot 2] At 00:04.000, the shot cuts to a side-on cut-paper cross-section of a hillside as paper raindrops fall in stepped increments, collect into a torn blue river shape that grows downward across the frame, and flow into a paper sea while the camera pans right with small amplitude at slow speed to follow the water.

What came back, in full
subject_definitions:
<Subject 1> is the paper cutout Earth, featuring torn green paper continents over a blue paper ocean, set against a solid brown background in <Video 1>.
<Subject 2> is the paper cutout landscape, featuring a brown sloped hill on the left and a darker blue body of water on the right, set against an orange-brown sky background in <Video 1>.
<Audio 1> contains a calm female voiceover, a light acoustic musical track, and water sound effects, serving as the exact audio track for the target video.

summary:
[reference generation + audio reuse] The target video is a stop-motion paper cutout animation illustrating the water cycle. It opens with a view of <Subject 1>, the Earth, as paper clouds and a sun smoothly enter the frame. The video then cuts to <Subject 2>, a sloped landscape where raindrops fall and quickly form a stream that flows into a larger body of water. <Audio 1> provides a continuous acoustic music track, a female voiceover discussing rain, and the clear sound effects of flowing water.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the paper cutout Earth, its continents, and its oceans are faithfully preserved exactly as seen in the reference.
<Subject 2> (appears in [Shot 2]): fully_preserved - the paper cutout brown landscape, sloped hill, and blue water body are preserved exactly as seen in the reference.
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track without any additional layers.

detailed_description:
The target video is in a stop-motion paper cutout animation style, featuring flat, beautifully arranged layers of vibrantly colored paper with authentically torn and textured edges. The lighting mimics a clean studio setup with a soft overhead light that casts subtle, realistic drop shadows beneath each overlapping paper element, enhancing the three-dimensional, handcrafted aesthetic.

[Shot 1] The video opens with a static shot of <Subject 1>, the paper cutout Earth, positioned perfectly in the center against a solid, warm brown background. The Earth itself is constructed from a circular light blue paper base representing the oceans, overlaid with meticulously torn green paper shapes forming the recognizable continents of the Americas, Europe, and Africa. The off-screen narrator voice (S1), a calm and clear female voice, speaks at a gentle pace, <d>[English] Every drop of rain has already been somewhere else.</d> As she begins her line, crisp white cutout clouds smoothly slide into the frame from the left and right edges, slightly overlapping the sides of the Earth. Simultaneously, a bright, textured yellow cutout sun smoothly rotates into view at the extreme top-left corner, bringing a cheerful warmth to the composition. The sun and clouds settle into their final resting positions around the Earth exactly as the spoken line concludes. The camera maintains a completely static composition throughout this opening sequence, allowing the viewer to fully appreciate the tactile, fibrous appearance of the paper layers and the gentle acoustic background music that plays underneath the voiceover.

[Shot 2] At 00:03.500, the shot cuts to <Subject 2>, the paper cutout landscape. This scene features a rugged, brown sloped hill descending dramatically from the left side of the frame down to the right, meeting a darker blue, unevenly cut body of water. The background sky is a solid orange-brown paper layer. The camera remains a perfectly static shot, maintaining the exact same scale and stop-motion style. Light blue, teardrop-shaped paper raindrops begin to fall continuously from the top edge of the frame, landing squarely onto the upper left section of the brown hill. As the first raindrops hit the ground, the narrator's voice has completely stopped, leaving only the continuous acoustic music and the newly introduced, crisp sound effects of falling rain and rushing water. Reacting to the rainfall, a jagged, continuous strip of light blue paper rapidly materializes, representing a newly formed stream. This stream smoothly but swiftly flows down the brown slope, closely following the uneven, stepped contour of the paper hill. It snakes its way downward and seamlessly merges into the darker blue body of water on the right. The stream continuously cascades down the slope, simulating the energetic flow of water through rapid stop-motion replacement animation. The tactile paper textures, torn edges, and consistent drop shadows remain brilliantly stable, grounding the entire sequence in its charming physical medium until the video smoothly concludes.

overall_soundscape:
The overall soundscape is defined by <Audio 1>, featuring a calm female voiceover in the first half and the crisp, natural sound effects of falling rain and rushing water that take over in the second half.

non_diegetic_music:
<Audio 1> provides a light, cheerful acoustic background score featuring plucked strings and a steady, mid-tempo rhythm that plays continuously beneath the narration and sound effects.

Unedited output from MiniMax H3-Context-IR, 2026-08-19.

What survived the round trip, and what did not

4kept

Read back from the footage accurately

1drifted

Close, but off by a measurable amount

1lost

In the source, smoothed out of the read

  • kept

    The narration, word for word

    The line "Every drop of rain has already been somewhere else." came back verbatim, wrapped in the same dialogue tag and language marker the prompt format uses.

  • kept

    Paper-craft style, torn edges, drop shadows

    Recovered as stop-motion paper cutout with torn, textured edges and soft drop shadows — more material detail than the original prompt specified.

  • kept

    Two clouds, and the sun in the top-left

    The original prompt wrote one cloud sliding in from the left. The footage shows clouds entering from both edges — and the reader described the footage, not our prompt.

  • kept

    The score: plucked strings at a steady tempo

    We asked for plucked ukulele with brushed percussion; the reader heard "plucked strings" at a mid tempo. Generalised, but this clip has a score and it said so — on a silent clip it returns N/A instead.

  • drifted

    The cut lands between 3.90s and 3.95s

    Recovered as 00:03.500 — about 0.4s early. Same direction and size as our first round trip on a different clip: close enough to edit against, not to trust blindly.

  • lost

    The stepped, per-drop paper taps

    The original soundscape marks each raindrop with a dry stepped tap. Recovered as natural "falling rain and rushing water" — the paper-foley character got smoothed out.

Measured once, on 2026-08-19, on the 8s clip above: 250 seconds end to end. The yardstick comes from the file, not from the prompt we wrote — the cut point was frame-checked, the narration span read off the audio envelope. Quoting our own instructions back as if they were results would be marking our own homework.

Six sections, not one paragraph

The output follows the shape MiniMax H3 actually reads, so nothing has to be reworded before you can use it. If you want to write these fields by hand instead, the MiniMax H3 prompt guide covers each one with worked examples.

subject_definitions
Who and what is in the shot
Every person, object and location gets a stable label, so later sections can refer back to them instead of describing them again.
summary
The clip in two sentences
What happens, in order. The part you skim before deciding whether the rest is worth reading.
retention_analysis
What carries over, what does not
States, per subject, whether it is preserved exactly or only loosely — including whether the audio track is reused as-is.
detailed_description
Shot by shot, with the cut times
Framing, camera movement, lighting, action and the timestamp of every cut. The longest section, and the one that does the work.
overall_soundscape
The sound, as a structure
Not a one-line sound tag — the ambient bed, the events on top of it, and the order they arrive in.
non_diegetic_music
Score, if there is one
Left as N/A when the clip has no music. It does not invent a soundtrack that was never there.

Under the hood of a video to prompt read

There is no frame-sampling shortcut here: the clip goes to MiniMax’s H3-Context-IR endpoint whole, picture and audio track together, and comes back in the model’s own six-section format. That is also why a video to prompt read takes minutes rather than seconds — the numbers below are measured on our own runs, not estimates.

8s
The clip we fed in — our own footage, original prompt on file
250s
End to end on the task record, submit to done
6
Sections in the output, in the order H3 reads them
2
Round trips measured so far, misses published both times

What video to prompt is actually for

A video to prompt read is a working asset, not a party trick. These are the four jobs people bring clips here for — each one ends either in the editor or in the generator.

  • Recreate a clip you admire

    Paste a reference and get the prompt that would rebuild it — subjects, framing, cut timing, sound. Then change the parts you want changed and generate your own take.

  • See why a video works

    The shot table lays out every cut with its timing, framing and sound event. Study a reference shot by shot instead of scrubbing back and forth guessing.

  • Turn footage into a reusable prompt

    A read is not a one-off answer — it is a prompt you keep. Edit one field, render again, and the same reference becomes a family of variations.

  • Recover a prompt you lost

    Generated something good and lost the text? Feed the render back in. That round trip is exactly the experiment above — you can see how much survives.

Against a typical video to prompt generator

We looked at every tool ranking for this query before building this page. The differences below are structural, not benchmarks — they follow from one fact: the other tools stop at the prompt, and the prompt is the half of the job that was never the point.

Where the prompt runsCopied out to Sora, Veo or Kling — the result happens elsewhereGenerated right here, with the clip references intact
Shape of the outputGeneral-purpose prose, tuned for no model in particularThe six-field structure MiniMax H3 natively reads
How sound is handledA caption-length sound note per sceneA structured soundscape the model acts on — H3 writes audio in the same pass
How accuracy is shownClaimed in the marketing copyMeasured on our own clip — kept, lost and drifted, misses included

Video to prompt is not captioning

Three different jobs hide behind “turn my video into text”, and they are not interchangeable. Only one of them produces something a video model can act on.

Video captioning

One sentence naming what the clip shows

Search, accessibility, alt text. It names the clip; it cannot rebuild it.

Transcription

The spoken words, with timestamps

Subtitles and quotes. Says nothing about the picture, the camera or the score.

Video to prompt

Subjects, framing, cut times, soundscape and score — structured

Generating. The output is written to go back into a video model, not to be read about.

How it works

  1. Point it at a clip

    Paste a direct link or upload the file. Two to fifteen seconds is the sweet spot, and the audio track matters as much as the picture — keep it in.

  2. Get the structured prompt

    Six sections, including the cut times and the full soundscape, plus a shot table you can export. It takes a couple of minutes — keep the tab open while it reads.

  3. Generate from it here

    The prompt refers back to the clip you gave it, so it is at its most useful in the generator on this site.

The prompt refers back to the clip you uploaded, so it is most complete where that clip is still on hand — that is, in the generator here. Rendering the video is the paid step; see what a second costs.

The video to prompt checklist

Five checks before you run a video to prompt read. Each one traces back to the API limits or to something our round trips actually measured — none of it is folklore.

  1. Keep the audio track in

    The soundscape and music sections are written from what the model hears. A muted clip still reads, but the sound half of the prompt comes back empty.

  2. Trim to the 2–15 second window

    That is the range the API accepts. Longer footage needs cutting first — send the stretch you actually want to recreate, not the whole video.

  3. Send a file or a direct link

    MP4 or MOV, up to 32MB here. A YouTube or TikTok page is not a direct video file and will not resolve yet.

  4. Leave spoken lines audible

    Dialogue survives the round trip best — our test clip's narration came back word for word, dialogue tags included.

  5. Frame-check timings before cutting against them

    In both of our round trips, the reported cut point landed about 0.4 seconds early. Treat timestamps as a draft to verify, not as ground truth.

What you can upload

Formats
MP4 or MOV, H.264 / H.265 video, AAC / MP3 audio
Length
2 to 15 seconds per clip
Size
Up to 50 MB per file
Frame
256 to 5760 px, aspect ratio between 0.4 and 2.5

Limits published by MiniMax, not ours — MiniMax H3-Context-IR API reference. The model behind it is documented on our MiniMax H3 page.

Frequently asked questions

A video to prompt generator watches a clip and writes the text prompt that would recreate it in an AI video model. This one reads the picture and the audio track together and returns MiniMax H3's own six-section format — subjects, a shot-by-shot description with cut times, the soundscape and the score — so the result can be generated on this site without rewording.

Reading a clip and writing the prompt is free once you sign in — it is a far lighter job than rendering video, so we can afford to give it away. Generating a video from that prompt is the paid part, priced per second in credits and shown before you submit. Sign-in is required because any anonymous free tool gets drained by scripts within a day.

On our own round-trip test it kept the subject, clothing, set, props, framing and every audio event, including details the original prompt never specified. Two things slipped: repeated actions lose their count — three hammer strikes came back as a steady rhythm — and cut timings drift by a few tenths of a second. Close enough to edit against, not close enough to trust blindly.

Partly. The prompt refers back to the clip you uploaded, so the reference and audio-reuse instructions only resolve on a model that accepts that clip alongside the text. The shot-by-shot description reads fine on its own and transfers to other tools. The complete version, references intact, is the one you generate from here.

MP4 or MOV, H.264 or H.265 video with AAC or MP3 audio, two to fifteen seconds, between 256 and 5760 pixels on a side. The model accepts files up to 50MB; uploads on this page are capped at 32MB. Keep the audio track in — the soundscape section is written from what the model hears, and a silent upload gives you a weaker prompt.

Time and sound. An image to prompt tool reads one still — subject, composition, light. Video to prompt adds everything that happens across frames: the action, the camera movement, the timestamp of each cut, and the audio track, which a still simply does not have. If your reference is a single frame, an image tool is enough; if the pacing or the sound is the point, you want the video read.

Not yet. A platform page is not a direct video file, so those links need a download step we have not shipped. Today the tool takes direct MP4 or MOV links and file uploads; if your clip lives on a platform, download it and drop the file in. Platform links are on the roadmap.

A couple of minutes for a short clip — our measured run on eight seconds took just under three. Other tools advertise thirty seconds; we would rather quote the number we actually measured. Keep the tab open while it reads; your last few reads stay available on this page afterwards.

Yes, and separately. One section covers the soundscape as a structure — the ambient bed, the events on top of it and the order they arrive in — and another covers non-diegetic music, left as N/A when the clip has none. This matters because MiniMax H3 generates audio in the same pass as the picture, so a prompt that describes sound is a prompt the model can act on.