Picture, dialogue and score, in one pass
One prompt in. Picture, voices and score out — together.
A 33B open-weights omni model that packs video and audio into a single sequence, so a five-to-fifteen second clip arrives at 768p with 32 kHz stereo already inside it, cuts where you asked for them, and mouths that match the words. Make it in the browser, or through an API that matches MiniMax field for field.
Three shots, one generation
asked for cuts at 2.00 s and 3.80 s; scene detection on this file finds them at 2.00 s and 3.79 s
Made on LibtvsAI
Every clip here came out of this service
Same API, same machines, and the prompt on each card is the one that was sent. The sound is in the file — turn one on.
- Text to video6.6s
Three shots, one generation
[Shot 1] Aerial over a snow ridge at sunrise, pushing in slowly. [Shot 2] At 00:02.000, cut to crampons biting into blue ice, low angle, close. [Shot 3] At 00:03.800, cut to the climber cresting the ridge against the sun. Soundscape: Wind over rock, crampons on ice, breathing close to the microphone. Non-diegetic music: A sparse drone, one sustained cello note.
asked for cuts at 2.00 s and 3.80 s; scene detection on this file finds them at 2.00 s and 3.79 s
- Resolution
- 1376×768
- source
- Generated on LibtvsAI
- Text to video5.2s
A line to camera, spoken
Vertical portrait framing. A young barista behind a wooden counter looks straight into the camera and speaks, warm morning light from a window to her left, an espresso machine steaming behind her. <d>[English] We roast this one ourselves. Come by before nine and it is still warm.</d> Soundscape: A steam wand hissing, cups on saucers, quiet room tone.
14.8% silent frames — the gaps between words — against 0% on every clip here without dialogue; L/R correlation 0.988, a voice sitting centred
- Resolution
- 768×1376
- source
- Generated on LibtvsAI
- First / last frame5.2s
Two stills, and the middle invented
One unbroken move: the camera lifts off the sunrise ridge, the clouds thin into black sky and starfield, and the frame resolves into the flooded white interior of a starship bridge. Soundscape: Wind thinning to vacuum, then interior room tone. Non-diegetic music: One held synth note carrying across the transition.
first frame 3.5/255 from the still supplied for it, last frame 3.6/255 from its own — and 86 and 88 against the other end, so both are held and the middle is new
- Resolution
- 1376×768
- source
- Generated on LibtvsAI
Capabilities
One context, every modality — and a number under every claim
Claims about generated media are cheap to write and hard to check, so each one here carries the measurement it came from. What the model card asserts and nobody here has verified is not on this page.
Picture and sound in one pass
Video rows and audio rows are packed into a single sequence and attend to each other directly — no cross-attention, no masking. Sync is not something we align afterwards; there is no afterwards. The usual chain of text-to-speech, lip-sync, scoring and mixing is four chances to drift, and none of them are here.
32 kHz stereo, 24 fps, generated together
Cut on a timestamp
Write `At 00:02.000, cut to…` in the prompt and the cut lands there. Several shots come out of one generation, already assembled, and the sound cuts with the picture rather than being laid under it.
asked for 2.00 s and 3.80 s, measured 2.00 s and 3.79 s
Pin the first frame, the last, or both
Supply an image and the clip starts there. Supply a last frame instead and it arrives there. Supply both and the model fills the middle. A last frame on its own is a real mode, not an afterthought — it is how you generate up to a shot you already have.
anchored end 3.9/255 from the still supplied; the far end 53, so it holds and then moves
One portrait fixes the person
A single reference photograph, and the same face turns up in scenes that share nothing else. Features, hair and bearing hold while the setting, clothing and light follow the new prompt. Add dialogue and that person speaks it.
one portrait reference, two unrelated scenes, same identity
Borrow a camera move
Give it a clip and the new shot moves the way that one did. The reference conditions motion, not content — you are lending it the camera, not the scene.
inter-frame motion 1.25 in the reference, 1.70 with it, 14.68 without
Diegetic sound and score, separately
The prompt distinguishes what the scene makes from what plays over it, and the model keeps them apart. The result is a stereo field rather than one signal copied to two channels.
L/R correlation 0.62 on a scene; 0.988 on a piece to camera, where the voice belongs centred
Three ways in
Text, keyframes, references — one request body
There is no mode switch. The mode is inferred from what you put in content, exactly as the hosted API does it, so a prompt structure and your own media stay a single call.
Text to video and audio
One prompt, one generation, and the whole thing comes out assembled: several shots with the cuts where you put them, room tone and effects from inside the scene, and a score over the top. Nothing is a second pass.
- A six-second short with three cuts, no editor
- An advert idea, shot and scored, before the pitch
- A piece to camera where the mouth actually matches
content: [text]
Start on your frame, end on your frame, or both
Give an image and the clip starts there. Give a last frame instead and it arrives there — a real mode, not an afterthought, and the way you generate up to a shot you already have. Give both and the model fills the middle.
- A product photo that starts moving
- A clip that lands on your logo or your last still
- Before and after, with the transition generated between them
role: first_frame / last_frame
Bring a face, a camera move, or a voice
References condition the generation without dictating the scene. A portrait holds the identity while the setting changes completely; a clip lends its camera movement rather than its content; a recording sets the voice. Up to nine images, three videos and three audio tracks in one request.
- The same character across a whole set of scenes
- A camera move you liked, applied to a new shot
- A presenter who sounds like the same presenter every time
role: reference_image / _video / _audio
How it works
Describe it, walk away, come back to a clip.
- 1
Describe
Shots, dialogue, the sound the scene makes and the music over it are four separate fields. Fill in as many as you like; only the shots are required.
- 2
Walk away
Generation is real work, so it does not block the page. Leave if you like — finished clips are waiting when you come back.
- 3
Preview and download
Play it in the browser with sound, then download the mp4. Every clip you have made stays in your library.
Or call it from your own code
The same generator behind an API whose request body matches platform.minimax.io field for field. Point an existing MiniMax client at our base URL and it works.
curl https://api.libtvs.io/v2/video_generation \
-H "Authorization: Bearer $LIBTVS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMax-H3",
"duration": 6,
"resolution": "768P",
"ratio": "16:9",
"content": [{ "type": "text", "text": "A sunrise ascent on an alpine ridge." }],
"shots": [
{ "at_seconds": 0, "description": "Aerial over a snow ridge, pushing in" },
{ "at_seconds": 2.0, "description": "Cut to crampons biting into blue ice" }
],
"audio": { "soundscape": "Wind over rock, crampons on ice" }
}'
{ "task_id": "424010985738629" }Questions we get asked
Where the honest answer is narrower than the model card, the honest answer is the one written here.
What is MiniMax-H3?
An open-weights 33B omni-modal diffusion transformer. Video rows and audio rows sit in one sequence and attend to each other directly, so a single generation produces the picture and the sound together rather than a video model handing off to a soundtrack.
Is the audio really generated with the video?
Yes — 32 kHz stereo, in the same forward pass, with no text-to-speech, lip-sync or mixing stage anywhere. Two checks we run: a piece to camera has 25.3% silent frames when there is dialogue and 0% when there is only ambience, so the mouth is being driven; and a scene measures 0.64 left/right correlation, which is a real stereo field rather than one channel copied twice.
How long can a clip be?
Five to fifteen seconds at 24 fps. The visual VAE decodes in blocks, so the frame count rounds up to the next value it can produce — ask for 10 s and you get 10.13 s. The overshoot is free.
What resolution, and which shapes?
768p on the short edge, in 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16 — and `adaptive`, which takes the shape from the image you supplied. Higher than 768p needs a module MiniMax has not released, and generating straight at 2K is outside what this checkpoint was trained on.
Can I control where the cuts land?
Write the timestamp and the cut goes there — up to twelve shots in one request, assembled, with the sound cutting alongside the picture because they were generated together. Asked for 2.00 s and 3.80 s, we measured 2.00 s and 3.71 s.
Can I use my own images, videos and voice?
Three ways. As keyframes — a first frame, a last frame, or both, with the model filling the middle. As references — up to nine images, three videos and three audio tracks, twelve files in total, conditioning identity, camera movement and voice without dictating the scene. Or both kinds of prompt structure at once: shots, dialogue and the two sound layers are separate fields.
Which languages does the dialogue support?
The model card claims eleven. English is the one we have measured on our own hardware, so English is the one we will stand behind. The language tag is written into the prompt verbatim rather than inferred, so the others are there to try.
Is there an API?
The request body matches `platform.minimax.io` field for field, so an existing MiniMax client works by changing the base URL. Every deviation is a restriction on a value rather than a change of shape, and all of them are listed in the docs. There is also an OpenAI-shaped surface for clients already driving a local deployment.
How long does a clip take?
Minutes, not hours — and it depends on the length, the shape and how busy the fleet is, so the console shows the elapsed time rather than a promise. Generation does not block the page: leave, come back, the clip is in your library.
Can I use what I make commercially?
Yes. You keep the output; we keep inputs for seven days and outputs for thirty, then delete them. Billing is per second of output you asked for, with no subscription and no minimum — the pricing page has the rates.
Start creating
Create an account and generate your first clip in a couple of minutes.