Feedback
AI Ad Video Example
Loading...
minimax h3 video model
Produce 2K clips with built-in stereo sound in the minimax h3 video model — it reads text, stills, clips, and audio to shape shots of up to 15 seconds.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Veo3.1
Create Stunning Videos with Veo3.1
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Sora2 AI
Advanced AI Video Generator for High-Quality Videos

Suno AI Music Generator
Create Professional Music with AI
price
My Creations
Why Creators Choose the minimax h3 video model
Built by MiniMax as an open-weight omni-modal engine, the minimax h3 video model runs on fal.ai from day one. A single context absorbs text, stills, footage, and audio, then renders 2K clips with original stereo sound lasting up to 15 seconds. Expect targeted region edits, crisp on-screen text, and as many as 12 reference inputs per run.
- Text, Stills, Clips, and Sound TogetherFeed as many as 9 stills, 3 video clips, and 3 audio tracks into one run — the minimax h3 video model keeps character identity, acting, camera language, and sound aligned in a single coherent output.
- Sound Born With the PictureEach render ships with its own score, spoken lines, foley, and room tone locked to the cut. The minimax h3 video model can also carry over or clone a voice from a supplied recording.
- Surgical, Region-Level EditsSwap a product, reletter a sign, redub a line, or flip daylight into night — the minimax h3 video model touches only the area you name, leaving the surrounding frame untouched.
Running the minimax h3 video model in Three Steps
Three quick steps take you from an API key to a finished 2K clip with matched audio.
Capabilities of the minimax h3 video model
Three endpoints, a shared multimodal context, original stereo sound, region-targeted edits, sharp typography, and usage-based billing — the minimax h3 video model covers the whole 2K production chain through fal.ai.
Three Endpoints, One Model
Route your idea through text-to-video, image-to-video with first- and last-frame control, or reference-to-video — the minimax h3 video model meets most workflows where they start.
A Dozen Reference Slots
Mix 9 stills, 3 clips, and 3 audio tracks, and the minimax h3 video model pulls face, acting, camera movement, framing, and cutting pace straight from your references.
Crisp Type and Live UI
Lay down legible captions, end cards, and logos, or bring real screens to life — landing pages, game menus, and HUDs — all animated by the minimax h3 video model.
Room for 7,000-Character Prompts
Drop an entire shot list into one box. The minimax h3 video model reads prompts up to 7,000 characters long, so you keep full control of the scene.
2K Output at 24fps
Clips land at 2K with a 1440px short edge, running as long as 15 seconds at 24fps, and you can pick from six aspect ratios or let the minimax h3 video model adapt.
Usage-Based API Pricing
Access the minimax h3 video model serverlessly and pay only for what you render — no minimum spend, no subscription, and commercial rights on the results.
Questions About the minimax h3 video model
Straight answers to what people ask most about the minimax h3 video model on fal.ai.
What exactly is the minimax h3 video model?
It is an open-weight, omni-modal engine from MiniMax, available on fal.ai from launch day. Within one context it takes in text, stills, footage, and sound, and returns 2K video carrying its own stereo audio, up to 15 seconds long.
Which endpoints can I call?
Three are available: text-to-video, image-to-video with optional first- and last-frame control, and reference-to-video, which pins down subjects, styles, motion, camera work, and voices from your uploaded materials.
What output size and length can I get?
Expect 2K resolution — a 1440px short edge — at 24fps, with clips running 5 to 15 seconds. Frame shapes include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an adaptive option.
Does audio come out of it?
Yes. Each render carries its own stereo track — music, dialogue, foley, and ambience matched to the cut — and the minimax h3 video model can transfer or clone a voice from a reference recording.
How many reference files are allowed?
Twelve in total — 9 images, 3 video clips of 2 to 15 seconds each, and 3 audio tracks of 2 to 15 seconds each. Note that any audio must be paired with at least one image or clip when you call the minimax h3 video model.
Can I use the results commercially?
Yes. Anything you produce through the fal.ai API with the minimax h3 video model can go into commercial work, subject to fal.ai's terms of service.
Put the minimax h3 video model to Work Today
Send one request to the minimax h3 video model and receive a 2K clip with original stereo sound — multimodal inputs, targeted edits, and pay-as-you-go pricing on fal.ai.
