MiniMax H3 Video Model
Send text, stills, footage, or voice to the minimax h3 video model API and collect a 2K clip with matched stereo audio
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

minimax h3 video model

Produce 2K clips with built-in stereo sound in the minimax h3 video model — it reads text, stills, clips, and audio to shape shots of up to 15 seconds.

All Tools

Discover our comprehensive AI-powered animation toolkit

Why Creators Choose the minimax h3 video model

Built by MiniMax as an open-weight omni-modal engine, the minimax h3 video model runs on fal.ai from day one. A single context absorbs text, stills, footage, and audio, then renders 2K clips with original stereo sound lasting up to 15 seconds. Expect targeted region edits, crisp on-screen text, and as many as 12 reference inputs per run.

  • Text, Stills, Clips, and Sound Together
    Feed as many as 9 stills, 3 video clips, and 3 audio tracks into one run — the minimax h3 video model keeps character identity, acting, camera language, and sound aligned in a single coherent output.
  • Sound Born With the Picture
    Each render ships with its own score, spoken lines, foley, and room tone locked to the cut. The minimax h3 video model can also carry over or clone a voice from a supplied recording.
  • Surgical, Region-Level Edits
    Swap a product, reletter a sign, redub a line, or flip daylight into night — the minimax h3 video model touches only the area you name, leaving the surrounding frame untouched.

Running the minimax h3 video model in Three Steps

Three quick steps take you from an API key to a finished 2K clip with matched audio.

Capabilities of the minimax h3 video model

Three endpoints, a shared multimodal context, original stereo sound, region-targeted edits, sharp typography, and usage-based billing — the minimax h3 video model covers the whole 2K production chain through fal.ai.

Three Endpoints, One Model

Route your idea through text-to-video, image-to-video with first- and last-frame control, or reference-to-video — the minimax h3 video model meets most workflows where they start.

A Dozen Reference Slots

Mix 9 stills, 3 clips, and 3 audio tracks, and the minimax h3 video model pulls face, acting, camera movement, framing, and cutting pace straight from your references.

Crisp Type and Live UI

Lay down legible captions, end cards, and logos, or bring real screens to life — landing pages, game menus, and HUDs — all animated by the minimax h3 video model.

Room for 7,000-Character Prompts

Drop an entire shot list into one box. The minimax h3 video model reads prompts up to 7,000 characters long, so you keep full control of the scene.

2K Output at 24fps

Clips land at 2K with a 1440px short edge, running as long as 15 seconds at 24fps, and you can pick from six aspect ratios or let the minimax h3 video model adapt.

Usage-Based API Pricing

Access the minimax h3 video model serverlessly and pay only for what you render — no minimum spend, no subscription, and commercial rights on the results.

FAQ

Questions About the minimax h3 video model

Straight answers to what people ask most about the minimax h3 video model on fal.ai.

1

What exactly is the minimax h3 video model?

It is an open-weight, omni-modal engine from MiniMax, available on fal.ai from launch day. Within one context it takes in text, stills, footage, and sound, and returns 2K video carrying its own stereo audio, up to 15 seconds long.

2

Which endpoints can I call?

Three are available: text-to-video, image-to-video with optional first- and last-frame control, and reference-to-video, which pins down subjects, styles, motion, camera work, and voices from your uploaded materials.

3

What output size and length can I get?

Expect 2K resolution — a 1440px short edge — at 24fps, with clips running 5 to 15 seconds. Frame shapes include 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an adaptive option.

4

Does audio come out of it?

Yes. Each render carries its own stereo track — music, dialogue, foley, and ambience matched to the cut — and the minimax h3 video model can transfer or clone a voice from a reference recording.

5

How many reference files are allowed?

Twelve in total — 9 images, 3 video clips of 2 to 15 seconds each, and 3 audio tracks of 2 to 15 seconds each. Note that any audio must be paired with at least one image or clip when you call the minimax h3 video model.

6

Can I use the results commercially?

Yes. Anything you produce through the fal.ai API with the minimax h3 video model can go into commercial work, subject to fal.ai's terms of service.

Put the minimax h3 video model to Work Today

Send one request to the minimax h3 video model and receive a 2K clip with original stereo sound — multimodal inputs, targeted edits, and pay-as-you-go pricing on fal.ai.