Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Try the comfyui minimax h3 workflow free: turn text, photos, or clips into 2K video with synced stereo audio — no install, full node control.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Veo3.1
Create Stunning Videos with Veo3.1
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Sora2 AI
Advanced AI Video Generator for High-Quality Videos

Suno AI Music Generator
Create Professional Music with AI
price
My Creations
What the comfyui minimax h3 Node Pack Brings to Your GPU
Built on MiniMax's omni-modal generation model and shipped as open weights, the comfyui minimax h3 workflow reads text, stills, footage, and audio as one combined context. Voice, effects, and background music are modeled alongside the picture, so a single MP4 leaves the graph already mixed in stereo. Clips run up to roughly fifteen seconds at 2K and 24fps, and every dial — resolution, length, sampling — stays exposed on the canvas.
- Stereo Sound Baked InSpeech, effects, and score arrive in the same MP4 as the picture, so the comfyui minimax h3 graph hands you a clip that is already mixed and in sync.
- Open Weights, Local RunsKeep the model on your own machine and tune resolution, clip length, and every sampling setting yourself, with no API ceiling in the way.
- Mixed Reference InputsFeed prompts, stills, footage, and audio into the same run to hold a face, a look, a movement, a camera path, or a voice steady across the generation.
Running the comfyui minimax h3 Graph: Three Steps
Go from a fresh install to open-weight video with synced sound in three short steps with the comfyui minimax h3 workflow.
Capabilities of the comfyui minimax h3 Node Pack
From three ready-made ComfyUI templates to open-weight multimodal generation, stereo audio, reference-driven control, and optional Sage Attention acceleration, the comfyui minimax h3 node pack covers a full local video pipeline.
Three Ready-Made Templates
ComfyUI ships the comfyui minimax h3 graphs in text-to-video, image-to-video, and reference-to-video flavours, so each generation mode works straight after loading.
One Shared Context
Text, stills, footage, and sound are read together by the model, letting every reference type shape a single run.
References That Hold the Details
Pin down a face, an aesthetic, a movement, a camera path, or a voice from supplied material — as many as 9 stills, 3 clips, and 3 audio files through the R2V node.
Clean Text and Brand Marks
On-screen words and logos come out crisp, and you can describe how references relate to each other in plain language.
Faster Renders with Sage Attention
Drop the Patch Sage Attention KJ node into the graph and render around twice as fast with almost no drop in fidelity.
Resolution and Duration Grid
The Resolution Selector derives width and height from your aspect ratio and megapixel target, then snaps them to the model's 32-multiple grid and 17-frame blocks at 24fps.
comfyui minimax h3 — Frequently Asked Questions
Everything people ask before running the comfyui minimax h3 model inside ComfyUI, gathered in one place.
What exactly is the comfyui minimax h3 workflow?
It is ComfyUI's built-in integration for MiniMax H3, an omni-modal generation model MiniMax released as open weights. The graph turns prompts, stills, footage, and audio references into video with stereo sound in one forward pass.
How high does the output resolution go?
Clips from the comfyui minimax h3 workflow reach 2K at 24fps for roughly 15 seconds. The native canvas keeps a 768px short edge, tops out at 768x1344, and rounds dimensions to a multiple of 32.
Which generation modes come bundled?
Three examples ship with the template set: text-to-video, image-to-video with optional first- and last-frame guidance, and reference-to-video, which pins a character, style, motion, camera move, or voice.
Is audio generated too?
It does. Speech, effects, and music are produced as native stereo by the comfyui minimax h3 model, generated in the same pass as the picture and delivered in one synced MP4.
What do I need to get started?
Bring ComfyUI to 0.30.0 or newer, head to Template Library > Video, select one of the graphs, and accept the prompt that pulls the weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Can generation be made faster?
Install SageAttention alongside the KJNodes custom nodes, then slot a Patch Sage Attention KJ node between UNETLoader and BasicGuider — render times drop by about half.
Put the comfyui minimax h3 Workflow to Work
Keep MiniMax H3 on your own hardware with open weights, stereo sound, and every parameter in reach — text, image, and reference video graphs are waiting for your first prompt.
