Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Render open-weight video with matching stereo sound using the comfyui minimax h3 workflow — start from a prompt, a photo, or a short clip.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Sora2 AI
Advanced AI Video Generator for High-Quality Videos

Suno AI Music Generator
Create Professional Music with AI
Meet the comfyui minimax h3 Workflow for Open-Weight Video
Built on MiniMax H3 — a general-purpose, omni-modal model released with open weights — the comfyui minimax h3 workflow brings that engine straight into ComfyUI. Text, pictures, footage, and sound are interpreted together in one shared context, so speech, effects, and music are produced alongside the picture in a single pass. Clips stretch to roughly 15 seconds at up to 2K and 24fps, and every node parameter stays under your control.
- Built-In Stereo SoundSpeech, effects, and background music are produced in the same pass as the picture, then delivered as one synchronized MP4 from the comfyui minimax h3 workflow.
- Open Weights, Total ControlKeep the comfyui minimax h3 model on your own machine and tune resolution, clip length, and diffusion settings however you like — nothing is capped by an API.
- Mix Text, Image, Video & AudioFeed prompts alongside picture, footage, or voice samples in a single run, pinning down a face, an art style, a movement, a camera path, or a specific voice through the comfyui minimax h3 nodes.
Running the comfyui minimax h3 Workflow in 3 Steps
Three quick steps take you from a fresh ComfyUI install to open-weight clips with built-in audio.
Capabilities Built Into the comfyui minimax h3 Workflow
From three ready-made ComfyUI templates to multimodal open-weight inference with stereo sound, reference-driven guidance, and optional Sage Attention acceleration — the comfyui minimax h3 workflow is a complete local video studio.
Three Ready-Made Templates
Text-to-video, image-to-video, and reference-to-video samples all ship inside the comfyui minimax h3 template library, so every generation mode works the moment you open it.
One Shared Omni-Modal Context
Prompts, stills, footage, and audio are all interpreted inside a single context by the comfyui minimax h3 model, letting every reference type feed the same run.
Guidance From Your References
Pin down a face, an aesthetic, a movement, a camera path, or a voice using source material — the comfyui minimax h3 R2V node accepts as many as 9 images, 3 videos, and 3 audio clips.
Crisp On-Screen Text and Logos
Words and brand marks come out sharp with the comfyui minimax h3 model, and you can spell out how references relate to one another in plain language.
Sage Attention Acceleration
Drop a Patch Sage Attention KJ node into the comfyui minimax h3 graph to roughly halve render time while barely touching output quality.
Resolution and Duration Controls
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, rounded to the model's 32-pixel grid, with duration counted in 17-frame blocks at 24fps.
comfyui minimax h3 — Frequently Asked Questions
Answers to the questions people ask most about running MiniMax H3 inside ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It is the official ComfyUI integration of MiniMax H3 — a general-purpose, omni-modal model that MiniMax shipped with open weights. One forward pass turns text, pictures, footage, or audio references into a video that already carries stereo sound.
How high can the output quality go?
Clips from the comfyui minimax h3 workflow top out around 2K at 24fps and run roughly 15 seconds. The native canvas keeps a 768px short edge, caps at 768x1344, and snaps dimensions to multiples of 32.
What generation modes come bundled?
Three samples ship in the comfyui minimax h3 template library: text-to-video (T2V), image-to-video (I2V) with optional first- and last-frame guidance, and reference-to-video (R2V) for pinning a character, style, motion, camera move, or voice.
Does the model produce its own audio?
It does. Voice, sound effects, and music are generated natively in stereo by the comfyui minimax h3 model, produced in the same pass as the picture and muxed into one synchronized MP4.
How do I get set up?
Install ComfyUI 0.30.0 or newer, go to Template Library > Video, load any comfyui minimax h3 workflow, and accept the pop-up that pulls the weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to render faster?
Add SageAttention plus the KJNodes custom pack, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 graph — render times drop by about half.
Your Next Video Starts with the comfyui minimax h3 Workflow
Keep MiniMax H3 on your own hardware, with open weights, stereo sound, and every knob exposed — text-to-video, image-to-video, and reference-to-video graphs are waiting for you.
