comfyui minimax h3 Video Generator
Type a prompt or drop in a reference — the comfyui minimax h3 workflow handles the rest, audio included
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Render open-weight video with matching stereo sound using the comfyui minimax h3 workflow — start from a prompt, a photo, or a short clip.

All Tools

Discover our comprehensive AI-powered animation toolkit

Meet the comfyui minimax h3 Workflow for Open-Weight Video

Built on MiniMax H3 — a general-purpose, omni-modal model released with open weights — the comfyui minimax h3 workflow brings that engine straight into ComfyUI. Text, pictures, footage, and sound are interpreted together in one shared context, so speech, effects, and music are produced alongside the picture in a single pass. Clips stretch to roughly 15 seconds at up to 2K and 24fps, and every node parameter stays under your control.

  • Built-In Stereo Sound
    Speech, effects, and background music are produced in the same pass as the picture, then delivered as one synchronized MP4 from the comfyui minimax h3 workflow.
  • Open Weights, Total Control
    Keep the comfyui minimax h3 model on your own machine and tune resolution, clip length, and diffusion settings however you like — nothing is capped by an API.
  • Mix Text, Image, Video & Audio
    Feed prompts alongside picture, footage, or voice samples in a single run, pinning down a face, an art style, a movement, a camera path, or a specific voice through the comfyui minimax h3 nodes.

Running the comfyui minimax h3 Workflow in 3 Steps

Three quick steps take you from a fresh ComfyUI install to open-weight clips with built-in audio.

Capabilities Built Into the comfyui minimax h3 Workflow

From three ready-made ComfyUI templates to multimodal open-weight inference with stereo sound, reference-driven guidance, and optional Sage Attention acceleration — the comfyui minimax h3 workflow is a complete local video studio.

Three Ready-Made Templates

Text-to-video, image-to-video, and reference-to-video samples all ship inside the comfyui minimax h3 template library, so every generation mode works the moment you open it.

One Shared Omni-Modal Context

Prompts, stills, footage, and audio are all interpreted inside a single context by the comfyui minimax h3 model, letting every reference type feed the same run.

Guidance From Your References

Pin down a face, an aesthetic, a movement, a camera path, or a voice using source material — the comfyui minimax h3 R2V node accepts as many as 9 images, 3 videos, and 3 audio clips.

Crisp On-Screen Text and Logos

Words and brand marks come out sharp with the comfyui minimax h3 model, and you can spell out how references relate to one another in plain language.

Sage Attention Acceleration

Drop a Patch Sage Attention KJ node into the comfyui minimax h3 graph to roughly halve render time while barely touching output quality.

Resolution and Duration Controls

The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, rounded to the model's 32-pixel grid, with duration counted in 17-frame blocks at 24fps.

FAQ

comfyui minimax h3 — Frequently Asked Questions

Answers to the questions people ask most about running MiniMax H3 inside ComfyUI.

1

What exactly is the comfyui minimax h3 workflow?

It is the official ComfyUI integration of MiniMax H3 — a general-purpose, omni-modal model that MiniMax shipped with open weights. One forward pass turns text, pictures, footage, or audio references into a video that already carries stereo sound.

2

How high can the output quality go?

Clips from the comfyui minimax h3 workflow top out around 2K at 24fps and run roughly 15 seconds. The native canvas keeps a 768px short edge, caps at 768x1344, and snaps dimensions to multiples of 32.

3

What generation modes come bundled?

Three samples ship in the comfyui minimax h3 template library: text-to-video (T2V), image-to-video (I2V) with optional first- and last-frame guidance, and reference-to-video (R2V) for pinning a character, style, motion, camera move, or voice.

4

Does the model produce its own audio?

It does. Voice, sound effects, and music are generated natively in stereo by the comfyui minimax h3 model, produced in the same pass as the picture and muxed into one synchronized MP4.

5

How do I get set up?

Install ComfyUI 0.30.0 or newer, go to Template Library > Video, load any comfyui minimax h3 workflow, and accept the pop-up that pulls the weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Is there a way to render faster?

Add SageAttention plus the KJNodes custom pack, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 graph — render times drop by about half.

Your Next Video Starts with the comfyui minimax h3 Workflow

Keep MiniMax H3 on your own hardware, with open weights, stereo sound, and every knob exposed — text-to-video, image-to-video, and reference-to-video graphs are waiting for you.