← Back to news

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

blog.comfy.org|168 points|50 comments|by vblanco|Aug 3, 2026

MiniMax H3: Day-0 Integration in ComfyUI

Open Weights, Native Audio, and 2K Resolution

ComfyUI Newsletter By Rob and Alexis Rolland | August 03, 2026

The AI landscape just shifted with the release of MiniMax H3. This omni-modal video powerhouse has arrived with open weights and immediate Day-0 support within ComfyUI. Remarkably, this model is optimized enough to run locally on hardware as accessible as an RTX 3060.

Model Evolution

MiniMax H3 represents the third iteration of the company's video generation journey.


🛠️ Technical Specifications & Capabilities

The model can process a variety of inputs—text, imagery, video, or audio—to produce high-fidelity clips.

FeatureSpecification
Max Resolution2K
Clip DurationUp to 15 seconds
AudioNative Stereo Sound
Hardware Req.Local execution possible on RTX 3060
Weight StatusOpen Weights

Core Functionalities

  • Text-to-Video: Pure prompt-based generation.
  • First-and-Last-Frame: Define the start and end points; the model interpolates the middle.
  • Reference-to-Video: Use existing audio, video, or images to maintain subject, voice, or motion consistency.
  • Native Audio: Audio is not bolted on as a post-process \rightarrow Audio is a fundamental property of the model generated in a single pass.

🧠 Multimodal Intelligence

The standout feature of H3 is its multimodal context understanding. Rather than treating different inputs as separate layers, H3 collapses five distinct tasks into one unified process.

"H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate."

Mathematically, the model's logic can be viewed as: Video Output=Model(PromptImagerefAudiorefVideoref)\text{Video Output} = \text{Model}(\text{Prompt} \cup \text{Image}_{ref} \cup \text{Audio}_{ref} \cup \text{Video}_{ref})

Motion Transfer & Iteration

For those working with complex graphs, motion transfer is the critical tool. You can extract a specific camera movement or performance from a reference video while sourcing the style and subject from different inputs. This, combined with in-place editing, allows for rapid shot iteration.


🎬 Example Prompting & Outputs

Below are the detailed prompts used to showcase the model's versatility.

Example 1: The Comic Book Aesthetic

Style: Bold comic-book ink, heavy linework, red/blue-black palette, night city. References: Picture 2 (Start), Picture 1 (End), Audio 1 (Exact).

CUT 1: Top-down view of a young superhero boy on a rooftop. Red cape fluttering, hands on hips, cocky grin. 
Camera descends as he speaks; comic-book text overlays "GET READY TO - MEET — YOUR — MAKER" in sync with voice. 
Text is jagged, white with black outlines and red shadows.

TRANSITION: Violent WHIP PAN smearing the text away.

CUT 2: Low hero angle of a massive black mech-kaiju. It rears back and ROARS. 
Fangs visible, red eyes/chest flaring, blue lightning arcing. 
Shockwaves ripple dust and rattle windows with comic-style speed-lines and ink splatter.

Example 2: High-End Product Cinematography

Subject: Transparent gaming mouse from Picture 1. Setting: Pitch-black studio void, reflective surface, duotone blue and neon orange rim lighting.

SHOT 1: Starts on Image 1. Blue/orange lights pulse, refracting through acrylic as camera pushes in on circuitry.

SHOT 2: Extreme macro profile of scroll wheel and internals. Orange light sweeps across metallic textures against blue ambient glow.

SHOT 3: Low-angle beauty shot. Mouse levitates and rotates in a slow orbit; lighting flares on glassy edges before fading to silhouette.

AUDIO: Deep pulsing sub-bass, tactile mechanical clicks, glassy whooshes on cuts, and a rising electronic swell.

Example 3: High-Fashion Editorial

Style: Luxurious slow motion, soft gradient studio sky. Audio: Cinematic score (Taiko drums, koto plucks, modern sub-bass).

SHOT 1: A broken mask (Picture 2) hangs in floating shards; gold kintsugi seams are dim.

SHOT 2: THE ASSEMBLY. Gold seams ignite with molten light; shards snap together rapidly with gold flares and shockwave ripples until the mask fuses.

SHOT 3: The golden dragon (Picture 3) swoops through in a serpentine motion, red glass antlers first, dragging crimson liquid into a spiral.

SHOT 4: Mask magnetically rips onto the woman's face. Gold veins spread from the mask down her neck and across her sunset jacket embroidery.

Feature Image


Contributors: Rob Rob & Alexis Rolland Alexis Rolland