MiniMax H3 β€” ComfyUI Workflows

Drag-and-drop ComfyUI workflows for MiniMax H3, the omni-modal video model that generates video with native stereo audio (voice, SFX, room tone β€” all in one pass). Shared by @BennyDaBall_OG.

Workflows

File What it does
MiniMax_H3_Turbo_v4_8step_T2VA_BennyDaBall.json Text β†’ video+audio with the Turbo v4 LoRA β€” 8 steps instead of 20, ~1.9Γ— faster, audio stays clean

Turbo v4 Β· 8-step T2VA

100% core ComfyUI nodes β€” no custom node packs. Verified with a same-seed A/B against the stock 20-step recipe: same choreography, healthy audio (peak β‰ˆ βˆ’9 dB, zero clipping).

The recipe

Setting Value
LoRA minimax_h3_turbo_v4_step600_ema_pruned_comfyui @ 1.0 (built-in LoraLoaderModelOnly)
Steps / sampler / scheduler 8 Β· euler Β· beta (no sigma-shift node needed)
Guidance CFG-free (BasicGuider) β€” write a Negative: line inside the prompt itself
Resolution 1344Γ—768 @ 24 fps
Length frames on a 17k+5 grid: 124 β‰ˆ 5 s, 243 β‰ˆ 10 s, 362 β‰ˆ 15 s (model max)

The bundled demo prompt ships with a fixed seed so your first render reproduces the demo clip β€” flip the seed widget to randomize for new takes.

Getting started

  1. Update ComfyUI to β‰₯ 0.30 (MiniMax H3 nodes are built in).
  2. Drag the .json onto the ComfyUI canvas.
  3. ComfyUI will offer to download any missing models automatically (every loader node carries its download URL). Manual links below.
  4. Queue. First run: the console should show zero lora key not loaded warnings.

Models

Tested / not tested

  • Verified on 5-second clips at 1344Γ—768. Higher res (1920Γ—1088) and longer clips are untested with this LoRA β€” earlier turbo versions fell apart there, so treat that as experimental.
  • Few-step audio blowout is the classic turbo failure mode; this v4 @ 8 steps is the combo that passed. If you drop steps further, check your audio peaks.

Prompt pattern that works

Style contract β†’ timed beats (0.0-0.6s: ...) β†’ camera lock line β†’ an audio timeline starting with room tone at 0.0s β†’ Negative: line. Dialogue inside <d>[English] ... </d> is spoken verbatim with lip-sync. Give the first spoken line β‰₯ 1.2 s of lead-in and anchor t=0 with room tone. The demo prompt in the workflow shows the full pattern.

Credits

License note: this repo covers the workflow files only. The H3 model weights and the LoRA carry their own licenses β€” check the linked repos before commercial use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support