Summary
This Veteran AI tutorial provides an end-to-end guide to MiniMax H3’s video editing capabilities through its Ref2VA base model, 10Eros Max — a system designed not for text-to-video generation but for replacing characters, outfits, and props within existing footage while preserving original motion, camera movement, and scene environment. The workflow is built in ComfyUI and mirrored on RunningHub for cloud-based access.
The pipeline uses a two-stage sampling approach: an initial pass at 0.2 megapixels with the base model and a special attention mechanism, followed by a 3D latent upscaler and a second high-resolution sampling stage. Setup requires cloning the MiniMax H3 Latent Upscaler into ComfyUI’s custom_nodes folder and downloading the corresponding model weights. The tutorial walks through aspect ratio configuration, frame count calculation (approximately 124 frames for a five-second portrait video), and reference image vs. reference video handling — importantly, reference images need not be resized since they provide visual features only, while reference videos must match the output dimensions exactly.
Prompt construction is covered in detail using the official six-section Ref2VA structure: subject definitions, task summary, preservation constraints, detailed action descriptions, and additional parameters. Testing escalates from simple character-walking replacement to complex multi-step sequences (blush application, lipstick, glasses, waving). The video honestly documents a failure case where a brief prompt caused the model to skip the lipstick step and alter a one-handed wave to two-handed — demonstrating that action complexity requires proportionally more explicit prompt descriptions.
📺 Source: Veteran AI · Published August 28, 2026
🏷️ Format: Tutorial Demo







