Summary
SCAIL-2 represents a significant upgrade over the original SCAIL model, moving from skeleton and pose map intermediaries to a more end-to-end approach where the model learns motion, character, and region relationships directly from a reference video and reference image pair. This tutorial from Veteran AI breaks down a complete ComfyUI workflow and stress-tests it across six scenarios to show exactly what works, what needs tuning, and what commonly fails.
The workflow is built on the Wan 2.1 architecture using an FP8 SCAIL-2 checkpoint, combined with CLIP Vision for image encoding and SAM 3.1 for subject segmentation. Generation runs in two stages: a 65-frame base clip followed by an 81-frame extension with a 5-frame overlap for seamless stitching, for an effective 141 total frames. A critical and easily missed parameter is the `max_object` setting inside the SAM tracking node — leaving it at the default of 1 will silently drop one subject in any two-person scene. The video covers the SCAIL2ColoredMask node that converts segmentation output into the colored-region format SCAIL-2 requires, and explains the WanSCAILToVideo node’s special inputs including pose video, reference image mask, and replacement mode toggle.
Test scenarios demonstrated include basic full-body motion transfer, close-up facial expression driving (where the model’s performance is more variable), and character replacement mode. The tutorial is also hosted on RunningHub for users who prefer a browser-based ComfyUI environment, with the workflow available directly in their community library.
📺 Source: Veteran AI · Published June 15, 2026
🏷️ Format: Tutorial Demo







