Summary
ByteDance has open-sourced Bernini-R, a versatile video and image editing model, and this tutorial from Nerdy Rodent walks through a complete ComfyUI workflow covering its full range of capabilities. The video follows the channel’s “rodent method” — clean layouts with named groups — making the pipeline easy to follow and modify. Key technical components include dual transformer models for high- and low-noise denoising, normalized attention guidance (NAG) from KJ Nodes to prevent prompt distortion, skip layer guidance for fixing broken shapes and unstable motion, and VRAM-saving options including Sage Attention patches and WAN chunk feed-forward nodes.
The tutorial covers a unified single-workflow approach that handles text-to-image, image-to-video, video editing, and subject swaps. Bernini-R defaults to 480p at 16fps but supports up to 720p, and its conditioning node accepts up to eight reference images. A two-node 2× VAE upscaler from VAE Utils can bring image outputs to 2560×2560. A practical tip shared: including a reference image even during text-to-image generation tends to improve style consistency.
Viewers will come away knowing how to configure the full loader stack, set up prompting with optional negative prompt support, resize inputs with KJ Nodes resize image v2, and switch between task modes via a combo-box-driven string replacement system. Those with limited GPU memory can also use the 1.3B tiny Bernini-R model available on the official ByteDance GitHub repository.
📺 Source: Nerdy Rodent · Published June 18, 2026
🏷️ Format: Tutorial Demo







