LongCat Video Avatar 1.5 – Make Any Image Talk With Your Voice Locally for Free

LongCat Video Avatar 1.5 – Make Any Image Talk With Your Voice Locally for Free

More

Summary

LongCat Video Avatar 1.5, released by Meituan — the Chinese tech giant behind food delivery — is a talking avatar generation model that animates a single reference image in sync with a provided audio clip. In this walkthrough, Fahd Mirza installs and tests the model locally on Ubuntu, running into a real-world out-of-memory error on his 48GB NVIDIA RTX A6000 before upgrading to an 80GB H100 via Vast Compute, where the model consumes roughly 54GB VRAM at full load.

The key architectural upgrades in version 1.5 include replacing the older Wave2Vec audio encoder with Whisper Large for noticeably smoother lip-sync, and applying DMD2-based step distillation to cut inference down to just eight steps. The model supports three core generation modes — audio-only, audio plus reference image, and video continuation — and handles single-person, multi-person, anime characters, animals, and object-interaction scenes. Released under an MIT license, it is commercially usable, which distinguishes it from many closed alternatives in the space.

Demonstrations in the video show photorealistic and anime character outputs, with improved eyebrow and facial expression movement compared to prior versions. The lip-sync quality on anime characters is highlighted as particularly strong. Viewers interested in running the model locally will need significant GPU resources, but the video provides step-by-step setup instructions using the public GitHub repo and a Streamlit-based interface.


📺 Source: Fahd Mirza · Published May 24, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels