Dots.TTS SOAR – State of the Art Speaker Similarity, Runs Fully Local

Dots.TTS SOAR – State of the Art Speaker Similarity, Runs Fully Local

More

Summary

Fahd Mirza walks through the installation and hands-on testing of Dots.TTS SOAR, a fully open-source zero-shot voice cloning model that runs entirely on local hardware. The model, built on a 2-billion-parameter architecture, supports 107 languages and is demonstrated cloning the presenter’s own voice to narrate a David Attenborough passage — with results showing strong speaker similarity and natural prosody despite a few minor pronunciation issues. The system runs on approximately 6 GB of VRAM and is accessed via a Gradio interface.

The video also covers multilingual voice cloning tests across Portuguese, Arabic, German, Slovak, and Hindi, using audio clips from Google’s Fleurs dataset as reference voices. Results across languages are generally strong, with consistent voice identity preserved.

Mirza provides a plain-language breakdown of Dots.TTS SOAR’s architecture: a reference audio clip passes through an audio VAE and a CAM++ speaker encoder to produce a voice fingerprint, while input text flows through a semantic encoder into a Qwen 2.5-based LLM. A flow-matching head (acting as a mini diffusion model) autoregressively generates audio patches, which are decoded into a 48 kHz waveform. The “SOAR” suffix refers to Self-Corrective Alignment Refinement, a post-training stage that improves adherence to both the source text and target voice. The model is available on Hugging Face with a public playground for browser-based testing.


📺 Source: Fahd Mirza · Published June 18, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels