BreezeTTS2 – 100% Local Real-Time Voice

BreezeTTS2 – 100% Local Real-Time Voice

More

Summary

Sam Witteveen takes BreezeTTS2 — a brand-new open-weight text-to-speech model from Chinese startup Breeze Blue — for a hands-on local deployment test just six days after its release. The model currently tops the open-weight TTS leaderboard and supports three distinct generation modes: voice design (creating a voice from a text description alone), voice direction (cloning a reference voice and steering its emotion and delivery), and vocal events (laughs, coughs, sighs, and other paralinguistic sounds). It also handles 50 languages.

Witteveen runs the model entirely locally on a Dell T2 Pro Max with an RTX Pro 6000 GPU, streaming inference via Tailscale, and walks through live demos of each feature — including cross-lingual voice cloning from just four seconds of reference audio. The results show convincing accent reproduction, emotion transfer, and real-time generation speeds competitive with closed-source alternatives like ElevenLabs.

The major caveat is licensing: BreezeTTS2 is released under a research and non-commercial license that prohibits commercial output and model distillation. Witteveen is candid that this restriction significantly limits its practical reach — without it, the model would be a strong default choice for game audio, AI assistants, and production pipelines. For personal or research use, however, it is currently the most capable freely available TTS model, making this video a useful benchmark for where open-weight voice synthesis stands in mid-2026.


📺 Source: Sam Witteveen · Published August 31, 2026
🏷️ Format: Review

1 Item

Channels