Summary
Fahd Mirza installs and evaluates Audio8 TTS, a newly released open-source text-to-speech model distinguished by its unusually small 0.6 billion parameter footprint. Despite being the lightest model in its class, Audio8 is positioned by its developers as competitive with much larger systems including Moss GTS at 8.5 billion parameters and Higs Audio V2 at 4.7 billion parameters.
The model uses a dual autoregressive architecture adapted from Fish Audio: a slow transformer generates one semantic token per audio frame while a fast transformer fills in acoustic detail, producing 44.1 kHz output. It supports zero-shot voice cloning across approximately 11 languages including English, Spanish, Japanese, Korean, German, and Polish. Mirza runs the Gradio demo locally on a single GPU, observing peak VRAM consumption of approximately 32 GB with default KV cache settings โ adjustable downward for hardware-constrained setups. The model is served via SGLang.
Mirza’s evaluation is candid and language-dependent. Spanish and German outputs are notably stronger than English, which he characterizes as carrying a persistent artificial quality that lags behind the current state of the art. Zero-shot voice cloning captures speaker texture reasonably well but lacks emotional range and sounds flat. His overall verdict: Audio8 is a technically interesting achievement at its parameter count and worth monitoring, but Fish Audio โ which uses the same underlying architecture โ currently produces more natural-sounding results, and Audio8 needs further refinement before it can compete at the top tier of open-source TTS.
๐บ Source: Fahd Mirza ยท Published August 06, 2026
๐ท๏ธ Format: Review







