Nemotron 3 Diarization – Who Said That?

Nemotron 3 Diarization – Who Said That?

More

Summary

Sam Witteveen looks at Nvidia’s newly released Nemotron 3 Diarization, an open-weights model that answers the question speech recognition alone cannot: who said what. Transcripts from Whisper, Parakeet or Canary tell you the words, but meeting notes, podcast transcripts and AI agents also need to attribute each quote to the right speaker.

The model has roughly 100 million parameters, runs on a GPU with about 4 GB of memory, and is licensed for commercial use. It handles up to eight speakers, compared with three or four for most earlier systems, and copes with overlapping speech. It works in both offline and streaming modes, is language agnostic, and is designed to pair with whatever ASR model you already use. Witteveen explains diarization error rate, which combines missed speech, false alarms and speaker confusion, and shows the model beating Nvidia’s previous Sortformer, which had over 300,000 downloads in August alone.

He also places the release within Nvidia’s wider speech family, including Parakeet, Canary, Nemotron ASR, the Magpie TTS model and PersonaPlex. The demo runs on a DGX Spark with NeMo in Docker, a FastAPI backend reached over Tailscale, and a Next.js app on a Mac, turning a raw podcast recording into a transcript with every line attributed to its speaker.


๐Ÿ“บ Source: Sam Witteveen ยท Published September 23, 2026
๐Ÿท๏ธ Format: Tutorial Demo

1 Item

Channels

1 Item

Companies