Summary
NVIDIA has released NeMo-Tron 2 Tower 30B A3 Base, a 30-billion parameter model that rethinks how large language models generate text at the architectural level. Built on top of the existing NeMo-Tron 3 Nano backbone, the model introduces a two-tower design: a frozen autoregressive “context tower” that reads and remembers all prior tokens, and a separate “diffusion denoiser tower” that generates entire blocks of new tokens in parallel rather than one token at a time. The two towers share knowledge layer-by-layer via KV cache and MoBA states, allowing the denoiser to generate text with full context awareness.
The masked diffusion process works by starting with a block of masked positions — say, 16 tokens — scoring each in parallel, committing the highest-confidence tokens first, then iterating until the block is complete. This yields multiple token commits per step versus the single-token-per-step constraint of traditional autoregressive generation, producing a reported 2.4x throughput increase. Because only the right (denoiser) tower was trained while the left tower remained frozen, training required roughly 10% of the original pre-training compute.
Benchmarks show the model retains 98.7% of the original NeMo-Tron 3 Nano’s quality across general knowledge, code, math, common sense, and multilingual tasks — with math dropping ~3 points but multilingual performance actually improving slightly. The channel host, Fahd Mirza, argues this architecture is meaningfully different from bolt-on speed techniques like speculative decoding or multi-token prediction because the speed gain is baked into the model itself. The model is released for commercial use.
📺 Source: Fahd Mirza · Published July 02, 2026
🏷️ Format: Deep Dive







