NVIDIA’s Two-Tower Model Generates Text 2.4x Faster Without Losing Quality

NVIDIA’s Two-Tower Model Generates Text 2.4x Faster Without Losing Quality

More

Summary

NVIDIA has released NeMo-Tron 2 Tower 30B A3 Base, a 30-billion parameter model that rethinks how large language models generate text at the architectural level. Built on top of the existing NeMo-Tron 3 Nano backbone, the model introduces a two-tower design: a frozen autoregressive “context tower” that reads and remembers all prior tokens, and a separate “diffusion denoiser tower” that generates entire blocks of new tokens in parallel rather than one token at a time. The two towers share knowledge layer-by-layer via KV cache and MoBA states, allowing the denoiser to generate text with full context awareness.

The masked diffusion process works by starting with a block of masked positions — say, 16 tokens — scoring each in parallel, committing the highest-confidence tokens first, then iterating until the block is complete. This yields multiple token commits per step versus the single-token-per-step constraint of traditional autoregressive generation, producing a reported 2.4x throughput increase. Because only the right (denoiser) tower was trained while the left tower remained frozen, training required roughly 10% of the original pre-training compute.

Benchmarks show the model retains 98.7% of the original NeMo-Tron 3 Nano’s quality across general knowledge, code, math, common sense, and multilingual tasks — with math dropping ~3 points but multilingual performance actually improving slightly. The channel host, Fahd Mirza, argues this architecture is meaningfully different from bolt-on speed techniques like speculative decoding or multi-token prediction because the speed gain is baked into the model itself. The model is released for commercial use.


📺 Source: Fahd Mirza · Published July 02, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies