Summary
Sam Witteveen reviews NVIDIA’s NeMo Tron 3.5 Lightning, a 30B-parameter mixture-of-experts model with only 3B active parameters, purpose-built for the execution layer of long-running AI agents. Rather than competing on reasoning or intelligence benchmarks, Lightning targets the unglamorous work that consumes the majority of agent token budgets: tool calls, output validation, retrieval, RAG pipelines, summarization, and classification.
Speed is the headline claim. NVIDIA reports 4x throughput versus comparable models, with Pinch Bench results across 10,000 tasks showing Lightning runs 30–35% faster than the similarly sized Qwen 3.6 MoE. The gains come from a hybrid Mamba-transformer architecture, a multi-token predictor baked in during continued pre-training on the Nemotron 3 base, D-Flash speculative decoding, and D-Spark — a DGX Spark-tuned variant derived from DeepSeek’s published speculative decoding methodology. Both BFloat16 and NVFP4 checkpoints are available, with the latter tuned for NVIDIA RTX and DGX hardware.
The more interesting story may be customization. CrowdStrike fine-tuned Lightning to match Nemotron 3 Ultra accuracy at roughly one-fifth the cost. CodeRabbit and Base10 trained a task-specific version in a single epoch — under three hours, at approximately $100. NVIDIA has released the post-training recipes and datasets openly, and Unsloth has published fine-tuning scripts that work on consumer RTX GPUs, putting local customization within reach of individual developers.
📺 Source: Sam Witteveen · Published August 11, 2026
🏷️ Format: Review







