Summary
Mozhgan Kabiri chimeh, Developer Relations Manager at NVIDIA, presents empirical benchmarking results from running large language models locally on the NVIDIA DGX Spark — a desktop system powered by the GB10 Grace Blackwell superchip featuring 128 GB of unified memory and support for models up to approximately 200 billion parameters. The talk is framed explicitly around reproducibility and practical developer workflows rather than theoretical ceilings.
The benchmarking harness is rigorous: all models are served via vLLM inside NVIDIA-optimized Docker containers, with mandatory three-run warm-ups and GPU metrics logged at one-second intervals. Models tested range from 1.5 billion to 14 billion parameters. Key results: the 1.5B instruction model delivers 61.73 tokens per second; a 14B model using NVIDIA’s NVFB4 4-bit floating-point quantization achieves 20.19 tokens/sec — still above average human reading speed — while the same 14B model in base precision drops to just 8.40 tokens/sec. The data makes a strong case that quantization format selection is as consequential as hardware choice on Blackwell architecture.
The session also covers time-to-first-token (TTFT) measurement methodology, explaining why it is the dominant UX metric in streaming LLM applications, and discusses how the DGX Spark’s unified memory architecture and shared CUDA software stack with data centers reduces friction when moving workloads from local development to production deployment.
📺 Source: AI Engineer · Published April 10, 2026
🏷️ Format: Benchmark Test







