Summary
Alex Ziskind pits a 256 GB M5 Ultra Mac Studio against a cluster of two NVIDIA DGX Sparks, each with 128 GB, linked by a 200 Gb QSFP cable running RDMA over Converged Ethernet (RoCE). The Mac as configured costs about $14,000 with an 8 TB drive, while the Sparks now run close to $5,000 each, so the pair is cheaper even with the cable.
The tests run DeepSeek V4 Flash and a Qwen model at four-bit quantization. The Sparks use vLLM with tensor parallelism, splitting every layer across both boxes, while the Mac is tested with both llama.cpp and MLX. On token generation the two are nearly even, at 38.7 tokens per second for the Mac versus 34.3 for the dual Sparks in the opening run.
The gap shows up in prompt processing. At a 32,000-token codebase prompt, the Sparks begin streaming after about 17 seconds, while the Mac takes around 50 seconds. Across other tests the Sparks were roughly 2.4 to 3 times faster at reading prompts, and they handled a 128,000-token prompt in 72 seconds. The video explains the prefill and decode phases so developers running large models locally or serving a small team can judge which hardware suits their workloads.
📺 Source: Alex Ziskind · Published October 01, 2026
🏷️ Format: Benchmark Test







