Summary
NetworkChuck gets hands-on with Apple’s M5 Ultra, on loan from Apple, and puts it head to head with his M3 Ultra to see how much faster it really is for local AI. Both machines run in his server room with no cloud services involved, so every result reflects on-device performance.
The tests start with a Qwen 27B 4-bit model summarizing a 32,000-token essay. The M5 Ultra reaches the first token in about 21.7 seconds versus 82.9 seconds for the M3 Ultra, a 3.8x gain that he credits to the new per-core Neural Accelerators. Generation speed rises from 28.1 to 43.3 tokens per second, which he ties to memory bandwidth climbing from 819 GB/s to 1.2 TB/s. He explains why prompt processing and token generation are limited by different hardware.
He then tests practical creator workloads. OpenAI’s Whisper models transcribe a two-hour recording in roughly 24 seconds on turbo, about 2.3x faster than the M3 Ultra. A Qwen3-VL 32B model analyzes video footage 2.6x faster, an alternative to the Gemini video API he uses today. The video also covers image and video generation and asks whether a local model can drive DaVinci Resolve to edit footage. It is a useful reference for anyone weighing Apple silicon as a local AI server.
📺 Source: NetworkChuck · Published September 23, 2026
🏷️ Format: Benchmark Test







