M5 Ultra… Apple Wasn’t Messing Around

M5 Ultra… Apple Wasn’t Messing Around

More

Summary

Alex Ziskind puts Apple’s new M5 Ultra chip through a rigorous set of benchmarks against the previous-generation M3 Ultra, focusing heavily on what the upgrade means for running large language models locally. Testing a maxed-out configuration with an 80-core GPU, 256GB of memory, and 8TB of storage priced at $14,299, he verifies Apple’s marketing claims using tools like Llama.cpp, MLX, and the STREAM memory bandwidth benchmark rather than taking spec-sheet numbers at face value.

The video breaks down real performance gains across CPU compilation tasks, storage speed, and — most notably — local AI inference, measuring both prompt processing and token generation on models including DeepSeek V4 Flash at 284 billion parameters. Ziskind finds token generation roughly 1.5x faster on the M5 Ultra, closely tracking the measured 1.4x increase in GPU memory bandwidth (1,039 GB/s versus 724 GB/s), and explains why the new per-core neural accelerators represent a bigger architectural shift than the raw core count suggests.

For developers and AI enthusiasts considering high-end Apple Silicon for local model hosting, the video offers concrete, reproducible numbers rather than marketing claims, making clear which improvements matter for compute-bound versus memory-bound AI workloads.


📺 Source: Alex Ziskind · Published September 22, 2026
🏷️ Format: Benchmark Test

1 Item

Channels

1 Item

Companies