Summary
Alex Ziskind puts the Camino Grando, a liquid-cooled workstation packing eight Nvidia RTX Pro 6000 Blackwell GPUs and 768GB of total VRAM, through a battery of local LLM benchmarks. He details the hardware itself, including an AMD Epic 9474F CPU, 512GB of DDR5 RAM, custom copper water blocks, and a 6.5kW power draw, before testing how it handles massive open-weight models like GLM 5.2, GLM 5.3, Qwen3 235B, Qwen3.8 Flash Next, and DeepSeek V4 Flash.
The video walks through generation speed, prompt processing throughput, and time-to-first-token across prompt sizes from 2,000 to 128,000 tokens, showing GLM 5.2 hitting 48 tokens per second at 433GB in size while Qwen3.8 Flash Next processes prompts at over 12,000 tokens per second. Ziskind explains why multi-agent coding workflows, running several coding agents simultaneously rather than a single chat session, are pushing local hardware requirements far beyond what a single developer needed a year ago.
He also covers practical constraints like PCIe lane allocation, the lack of NVLink, and running out of storage mid-test. For developers or teams evaluating whether a local multi-GPU rig can replace cloud inference for agentic coding workloads, this offers concrete, reproducible performance numbers rather than marketing claims.
📺 Source: Alex Ziskind · Published September 16, 2026
🏷️ Format: Benchmark Test







