Summary
Alex Ziskind and guest Wendell put two high-end local AI machines to work: a system with eight RTX Pro 6000 GPUs and an Nvidia DGX Station. Together they draw around 6,000 watts, run headless, and are managed remotely. The goal is to see how they handle real developer workloads using Turnstone, a tool for orchestrating coding sessions across a fleet of machines and models.
The video demonstrates setting up models in Turnstone, using built-in personas such as engineer, researcher, and orchestrator, connecting MCP servers, saving memories, and running multiple tasks in parallel on one machine. It also includes a candid moment when an agent deletes a home directory, a reminder of the risks of giving agents broad permissions.
The hardware comparison is where the numbers appear. The DGX Station has 252 GB of HBM3e at about 7 TB/s plus 748 GB of total unified memory with roughly 600 GB/s LPDDR5, while the RTX Pro 6000 setup communicates over PCIe Gen 5 at about 64 GB/s each way. The hosts compare results on models such as GPT-OSS 120B and DeepSeek V4 Flash, including one run at 307 tokens per second, and discuss why mixture-of-experts models suit this memory design.
📺 Source: Alex Ziskind · Published October 07, 2026
🏷️ Format: Benchmark Test







