Local AI models explained: How to run a fleet of Mac Studios and GPUs at home

Local AI models explained: How to run a fleet of Mac Studios and GPUs at home

More

Summary

How I AI host Claire Vell interviews Alex Finn, who has built one of the more elaborate home AI infrastructure setups in the practitioner community: three Mac Studio 512GB units (reselling around $30,000 each), an Nvidia DGX Spark, and a custom desktop built around an RTX 5090. Every machine runs continuously, burning tokens around the clock on what Finn calls ambient AI — a philosophy of having local models handle ongoing life and business tasks without relying on cloud APIs.

Finn traces his journey back to discovering Open Claw in January, which sparked an obsession with running models locally and building a personal bond with on-device agents. He now uses Open Claw and Hermes as the orchestration layer across his fleet, with Tailscale providing a private network that lets agents jump between machines, load models, and execute tasks without manual configuration — including from a phone on the go. This setup enables a software factory with parallel build and review loops running in Claude Code: a build agent works through a task queue continuously while a separate review agent validates completed work and pings Slack with a summary; a rocket emoji from Finn triggers an automated merge.

The conversation digs into why pure cloud ROI misses the point — unlimited local inference unlocks use cases that would be cost-prohibitive at API pricing — and walks through practical Tailscale configuration, Ollama for model management, and the workflow design behind Finn’s self-improving agent system. A detailed, firsthand look at what serious local AI infrastructure looks like in 2026.


📺 Source: How I AI · Published July 13, 2026
🏷️ Format: Interview

1 Item

Channels

1 Item

People