Summary
This tutorial walks through building a fully local AI assistant on a 16GB Apple Silicon Mac Mini using Google’s Gemma 4 12B language model loaded in LM Studio, paired with Hermes Agent for task automation. The creator uses Claude Code running in Opus 4.8 Ultra mode to orchestrate the entire setup — installing Hermes Agent via terminal, configuring LM Studio, adjusting Mac power settings to keep the device always-on, and wiring Gemma as Hermes Agent’s default backend through LM Link.
The video covers the practical RAM management challenges of running a large model on constrained hardware, including how enabling Flash Attention and dropping to Q8 quantization frees up enough VRAM to push the usable context window to approximately 67,000 tokens — a meaningful jump for longer agent tasks. The creator also discusses the newer QAT (Quantization Aware Training) variant of Gemma 4 12B at 6.66GB as an alternative to the standard 7.04GB release.
By the end, viewers have a working local agent pipeline that uses Chromium browser mode for free, private web search without any paid API dependencies. The video is a practical reference for anyone wanting to run a self-hosted AI assistant stack — combining open-weight models, a local inference server, and an automation layer — entirely on consumer hardware without cloud costs.
📺 Source: Bart Slodyczka · Published June 15, 2026
🏷️ Format: Hands On Build







