I Built a Local AI Cluster and My AI Agents Are Free Now

I Built a Local AI Cluster and My AI Agents Are Free Now

More

Summary

Creator Magic demonstrates Buzz’s newly shipped “Shared Compute” feature, built on an open-source project called Mesh LLM, which turns a community of personal computers into a distributed AI inference cluster. The core idea: one community member runs a local model on their machine, and every other member’s Buzz agents can route queries to it — encrypted peer-to-peer, bypassing OpenAI, Anthropic, and any cloud API entirely. The Buzz relay server acts only as a phone book of who is sharing and which models are available; it never sees the prompt content.

The setup walkthrough is detailed: enabling Shared Compute in Buzz settings automatically selects a model based on available RAM — Gemma 4 26B (17GB) for machines with 64GB or more, a lighter 4B variant for lower-spec hardware. Agents are then pointed at the “Buzz Shared Compute” LLM provider instead of a commercial API, with no API key required. The demo runs across Apple Silicon hardware including a Mac Studio M3 Ultra with 256GB RAM, which can serve Llama 3.3 70B to the entire community pool. Three operational modes are explained: simple model sharing, Mixture of Agents (querying multiple models simultaneously and merging answers), and model parallelism for trillion-parameter models split layer-by-layer across several machines.

The video closes with a preview of an unreleased UI redesign — shown as a working mock — that will surface live peer status, connection proximity rings, and compute capacity directly in the Buzz sidebar. The privacy tradeoff is addressed directly: the machine answering a query can read the prompt, making trust and community membership gating essential design constraints.


📺 Source: Creator Magic · Published August 07, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels