FreeToken Setup Guide: Run 290B+ Frontier Models Locally on Your Gaming PC

FreeToken Setup Guide: Run 290B+ Frontier Models Locally on Your Gaming PC

More

Summary

Fahd Mirza demonstrates FreeToken, an open-source inference tool that makes it possible to run frontier mixture-of-experts models like DeepSeek and GLM — models with hundreds of billions of parameters — on a single consumer GPU by keeping most weights in system RAM and streaming only the experts each token actually needs onto the GPU on demand. The core insight is that MoE models never activate their full parameter count per token, so FreeToken’s expert cache plus LRU eviction means you only ever pay VRAM for the handful of experts currently in flight.

The video walks through installation via UV, downloading a model from Hugging Face, and running FreeToken’s built-in benchmark to automatically determine whether the workload should use offload mode (streaming experts over PCIe) or hybrid mode (splitting compute between CPU and GPU). On an NVIDIA RTX A6000, the tool auto-selected offload for the NVFP4 quantized model, loaded in about 30 seconds, and reserved over 21 GB for the KV cache — with actual VRAM consumption landing at roughly 44 GB total.

Mirza then shows the live terminal shell (ft shell), walking through the real-time status bar displaying tokens per second, expert cache hit rate, KV utilization, and VRAM consumption. He also notes FreeToken’s compatibility with six coding agents including Codex, Claude, and OpenCode, positioning it as a practical bridge for developers who want to run frontier-class models locally without a multi-GPU rig or API spend.


📺 Source: Fahd Mirza · Published August 25, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels