Summary
Tina Huang runs local AI models on every tier of hardware she owns, from a bare Arduino R4 Uno microcontroller up through Raspberry Pi and full computers, to show viewers exactly what’s feasible at each memory and compute level. On the tiniest ESP32-S3 board (8MB RAM, roughly $2-20), she runs “Tiny Stories” models as small as 260KB, while a Raspberry Pi 5 chains together Whisper for speech-to-text, a Qwen 3 1.7B language model, and the Piper text-to-speech model to hold full spoken conversations.
Along the way, she explains core AI hardware concepts using a kitchen analogy—RAM as prep counter, memory bus as conveyor belt, CPU/GPU/NPU as chefs and line cooks—to clarify why devices with identical RAM, like a Raspberry Pi and an iPhone, perform so differently due to memory bandwidth and the presence (or absence) of GPUs and NPUs.
The video is a practical reference for anyone curious about self-hosting AI on constrained devices, covering which models fit where, what tasks are realistically achievable (from simple text prediction to full multimodal conversation), and why capabilities like image generation require dedicated GPU hardware rather than just sufficient memory.
📺 Source: Tina Huang · Published September 28, 2026
🏷️ Format: Hands On Build







