Summary
In this hands-on build video, developer sentdex continues a series exploring how large language models can directly control robots, this time upgrading a small XGO Mini robot to use function calling powered by GLM 5.3 Flash instead of hardcoded, rule-based logic. Rather than relying on manual HSV color masking to detect objects like a green cube, the video walks through replacing that vision pipeline with GLM 5.3 Flash’s own object detection capabilities, run through a self-hosted VLM server with adjustable reasoning levels.
The video also details a voice-driven development workflow, using OpenAI’s Whisper model for speech-to-text so the creator can talk through coding changes with an AI coding assistant instead of typing, along with a discussion of why verbal prompting often produces better context than manual typing. Viewers get a candid look at real debugging moments, including inconsistent detection loop timing and tradeoffs between polling speed and request latency when controlling a physical robot in real time.
For anyone interested in practical robotics, vision-language models, or agentic function calling, the video offers a transparent, unscripted look at building an LLM-controlled robot from the ground up, including the tools, hardware, and voice-to-text setup used throughout the process.
📺 Source: sentdex · Published September 18, 2026
🏷️ Format: Hands On Build







