Summary
Web Dev Cody walks through how he embedded Whisper.cpp — a C++ port of OpenAI’s Whisper speech recognition model — directly into Mission Control, his free MIT-licensed Electron application for developer workflow automation. Rather than routing audio to a remote API, the app bundles the Whisper base model locally, ensuring voice data never leaves the user’s machine. The practical cost is significant but manageable: approximately 142 MB of additional disk space, 388 MB of RAM, and around 80% transcription accuracy — with the model occasionally truncating speech if the user talks too quickly.
The push-to-talk flow works by capturing microphone audio as a WAV file when a hotkey is held, passing it to a locally running Whisper Server process, and piping the resulting transcript through a command router that combines regex pattern matching with fuzzy string matching to identify the intended action — switching projects, launching browsers, opening sessions, or triggering custom scripts like a deploy pipeline. Cody considered routing transcripts to an LLM via MCP for more robust intent parsing, but measured the round-trip latency at five to six seconds even with fast models like HiQ or Composer 2.5, which he deemed too slow for a responsive voice-control experience.
The video also demonstrates Mission Control’s configurable voice command vocabulary, allowing users to define custom trigger words in settings. Mission Control is available free at agentsystem.dev.
📺 Source: Web Dev Cody · Published July 02, 2026
🏷️ Format: Hands On Build







