Summary
Sam Witteveen breaks down Microsoft’s Windows and Surface event in San Francisco, focusing on what he sees as the most important idea on show: hybrid intelligence, where Windows decides which AI tokens run on the laptop and which go to the cloud.
The centerpiece is a GitHub Copilot demo in which an auto model picker routed simple cleanup work to a local Microsoft AI Code Flash model, then used a cloud OpenAI model for planning while three local sub-agents read through Windows Terminal issues. The session sent about 1.6 million tokens to the model with only around 10,000 coming back out, at no token cost because the work ran on-device. The updated Copilot app is slated to ship on October 15th.
Witteveen also looks at the hardware side, including the NVIDIA RTX Spark and its memory configurations, and at the quantization claims behind the local models, such as a three-bit coding model and DeepSeek-4 running at an average of 1.6 bits in about 60 GB. He raises practical questions about coding pass rates, long-context behavior, KV cache memory, tool-call reliability, and what happens when a router sends a task to a local model that cannot handle it.
📺 Source: Sam Witteveen · Published October 08, 2026
🏷️ Format: News Analysis







