Summary
Sam Witteveen takes a close look at the newly released Qwen 3.8 27B model from Alibaba’s Qwen team, positioning it as the natural successor to the widely adopted Qwen 3.6 27B that became a go-to choice for local coding and agent workflows. Official benchmarks show substantial gains over the 3.6 model and Meta’s Llama Maverick, with Artificial Analysis placing Qwen 3.8 27B at an Intelligence Index score of 52 โ ahead of most locally runnable open-weight models and competitive with larger closed-source systems on agentic tasks.
A large portion of the video addresses which version to actually run: the BF16 full-precision release, the FP8 quantization, or the NV FP4 4-bit builds from Unsloth, with notes on GPU compatibility caveats for each. Witteveen also systematically tests all four thinking modes โ no thinking, low, medium, and XHigh โ measuring real token consumption across repeated runs. Low thinking consistently produces around 512 thinking tokens, while XHigh regularly consumes 17,000โ22,000 tokens per query, creating serious latency and context window pressure. Practical tasks including website generation and an SVG pelican drawing test (drawn from Simon Willison’s benchmark suite) illustrate these tradeoffs across quantization levels, making this a useful decision guide for anyone choosing how to deploy Qwen 3.8 27B on consumer or prosumer hardware.
๐บ Source: Sam Witteveen ยท Published August 18, 2026
๐ท๏ธ Format: Review







