Which is The Best Qwen3.8-27B?

Which is The Best Qwen3.8-27B?

More

Summary

Reasoning models often reach the right answer only after a long, expensive detour of thinking tokens. This video from Sam Witteveen compares several community fine-tunes of Qwen3.8-27B that aim to cut that overhead while keeping accuracy intact.

Witteveen starts with the base model’s built-in reasoning effort settings (low, medium, and X high, with X high as the default) and explains why shorter reasoning matters for local use: on a single GPU, decoding is the bottleneck, so fewer tokens translate directly into less waiting. He also notes that multi-token prediction heads used for speculative decoding work better on shorter traces, with Thinking Cap reporting about 2.6 tokens per step at X high versus more than three at medium and low.

The video then walks through the family tree of variants, including Thinking Cap from Bottlecap AI, Swift 1.5 from UKIS AI, and Qwen Pi, a version aimed at coding agents. Swift 1.5 claims up to 58.5% fewer thinking tokens, with LiveCodeBench rising from just under 77% to 81.7% while output drops from roughly 11,200 to 8,400 tokens. Witteveen covers how each was trained, their licenses and available formats, and runs them against the base model to test whether the claims hold up.


📺 Source: Sam Witteveen · Published October 04, 2026
🏷️ Format: Comparison

1 Item

Channels