Qwen3.6-27B Fable Fusion 711: Strongest, Smartest Fine-Tune: Tested Locally

Qwen3.6-27B Fable Fusion 711: Strongest, Smartest Fine-Tune: Tested Locally

More

Summary

Fahd Mirza installs and thoroughly benchmarks Qwen3.6-27B Fable Fusion 711, a community fine-tune that claims to be the first model of its size to breach 700 on ARC-C in both 8-bit and 4-bit quantization — hence the “711” in its name. The model starts from Qwen 3.6 27B, layers on fine-tuning using the Polaris and F451 datasets along with frontier model traces, applies refusal ablation via arbitrary-rank ablation, and is merged in multiple stages, with benchmarks taken at every step.

Running on an NVIDIA A6000 with 48 GB VRAM and serving through llama.cpp, Mirza conducts a controlled head-to-head comparison of multi-token prediction (MTP) versus standard inference. MTP — where extra prediction heads baked into the weights draft two to three tokens per forward pass without a separate draft model — delivers close to double the tokens-per-second on his hardware, far exceeding the roughly 20% gain cited on the model card for a different GPU/quant combination.

For qualitative testing, Mirza uses the Hermes agent framework to have the model enforce a food-safety temperature rule against a real SQLite database of 24 rotisserie chickens, requiring it to write and commit updates, then prove the changes with a SELECT and a count. The model navigates the task correctly despite planted traps in the data. Mirza also explains MTP architecture in accessible terms, making this a useful reference for anyone evaluating local inference options for large open-weight models.


📺 Source: Fahd Mirza · Published August 07, 2026
🏷️ Format: Hands On Build

1 Item

Channels