08:28 Coding & Builds4 months ago Qwen3-8B at 74 tok/s with RedHat DFlash Speculator on vLLM Locally Fahd Mirza walks through running Red Hat's DFlash speculative decoding implementation on Qwen3-8B using vLLM, achieving 74 tokens per... 0 comments 1.7K views
11:12 Benchmarks4 months ago Qwen3.6 27B Gets 20% Faster with MTP and llama.cpp Locally Fahd Mirza demonstrates how to enable multi-token prediction (MTP) on Qwen3.6 27B using ik_llama.cpp — a community fork of the popula... 0 comments 3.4K views