PFlash + Qwen3.6-27B-DFlash: 10x Faster Prefill on a Single GPU: Run Locally

PFlash + Qwen3.6-27B-DFlash: 10x Faster Prefill on a Single GPU: Run Locally

More

Summary

Fahd Mirza builds and benchmarks PFlash, a prefill acceleration tool that dramatically reduces the blank-screen wait time when feeding long documents to locally hosted large language models. On a standard llama.cpp setup, processing a 128,000-token document with a 27-billion-parameter model takes approximately 4 minutes before the first output token appears; PFlash reduces this to roughly 25 seconds—a 10x improvement—by using a two-stage sparse attention approach.

PFlash works by first running a small 6-billion-parameter draft model that scores every token in the prompt for importance, selecting the top 5% for the main model to process. The main 27B model then prefills only those highlighted tokens using Block Sparse Attention (BSA) CUDA kernels. Combined with DFlash speculative decoding for the generation phase (74 tokens per second), the full three-model pipeline—drafter, main model, and generation draft—fits within 24GB of VRAM by loading each model sequentially rather than simultaneously.

The tutorial covers the complete setup on Ubuntu with an NVIDIA RTX 6000 (48GB VRAM, SM_86 architecture, CUDA 12.4): cloning the LooseHub repository, compiling llama.cpp with BSA kernels enabled, downloading and converting the Qwen-based draft model from HuggingFace GGUF format, and running a needle-in-a-haystack benchmark to verify the prefill speedup on a 128K token test prompt. The build takes 25–30 minutes to compile due to the CUDA kernel compilation step.


📺 Source: Fahd Mirza · Published May 02, 2026
🏷️ Format: Hands On Build

1 Item

Channels