05:44 News & Opinion1 month ago Meta’s new model wants “deep access” to your personal life… Meta has released Muse Glimmer, a 30-billion-parameter open-source agentic model under the Apache 2.0 license, marking a significant... 0 comments 349.4K views
10:50 Reviews & Comparisons2 months ago Laguna XS 2.1: Poolside’s Local Coding Agent Tested – Nine Languages Fahd Mirza puts Poolside's newly released Laguna XS 2.1 through a live evaluation using the Hermit agentic framework. The model is a... 0 comments 1.2K views
10:06 Coding & Builds4 months ago DFlash Leaves Qwen Territory – Gemma 4 31B Now Runs 5x Faster with Speculative Decoding Fahd Mirza demonstrates the first end-to-end deployment of Llama Box DFlash with Google's Gemma 4 31B model, following the merge of P... 0 comments 3.5K views
08:43 Tutorials4 months ago DFlash Drafter for Gemma 4 26B – Official Speculative Decoding is Here: Run Locally ZLab, the UC San Diego research team that invented DFlash speculative decoding, has released the first official drafter model paired... 0 comments 572 views
08:41 Tutorials4 months ago Gemma 4 31B at 196 tok/s with RedHat DFlash Speculator Locally This hands-on tutorial from the Fahd Mirza channel demonstrates running Google's Gemma 4 31B model locally at 196 tokens per second u... 0 comments 2.3K views
08:57 Benchmarks4 months ago Google Releases Gemma 4 MTP Drafters – Run Locally and DFlash Comparison Fahd Mirza demonstrates Google's newly released MTP (multi-token prediction) draft models for the Gemma 4 family, running live tests... 0 comments 5.3K views
15:31 Coding & Builds4 months ago PFlash + Qwen3.6-27B-DFlash: 10x Faster Prefill on a Single GPU: Run Locally Fahd Mirza builds and benchmarks PFlash, a prefill acceleration tool that dramatically reduces the blank-screen wait time when feedin... 0 comments 3.8K views