10:06 Coding & Builds5 months ago DFlash Leaves Qwen Territory – Gemma 4 31B Now Runs 5x Faster with Speculative Decoding Fahd Mirza demonstrates the first end-to-end deployment of Llama Box DFlash with Google's Gemma 4 31B model, following the merge of P... 0 comments 3.5K views
08:53 Reviews & Comparisons5 months ago $400 Chinese GPU That Wants to Dethrone NVIDIA Fahd Mirza takes a close look at the Lision LX7G 100, a roughly $485 consumer GPU developed entirely in China without CUDA, AMD archi... 0 comments 3.2K views
29:59 Interviews5 months ago ⚡️ Google’s Open AI Strategy — Omar Sanseviero, Google DeepMind In this Latent Space podcast interview, Omar Sanseviero from Google DeepMind walks through the technical decisions and strategic thin... 0 comments 395 views
08:11 News & Opinion5 months ago Weekly AI Recap – Qwen3.7, MTP in llama.cpp, SANA and More | May 2026 Fahd Mirza's weekly AI recap for May 2026 covers the most consequential model releases, infrastructure updates, and industry deals of... 0 comments 0.9K views
09:01 Tutorials5 months ago Llama.cpp Router Mode: Switch Models Instantly: Hands-on Local Demo Fahd Mirza demonstrates llama.cpp's built-in router mode, a native feature that enables instant model hot-swapping without third-part... 0 comments 2.2K views
10:48 Tutorials5 months ago LM Studio Just Got MTP — Qwen3.6-27B Runs 63% Faster with One Toggle Fahd Mirza demonstrates how to enable Multi-Token Prediction (MTP) speculative decoding in LM Studio's new beta release (version 0.4.... 0 comments 5.9K views
09:45 Tutorials5 months ago Llama.cpp Just Got MTP – Qwen3.6 27B Runs 2x Faster Locally with Two Flags Multi-token prediction (MTP) support has officially merged into the mainline llama.cpp repository—not a fork or custom branch, but th... 0 comments 3.1K views
14:24 News & Opinion5 months ago AI Dev 26 x SF | Anush Elangovan: Impact of AI on Software Anush Elangovan, VP of Software at AMD, delivered a keynote at the AI Dev 26 x SF conference hosted by DeepLearning.AI, sharing how h... 0 comments 292 views
15:56 Reviews & Comparisons5 months ago MiniCPM-V 4.6: The Agent Vision Model Sam Witteveen examines MiniCPM-V 4.6, a 1.3 billion parameter vision-language model released by OpenBMB—a joint initiative between AI... 0 comments 2.4K views
08:06 Reviews & Comparisons5 months ago MTP vs DFlash — Speculative Decoding Explained Simply This video by Fahd Mirza offers a clear, structured comparison of two speculative decoding techniques — Multi-Token Prediction (MTP)... 0 comments 1.2K views