20:28 Benchmarks1 month ago Muse Glimmer 30B GGUF + DFlash: 3x Faster Local Inference Fahd Mirza follows up his full-precision Muse Glimmer 30B video with a focused benchmark of the GGUF quantized version running with D... 0 comments 1.6K views
09:22 Reviews & Comparisons1 month ago Nemotron Lightning – NVIDIA’s Super Fast Agent MoE Sam Witteveen reviews NVIDIA's NeMo Tron 3.5 Lightning, a 30B-parameter mixture-of-experts model with only 3B active parameters, purp... 0 comments 3.6K views
59:10 Interviews2 months ago The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO Latent Space hosts Modal CTO Akshat Bubna for an in-depth conversation on how Modal evolved from a general serverless runtime into a... 0 comments 2.1K views
09:39 Coding & Builds2 months ago DeepSeek DFlash on Gemma 12B Locally: Up To 5x Faster Fahd Mirza demonstrates how to run DeepSeek's DFlash speculative decoding method locally, pairing the open-source DeepSeek drafter mo... 0 comments 1.8K views
08:49 Deep Dives3 months ago DSpark – DeepSeek Just Made Inference 85% Faster DeepSeek has released DSpark, a speculative decoding system that makes their models generate text 60 to 85% faster without any change... 0 comments 3.6K views
09:40 Benchmarks3 months ago DFlash Just Got Faster: 4x Speed with 160 tok/s Locally Fahd Mirza benchmarks DFlash with SGLang's new SpecV2 overlapping scheduler on an NVIDIA H100 80GB GPU, demonstrating a 4.3x throughp... 0 comments 2K views
12:30 Coding & Builds3 months ago Luce KVFlash: Fit 256K Context on a Small GPU – Local Hands-On Guide KV Flash is a new memory management engine for local LLM inference that keeps only the most relevant tokens on GPU VRAM while paging... 0 comments 2.3K views
13:13 Tutorials3 months ago Adaptive PFlash + Hermes Agent – Self-Tuning Prefill on a Single GPU Locally Fahd Mirza demonstrates the newly shipped adaptive compression feature in PFlash, the prefill-acceleration component of the open-sour... 0 comments 2.2K views
08:41 Tutorials4 months ago Luce DFlash Meets OpenClaw – Local AI Agents at 2x Speed with Qwen3.6-27B Fahd Mirza walks through a complete, reproducible integration of DFlash — a speculative decoding inference engine — with OpenClaw, an... 0 comments 0.9K views
09:45 Tutorials4 months ago TurboQuant + DFlash: Supercharge Local LLM Speed Fahd Mirza demonstrates the practical integration of two recently released local inference tools: Google Research's TurboCore KV cach... 0 comments 2.5K views