10:06 Coding & Dev Tools3 months ago DFlash Leaves Qwen Territory – Gemma 4 31B Now Runs 5x Faster with Speculative Decoding Fahd Mirza demonstrates the first end-to-end deployment of Llama Box DFlash with Google's Gemma 4 31B model, following the merge of P... 0 comments 3.5K views
08:53 Research & Benchmarks3 months ago $400 Chinese GPU That Wants to Dethrone NVIDIA Fahd Mirza takes a close look at the Lision LX7G 100, a roughly $485 consumer GPU developed entirely in China without CUDA, AMD archi... 0 comments 3.1K views
01:47:27 Interviews3 months ago we are NOT PREPARED for the end of 2026 Wes Roth and co-host Dylan deliver a wide-ranging AI industry podcast covering the most significant developments from the week of Goo... 0 comments 17.5K views
18:25 Coding & Dev Tools3 months ago Your Coding Agent Should Do AI System Engineering — Ben Burtenshaw, Hugging Face Ben Burtenshaw, an engineer at Hugging Face, makes the case that coding agents have crossed a capability threshold where they can now... 0 comments 3.2K views
09:52 Benchmarks3 months ago Luce Megakernel — 25x Faster Than PyTorch on a Single GPU – Test Locally A new open-source project called Luce Megakernel is challenging long-held assumptions about GPU inference efficiency by fusing all 24... 0 comments 3K views
09:01 Coding & Dev Tools4 months ago Running a 27B model at 130 tokens sec on a single GPU Locally with Luce DFlash LlamaDeFlash is a custom inference engine built from scratch in C++ and CUDA — no vLLM, no llama.cpp, no Python in the critical path... 0 comments 7.7K views
10:01 Foundation Models4 months ago The Hidden Engine Behind DeepSeek V4 – DeepEP V2 and TileKernels Explained While most coverage of DeepSeek V4 focuses on benchmark scores, Fahd Mirza goes a level deeper to explain the two open-sourced infras... 0 comments 562 views
10:33 Coding & Dev Tools4 months ago Kimi FlashKDA: 2x Faster AI Prefill — Installed, Explained and Tested Locally Fahd Mirza walks through the live installation of Flash KDA, Moonshot AI's open-source CUDA kernel that accelerates the prefill phase... 0 comments 1.4K views
01:43:13 Interviews4 months ago Jensen Huang – TPU competition, why we should sell chips to China, & Nvidia’s supply chain moat Dwarkesh Patel sits down with Nvidia CEO Jensen Huang for one of the most substantive interviews the company's founder has given on N... 0 comments 15.2K views
16:23 Tutorials4 months ago 3 Steps to Train Perfect LTX 2.3 Video LoRAs|How to Train Custom LTX 2.3 LoRAs (Video + Audio!) Veteran AI presents a detailed three-part guide to training custom character LoRAs for LTX Video 2.3, covering dataset preparation, t... 0 comments 2.2K views