01:19:25 Interviews2 days ago Why companies are becoming a series of loops | Anish Acharya (a16z) Anish Acharya, general partner at Andreessen Horowitz where he focuses on consumer investing, joins Lenny's Podcast to explore how AI... 0 comments 6.7K views
08:11 Benchmarks7 days ago JetSpec Locally: Breaking the Speed Ceiling of LLM Inference – Up to 9x Fahd Mirza installs JetSpec — a speculative decoding framework from UCSD — and benchmarks it locally on an H100 GPU running Qwen 3 8B... 0 comments 7.4K views
18:25 Coding & Builds4 months ago Your Coding Agent Should Do AI System Engineering — Ben Burtenshaw, Hugging Face Ben Burtenshaw, an engineer at Hugging Face, makes the case that coding agents have crossed a capability threshold where they can now... 0 comments 3.2K views
08:28 Coding & Builds4 months ago Qwen3-8B at 74 tok/s with RedHat DFlash Speculator on vLLM Locally Fahd Mirza walks through running Red Hat's DFlash speculative decoding implementation on Qwen3-8B using vLLM, achieving 74 tokens per... 0 comments 1.7K views
08:37 Deep Dives5 months ago Ternary Bonsai: The Tiny Model That Should Not Be This Good Fahd Mirza covers Prism ML's Ternary Bonsai, the latest release from the team behind the one-bit Bonsai model, which pushes ultra-low... 0 comments 1.6K views
08:33 Coding & Builds6 months ago Qwen3 Speculator Eagle: Red Hat Made Qwen3-8B 6x Faster: Full Hands-on Guide Red Hat has quietly entered the AI inference space with a significant technical contribution: a speculative decoding model that makes... 0 comments 7.8K views