01:16:25 News & Opinion2 months ago Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club This YC Paper Club special edition gathers Stanford researchers and senior engineers for a dense technical panel on GPU kernel optimi... 0 comments 4.2K views
09:16 Coding & Builds2 months ago DeepSeek V4 Flash Fully Local — 32 tok/s on a Single Chip Fahd Mirza demonstrates running DeepSeek V4 Flash, a 284-billion parameter mixture-of-experts model, entirely on a single AMD Ryzen A... 0 comments 5.2K views
09:20 Coding & Builds3 months ago Nanbeige4.2 – 3B Model That Beats 9B Models | Local Install + Real App Test Fahd Mirza tests Nanbeige 4.2, the second release from Nan Beach and a 3-billion-parameter model built on a "looped transformer" arch... 0 comments 1.6K views
08:12 Tutorials3 months ago Gemma 4 Just Got a Massive Update (Tested Live Locally) Fahd Mirza puts Gemma 4's latest patch update under the microscope in a live session running on an Nvidia H100 with vLLM on a fresh U... 0 comments 4.3K views
49:44 Interviews3 months ago The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO Dan Biderman, co-founder and CEO of Engram — the AI memory startup that raised a $98 million seed round — joins the Latent Space podc... 0 comments 802 views
08:13 Deep Dives3 months ago China’s AI Chips: What’s Real and What’s Just a Slide Fahd Mirza cuts through the hype around China's AI chip industry, offering a technically grounded breakdown of what current domestic... 0 comments 1.6K views
13:00 Coding & Builds3 months ago Tess-4-27B + EAGLE-3: Local Reasoning, Nearly 2× Faster Fahd Mirza demonstrates running Tess-4-27B — a reasoning-capable fine-tune of Qwen 3.6 27B by Miguel de Icaza — locally on an NVIDIA... 0 comments 1.2K views
09:19 Coding & Builds3 months ago Lift: Schema-Based PDF Extraction Tested Locally on 10 Languages Fahd Mirza tests Lift, a 9-billion parameter model purpose-built for schema-constrained PDF extraction, deployed locally on Ubuntu wi... 0 comments 0.9K views
08:08 Reviews & Comparisons3 months ago $2000 96GB Huawei GPU vs Nvidia — Is This The End of the Monopoly? Fahd Mirza investigates the Huawei Atlas 300I Duo — a 96 GB AI accelerator currently listed on Alibaba for $2,600–$2,800 — which has... 0 comments 3.9K views
11:07 Coding & Builds3 months ago NVIDIA Puzzle 75B: A 120B Model Squeezed onto ONE GPU NVIDIA's Nemotron Lab 3 Puzzle 75B compresses the 120B-parameter Nemotron Super model down to 75B total parameters — with only 9.3B a... 0 comments 2K views