11:00 Coding & Dev Tools3 days ago Gemma 4’s Big Update in 4-Bit QAT — Low VRAM, Local Fahd Mirza covers the most significant update to Google's Gemma 4 model to date: 11 targeted fixes spanning tool calling, chat format... 0 comments 2.1K views
11:07 Coding & Dev Tools2 weeks ago NVIDIA Puzzle 75B: A 120B Model Squeezed onto ONE GPU NVIDIA's Nemotron Lab 3 Puzzle 75B compresses the 120B-parameter Nemotron Super model down to 75B total parameters — with only 9.3B a... 0 comments 1.9K views
09:52 Coding & Dev Tools3 weeks ago Run DeepSeek DSpark on Qwen3 Locally and Reproduce the Speedup Fahd Mirza walks through a complete hands-on reproduction of DeepSeek's DeepSpark speculative decoding results on a single NVIDIA RTX... 0 comments 2.9K views
16:47 Research & Benchmarks1 month ago Google QAT vs Unsloth QAT + MTP – Which Gemma 4 12B Is Actually Better? This video pits two quantized versions of Google's Gemma 4 12B against each other in a practical, locally-run benchmark: Google's own... 0 comments 3.1K views
13:08 Tutorials1 month ago Gemma 4 12B QAT + MTP on llama.cpp Locally – Twice the Speed, Same Quality? This video by Fahd Mirza walks through running Google's newly released Gemma 4 12B QAT (Quantization-Aware Training) model alongside... 0 comments 2.8K views