09:01 Foundation Models4 weeks ago NVIDIA’s Two-Tower Model Generates Text 2.4x Faster Without Losing Quality NVIDIA has released NeMo-Tron 2 Tower 30B A3 Base, a 30-billion parameter model that rethinks how large language models generate text... 0 comments 1.1K views
08:18 Benchmarks4 weeks ago Qwopus 35B + MTP: The Coder That Fixes Its Own Bugs at 160 tok/s Fahd Mirza tests Qwopus Coder, a 35-billion-parameter mixture-of-experts coding model built on the Qwen 3.6 architecture (3B paramete... 0 comments 1.7K views
08:11 Business & Strategy2 months ago Weekly AI Recap – Qwen3.7, MTP in llama.cpp, SANA and More | May 2026 Fahd Mirza's weekly AI recap for May 2026 covers the most consequential model releases, infrastructure updates, and industry deals of... 0 comments 802 views
10:48 Tutorials2 months ago LM Studio Just Got MTP — Qwen3.6-27B Runs 63% Faster with One Toggle Fahd Mirza demonstrates how to enable Multi-Token Prediction (MTP) speculative decoding in LM Studio's new beta release (version 0.4.... 0 comments 5.7K views
09:45 Tutorials2 months ago Llama.cpp Just Got MTP – Qwen3.6 27B Runs 2x Faster Locally with Two Flags Multi-token prediction (MTP) support has officially merged into the mainline llama.cpp repository—not a fork or custom branch, but th... 0 comments 3K views
08:06 Research & Benchmarks2 months ago MTP vs DFlash — Speculative Decoding Explained Simply This video by Fahd Mirza offers a clear, structured comparison of two speculative decoding techniques — Multi-Token Prediction (MTP)... 0 comments 1.1K views
11:12 Benchmarks3 months ago Qwen3.6 27B Gets 20% Faster with MTP and llama.cpp Locally Fahd Mirza demonstrates how to enable multi-token prediction (MTP) on Qwen3.6 27B using ik_llama.cpp — a community fork of the popula... 0 comments 3.3K views