16:41 Tutorials1 month ago Run GLM-5.3 Flash Locally on CPU + RAM Fahd Mirza walks through a complete, reproducible tutorial for running GLM-5.3 Flash — ZAI's 320-billion-parameter sparse mixture-of-... 0 comments 8.4K views
11:07 Reviews & Comparisons1 month ago Qwen3.8-Flash-Next: Local Install, Serve, Break, Fix Fahd Mirza takes Qwen3.8-Flash-Next for a hands-on spin the day it drops — an open-weight preview of the upcoming Qwen 4 architecture... 0 comments 3.3K views
08:43 Deep Dives2 months ago Qwen3.8-27B Obliterated: Don’t Use this Model in Production Security-focused AI researcher Fahd Mirza examines "obliterated" open-weight models — specifically Qwen 3.8-27B Obliterated, availabl... 0 comments 4.2K views
11:12 Benchmarks2 months ago Qwen3.8-4B Distilled: Q4 vs Q6 vs Q8 — Which Quant Actually Wins? Fahad Mirza benchmarks the Qwen 3.8 4B distilled model — built by the Emperor team, not Alibaba — across three quantization levels (Q... 0 comments 2.9K views
09:20 Benchmarks2 months ago DFlash 2: Qwen3.8-27B at 2× Speed – Live Benchmark Locally Fahd Mirza benchmarks DFlash 2, a community-developed speculative decoding enhancement for the Qwen3.8-27B model, running live on an... 0 comments 6.5K views
08:06 Coding & Builds2 months ago Loops in Hermes Agent – Hands-on Demo with Qwen3.8 27B Fahd Mirza demonstrates the newly released `/loop` command in Hermes Agent, an open-source agentic framework that gives language mode... 0 comments 1.6K views
19:32 Tutorials2 months ago Qwen3.8-27B GGUF: Run It Local with llama.cpp, Ollama & LM Studio Fahd Mirza walks through the complete process of running Qwen3.8-27B in GGUF quantized format using three local inference tools: LM S... 0 comments 2.2K views
09:28 Tutorials2 months ago DeepSeek Harness + Ollama or Any Other Provider Locally Fahd Mirza walks through the installation and configuration of DeepSeek Harness (DSH), an open-source agent framework quietly release... 0 comments 1.9K views
20:28 Benchmarks2 months ago Muse Glimmer 30B GGUF + DFlash: 3x Faster Local Inference Fahd Mirza follows up his full-precision Muse Glimmer 30B video with a focused benchmark of the GGUF quantized version running with D... 0 comments 1.6K views
08:13 Coding & Builds2 months ago Neutrino-8B Locally: Extreme Quantization Done Right Fahd Mirza takes a close look at Neutrino-8B, a model from Fermi-on that pushes ternary quantization to its logical extreme — every w... 0 comments 2K views