11:12 Benchmarks2 days ago Qwen3.8-4B Distilled: Q4 vs Q6 vs Q8 — Which Quant Actually Wins? Fahad Mirza benchmarks the Qwen 3.8 4B distilled model — built by the Emperor team, not Alibaba — across three quantization levels (Q... 0 comments 2.8K views
10:30 Benchmarks4 days ago Ornith-1.5-9B: Great on Paper, Struggled in My Tests Locally Fahd Mirza tests the Ornith 1.5 9B model — the smallest and densest member of the Ornith 1.5 family — on a single NVIDIA A100 with 80... 0 comments 1.5K views
17:35 Benchmarks4 days ago I Tested Ornith’s New 35b MoE on Mac — Here’s How It Goes Bart Slodyczka runs Ornith 1.5's 35B mixture-of-experts model through a series of practical one-shot tests on an Apple M3 Ultra Mac S... 0 comments 3.7K views
09:20 Benchmarks5 days ago DFlash 2: Qwen3.8-27B at 2× Speed – Live Benchmark Locally Fahd Mirza benchmarks DFlash 2, a community-developed speculative decoding enhancement for the Qwen3.8-27B model, running live on an... 0 comments 6.5K views
20:28 Benchmarks2 weeks ago Muse Glimmer 30B GGUF + DFlash: 3x Faster Local Inference Fahd Mirza follows up his full-precision Muse Glimmer 30B video with a focused benchmark of the GGUF quantized version running with D... 0 comments 1.6K views
11:04 Benchmarks2 weeks ago Maple-Preview Tested: Ternary Weights, Better on CPU than GPU, Honest Results Fahd Mirza independently tests Maple-Preview, a 20-billion-parameter sparse mixture-of-experts model from Deep Grove (an independent... 0 comments 3K views
08:24 Benchmarks3 weeks ago OpenAI vs Claude Polymarket AI Prediciton Battle Week 2 All About AI returns with Week 2 of its ongoing head-to-head prediction battle between OpenAI's GPT-5.6 and Anthropic's Claude Opus 5... 0 comments 263 views
08:29 Benchmarks3 weeks ago NVIDIA Nemotron 3 Embed 1B Cuts Agent Token Costs by 31% Fahd Mirza puts NVIDIA's newly released Nemotron 3 Embed 1B to a practical cost test, pitting it against Nomic Embed Text in a live a... 0 comments 1K views
13:57 Benchmarks3 weeks ago Qwen3.8-Max: Can It Actually SEE? (Vision & Video Test) Fahd Mirza puts Qwen 3.8 Max through a dedicated multimodal evaluation, moving beyond text benchmarks to test what the model can actu... 0 comments 1.1K views
14:36 Benchmarks3 weeks ago Whale Wokeup: DeepSeek V4-Flash Is Out of Preview — And It’s Brutal Fahd Mirza puts DeepSeek V4-Flash — freshly released from preview to general availability — through a demanding real-world test using... 0 comments 2.5K views
16:29 Benchmarks4 weeks ago Opus 5 vs GPT-5.6 On Polymarket Predictions — Week 1 This video launches a recurring head-to-head series pitting Claude Opus 5 against GPT-5.6 on real-world prediction market events sour... 0 comments 533 views
11:15 Benchmarks4 weeks ago Single Photo vs. Character Sheet: The LTX 2.3 Best Face ID Secret This video from Veteran AI provides a thorough hands-on evaluation of Best Face ID, a character-consistency LoRA designed for the LTX... 0 comments 2.7K views
13:14 Benchmarks1 month ago Qwen-Audio-3.0-TTS Tested: 16 Languages, Instruction Control & Emotion Tags Fahd Mirza puts Qwen Audio 3.0 TTS — released by Alibaba's Tongyi team in Flash and Plus tiers, built on a 12.5 Hz low-frame-rate spe... 0 comments 2.1K views
21:31 Benchmarks1 month ago Is Kimi K3 Really That Good?! (Don’t Just Believe The Hype) Cole Medin takes a critical look at Kimi K3, Moonshot AI's recently released open-weight model that published benchmarks claim rivals... 0 comments 1.2K views
10:49 Benchmarks1 month ago Ling 3.0 Flash: A Production-Scale Coding Agentic Model Fahd Mirza tests Ling 3.0 Flash, a 124-billion-parameter mixture-of-experts model from Ant Group (Alibaba's financial subsidiary), as... 0 comments 1.9K views
08:48 Benchmarks1 month ago Catmind-1.2b: A Reasoning Model that Thinks in Cat Stories Fahd Mirza tests CatMind-1.2B, a fine-tune of the LFM 2.5 1.2B thinking model in which all reasoning traces have been replaced with i... 0 comments 863 views