12:50 Benchmarks1 month ago What’s the Best AI Model for Quick Questions? Web Dev Cody runs a structured experiment to determine which AI model is best suited for quick coding questions inside his custom Neb... 0 comments 5K views
19:02 Benchmarks1 month ago GLM-5.3-Flash: The Ox Alpha Mystery, Finally Tested GLM-5.3-Flash — the model that spent a week circulating anonymously as "Ox Alpha" on OpenRouter — gets its first systematic public ev... 0 comments 1.4K views
11:53 Benchmarks1 month ago Granite 4.2 (3B vs 8B): IBM’s New Reasoning Models, Tested Locally Fahd Mirza runs IBM's newly released Granite 4.2 model family through a practical local benchmark, testing both the 3B and 8B paramet... 0 comments 2.5K views
11:12 Benchmarks2 months ago Qwen3.8-4B Distilled: Q4 vs Q6 vs Q8 — Which Quant Actually Wins? Fahad Mirza benchmarks the Qwen 3.8 4B distilled model — built by the Emperor team, not Alibaba — across three quantization levels (Q... 0 comments 2.9K views
10:30 Benchmarks2 months ago Ornith-1.5-9B: Great on Paper, Struggled in My Tests Locally Fahd Mirza tests the Ornith 1.5 9B model — the smallest and densest member of the Ornith 1.5 family — on a single NVIDIA A100 with 80... 0 comments 1.6K views
17:35 Benchmarks2 months ago I Tested Ornith’s New 35b MoE on Mac — Here’s How It Goes Bart Slodyczka runs Ornith 1.5's 35B mixture-of-experts model through a series of practical one-shot tests on an Apple M3 Ultra Mac S... 0 comments 3.8K views
09:20 Benchmarks2 months ago DFlash 2: Qwen3.8-27B at 2× Speed – Live Benchmark Locally Fahd Mirza benchmarks DFlash 2, a community-developed speculative decoding enhancement for the Qwen3.8-27B model, running live on an... 0 comments 6.5K views
20:28 Benchmarks2 months ago Muse Glimmer 30B GGUF + DFlash: 3x Faster Local Inference Fahd Mirza follows up his full-precision Muse Glimmer 30B video with a focused benchmark of the GGUF quantized version running with D... 0 comments 1.6K views
11:04 Benchmarks2 months ago Maple-Preview Tested: Ternary Weights, Better on CPU than GPU, Honest Results Fahd Mirza independently tests Maple-Preview, a 20-billion-parameter sparse mixture-of-experts model from Deep Grove (an independent... 0 comments 3.1K views
08:24 Benchmarks2 months ago OpenAI vs Claude Polymarket AI Prediciton Battle Week 2 All About AI returns with Week 2 of its ongoing head-to-head prediction battle between OpenAI's GPT-5.6 and Anthropic's Claude Opus 5... 0 comments 308 views
08:29 Benchmarks2 months ago NVIDIA Nemotron 3 Embed 1B Cuts Agent Token Costs by 31% Fahd Mirza puts NVIDIA's newly released Nemotron 3 Embed 1B to a practical cost test, pitting it against Nomic Embed Text in a live a... 0 comments 1.1K views
13:57 Benchmarks2 months ago Qwen3.8-Max: Can It Actually SEE? (Vision & Video Test) Fahd Mirza puts Qwen 3.8 Max through a dedicated multimodal evaluation, moving beyond text benchmarks to test what the model can actu... 0 comments 1.1K views
14:36 Benchmarks2 months ago Whale Wokeup: DeepSeek V4-Flash Is Out of Preview — And It’s Brutal Fahd Mirza puts DeepSeek V4-Flash — freshly released from preview to general availability — through a demanding real-world test using... 0 comments 2.6K views
16:29 Benchmarks2 months ago Opus 5 vs GPT-5.6 On Polymarket Predictions — Week 1 This video launches a recurring head-to-head series pitting Claude Opus 5 against GPT-5.6 on real-world prediction market events sour... 0 comments 577 views
11:15 Benchmarks2 months ago Single Photo vs. Character Sheet: The LTX 2.3 Best Face ID Secret This video from Veteran AI provides a thorough hands-on evaluation of Best Face ID, a character-consistency LoRA designed for the LTX... 0 comments 2.8K views
13:14 Benchmarks3 months ago Qwen-Audio-3.0-TTS Tested: 16 Languages, Instruction Control & Emotion Tags Fahd Mirza puts Qwen Audio 3.0 TTS — released by Alibaba's Tongyi team in Flash and Plus tiers, built on a 12.5 Hz low-frame-rate spe... 0 comments 2.2K views