17:49 Benchmarks1 month ago Kimi K3 vs Fable 5 vs GLM 5.2 – An Unforgettable Showdown Fahd Mirza runs three frontier AI models through a demanding single-file coding gauntlet: simulate a full concrete plant network—truc... 0 comments 1.3K views
10:03 Benchmarks1 month ago Qwythos-9B-v2: The Looping Bug Gone – FTPO Explained + Local Testing Fahd Mirza tests Qwythos 9B v2 (also called Qutoz) — a 9-billion-parameter fine-tune of Qwen 2 with deep chain-of-thought reasoning —... 0 comments 1.2K views
08:51 Benchmarks2 months ago Qwen3.6-27B with Thinking Cap on: Same Accuracy, 36% Less Thinking Fahd Mirza runs a live, unscripted comparison between Qwen3.6-27B and a community fine-tune called Thinking Cap — a variant of the sa... 0 comments 4K views
10:15 Benchmarks2 months ago How-To Use Fable 5 Cheaper Anywhere in World (82% Savings) Hands-on Demo Fahd Mirza demonstrates two practical strategies for cutting the cost of Claude Fable 5 by up to 82%, running live comparisons across... 0 comments 1.7K views
08:18 Benchmarks2 months ago Qwopus 35B + MTP: The Coder That Fixes Its Own Bugs at 160 tok/s Fahd Mirza tests Qwopus Coder, a 35-billion-parameter mixture-of-experts coding model built on the Qwen 3.6 architecture (3B paramete... 0 comments 1.7K views
25:57 Benchmarks2 months ago I benchmarked the NEW Sonnet 5. The results shocked me. How I AI introduces the Howi AI Bench — a repeatable, multi-dimensional evaluation framework built with Claude Code — and runs Claude... 0 comments 2.9K views
30:52 Benchmarks2 months ago Frontier results, on device – RL Nabors, Arize Rachel Lee Nabors — formerly at Mozilla on Firefox DevTools, the W3C, Microsoft Edge, and the React team, now at Arize — presents a p... 0 comments 2.1K views
13:57 Benchmarks2 months ago Can Krea 2 Turbo Really Make Great Images in 8 Steps? ComfyUI Test Veteran AI runs a structured eight-category evaluation of Krea 2 Turbo — the eight-step distilled image generation model released by... 0 comments 1.1K views
14:08 Benchmarks2 months ago Qwythos 9B: When You Train a Small Model on Claude Traces: Run Locally Fahd Mirza introduces and benchmarks Qwythos 9B, a reasoning-focused open-source model fine-tuned on over 500 million tokens of Claud... 0 comments 2.9K views
09:36 Benchmarks2 months ago Qwen3.6 (REAP 90pct GGUF): The Brain-Damaged Model Fahd Mirza takes a deep look at an aggressively pruned variant of Qwen 3.6 — a 35-billion-parameter mixture-of-experts model — compre... 0 comments 2.8K views
18:17 Benchmarks2 months ago VibeThinker 3B – Taking on Giant Models Sam Witteveen digs into VibeThinker 3B, a small language model from Waybo AI Lab — the AI research arm of the Chinese social network... 0 comments 4.1K views
08:20 Benchmarks2 months ago LoopCoder – The 7B Model That Thinks Twice – Does it Beat Others? LoopCoder V2 is a 7-billion-parameter open-source code model built on an unusual architectural idea: instead of stacking more transfo... 0 comments 2.3K views
09:40 Benchmarks2 months ago DFlash Just Got Faster: 4x Speed with 160 tok/s Locally Fahd Mirza benchmarks DFlash with SGLang's new SpecV2 overlapping scheduler on an NVIDIA H100 80GB GPU, demonstrating a 4.3x throughp... 0 comments 2K views
31:25 Benchmarks2 months ago Claude Fable 5 BANNED: The First Model Agentic Engineers DON’T NEED IndyDevDan covers two intertwined stories in this video: the sudden federal suspension of Claude Fable 5 and Mythos 5, and a detailed... 0 comments 7.6K views
09:03 Benchmarks2 months ago I tried to prove AI trading is BS and it backfired The Algovibes channel set out to definitively disprove AI-powered crypto trading — and ended up with results more interesting than ex... 0 comments 2.4K views
09:13 Benchmarks3 months ago I Tested 100,000 Trading Strategies on 1,000 Stocks Algovibes presents a large-scale systematic backtesting study covering the full Russell 1000 universe — 1,014 stocks, 66 technical tr... 0 comments 1.2K views