23:35 Deep Dives3 months ago Stop Evaluating Models Like It’s the 50s – Alejandro Vidal, Mindmakers Alejandro Vidal, founder of Mind Makers, argues that the AI industry's standard approach to LLM benchmarking—summing correct answers... 0 comments 336 views
23:00 Deep Dives3 months ago Stop Evaluating Models Like It’s the 50s – Alejandro Vidal, Mindmakers Alejandro Vidal of Mindmakers presents a conference talk arguing that LLM benchmarks are fundamentally flawed because they treat ever... 0 comments 0.9K views
02:11:20 Interviews3 months ago Exploring the J-Space The Cognitive Revolution hosts Nathan and Pash spend a live morning session working through Anthropic's major new J-Space interpretab... 0 comments 251 views
25:57 Benchmarks3 months ago I benchmarked the NEW Sonnet 5. The results shocked me. How I AI introduces the Howi AI Bench — a repeatable, multi-dimensional evaluation framework built with Claude Code — and runs Claude... 0 comments 3K views
18:17 Benchmarks4 months ago VibeThinker 3B – Taking on Giant Models Sam Witteveen digs into VibeThinker 3B, a small language model from Waybo AI Lab — the AI research arm of the Chinese social network... 0 comments 4.1K views
01:45:50 Interviews4 months ago Radically Better Reasoning: Elicit’s Andreas Stuhlmüller & Jungwon Byun on World Models for Research Andreas Stuhlmüller and Jungwon Byun, co-founders of Elicit, join the Cognitive Revolution podcast to discuss how their AI platform f... 0 comments 325 views
12:46 Coding & Builds4 months ago VibeThinker-3B: 3B Model That Challenges Claude Opus? Test Locally Fahd Mirza installs and tests VibeThinker-3B — a reasoning model released by Weibo, the Chinese social media giant — directly on an N... 0 comments 2.8K views
19:34 Tutorials4 months ago From Transcription to Live Music: Gemini’s Audio Stack — Thor Schaeff, Google DeepMind Thor Schaeff, developer experience lead on the Gemini API and Google AI Studio at Google DeepMind, walks through the current state of... 0 comments 540 views
28:36 Interviews4 months ago What Happens After A 1,000,000x AI Compute Leap? | Jeff Dean In this Two Minute Papers interview, host Károly Zsolnai-Fehér sits down with Jeff Dean — Google's Chief Scientist, co-creator of Map... 0 comments 18.6K views
14:21 News & Opinion4 months ago The Biggest Lie You’ve Been Told About Hermes Agent Craig Hewitt, who runs two Hermes agents in his own business, pushes back against what he describes as misleading YouTube tutorials a... 0 comments 268 views