29:57 Reviews & Comparisons3 months ago New #1 open-source AI model is here! AI Search puts GLM 5.2 — the latest release from ZAI — through a series of demanding real-world tasks, claiming it currently tops ope... 0 comments 162.3K views
46:59 Interviews3 months ago China’s New AI Breakthrough Has Silicon Valley Nervous | China Decode China Decode hosts Alice Han and James King examine a significant milestone in the global physical AI race: Chinese startup Spirit AI... 0 comments 22.2K views
28:17 Reviews & Comparisons3 months ago MYTHOS MYTHOS MYTHOS Matthew Berman shares firsthand observations from early access to Claude Fable 5, which Anthropic describes as a Mythos-class model m... 0 comments 25.5K views
19:04 Deep Dives4 months ago Evals Are Broken, Use Them Anyway — Ara Khan, Cline Ara Khan, an engineer on the Cline team, delivers a pointed critique of how the AI industry uses evaluation benchmarks — and why most... 0 comments 1.4K views
23:25 Deep Dives4 months ago The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI Vincent Chen, research fellow and co-founder at Snorkel AI, took the stage at AI Engineer to share meta-level lessons on what separat... 0 comments 427 views
20:40 Deep Dives4 months ago Task Fidelity Scaling Laws — Kobie Crawdord, Snorkel Kobie Crawford, developer advocate at Snorkel AI, presents original research from the company's frontier AI data lab quantifying how... 0 comments 356 views
23:46 News & Opinion4 months ago First Impressions of the New Opus 4.8 The AI Daily Brief leads with a striking enterprise story: Kirkland & Ellis, the world's largest law firm with $10.6 billion in 2025... 0 comments 9K views
09:39 Coding & Builds4 months ago Step 3.7 Flash – 198B Open Source Model That Does Everything; Does it Really? Step 3.7 Flash is a 198 billion parameter sparse mixture-of-experts model from Step One, activating only 11 billion parameters per to... 0 comments 1.8K views
25:27 News & Opinion4 months ago The Annual AI Slowdown Panic Is Here The AI Daily Brief examines a new coding benchmark called DeepSWE from a company called Data Curve, which is drawing wide attention f... 0 comments 2.9K views
12:26 Reviews & Comparisons4 months ago Everyone Is Sleeping on Composer 2.5 Web Dev Cody shares a hands-on assessment of Composer 2.5 after integrating it into real development work on his Mission Control proj... 0 comments 1.5K views