35:25 Interviews4 months ago Fable 5 Raises the Bar for AI Ambition The AI Daily Brief delivers a detailed breakdown of Anthropic's Claude Fable 5 launch — the first model in Anthropic's new Mythos cla... 0 comments 7.7K views
27:56 Reviews & Comparisons4 months ago Claude Fable 5 is here! The AI Search channel puts Claude Fable 5 through a broad capability gauntlet, opening with a hidden-frog image challenge that the mo... 0 comments 83.4K views
56:19 Interviews4 months ago How This Ex-Meta L8 Engineer Ships 40 PRs a Day with AI Agents | Kun Chen Peter Yang interviews Kun Chen, a former Meta L8 engineer turned solo AI builder, who has developed a systematic method for shipping... 0 comments 1.6K views
29:05 Interviews4 months ago Emergent: How Six Months of Tinkering Led To A $100M ARR Company In this Y Combinator interview filmed in India, Emergent co-founder and CEO shares the story behind one of the fastest-growing AI com... 0 comments 5.7K views
19:04 Deep Dives4 months ago Evals Are Broken, Use Them Anyway — Ara Khan, Cline Ara Khan, an engineer on the Cline team, delivers a pointed critique of how the AI industry uses evaluation benchmarks — and why most... 0 comments 1.4K views
16:30 Deep Dives4 months ago SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius Ibragim Badertdinov, an AI researcher at Nebius with an unconventional background—a trained dentist turned NeurIPS and ICML author—pr... 0 comments 0.9K views
23:25 Deep Dives4 months ago The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI Vincent Chen, research fellow and co-founder at Snorkel AI, took the stage at AI Engineer to share meta-level lessons on what separat... 0 comments 453 views
15:12 Benchmarks4 months ago Can LLMs generate Enterprise Quality Code? — Prasenjit Sarkar, Sonar Prasenjit Sarkar from Sonar presents an enterprise-focused LLM code quality evaluation that goes substantially beyond standard SWE-be... 0 comments 648 views
25:27 News & Opinion4 months ago The Annual AI Slowdown Panic Is Here The AI Daily Brief examines a new coding benchmark called DeepSWE from a company called Data Curve, which is drawing wide attention f... 0 comments 2.9K views
20:03 News & Opinion5 months ago Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind Nicholas Kang, product manager for Kaggle Benchmarks at Google DeepMind, and Michael Aaron, a Kaggle software engineer, present the c... 0 comments 796 views