01:53:27 Foundation Models4 months ago Why AI’s “12-Hour” Task Number Is a Mirage — Beth Barnes & David Rein Machine Learning Street Talk hosts an in-depth technical conversation with Beth Barnes and David Rein, two researchers at Metr — the... 0 comments 2.8K views
01:46:33 Interviews5 months ago Success without Dignity? Nathan finds Hope Amidst Chaos, from The Intelligence Horizon Podcast Nathan Labenz of the Cognitive Revolution podcast appears as a guest on the Intelligence Horizon podcast, hosted by Yale seniors Owen... 0 comments 421 views
01:18:46 Foundation Models5 months ago AI Scouting Report: the Good, Bad, & Weird @ the Law & AI Certificate Program, by LexLab, UC Law SF Nathan Labenz of The Cognitive Revolution podcast delivered this wide-ranging AI landscape survey at UC Law San Francisco's Law and A... 0 comments 614 views
01:05:12 Interviews6 months ago Measuring Exponential Trends Rising (in AI) — Joel Becker, METR Joel Becker from METR (Model Evaluation and Threat Research) joins Latent Space to discuss his organization's work quantifying AI cap... 0 comments 5.5K views
24:44 Foundation Models6 months ago the SCARIEST chart in AI Wes Roth breaks down what he calls \"the scariest chart in AI development history\" — a METR (Meter Research) benchmark tracking AI a... 0 comments 80.7K views
42:15 Foundation Models6 months ago The 5 Levels of AI Coding (Why Most Won’t Make It Past Level 2) Nate B Jones builds an extended analysis around the \"five levels of AI coding\" framework published by Glowforge CEO Dan Shapiro in... 0 comments 229.2K views
01:15:52 Foundation Models7 months ago How METR measures Long Tasks and Experienced Open Source Dev Productivity – Joel Becker, METR Joel Becker from METR (Model Evaluation and Threat Research) presents the organization's framework for measuring AI agent task horizo... 0 comments 9.4K views
21:22 Foundation Models8 months ago Why Agent Hype can fall short of reality – Joel Becker, METR Joel Becker, a researcher at METR (Model Evaluation and Threat Research), presents two empirical studies that together expose a strik... 0 comments 7.6K views