Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

More

Descriptions:

Mahesh Sathiamoorthy, co-founder and CEO of Bespoke Labs and formerly a researcher at Google DeepMind, presents the company’s open-source work on data and RL environment curation for post-training LLMs at AI Engineer. The talk centers on two flagship contributions: Open Thoughts, a curated reasoning dataset whose pipeline has been cited by John Schulman and used internally at Thinking Machines, and Terminal Bench, a coding agent benchmark Bespoke has co-developed.

Sathiamoorthy walks through the full Open Thoughts curation pipeline: sourcing prompts from existing datasets, using LLMs to score question quality and difficulty, generating answers via teacher models including DeepSeek, Qwen, and Gemini variants, and running systematic ablations at each stage. A counterintuitive finding: sampling multiple answers per question consistently outperforms acquiring more unique questions at the same data budget — a result he says surprised even experienced practitioners. The pipeline demonstrates clear scaling behavior, with benchmark performance improving predictably as dataset size grows.

The broader argument is that the field has shifted from evaluating what models know (static benchmarks like MMLU) to evaluating what agents can do autonomously over extended periods (SWE-bench, Terminal Bench). Sathiamoorthy frames data quality — including RL environment design — as the primary bottleneck separating frontier labs from enterprises, and positions carefully designed synthetic pipelines as the most practical path to closing that gap.


📺 Source: AI Engineer · Published July 31, 2026
🏷️ Format: Deep Dive

1 Item

Channels