Robot-Use Agents: Why General-Purpose Models May Win in Robotics

Robot-Use Agents: Why General-Purpose Models May Win in Robotics

More

Summary

This Y Combinator conversation explores “robot-use agents,” the idea that general-purpose language models, much like coding agents, can control robots. It follows a recent essay by MIT professor Phillip Isola and features founders from two startups: Waddle Labs, which builds a harness that lets LLMs control robots and trains better models on the resulting data, and RoboCurve, an evaluation company for physical AI.

The guests trace the research that led here, starting with Google’s RT-2 paper, which established vision-language-action models by fine-tuning a web-trained model to output end-effector poses. They compare that with today’s frontier models, which can control hardware out of the box, and draw parallels to the chain-of-thought moment in math reasoning and to the “bitter lesson” that more autonomy and better data matter more than architecture.

A large part of the discussion covers in-context learning as a low-cost way to adapt a robot to a new task. The hosts note that gains can be non-monotonic and saturate after roughly 20 to 40 examples, and that they are limited by the context length the model was trained on. The conversation also considers retrieval, LoRA and full fine-tuning, and how harnesses and evals are used to measure performance across arms, grippers, humanoids and quadrupeds.


📺 Source: Y Combinator · Published September 26, 2026
🏷️ Format: Interview

1 Item

Channels