Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI

Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI

More

Descriptions:

At the AI Engineer conference, Lee Robinson — machine learning engineer for model behavior at Cursor — walks through how Cursor trains its own models and introduces the concept of recursive model improvement: a self-reinforcing loop where deployed models generate feedback data that trains better models, which in turn create harder problems for the next training run.

Robinson describes two nested loops. The outer loop collects user feedback (thumbs up/down in the product), runs A/B tests across model checkpoints, and translates that signal into higher-quality evals and more ambitious training environments. The inner loop uses reinforcement learning at scale, with Cursor generating synthetic environments by taking complex codebases, deleting features or files, and asking models to restore passing tests — providing a verifiable reward signal. A key technique called textual feedback lets a teacher model zoom into specific moments in a long RL rollout (which can span hundreds of thousands of tokens) and nudge the model’s probability distribution at precisely the decision point that went wrong, rather than assigning credit at the end of the full trajectory.

Robinson reports that Composer 2.5, released in May, is now Cursor’s most popular model — valued by users for combining speed, intelligence, and cost-effectiveness. The team’s ambitions for the next version include training a full pre-train from scratch (moving away from the Kimi open-source base), scaling up both data and compute, and building toward a more general-purpose model beyond pure coding tasks.


📺 Source: AI Engineer · Published July 15, 2026
🏷️ Format: Keynote Launch

1 Item

Channels

1 Item

Companies