Summary
Ronak Malde, co-founder of Trajectory and formerly the research lead at Windsurf who helped build the Sui 1 model that preceded the company’s $2 billion acquisition by DeepMind, delivers a conference talk on the technical barriers to scaling continual learning and the algorithm Trajectory has built to address them.
Malde’s critique of existing training paradigms is systematic: SFT lacks on-policy sampling, DPO/RLHF recovers online task distributions but uses sequence-level reward, and GRPO — the dominant current approach — achieves strong on-policy rollouts at the cost of off-policy task distributions, massive parallel environment infrastructure, and scalar reward signals that lose per-token information. His proposed solution, Online Policy Self-Distillation (OPSD), replaces the standard RL rollout group with a teacher-student distillation step where the teacher is given the golden solution as guidance, making log-probability supervision tractable from a single example rather than requiring eight parallel rollouts. The result, he argues, is richer per-token feedback, elimination of the environment infrastructure bottleneck, and the ability to shift entire token probability distributions rather than merely sharpening them.
On LiveCodeBench, GRPO saturates around Claude Sonnet-level performance while OPSD pushes into new territory, with token efficiency improving significantly on short-horizon tasks. Malde notes OPSD can be tested today via open-source projects including OpenCLAW RL, and frames continual learning from real-world inference data as the next major scaling frontier after pre-training and RL on benchmarks.
📺 Source: AI Engineer · Published August 12, 2026
🏷️ Format: Keynote Launch







