Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning

More

Descriptions:

Ross Taylor, CEO of General Reasoning and former reasoning lead at Meta AI, co-presents with co-founder Chengxi at AI Engineer on what it takes to scale AI agents to long-horizon tasks. Taylor opens with a candid first-person account of Galactica’s 2022 launch — a scientifically capable base model that failed publicly just weeks before ChatGPT demonstrated precisely why RLHF post-training matters. The contrast between a strong base model and a human-preference-aligned product, Taylor argues, was the field’s first real natural experiment proving RL’s practical value.

The talk then reveals an unpublished Meta recipe called RSP — combining continued pre-training of Llama 2 on math and science data with PPO using verifiable rewards and a value model initialized from a strong outcome reward model. Internally this achieved state-of-the-art math results, yet lacked the reflective backtracking behavior that later emerged in DeepSeek-R1 and OpenAI o1. Taylor attributes that gap to the bitter lesson in its purest form: better base models, larger context windows, and more RL compute were necessary prerequisites that simply weren’t available at the time.

Chengxi closes with a forward-looking section on General Reasoning’s current work — covering algorithms, environment design, and compute strategies needed to push agents beyond short-horizon tasks toward multi-hour and multi-day autonomous operation. This is an unusually candid insider account from researchers who were present at the origin of open-weight model post-training.


📺 Source: AI Engineer · Published July 31, 2026
🏷️ Format: Keynote Launch

1 Item

Channels