Descriptions:
Varun Singh, pre-training lead at Arcee AI, argues that the traditional base model — defined by massive web-text ingestion as in GPT-3’s mix of Common Crawl, WebText-2, and Wikipedia (roughly 85% of tokens) — is giving way to a fundamentally different training paradigm. The talk traces the historical arc from GPT-3 through Llama 3 (still ~50% general web text) to the present, where reinforcement learning is no longer a cosmetic post-training step but a core capability driver.
The shift was catalyzed by OpenAI’s o1 and DeepSeek R1, which demonstrated that RL could confer deep reasoning capability rather than merely shape conversational style. This raises the question of what an ideal pre-training prior looks like for large-scale RL — and Singh presents two contrasting answers from recent papers. MAI Thinking 1 deliberately avoided all synthetic data, filtering web crawls to preserve human-generated signal. NeMo Tron 3 Ultra took the opposite approach, pulling SFT-format conversational data back into the pre-training phase, with the top data buckets in its mix labeled with an SFT prefix. Kimi K2 applied similar synthetic rephrasing at scale.
Singh’s own work on Arcee’s Trinity Lodge model used large-scale synthetic rephrasing — taking seed data items and upsampling them by generating multiple paraphrases — to improve token quality and task-shape alignment from the earliest training stages. The talk frames this as an industry-wide inflection: code now dominates pre-training data mixes where it barely appeared in GPT-3, and the question is no longer how big to make the base model but how to structure the entire training continuum for agentic, reasoning-heavy downstream use.
📺 Source: AI Engineer · Published July 31, 2026
🏷️ Format: Deep Dive







