Summary
Sam Witteveen introduces Ornith 1.0, a new family of open-weight models from Deep Reinforce that takes a fundamentally different approach to agentic coding. Rather than relying on human-designed scaffolding, Ornith models are trained to write their own task-specific harnesses on the fly — a concept Deep Reinforce calls “self-scaffolding LLMs.” The idea extends the 2022 PAL paper’s insight (having models write and execute Python for math tasks) to full agentic scaffolds, with the model learning to do context engineering that developers would otherwise write by hand.
The Ornith 1.0 family comprises four publicly released models fine-tuned from Qwen 3.5 and Gemma 4 bases: a 9B, a 31B, a 35B MoE, and a 397B MoE — all available on Hugging Face. Training uses a two-stage iterative RL process with GRPO, incorporating three layers of reward-hacking prevention: an immutable sandbox environment, a deterministic monitor that penalizes out-of-sandbox behavior, and an LM-as-judge that can veto suspicious rollouts. Benchmarks show the 397B model competitive with Claude Opus and outperforming Qwen 3.7 Max on several coding tasks, with the 9B holding its own against models three times its size.
Witteveen live-tests the 35B MoE, running SVG generation (the classic pelican test), multi-step coding tasks, and comparing outputs with other frontier models. He argues that the self-scaffolding framing — treating the harness as a learnable object rather than a human-authored artifact — represents a meaningful architectural bet on how agentic systems will evolve as models grow more capable.
📺 Source: Sam Witteveen · Published June 26, 2026
🏷️ Format: Review







