Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

More

Summary

The co-founders of Theta Software — one previously a founding engineer at Deep Silken working on ternary models — present a framework for designing RL environments and evaluation systems for long-horizon AI agents at AI Engineer, addressing one of the field’s hardest open problems: how to verify whether an agent’s work is actually correct when tasks extend over hours or days.

The talk opens by reframing ‘long horizon’ as a scalar rather than binary category, using the Meter benchmark’s human-time threshold methodology (e.g., 50% success rate on tasks taking a human 16 hours) alongside token count and trajectory length as complementary proxies. The speakers then introduce a three-dimensional model capability framework: breadth (how much of the environment the agent can explore in parallel, where multi-agent spawning helps), sequential complexity (how much a bad early decision cascades into downstream failures), and ambiguity (how underspecified the initial task and artifacts are). Increasing ambiguity better mirrors real human work but makes standardized evaluation significantly harder because correct solutions become non-unique.

The second half focuses on verifier design — the core bottleneck for moving from math and competitive coding to economically meaningful software work. Unlike provable domains where test suites or proof checkers suffice, file-editing and system-configuration tasks require judge or critic models, and the speakers walk through tradeoffs between deterministic and learned verifiers across varying task complexity levels.


📺 Source: AI Engineer · Published August 01, 2026
🏷️ Format: Deep Dive

1 Item

Channels