Descriptions:
Web Dev Cody runs a structured experiment to determine which AI model is best suited for quick coding questions inside his custom Nebula agent interface — a different use case than deep planning or feature implementation, where speed and accuracy both matter without burning expensive compute.
Using Claude Fable 5 as an orchestrator, he dispatches Cursor’s agent CLI against every available model and effort-level combination, timing each run and scoring output quality with Claude Sonnet 5 as judge. The headline finding: Cursor Composer 2.5 Fast is the sweet spot — scoring 91 on quality while completing in roughly 27 seconds. Gemini 3.6 Flash is the fastest at 14 seconds but hallucinates frequently. Opus 5 takes 90 seconds and scores no better. One notable behavioral observation: older models like Claude Sonnet 5 and Opus 5 ignore explicit instructions not to ask clarifying questions, while the more recent frontier models follow the constraint reliably.
The video also demos Nebula’s new agent preset system, which lets developers bind a model, effort level, and prompt prefix/suffix to a named workflow — switching between a quick-question preset and a full-implementation preset with one click. The experiment methodology and results are shown on screen, giving viewers a replicable framework for running similar evaluations against their own codebases.
📺 Source: Web Dev Cody · Published August 29, 2026
🏷️ Format: Benchmark Test







