I Tested Opus 5.5 vs. GPT-6 Astra on 12 Real Use Cases

I Tested Opus 5.5 vs. GPT-6 Astra on 12 Real Use Cases

More

Summary

Nate Herk puts Anthropic’s Opus 5.5 and OpenAI’s GPT-6 Astra head to head on 12 real-world use cases, running identical prompts in each model’s own harness at high effort. The tasks include building websites, editing video and Instagram reels, planning trip itineraries, producing branded slide decks, spreadsheets and P&Ls, generating carousels, testing browser use, and creating playable 3D games and learning experiences.

For every experiment, the video reports how long each model took and what it would have cost at API rates, even though the tests ran on subscription plans. Astra is about 2.5 times more expensive than Opus 5.5, which frames the central question: does it deliver 2.5 times the results?

In the examples shown, Opus 5.5 often came out ahead. On a landing-page build, Herk found the Opus design felt more premium and told a better story, while also finishing about seven minutes faster for roughly a dollar more. In a spreadsheet task, Opus used far more formulas. A browser game challenge, escaping a miniature museum, shows both models attempting lighting, puzzles, sound and story.

Herk notes that quality judgments are subjective, and he corrects a flaw in his previous Opus 5.5 versus GPT-6 Soul comparison by isolating the agents in separate work trees this time.


📺 Source: Nate Herk | AI Automation · Published September 23, 2026
🏷️ Format: Comparison

1 Item

Channels