Sakana AI’s New “Fugu Ultra” Beats Claude Fable 5 (Sakana Fugu)

Sakana AI’s New “Fugu Ultra” Beats Claude Fable 5 (Sakana Fugu)

More

Summary

Sakana AI has launched Fugu, a multi-agent orchestration system that presents itself as a single API endpoint while internally routing tasks across a dynamic pool of specialized models — including recursive instances of itself. TheAIGRID walks through how Fugu works: a core LLM trained to delegate subtasks, verify outputs, and synthesize results, with Fugu Ultra tuned for maximum quality on hard multi-step problems by coordinating a deeper expert-agent pool.

Benchmark comparisons are the centerpiece of the video. Across LiveCodeBench (designed to prevent training contamination via release-date tagging), GPQA-Diamond, Chalk’s Civ scientific chart reasoning, and SWE Bench Pro (Scale AI’s long-horizon software engineering benchmark), Fugu Ultra consistently outperforms Claude Fable 5, Gemini 3.1 Pro, GPT 5.5, and Opus 4.8 — with the notable exception of SWE Bench Pro, where Fable 5’s design for long-running agentic tasks gives it an edge. The video positions this as evidence that performance gains are increasingly coming from orchestration architecture rather than raw base-model scaling.

Practical use cases shown include autonomous ML research — Fugu Ultra ran over 100 experiments across 14 hours on a single H100 GPU to iteratively improve a GPT training recipe — and a financial time-series prediction test where it grew a simulated $10,000 portfolio to $11,943 (20%) versus under 15% for competing models. Sakana also pitches Fugu’s ability to reroute around unavailable models as an AI sovereignty benefit in the context of ongoing export controls.


📺 Source: TheAIGRID · Published June 23, 2026
🏷️ Format: News Analysis

1 Item

Channels

1 Item

Companies