Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
| Source: MarkTechPost
Tags: Sakana AI, Fugu, multi-agent, orchestration, NVIDIA Nemotron, agentic AI, LLM routing
Sakana AI released Fugu Max ($2/$6 per 1M input/output tokens) and Fugu Ultra v2 — both orchestrators that route tasks across a pool of AI models rather than running a single foundation model. Fugu Max claims top scores on 6 benchmarks at 40-60% lower cost than Sonnet 5 and GPT 5.6 Terra, per Sakana's own tests.
Details
Sakana AI expanded its Fugu family with two releases targeting opposite ends of the cost-performance curve. Fugu is not a traditional foundation model: it reads a query, builds an agentic scaffold on the fly, and routes subtasks to the cheapest capable model in its pool. The new Fugu Max widens that pool to include NVIDIA's Nemotron family, through a Sakana-NVIDIA collaboration. Fugu Max is priced at $2 per 1M input and $6 per 1M output tokens — 40-60% cheaper than Sonnet 5, GPT 5.6 Terra, and Kimi K3 by Sakana's own measure. Sakana reports it leads on 6 benchmarks: Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench, and SWEFish. SWEFish is a Sakana-internal coding benchmark; treat that result with extra caution. Fugu Ultra v2 targets the quality ceiling on complex multi-step reasoning tasks; pricing was not disclosed. The orchestration architecture is grounded in two ICLR 2026 papers: TRINITY uses a lightweight evolved coordinator that assigns Thinker, Worker, or Verifier roles per turn; Conductor uses RL to learn natural-language coordination strategies. Training combines fine-tuning, evolutionary algorithms, and RL. For teams building agentic pipelines, Fugu Max is worth evaluating against dedicated frontier models on cost-sensitive workloads. Key constraints: no open weights, no EU/EEA availability, and all benchmark results are vendor-reported.