Dashboard

Sakana Fugu Max and Ultra v2: Orchestration as a Model

Sakana AI launched Fugu Max and Fugu Ultra v2, orchestration engines sold as a single model. What the no-frontier-models claim actually proves.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
12 September 20261 min read

Sakana Fugu Max and Ultra v2: Orchestration as a Model

Sakana AI announced Fugu Max and Fugu Ultra v2 on 11 September 2026. Both are orchestration engines that present as a single model behind one API endpoint: you send a request, and a coordinator picks and combines models from a pool to answer it. Fugu Max is priced at $2 per million input tokens and $6 per million output tokens.

The claim worth reading twice is this one, from Sakana's own announcement: Fugu Ultra v2 reaches its scores without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.

What that claim does and does not prove

It does show that a well-routed pool of smaller and open-weight models can land on benchmarks near where single frontier models land. Sakana reports Fugu Ultra v2 at 48.3 on Chartography and 74.3 on DeepSWE, with best or joint-best results on GDP.pdf, Chartography, SWEFish, DeepSWE and Toolathon. Fugu Max takes the best overall scores on Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish.

It does not show that orchestration beats a frontier model on your work. Benchmarks are a fixed set of tasks with known shapes, which is exactly the situation a router is best at: there is a right specialist for each question, and the coordinator has had every opportunity to learn which. Your production traffic is messier and less evenly distributed.

It also does not tell you the variance. A single model gives you one distribution of answers. A pool gives you a distribution over routing decisions on top of a distribution over models, which tends to mean a better average and a longer tail of surprising results. Benchmarks report the average.

Where this is genuinely useful

  • Mixed workloads. If your app does document extraction, some code, and some chart reading, a router covers all three without you maintaining three integrations and three prompt styles.

  • Cost pressure at the top end. Fugu Max at $2 and $6 per million tokens sits well below frontier pricing while claiming frontier-adjacent results on several benchmarks.

  • Avoiding single-vendor exposure. A pool that explicitly excludes the big three proprietary models is a different risk profile from a direct dependency on one lab.

The question to ask before you route production traffic through it

An orchestration layer is a subcontracting arrangement. Your prompt reaches a coordinator, and the coordinator forwards some or all of it to a model run by somebody else. That is fine, and it is also a data question you should get answered in writing rather than inferred from a pricing page. We go through the specifics in who sees your data when an AI router picks the model.

The second question is stability. If the pool changes, your outputs change, and pool composition is the vendor's lever to pull for cost. Ask what notice you get.

How to test it honestly in an afternoon

  1. Take 20 real requests out of your logs, weighted the way your traffic actually is, not the way you wish it were.

  2. Run them against your current model and against Fugu, same prompts, no tuning for either.

  3. Grade blind. Shuffle the outputs, strip the labels, score against a rubric you wrote before you looked.

  4. Count the tail, not just the mean. How many answers were unusable? That number decides whether a router is workable for you, and it is the number benchmarks hide.

  5. Price the real mix. Input-heavy and output-heavy workloads land very differently on a $2 and $6 split.

If that sounds like work, it is roughly four hours, and it is the difference between a decision and a vibe. How to test a new AI model before switching has the longer version of the method.

FAQ

Is Fugu a model or a product?

Functionally it is a system presented as a model. Sakana describes Fugu as a multi-agent system delivered as one model, reached through a single OpenAI-compatible endpoint. For your integration it behaves like a model. For your risk register it behaves like a vendor with subcontractors.

What is the difference between Fugu Max and Fugu Ultra v2?

Same orchestration architecture, tuned for different priorities. Max optimises for cost and integrates a large number of open-weights and specialist models, including the NVIDIA Nemotron family. Ultra v2 optimises for capability on harder reasoning work.

Does excluding frontier models make it safer?

It makes it different, not safer. You trade dependence on one large lab for dependence on a coordinator plus a shifting pool of providers, each with its own data handling. Fewer eggs in one basket, more baskets to check.

Is this the same thing as model routing?

It is model routing with the routing hidden behind a single endpoint and priced as one product. What is model routing in AI covers the underlying idea, and what is multi-agent orchestration covers the coordination layer on top of it.

Independent coverage of the launch: MarkTechPost. If you are trying to work out which releases deserve your attention at all, how to tell if an AI product launch actually matters is the filter we use, and keeping up with AI releases without reading everything is the wider routine.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.