Hook

This beats Mythos and Fable 5. Fugu explained. This Japanese AI lab just released something that claims to beat Mythos and Fable 5. Our "Fugu Ultra" model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls. Try it: sakana.ai/fugu. And it's going absolutely viral with 24 million views. The name is Sakana AI. Building Frontier AI in Japan. And it's not just a random startup. It's a startup with the pedigree. The co-founder is one of the original authors of the Transformer paper that made that ChatGPT possible. Their bet is that the next big jump in AI is not gonna come from training one bigger model, but from learning how to combine different ones together. And they prove it by releasing something called Fugu. Fugu is like an orchestrator that combines ChatGPT, Claude and Gemini together to answer your question. And it scores higher than any of them alone in most benchmarks. If that's true, that's gonna be really bad for Anthropic and OpenAI because their entire pitch is that there's only gonna be one winner, one model that beats all the others. But the coolest part is that this is not a black box. In the past two months, Sakana has published two papers that explain how the orchestration system works. We can take a look inside. The first paper, it's called Trinity. It's a 0.6 billion parameters model. That is smaller than anything that runs on your phone. Trinity is basically a dispatcher. Its superpower is to decide which model is the best for a given task. It does this in what we call a hidden state. So it keeps costs and latency low. The second paper published just a few weeks later, it's called Conductor. It's a 7 billion parameters model that acts like a project manager. The conductor job is literally bossing around ChatGPT and Claude. It writes out an actual workflow and then assigns tasks to each model. Fugu is basically those two papers combined into one product. Regular Fugu is Trinity and Fugu Ultra is the conductor. I think it's really cool watching actual research getting productized. And the conductor job is literally bossing around ChatGPT and Claude. It writes out an actual workflow and then assigns tasks to each model. Fugu is basically those two papers combined into one product. Regular Fugu is Trinity and Fugu Ultra is the conductor. I think it's really cool watching actual research getting productized. Sakana Fugu surprisingly performed near GLM 5.2 but 17x more expensive. We gave the same prompt to 4 models: build a complete live Trader Desk with both frontend and backend components, real-time market data fetched from external APIs, tools, and a custom dark-theme UI. Outputs: Fugu Ultra - 22,225 t, $0.51. Opus 4.8 - 15,802 t, $0.31. GPT-5.5 - 11,474 t, $0.26. GLM 5.2 - 13,677 t, $0.03. Fugu created the most polished and feature-rich trading desk in the run. GLM 5.2 was very close behind, with a similarly complete multi-panel interface and live data, but at a much lower cost. Opus and GPT also performed well, delivering solid results with a better balance between quality and cost. An Ultra is slower and more expensive than just using Opus or ChatGPT. I love the direction. I'm not sure about the execution yet. If this is new to you, chances are it's new to most people. So share this video and hit follow for more well-researched AI content.
Their other posts in the index, biggest breakout first.