← Back to search
Sakana: BETTER Than Fable 5?
AI News Today | Julian Goldie Podcast · 2026-06-25 · 14 min
Show full episode description
Sakana Fugu Ultra vs Fusion vs Claude Opus 4.8 (42 Builds + Goldie Bench): Is It Fable 5-Level?The video reviews Sakana Fugu Ultra after several days of release, testing whether it matches “Fable 5-level intelligence” claims by running 42 builds using the same prompts across Fugu Ultra, Fusion (OpenRouter), and Claude Opus 4.8, plus Goldie Bench leaderboard comparisons. The script explains Fugu as a multi-agent orchestration system that fuses outputs from closed and open models, contrasting it with Fusion’s pay-per-token API pricing versus Fugu’s flat subscription with token limits and frequent lockouts, slow responses, and one-shot workflows. Side-by-side build demos (solar system, runner, fireworks, dungeon crawler, open world) show mixed results but generally stronger outputs from Fugu Ultra, with Fusion often close or better on some examples and Opus 4.8 frequently failing. Benchmarks cited include Fugu leading LiveCode Bench but trailing Fable 5 on SWE Bench Pro and Terminal Bench, and a brief comparison of Fugu Ultra vs cheaper Fugu Mini. The episode ends by promoting the AI Profit Boardroom agent operating system integrating Sakana, Fusion, Claude, and other tools.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Tests whether Sakana's Fugu Ultra multi-agent model truly matches
Fable 5 level intelligence on real builds.
Benefits
- Multi-agent LLM pool fuses closed and open models
- Flat subscription pricing instead of per-token billing
- Generates expansive, near-endless game worlds
- Outperforms Opus 4.8 on coding builds
- Fugu Mini offers cheaper, faster alternative
Use cases
- Built 42 different things with Fugu Ultra using same 42 prompts across three models
- Built solar system, voxel tempo runner, fireworks display, Nordic dungeon crawler, flight simulator
- Head-to-head Fugu Ultra vs Fugu Mini vs Fusion vs Opus 4.8 comparisons in agent OS
KPIs / results
- LiveCode Bench 93.2 vs Fable 5's 89.8
- GPQA Diamond 95.5
- Fusion top of Goldie Bench leaderboard
- Lockouts up to 5 hours, 15-20 min per response
Tools / build
- AI Profit Boardroom agent operating system
- Goldie Bench leaderboard
- Sakana Fugu and Fusion integration
- Hermes Oracle and lead generation tool
📑 Chapters — tap a time to jump there
00:00
Fugu Ultra Overview
- Fugu Ultra claims Fable 5 intelligence; 42 builds tested
00:45
Testing Setup and Limits
- Token limits caused slow responses and lockouts
01:26
Models and Method
- Compared Fugu Ultra, Fusion, Opus 4.8 on Goldie Bench
02:02
What Is Fugu
- Fugu orchestrates multi-agents, fusing closed and open models
03:00
Leaderboard and Pricing
- Fusion tops leaderboard; Fugu subscription vs Fusion per-token
04:12
Build Demos Part 1
- Solar system and voxel runner builds; Fugu leads
06:21
Build Demos Part 2
- Fireworks and Nordic dungeon crawler; huge endless worlds
08:33
Open World Comparison
- Open world game; Fugu beats Fusion and Opus 4.8
09:23
Benchmark Breakdown
- Fugu wins LiveCode Bench; trails on SWBench Pro
10:16
Fugu Mini vs Ultra
- Fugu Mini cheaper, faster but weaker than Ultra
12:03
Final Verdict
- Verdict: Fugu Ultra a real step up from Opus 4.8
12:13
Boardroom and Outro
- Full setup available in AI Profit Boardroom with Hermes
So Sakana Fugu Ultra has been out for a few days and this is, according to benchmarks, able to achieve Fable 5 level intelligence. Now does it actually match that standard? How does it perform? Where's it up to? Is it better than Fusion as well? I'm going to cover all of that in this video today. I'm going to show you what we've built with it. We've actually built out 42 different things with Fugu Ultra. So you get a really comprehensive idea of how powerful it is, whether it's worth using, the best ways to use it and we've tested it on Goldie Bench as well. I'll show you the benchmarks and the leaderboards in a second in terms of how it performs versus everything else. Also how it performs side by side and what we've built with it so far. So we have tested it absolutely relentlessly using our system here. So we have Sakana Fugu plugged into our agent operating system like you can see. And then what we can do with this is we can build cool new stuff with it and then test it out side by side. The only problem was that when it first got released, which was just a couple of days ago, you were kind of limited by the number of tokens you could use. And so the problem with that is like, number one, it was super slow. And number two, you couldn't really test it as much as you can without getting like locked out for five hours. So we've built 42 different things with it today. I'm going to just guide you through what it is, how to use it, and also how it performs versus Fable 5, Fusion, and Opus 4.8. So let's get straight into this. We ran 42 builds to find out. And we used the same 42 prompts with three models. Fugu Ultra, which is from Sakana.ai. We've got Fusion, which is from OpenRouter. Both of these on benchmarks are claiming to achieve Fable 5 level intelligence. And then we have Opus 4.8, which is the only or sorry, the most powerful model you can get from Claude right now. And we'll compare these side by side with our Goldie benchmarks. And we've tested it, you know, with real builds that you'll see in a second right here. By the way, if you're wondering, okay, what is Sakana Fugu? How does it work? So it's basically a way of orchestrating multi-agents together, right? So you have Sakana Fugu, you have closed and open models, and that's an LLM pool that basically fuses an answer together. So the idea here is like, if you have multiple agents working together, multiple models, closed and open together, then you tend to get better results than just one single output from one single model. And so on their benchmarks, they claim to achieve better level intelligence than Fable 5. So you can see the benchmarks right here. And this is compared against Mythos preview, for example. And you can see Fable 5 here too on the benchmarks. And on a lot of the benchmarks from this Japanese AI Fugu, which is a tiny little AI lab, they are outperforming Fable 5, right? Now let's see how it performs in reality on the actual builds that we've created right here. So on the build, I'm going to be 100% honest with you, on our leaderboard with Goldie Bench, you can see here that Fusion is actually top right now. It's actually crushing everyone else. Fugu Ultra is not doing too badly, but just have a look at the builds yourself today and see what you think, and see what you think about in terms of what you can get out of it. Now one thing to note as well is like, if you're using Fusion, which is from Open Router, follows the same idea, use multiple AIs together. With Fusion, the difference is that you pay per API. With Fugu, you use a subscription. And so with the subscription, it's a flat plan. With Fusion, you pay per token. So that's one of the biggest differences. Also, Fugu Ultra is from Japan, whereas Fusion is from Open Router, which I think it's the US. And also with Fugu, if you're on a flat plan, you can obviously run out of tokens, just like you would with a Claude subscription. Whereas the difference with Fusion is that with Fusion, you don't run out of tokens because you're paying per API, if that makes sense. So let's test them out and see what we've got here in terms of builds and how they perform. So first of all, we've got the solar system test right here. And you can see that we have Fugu Ultra, Fusion, and Opus 4.8. So we tested all three together. So this is the version from Fugu. Fugu, this is Fusion, and this is Opus 4.8. Now, if I had to pick one, I'd probably go with this one. I think it's the most sophisticated and elegant version. I do quite like the output from Opus 4.8. I would say that's coming in second. And then Fusion comes in last on that particular example. Let's have a look at the next one. So this is a voxel tempo run, like a runner game, as you can see. Pretty nice outputs from Fugu again. And it can build some super nice stuff, as you can see right here. If we compare that versus Fusion, I would say probably Fusion's created a better example. And then if we compare this versus Opus 4.8, wow, that is bad. It's just so basic, isn't it? It's just not very useful at all. Now, I will say the most frustrating thing about testing Fugu was that, you know, we just kept getting locked out for five hours. So bear that in mind. It's like, it's pretty cool. It's created some awesome stuff. However, not that great if you want to test it a lot. If you want to use it a lot, yeah, it just, you know, it would take sometimes 20 minutes as well for a response. And then sometimes it would fail as well. So that's something to bear in mind. Also, one thing to bear in mind with like both these setups, for example, Fusion and Fugu Ultra, is at the one shot because you wait 15 to 20 minutes for a response. And that's it, you know, you get multiple answers, but you're not going to go back and forth with it like chat chibit or something like that. So this is the output from Fugu Ultra. This is Fusion, which is a bit laggy, but not bad. I would say Fugu's game was better. And this is Opus 4.8, which kind of just feels a little bit retro for me. It's like not even full screen. It's probably, it's slightly smoother than Fusion's output. But I would say again, Fugu Ultra is winning. So, you know, on all the tests so far, pretty much Fugu is winning on all of these. Let's have a look at the next one. So this is a fireworks display. Basically, and what you can do is like, you can click and to generate the fireworks, and then it'll follow you around based on where you put your mouse, right, which is pretty nice. So it looks really cool. Let's compare that versus Fusion. Fusion just literally didn't work. And then we have Opus 4.8. So Opus 4.8 is outperforming Fusion, but it's nowhere near the same level as Fugu. Next one, this is a Nordic dungeon crawler. So let's have a look here. Wow, this is pretty nice. This is from Fusion again. I mean, it looks super nice, right? Nice 3D sort of setup. Walking around this Norwegian dungeon. Super nice, like 3D style. I like that. I like that. It feels huge as well. I don't know where this world ends. That is one of the things I found with like, when we're creating games with, for example, like Fugu, the games feel like endless, if that makes sense. They feel absolutely huge. So let's just go over here and see if we can get those coins. See if actually, oh no. Oh no, Sunshine. Who's this guy? All right. So yeah, it's pretty cool. Now let's have a look at that one. So this is from Fusion. So I mean, I don't really know what's going on here, to be honest with you. Like it's very difficult to see anything or to understand what's going on, right? It's just like, it's so dark. I mean, it looks, it's kind of cool, but I don't know. I literally don't know what's going on there. And then we have Opus 4.8 where you can't actually get through the start screen, right? It totally failed. So you see the level difference here. You're seeing how different it is between each of these models. This is not like, for example, like a close race at all when it comes to being compared to Opus 4.8. It is clearly at a massive, massive level above Opus 4.8. And that's why we would say it is comparable to Fable 5. I wouldn't say it's as good and I wouldn't say you can go back and forth with it as much, but it can create like comparable stuff as you can see. Now this is an open world game from Fugu. As you can see, this one just feels a little bit more basic to me. Like I think you need better graphics. Again, the problem with this is like, you can't go back and forth with these games. So if you create something, that's it. Unless you get called Opus 4.8 to tweak the code base later, you're kind of stuck with the outputs. Whereas for example, if you create something like this with Opus 4.8, you can say, okay, change this, change that, right? It's not just one shot. Let's have a look at Fusion. Fusion's, sorry. Yeah. Fusion. This is even worse. I would say super basic. Feels quite slow when you're walking around. Graphics are not nice at all. And then we have Opus 4.8. And again, we can't get through the enter screen. Like it's literally just stuck there. Opus 4.8 seems to do that a lot. Let's see as well how it performs on the benchmarks versus Fable 5. So LiveCode Bench outperforms Fable 5. SWBench Pro. Fable 5 is outperforming Ciccana. And then on Terminal Bench, Fable 5 is outperforming Ciccana as well. So just to be clear here, like Fugil Ultra wins on LiveCode Bench 93.2 versus Fable 5's 89.8. It tops GPQA Diamond at 95.5 above Ciccana's own Mythos preview number right at the saturated frontier. And then it trails on SW approach. So it is pretty close. And again, it's a, I mean, Ciccana is just crushing, crushing Opus 4.8. Even Fusion is as well. But from what I've seen, Fusion is not as good as Fugil Ultra when I test it. Also something to bear in mind here as well is that there's Fugil Ultra and there's Fugil Mini, right? So if you want to see some stuff that we've created with Fugil Mini, let me pull this up and we'll have a look at the models here. So this is Fugil Mini. So it created some interesting stuff. Now, if you want to see versus other models, we can go to the compare section here and we've got head to head comparisons of all these models. So if you want to see Fugil Ultra versus Fugil Mini, we've got that right here and we can see how it performs. So, you know, Fugil Mini is a cheaper, faster version of Fugil Ultra. So this is kind of like a flight simulator game, as you can see from Fugil Ultra, which is quite hard to control, but pretty fun. And then we have Fugil Mini here, which actually feels a bit smoother, to be honest with you. It's not that, not, I mean, the graphics are not as nice, but it does look like more pleasant on the UI. Got a couple of examples here. So we recreated the same game with both. Now it says you can use WASD to move, but you actually can't on that. And then Fugil Mini literally can't get past the start screen here. That was the other problem as well we found was like sometimes the tokens run out in the response. This is Fugil Ultra, which is pretty insane. Look at that. Wow. That's pretty nice. If you look at Fugil Mini, not quite as good. Look at that snow. That snow is nowhere near the same standard as Fugil Ultra. So there's a huge difference even between those two different models. I think Fugil Mini is kind of like Opus 4.8 for Sakana basically at this point. So yeah, there's a big difference between them as well. And overall, I would say, you know, something that's just a better step up from Opus 4.8. Yeah, for sure. I would go with Fugil Ultra. Like it performs really good. It's created some awesome stuff. What we've actually done is if you want the full setup, you can get that inside the AI Profit Boardroom. So this full agent operating system that you see here with Sakana and Fusion built in, so you can chat to either of them. And you can see what we've built inside the workspace here. So everything you create gets saved inside the workspace too. This is all inside the AI Profit Boardroom. And this is an agent operating system where, for example, you've got a memory plugged into every agent you use. You can orchestrate all your agents with paperclip and a group chat. You've also got a pipeline where you can go from idea to implementation really quickly. And we have a Kanban board for orchestrating new agents too. We've even, for example, built out some awesome custom stuff with Hermes Agent. Like you see, include the Hermes Oracle and all sorts of cool stuff for outreach and lead generation tool as well for this sort of thing. So if you want to get that, that's all inside the AI Profit Boardroom. Link in the comments description or just go to the AI Profit Boardroom.com. You can get the agent operating system with Fusion, Sakana, and Claude all built in along with Hermes right here. We update it daily and you can see the last update date. I have a video tutorial and a guide on how to install it and with the resources over here as well. And then we drop new daily tutorials based on what's just been released. Inside the community, you can ask questions and I create video tutorials helping you every day on everything that you ask. And then inside the classroom, you can go from beginner to expert in just six weeks with our course right here with AI. And inside the calendar, you can jump a weekly AI automation coaching course, ask questions, meet the community, share your screen, ask questions about your setup as well. Inside the map, you can meet people in your local area who are building with AI agents. And that's all inside the AI Profit Boardroom. Link in the comments description or just go to the AI Profit Boardroom.com to get access. Thanks for watching.