← Back to search
Kimi K3 VS Fable 5 VS GPT 5.6: Who Wins?
AI News Today | Julian Goldie Podcast · 2026-07-17 · 15 min
Show full episode description
Kimi K3 VS Fable 5 VS GPT 5.6: Who Wins?
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Which model —
Kimi K3,
Claude Fable 5, or GPT 5.6 — actually builds the best output when tested head-to-head, and how practitioners combine them.
Benefits
- Side-by-side game builds reveal real coding/design differences
- Kimi K3 wins on detail, ambience, and front-end polish
- Multi-model harnesses (Fable planning, Sol executing) beat single models
- Expensive models are affordable for simple tasks per-token
- Long tasks can now be trusted to run unattended overnight
Use cases
- Tested Kimi K3 across 50 different tasks within 24 hours of launch
- Built Skyrim-style open-world, Dragon Realm, and racing games with all three models
- tmux harness: Sol runs a goal while Opus/Fable review the last 200 lines every 45 minutes to steer
- Doom rendered in ~2,000 lines of SQL, essentially one-shot by Sol in about an hour
- Luna used for massively parallel simple tasks across all files
KPIs / results
- 50 tasks tested in under 24 hours
- Doom-in-SQL: ~2,000 lines, one shot, ~1 hour
- GPT 5.6 controls backwards on Skyrim test; Kimi K3 and Fable 5 best outputs
Tools / build
- Skyrim-style open-world game builds
- Dragon Realm game benchmark
- tmux dual-session Sol + Fable review harness
- Doom clone written in SQL
- intern.md two-file agent steering setup
Kimi K3 VS Claude Fable 5 VS GPT 5.6 Wins Today, we're going to be putting them side by side of the test now. So Kimi K3 is super impressive and I've been blown away by what it can create so far. This is actually kind of like a Skyrim open world style game that we created and it's absolutely massive. Look at the size of this, like how big the full open world game goes. Then we have Fable 5, which is honestly my favorite before Kimi K3 came out. But let's see what happens after. And then we have GPT 5.6 Soul as well. So we're comparing all three side by side today to see which one creates the best outpost. Let's kick it off with this Skyrim style game. So if we have a look here, for example, this is really, really nice. Feels smooth. I love like the light in the sky, the details in the sky. I like how open wide it looks as well. It's got kind of that like magical feel to it as well. Let's have a look. If we go inside the village over here, you can see it comes up with like the screens as well of where we are and that sort of thing. And we've got like all these little houses. It just looks really cool. Looks super nice. I mean, that has done a great job. If we have a look at Skyrim from Fable 5, I would say the ambience of this just feels a bit nicer. Also, you can see more of the first character when you're going through. It is a little bit buggy in parts, but I would say that Kimi K3 and Fable 5 have done a great job there. When I actually look at the outputs from GPT 5.6 Soul, it's not bad, but the buttons are totally backwards. So if I try to go forwards, it's actually the button for going backwards and you can't really go side by side. So I would say that if I look at all three of those, Kimi K3 and Fable 5 by far did the best outputs. Now, let's have a look at a similar sort of game called Dragon Realm that we test with all of our agents. By the way, if you're wondering, okay, how relentlessly did you test this model? So I mean, it dropped less than 24 hours ago and we've actually tested it out across 50 different tasks so far. So we have tested it relentlessly. It does create some awesome stuff and I'll come on to it in a second. So let's have a look at this. This is the output from Kimi K3, which looks absolutely awesome, doesn't it? Like the snow falling, the graphics, look at the sky. It's wild. Moving around is pretty easy, easy to control. And we've got this dragon over here, which I've never seen any of my AIs create that style of dragon on a normal test like this. So that's the first time I've seen an AI model complete the test like that, which is pretty impressive in itself. Let's have a look at Fable 5. Now, this does feel a little bit more basic than Kimi K3 in terms of the details and nothing else. Genuinely, I would say K3 did a better job than Fable 5 here. And then if we have a look, for example, we've got GPT 5.6. GPT 5.6, I would say probably did slightly better than Fable 5, especially with adding the enemies and that sort of thing. So I'm going to go with K3 for winning on detail and ambience. GPT 5.6 did an awesome job in just making the gameplay a bit more interesting. And then Fable 5, I would say we pretty much struggled on that test, to be honest. Now we have a racing game, and we can just compare these side by side here. So if we open this all up, we've got this one from Kimi K3. Looks awesome, feels nice. I just love the colors and the vibe and the gameplay is smooth. I'm noticing that on every single game. It's so hard to describe, but it feels smooth. It feels nice to use. And I think that would be amazing for creating websites as well. Now, if we have a look at Fable 5, this has not done a bad job, but it just feels a bit more basic. It's a bit more like a block sort of thing. So the graphics over here are nicer. The details over here are nicer from K3. And then if we have a look, we've got GPT 5.6, which I would say is done better than Fable 5. It's not as interesting to play, but the graphics themselves look nicer than Kimi K3. But again, there's nothing to dodge. The gameplay is not really there. It's not really that exciting and it slows down a little bit as well. So I'm going to go with K3 winning on that one too. So let's have a look. We've got the next one, which is Neon City. We've got a lot of neon tests coming up. So we have a look at this. I mean, this is insane. This is K3 over here. Looks awesome. Feels fun to drive. Pretty intense, but we love it. Love the detail of all the city blocks and everything else, the colors, everything else. If we have a look at Fable 5, it just feels more basic. It almost feels like a generation behind on that one. And then if you look at GPT 5.6, it's okay. It's better than Fable 5, but it's not quite as good as K3. So K3 is really, really powerful at this point. And it can create some amazing stuff. Now, if we look at the crit game, this kind of feels a bit more like a maze. There's some nice lighting. I like the graphics. I like the details. It feels nice to use it. If we have a look at the version from Fable 5, this is a bit more linear, but it's a bit more interesting in terms of gameplay. So when we're doing the tests here, you can see it's also quite buggy. Breaking a lot. Let's have a look at the next one. This is similar sort of vibe, but from GPT 5.6, which I would say again is beaten Fable 5 there. So I think in terms of graphics, K3, in terms of gameplay, GPT 5.6. That's quite interesting. You know, the AIs are good at different things, aren't they? One can be great at graphics. One can be great at gameplay. You take your pick and decide which one you prefer. I think on this one, K3 actually failed, so we might have to regenerate that later. You can see this just does some work. If we have a look at GPT 5.6 and Fable 5 over here, very basic output from Fable 5, but it's playable. It can do the job. The one from GPT 5.6 looks 10 times better than anything else here. Like, look how cool this is. The graphics, the colors, the gameplay, everything is much better. Now, this is a black hole simulation. I would say definitely K3 has created the most visual interesting one here. These two just feel a little bit not quite right, like something that's broken inside there or it doesn't understand the problem properly or something like that. Whereas if you look at K3, like super visual, nice. We can move it around. We can preview it side by side. Looks way nicer than the black hole simulation from GPT 5.6 and Fable 5. So Kimmy K3, you know, I think we're really at the point now where open source models are right at the frontier, you know, and they're only going to improve faster. I would expect to see a new model from Kimmy and GLM every single month for this rate. The rate that they've improved, the progress over the last few months, they're evolving way faster. Whereas you look at something like K3, you look at GPT 5.6 and there's a lot of back and forth, number one, in terms of getting the models out there. Sometimes they get taken down, then they come back out. And also bear in mind that GPT 5.6 and Claude Fable 5, you have to be very careful with tokens. So for example, when I was running the Goldie bench experiments, we had to use the API on a lot of the tests because it just ran out of tokens on the subscription with chat GPT and Claude. That has not happened with Kimmy K3. And we're just on like the $39 per month plan. We've had to regenerate some of the tests and it still was just steaming a lot. So the thing is not just that the quality is right up there with both of these models, but also the Kimmy K3 is open source, it's from China, and additionally is way, way cheaper on the token plan. So if you get the coding plan, you can actually plug it into your Hermes agent, which you can do with chat GPT, but you can't do with Claude Fable 5. Unless you use the API, which gets expensive. Now we have this fluid in a box, test looks pretty crazy from Kimmy K3. Fable 5 just created something really basic there. GPT 5.6 did something interesting. But for sure, like Kimmy K3 is just crushing on these tests. Now let's talk about benchmarks. By the way, for me personally, I don't really pay attention to benchmarks so much. I like to test this stuff out myself, particularly when it's not a company that I'm massively familiar with. So I like to test out myself. And you've seen on the benchmarks, actually, in reality, K3 is outperforming a lot of these different models. But if we have a look over here, so if we look at this, we've got, for example, Kimmy K3, and that's being outperformed by Fable 5 and GPT 5.6, so on Deep SWE. And again, like I just pay attention to my own benchmarks. That's why we created Goldie Bench, because I just don't think that these benchmarks are realistic. Sometimes they're too optimistic. Sometimes they're not optimistic at all. Either way, test yourself out for yourself. Don't even listen to me. You know, just make up your own mind. If we have a look over here as well, terminal bench 2.1. So we've got Kimmy K3 versus GPT 5.6. So Fable 5 is way down here at 84.6. Kimmy K3 scores 88.3. GPT 5.6 SOLD scores 88.8. These are according to the benchmarks that have been released by Leo over here, thanks to Leo. I've not seen the official benchmarks. Let's see if we can find them. So they've put the API documentation here, as you can see, but there's not much information on the benchmarks. I can't find them on the website, either. However, we can have a look and compare them on Router. So let's compare Fable 5 and also GPT 5.6 SOLD. By the way, SOLD is a model that you want to compare against all of these. So if we have a look here, they've all got the same context window. Moonshot AI is a Chinese lab, which created Kimmy. Anthropic is the founder of Claude. OpenAI is the founder of GPT 5.6. And they've all got the same context length. They're all reasoning models. They all have very similar input and output modalities as well. The price is a big, big difference here. Right now, actually what's surprising here, I guess it's not that surprising because models have just improved so much. But if you look at the comparisons, Kimmy K2.7 was actually a lot cheaper. So let's add that to the list as well. So yeah, I mean, look at the difference in price here, like Kimmy K2.7, which is the previous generation of Kimmy K3 was way, way cheaper. For example, input tokens for Kimmy K3 is $3 per million tokens. Whereas for example, Kimmy K2.7 is 0.72 cents. However, the step up in quality is probably about two to three times better. Like it's way more powerful. It would also be very interesting to use Kimmy K3 with GPT 5.6 as a mixture of experts models. So for example, we've got a mixture of agents with Fable 5 and you can compare them side by side, but I would say that would be an amazing way to get better outputs, like fusing the models together and just seeing how the agents perform when they work together as a team. In terms of latency, Kimmy K2.7 code is a lot faster than K3. And also bear in mind, Table 5, GPT 5.6, not open source. Kimmy K3 is open source as well. So let's have a look at the benchmarks now. So Kimmy K2.7 is a huge gap in terms of the improvement here. So you've got Kimmy K3 holding its own with all the Frontier models. Kimmy K2.7 was nowhere near. And what actually surprised me here is like how quickly Kimmy K2.7 improved. It was only like last month that it came out. Now, it would also be interesting to see what Kimmy K3 is being used inside. So for example, where are people using it the most? I can imagine it's going to be Hermes is number one. Ah, Claude Code. So if we have a look at this, Hermes Agent is actually third on the list. Claude Code is the number one place and the number one app that people are using Kimmy K3 in. So that will be an interesting experiment as well. It's like plugging Kimmy K3 into Claude Code and just seeing how it performs inside Claude's agent harness, which you can easily do. And oh my pie. This is not an app that I've tested out much, but I'd be excited to test out with Kimmy K3 as well. Might be something in the future that we do. And I think one of the best ways to get the most out of all of this stuff, so we've got Claude Codex and Kimmy Code all plugged into our agent operating system. And that means we can combine them both. We can have a group chat where all of our agents, including Kimmy K3, speak over here. We have, for example, Paperclip where we could orchestrate all three agents together as a team inside a company. And then for example, we have the memory system. So as soon as Kimmy K3 came out, we were not starting from scratch because it has all of that memory and all these different pieces of context about me inside a beautiful memory galaxy using Obsidian. So I think that's one of the best ways to use it. We also have a way of basically switching between GPT 5.6 Soul and GLM 5.2 inside Claude Code. And I'm probably going to add Kimmy K3 inside there as well. So those are some interesting ways you can use this. As soon as the rest of the results from Kimmy K3 come out on Goldie Bench as well, we'll show you those. They're just being built in the background. That's why they're not done. Bear in mind, this has only just dropped 24 hours ago. Something else that you can do with Kimmy K3 is you can build out these beautiful videos. So this is a video that we actually created using Kimmy K3 with ReMotion as a skill. And it created like a really nice video, nice camera angles, nice animations, etc. Fully automated in the background. It actually moves as a background when we move our mouse, which I've never seen with a video before. So it's kind of like overlaid with a website and a video at the same time. So which one would I pick overall? I would say for 3D games, obviously like K3 outperformed all of those different games. I also think that with the coding plan, it's a lot cheaper to use K3. For me personally, if I'm building out something big like the agent OS where I can't make any mistakes, really, then I'm going to stick with Claude because we have our systems built into that. And also it's fully trained on all the skills that we need. But I think K3 is a fantastic, cheaper alternative to Fable 5 and GPT 5.6. So if you're wondering about the best way to get the most out of all of them, I would make sure they have an agent operating system with all of these agents side by side. So if you want to get our agent operating system, you can get that inside our AR Profit 1 community. This is a place where you can learn, grow and scale with AR automation. We have all of our best trainings inside here. You get our video tutorial for the agent OS system. You can see that it was just updated today with K3. And then you can also get the zip file to install it. And we had new daily tutorials based on what's just dropped. Additionally, inside the community, you can ask questions, get help and support in real time. And I personally answer all the questions inside the community with a video tutorial. And then inside the calendar, you can jump up with coach goals, get help and support in real time, share your screen, meet other members. Inside the map, you can actually meet people in your local area who are building out with AR agents just like you. And that's all available inside the AR Profit Boardroom. Link in the comments description or just go to theairprofitable.com. Thanks for watching.