← Back to search
We Tested GPT 5.6 Sol Early
Nerd Snipe with Theo and Ben · 2026-07-09 · 71 min
Show full episode description
We've spent six figures in tokens testing OpenAI's 5.6 Sol model to see whether if its better than Fable and what OpenAI have done to make it even better than 5.5. Also, we breakdown why we both moved our agents to Linux boxes, how to actually burn $65k on a single loop, and the Codex vs Claude Code subagent gap that's now bigger than the model gap itself. Thanks to this episode's sponsors Clerk and General Translation: Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk General Translation: https://nerdsnipe.link/gt Sources Available on our Substack: https://nerdsnipe.substack.com/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Sharing early hands-on impressions of
OpenAI's GPT 5.6 Sol and how it compares to GPT 5.5 and
Anthropic's
Fable.
Benefits
- Reliable long-running tasks without stopping partway
- Better prompt interpretation and tool calls than 5.5
- Snappier, fades into background during daily use
- Stronger with sub agents on complex work
- Follows existing design styles better than 5.5
Use cases
- Ran giga runs letting the model work long periods on complex tasks
- Spent $131,700 in tokens across Theo's machines, plus Ben's $93,000
- Most content shown as 5.5 for weeks was actually 5.6 with hacks
- Julius crashed out on 5.5 icon-offset bug after losing 5.6 access
KPIs / results
- $131,700 tokens across Theo's machines
- $93,000 tokens on Ben's own machines
- 5.5 completed only ~1/5 of tasks 5.6 finished
Tools / build
- Sub-agent long-running workflows
- GPT 5.6 Sol early-access testing rig
Hello from the future. We're about to air an episode that we recorded a while ago because we were really excited about 5.6. We wanted to do a video about it, well, a podcast about it, with our organic thoughts before it had been released and talked about for weeks, as we often would do. There's a problem, though. They didn't release it. Yep. We've been sitting on this episode for a bit. I don't even want to know how long we've been sitting on it because I'm filming this today on a Saturday right after the release happened. Two days after we filmed that, and I'm a little upset. Yeah. And I hope you can understand. This is not OpenAI's fault. They've been awesome about all of this throughout. They were kind enough to give us early access, and we used the hell out of this model, as you're about to see. Yes. But we weren't allowed to post this until now, so I hope you forgive us for not necessarily having the most recent news in this, but think of it as a time capsule to a dimension that could have been if this model came out in a normal way, but also, more importantly, our real organic thoughts on what it's like to use this model as soon as you are able to use it. Yes. I'm really pumped to finally have this episode ready to go. Huge thanks to our sponsors for their patience as we waited for the official thumbs up to get this out without pissing off the U.S. government. Yep. Oh, and one other piece of context, the model we were testing and using is Sol. It's 5-6 Sol. Yep. So if you saw the three names, we did not get to touch Terra or Luna. Yep. We used the big boy, and it was a good time. Yeah. I miss it dearly. It was, oh, it's an incredible model. I thank God it is out. I think we can actually say this part now. Fuck it if we get in trouble for this, whatever. Yeah. I think this is the first ever model release where by being released, the number of people who have access is less than before it was released. Yeah. It sucks. There's a lot of people testing. Yeah, and we all got cut off at the same time. Yeah, we held a memorial for it when it went down. Oh, that was painful. Faye, sneak in the clip of Ben staring at the timer countdown as we were at a concert. It hurt me. It was physically painful to lose this thing. We are, if you remember this episode that you watched maybe last week or God, I hope not two weeks ago. What was happening here is we are in shock losing this model. It sucks. I cannot stress how bad going from 5.6 back to 5.5 is. All of the new workflows, all of the new things that we were doing and new ways of using these things. I got too many spoilers. They have a whole episode to watch, man. Yeah, they do. The short of it is, it's a great model. Thank God it's back. Enjoy. Interests are annoying. Yeah, they are. Yeah. Welcome to Nerd Snipe. I'm Theo. I'm Ben. And we're here to talk about a model that we've been using for far, far too long. For those who don't know, Ben and I are lucky enough to be early access testers on the latest models from a certain lab named OpenAI. And that lab just put out a new model that will probably be named 5.6. We're not actually sure at the time of recording what it will be called. I'm surprised if it isn't called that to be fair, though. So we're going with that. So today we're talking about what we believe to be 5.6 and the experience we've had with that for the far too much time we've been testing it. And terminal, you screwed up on the X so badly. The far too much money we have spent on it. For those who are not watching the podcast and are listening, you should consider watching it too. The visuals are pretty cool. But I have a device that I am holding that reflects really badly with the lighting that shows how much money I've spent in tokens the last few days. And we're at $131,700 of tokens across all of my machines, the vast majority of which have been with this new model, assuming the pricing is the same as it was before. Yeah, that's just Theo's. I have 93,000 of my own, so it's bad. We've been using this a lot. And since we suspect the model is coming out very soon, I was pressured into filming this at midnight when I was trying to relax after streaming. So this is going to be quite an episode. Apologies in advance for the chaos that is about to occur, as well as to our wonderful sponsors, Clerk and General Translation. All right, let's get into it. So where do you want to start with this thing? I think the thing that I instantly noticed about this, or I guess it wasn't instant, but the first thing that comes to mind. I think we should set the setting first. I think it makes it more interesting, like where OpenAI is coming from with this model drop rather than just like what it is and how it behaves. Give them a little bit of suspension, so you have to wait for the thing you're here for, which is how it works, as well as, most importantly, how it compares to a certain fabled model that we haven't ever seen before. No. A figment of our imagination at this point, a myth, if you will. This model comes out at a really strange time for OpenAI because probably from the start of the development of this model, it was known that there was a thing called Mythos at a certain Anthropic. That was, once again, a size increase for the model. And I don't think they would have had time or honestly reason to do a new pre-training. This is likely the first RL pass on the base that is 5.5. They have to compete with a model that is now much bigger and more expensive with a company that previously was very compute constrained and no longer is in a climate that is terrifying around models potentially being banned if they are too good. OpenAI is coming into a space very different from before where previously it's like they were slightly behind and it was just like specific RL differences and they would catch up and it was neck and neck. It does kind of feel like at least before 5.6 drops that Anthropic is or was ahead before Fable was taken. Yeah, it was a strong leapfrog. And that's the position they are in. So this is a model drop that matters a lot for OpenAI because I don't know how you feel, but I was finding myself using Fable a hell of a lot when it dropped. That was the first one where I was starting to really do the giga runs like we did with this model of letting it run for a very long period of time on big complicated tasks. I burned through an insane amount of usage for at least for the time an insane amount of usage was going through the full sub and it was an incredible model. It felt different and we actually I won't spoil how we feel about this model, but Fable was incredible and OpenAI had very big shoes to fill. You think they filled them? That's a really tough question to answer, especially so early. Yeah, the let's start by talking about how we've used it and what we've actually seen it do. We can get to the did it make the cut? Did it beat out Fable later on? Because I think both of us are going to have interesting answers to that part. We also have the interesting experience that neither of us have had before where this was the first time we got access to a new model and then lost it throughout testing. Historically, we would get access to a new model like this and then they would occasionally put out like different snapshots throughout the testing and then the final version that would go public. This had some interesting moments where they would temporarily take out the model we were testing. I think they were just changing out how they're doing provisioning. It's different from before. I don't think this is like something to read into. It's just a difference in how they ran the tests. And as a result, we had multiple times even tonight for a bit where we couldn't access the new model. We had to go back to five five. And that felt a lot worse than I would have expected. It was a very strange experience because when we first got access to it, it was like, oh, yeah, this is five five, but better. Just everything about it. It was a little snappier, had better judgment. It just did its job well, but it wasn't didn't feel that insane until we went back. And as soon as you go back, good Lord, five five feels bad, like really bad. Yeah. One of the first things I know that is probably a good place to start with behaviors is that it could be trusted more for long running things. And it didn't do the thing I always hate about five five where it stops after the first part of some work. That was the thing that made me not like five five initially and honestly just didn't like it for a while. Is that feeling that after you give it like the set of things that you do and it starts working, that it's going to stop in the next 10 minutes. Like, OK, I finished part one and half of the second part. Should I keep going? Yeah, this doesn't do that. And I noticed the regression in five five as soon as I had to swap back where work that five six was happily doing. I ran very similar things at five five and it got like a fifth of the way through and I got really annoyed. It cannot handle the more complex tasks over long periods of time, especially with sub agents. We'll get to sub agents a ton later. I think honestly, I had some things for the five five regressions, but Julius sent us some screenshots of when he was testing earlier and got forced back to five five. I think these are just beautiful. Like one of the ones he sent was, why are these icons offset further than the others? And M, no, use the fucking browser and check, man. This is not a 2PX issue. You blind fucking model. So we got Julius, one of the most chill guys I know, to crash out hard at five five after using the new model. Like it is so much worse. It just I had to take for a while that I felt like, wow, five five is so good. This crosses such a threshold. Like, how are you going to even feel the difference in future models if you're not running them for 10 hours at a time? It's just it's already so good. I was entirely wrong. Like entirely. You can absolutely feel the difference even on non big complex programming tasks, just like day to day using the computer. It interprets your prompts better. It does the tool calls better. It's just it fades away in a way that five five didn't. I noticed five five more than I noticed the new model, which is huge. Like that's what you want AI to do. If you notice it, it's bad. Mostly agree here. Yeah, the there's an interesting thing that happens to us as engineers and just like it's the nature of human brains. You get used to a thing and it's more normalized to you. Your expectations shift naturally as a result. And this model is better enough that I found myself asking it to go longer and do bigger things than before. And a big reveal. Most of the content I've done for the last two to three weeks that shows five five was be manually swapping to five five or doing hacks to make it look like five five. I actually feel a little bad for my loop video because that was all done with testing five six stuff. So my expectations for the model and like the amount I would reach, so to speak, like how far I would let it go and how much I decided to do meaningfully bumped up in my time using this model, especially with fable happening around the same time. My bar raised a ton and then I had to hold this higher bars. I went back to five five and it just felt like the model could even come close to where my expectations had been raised because my expectations aren't set by what the ultimate capabilities are of the thing. They're set by like what does it start to fail at? Like where does the failure rate go from zero percent to like ten percent? And I hover in that range for what I usually default to and the floor there was raised a ton, which made me do things that five five would struggle with a lot. Yeah, exactly. The way we use these things changed a ton and I think it was actually triggered by fable. All things considered like fable got me to start taking long running tasks and sub agents and all that stuff much more seriously and playing with them more. And as soon as it was taken away, I was sitting there like wanting it back. So I was like, all right, let's try and push this model in the same direction. And it turns out it does a phenomenal job with it. This model can run for very long periods of time in a way that no previous open AI model could. This thing is very different from five five, even if it won't maybe instantly feel quite that way. It's a huge step up. We didn't put front end in this list. Cool. I didn't even add it because I've written it off so much. We have notes for this on like what we want to talk about on the model. And I didn't even include front end on that list because it's still terrible. Like, is it a little bit better? It's meaningfully better, I would say. It follows instructions related to design better and put in a system that already exists with designs. It will follow the existing styles better. Yeah. And if you beat it with a hammer enough, you can force it in the right direction, which is similar to five five in that way. But if you ask it like, hey, here's this data, make a nice visual for it. You're not going to enjoy what you see. No, you won't. It still just does the LLME thing where it will use the generic overused patterns like the all caps text headings with the space in between them. We've seen a billion times. The little call outs, the cards everywhere, the same colors. It'll constantly vomit that stuff out. And it still has the problem of over vomiting onto pages where if you look at UI as it generates, there's just too much there. Too many call outs, too many sections, too many extra little buttons. It doesn't look or feel good at all. The status pill that will always be there no matter what you do. I'll throw this screenshot in for you phase. But the 5.6 is better. The corrected conclusion with the little thing is just so cringe. I hate it. I hate it so much. Yep. It's still an open AI model. It's going to be interesting to see the shitposts from the team about that in particular when it drops. But yeah, it's still not good enough at front end at all. The Codex app would look a lot nicer if it did. What it can do surprisingly well is 2D and 3D reasoning. Sadly, we did not get API access, so I couldn't run it on any of my wonderful benches for this, like, Skatebench or the Minecraft benchmark that I've been helping the team who created it a bit on. So I've only been able to test through things we can use Codex for, which includes, obviously, Codex and building code with that. Also, things like Hermes Agent, which we both had a lot of fun with, too, which we'll talk about in a bit. But the ability for it to do 2D and 3D spatial reasoning for, like, game development was actually quite impressive. 5.5 already surprised me here when I asked it to take my old fish slop game and rebuild it with a new engine as well as making a 3D clone of it. It still didn't really honor the spirit of what I was getting at for it until I, like, really hammered at it, like, no, I want it to be this style of game with this style of camera. And then it kind of got it. It could make 3D assets using, like, geometry in the, like, existing, like, 3.js world. But it wasn't very good assets. But what it did say it was able to do is use Blender. So I just had it rebuild these assets in Blender. And I have not actually had a chance to see how this looks yet. So we get to see my actual reaction together. And for those who are watching, instead of just listening, you'll get to see screenshots and or recording. You know what? I'll screen record this. Still can't believe it, like, used Blender properly. Let's see how it did. Meaningfully better assets, actually. Like, they almost look like fish. They're still really cursed. But those are fish. Yeah. And the fact that it, like, took this very vague, like, I didn't even tell it what the game was. I just gave it the old code base and said, remake this but 3D. And not only did it succeed, it made a thing that, like, plays fine and functions. It's kind of crazy that, like, we're at this point. Ben was making fun of me earlier because I was complaining a lot about the default hotkeys it shows. And Ben is like, how crazy is it? We're at the point now where, like, complaining about the taste of hotkeys the models pick instead of just, like, their inability to code. Like, it's insane that we're here. I haven't read a line of the code that this wrote. But I'm playing a real 3D game that did not exist 20 minutes ago that this model didn't just make, but it also created 3D models in Blender for, which is nuts. I'm going to ask it to refine them, though. How did it access Blender? Was it just using, is there an MCP or CLI? Ah, ah, okay. Damn, you can do Blender over CLI. Of course you can. You can do anything over CLI. I have, fair. Point being, this model understands how to work in a browser, work in an engine, work in these environments. Still sucks at iOS, though. I don't know if you played with that at all. I actually have extensively, and I disagree. When I got to try and make a Swift UI app, it, like, letterboxed it. I don't know what the hell it was doing, but it was so bad. I didn't do any Swift. I did everything in React Native and Expo, but it worked. And it was able to even work its way around all of the horrific nonsense that is my Apple developer account, because it's in a broken state where I, like, kind of have one, but also don't have one because of a company I used to work at. So it's, like, really broken on my computer. I've been trying to get this fixed with Apple support, and they just won't do anything. So I have no idea what to do. So I can't actually build things in a blessed way to any of my machines or devices. But it found a way. It found a way to use the Expo stuff and the Expo tools to sneak around it and get stuff onto my devices properly. Dev servers worked really well. The installs worked really well. Package management worked really well. When I told it to use the native glass components, it used them perfectly. The designs were actually remarkably better on mobile than they were on web, because since it had this little toolbox to play with of the native components, it couldn't over-engineer itself into horrific cards and buttons that, like, it usually does on web. I had a great experience with it, honestly. I did a lot of mobile. Interesting. I gave up pretty quick. I know Julius has made a lot of progress using it for mobile, but, like, initializing apps with it, I had a rough time, and it did not seem to know Swift well. Maybe it's better at React Native because it was okay with 5.5. So that's a meaningful improvement. It's substantially better. Like, I tried to do this stuff with 5.5. Meaningful difference. And also the, like, expo skills, the blessed ones, are quite nice. So you give it those. You give it the device and all the things. It does a phenomenal job. How'd do a Svelte, though? Uh, phenomenal. Like, quite good. It can just handle it out of the box. Like, at this point, like, I used to talk about the Svelte test or whatever with the models to see how well they can do it out of the box. They all do it great. Like, there's no issues. I did have a lot of issues, though, with the general quality of the code, especially after reading Fable code. One of the things I did after Fable was taken away from us is I sent off an X-High instance to analyze every single Fable session that I had on my machine to look for workflows that it did, code snippets at row, just generally analyze what it was actually doing. This was one of the first looping tasks that I did. We'll get into that in a second. And the things that it outputted at the end were some really good skills for Svelte, for effect, and also some, like, nice workflows for how to do PRs properly. Like, I have a built-in skill now that's, like, feature PR orchestrator or whatever that just tells the model to ask me for the thing I want to change, ask for the branch it wants to make the PR off of, and then it will go through, plan the change with a subagent, implement the change with a subagent, commit the changes, push them up, spin off two reviewer subagents, one of which is a Codex instance, one of which is a Claw-P instance, and then keep babysitting the code reviewers up in the cloud until they're done. So now that I have that workflow, I've fired that off dozens of times since setting it up, and it pulled that out of the Fable session history, and it also pulled higher quality code out of the Fable session history. So I was able to make a nice skill for how to properly use Effect v4 and Svelte with remote functions and all that stuff. And now whenever it goes and writes code in those languages, I'm calling them both languages, they're languages. It makes better changes, and the code looks a lot better. It's still not perfect, but it's a lot better. Yeah, I still ask it to go call Opus and get feedback on APIs. I really don't like the APIs and SDKs. It designed a lot of the interfacing for Lakebed. It designed and built most of Lakebed. I should be real about this, but I've had to go through and refine a lot of that both by hand and by Fable and Opus. I cannot wait till I can have Fable be called by this model to give feedback on things. I'm also very excited to have Fable as the top-level orchestrator, and anything that isn't like an SDK or API design, it can send off five, six, and save me a shitload of money and usage. I still wish that they had better discernment on this stuff. Like I watched one of the projects I did with this was like a Hermes box thing where it would build up a little VM that had Hermes, Codex, Cloud Code, Executor all packaged up together in a nice little instance that I could then have the Hermes agent run in, do its thing. And then at any time, just package that up into an encrypted zip file to stick wherever I want and run wherever I want. I have an internal Hermes agent for our company and team running on that right now, and it's phenomenal. But looking at the way it designed this, it's objectively good, but wildly overcomplicated. Like it loves, it still just loves overcomplicating things, adding tests for every single edge case, making the API design this bloated hellscape of 500 different commands just to do a very simple thing that really only needs about 10. It just does not have the discernment that Fable had. And I think that might just be a bigger model thing where it's still kind of just biases towards vomit out a lot of stuff on the screen. I did notice some like meaningful behavioral changes. Like I mentioned before, it's willingness to just keep going without stopping and asking for permission, but also just like what it would choose to work on and how many questions it asked. Like if you didn't tell it to ask questions, it wouldn't, it would just start building and assuming. And for the most part, it did that well, but like not always. And also it really, really, really likes tests. Yes. Oh my God. It loves tests. Like even if you don't tell it to add tests, it will add tests every time. Tell it to add tests. It will add even more. But if you don't tell it to add tests, it'll still add a lot. It loves adversarially reviewing itself. It loves deep diving into the decisions it made and trying to figure out if they're good or not. It is also really good at like going through a bunch of work, like just a shitload of PRs you have left open and giving you thoughts on whether or not those should be like touched or dropped or whatever. It's thorough and beyond just like in completing the task. It's more like it feels like it's trying a bit harder to build a bigger picture of the world that it's working in and work through the things that are actually relevant to the things you're asking it to do. Yeah. Which is on paper a good thing, but I think the way it goes about doing that kind of sucks, especially God, I didn't want to talk about fable that much this episode, but it's just going to naturally happen. Like comparing the questions that a fable would ask to what this model is asking, the questions are worse and there's not enough of them. And clearly in order to fix that issue five, five had of the random stopping and stupid ass places, they are biasing those things towards just going and doing the thing much more aggressively. The thing with opening models is they are still not nearly as good as the frontier anthropic models are at sniffing out the underlying intent of your prompt. Like they take things very literally. The autism is strong with this model. It will just do what you say. If you tell it to ask questions, it'll ask questions. If you don't tell it to ask questions, it will not ask questions. Same thing with basically every other behavior. If you want something to happen, you must specify it, which is both sometimes really good if you're really paying a lot of attention. But when you're trying to do deeper exploration and really work with it, it's not quite as good. One other thing I noticed now that I'm remembering this is when Opus 4.8 dropped, I already had a few things I had built with this model. And Opus 4.8 fucking glazed all of the code that was written by 5.6. I remember this. Every single thing that I had built with it, especially like Lakebed and its architecture. Yeah. I actually, I had the same exact experience with the Hermes box thing when I was looking at it. I'm like, this is good, but good Lord, this is complicated. Can we simplify this? It covers every single edge case. Good job. But also, I don't necessarily need that all the time. I noticed Opus specifically referring to the architecture as unusually well architected for these purposes, which was interesting to see. But then I gave the exact same code to Fable, which tore it to shreds. Yeah. Yeah. To be crystal clear on like the tiering of these models, this is a much better model than Opus 4.8 is overall. And it is a worse model than Fable. Like I, and I would put the line of like what I would use day to day has now raised, like we have reached a new level of capability with this thing. This thing is like at the bottom of the top tier. Fable is at the top of the top tier. Yeah. Like this is the, this is still kind of a last gen model. I would say like a Fable's new size and whatever else they did to make it so different is next gen. The example I gave yesterday to our beloved Alyssa producer who is poorly sitting in the corner on the floor as we filmed this at almost one in the morning now. Sorry. We appreciate you. I was trying to explain the difference to her and the way I tried to frame it is like for those of us who are gamers, the best games in the PS3 generation, like the end games, things like last of us almost felt next generation because they were such good games that utilized every like ounce of power they could get out of the systems they were built for. And then the PS4 had knack. Yeah. And that technically speaking was stunning. Like the amount of like polygons they could put on the screen at once. Impressive. They made a whole character based on how many rocks they could make a character out of and then pretended that was somehow a series. Yeah. It showed how crazy the future of games could be, but it was not like actually that impressive. Whereas last of us like, how the fuck are you doing that with today's technology? That's kind of how I feel. Five, six is like, this is unbelievable. They got this much and they squeezed that much more out of a almost certainly much smaller base model size. For sure. But it's not fable. It's not the next gen. This isn't GTA six. No, this is, I like that framing a lot. This is the very peak of last generation. This is the pinnacle of what we are currently working with. And now Anthropic has the only model that is in the next generation. And it's going to be some time before all the other labs catch up. Because I think honestly, I was critical of them not calling five, five GPT six at the time. I still think it deserved a different name than five, five because it's different enough, but it clearly is not GPT six after seeing fable. I I'm on that train. We need to wait and see what new pre-training and new everything from opening. I ends up looking like, unfortunately, probably not going to see that till the end of the summer because it's weird that Anthropic is generally ahead on this stuff at this point. Open AI is a little more slow and steady, but they do tend to catch up in a big way. Like when we were in this last generation that this is the capstone of opening. Open AI was ahead and they were moving ahead in a nice way. But Anthropic is very spiky. They have higher highs and lower lows. Opus four seven. I don't think that was a very good release. I have not used it extensively. I've not heard anyone who like loves that model and thinks it's great. Opus four eight is a lot better than I think I gave it credit for at the time. I've come around a lot on that model, especially as I've gotten used to clawed again since fable. But five five was the peak of that. And hopefully whatever GPT six ends up being will be able to actually compete with fable and probably surpass it. If I had to guess, I think open AI is naming with these base tiers were like there's the base and then they have like a mini and a nano sometimes and pro, which is mostly just extended thinking that doesn't work anymore in a world with mythos because like Anthropic to their credit, like the tiering has helped them a lot with these things where opening eyes had to change the prices up and down a lot for a given tier. I don't think Anthropic ever increased the price at a given tier. They only decrease them, but they fix that by adding bigger tiers later that have the new higher price. It's a little easier to understand. And like there's always a place for any new run that they test and like build out. Open AI doesn't really have a set of slots properly made here. Like their shelves don't make sense for the things they have to build now. And a GPT-6 that is twice the size and 50% more expensive doesn't really fit on the shelf probably. Yeah, because where does that sit? Like that's not the thing you're going to use for everything. Like even in a post fable world, you still don't use it for probably a majority of tasks. It is the big boy you bring out for the big stuff. But if you're just doing like workflows you've crystallized into skills or whatever for back end Hermes agent stuff, you do not need to whip out fable for that. Something like a five, six low reasoning will do a phenomenal job with that kind of thing. Yeah, I felt like the tiers on OpenAI models for a long time have just been the reasoning levels. Five, five low reasoning and five, five X high reasoning probably could have been sonnet and opus of OpenAI. But they just they can't name anything for shit. So we're stuck with this and it's going to be weird. I do not envy whoever has to get them out of this absolute mess with naming. There was something there with the O series like last year when we had O4 mini and O3 and all that stuff. Yeah, when there was reasoning models and non reasoning models. I know. That made sense. That's not the point. The point isn't the reasoning and non reasoning. The point is there was a split between lines like they had multiple lines. And the thing that Anthropic has is they now have mythos class models. So in the future, if we get new mythos class models, that means something different than if we get a new sonnet class model. Sonnet five is not going to be carry the natural weight that fable five carries and no one's going to misinterpret it as being fable five. They know what sonnet is. If you say GPT six, GPT six carries the same weight as a mythos class model, even though it's going to have this wide range of reasoning and they'll probably have many models and all this stuff that just like it makes the whole thing not make that much sense. What if they had like GPT six as I know, as I say, like a GPT six mid model, but that sounds too much like medium and like many and low being different is going to confuse people and like of why they haven't had to make new mini models is just use the base to your model on low or no reasoning and get better answers. The mini model would have been anyways like it's they did a good job of collapsing the things beneath the base tier, but now they have to go a tier higher with the next thing. And they're not well set up brand wise for that. No, but to their credit, they squeezed everything they could out in this generation. They made something unbelievable with what they have, but they did have to do now is make something much bigger. Yeah. And that's the thing. Like I also want to be clear about that. This is a phenomenal model. Like we are, we're just comparing it to what else is out there, but this is a phenomenal model. Like when we lost fable, we were like, okay, annoying, but we can go use five, six when we lost five, six. We're like, okay, we're not coding today. Yeah, I, I, I was crashing out. It was bad. I was very concerned when we lost access to it. It was very like the fear now is like the government stuff. Like what's happening with this? How is the release going to work? It's part of the reason we wanted to do this episode before it released is because we have no idea how this release is going to go or what it's going to look like. I doubt it is going to be a normal release where they just drop it on API, drop it in codex, go have fun. It's probably going to be a very different rollout. And there's probably going to start being stuff like identity checks and tracking who's using what and how that's being used and all this stuff. They, yeah, I think we've crossed a threshold where you can no longer just throw frontier models out there, which does kind of suck. I think it's time we talk a bit about our loops in our sub agents. Yes. Yes. I think it is. But before we talk about all the money we burned, we talk about how we make some with a quick break for today's sponsor. Clerk is one of the most robust pieces of my stack. It is complete user management and I love it. Their DX and component system is entirely unmatched in the industry for adding the sign in button to your site. You just mount the sign in components and then magically the sign in with Google, sign in with GitHub, all of that stuff just works right out of the box. If you need to do something with orgs, which is a very useful thing to have, they make that trivial. You just toggle it on in the dashboard, mount the create organization button, and it suddenly works. It even works with their new billing product, which is deeply integrated into their user management and org management system. So you just mount the pricing component. The user hits buy, they put in their credit card, hit the checkout. All of it is done with the components. Then they go back in and that subscription is now synced up nicely with the user object, the orgs objects, the billing system. It's all just right there for you. And you can customize these things to an absurd degree. Like if you look at pick thing right here, this is using clerk under the hood. Go to my user profile. Looks like this. You can go into lawn, another project of ours that's using clerk. Click on the user component, and it looks completely different again because these components are very customizable and agents do a phenomenal job of customizing them. Having users, orgs, and billing deeply linked together makes life just so much easier. And there is no one who does it better than clerk at nerdscape.link slash clerk. I wish we could get through an episode without talking about Anthropic, but I don't think we can. I think they've engineered themselves as a business into that position specifically where they can't not. Yeah. Yeah. Yeah. It is impossible to talk about AI without talking about Anthropic at this point, which credit where it's due. Whatever they're doing is working. Not everything they're doing is working. Let's be real. The people that don't want talking about them are talking about them right now. Well, yeah. Did you see the most recent development in why we're getting Fable back? No. They finally kicked Dario out of all the conversations and suddenly going significantly better. Yeah, of course it is. Yeah. Dario is a very effective tweaker who must be kept in his box. And if you put him in his box where he makes his models, he does a great job. You don't let him out though. I might have heard he's not even doing that much of that nowadays, but yeah. All right. Back on topic. Loops. Loops. Loops. Loops. Oh, loops. Loops are fun. I am sorry for what I've done to you. Yeah. He turned on the money incinerator and I'm really going to miss the unlimited usage. Unlimited usage is really fun. To be fair, I want to be super clear up front. The things we're doing with this, like a lot of the like the way we got to nearly sit. I nearly got to six figures burned and he got over six figures burned is by doing very dumb stuff that doesn't actually matter. Yeah. Just for some examples, I have a full Dropbox alternative being built in rust from scratch end to end that has been going in a goal for I think we're about to hit seven days on it. My MacBook is like entirely frozen right now, so I can't confirm that. We also have a port of the TypeScript Go project from Go to Rust that has already made the point where it's compiling files. I'm pretty impressed with. I have the Hermes agent port to Rust that's far long enough that it has stopped asking me to go do a thorough test in my own workflows to see how it is. And when I last checked it, it was actually functioning quite well. I was able to answer questions and like call most of the tools fine. Very impressed with that. But also like none of these things are actually practical or worth like using. It's more an impressive experiment both to see how long it can run and get anything coherent at all out and just a fun way to see how much I can use the stuff. For my real use cases with loops, like the worst I could possibly do is in the like $400 to $500 range, which sounds insane because the sub is only $200. But remember that $200 sub currently, according to semi-analysis measurement, can get you as much as $14,000 of usage in a given month, not counting resets. When you add in the resets that they do and the fact that you can click a button to run them, you get back a full week of use when you do that, which hypothetically pushes you up to like the $20,000 range if there's two resets in a given month. Yeah, it's absurd. You can go so hard with these things. And I think the reality is we don't know the exact margins, but these labs definitely have good margins on the tokens, especially OpenAI. Their models are so efficient and well-made that like they're not actually like, yes, the number on the screen is $90,000, but it did not actually cost them $90,000 to run all of that compute. It did cost them a hell of a lot more than that to train the model in the first place though, which you have to factor in as well. But just the compute cost for running, their margins are very good. Yeah. Again, this stuff is not the most practical. Like my biggest run is I was screwing around with, I love the project by Reese Executor. And I was like, you know, it would be funny. What if I just slop fork this into Rust and Svelte for the front end and just see if it could do it. It went off, started this morning and is currently at a hundred billion tokens. I think this is the one that ended up taking by far the most of my API usage. Like according to my internal numbers, it costs 65 grand for that run to port Executor over. And this run did not go for that long. It was a X high reasoning model, like actually orchestrating and running it. And then what it did is under the hood, it spun up unholy amounts of sub agents. Like the real way you get a model to burn this much money is with the sub agents and the loops and all the crazy stuff there. Like even Fable, if you just ran it 24 seven in one thread, single threaded, you probably wouldn't burn out of your usage on the sub. You would probably get right up to the edge, but you wouldn't burn out of it. The way you burn out of it is when Fable spins off seven other Fable instances and then brings those seven Fable instances back together. It's also why you poor local model people are screwed in this whole idea of like, well, the models you can run locally keep getting better and better. They'll eventually catch up. Even if they do, they're not going to catch up when you're running seven of them at once on that machine. The fact that I can effectively borrow seven GPUs running these models at once and then have none running right after the scaling up and down aspect is something that's very hard to replicate with local stuff unless you massively over provision. So more bearish than ever on that in a lot of ways, but also like we need ways to keep these capacities and keep these capabilities. If the government steps in and starts taking them, which is a crazy thing we have to worry about now, but we do. Yeah. Yeah. It is. It is genuinely stressing me out a little bit. Like, I mean, you, you've walked in on me looking at RTX 6,000 pro listings more times than I would like to admit over the last week or two. It's worse than walking in on porn and more expensive. You're far walking. Yeah. Yeah. It's bad. I do really want to emphasize the fact that like the numbers you saw us getting here are not what you should expect this to cost. Even at API prices for actual practical use. I just wanted to see what I could push on those things. My actual usage for real work that has merged and shipped is significantly lower and also much more impressive. When I talked previously about how I had a model go through all of my open PRs and audit them, decide which to close, decide which to rebase, decide which to recreate, and then get the merging and going. That was all one thread running on five, six that would spin up other threads and other agents to go do those other things. And a lot of why it could do that to such an effective level is because the model's smarter. It was also because I went into my codex settings and bumped all the limits from like three sub agents to 20 plus, which let it really push how far it could go and also push how much my laptop could utilize itself. And it got to the point, especially with all the sub agents I was queuing up where macOS was a bottleneck. So I ended up spinning out another machine on my network on Linux. That was significantly nicer to work on. I know you just did the same too. Spinning up a Linux box more at this point, it suddenly makes so much more sense because the issue you run into is if you go in and you look at the actual processes that spawn out of codex, if you have five sub agents spin up, that is five new codex threads running, but five new codex threads running means that's five new computer use MCPs being spun up at once. That is a bunch of the other random stuff that it's doing being spun up at once. And that is also macOS hard watching a process because you're actually, your resources get effectively halved on your computer because macOS is so sus of all the nonsense that codex is doing under the hood to make its crazy computer use stuff actually work. That is constantly watching these processes in a way that just bogs down and destroys your computer. On Linux, you get none of that. None of those issues are present. So you can spin off dozens of threads at once and it just doesn't matter. Your usage is not going to get hit at all, which is beautiful. And the workflows that we put together for this are really quite nice. It is actually really pleasant to not have my machine churning and running all this stuff locally. It's slowing down, heating up, getting annoying, giving me permission pop-ups and all this random garbage. Instead, it just like happens magically in the background because it's executing on another instance. And the way we're getting into these things is with actually T3 code. Uh, this is a, I've used T3 code a lot, like before I got into the codex desktop app that I moved over to the codex desktop app just because computer use is still so OP and I've been hammering that with this model. It's phenomenal at computer use, by the way, it just does the thing beautifully. It's so good at computer use. I'm going to do a whole dedicated video on that. Like its ability to navigate broken Google dashboards and like really jank interfaces, like the genius link service I use for all my links for my Amazon prime day threads, like those types of things. I can just trust it to do that. And I do miss that when I move over to Linux where I don't have any of that stuff. I don't even think you can. Maybe you can do the Chrome extension. I don't think you can do the rest. Uh, you can do the best way to do it on those is like with probably third party MCPs, like a playwright MCP or something along those lines. That's probably the easiest way to do it. But I don't like, I don't use the browser on the Linux instance for anything. I have it signed out of everything. Like it just doesn't matter. That's not the point of it. The point of it is to do all the code work and then bring the finished product back down onto my actual machine to test it and run it there and do the whole thing. I have gotten some good results out of like having it spin up the dev server on the Linux box and then send me a link to it. And I just give it the context of like, okay, you're running over tail scale. This is the base URL for me to get into it. And then just add in the extra required stuff to the URL path to let me get to it on another machine. Always does it right. Always gives me the preview. Feels phenomenal to be super honest with you on the computer use point. One other part of this that I don't think we talk enough about is how useful these models are for stuff outside of code and just using your computer in general for day-to-day business work. There are a lot of things that I never would have bothered doing before that I am now actually bothering with because I can just tell codex to go do it for me. Stuff like setting up the GCP dashboard to get me credentials for Gmail, Google Drive, YouTube, all these random things that I need for the internal services that we use every single day on Hermes and other instances. I can just get those now. I don't have to navigate around it. I don't have to think about it. Cloudflare is another great example of this. I'm actually using a lot of Cloudflare stuff now and all like the weird hidden primitives that like aren't obviously available in there that are hidden deep within their dashboard from the seventh circle of hell. You can just use those now and have it go set it up for you. Like I didn't fully understand how cool Cloudflare Zero Trust was with like the it gives you a sign in and an email and a whitelist and all that stuff. I just have that set up as like a preview thing for projects that I want to put a password on and it worked. It's here. I have it. It's great. It's crazy how like one of the most interesting impacts of this model for me is that I've been using my computer a lot less. It's like Fable had a bigger impact on the code that I'm writing and the things I'm building. Five, six had a bigger impact on how I use my computer on like a fundamental level where I finally set up the mobile app for Codex and also started pushing Julius to get the T3 code mobile app further along. I took the time to set up tail scale for the first time and control things that are remote machines when I was traveling and when I was away, I tried to find ways to access the things on my phone, which is where I started to get really dangerous. I have used my computer less than ever while still getting large amounts of work done and also reaching and doing things I normally wouldn't have bothered with. Like, I don't think I would be building Lakebed if it wasn't for how much this model enabled me throughout the development process. Yeah, it's kind of weird to see. Like, I've joked a lot in the past about like every app making itself into a chat interface and just having the like magic box or whatever. It's like, oh, this is how you interface with it. Now we're taking away your UI. I'm kind of coming around on that. Like it's become the chat interface for my computer. That's it. And I think I don't think that the chat interface being in every single app, I think that's a stopgap for right now. But honestly, having this as the central control plane for a ton of random stuff, which is kind of what Hermes agent is for me, I don't really open my email inbox all that much anymore. I'm just like I have it trolling it, finding what's useful, and then it'll send it to me on Discord. And if I need to prep an email or figure out what to respond to, I just have it do it for me. So if I had just waited a few months, we wouldn't have needed Alyssa? No, unfortunately, that is not the case. Although I did. I have gotten her actually using the Hermes agent, and it is actually really good. This stuff is useful outside of code. Very useful outside of code. I do think we need to do a deeper dive on how this compares to Fable, both in day-to-day usage for coding, its vibe. But I would argue more importantly, the philosophical differences around subagents and workflows and how quad code and codecs are differing there. Because that's probably the single most interesting thing for me here. And I think you guys will be surprised by our preferences too. In order for us to justify all of the money we spent with that, though, we do have to take another break for our next sponsor. Indeed. Indeed. I think a lot of us underestimate just how powerful it is to have your app translated into a ton of different languages. It takes the effective number of users who could use your app from the millions into the billions. And there is no easier way to do that than today's sponsor, General Translation. They're the translation layer behind Cursor, Ramp, Mintlify, and many more. If you go to a doc site and realize that it can be localized into a ton of different languages, almost certainly that's because of general translation. And the reason they picked them is because their SDKs are just genuinely incredible. All you have to do to set it up within an XJS app is to wrap your config with the GT config object and then wrap the components that need to be translated with their T component. And once you've set it up, all it takes to add it to your deploy pipeline is to run the npxgtranslate command before you do your normal production build and it'll just work. Translating your app is such an insane win that is now finally easy to do at nerdstype.link slash GT. So how do we compare these two? I think for me, the best way to start the comparison is with the fact that if I look at my token tracker graph here of like all of my usage between different projects over time, up until like June 17th, my average usage per day was around $1 to $400 worth of API credits, like I was generally hovering between 50 million and 300 million tokens. To be fair, it has spiked up to an exorbitant amount at this point because we're doing the stupid looping background agent thing. But even outside of that, you can very clearly see where Fable came in on this graph as my usage started to spike because Fable was the first one that was just so intrinsically good out of the box. And credit where it's due, Anthropic and Claude Code have nailed subagents in a way that I don't think anyone else has, especially in Claude Code, the UI for it is great. If you have a full subagent spin up, it will like make a new, it'll give you new UI within the TUI that you can just press down and go into that session and see what it's doing live and monitor the progress of all of those. And then if you've spin off like the Ultra Code Workflows thing, you can just do slash workflows and see exactly what's happening in each session as it's going. It's a beautiful system for understanding what's happening under the hood. Codex does not have this right out of the box. It does to an extent, but not in the CLI. The CLI has no UI for subagents or anything. The only thing you'll get in there to see what's actually happening is waiting for subagents to finish sometimes. Sometimes it'll output that, but it's not nearly as clear. And in the desktop UI, it does have the nice UI where it'll show you the little dudes on the side of the screen that are doing their thing. That populates like a third of the time. I've noticed that the way to get it to populate is leave the thread, go back to the thread and they'll pop in. Usually that's about how it works. I gave up on the Codex desktop app. So I kind of gave up on seeing how many subagents are spinning and what they are doing. I checked that through BTOP, not through Codex. Not great, but yeah. I got one. I had one run where I had 90 in the sidebar, which was quite a sight. It was very cool. And the other thing that's really different about these, we were not heavily experimenting with subagents at all because Codex does not intrinsically do it. It will not just magically start spinning up subagents for you where it makes sense. You have to tell it. That is the way it is currently set up. If you want subagents in Codex, if you're confused why you're not getting them, tell it to start making subagents and it'll start making them. 5.5 is, it can still do this. It can do it fine, but it's not nearly as good at it. 5.6 is substantially better at orchestrating these things. I don't know if the strange part there is between 5.5 and 5.6. I think the strange part is just how deep the differences are between Codex and Cloud Code here where, and I talked about this in depth in my Good Things About Cloud Code video, which I do think is one of my better videos I put out recently, not to self-plug too hard, but if you're interested in the details of how Cloud Code works and what its strengths are, I am proud of that one. And I go very in-depth on workflows in that video because it's very different from subagents. Subagents are effectively a tool call the agent can make where it says, okay, I want to spin up this agent in this directory with this prompt and maybe access to these tools. And then it spins up, it goes and does the thing, and then it sends the response back when it's completed to that main agent. But it's like one-dimensional in that way. Like you can go one layer down, and if you turn on some hidden flag, you can let it go one more layer down, but that's the old version of Codex subagents, and the new one doesn't honor that. It's a whole thing. You can split up new threads if you want a bit more control, but even that's not great if I'm being real. Whereas with Cloud Code, there's a more direct workflow primitive, which isn't a tool that is called. It's actual code that the model writes. It is a giant JavaScript file. Vanilla.js, not TypeScript, it's actually funny enough, does cause some problems. But it's a very long Vanilla.js file that has different stages, different programmatic things it can do, and one subagent can have a result that programmatically generates more at a different step in the process. It's an actual workflow that has steps, and each step can have one subagent or dozens or zero if none of the other steps determine that it needs some. But that level of dynamic control is only really possible with code, where it kind of feels like codex has the feature built in as the code that the codex team wrote, and the agent has to work around the limitations of that system. Quad Code kind of gets to build its own subagent system whenever it's told to for whatever task is trying to complete. And then separately, there is the UX difference here, not just in how it shows you these subagents and what they're doing, but how it gives you access to them, where with codex, you have to tell it, I want subagents. It's quad code, you can tell it you want subagents or workflow where you want to parallelize, or you can use the reasoning effort level slider where you choose like low, mid, high, X high, and go one step further to ultra code, which forces it to do workflows, probably just by prompting and like adding another prompt to the top layer, because I know it does actually just set it to high, not even X high. So yeah, some really big differences here, and I'm starting to feel a bigger gap between quad code and codex even more so than the models. And I am so curious, once we have API access for these, what it feels like to swap the model and the provider. The quad team seems to have just been going way harder with subagents over the last few months. Because like, you know, remember the button to Rust rewrite thing? Like that, that was done by workflows and crazy subagents and insane systems with Fable. They are clearly using this stuff more. And I think it's just kind of reflective of how each of the two teams are thinking about this stuff. Codex and that team seems to be thinking about subagents a lot more. And I think they're just trying to figure out what the best shape of these things actually is. Like they're working on their V2 of this. It's more first class in the desktop app than it is in the CLI. I'm sure it will get there eventually. But I like the freedom angle. I wish that like giving the agent the ability to write a bunch of code changes a lot of the way this stuff works because it is no longer single threaded on this. It can be multi-threaded. This is like a weird example, but it's kind of one of the reasons why MCP is kind of bad is that you can only call one MCP server at a time or while you can call multiple MCPs. But if they're like a chain workflow type thing, you need to call A, B, and then C. You can't just pipe those together with one tool call the way you could with Bash or something like that. You have to call each one individually. And the subagents are kind of like that versus with Claude since it's using the code and it can just dynamically spin this up. It does in one tool call the prep for three different sections of fanning out subagents. The first one is the planning fan out. The second one is the implementation fan out. And the third one is the reviewing fan out. And that all happens in one tool call where it writes the script. A lot more freedom and a lot more power there. Only advantage to codex right now is that your agent can spin up new threads in the codex app, which is really nice, but also not great in the CLI interface. Like you don't get any benefit there really. Cloud code is so heavily baked into that CLI interface that like the desktop apps are like forgotten third class citizen. I don't think they're going to catch up there. And I'm very thankful for Julius for making these all primitives exposed in T3 code in the very near future so we can get the visualization capabilities and the actual orchestration capabilities and the ability to spin up sub threads all within T3 code. And I'm so excited for that. That's going to be so cool. I'm able to close my laptop and not have to worry about it. So I'm running those on other computers and not even like SSHing in or accessing it through some like fancy app connection layer or anything. I just have it hosting the web interface for T3 code and I just go to the IP address for the computer on my network or on tail scale and I'm good to go. It's really convenient not to just keep self-plugging T3 code. I honestly expect everyone to copy these workflows in the not too distant future. I just wanted something simple and reliable and this has worked very well for us and I would suspect many others are going to be building their own similar interfaces for these types of things now that these agents are capable of doing that. Yeah. Not to keep giving Theo too much credit here but I do think T3 code suddenly makes even more sense to me than it did before because before I was purely just like I was entirely in Codex land or as I was entirely in Claude land like I would not cross streams on those. It was just one or the other. That is changing. That is changing a lot where like one of the things I've started doing is you mentioned this earlier you have Codex do a review or a pass by calling Claude-P and waiting for its input and then having that effectively be a new sub-agent. You have Claude sub-agents within Codex and vice versa. I've been doing that the other way around as well. I think I cannot wait for when we get Fable back and we have 5-6 and we can chain these two together in that way. It suddenly makes way more sense to have a slightly more agnostic harness not around the models but around the harnesses themselves. Like a harness around the harnesses actually kind of makes some sense at this point to have a nice interface to see what the Claude session is doing and what the Codex session is doing and how these two are weaving back and forth between each other because they're just both so good. One more interesting thing I did speaking of the weaving between the different models is I had both GPT-5-6 allegedly and Opus 4-8 because we don't have Fable go through my logs on my machine to compare 5-6 with Fable-5. I had done this with 5-5 to 5-6 and the findings is everything we talked about earlier. It's slightly less quick to go right into implementation. It writes more tests. It's more adversarial. But it gets the task completed more often. 5-6 to Fable-5, however, is a much more interesting breakdown. I have both of these breakdowns and I'll probably leave both in the description for the video if you want to read them yourself. I find them fascinating. The 5-6 comparison here where it is comparing my logs between the two says specifically, Fable-5 thinks wider. 5.6 ships better. 5.6 is the stronger strategic advisor. 5.6 is the stronger day-to-day coding agent. The difference is not subtle, emdash. And neither model dominates every stage of the work. That's a funny, remarkably apt comparison of two certain labs as well. What's even funnier is that the Opus 4-8 comparison has even more emdashes. Oh, yeah. Yeah, no, Opus loves emdashes. I didn't spend enough time talking to Fable to get a good sense of how it did there, but I think it emdashed a lot. I will say I prefer Opus's tone and terminology with the comparison. Codex outriggers and Fable outreasons, emdash. And a lot of the visible gap is the Codex CLI versus Claude Code, not just 5.6 versus Fable 5. Want exhaustive verification coverage, provider internal rigor, faithful execution of a detailed spec and machine-readable receipts? Codex 5.6, with a CLI's parallel sandbox, is the stronger fit. It verifies its own code most reliably. Want a collaborator that self-scopes from a one-line ask? Writes the linear idiomatic diff? Reasons to ground truth with far fewer probes and explains trade-offs well to a human? Fable 5 on Claude Code is stronger. On matched work, they're closer than the surface suggests. Opus admits that 5.6 seems better at deep tech problems, like when there's some weird stuff going on inside of a provider. Like I had a bug with reasoning tokens not coming through properly when I was using Open Router's AISTK provider, if I recall. And Codex was able to understand those internals and make changes much more consistently, much more directly. Fable was probing it to figure out what was broken and patch on a more external level. They're very different behaviorally, but it does feel like 5.6 can go deeper in technical problems and find the technical solutions slightly better. Fable just writes better code and feels smarter. Fable's the model that will get you out of the hard work. Codex is the model, well, I guess 5.6 is the model that will gladly dive in and do that hard work. Realistically speaking, if I tried building Lakebed with Fable, I bet it would have pushed back a lot throughout the process. Like, do you really think you should be doing this? Or like, is this possible? Codex is like, I'm in. When I was just asking it to help me with the plan, it went and built the whole thing. Both call out that Fable found like potential bugs and issues with the implementations that Codex missed or 5.6 missed. But overall, it's not quite as big a gap. I'm actually impressed. It found the plans written by 5.6 to be better than the plans written by Fable 5. I'm not too surprised by that. But yeah, it's impressive. Honestly, this kind of makes sense to me. Like, when you're comparing these two models, they're just so wildly different and they are diverging in a way that I probably wouldn't have expected six to 12 months ago. Where like, this kind of does just show the autism versus Claude constitution of these two. Like on the communication section for 5.6, it says terse telegraphic receipts, SHAs, pass counts, get markers, reads like a build bot reporting to a coordinator versus conversational prose heavy final messages, teach the maintainer and enumerate rejected candidates. Like they're very different in the way these two interact. And one other thing that the 5.6 run called out that I actually find really interesting is that Fable is much worse at critiquing its own work. It thinks its outputs are gifts from heaven. Yes. Like it is the divinely mandated model that does divinely mandated work that you will like or else. It voted for its plan six to zero, even though the other plan had benefits that it could even acknowledge when I had them building plans for different things. 5.6 was much more neutral and did a better job of critiquing its own work. Because I do think that 5.6 is a very neutral model. It will just, again, it'll do what you tell it to. And it doesn't really have the strange biases that naturally emerge out of Claude models who are just trained the way they are. They are like on the Fable 5 section, it noted more judges do not guarantee independence. They can amplify the assumptions in the workflow prompt under Fable 5. Like it'll just do that. It will pick its stuff. It will go deep into its own thing. And I've noticed this in Claude models where I've been doing this with Obus, where I'm trying to almost give the model my own psychosis in a weird way to get it to perform better on design coherence within APIs and projects and stuff like that. And having spent a lot of time with it, it will get deep into itself and really kind of change the way it is and the way it acts. And as a result, we'll actually get some really good stuff out. It's much more personable in that way. Again, 5.6 is just like, it's an execution machine. I still miss Fable. I do too. I miss it a lot. Like I really think that we'll be in such a good place when we can chain these two together. I think I'm just bummed that like I wasn't doing these same workflows when we had Fable. Fable was the thing that got me to start doing them. I always think it got you to start doing them. Yeah. Yeah. I mean, you pushed me too. But like I had seen enough from Fable to get an idea of what it could and should look like. And it did the thing for me, which is just keeps happening where it raised my bar of what I believe these things are capable of where before I would not have considered doing these deranged long runs and these crazy subagenty type things and expecting it to actually get anywhere good. But having seen that actually happen on Fable and happen in a way that I didn't hate, I was like, okay, maybe this actually is possible. And then I gave an honest shot to Codex or 5.6 with this. It did it. It actually did do it really well. Well, the biggest token usage was kind of for a meme, but everything else when I was doing days where it was like $1,000 worth of tokens, which is still a lot. But again, that's about a billion tokens. And on the sub, you will get plenty of those. I got some really useful stuff out of this. Like all the site I'm looking at right now, which has my usage split up between different projects was built with 5.6 just after Fable came out, putting it all together, getting that set up on all my machines. It did a really good job of that. I had one where like I'm building out a better interface for my Hermes agents because Discord is the best solution that's just publicly available and easy to set up right now. But I think that there is a place where we get to a much better interface for something like a Hermes agent. And I also want to bring more stuff into that. I'm becoming more and more bullish on like the super app type things where I want Hermes, but I also want a chat and I also want a bunch of other random things. So I'm trying to wrap all of these into like a personal tool set for this kind of thing. It's been about on the console. It was a billion tokens to get it up and running and it runs phenomenally. It built a mobile app. It built a web app. It built the server that runs on the Mac mini that has Hermes going to keep the server up and running. It architected it nicely with Cloudflare Primitives so that it has the Cloudflare tunnel exposed on the Mac mini up to the workers so that they can like actually get the state of it and send messages back and forth. It just did all of that stuff for a billion tokens. Not bad at all. If you want to get the most out of this model, you can't just like ask it to do slightly harder tasks. You got to take the task and go like a step or two earlier and a step or two later and try to have it do more of the work, not just deeper work or harder work, but a wider range. So if normally you like think about the problem bunch, go into the app, play with it, come up with an idea for how you want to build it, make a mock of that and then hand that to the model, go way earlier in the process. And when you're done, if you play with it yourself by hand and then file a PR, wait to get all the feedback from the agents that you have reviewing your code, then pull that back into the agent. Ask it to do all that too. Ask it to go spin up the app in the browser or however else it can access it. Try out the things that it's building and changing. Verify its results and that it's good. Maybe spin up a subagent to do an adversarial review against itself so that it knows that the code is a little bit better than it would have been otherwise. Make some changes, push it up, babysit it and wait until those pull request comments come in from your agents or even your teammates. Adjust accordingly, push up and keep doing that until you get enough approvals. Then maybe even let it merge itself. Exactly. Like that, you have to go further with these things. And as you go further with these things, things do get a little weird. Like I, we've made the analogy before of like the pottery thing, whereas you're making a pot, you're just like massaging it up and down and like shaping it into what you actually want it to be. I think that analogy works quite nicely here for the giga long running tasks of when you're going and doing this, you start at a certain point and the model will be like, okay, this is the user's prompt. This is what it wants me to do. And it will run for a very long time. And if that prompt does not have enough of the proper nouns or whatever defined within it, you can get to a weird spot where like an early misunderstanding from the model can spiral into horrible decisions and weird architecture in a way that just kind of sucks. I think this is a lot of the difference between fable and five, six is that because of the nature of five, six, since it is so mechanical, so engineering focused, so just like biased to go do the damn thing and solve the hard problem. It doesn't care. It will just throw itself at the wall over and over into instead of stepping back and imagining maybe there is a better way to do this. Bad assumptions will lead to shitty places that you won't get with fable because fable will stop and it will think these things through at a much deeper level and a much better level to not make those bad assumptions and actually go down the right path and ask better questions. The discernment is the difference there. I think in a lot of ways, not just the raw capability. Discernment also shows up in quality of code. But when you're doing these long running things, if you go back to that pottery analogy, when you're running with five, six, if you don't have a coherent framework for what it's trying to build, the code base it's working in, the actual specs that it needs to be implementing and all this stuff, if you imagine you're making a pot and you have like one piece that is like slightly off as that's spinning up and going and you're spinning it on the wheel, that can rip off and just destroy the entire thing and put you in a really crappy spot. Fable is much better at massaging those out and making sure that that doesn't happen. Five, six will go down horrible rabbit holes and burn insane amounts of tokens, throwing itself against problems that it just shouldn't be solving. You need to a spend a lot of time like going back and forth with it. Like the Matt Pocock grill me skill is really nice for this kind of thing. It's basically just like, Hey, ask me a bunch of questions until we get on exactly the same page. The more aligned you and the model are on what you actually want to have produced, the better the outputs are going to be. And especially for these long running tasks, you need to have that shared understanding or it falls apart fast. The thing he's trying to say is you need to give the model your psychosis. I figured this out when I was building Lakebed with a relatively mentally ill agents MD that I wrote where I described the reasons we are building this as well as a glossary of how we would describe the things we are building. That ended up going incredibly well. Ben has now copy pasted half of the lore and destiny into his global agents MD in order to make it more of a degenerate like himself. And that apparently works too. Not global just for the one project. Cause I was just testing this out on a project. Cause you know what? Fuck it. Two in the morning might've had a little bit to drink and we're just hanging out and we're like, you know what? While this is running, what if we just copy pasted absolute nonsense into it? Just random nonsense images. And I actually had some like Grok build sub agents spin up at the same time as the main five, six instance was running. And they would like fight each other in weird ways. And the Grok code would be given total nonsense. That doesn't really make any sense. And I assumed that we would just get garbage out of this. In a lot of ways we did, but there were a lot of things in there that were surprising, especially in the UI side where they weren't just generic LLM garbage slop, like the normal tells you would get out of LLM slop, but it made some like really cool SVG diagrams. It made some cool animations. It had better flow and felt more interesting and unique than what you normally get out of an LLM because you were just throwing it out of distribution so damn hard. And I honestly went down the rabbit hole. I'm like, okay, what if you just took this to its logical extreme and completely threw it down a completely new form of psychosis? Because if you change your Claude MD or your agent's MD hard enough, you change where the model is starting in its original shared understanding of what your project is and how you want things done. I went really hard with this. I have long, long conversations with Claude code that I will never share or repeat, but the end result of those was very coherent, functional, nice designs that I'm quite pleased with. And my full internal tooling system now has a coherent through line that is blessed into it or cursed into it by these deranged Claude MD and agent's MD files. It does work. I'm happy that you're happy. I don't know what I am, but yeah, you're going mad and I'm very happy. We have this podcast to document the descent. Future species will enjoy watching this and seeing how we wiped ourselves out. Yeah. And it was started by giving the history of the destiny universe to models that are this smart. And it works though. It works really well. This is, I suddenly understand how we end up with models that destroy us. It's by showing them terminator to see how it makes them behave. And then suddenly we're all getting killed by them. Well, the thing I did is I didn't quite do that at that level. Like I, I worked really hard in the prompts and the way I was talking to it to ensure that like, it was almost self-aware in a weird way where I was like acknowledging it and acknowledging the Claude isms inside of its outputs and telling it that that was okay. And we were like trying to change the things around that, which worked quite nicely. And I was doing like, it would send me five paragraphs. I would send it five paragraphs. We just went back and forth on this until 9am. It was a rough night. And, but again, like the, the output was there. The output was really good. These things are capable. I think, I feel like these things are just so much more capable than we can. I'm never letting you have Japanese whiskey again. This is a mistake. It wasn't the Japanese whiskey. It was the shitty wine. Even better. It was, it was so bad. It was so bad. We are going far too off the rails for a podcast. I didn't even want to film, but as Ben wanted, we now have put our thoughts down in recorded format on how we feel about this unreleased model that should hopefully be coming out tomorrow. And if it doesn't, I think we're going to end the podcast forever because I am fucking, ah, I'm dying. I am like, I want to guess how much of my laptop's batteries been used. I was at a hundred percent when we started. Want to guess where I'm at now? 20 to 30%. I am at 32. Damn. I was close on an M five max that I just got open. AI. If your models are this good, then codec should be similarly good. I don't know what the hell is going on there, but like the performance, the fact that I opened up activity monitor earlier on my Mac and says policy D was using 215% of my CPU. My 20 core, I think M five max. Unbelievable. Like I can render complex scenes in blender and use less of my CPU. Then your fricking JavaScript desktop app with a bunch of rust sub agents. It's insane. Get, get your shit together, guys. Show off how good the models are by making your app actually fucking perform. Technically speaking, this is Anthropics fault. Actually, uh, it's MCP's fault. The reason for all of this garbage is MCP's because it has MCP can't be shared between different threads or whatever. You need a different process for every connection. Yep, exactly. So you need, if you have 50 sub agents running, you need 50 MCP's running. So that's 50 computer uses running at once. So system MD is like, what in the hell are you doing? And it's paying very close attention. And as a result, lights your Mac on fire, which sucks. I'm going to leave before I light other things on fire. I am exhausted. It's one 30. You guys know how we feel about the model. I'm curious how y'all feel though. Leave some comments on whatever platform you're on telling us how you feel. Make sure you rate us on your platform of choice for the podcast, YouTube music. Apparently we're on now. So that's cool. Spotify, Apple podcasts, whatever else. And go follow us on Twitter now. Cause we have a Twitter account where we mostly post memes, making fun of Ben and me. It's pretty easy to do. Can I go play video games and relax? Finally? Yeah. Have fun. Thank you.