← Back to search

5 different models dropped last week & the GPT-5.6 usage limits are brutal

Nerd Snipe with Theo and Ben · 2026-07-14 · 126 min
relevance 76 26071 words Episode page ↗ Audio ↗
Show full episode description
This week Grok 4.5, Muse Spark 1.1, GPT-Live, and GPT 5.6 all dropped. ⁠Theo⁠⁠⁠ and ⁠⁠Ben⁠⁠ break down which models you should care about, how OpenAI fumbled the Codex to ChatGPT app transition, the abysmal usage numbers for 5.6, and the latest drama surrounding OpenAI and Sam Altman. Plus, how many satellites would it take to trap humanity on Earth, and how many data centers to heat up the ocean? Thank you to this episode's sponsors: Composio & WorkOs. https://nerdsnipe.link/composio https://nerdsnipe.link/workos Sources available on Substack: https://nerdsnipe.substack.com/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Five new AI models dropped in a week while punishing usage limits disrupt everyday coding workflows.
Benefits
  • Workarounds for hitting Codex usage limits fast
  • Guidance on which new models actually matter for devs
  • Real hands-on impressions of GPT Live voice mode
  • Context on Grok 4.5 and Muse Spark 1.1 quality
Use cases
  • Simon Willison using GPT Live as a talking buddy via Meta glasses
  • Live voice used as a driving assistant linked to email and tools
  • Real-time bilingual conversation partner between two languages
  • Attempting to build a full paid service end-to-end by voice
KPIs / results
  • 5 models dropped in one week
  • Hit Codex limits ~10 times in a few days
  • One reset burned at ~40% usage
Tools / build
  • GPT Live
  • Codex
  • Grok 4.5
  • Muse Spark 1.1
  • Whisperflow
0:00 / 0:00
What's up, Nerds? I'm Theo, he's Ben and we did two episodes of the podcast last week because there's just been that much going on and believe it or not, it hasn't slown down. No, we ended up with how many models? Three, four models last week? Was it an insane week? Five at least. Five? Yeah. I guess if you count the real ones. Terra, Soul, Luna. Oh, well that's cheating, but yeah, fair. They are different models. They are quite the models, if I'm being real. If that was all we had to talk about, it would be a full episode, but we also have to talk about all of the chaos that we've been causing, all of the chaos that Sam Altman's been causing, the lawsuit between Apple and OpenAI, Elon crashing out over all of that, and what else do we have? Oh yeah, the chaos of everybody hitting their usage limits. I've basically never hit one before and I've hit the limits on Codex like 10 times in the past few days. It's bad, but we have advice on how to work around that. How many times have you hit the limits? Twice. I hit them once about five minutes ago. I just had to use one of my resets to fire off the jobs that we run before we go. How much was it when you hit that? About 40%. Okay, it's not too big of a loss. It's not great, but it's fine. Thank you to today's sponsors, Composio and WorkOS. Let's get into it. Let's do it. Shall we start with the models in order of least to most exciting? Yes, absolutely. Cool. So that means we're starting with the new Muse Spark 1.1, if I believe. 3.1. 3.1? I think so, yeah. Why would they only had one version? Oh, it's not on OpenRouter? It's not on anything. It was 1.1. I was right. I don't know. I'm clearly hallucinating something, but yes, it is 1.1. Yeah, it's only through their API so far. It's seemingly, depending on who you talk to, pretty good to find. I'm curious what you've heard about it so far. Supposedly, it looks more competitive than anyone was expecting. It seems about around that Opus tier, but I haven't heard anything specifically incredible about it, and I've not seen anyone serious using it. I'm trying to remember who I heard this from, but somebody I trusted either tweeted or hit me up about it and said that it has the weird, like, uncanny feel that some of the earlier Openweight models did, where it, like, given a very well-specked-out task of work that's been done before, it's relatively good at, like, dealing with that. But outside of that, when you're, like, trying to do anything more traditional, like a vague request or go explore, it loses its shit really fast and has, like, the Gemini Flash style of, like, looping on things that aren't correct. Yeah, like that weird Quenny behavior where it'll just reason in circles over and over again and probably duplicate reasoning tokens for bizarre reasons. I wouldn't be surprised. I, this did not seem to be all that exciting of a drop. The only thing interesting here is that at least they now have a model that benches as if it is a more serious model, and perhaps this is the beginning of them potentially being competitive, but time will tell. I did see one great joke about this release. Meta's been training their models on the screen recordings and other data they've been collecting from all of their employees as of recent, which means this model is state-of-the-art at one thing, applying for jobs at Anthropic. Aren't there a lot of Facebook people at OpenAI, though? I feel like a lot of them have gone there. I don't know that many. I know, like, one or two, but I know a lot more, especially on the ML side than it's at Anthropic. Everything we've talked about in the past, it does not sound like a good place to work. 50 plus percent of their engineers are just labeling data at this point. They're just grabbing old Claude code sessions and trying to label it into something competent that they can try and make a model out of. Wait, do you think engineering and data labeling are different? I thought those were the same. In 2026, God only knows. I believe that is definitely the least exciting one. Maybe they'll do something in the future, but the one that is more interesting is definitely Grok45. Do you think it's less interesting than GPT Live 1, though? Because we're going in order of least to most interesting here. Uh, well, as far as models go, I think Live is far more interesting. But as far as implications, 4.5 is more interesting. So we'll talk about Live first. Implications for devs, to be clear. Yes, of course. Because Live 1 is much more interesting for, like, my parents. Oh, absolutely. And just the general future. But let's talk about Live first. Talk about Live first? Okay, we can do that. Yeah, let's do it. So, GPT Live is a new model from OpenAI that is really focused on the conversational aspect, making it easier to talk to the model and have it talk back. There have been a lot of funny clips going around of people pushing the limits of what OpenAI's voice models can do in the ChatGPT app. The one that I absolutely love is the guy trying to get three of them at once to count to 100, and they keep talking over each other and interrupting each other. That's what Live 1 is actually built to be good at, where it can interrupt you and talk and understand when you're talking and it's finishing and really feels to understand context a lot better. I only briefly tried it because I'll be real. I'm normally surrounded with my team. Hi, Alyssa. Hi, Jeff. So I don't get to talk to my phone particularly often. I barely even get to use Whisperflow as much nowadays. I love it when I can, but I just don't want to talk to my devices. But like if I was driving or walking around town, this is really cool. And from the little bit I tried, it is actually really, really impressive. It feels much more conversational. It also reminds me of the difference from going to like a bad WebRTC app where you're like trying to do a phone call on Zoom or something with someone and the latency is really bad over to something where latency is suddenly really good. Or in person, you just stop stumbling and talking over each other and having those awkward pauses. It just feels better. And the model is smarter because it's trained on all of their more recent shit. I had a very similar experience. It's the UI that they have in the mobile app for it is gorgeous. It is really, really cool. And it feels quite nice to use. I'm just not a voice guy personally. Like I like voice to texting on like a Whisperflow type thing when I'm at home alone in my office just chilling. But outside of that, I'm generally speaking typing. And the way I use these things is just not that conversational. Generally speaking, you're typing. Shut up. I can't help myself. Shut up. Shame. No, we're not going to do that gag. Whatever. Well, we'll go back in. But yes, as I was saying before I got rudely dad joked, the model is good. If I don't like what would you say the use case for this is though? Oh, there's a ton. And it's like similar to how we are now getting so creative with using all these LLMs for code tasks and finding different ways to chain them and use them together. I'm seeing people do really creative stuff with GPT Live as well. I know Simon Willis has been using it a ton just as like his talking buddy when he's out and about. I think if I recall, he set up his meta glasses or something so you can talk with it through that. And it's really good for that type of thing. I've seen a lot of people who use it for like their assistant when they're driving, having it linked up to all their fancy tools and email and stuff where they can just talk to it and it will talk back. One of the more interesting examples I saw was somebody using it for translation, not just the usual. I talk to it. It translates. They talk and it translates back. But as like a conversational partner that understands both people's languages and can like know how to interrupt and intervene between things. Someone shared in our early access channel that they were using it with their girlfriend who speaks a different language and they said, I love you in English. And when it told them that that's what it said, the model inserted, I heard in their tone that they really meant it. It's clear this guy really loves you. And I thought that was kind of cool that it like isn't just blindly following the task. It's much more natural, which is interesting and definitely won't have more psychosis problems going for it long term. Yeah, wasn't the in that legendary psychosis video from Eddie Burback. He used a ton of voice, the voice mode stuff to teach him to talk to rocks because 4-0 is a beautiful model. I don't know what the because this is a full separate model. Yeah, it's not just using five six. You've effectively been talking to 4-0 if you were using the voice in the chat GBT app like still to this day. He's so now it's like a huge leap. So anybody who was using voice mode just had a like a generational shift in its capability. So that probably means it can actually do tool calls because in my brief testing of it, it was able to pull emails and search the web and do that kind of thing competently in a way that like I mean, we went back and tried out 4-0 mostly as a gag a couple weeks ago just to see how bad it actually was. It is so much worse than I remembered it. It is horrifically terrible. I do not know how anyone could ever have used that model for anything. It can borderline not call tools at all. So if tools are now available in voice, that'll be interesting. Something that I've been considering doing is trying to build like a full service from scratch only using my voice to see how far I can get with it. Like can I actually spin up a service use computer use and do all these things from my phone with my voice to like build a full thing end to end? Well, what are the parameters on this? Like are you allowed to just use whisper flow and then remote app into your like codex desktop because I feel like that's I want to use live one from codex on my phone or from chat to me. Codex is dead, which we'll talk about later. Yes, from chat to me on my phone. I want to be able to just use live one and talk about what I want and see how far it can get. Like can a computer use to set up stripe and everything like an actual paid service from scratch just talking. That would be cool. If it has the tools probably like I just don't know if it would be able to actually interact with any of the things you would actually need to interact with there. Like can it send off codex threads because it's not going to write the code if it has any support for plugins that I can probably get it to make a plug in. Oh, I guess that's fair. Yeah, work through it from there. I'm almost certain that they have them like you can add custom MCPs to chat GPT dot com and those custom MCPs are just new tools that the model can then call within the normal harness. And I think the tools that you get from voice are just the normal chat GPT tools. So theoretically, that should be possible. It's an easy way to figure out. Yeah, do it. I'm curious. Are you able to do computer use work so you can access my computer through the voice to text model that I'm using right now? I can't access or control your computer. Are you sure? Yeah, I'm pretty sure. In this chat, I can talk you through steps like but I can't actually click type or view your screen. Disappointing. If I was to ask you to build an app, do you have the capability of doing that in some way? I can definitely help you think through the architecture and help you plan it out, even help with some code if that comes down to it. And debugging advice as well. But I can't actually run the code or click around on your machine. I can be your death partner, but not really your hands on keyboard. Do you have access to any tools or skills that you can use that might make you more capable as a voice assistant? Good question. I do actually have access to some tools and skills. It just depends on where I'm running. So here in voice, for instance, I can help you plan, write code, you paste and debug with you. But I'm not actually able to run or interact with your applications. I'm unimpressed. Yeah, I think it does have more tools than that, but it definitely does not seem to be all that intelligent. The the steering on that was rough. It clearly got stuck on the I'm trying to build an app thing. It was on intelligence level instant things I will play with more in the future. The way it handled pauses itself was not as good as I was hoping to, especially like considering the demos I saw. It felt awkward in a way. Yeah, you could interrupt it pretty well. The the cool one that I heard a lot of people talking about was like telling it to interrupt you repeatedly at random times when it thought that it should. And that worked better than I expected. Yeah, but yeah. Overall, this could be very cool in the future, but at least for right now. Not they're really good model for people who don't have very many friends. Yeah. Unfortunately, that will probably be the use case until it gets competent at running tools like most models. Speaking of models for people who don't have many friends. Grok 4.5, the first really good Grok model we have gotten since Grok Code Fast. Did they pay you to say that? Nope. I do this for free. I do it for the love of the game. Is somebody paying us? Hopefully today's sponsor. The agents we have access to now are absurdly powerful, but the problem is they can't connect into all the different services we use. What if you wanted to be able to see Notion and Figma and GitHub and Linear and Slack and so many others? Composio is the easiest way to do this by far. If you're building an agent, their platform and SDK makes it insanely easy to hook this stuff up with managed auth, a full trigger system that makes it so that as something changes on a data source, your agent gets that information. It's bidirectional communication. And since this is an SDK, it's fully model and framework agnostic. You can use this with Cloud Code, Codex, AISDK, PI, whatever you need to use. Composio fits inside of it. And if you're not building a new agent, you're just trying to hook tools into your agent. Composio for you is one of the best ways to do that. You add one new connection, MCP or CLI, whichever you prefer, and you suddenly get access to over a thousand different integrations. They handle one of the most painful parts of this for you, which is multi-account auth. Well, oftentimes if you're building a really useful agent, you want it to have access to a lot of different Gmail accounts. Composio just handles that for you under the hood. The auth scopes and rules and everything lives on Composio's side. So there's no painful process of getting this stuff connected in. You just sign into it once and now every single agent you have just gets full access. Your agents will get so much more powerful when you give them access to all of your tools at nerdstype.link slash Composio. Sorry about the interruption. I hope you guys understand we have a lot of subscriptions to pay for now with all of these drops. And now it looks like I'll be using cursor and I guess my free Grok subscription I get through my Twitter premium. Sorry, my X premium. Is that what it's officially called now? They change the name every other week, I swear. Regardless, it comes with Grok and now you can use that in Grok build, which isn't too bad. But more importantly, it's Grok 4.5, which is actually. Yeah. Last week we were, we got a message from cursor and they invited us to early access test one of their new models. And we were very excited to see what this was. We thought it was going to be composer three until the next day we realized that we were actually testing a Grok model. It was a pretty good Grok model. I'm not going to lie. It, there were a lot of weird things that I ran into during the testing. Specifically, it would like malform pushes and commits and pull requests and stuff like that. Like it, I watched it try and make a PR 10 times in a row and just fail. I don't know if that was a model problem or a cursor problem or just an early access problem. I think that was an early access problem because I haven't had any of that since. And I also had a lot of weird problems. Also like disconnects in it, like failing to generate responses and things. All that disappeared when it stopped being early access for me, at least. I have had a similar experience. I just haven't had the volume of usage I've had on the models we're going to talk about later. But from what I have used, I probably did five or six PRs on this thing. UI wise, it is nothing particularly special. I don't think it's any better than anything else that's out there right now. But capability wise, it ran sub agents really well. It handled the like loops, especially the built in cursor loops really, really well. The speed is excellent. Like even just on open router right now, I'm looking at it. It's currently averaging one Oh two TPS and the price. The price I feel like is the most interesting part. It's $2 in $6 out. The price is interesting. I think there's a different much, much, much more interesting part here. This is the only model I've ever used. That's not from one of the frontier labs that actually stays on task and can like be given three things in one message and do all of them and not get like confused and lost. Even certain models from like the top two labs. I've had problems with this in the past, like same GPT five five. If I gave it three somewhat disparate things, it would get stuck and lost on them. Yep. It would do like I. I had one thread I did with Grok four or five where I had it like monitoring a couple of PRs and giving feedback on them. It wrote a bunch of feedback. I responded to like three of the points asked it to make these two PRs and then go make this change to this other thing. And then, oh yeah, I just realized we have this other problem too. Can you make that a change separate on another branch as well? All in one message and it got all of it and like knew how to handle all that. Yep. That's a skill that only open AI and anthropic have made models that do. This is the first one that's not just like usable in the sense that like it's good at code and it benches well or it's cheap or whatever. It's usable in the fact that it's like actually fucking usable. Yes. As a developer, like this is the first time I used a model that wasn't from the big two where I didn't feel like I was compromising in my like user experience. Obviously, it's not as smart or quite as capable. It does have weird failures here and there, but it doesn't get lost the way I expect from like a Gemini model or an open weight model. Now it is substantially better like 100%. This is a model that if I was forced to use, I would not hate everything. Like honestly, if I had to pick between five, five and Grok four, five, I would pick rock four, five all day. If I only had one model, I think I would too. The Opus plus five, five dual wield is reasonable. Yeah, but Grok four, five is like it's a kind of an in-between of the two because it's much more token efficient and cheaper. So it's like five, five in that way, but it could do longer running things without having to be reminded to keep going like five, five needs more like Opus in that way. It's an interesting model and you can you can smell cursors RL and training on it for sure. Oh, absolutely. But it like everything you said I agree with. I had a similar situation where I had two PRs that it made and then I wanted to merge them into one PR and then also split that merge PR into like one goes in first and then another goes out like weird PR manipulation stuff. The kind of thing that I've really only been doing with fable and five, six failed horribly on five, five. It did it very well on Grok four or five. I believe this. It's a good model. It's a really, really good model. And I think like the main point I wanted to make on price is that we this is a model that can get not at the frontier. It's nowhere close to fable and five, six, but it is a really good model that you can get crazy amounts of usage for a very, very low price. Like we are burning insane amounts of API priced usage on the big frontier models. This guy I've had like the I added the Grok sub thing, whatever it's called to my vibe proxy usage tracker thing to just see how much I was using on it. I burned a ton of tokens and barely touch my limits. It is not an expensive model and I would. It's like on the $20 tier, isn't it? I know that this is the one through the X premium subscription. I honestly do not know. I think that's like 10 or 20 a month. The Grok quota, but like the subscription through X is only 10 or 20 bucks. No, it's it's more than that. It's it says that the monthly credits are you get $200 plus a weekly limits. The weekly limit. What are you paying for that though? I just it's included with your Twitter sub. Yeah, so that's what I was saying. There's like your Twitter sub is like 20 bucks. You're getting all this usage for that. Oh, yeah, yeah, yeah. Exactly. Sorry. I thought you meant you were getting $20 of the sub. No, no, no, no. It's a good model. You know, finally found a way to use this GPUs out of it. Just reselling them. Yeah, no, this thing I think will actually get real use and another. I did enjoy using this a lot through the have you tried the Grok build CLI at all? No, not yet. It's good. It's honestly one of the best two ways I've seen. It's very, very clear that it is taking everything that they liked about Claude code and codex and merging them together. Like it feels I can see the inspiration there. A lot of the UI looks similar. A lot of the patterns work similar. It's workflow view looks similar. Like it is clearly inspired off of these, but it is polished in a way that the other ones just really aren't. It's fast. It's snappy. The notifications are good. The UI looks really good. It's a great harness and the model does well in it. I have a hot take on this. I think one of the reasons that things like cursor and also like the codex CLI and the like entire quad code experience are starting to get like bloated and rough to use is that a lot of those experiences were built before the models were this good. And they have this giant corpus of tech debt. The models can help with some amount, but at a point having the starting point be dumb, like a project that was started by a junior engineer is almost necessarily going to be less maintainable than one that was started by a much more experienced engineer. So companies like XAI are at a bit of an advantage here building rock code been rock build now because they don't have this old decrepit code base. They have to maintain and like fix and it's nice. Like I know I've had a lot more fun like deleting projects and rebuilding now than ever and getting like the starting point to a better spot is way easier than it's ever been. But if you have a shitload of users like millions upon millions that have specific integrations with certain plugins and things that's not going to be as maintainable or easy to get out of as somebody starting something from scratch. So, yep, they have a huge advantage on that. And also they can learn from everything that has been done before. Like they can see what is and isn't good from all of the existing two ways and just learn from it and not make the mistakes that they all went through the starting scratch thing I think is a huge advantage for them. And they're also just like XAI for all of their grokism sometimes. It is a remarkably good engineering org. Like their engineering is rock solid and that shows through all of this. I feel like that's probably going to be their play long term is just like grind it out on the engineering side. Like they're they're not researchers. They're not going to out research anthropic or open AI. That's not going to happen. Maybe XAI can teach the cursor guys how to have better more reliable software engineering and the cursor guys can teach XAI how to train models that don't suck. It really is a match made in heaven. Like it truly is because cursor their problem has always been their surfaces suck to use like their engineering discipline is trash. Yeah, they have terrible engineering discipline remarkably good like research and models XAI remarkably good engineering remarkably terrible models. It's a good fit. Clearly this acquisition is working. And as always, I have to rub in the fact that my favorite saying of it's better to be late than early is being proven out once again. There's a lot of you who built CLIs before cloud code that have been entirely forgotten. Rest in peace, Ader. Yeah, there was a lot of you trying to do that. Cloud code kicks are this new wave and the ones that matter now are not the ones that happened then. It's the ones that are happening now like this. Yep. Yeah, the all of the ones that I'm currently using it's I still use pi a lot and that is more recent than most of these. I'm actually using the grok one a decent amount. That's also a new one and then codex and cloud code have effectively been rebuilt multiple times with the zombified compatibility layers for their old plug-in systems and shit like that that they need for all the enterprises they have using it. Yeah, I've been digging probably too deep in the codex code base recently debating whether or not I have to do a slop fork. Dude, there's nests in there. You showed me some horrible stuff. We will what we're gonna have a long discussion about that later. I think I don't know. I don't think I have too much more on grok four five other than didn't Elon post that they're gonna be doing like monthly model drops at this point. That's their goal is to just be churning out models. Elon promises a lot of things. He promises a lot of things. Will that actually happen? I kind of doubt it. But if that's the direction they're trying to go, they're probably just gonna try and grind this out against the labs because that's the place where they win is they win by driving down prices constantly iterating and having really just good data good RL do it over and over and over again and try and beat down the other ones because I just I don't see them out researching them and like having something that can overtake them to be honest. Maybe you disagree. Maybe you disagree, but I really think we're at the point where if the exponential quote unquote happens the way the labs think it will especially open eye and anthropic they're just gonna keep hitting escape velocity and it's gonna be really hard to catch up with them. And researchers talk they talk a lot. A lot of the secrets that are allowing for the models to be as good as they are now are starting to get out there. I see a much clearer path to how they can get ahead, but I don't see it as like by being 5% better you win 100% of the money. I think that there's always gonna be like this neck and neck race between the options. It it's gonna look less like iPhone versus Android and more like the 15 PC manufacturers that are constantly fighting is my my guess is where we going. I agree. It's 100% the thing is just gonna be if we get to the point where they are neck and neck and you are now deciding between Grok 6 GPT 7 and Fable 7 and Fable 7 and GPT 7 are both very expensive big slower models, but Grok 6 is a lot faster and cheaper and it's comparable performance. They're just gonna win on the faster cheaper angle. I don't know how much strength they have there. I think open AI has the best angle for faster and cheaper now with their partnership with Cerebris like everyone else is kind of stuck on Nvidia. If that works I that is a huge if I am not like I want to see it. We haven't seen it. I am very skeptical that they can make that work maybe long term they can but in the short term every single Cerebris model I've used has been it's a remarkable demo. And when it works it's magical but like the Cerebris schizophrenia that happens to the models where you get a really smart 12 year old that is just spazzing out on meth. It's really bad like the models perform far far worse at that speed. Yeah, they do. I have heard enough rumors about like the weird architectural things they are doing and how they're gonna use like nine Cerebris chips to serve the model for one request and stuff that like it's gonna be interesting. I I trust they wouldn't have said they're putting soul on it if you weren't putting soul on it. They would have called it something else like they did before with spark. Yes, they are taking the actual model. We use the weights that we're used to and putting it there. It would be really good content if it fails. So I'm like straight up 50 50 which one I want to have happen. I it would be so cool if it did actually work, but uh, oh, it would be funny if it just spazzed out and destroyed everything. Yeah, the only lab that's even close to having multi architecture working well is still anthropic because they have their models working on Nvidia stuff because what the researchers use. That's what they own as well as working on Amazon Tranium through bedrock as well as on Google's TPUs through the Google vertex crap. Yep, so they know how to do this open AI does not know how to do multi architecture yet, but they seem to be trying to figure it out. And then Elon has no interest in doing this because he has a shitload of Nvidia GPUs. Yep, so it's interesting here is that Google should have been the player to be able to win price to performance and like using data to catch up. Google is too busy like sticking their own. I didn't yeah, I don't want to get canceled for the things I say about Google right now. But you both it's it's bad. It's pretty bad. Is your Gemini 3 5 Pro got delayed again? I did dude. It is so funny. I like I did my quarterly Google sucks lol post because like it's just it's tradition at this point. You got to do it and you know every time I've done it over the last I think I've done this four or five times now the first time I did it got a bunch of people being like bro. You don't understand they have all this data. They're the sleeping giant. You're just coping. They're going to crush everyone and you would just get a bunch of those comments and just slowly but surely every time I make the post I get less and less and less and less and at this point like I got like one or two but like their heart wasn't in it anymore. Like they're just they're beaten like they were just kind of sad like all of the no man trust Google's gonna get there. It just felt like it felt like a defeated man coping like it's just rest in peace. I am so excited in a few years probably five to ten years to figure out what the fuck actually went wrong. I don't think it's so much that like a thing went wrong rather nothing ever went right and you need something to go like novelly right to win right now. That's what happened with XAI. They had the novel great thing of acquiring cursor for an amount of money that seemed insane but it's clearly working now. Yeah, so when you consider the like weird pseudo aqua hire that Google did for windsurf contrast. They got the worst parts of windsurf like the windsurf team is killing it at cognition cognition is doing great. That seems like a very strong company right now. All of their products are excellent. They're improving rapidly. That was a good buy is their 1.7 model. They just put up. I forgot that also came out last week. Yeah, there's even more. But yeah, I did that model do better than Gemini models for actual code work. I haven't tried it yet. I have not tried it either. I haven't looked at the benchmarks. I I feel kind of bad because those models always get kind of forgotten about because I think that they're locked in windsurf jail like you can only use it there, which like cognition. I love you guys. Please free them from windsurf jail like cursor can get away with it because their cursor you guys can't like you have to put it on API and you have to make it exposed in other places. We're just like I'm not going to use windsurf. I hate to say it. It's just not going to happen. If you guys want it in T3 code, let me know we can figure out something that that would make me use it like I need to be able to use it in one of the things that I don't know. I don't know if I'm using one of these other surfaces right now. I just don't see myself using another VS code fork right now. That is the correct solution for a year ago. It is not the correct solution for now. Year and a half ago. Yeah, I guess that's fair. I think I was holding on to it longer than most worse. That's fair. I remember sitting you down and making you change your alias for cursor on your computer, have a pop up come up and make you think about whether you really want to commit to opening this or not. Yup. Yeah, you took away my cursor and it was a net positive. The start of the slow descent into madness I've been going through. Speaking of the things we love being taken from us, shall we pour one out for our codex? Oh, rest in peace codex. You were gone but not forgotten. For those who aren't aware, the codex app, which is a surface that I loved so dearly that I ended up building my own open source alternative, specifically because I was afraid that they might have more performance regressions or something could go wrong with it. Or I wouldn't be able to use it with other models from other labs. So I loved the codex app so much that I built my own with T3 code that is open source. Still very nice. Not trying to self plug. Just know that the reason I made this is I was scared of something like what just happened happening. So what happened exactly? Well, when I go to my computer and I try to open up my beloved codex app, it doesn't open up codex because the last time I pressed the update button in codex, it got overridden with the chat GPT desktop app. But Theo, there was already a chat GPT desktop app. Oh, do you mean chat GPT classic? Because that's what that got renamed to. And all of your history and threads from before in the actual chat GPT app are now folded into a weird pop up in the new version. Oh, it's terrible. The like chat pop up is so bad. I get why they have like the work codex delineation. The chat rest in peace. Like if you want to use chat and chat GPT, go to the website. I would not use it on desktop anymore. But if you're going to the website, you better not be going to it through Atlas. Oh, yeah. Yeah. Atlas is well and truly dead. I think they did definitely learn some stuff from it. And now it seems like Atlas has been folded into the chat GPT super app because I think that's what this is. This is the super app. This is the merging everything into one thing. And now it is codex. The codex mode in chat GPT. Kind of. It's more that. How do I put this? The codex harness was one of the most powerful things that opening. I had it allowed the models to do way more and they started folding the codex model behaviors into the GPT models directly because remember codex wasn't just the app or the CLI. It was also a model family. It was also an experimental model. They made back in 2022 for the original co-pilot. There have been like 18 different codexes throughout the history of the codex naming. In fairness, that's just opening eyes sucking at naming things. Most of these things should not have been named codex. Soul, Terra and Luna aren't much better. At least they kind of make sense. I agree. They're weird, but it makes more sense. Dropping all three at once is also mistake. They should have spaced it out so we can get used to the new names. This is a separate issue. I want to complain about codex for now because the codex brand was finally getting better. The problem that I have and I've sat down multiple people pretty high up at open AI and explain this to them. The problem with codex is branding is nobody knows what you mean. If I told you right now I'm using cloud code, you know what I mean? I'm using the cloud code CLI because nobody including the employees use the desktop app and cloud code for web is cloud code for web, which means that you're kind of stupid because no one uses that. It's a it's so bad. So cloud code means the CLI always does when I say I'm using codex. Am I using the codex model? Am I using it inside of cursor or using the codex CLI or maybe using the codex app? What am I using? Maybe I was still using the old copilot back in the day that would explain a lot of your guys takes on AI stuff recently, but yeah, it's it was always bad. But codex was starting to build loyalty both because the app was really good and the models were good, but also because they were really forming a team around it. Codex finally had become a team at open AI and was building more and more trust with the engineering community within themselves and getting more buy in from the company and almost feels like their reward is getting destroyed. I think this will be a fun one because I actually disagree with you on this one. I agree that it was getting way better. Like I think before this happened when someone said codex I generally meant thought the CLI and when someone said codex desktop like anytime I would be talking to someone they big oh yeah, I'm using the codex desktop app. We all knew what that meant. They got the branding fixed. Everyone knew what that was, but I think the thing opening eye is really trying to do here is they are trying to put codex in front of more people rather than just leaving it stuck in the developer world. It's a gamble. It is definitely a gamble here to take away that branding that has been worked very well on devs, but I don't think has worked all that well on the wider world like how I have a lot of very, very good engineer friends who I know who whenever I go back home back to Ohio and I'm talking to them about the stuff. Most of them have big like they might have heard of codex, but they've never tried it. They've never done anything with it. Most people are still just using cloud code like cursor was the big thing for a while and now it seems like the general populace is on cloud code. So they need some way to get it in more people's faces. I'm going to disagree on all of this in two particular ways. The first one is he said the codex desktop app is the codex desktop app. I don't agree. I had this pushback from one of the people I talked to at open AI and I said no because that's like you're claiming this is what you call it. But when I command tab over and I hover over it, it says codex, not codex app, not codex desktop. It just says codex. You guys named it codex, but it doesn't say that anymore because now it says chat GPT even with the codex icon. Yeah, what what does what is the cloud desktop app called? Is it cloud? Cloud? It's just called cloud. Yeah, but like it's the cloud desktop app. And like when I say cloud desktop app, that's what I mean. But when I say cloud code, I mean the CLI. Yeah, that's what everyone knows that. When I say codex, what do I mean? It is definitely more ambiguous. I personally and everyone I know has been very clearly made that distinction. Like I everyone whenever I talked to I was like, yeah, I'm using the codex desktop app. I had a bad experience at a conference recently where I told somebody they should use codex. And they hit me up a week or two later saying, yo, I've been using GPT five three codex in cursor. Is there a newer model I should be using? Oh, I was compiling a cursor. But yeah, I told them they really should be trying out codex more because it's gotten really good. And they're using five three codex in GitHub copilot. Yes. I mean, at some point you like not everyone can be saved. Like the you realize that this is what they're trying to do with this, though. It's like they're trying to like help the lowest common denominator understand these things because the branding sucked. And they're doing that by making it even more confusing. Now, if you want to install code, there is one aside here that I like. When I would Google to install the codex CLI, it would always bring me to the codex app. Even on a Linux bot where I can't install it. And like hunting through the docs to find the codex CLI was annoying. Now, when you search how to download codex, only the CLI comes up. Yes. Which means all my old videos about going to install codex are now irrelevant. But they did make one change. I don't know if you saw this. It no longer says chat GPT work in chat GPT codex. They shorten that to codex. So when I'm in the chat GPT app, it says codex. Yeah, it does. This is a fumble. I'm dying on this hill. This will be remembered as one of the bigger mistakes they made. Yeah, it's a gamble. I could see it going either way. I don't know. Well, if you're committing, you got to commit. You got to do it the Nintendo way. You put out the Nintendo DS effectively sunsetting one of the greatest brands of all time with Game Boy. You lie to your fans and say, don't worry. The Game Boy line is not going anywhere. We're just trying this other thing with the DS. And then DS is exactly well enough that you kill Game Boy quietly and no one notices. That's what they should have done here. They didn't have the balls and they're trying to live in between the two in a way that's just cringe. Okay, so if you go to chat GPT.com, there is nowhere that it says codex. Like chat GPT.com, there is no mention of codex anywhere. There is chat and work now. So like you can go to the work tab, but that is not the codex tab. Like there is an open desktop app. So I guess the funnel here would be you go from chat GPT.com to work to open desktop app to go see the thing and do the thing in there. And then maybe you find codex, but if they are not funneling you in from there, then yeah, I guess. Yeah, I don't know either how to get to the codex page, which means I don't know how to get to our usage anymore either. Oh, I just figured it out. That's really funny. When you go to settings, the usage option appears latently in the UI because they had a feature flag that populates that. Oh, no. They are stuffing everything into one place. And now it's not even a URL. Now it is just like a bad overlay. Oh, they're tan stack routering. I can see it. This pattern. Oh, no. Actually, I like tan stack router. I don't like this particular pattern from it. And you can. I don't know well enough to know the pattern. Claude has a problem where they don't honor the appended URL. So when you refresh the usage page, it doesn't have the usage come back up. But here it does still. So they did do that part right at least. But this is a whole new UI. They finally aren't showing me the 5.3 Spark usage on this one. But this is sad because it's chatgbt.com slash codex is now hidden. You have to like go hunt for it. I knew this one would happen because Mark was complaining about how hard it was to get to Codex Cloud on the chatgbt app. It used to be easy. And they hit it many layers deep when they added their remote control stuff. And Mark uses it a lot for filing random PRs on the go. Yeah. Because this is the sub that he has. He can't do that anymore. Okay. That is bad. The fact that on chatgbt.com there's a chat tab and a work tab. Totally fine. There should now be a Codex tab. Like if they're going to lean into this and they're going to fold them all in there. If you want like because the thing that they cite is like wow there are billions of people using chatgbt or whatever. And a very tiny percentage of them have ever tried Codex and they want more people to try Codex. I don't think that just turning the desktop app into Codex is going to do enough because I don't know how many people are using the desktop app. I know a bunch of people who use chatgbt.com who are more normie types who would benefit from the crazy computer use stuff and using it for more complicated tasks. They're not going to find it on the desktop app because they're not using the desktop app. They need to get it from the website and the website does not show it anywhere. Don't worry though when you click the work tab on the website it has the big open desktop app button in the corner that doesn't give you any reason or incentive to do it. Yeah exactly and it's hidden and yeah no one would ever click that. That sucks. I also think just like I'm staring at this screen. This is not good for users at all. Previously the chatgbt screen was quite simple. They had the model tucked away in the corner in a way that it wasn't particularly suspicious and it would just say like latest or something. Now it's in the chat box and it says 5.6 soul and then in a lighter font light. No one other than us knows what that means. Yeah absolutely. Below it there's this weird bar with choose project. Odd amount of space plugins. Weird amount of space plus github. Giant space open desktop app. A bunch of random things here with no additional context. There's way too much going on here. So if this is their attempt to like get a hold of normies they need to stop letting slop design their UI. No yeah this does not make any sense to anyone. Like this. How many fucking plus buttons are there in the UI right now? Okay that one's a lot of pluses a new chat. But there's the plus there. The plus there. The new chat there. And when you hover there's a new chat there. This isn't good. No. Okay that's all one button. And it's a drop down. It's not even a button. Well and what's weird about this too is this just like this looks like codex minus. Like this is codex with things taken away. I essentially clicked one of these things at the bottom. Uh huh. What do you think it does? I saw so I spoiled it. But good god. Yeah there's suggestions at the bottom of the like chat box. And when you click them it just populates the chat and then disappears them from the list underneath. But leaves the other ones there still. So you can get rid of them all by clicking them once. And then emptying them out of your chat box. Cool. Oh wait what does it reveal then? It's just a create image writer edit or look something up in the work tab. I'm creating images for work. And you can do all of this within the chat tab. Like what is the fundamental difference between work and chat? Oh it's the exact same as the menu in the chat app. Yeah. And like in the chat like the big difference I'm seeing. I guess you have a little more control over the model in the work tab. And I guess the plugins are more readily shown to you. So I guess that makes sense. But still why? This is a slop. This was a bad move. Yeah. Decent reasoning for why they wanted to do it. Implementation that I wouldn't put in front of anybody. I think this was. And I've noticed a few of these things. Anthropic is decent at copying open AI. Open AI is horrible at copying Anthropic. Oh my god. Yeah when we. I feel like we should crash out about the codex stuff when we get into the limits on 5.6 later. But yeah. No there are a lot of issues with codex that are not good. Especially with subagents. But the codex CLI was good. And that copied cloud code. No it didn't. The codex CLI was in progress before cloud code even was announced. One of my friends started it as an internal project. Because he was just so frustrated. And wanted a way to give the model access to his terminal. Codex CLI is its own genuine organic thing. That like they did out of their own interest. That cloud code just happened to come out before. They built it before then. So that's not them copying. A lot of things since have been. And those have not been great. Like the work mode. Like a certain U word we'll be talking about later. Yeah. I love you open AI. You know that. I have been called an open AI blind shill so many times. Yup. That's why I have to be harsher. We've had a lot of talks. We'll continue to have those talks. We're going to have a long one at the end of this. Yes we will. Well I guess before we get into the big model. We should talk about another open AI thing. With Apple suing open AI. I ended up looking into this more. And it's crazier than I thought. I'll ask you some questions. First I will give a quick overview to the audience. Of what the hell is happening. Open AI is building some hardware devices. We don't know what they are yet. But they are doing something that is hardware. That works with chat GPT. They've been hiring up a hardware team. With a handful of engineers from Apple. Apple is now accusing open AI. Of having an employee. That was continuing to access data from Apple. To use that in order to better open AI's position. In building hardware. And the lawsuit is kind of crazy. And they're doing a thing that Apple I've noticed. Is recently doing in lawsuits. Which is dropping names of ex-employees a lot. Like they put the names in over and over again. The main name that was coming up in this one. Was Tan Tang. Have you read who he is yet? No. Tan Tang's recent open AI person. Part of this like surge of ex-Apple people. Going to open AI. Traditionally when Apple is suing like this. And dropping employee names. It's like random employees that have access to things. And whatnot. How high up would you guess Tang was at Apple? Like probably mid-level. He was the hardware chief exec. Oh. He was working under Ternus directly. The new CEO. Oh. Oh. Oh. Oh. Okay. And opening AI poached him. And when they poached him. He brought 40 people with him. He was working on an entire team. That was working on an entire product. And he was bragging to other co-workers at Apple. That he still had access. Because of some broken shared folder stuff. And has been proven to have accessed it many times. Both on his way out of Apple. And since when he was at OpenAI. Why are tech guys so horrible at white collar crime? Good lord. Like if you're going to do IP thefts. Don't tell everyone about it. Yeah. Good god. I assume this is some random like lower level eng that just made some dumb mistakes when I first read the lawsuit. Because I was reading blurbs. Then I read the whole thing. Saw who he was. Read his history. It's like oh. Yeah. This is like when it turns his right hand guys. Betraying Apple. Jesus. Yeah. It's rare I side with Apple in lawsuits. I'm not going to blindly side with them until way more information is out here. But Tang Tan fucked up here. Yeah. No. That is atrocious. Like if that is true. Obviously I haven't read through it. I don't know the full details. And I'm sure there's a lot more that will come out in discovery. But if they are. Is it public what team he was on and what that team was building at Apple? Yeah. Because he was the hardware chief at OpenAI. Or hardware chief at Apple. He was leading all hardware. So all of hardware. Okay. Okay. So clearly there's a lot of trade secrets that OpenAI is going to really really want to have as this new device comes out. Yeah. He was the one who started IO, the weird company with. Joni Ive? Yeah. Joni Ive. If I recall him. Oh. Yep. Confirmed he was working with IO. Allegedly kept an Apple laptop and downloaded confidential files on unreleased products. Uh oh. The exec who was in charge of Apple's smart glasses work was also poached throughout all this. Yeah. Apple is real sensitive to poaching as well because like the cult nature of Apple. So they were already going to be on top of this. And then when they saw this exec asking things he shouldn't have, they went all in. Some amount of this is Apple trying to make employees scared of leaving. Yeah. But this one does seem more fucked up than I initially would have guessed. It's strange too because I know like culturally for a very long time, Apple was very much like you would go there and then you would spend your whole career there type company. Like they are not used to having people leave like this. And it seems like at least according to this over 400 former Apple employees are now at OpenAI. They're clearly poaching a ton. And I'm sure it's a different team. But clearly there is ex-Apple people involved with all of the crazy stuff they've been doing with like Codex Desktop. Do you not know like who that is and why that's so good? I do. Can we say it? Yeah, it's public information. It's been for a while. Yeah, the guy who made shortcuts at Apple is working on Codex. Yep. Yeah. He was a different company that Apple acquired to automate macOS. He didn't get to do as much as he wanted there. OpenAI reached out and now he's doing it at OpenAI instead. Yeah, because they are doing crazy system stuff. They are using and abusing macOS in ways that I don't think that they ever really intended for it to be used and abused. So yeah, it's the biggest reason to use Macs right now if you're not a creative using like graphic software like you and I are. One of the biggest reasons to stay on Mac and not make the move to Linux is that the Codex experience. Sorry, the ChatGPT work and Codex sub button experience. Yeah. Happen to be really good on Macs because that's what they are focusing on. And if you have a bad OS that's hard to automate, but you have some of the best engineers in the world trying to automate it and you compare that to a good operating system that's easy to automate but has nobody really working on it. Yep. The bad one's going to win. Yeah, absolutely. And it's also just like everyone generally speaking is on Apple, like especially if you're a developer, you're probably on a Mac. That's just what we're all using. So they're going to target that first. I don't love that, but you're not wrong. This is this is the best hardware. This is the best laptop out there right now. I will die on that hill. This is the best laptop hardware on the market by far, and it's the one that I want to be using every day. But the place where I want to be running most of my stuff is Linux promoting it. But not this episode. Another episode. The Linux episode is coming sooner than you guys think. As soon as we have an even slightly boring news week, we'll just be ranting about Linux the whole day. I was hoping it would be this week. But since we have open source this week, it's probably going to be a huge news week. Yeah, we don't get breaks in nowadays. It is against the law. I went on vacation, quote unquote, earlier this week, right as 5.6 came out. So you're welcome. I took the bullet on this one for once instead of just me. Yep. It seems that the curse is spread. My favorite tweet I've seen so far about the Apple opening. I think was somebody that was kids tagging me and saying that he was going to do a four hour long video about this one. So I'm going to prove him wrong by doing an under one hour section in the podcast. Good work. We can wrap the lawsuit itself now to prove him even more wrong. But we have to transition to the funniest outcome of it, which is the Elon Sam Altman beef ramping back up, which is particularly funny because Elon has been much nicer to his competition recently as he's selling GPUs to them. He's been saying really nice things about Anthropic even lately. Oh, he's been so nice to them. And I can feel the pain in his tone when he talks about how he considers Grok 4.5 to be almost as good as Opus 4.8, which is even the best model for Anthropic. Which is a weird admission from him. Like usually he does not say stuff like that. I've never seen him be like, yeah, this is a good product, but be that realistic and chill with it. Like, no, it's not the true frontier, but it's good. And it is now you can use it. That's why I don't believe him that he's going to be putting out models every month, by the way, because he doesn't like talking that way. He wants to be number one in the world. He's not going to admit that he isn't again. Yep. And as such, he will never admit that he is bad at communication. And that's why he is going ham replying to Doge designers news post about the new Open AI Apple lawsuit, which, by the way, those photos are so AI generated. It hurts. Yeah. But Elon replied with scam Altman strikes again. Oh, boy. He loves his puns. Like, fuck. Do I have to rethink my pun stance? Yes, but for different reasons. Yeah. I am sad because Elon and I have been on good terms recently. He even signal boosted my review of Grok 4.5. I forgot about that. Yeah. Yeah. He ripped your video and reposted the entire thing. He did the built-in feature where you can post someone else's video, so I still got the views for it. Oh, cool. Yeah. He did it right. Yeah. Credit words, too. He'll admit when people he doesn't like say things that he does like. Yeah. It was a good review. I was thorough as fuck with that model. Like, I had to delay testing for things I liked, and I basically stopped using Fable for the day just so I could go all in on Grok 4.5, and it was good. Yeah. But, yeah, going after Sam Altman this way, just, it feels petty, Elon. Come on, man. Well, and he did it twice. He replied with scam Altman strikes again, and then he also quote tweeted that same tweet with, he takes scamming to a whole new level. Or, no, actually, that was a different tweet. It was Doge designer taking the reply and making a fancy AI-generated image with it. Fucking Doge designer. Please stop being cringe. I am begging you for five seconds. Just XAI. I want to like you guys a lot. I do like you guys a lot. I love the team, but God, please stop being cringe. Then he followed up with the image of Sam saying, I'm doing this because I love it, which love and hate Sam Altman. He still doesn't have equity in open AI as far as it is publicly known. I think of that change, that would be very, very bad for him. You'd have to admit it because of all the things he said. I think he is actually in this for the love of the game. He has enough money to never have to worry about it again. He's just, this is how Sam is. Elon quote tweeted with this image saying, by this, he means scamming. So now we are three plus in. And then he replied to himself with, he might literally love scamming more than any human alive with an exclamation point. Good God, that's four. That is four within. Oh, we have a timeline here. We have at 10 p.m. on July 11th and 1030 on July 11th and then midnight on July or not midnight around 1 p.m. on July 11th. Those timestamps are wrong. I don't know where they came from. It was 313 a.m. Our time when he sent the first one. And then it was 337 a.m. When he did the quote tweet with the scamming to a whole new level. Then it was 552 a.m. When he did the next image and 553 a.m. for the one after those timestamps make way more sense for these posts. So he was up way too late crashing out about this credit where it's due. I've been up late crashing out, but I avoid doing it publicly. Yeah, I have a whole team of people that are there to help stop me from doing this. Yep. Sam quote tweeted him finally after all of this at 930 this morning with homeboy. You're the one selling public market investors on short term space data centers. He went after the space data centers. No, you want to buy it? We'll start flying them next year. Maybe you can come see them if your parole officer. Mark approves. Okay, that's a banger. That's a banger. After stealing an open source AI charity, you then stole all of Apple's phone technology. Wow. What do you plan for an encore? That's going to be tough to beat. God, I wish over the day I was doing a phone. I will put money on it that they're not. No, I don't think they are. I wish they would do a phone. Oh, dude. Dude. Then Sam starts posting through it because of course he does with a there's a lot of benchmarks that suggest 5.6 soul is the best model in the world right now. But the most reliable way to tell is that Elon is obsessed with me again. Yeah. He does seem to really, really hate open AI to the point where he will be insanely nice to Anthropic just to give them the finger. And there's this wonderful thread here where I like Tesla's posted. Sam Altman wasn't afraid of Elon, but he is terrified of Apple. You can tell by all his posting today. Sam replied, I'm not afraid of Apple, but I have tremendous respect for them. S tier company, which is weird because he used to say the same about Elon and his like back and forth. I have so much respect for Elon. It's so disappointing. He's acting this way. He no longer acts that way. That is true. Yeah. Yeah. Something flipped here. What flipped was the lawsuit. Yeah. Once that went as far as it did. Yeah. It's over and Elon got nothing out of it. So. Yeah. So he can say whatever he wants now. Yep. Nikita replied to Sam with incredible trade secrets as well. Some of the best. Nikita can post pretty well when he feels like it. Elon did a laugh emoji to that. I think that's the end so far. Yeah. What do you think about space data centers? I prefer to not think about space data centers. China's doing underwater ones now, which makes a lot more sense. I guess kind of water cooling with infinite water. Does saltwater cause issues though? Like I see you would probably, you have to have a filtration filtration system, right? Depends on what water you're letting touch and not touch. If you have a like layer between the two, like similar to how like you're not letting water touch your CPU. Yeah. You have a block between them. Okay. Yeah. If you have that. Another thing like that. Yeah. If you have something that's resistant to it, then yeah, I can see it at least theoretically. Not that I know anything about underwater data centers, not my area of expertise. Depending on how deep you put them, it seems more valuable than the space stuff. Space is that. That one. I have so many issues with space data centers. I also, I am a person who is very concerned about space junk and making it impossible for us to get out of our own planet. I think that's likely one of the failure modes that Fermi paradoxes us. It's certainly possible. We already have problems with this. Like there's a, if you don't time your launch rate, you're going to get hit with some random space debris and your thing's going to be dropped. Because even a speck of like sand is worse than like getting a bullet shot through you. Yeah, of course. Because you're moving so damn fast. Yep. Yep. Of course. Yeah, that, that'll be a problem. I, and even if like, even if magically that wasn't an issue and you could just shoot them off, there's still just the reality of like, okay, if I am doing a codex session and I am streaming that down from a space data center, there's, I don't know how the latency is going to be on that. They're going to use it more for training, I think. For training, it would make sense. Or for like gigalong running stuff. Like if you basically packaged up a job and it was Linux boxes plus, plus, plus, plus type thing. There are ways they can do it. I just don't think that's the direction they're going to go. This is more like put a bunch of GPUs somewhere where they'll have more access to electricity and less geopolitical problems. Yeah. I honestly think it's more politics thing than anything. Oh, absolutely. Because once it's in space, you can't regulate it. Yeah, that's the thing. Because like land, like we have enough space to put the data centers down here, at least for a very long time. But there's just politically, that's going to get to be harder and harder. And in the US right now, the politics around like getting more electricity is really bad too. Like no one's going to stop you from collecting infinite solar once you've left the fucking solar center. Like once you've left our orbit. No one cares. No one can do anything. Like the only ones who could do anything is like the US and the US isn't going to stop them. It's not going to happen. Yeah. And the risk of shooting it down again with space debris, my favorite thing to be paranoid about. It'd be really funny if all of Elon's efforts actually end up being the thing that doom us to stuck on Earth forever. It's possible. But also like if you get to the point where you can just do interstellar travel, you probably have the God machine at that point. And the God machine can probably just laser it down somehow. Like problems can be solved. I had to go look at my numbers to be sure about this. As few as 10 satellite collisions would be enough to meaningfully impact our confidence in doing launches. And once we hit 100, it will likely cause a domino effect that effectively destroys all of our satellites and prevents us from being able to launch going forward. Like it doesn't take much. Yeah, that's a fun, delicate balance. When planes crash, they come back down. When satellites crash, they break into lots of pieces out there. And any one tiny piece hitting another satellite. Yeah, because it's an absurd number of pieces and it is so damn hard to model and track where all of these pieces are going to go. Like I know that this is a real field of like science and defense technology effectively where they're just trying to map what the fuck is out in space. And it's really hard to do. Borderline impossible. Yeah. I will continue to try and not be paranoid about us Kessler-ing ourselves. Data centers in the ocean. It's basically boiling the ocean. No, it's not. You can't boil the ocean like that. It's too big. If you ran enough G-Stack instances, I think that could boil the ocean. But that's the prerequisite. How many data centers are we going to put in the ocean? However many it takes to make G-Brain universal. Yeah. However many it takes. Whatever it takes. I will boil the damn Atlantic to get enough G-Stack in this world. We're all connected. It's just one body of water. Yeah. But like it's funnier to just say the Atlantic. And like Gary could boil the ocean. I believe in him. He could do it. He could do it. We need to have him on as a guest. We've gone too far. It's such a good meme. He's not a meme. He's well, that's why he's my favorite. He is the most useful meme I've ever seen because it's so funny and it's so correct at the same exact time. It's glorious. Everything about it is just beautiful. He is the vessel by which Claude has come into our world. And it's magical to see. He's definitely had more of an impact on Anthropics income than most individuals that don't work there. Think he'll end up on the Claude code team someday? Orange to orange. Oh, God. I hope so. That would save Claude code. I would cancel my codex sub if they brought in Gary. It would be a really good build-in plugin. Imagine G-Brain comes with Claude code. It already comes with Conductor. Well, no. G-Stack is in Conductor. Okay. G-Brain is a different project. Don't cross your Gs here, my friend. So the real reason Ben is saying this, by the way, is because he tried out and explored G-Stack out of interest, but he genuinely uses G-Brain. I do. It's great cold storage. It is, like, genuinely the best open source memory management for taking a bunch of crap, saving it to cold storage, and then searching it. This is your fault, Alyssa. You brought up the ocean. Yeah. You brought up boiling it. Yeah. We didn't even make the connection. You did this. We can't trust him with the G. Yeah. You said it. One day, y'all are going to accidentally say G-Spot, and I'm not going to let Jeff edit it out. You think he would ever accidentally say G-Spot? Are you looking at this, man? Unless. Not only can he not find it, he can't pronounce it. Unless Gary makes a project called, well, actually, I can't speak that name until Gary makes one. You said we weren't unhinged enough earlier, Alyssa. Yeah. Are you feeling good about your phrasing earlier? You've done this to yourself. On a scale of 1 to 10, we're maybe at a 3.5-4. Is this a challenge? Yeah. Do not challenge us. I may not boil the ocean, but I could absolutely boil this podcast. The good news is WorkoS and Composio are chill, so I think we're okay. Thank you to them. Yeah. We are about to talk about saving a ton of money with the new 5-6 Soul model, though, because we've seen a lot of people burning way more uses than they mean to. I've done even more, as I talked about in the previous episode, which, by the way, still relevant even if you are watching this one now or even if you're watching this one later. The 5-6 episode we did was really fun because we filmed it before we had any real info beyond just using it ourselves. We didn't even know there would be two other models. We just had the one Soul model we were using. I think it's really cool to watch in retrospect now that we have all this additional info and other people using it, especially with how we compared it with Fable, with the limited Fable we had back then. Yeah. That all said, I've done, like, 200K plus tokens through this thing. I know how to make it burn. I know how to make it stop. And I know how to get good, useful, like, outputs from it without having to waste all of your money and sanity in the process. So if you're the type that's been hitting those usage limits or is scared of it, you definitely want to stick around. But as you can guess, this cost me a lot of money to figure out, so I hope you can pardon me for a real quick sponsor break. You've probably already heard of today's sponsor, WorkOS. They are the best way to authenticate and onboard enterprise customers onto your product. But that's not what I want to talk about. I want to talk about their pipe system. This is a really cool feature that I don't think that they talk about nearly enough. It's a system for letting end users authenticate whatever random service they need to, like Airtable, Dropbox, Jira, Zoom, MailChimp, Notion, GitHub, Gmail, whatever you need, into your app. It's incredibly simple to set up. You just mount the WorkOS widgets component into your app, and then the users can sign in there. Then on the back end, you call the get access token function, which will allow you to then call the APIs of these third-party services. And all of this stuff is stored within WorkOS's vaults. Their vault product is a great way to store and manage secrets, and it's a perfect fit for this. Integrations have always been very useful, but especially now that agents have become such a big thing, it is so valuable to allow an agent to get access to Gmail or Notion or something like that. And this is one of the easiest ways to build that into your product in a safe, secure, reliable way. And that's just one of many things that they do. They have great user management, the best MCP auth solution I've seen from anybody. Their admin portal is the best way to onboard any enterprise into your product. Multifactor auth in so much more. This is the ultimate auth platform for going from starting a new project all the way up to IPO at nerdstype.link slash WorkOS. Time to save some money for our fans. Yes, sir. It is time to talk about the models. Is there anything? We haven't done an episode since they actually dropped, right? No, we haven't. I mean, we've both done videos on it, and I think basically everything we said about them before release pretty much still holds true. For reference, we were talking about in that previous episode what turned out to be 5-6 Sol, the big boy. The other two that we have not talked about are Terra and Luna. I have been really impressed with Luna specifically, and I think it has a lot of really good use cases. I like it a lot for kind of the way I'm thinking about small models at this point is any prompt that you're going to run repeatedly makes a lot of sense for a small model where it's effectively a function. Classifying something, generating a title, running like a workflow type thing that you make, like a markdown file that's a program. Smaller models make a lot of sense for that kind of thing. Like, one of the things that I'm currently using it for is in my Hermes agents, they have the permission checks that come up where you'll get that Discord notification of, hey, is this okay? Can I run this gigantic bash command? That makes absolutely no sense because it's GPT bash commands. What you can do is set up 5-6 Luna to be the reviewer that will give a thumb up or a thumb down on those and then just let it go on its own like an auto mode type thing. It's great for that. Have you heard how I frame this? So I don't think of Luna as a thing that you should call on tasks you repeat. I think of it as a thing that you should call programmatically without other models involved. Like a model shouldn't call a tool to trigger Luna. I think that you should have code that for reasons that you determine like you're checking permissions for things or you're generating title. Like if you build the model in that you're calling via API in your code, Luna is one of the best options ever at this point. And it's actually really good for that. Hard to agree. That's exactly what I was ending at. Yeah. Yeah. But I wouldn't use this inside of like Hermes agent as one of the steps in a Hermes process because those like either I'm going to have Hermes help me write code to execute things and those things executing might call Luna. But generally speaking, I'm just going to let it call itself and it's going to call itself with Sol. No, no, no. That's not what I'm talking about. This is not a sub agent that 5-6 like because I'm using 5-6 Sol as the agent. It is not calling that. This is a built-in function within Hermes. So like they have regex that go through and do like the permission getting stuff. But it's their code that is executing this. Correct. Yeah. Their code is triggering it. I am not triggering a sub agent with it. When I trigger sub agents, they are other 5-6 Sol instances. Yeah. Absolutely. Yep. They just dropped Sol and then a week later, they dropped Luna, this super interesting small model. And then they waited a bit and put out Terra when it was slightly better to like be more traditionally like Sonnet 5. They had a way to roll this out that I think could have made more sense. But now they dropped everything at once, which like credit to them. They're all good models. They are. It just causes a lot of confusion. And I think that there could be similar confusion with Sonnet 5 versus Opus 4-8 versus Fable 5. But since they came out on a cadence instead of all at once, people are asking less questions. Well, and it's also more, it's partially just what gets covered and what people actually try. Like I think Terra feels like the forgotten child of this release. Like I have spent very little time with it. No one I've talked to has spent a huge amount of time with it. I've seen basically no one talking about it online. I've seen discourse on Luna a lot less. But like the primary discourse here is just on Sol. And I think putting them all out at once just made Sol be the one that everyone cares about and talks about. And the others just kind of got forgotten. And a lot of the point of the new naming was to try and get them to not be forgotten. Because the Mini and Nano models, those names imply, they sound worse than something like a Sonnet does. Like Sonnet sounds a lot better than Mini. So I get why they renamed them this way. But unfortunately it seems like they're still the forgotten children. Yeah. I just think they changed too many things at once. And this is affecting them both like in the implementation details where some of these things just aren't working on they're supposed to. Yeah. It's affecting their branding where they're sunsetting the Codex brand piece by piece. And it's making the model names confusing. Like I screw up Terra and Luna all the time. Oh yeah. It's just it's too much at once where like we can keep up on this because this is all we fucking do. Like when we're done with this we're going to go sit at our desks and vibe code for another three plus hours probably. It's just how we are. Yeah I got Fable to burn before it goes away tomorrow. And even then it's tomorrow. Fuck. Yeah it's the 12th. God I really hope OpenAI bullies them into keeping Fable around a bit longer. I do too. Yeah. We are still Fable versus Sol here. I'm team Fable at this point. Oh I am entirely team Sol. I am fully fully open AI. We should do a dedicated Fable versus Sol at some point. I agree. I think that would be fun. That's what we're here to talk about now. We're here to talk about all the things OpenAI screwed up with the Codex, API, CLI, app server, app, whatever. And how these things are costing you almost all of your usage. Absurd amounts. I actually took the time to write an article about this because I'm so goddamn annoyed that I just felt like I had to vent a bit and like also give people the advice. I will do a video eventually on how to maximize your Sol usage but I didn't want to have people wait especially with how many videos I have in the queue right now. So I just dropped this as an article to do my best advice. Shall I just go through this or do you want to kick us off for it or push back on the things I say? As I told you earlier, I generally speaking agreed with all of it. What's the easiest place to start? I think we should start with a thing that has the least to talk about which is fast mode. I think fast mode is a pretty simple one where... You should turn it off. Yeah, turn it off. Like I've been using it for everything for quite a while and especially as they were more in the foreground, it made a little bit more sense. It makes absolutely no sense right now. It is destroying your usage since turning it off. My usage has been lasting like 4x longer. I told Ben in particular that he needs to stop using it. I, on the first day we had access again. Before we had our access cut with the government stuff, we didn't really have it count towards our usage. We just kind of burned. Once we had it unlocked for us, suddenly it was actually affecting our usage. And I hit the 5 hour 20 minutes after we got the DM that we had access again. Because I was trying out Ultra for the first time and I tried it with fast mode on because it was just on. So I was using a 5.5 earlier that day. I burned that limit immediately with one prompt. I was like, what the actual shit? And since then have spent a lot more time digging through my logs, my usage, and all of these different things. In my conclusions, the easy one is fast mode off. Because the biggest risk with fast right now is that 5.5 would stop. When 5.5, for any reason, whether it's like not sure what the next step is, it wants your approval for the thing it's going to do, Mercury's in retrograde or exists, it loved to stop. It just wouldn't keep going. I had to encourage that model more than like anything since like the sonnet three days. That's the thing I hated the most about 5.5. Remember back when I said 5.6 is going to fix all my problems and I'm going to be very happy? Yeah, I know. You were correct. More so than I thought it would be, honestly. But yeah, because of this nature, even without fast mode, the way I phrase it in here is that 5.5, the furthest I could see it use my five-hour window with a single message was like up to 2% of that window. 5.6, since it can go so much longer, can use way more. We're talking like 10% to 15% on a single message, even without altering things. Just like going for as long as it can is going to inherently burn more tokens. This means that if you multiply that number by 2.5, which is the usage difference that you get when you use fast mode, a single message now went from previously where it could be up to like 5% because of that multiplier. With 5.6, we are now talking 40% of your five-hour window with a basic prompt where it does a little more work than you expect using fast mode. Yep. You can't use fast mode right now. No. And that's not even the Cerebris 750 TPS one. This is only 50% faster, which sounds like a lot, but the model is so efficient that it spends way less time reasoning and way, way more time waiting on tool calls. So you don't really feel the difference as much. Yeah. And you save most of your usage. Yep. So that's the first thing. Yeah. The only place I still have it on is like we were talking about this earlier. You have the little fancy alias. I set mine to be CSF, but it just is an alias of my ZSHRC that will spin up a codex session or whatever. Your ZSHRC. Shut up. That'll spin up the codex instance with 5.6 Sol on low reasoning in fast mode because that is a very, very useful setup for just using and controlling your computer. Like if you don't remember the three commands you have to run, CCF is mine. CSF works as well. I just type that, press enter, say vaguely what I wanted to do in the model smart enough and fast enough with that setting that it's slightly slower than writing the commands myself. Yep. I like that flow a lot. But generally, though, other than that, fast mode is off for all of my usage now. Yep. But you also just talked about low, which means we need to talk about effort levels a bit. Correct. First and foremost, ultra is not an effort level. And I am so tired of this pattern. We'll talk about ultra in a second. So forget about ultra for a sec. We need to talk about from low to max. You have low, medium, high, X high, and max. 5.6 Sol offers multiple reasoning levels. You have low, medium, high, X high, and max and ultra, which isn't a reasoning level. The first thing I want to say is that once the reasoning levels stop being words that make sense and they start appending letters in front or trying to do the Apple-style naming with max and ultra, you know you're in for a bad time. Generally, I recommend avoiding those across all of these frontier models. Those kind of exist for bench-maxing is the way I would put it. Something I talk a lot about in my personal life, not in content, probably to do that more, is the absolute chaos going on in the processor world for computers. The CPU I have in my desktop is a modern top-of-the-line Intel chip. Its official TDP is 125 watts, which means that their expectation from Intel is that it'll run at up to 125 watts roughly. But it does have built-in turbos and boosts that can go a little bit higher. My motherboard stock ran that chip up to 350 watts. That is scary for a bunch of reasons, but hypothetically speaking, should allow it to perform meaningfully better. It does, by like 5-8%, which allows this combo to score as high as possible in benchmarks, which is why the motherboard manufacturers and Intel are okay with just cranking these chips way further than they should be, even if it degrades the life of the chip. They do this because they want you to have the best number in a bench to look good against their competition, even if this is far from ideal for real-world usage. We are now at that point with models. X-high, max, and we'll talk about Ultra in a bit, are that. It is turning off the capability the model has to be incredibly efficient and smart in favor of letting it score slightly higher in benches. And when you look at the benchmarks, you can clearly see this between high and X-high, where high will be very reasonably priced per task and get really high scores. X-high will bump 1% to 2% higher at 2x the number of tokens. Max will be another 1% to 2% or even better, sometimes go down in effectiveness in one of these benchmarks at another 2x cost. So you're now four times more expensive than high for like a 4% improvement. That's not worth it, but they're there because they want the number one spot. So take that as you will. What I'm trying to say here is just use low, medium, and high. I have not found any tasks where high couldn't solve it and X-high could that were reasonable at all. I think X-high is cool, and when you have like crazy exploratory work where you really wanted to go investigate things, it's interesting. But think of it as more of an experimental thing rather than the mode you should be using. When I check my computer, it's on high right now. It tends to be what I use. I think this is the first time we are in like full agreement on a model reasoning effort thing because like I am fully high-pilled. High is great. Like my default for coding stuff is high. My default for dicking around using my computer is low. Those are the two that I generally speaking use. And then if I want subagents, I just tell high to make subagents, and it does a great job. Every once in a while, I'll break out like the max or the X-high if I'm doing some really weird stuff, especially with like the things I've been doing with TX-9 and Hermes agents, which is just like Docker containers to run them on my machine. Sometimes that requires enough weirdness, and I want it to be thorough enough that I'll give it the high reasoning. But generally speaking, it's just higher. I don't even think X-high and max are useful for those types of things that much. I just use high and haven't had any issues with it. It's been pretty solid. Another fun thing about the launch is the amount that our friends over at OpenCode have been shilling the hell out of Sol. It's just fully one-shot the entirety of the OpenCode company. And they're barely using Fable at all, and they're paying API prices, and they're picking Sol. Not just because it's cheaper, but they actually seem to like it a lot more. What was much, much funnier, though, is they were obviously using it through OpenCode, not through Codex. And they hit a problem I've hit before where reasoning levels aren't being passed through a layer properly to the official APIs. And because of this bug, they were using it on Medium for the entirety of their testing, comparing it against Fable on X-high and finding GPT 5-6-hole Medium to be better. On one hand, I will admit I see this as a bit of a level of incompetence for my friends over at OpenCode. I am so sorry, Ryan. I hope you learn how use-effect works someday. Okay. Ryan, I'll take the dunk on, but if Kit believes in the Medium, then I believe in it. I don't know how end-to-end pill they are. Like, do you know how bad the sub-agents are in OpenCode? Yeah, I've not tried OpenCode V2, so I don't really know what their current status is. It seems like they've made some good decisions with that, but I can't really comment on it yet. Are they still hard-coding the sub-agent types? I don't know. Cool. I will spin something up to figure this out for us. That's a good idea. Kit, who is at OpenCode, he is the legendary effect man who made the best videos on YouTube about effects. They are glorious. And he replied to one of my existential crash-outs earlier today about losing 5.6 with, I wouldn't wish losing 5.6 on my worst enemies. So, yes, they love the model. We love the model. And even on Medium Reasoning, it seems to do really damn well. I'm using your startup as a markdown file skill to explore the repo, and I will let you know momentarily whether or not they have sub-agents that function now. I have never pretended that's a startup. To their credit, Codex doesn't have sub-agents that function either. Nope. Yeah. Which, one of my favorite things we'll get to talk about in a bit, is that I am mostly using Sol now through Cloud Code. Yes. Oh, God. And it's blessed. Can I crash out about Ultra now? Let's do it. Okay. I don't think they should have shipped Ultra. There's a lot of layers to this one. I'm going to start with why it exists, because this is the core crash-out here. The only reason Ultra exists and has the shit UX it does, which I should probably actually start with what Ultra is, Ultra is a new reasoning level that isn't a reasoning level. It's in the effort selector in Codex, but Ultra isn't it thinking more. It is an addition to the system prompt that tells it, go spin up a shitload of sub-agents for this task. There are some problems, though. As I mentioned before, sub-agents are implemented questionably inside of Codex. There's two versions of this implementation. They have a V1 and a V2. The V1 is the default. It's what everything uses. The V2 is an opt-in flag that I opted into a while ago, because I wanted to see what it could do. There's a problem, though. V2 is not complete, and they keep making changes that have some meaningful regressions. V2 is opted in if you are using Sol or Terra. Luna still falls back to V1. They did this override in the models JSON, which is the file they have that describes all the models that the thing can use, and they added a way for a model definition to override your config and opt into agents V2. The sub-agents V2 implementation currently does not let you pass reasoning levels from the top-level agent defining the sub-agents. So if you tell the model, hey, I want you to go review all of these PRs, spin up sub-agents to break up the work, and you're using Ultra when it does that, every sub-agent is also Ultra. One of the many reasons why effort levels and these types of behaviors overlapping makes no fucking sense whatsoever, because even the model is picking Ultra. Not that it has a choice. There's a hidden config flag you can use to allow the main top-level agent to define sub-agents with effort level and model selected, but currently it doesn't do that by default. They might even have this fixed by the end of our recording session. I was talking to the team earlier. They are trying to, but goddamn, this is like a guaranteed token furnace, because Ultra is just a please, please, please use more sub-agents toggle that also forces you onto max effort. So you're getting two of the worst things for your burn. You're getting a reasoning level that is absurd and doesn't really benefit you, and a shitload of sub-agents spawned with that, which is why it burns so absurdly much. And the reason that this exists and also is in the effort selector is because Anthropic already made this mistake with Cloud Code. The Ultra Code mode is effectively the same thing. There are some meaningful benefits to Ultra Code, though, and I'll talk about the workflow side in a second. The big things I want to talk about are the actual thing it does. It also is effectively a system prompt append, where when you go to Ultra Code, it is telling the agents, okay, spin up sub-agents and orchestrate these things, in that case using a workflow, which we'll talk about in a second, but it also pins the effort to high, the right level for doing this type of thing. Ultra Code and Cloud Code pins it to high effort, not X, high or max, high, which is the right level to have all of these sub-agents orchestrated at, and it also has the ability to choose which models and sub-agent effort levels to use. Codex doesn't, and it pins it to max, so it's the worst of both. It pins it to a level that is too high, and it doesn't let the sub-agents get told what different model or levels to use. So the benefit of having these new models, Terra and Luna, being that they can be used for sub-agents, is literally not realized right now at all. In and of itself, enough to crash out for and make fun of them for. But we're just getting started here, because the methodology in which the sub-agents are called is even more fucking garbage. Would you like to do your best to explain how sub-agents are called in Codex? Cloudco calls them correctly. Codex does not. Codex is using the old system for calling sub-agents, which to my understanding is literally just, they have a tool call that is like, make sub-agent. They can just create a new sub-agent with a tool call and then send it off and have it do its thing. And that does not make a whole lot of sense at this point for the way good sub-agent workflows actually work. Do you know what v2 does different, though? I don't know the difference between v1 and v2. V2 allows for the sub-agents to effectively behave as threads with message passing between the two, and it also, by default, copies the entire message history over to the sub-agent, which I don't think is good at all. No. No, that's way too much context. It should just be like, if it wants to add more context in a longer message, it should just write that. You can have it send less of the messages when that happens. By hard-coding that in the config, you can hard-code the number of turns that are included. Brother. Really? It's a joke. So like, if you have a session that is like, you're on the 1 million, or I don't know how big the context window is, but say you have 200k tokens in the context window, you spent off six sub-agents, you just spun off six new threads with 200k context in there, and I assume there's no cash on those. There is cash on them. That's part of why they did it, is it reserves the cash. I think it's the stupidest thing ever. Eating a once-off read with way fewer tokens, considering how many turns these sub-agents do, one-time right to cash is more expensive than not doing it, especially now. It's another one of the things that makes it more expensive. Cash rates are finally billed, where previously they were free. Yep. The cash rate being billed is annoying. I get why they're doing it. It was about time. But that does mean that when these new sub-threads are spawned with different contexts, they have to cash rate once, just once, and then that doesn't get eaten at all going forward. It's not a big deal at all. People overreact. Even people I know that are doing crazy things like this, like I talked to Jared recently from the bun team now, obviously, at Anthropic. He had the same concerns about context in sub-agents, and I showed him the numbers, like rewriting this once is not a big deal, and I think he now agrees. Yep. Yeah, it makes total sense. It's an assumption that I see what people have because cash breaks. For those who aren't familiar with how the caching works, effectively, every additional token you have in your history is steering the model a certain way. It is configuring the parameters and the pointers to have the model be more likely to generate what you want next. But that's to do math and a bunch of weird matrix bullshit in order to get you to that point. So you could redo that every single time, or you can take where it is at a certain point in history, snapshot it, save it, and then when the next message comes in, restore that snapshot, which ends up being significantly cheaper and easier on the compute side, which is why they want this. It also makes it faster, which is nice too. Those cache reads are way cheaper than normal reads are, because you don't have to go through all of the context to regenerate the spot that you're at. And there's a catch. If you change something higher up in the history, everything from that point forward is nuked. Usually the whole thing is nuked. So if you had something innocent in your system prompt, like it is currently 1257 PM, and the next message is 1258, there goes your whole cache. And this has caused a lot of problems for certain sloppy harnesses that have things in the system prompt that seem smart and innocent and meaningful that aren't. For example, if it checks upstream to see if main has a new commit or not on it, it includes that recent commit. Someone pushing to main, when you're doing work, just nuked your cache. Getting these details right is not easy, which is why all the harness creators are really paranoid about their cache implementations and trying to get it right to the point where they're making really dumb decisions. Yep. Like I could crash out about this for like two hours. I'll do my best to not take all the time I want to here. It's bad. The message passing is stupid. It doesn't make sense in this era. And the fact that it has to define these subagents the way it does is just bad. No, it's terrible. Because the good way to do this, the way that CloudCode does it, and the reason why we have started using GPT-5-6-SOL in CloudCode is because when you define subagents, when CloudCode spins up subagents, it does it by writing out a crazy bastardized TypeScript function, I believe. Not by default. This is kind of a different thing. This is workflows, which is a method of doing subagents in CloudCode. You can also spin up traditional subagents, and it does do that a decent bit. But workflows are a different option. And UltraCode is a way to append to your system prompt, please, please, please use a workflow. Yeah. As you were saying about workflows, though. And workflows are, from my understanding, this gigantic TypeScript function it writes that defines all of the different calls that need to be made and the flow that needs to happen for each of these. So when you spawn it up, if you've seen it before, you end up with this UI that has three different bullet points in it. So maybe the first step would be explore, then plan, then implement, then review. You have those four steps. And then under each four steps, there are five agents queued within those. So then you are on the first step. Those five will run. Those all finish. You go down to the next step. They continue that workflow down, and it ends up creating a very powerful experience that is really, really good. That does not... You can't create that with tool calls. The tool call cannot create these deterministic, effectively, chains of agents that get spun up over and over again. Couple of corrections. First off, not TypeScript. Vanilla JS, very strictly. It will error hard if you try to use TypeScript syntax. Annoying, but not a big deal. Yeah. What is Claude if not a type checker? Yeah. Well... I have to make jokes about this. No comments. Second, the files are a lot smaller than you seem to think. Like this one's 190 lines for a pretty heavy workflow. The other interesting piece is since you're defining these stages, the model in a given stage can determine based on the result to pass or not pass another chunk of work to the next stage. Okay. So the stages are dynamic from the start where the initial model determines what are the payloads that we're going to be executing through this first stage, and then depending on what the outputs of those stages are because they all are required to do like... Or formatted outputs. Gotcha. So there's a schema that has to match on its output, and depending on the schema, it will programmatically decide if does this queue up into the next stage or not. Cool. So if it is like, let's say, you are doing this to review a bunch of PRs. It might determine that it will review a bunch of PRs, and the different states it's in at the end will determine whether or not it goes through the next stage. Gotcha. Maybe there's one stage that is close the PR if it's not in a good state, and it'll pass it to that stage, and it will close it. Maybe there's a stage prior that is analyze the effectiveness of this implementation and see if it is worth saving or replacing, and it will queue some to that stage instead. But based on the output of a given stage, it can programmatically determine which other subagents go into which other stages, which is one of the many parts of what makes this so powerful. I loved workflows as soon as I looked at the file for the first time and realized what it was doing. I've been saying for a while like one of the coolest parts of agentic coding isn't that we're going to replace all the code that we wrote with codes that agents wrote. We are going to do that, but that's an aside. What's much more exciting to me is we're going to be writing code to do things once. We're so used to code being a thing that's executed thousands upon thousands if not hundreds of thousands if millions of times. What if code was so cheap that we don't mind just throwing it away and writing a bunch of code that will be executed one time? This is a great example of that where code is being written for the explicit purpose of running once to orchestrate all of these subagents. It is really cool. It's really powerful. And I liked it so much that I found myself liking Claude Code more and I even went and did a video about all the things I like about Claude Code, which I never thought I would do. I'm like the anti-anthropic guy. We were incredibly positive about them and I think it was Opus 4.8 was the release for this one? Is that one? Yes. Yeah, that was a good time for Anthropic. That was a great release. I hated the way Ultra Code was marketed as a effort level though because not only has that plagued the minds of a lot of people using quad code, it's also now been plaguing the minds of our friends at Codex and now they've copied that specific really dumb thing. Hopefully you understand at this point in the rant that workflows and subagents have nothing to do with effort levels. They are a different, separate thing that does allow you to do longer running, harder work with your subagents and with your top level queries at the cost of insane token burn. That's not an effort level thing. That is a different way of using your agents and they put it in the wrong place in the UI and this is causing a shitload of confusion and the name Ultra also leads to even more of this confusion. Oh yeah. There is one more massive benefit of workflows and quad code that Codex just doesn't have though. UI. UI. No. Well, that's what I'm about. That's not the massive benefit I care about here. The massive benefit with workflows is they have an end. Ah. Since it is a code file that has stages, it eventually has to get to the end of it. It might take a while to get there but it will. With Ultra, you don't define the full workload up front. You spin up some subagents that are called as the work is being done rather than on completion of a task. One way to think of this is the way that the workflows work is an agent is spun up, it gets an input, it does some shit and then it has an output which is a JSON object that the top level code, not agent code, determines what to do with. With subagents in Codex, it determines a few steps in, I would like three subagents to confirm my work. So it spins up three subagents. All three of those might also decide they want to spin up subagents too. But it's happening dynamically throughout the workflow rather than being dynamically defined up front. So there's layers of static here. There's static in the sense that you can define like hard code subagents like this is my researcher subagent, it has this system prompt, it has these permissions, this is my implementation subagent, it uses this model, it has these permissions. That is the super static, super bad implementation that we are used to from a lot of things. If that's what you're used to as subagents, you're not using modern shit, I'm sorry, switch to a harness that does this right. The next level of static is what we're seeing with workflows I think is a really good balance where they're dynamic in the sense they're not hard coded ahead of time but they're static in the sense that once the model has decided at the top level how to go through this work and how to split it up, the amount of work that goes into each bucket might be different but it is defining the buckets and the styles of subagents on that, like in the moment on the fly. That is really cool because it keeps you bounded but also gives you the flexibility of it working for the specific flow you have rather than being forced to take your work and break it up into the buckets that already exist and then there is what Codex is doing which is what if every time a token was generated the model might decide that this is the token that's going to start the next subagent and that subagent can be called recursively. You can have five subagents spin off that it just decided to make but those five can each decide to spin off five from there and they can send messages to and from each other which makes it even more chaotic. They can just like yeah, they go insane. The UI for this is atrocious in everything except for the Codex desktop app. It's not perfect there but it's a lot better. In the 2i it is bizarre. There's no UI. There is no UI. It's bizarre and also like using this in T3 code it is wild watching like full responses just print out onto the screen over and over again. I see like four intermingling responses. We need to filter that better. Julius is working on his own orchestration that has fewer of these problems. More and more though I just want to expose what the actual hardasses are up for the Cloud Code version and I'm going to make it work with Codex as well. We'll see where we end up with there but that is a little bit on our end but we're not showing anything better or worse than what you're getting out of the terminal right now. That's the biggest issue is like there is no with Cloud Code it is quite nice being able to do slash workflows and see all of the workflows as they're running. It's great. It makes it so much easier to keep track of this stuff and you can like save the state of it or copy it into a new thread or even like go and inspect the thread as it's running. It's great stuff that is just not available in Codex even for the normal subagents that aren't workflows in Cloud Code you'll see a little thing pop up at the bottom that has a dot on it that you can like arrow down and view it. That one's a little broken though because if the workflows are long enough it pages off and you can't get to it so I don't do the down enter anymore. I did that for so long I didn't know slash workflows was a thing but once I found that command where I could go and actually look at the workflow my life got a lot better so don't do the down enter. I don't use that for workflows. For workflows I've always used slash workflows. I'm talking about like if you just have like a single subagent. Yes. A real subagent like subagent and workflow borderline different things. These terms are hard. Yeah. We need better phrasing. I do like workflows as a term. I think it is the right way to think of this and it makes sense. I wish that's what OpenAI chose to copy instead of the works part which is the effort level slider bit. It's weird because I understand why they kind of like why Clogcode probably wanted to go for that because it's a brand new term which is going to confuse a lot of people so they wanted to put it in something familiar. Codex probably like they should not have gotten baited by that and done the exact same thing in a worse way. It sucks. I think it is time at this point that like we need to make subagent and workflow real first class terms that mean things to people and show differently in the UI and are handled differently by the models. I want to do two things quick and then we're going to crash out about our Clogcode experimentation. Okay. First thing I want to do is summarize all of the ways that people can save money and save utilization so they can actually use a $200 a month plan and not have it run out constantly. Then I'm going to do a brief intermission where I talk to my friends at OpenAI with whoever else here is listening because I have some things that they need to hear and then we'll wrap up by talking about our chaos with using Clogcode with Sol which has been much more interesting than I ever thought it would be. I'm going to do a whole video about this. So first and foremost to summarize how to save your utilization effort levels between low and high X high and max are dangerous. Max is basically useless and ultra should not be touched at all right now. It might get fixed in the future. It's in a really rough spot right now. Just just avoid that fast mode off. Just trust me. It's not worth it. 50% faster for the generation side for 2.5 times higher utilization. Not worth it especially with how long these threads can go. Oh yeah. One other small thing that I forgot about that can be really useful. Make sure your prompts have a clear stopping point. Like instead of asking the model to explore this thing to see if you can add a feature tell it at the end of that sentence stop once you've explored and ask me questions. If you're asking it to end with the thing and also babysit PRs maybe tell it stop after you pass the first set of reviews so I can hop in and take a look. Give it a clear place to end or it will never end. This is how they RLed the model. They want it to go forever. The example I saw somebody post that I love is that Fable's a wise owl and Sol is a Rottweiler that will grab on and not let go of the task. Absolutely. Like it is and the GPT autism does kind of work here if you know how to wield it because you can tell it to do something and it will do that something. Fable's much better at if you don't tell it to do something it will naturally come to what is closer to the correct conclusion generally speaking. So with GPT you desperately have to do stuff like this. Give it an end point or it will run forever and I've watched it run forever. It sucks. I had to do a review earlier on a PR in ultra mode just to see how bad it would be. It burned 50% of my usage and it went for two hours and I think it ran 40 agents sub-agents during that time and that was just like top level sub-agents. I don't even know how many were happening inside of that. It's bad. And on that note one other thing is that it seems like Sol was RLed to just bring sub-agents in all the time. If you notice that it's doing that on places it shouldn't and you're burning more than you want I would recommend going into your global agents MD and putting a little line in there that says don't use sub-agents unless I ask you to or it's just one or two that do very specific things like guide it to not do that or you might end up burning more than you want to. I think that by default it is set so that if you are in any reasoning mode that is not ultra it will not bring in sub-agents at all. No it does. X High does bring them in pretty aggressively. Really? I've even seen High do it. Interesting. Okay I think when we were testing that was not the case but that is good to know. Yeah it's because they default to V2 now and V2 is much more aggressive. Okay gotcha. Yep it adjusts the system prop which I will crash out about the OpenAI system prop momentarily. I knew it was bad in codex. I did not know how bad it was until I read it recently. Nope. That cool. Yeah I think that's all I had for the pricing stuff. Yep. Which means it's time to do a quick letter to Tebow and my friends at OpenAI. Oh boy. Anyway we have a lot to talk about guys. I've already been communicating privately and I hate to do this publicly but there are some problems that we need to address. Generally speaking you are very bad at copying Anthropic. When Anthropic does a thing and people like it you copy the wrong parts and apply them in your systems in a way that not only don't honor the thing that made the Anthropic version good they actively hurt your users and you're losing a lot of trust as a result of this. The biggest thing you've now lost is this perception that Codex is effectively unlimited on the $200 plan which you agree pretty much everyone felt that way right? Oh absolutely. Yeah that's over now. You've fully lost that. You have lost the trust that you paid 200 bucks and can use the thing forever because you wanted to add a little toggle because Anthropic had it. Whatever the fuck led your brains to working that way you need to go to therapy and get it out. Like it's not going to work out long term. Your attempts to get ahead of Anthropic by copying them are putting you not just behind them they're putting you behind where you were before. It is making my life harder too as the person who has to talk to my audience about this. Normally you guys are good about these things. You don't just give me the good tools you give me the good resources I can use to explain it to my audience. I've had to make this shit up myself. I have become a better resource for users of Codex than anybody working at OpenAI and I hate that. That's not how this is supposed to work. You guys fucking screwed up this launch pad. Ultra is a disaster and I could have predicted this within the first five minutes of trying it. I don't know what's going on internally. At the very least you need to put some employees on the same rate limits that people have on the $200 plan so you can catch this shit ahead because it seems like you guys were caught off guard by the furnace that you built. Yep. On the note why is Fast Mode a toggle and Ultra is a slider? This was all very avoidable and on that note the current state of the sub-age implementation that you have in Codex neither are complete. You know neither are complete. That's why you have half of them used for one model and the new version being used by these two models of this weird hard-coded workaround. You gotta rethink the strategy here entirely. The way you guys are doing sub-agents is so piss poor that I'm having a significantly better experience hijacking Claude code for it which we'll talk about in a bit but god damn and I wanted to better understand this so I went and read through the system prompt in Codex and I think I'm the only person who has done that in a while because I talked to some people who were working on Codex and they had no idea just how bad it was. Let me read some lines from the front-end guidance section quick. The first subheading in front-end guidance is build with empathy. If working with an existing design or given a design framework in context you pay careful attention to existing conventions and ensure that what you build is consistent with the frameworks used yada yada that's all cool. You think deeply about the audience of what you are building and use that to decide what features to build and when designing layouts components visual styles on-screen text and interaction patterns using your application should feel rich and sophisticated rich and sophisticated that's a great phrasing for your UI design this is you make sure that front-end designs are tailored for the domain and subject matter of the application for example SAS CRM and other operational tools should feel quiet utilitarian and work focused rather than illustrative or editorial avoid oversized hero sections which clearly isn't working it still does it all the time decorative card heavy layouts in marketing style composition and instead prioritize dense but organized information note that this model's autistic so when you guys put this in your system prompt you just gave it all these instructions on what to not do if it's a SAS or a CRM but it now aggressively does those for everything else because it feels like you just gave it permission to use otherwise oh this makes so many things make so much more sense like we're not even in the fun section yet oh no go ahead the design instruction section is where I started spam texting my open AI friends with what the fuck why are you costing me money with this bullshit in for example you make sure to use icons and buttons for tools swatches for color segmented controls for modes toggles and check boxes for binary settings sliders slash steppers slash inputs three different things for numeric values menus for option sets tabs for views and text or icon plus text buttons only for clear commands parenthesis unless otherwise specified cards are kept at 8 pixel border radius or less unless the existing design system requires otherwise they are literally telling you to use cards here effectively you do not use rounded rectangular UI elements with text inside if you can use a familiar symbol or icon instead this is why it loves doing those icons without any indication that it's a button which I've seen so many of from codex you use lucid icons inside of buttons whenever one exists instead of a manually drawn SVG icon it's telling you to use a specific library for icons in the system prompt if there is a library enabled in text to describe the application's features functionality keyboard shortcuts styling visual elements or how to use the app you should not make a landing page unless absolutely required when asked for a site app gamer tool build the actual usable experience as the first screen not marketing or explanatory content it's full of garbage it tells you to use 3.js for 3D elements and make the primary 3D scene full blend or unframed you sorry for just writing about this but no this is so bad this is like sonnet 4 era prompting no this is literally like this is what you were supposed to do nine months ago and I remember as this to start pruning these out of your agent MD this reads like an agent MD that was written in early 2025 and has not been updated since the best part is that the section is half of the fucking system prompt it's half of it for all tasks so if you're working on rust and you're wondering why the model is referencing cards here you go is right here dude so sorry open AI if you're gonna burn my tokens you better audit what you're burning this is not a thing I should have had to be a 5 6 UI I have done a lot of UI stuff with this that looked closer to probably a cloud UI and you know where you generated that UI cloud code so the reason I went down this rabbit hole with the system prompt and codex is we'll talk about why I switch over to cloud code in a minute but I started using cloud code with 5 6 soul and I asked it to take some of the work I had done and make a nice looking HTML page for it like I normally didn't say make it nice looking I I make an HTML plan summarizing the stuff you I decided to go through the cloud code system prompt I read the whole thing it is cringe because it's anthropic doing prose it does not mention the words front end or UI a single fucking time the designs you get out of cloud code are the designs you get out of the model not out of cloud code and this confused the shit out of me because I knew that the designs I got out of codex were not good and the designs using the same model out of cloud code weren't great but they were a hell of a lot better so then I went and looked at this system prompt and realized that this isn't because cloud code system prompt has a bunch of good guidance for front end like I would have assumed it's because codex system prompt is full of terrible design guidance 100% this is like this is part of the reason why harnesses like pi have benched so well and are often so nice to use they're not polluted with all this extra crap and I honestly like I remember we were talking about way back when this felt like a cloud problem where the cloud code instance had way too much garbage in its system prompt and it was making the models worse but no it turns out codex was just as bad arguably worse I would argue worse like the cloud stuff as much as there is a lot of it seems to be pretty well thought through and they yeah it seems better yeah I'm at the point where I'm considering forking codex to make a better system prompt and fix the like do a slot fork of all the sub agent stuff and just see what I can do but in the interim I'm using cloud code and on my that I really liked the implementation of workflows in cloud code I wanted to see if 5.6 was capable of doing that because obviously they RLed these models to do sub agent work I didn't know how much that RL was tuned to the way that sub agents work in codex and wanted to see if its capability here and its knowledge so to speak here could be applied in other places and before did we talk about the podcast no I don't think we have I think we should probably talk about that the vibroxy stuff it's great the reason why I even bothered setting this up in the first place what it is it is a little open source app that you spin up and run on your machine that allows you to sign in with your codex or cloud code or a couple other ones I think they have anti gravity grok and kimmy of all things in there they ban aggressively and they ban your google account yeah I that it is safe and doing it correctly the reason why I'm routing my cloud code sub through this vibe proxy is because it allows you my auth literally didn't work in the codex desktop app until I turned it off so I am not using it for codex I didn't think I would need to use it for codex I thought it was fine that you can program into something like cloud code to have access to all of your stuff instead of just load balancing between my two cloud subs I also have the ability to put in a different model slug like GPT 5.6 soul and it works with that so what I ended up doing is I don't have worry about it because Tebow blessed my implementation I wish I was joking this is why we love open AI when I tweeted I was using cloud code with soul he replied pretty soon so fully blessed by him which implicitly also blesses CLI proxy API so you guys owe me anyways this please use a workflow when you do this with cloud it will usually do it if I give it a medium plus size task and I have ultra code on it will spin up a workflow and use that to break down the work and if I do a follow up after like hey can you summarize all this in HTML page so I so when I had it run a really big long workflow going through a ton of different things in a project that took like 40 plus minutes to run it had good results but I I surprised at how much I have been using this and how much I have been enjoying it because the outputs are genuinely really damn good it was able to get real work done and complete big tasks like auditing a giant code base finding what implementations exist for a given thing and what it would look like to rewrite it dealing with all my open PRs across multiple projects it can break this all down with the workflow and it does it really well and tastefully I've been impressed with how well it that escape has weird behaviors in here but if you just look at workflows here it's really nice I got five six soul things going off here doing its explorations doing its thing seems to be competent right now fuck it's good the binding I gave you had alter code on always you should go delete that or tell something to do it because just tell it to use a workflow it's better that's fair I understand why just the tone of it sucks interaction style you communicate respectfully focusing on the task at hand you always prioritize actionable guidance clearly stating assumptions environment prerequisites and next steps you avoid cheerleading motivational language artificial reassurance and general fluffiness you don't comment on user requests positively or negatively unless there is a reason for escalation like this is this just smells like there are so many things that you will just naturally end up having in there that are effectively worthless tokens that you're throwing in for no reason you don't need to tell it to be respectful and nice and talk to just getting to the point where the built in behavior is what you want and you need to tell them less you can be more goal oriented effectively just give it a vaguer thing that you so yeah it's this like I thought it was just a model thing but no it seems to just be this the prompt confuses persistence with permission continue until solved is useful but it should not mean assume every ambiguous conversation requires mutation oh that explains so much it contains impossible oh I forgot about this one I is in the prompt wow there's also a time never runs outline in there do not end while sessions are still running that probably makes some sense there there's too much repeated behavioral control the creature prohibition for goblins appears twice too many negative instructions compete for attention the identity is stale and unnecessarily specific yeah this is good feedback let's see how claude feels about it it's visibly a GBT 5 era artifact that's accumulated patch scar tissue and a fair chunk of it would be dead weight or actively counterproductive on newer models the front end sections are regression test suite not guidance the goblin clause harness references that may not survive a model swap yeah there's old tool called plumbing that doesn't actually exist anywhere that's still in the system prompt that's hilarious general over specification when I asked in codex with 5-6 soul it gave it a 6 out of 10 overall but 8 out of 10 values it's so bad and to be clear I really clear about this models do not write good prompts no they write okay subagent prompts given the right system and system prompts and tools and whatnot they cannot write a good top level prompt or a system prompt they don't know their own behavior well enough for that same thing with skills the best thing you can do is when you notice about doing something bad wrong or dumb ask it why and then make other changes it's similar to the classic users don't know what they skills and I was going through them like I wanted to release some of my random five six skills the closer I looked to the worse it was there is there are kernels of really useful stuff wrapped in mountains of garbage like you will get a really good code snippet which is the thing that said yes here's my consultation delete the whole thing I was very strict about that the reason I think we are good at this is we can have programmatic testing for it where we test different systems and different behaviors but it's only has to read those logs and we get a caucus of feedback like a judge panel and that panel should not just be codex models. But if we are not working at OpenAI, if we are working at OpenAI, we have to use OpenAI models for all of our work, obviously. You and I don't. We could have a caucus of Fable, Grok45, and Sol as feedback mechanisms testing different system prompts and get slightly better outputs in test results from all of that. It'd be interesting to see. Yeah, because to be clear, you can automate the testing of these things. You can automate the scaffolding of all this stuff. You don't have to handwrite all of this, but the actual words in the markdown file, you should bare minimum be taking what it generated, gutting it, and then writing out better sections within it. Do not just blindly generate these things. As nice as it is, it is so convenient to just use the built-in skill creator and let it go do the thing and not have to think about it. This is the place you can't be lazy with these. One other piece I can't help but notice is OpenAI and codex not supporting any programmatic modification of markdown files. It's not just a skill that is really, really bad. It sucks. This is one of the things I love about cloud code is a skill can have bash in it that executes on pulling in the skill in the system prompt and other things too. You can just tag a file in in your cloud MD and it will pull the content of that file into that, which is so useful. There's a bunch of things in this system prompt that I would argue shouldn't be there at all, but absolutely should be guarded under some catch. There's a whole section about using ripgrap, but if it doesn't exist, fall back. That section should just check with bash is ripgrap on the system. If no, drop it. Yeah. Yeah. You can get really advanced with skills. There's a lot to that spec. Well, I already got my allocated section on the thing earlier, so I won't say it, but you know what I'm thinking. Have you tried good old G stack with soul now that you can use it inside of cloud code? Oh, no, I haven't. That sounds incredibly entertaining. I'm going to do that later. We can wrap up now and you can do it. I don't have much left here. I have done my rant. I think cloud code has somehow went from my least favorite harness to my most. I know. Actually, yeah, this is the thing that has happened. I do still like Pi a lot and I've actually, I'm halfway through a bastardized sub-agent and workflow implementation in Pi, which I think could be very fun because then I get rid of all of this and I can just do it custom, but that'll be a lot of work. I'm not done with it yet. I think Boris beat us once again. Boris' task when he was starting cloud code was to build something not for the models and where they are now, where they could be going. And I think a lot of these features, a lot of these capabilities, a lot of this dynamic behavior, the models just didn't know how to use it before. They do now. Fable is a huge step up and finally, the things they added to cloud code make some sense because the models are smart enough to use them. Codex was built to patch the models as they existed at the time. The models are now smarter and all of those patches are holding it back. I was incredibly anti-sub-agent six months ago. I didn't like them. I didn't use them. They just were weird and would bloat things in dumb ways. Now I can't get enough of them. We have flipped. The models are in a different place. Cloud code was there for it. At least right now, it is the current best option. This has been a long one. Let's wrap up with some quick viewer questions. If you don't know, follow the NerdSnap account on Twitter. We put up an ask for questions every Saturday before we film. It's a good opportunity to piss off Ben. Okay. Speaking of which, Ben, I know you love tongue twisters. Want to start with Arian's question? Oh. Fine. I'll attempt this. I have such horrible enunciation. This is going to suck, but I will do my best. So we have this question. Here we go. If codex models are running in Claude, can Claude models run in Codex? Since codex models don't run well in Codex, so Claude models are worth trying over Codex models. Except Claude models don't need Codex because Claude models run well in Claude, which is why Codex models are running in Claude in the first place. So Codex models run in Claude because Codex isn't good enough for Codex, and Claude models run in Claude because Claude isn't good enough for Claude, meaning everything runs in Claude. Claude runs in Claude and Codex runs Claude Codex badly. So Claude runs Codex better than Codex runs Codex and nobody runs anything in Codex, not even Codex, which makes Codex a harness that's worse than Claude at running Codex models and Codex models run worse than Claude models at running in Claude. So the question is, run what exactly? And my answer is, I don't even know what I just said. I just said words. The real question is, if you run Codex models in Claude, how does it perform on the Convex bench? Because the Convex bench does a great job of filtering out the Codex models from the Claude models because the Claude models outperform on Convex compared to Codex. Codex seems to perform much worse in general for that type of task with or without guidance from the Convex team because the Convex skills bring a lot of Convex context into Codex, therefore making Codex better a bit, but Claude code does a better job with the context from Convex than Codex does. Holy shit, I one-shot that. That was very insightful. Thank you for that. I learned a lot from that sentence. I was listening to every word you said. Don't worry about it. Watch that one back later. It was actually intelligent. I'm sure it was lovely. The answer is, currently I like using the Codex models and auth in Claude code because it's a better harness right now. There's no reason to use Claude models in Codex because it is a bad harness. There is a catch though, which is that a lot of the Claude code system prompt is trying to get around Claudeisms, and the result of putting Codex models in Claude code there is their attempt to beat out Claudeisms ends up making the model a little more aggressive on certain things. It's easy to tune out, prompt around it, you're fine, but yeah. The next question was, what's the most token efficient orchestration setup that I would recommend, including other models and harnesses? I think we're going to disagree. I want to hear yours first. You first on this one for sure. You're not going to like my take. It is OpenAI codex sub using GPT-5-6 Sol low reasoning in Pi. My counter answer here is that if you're trying to be token efficient, you should not be trying to orchestrate. Period. Orchestration is the fastest way to burn tokens, so just light them on fire. Mm-hmm. I agree. That's why I said Pi, because Pi does not have subagents built in. It has four tools built in. Yeah, but then you're going to build your own orchestration layer or have it call itself. Well, I'm going to, but you don't have to. You don't have to do that. The question was about orchestration, Ben. Oh, well. What is the most token efficient orchestration setup? I guess that's fair. Yeah, orchestration setup then would actually probably be the same answer if you really want an efficient one because you can build it in such a way that you can smack it to be very, very specific. No offense to Nilesh here for the question, but do you sincerely think somebody asking this question is going to build a more efficient mechanism for orchestration? Fair point. Then the answer would probably have to be... It's definitely a GPT model, probably in Claude, not on UltraCode. UltraCode is just a trigger for workflows, though. It is a trigger for workflows, but if you have it in normal mode, you can tell it to use subagents but not workflows. That's fair. It's not as much orchestration as the way I think about it, but it makes sense. I would argue orchestration is literally just any loopy type task, which loops are another poorly defined term, which is just like you're calling a bunch of agents in a loop in subagent land type stuff. Like, do review the PR until it's green. Have one subagent review and then have another one check the GitHub PR bots and go in circles. I stand by my previous statement. Efficiency and orchestration are at odds. Agreed. We shouldn't be talking about what is the most efficient orchestration method. We should be talking about which ones are so inefficient they shouldn't be touched, and the answer to that is a resounding Ultra. Yes, absolutely. I think if you are very price sensitive and you are being like... If you're at a workplace where you have a very limited cash token budget, do not bother with orchestration. Just work on getting really good single-threaded workflows on a lower reasoning 5.6. Is Luna the new Haiku? This one's all you. Is it the new Haiku? I would argue it is closer to the new Sonnet, honestly. Like, Sonnet 5 is not a Sonnet model. I don't know what model that is. Are you no longer defending Sonnet 5? I still defend that it is like it's the demo for next generation models. It's like the early presentation beta build of like, this is what could be. Take a look at this. Isn't it neat? And then you never touch it again because it sucks. It is not what I remember Sonnet models being. Sonnet models were the fast things that you use for like in the loop, go make a change and then review it. That is what Luna is for. It is great at that. Model is so much more competent than I think most people realize. It is arguably, I have found it to be better at orchestration than 5.5 was in a weird way. And it's also remarkably good at just doing normal code stuff. Do you know how Luna was post-trained? Yeah, by 5.6 Sol. No, Terra. It was Terra? Yep. Seriously? I'm pretty sure, yeah. Jesus. Yeah, confirmed it was Terra. Okay, so I mean it's not, the problem with this analogy is that calling 5.6 Mythos, even though, calling 5.6 Sol Mythos, even though I do defend 5.6 Sol over Mythos, is not correct. Like it is not a fable or Mythos class model because it's not big enough. It's still missing the magic of the humongous model, but it's the closest they have and it seems like Terra is probably closer to an opus than it is to a sonnet. Like I think calling Luna a haiku model is a disservice to Luna. Same thing with calling Terra a sonnet model is a disservice to Terra. What is the proper Grok Code Fast replacement? Is it Luna or is it Grok 4.5? Oh, it's definitely Luna. It's 1000% Luna. Grok Code Fast is not a model. Grok Code Fast is an idea. It is not real. Like I, A, it's gone. It's disappeared from this world. We can never touch it again so we're left only with its memory. You missed out on the opportunity to use Grok Code Fast with GStack. I know I did and that will haunt me until the day I die but in order to live on, have its memory live on, the closest thing that feels the way it feels in my mind is Luna. Would you like for me to hit up my XAI contacts to see if they'll provision Grok Code Fast for just a day for you so you can use it briefly to see what it's like to use a GStack? No, because then it'll ruin its memory. Like you can't, it is greater, it is better to think about what it was in my mind at the time than to actually experience it now because if I experience it now it's going to be utter garbage like horrifically useless. This is a level of maturity I did not expect from you. I think it's time to wrap up and drink before we say something intelligent. Shit, that's a great point. We accidentally do that sometimes. It's bad. Our mistake. Yep. Anyways, thank you guys as always. Leave some rankings and reviews on whatever platform you're watching on and if you're watching on a platform you don't normally like check out the others. We're probably on there too. If Alyssa doesn't hear me say this I get whipped at the end of the show so hopefully I am now safe. I'm going to go spin up some agents. We'll see you next time.