← Back to search
Ox Alpha Revealed, OpenAI's Latest Pricing Updates, and Our Coding Model Tier List
Nerd Snipe with Theo and Ben · 2026-08-28 · 162 min
Show full episode description
Theo & Ben break down OpenAI's GPT-5.6 Sol price cut, breakdown the "Ox Alpha" stealth model we now know is GLM 5.3 Flash, then take a drink every time they say "Grok" while ranking every current AI model on a tier list! What could go wrong? Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Which coding models actually deliver the best real-world value amid
OpenAI's price war with
Anthropic and confusing token-cost comparisons.
Benefits
- Understand real end-to-end task cost vs misleading raw token prices
- See why GPT-5.6 Sol's efficiency beats cheaper-looking models
- Track OpenAI vs Anthropic pricing and growth dynamics
- Get a head-to-head coding model tier list from two developers
Use cases
- DeepSWE comparison: GPT-5.6 Sol on high reasoning ~15 cents cheaper per task than GLM 5.3 despite higher list price
- Sol medium 50% faster end-to-end than Gemini 3.7 Flash high, despite Flash's 340-380 tokens/sec output
- Enterprises defaulting to Opus 5 on Bedrock for ZDR compliance, despite a cache-write bug inflating bills to thousands vs ~$50
- OpenRouter, OpenCodeZen, and Vercel AI Gateway offering GPT-5.6 Sol at 50% off
KPIs / results
- GPT-5.6 Sol API price cut 20%: $20/mil out short context, $30/mil long context
- Ramp data: Q3 2026 API spend growth flipped to OpenAI 82% vs Anthropic 76%
- Gemini 3.7 Flash: ~200 tokens/sec on OpenRouter, ~13 cents in / $2 out
- Sonnet 5 max effort ~30% more expensive than GPT-5.6 Sol
📑 Chapters — tap a time to jump there
0:00
Intro
- Drinking-game intro for the model tier list episode
- Topics: OpenAI price war, delayed models, Ox Alpha stealth model
14:44
Anthropic’s delayed models
- Anthropic delaying releases as OpenAI delays its big model
- Opus 5 remains enterprise default on Bedrock for ZDR; costly cache-write bug discussed
20:53
Kimi K3 and open weights
- Kimi K3 and open-weights models: cheap on paper, token-hungry in practice
30:08
The Alpha stealth model
- Ox Alpha stealth model drops on OpenCode and OpenRouter
- Absurd throughput suggests the model is really cheap
48:26
Model tier list begins
- Ben and Theo begin head-to-head coding model tier list rankings
1:09:53
GPT-5.6 Luna
- GPT-5.6 Luna placed on the tier list
1:13:09
Opus and Sonnet
- Opus and Sonnet ranked; Sonnet 5 called poor value for its price
1:20:25
Gemini models
- Gemini models ranked, including frustrations with the Gemini Flash line
1:31:15
DeepSeek and local models
- DeepSeek and local models evaluated for the tier list
1:59:58
Fable vs. Sol
- Fable vs Sol debated: the hosts' differing picks for current best model
2:28:46
Final rankings
- Final tier list rankings locked in
Oh, I just came up with a fun idea for this episode. What do you want to do? A sip of a drink every time you mention Grok or GrokBot. I mean, we're going to get hammered, but like, no, yeah, drinking during the model tier list, I'm down for that. We'll wait to do it until that, but once we get to the model tier list, because I think we'll hit that pretty quickly, there's not a huge amount else to talk about. Is this an intro? What, teasing that we're going to get hammered at the end of this? That's more of a reason to keep watching, isn't it? Yeah. So, uh, welcome back to Nerd Snipe, the only developer podcast hosted by developers. I am... Drunk developers, to be clear. Yeah, very. Well, soon. I'm Ben, he's Theo, and we're going to be talking about the... Fuck, I need to retake this. I, uh... We got such a good transition in, man. We did, but like, my brain was not in intro mode. I'm not... I fucking... I hate scripts. Scripts suck. I can't do them. It is really hard to talk and read off bullet points. All right, you ready? Okay, so we're actually here to talk about models, as we often do, but I think the topics will actually be fun. First, we're talking about OpenAI's waged war against Anthropic. It's getting pretty intense, with some wild changes in price cuts that will affect most businesses and business use cases. I know that I've been impressed with how much cheaper things have been lately. Also, it seems like they've been delaying the release of a certain new big model, and that Anthropic's using the opportunity to delay theirs. We have a lot of thoughts there, too. But they also might have some new competition, with Aux Alpha, the stealth model that recently dropped in OpenCode and OpenRouter, with absurd performance and, more importantly, absurd throughput, which suggests the model might be really, really cheap. And finally, a thing that you guys have been asking for for a while, a model tier list. I already did one on my channel, but Ben has not done one yet. Nope. And I think it'll be fun to see how we go head-to-head here, because I happen to know we have different takes on what model is currently the best. We do. If you have a product, and that product has users, you're almost certainly missing out on a ton of other potential users, because you need to localize your product. And the only sane way to localize a product these days is with General Translation, the sponsor of today's episode. Localization usually takes an entire team and a huge amount of engineering efforts, but they make it dead simple. They have SDKs for every major framework and language. And most of the time, when you need to translate a section of, you know, like some React code, you just wrap it with their Magic T component, and suddenly it is translated into any language. They're also designed to work across your entire company. They have a shared context layer, so that if you maybe have some style guide or some specific way you want your translations to work, you make that change within the doc site, that same change will apply to the translations, which happen on your homepage and your dashboard and your mobile app. All of it fits together seamlessly. And this experience is the reason why Ramp, Cursor, Profound, Particle, Sierra, and many, many more are already translating their sites on General Translation. If you go to the Cursor homepage, you'll notice there's a lot of text on here, not just like normal heading tags, but even like their custom visualizations on here to show you what Cursor looks like. When you go to the French version of their site, you'll notice that all of this is translated. The headings, the buttons, and even this really complex demo they have built into the homepage, all of this gets translated because General Translation makes it insanely easy to do. They handle things like dynamic layouts for you where, you know, one language is left aligned, another is right aligned. They make all that easy to deal with. They handle currency conversion, number conversion, and even have search optimized routing for any locale. This is the biggest no-brainer possible. You need to get this set up for your company at nerdstype.link slash General Translation. All right. Well, thank you, General Translation, for sponsoring this episode. And let's get into it. Let's do it. All right. So I think I want to start with the OpenAI v. Anthropic stuff because OpenAI has been taking a lot of free wins over the last like six months just dunking on Anthropic being not the most pleasant company in the world to work with. And they've recently done a bunch of stuff, the biggest one being the price drop on Sol. They are currently... Have they already done this? Because what they're doing is they're dropping the API and credit pricing of GPT-5.6 Sol by 20% over the next three months. So I've seen some people talking about how this is already in effect some places, but I thought that this is just an ongoing thing they're doing to just improve this over time. I've also seen on OpenRouter a bunch of the different... I think OpenRouter, OpenCodeZen, Vercel.aiGateway, and a couple others have GPT-5.6 Sol up there for 50% off. So they're doing a lot to cut prices. Yeah. The 20% off is officially out now. It's $20 per mil out for the official version over API in the short context. Long context is still $30 per mil, but that's a meaningful discount now because previously it was very expensive. Input is still double. It goes from $4 per mil into $8 per mil, which is annoying when you're already over on the really long context. And cash rights are still way more expensive at $10 per mil. Regardless, that is a pretty significant price drop on a model that was already really efficient. I think part of why OpenAI is doing this is because people keep comparing the token cost, even though that's not the real world cost. It's actually very annoying that this continues to happen. Models like Kimi K3 or even Gemini 3.7 Flash look like a really good value when you look at the raw token cost. But when you see that they generate three to 10 times more tokens for most tasks, it just doesn't math out. And I'm getting increasingly tired of this part of the conversation because people keep looking at these numbers that don't mean anything and use it to dunk on like the people who like the current Frontier models. But the harsh reality is that Solon Medium is a really hard thing to beat because it uses so few tokens that ends up being way faster end to end, way cheaper because it's not doing as much work. And a surprisingly smart model because it's still a Frontier class model just on a slightly lower reasoning level. Yeah, the efficiency thing, I'm getting like, it is really annoying having to explain this every time because like a great example, I just pulled up DeepSWE. If you compare GPT-5.6 Sol and GLM 5.3, GLM 5.3 is substantially cheaper on paper than GPT-5.6 Sol. But on high reasoning, GPT-5.6 Sol is cheaper by what, 15 cents per task? And it is faster and it does fewer tool calls and it has, I believe, yes, way fewer output tokens. Even on max reasoning, GPT-5.6 Sol has far fewer output tokens than GLM 5.3. Like the efficiency wins that OpenAI has right now are pretty much second to none. Grok is, they were doing pretty well with this and then Grok 4.6 kind of fell off on that. Like it got a lot less efficient and hopefully over time that will continue to improve. The thing that actually matters these days is how many steps it takes the model to do the thing, not really the API billing. Yep. Although I will say the API billing difference here is pretty nuts too. I'm just comparing against models like the Gemini Flash line because I've been a little annoyed at Gemini Flash recently. How much, so I will say that Gemini 3.7 Flash on medium is cheaper than Sol medium, but high is more expensive than Sol medium by a decent bit too. It's like 40 cents according to artificial analysis for their intelligence index versus 37 cents for Sol medium. But much, much funnier is the output speed. If you look at the end-to-end response time, Sol medium ends up being 50% faster than Flash high. So the model that's whole thing is going fast, which is like, it is going fast. I'm just looking at these numbers once more here. Yeah, it is. 3.7 Flash can do 340 to 380 tokens per second, which is triple to, or it's like 3 to 5x, the token speeds that you can get out of 5.6 Sol. But again, end-to-end response time, how long does it take to actually complete a task? And is it 50% slower because that's not the number that matters, and I'm really tired of people pretending it is. Yeah, I'm looking at Open Router right now. 3.7 Flash is averaging 200 tokens per second over API, and the actual price of it, their effective number is effectively 13 cents in, $2 out. Like, it is nothing compared to GBT-5.6 Sol, and yet, not, it's, yeah. I'm just, I'm annoyed with this whole thing. I do think Anthropic needs to be more scared. Historically, they haven't paid much attention to what's going on outside of the, like, Anthropic bubble, and they are getting chewed out on price at, like, every corner. The only time they've impressed me with price stuff is when they dropped Fable, and the cost was meaningfully lower than they had originally announced. I think it was, like, half the price, because it was supposed to be, like, 1 to 125 per mil out, and it being 50, right? Yeah, that sounds about right. 125 is definitely correct. Yeah. So that was a nice surprise, like, a very, very nice surprise. But everything else, not so much. Like, Sonnet needs to be way cheaper than it is, because it's so token inefficient. It just makes, like, no sense at all to run. And actually put it in the running here. If I turn on Sonnet 5, yeah. Like, Sonnet 5 on max effort costs more than 5.6 Sol does by a meaningful amount. It's almost 30%, maybe slightly more, more expensive. So it's just not a good value model at all, and it makes no sense to have in the lineup the place that they have it right now. Anthropic just doesn't care, and they shouldn't have to. Just judging by the numbers and the revenue growth, companies are spending more and more on Anthropic models. They are more than happy to keep charging the rates they're charging. OpenAI wants to win, though, so they're going to keep fighting on the price side. And it shows in the numbers. They already were cheaper. I don't know if slightly more cheaper, the way they're doing it, is going to be the thing that wins. I almost think that the optics of the price drops is more valuable. Like, if they could hypothetically do a 50% price cut, they shouldn't do that all at once. They should do the 20% a few times so that this reputation of cost decreases over time, and OpenAI being the company that if you bet on them, things don't get more expensive as you use them. They actually get cheaper. Yep. Yeah, exactly. It is... They just need to keep fighting Anthropic. I have this pulled up. It's a post from Aura at Ramp. They did their... They have an analysis going on how fast both OpenAI and Anthropic are growing over API spend, which is basically the B2B spend for each company. And for a very long time, like since Q2 25, Anthropic has been destroying OpenAI, like 100% versus 52%. But just recently, this has now flipped in quarter three, 2026, to 82% growth on OpenAI and 76% growth on Anthropic. So OpenAI is starting to outgrow Anthropic, or at least the rate of growth is outpacing Anthropic. And I think this is going to continue. I don't really see a world where Anthropic models getting more expensive, slower, more restricted, and the only one that they clearly care about at this point is just Fable and whatever the biggest best model is. Both Sonnet 5 and Opus 5 are very silly models that no one serious is using for anything. I don't fully agree with that. There is a lot of serious people using Opus every day. There's a reason why it's like the model that someone like Matt Pocock fixates on. It is the default model for coding in a business right now, because it's the best Anthropic model that has EDR. It is the best model you can use on Bedrock reliably that has EDR. And there's just a huge bug you were talking about on Bedrock where it wasn't actually cash writing when it was supposed to, and it almost never would cash read. I saw somebody post that like when using Codex with Bedrock, the cash write to cash read ratio is like eight to one, where it was writing way too much. And the bill ended up being thousands of dollars when it should have been like 50. So for all of those reasons, I get why businesses are still using Opus 5. It is the best model they can use on Bedrock without having to like break ZDR and give all their data to Anthropic. Yep. A lot of companies are still on Opus for that reason. I think that's what OpenAI sees as the opportunity and the wedge. Well, and I think that's what matters right now is that it is more the direction this stuff is going, like especially the growth starting to slightly outpace. I think this is going to be very evident in six months time when there will probably ultimately end up being a shift. I would bet a lot of these companies have committed contracts with Anthropic for committed spend that they're just using and it's all they have. And at least right now, OpenAI and Bedrock really sucks. Like there's nothing good you can do with that. So until they get that right, it's not going to happen. But once we get to the point where OpenAI has all their stuff up on Bedrock, it's efficient, it works well, it's reliable, the prices keep going down, these contracts keep going up. I think OpenAI is going to grow a lot faster than Anthropic is. I don't see growth in the Opus market. That just doesn't, eventually these people are going to end up testing out stuff like Sol and then they're going to much prefer that. And if they can't use Fable for anything because the ZDR policies from Anthropic are insane, it's not going to work. I definitely have noticed the shift here where I feel like there will always be loyalists between Anthropic and OpenAI. Historically, the loyalists have like acknowledged one side, but been like, well, I don't really care though because I have all these things on the other and I'm fine. I've always felt like it was relatively even or skewed towards Anthropic where there would be more people saying things like, I don't care about codex because I really like Claude code and I really like Opus. Especially around like the Opus 4.5 era. There was basically no reason to be using OpenAI models until they caught up meaningfully with like 5.3 codex or so. That was a strange era, but I feel like we've flipped into an even stranger one where I know a lot of people that are using 5.6 Soul that just genuinely don't give a fuck about anything Anthropic anymore. Like they don't care. I know people who like tried Fable once were like, yeah, I don't really see a difference and then just stuck with Sol, which is strange to me because I still feel a meaningful difference, but I feel like I know more people who have become OpenAI loyalists with Sol than became Anthropic loyalists with Fable, even though Fable is the better model. As the people in this conversation who that description fits pretty aptly, other than I have done a ton of Fable use and I do, I really like the model. Yeah, like I'm not thinking of you when I say this to be clear. No, I'm not nearly as extreme in that position, but I am very much on the Sol side. Like I much prefer that model. And if you just told me tomorrow, like, hey, you can only ever use models from OpenAI or Anthropic again, I would pick OpenAI every day of the week. I would too because their portfolio covers my usage much better, whereas Anthropic really, as we said, is just focused on the like frontier frontier. Yes. As soon as they have a new best model, they just stop caring about everything below it. Meanwhile, OpenAI almost seems like they're thinking more about the smaller models to an extent right now. Like they're thinking a lot about Luna and they're thinking a surprising amount about like Sol, even though allegedly they have Astra, which is an even bigger model internally now. And I think that's an important note as well. Anthropic has a new snapshot of Fable or something along those lines internally. Yes. That they have made it pretty clear internally. They have no plan to release. There's a few reasons for that. The obvious one is that they don't want another ban to restrict employees from having access, but the bigger one is just that they already have the best model. Why would they have first and second place? Okay. They already think they have first and second because they think Opus is one and Fable five is number two. It's bullshit. They do by the benches have one and two. I, in reality, I think it's Fable is number one and Sol is surprisingly close at number two. And then Opus is wherever it ends up when we do our tier list later. I think they're holding because they're already in the lead and they want to see what OpenAI releases. But in that process, they're also choosing to hold their best model as a not workable for business model because again, the ZDR policies, like everything submitted to Fable is stored by Anthropic for research and making, not for training, but to make sure that it's safe. That is, it makes it untenable for most businesses. They don't have any real reason to release a new one. Like they, I would put a lot of money on it that Anthropic is making significantly more revenue on Opus five than Fable right now. I like a lot. Yeah, almost certainly. That would make sense. I wouldn't be surprised. So for all those reasons, they are choosing to wait. And if OpenAI puts out Astra and it's great, they will have to respond in some way. But in the interim, they are just going to sit on their hands and watch the industry run around in circles while they continue to strengthen their lead and make way too much money. Yeah. I mean, I think for Anthropic, you can just kind of look at the way they talk and what they do publicly. They just don't, like, yes, it's a business and yes, they need to make the money because they need to get the investor money so they can keep funding research, but they really do just care about the research. Like, they are just trying to build the biggest, best model and that's it. That's all that really matters. And things like, oh, we can't use Fable because of ZDR stuff. They're like, okay, cool. Don't use it. Whatever. We don't really care. As long as they have enough revenue and as long as they look good enough, it doesn't really matter. And I assume realistically what's going to happen is all of the releases we'll get from Anthropic. And to be fair, I don't think OpenAI is exempt from this. The releases we get publicly will always be at least a couple months behind what they have internally. And that's going to probably, that gap will just keep widening. They, this has always been my concern with the government regulation stuff is that these, it's not going to stop the models from being developed. It's going to stop them from being distributed and it's going to just give the labs this huge lever that no one else will get access to for God knows how long. Because like, no one is forcing them to release a new model as long as like, because I know that, I think there are rumors of a Fable 5.1 snapshot because there's that model. Model 2, yeah. I don't think Model 2 is Fable 5.1. I think those are separate things. No, Model 2 is like a whole new thing, I believe. It was part of their like investor updates thing that there was a bunch of reporting on. No, it's Model 2's Mythos class. So it's just a Mythos 5.1 type thing. Yeah. I've seen that. It was in their responsible scaling policy article, I believe. Model 1, which is broadly similar in capability to Cloud Mythos Preview and Mythos 5 and had a relatively small amount of internal deployment. Yada, yada, yada, details. Then Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that the model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use, but does not display capability jump to the degree observed from Cloud Opus 4.6 to Mythos Preview. You don't have plans to release this model externally. Yada, yada, yada. Based on the results of this review, we expect Model 2 to have broadly comparable capabilities and propensities to Mythos 5. Model 2's performance in Random Benchmark is slightly stronger than Mythos 5 and significantly worse than Mythos Preview. Hmm. I don't know. I mean, we're getting into crazy speculation land on like what Anthropic is going to release, but also like taking SynthWaved as a full source on this is, you know, take it with a grain of salt. He hasn't had many misses, to be fair. He doesn't, and he had a very recent one where he said that a successor to Fable 5, likely Fable 5.1, is now being tested for a subset of accounts with Fable 5 selected on Cloud Web and also potentially Cloud Code. So there are rumors going around of a Fable 5.1 being close to release and semi-ready for release, which would make sense. I'm assuming Model 2 is a different pre-train, a different type of model. Like this is probably something different. I don't think that this is a normal Fable successor. Well, I found a section at the end here. They have internal Model 1, internal Model 2, internal Model 1, legacy Mythos class model not yet externally deployed, internal Model 2, Mythos class model not externally employed. Interesting. So they are saying specifically this is Mythos class. That's what I thought. I was right. This is what could have been an internal snapshot for a 5.1. Yes, but that snapshot is not getting released almost certainly. Whatever we do get will be different and whether or not they are still using that internally, no one has any idea. So again, they're taking their time because they don't have any pressure until OpenAI releases. And honestly, they could do a little bit of cleaning up on this Model 2 snapshot and put it out, especially if they make it any cheaper or finally kill off the ZDR requirements on it. And now they are back in the running. They are in first place and they have no interest in showing their hand until they have any reason to believe they won't be in first place again. Definitely. And I feel like eventually they might fold on the ZDR thing. OpenAI just put out their article about offering zero data retention for frontier models. They are going to be doing that and future, I think 5.6 Sol, you will be able to get deployments of that that are fully ZDR'd so that all the enterprises can use it. They're doing all the things they need to to get as much of this adoption as humanly possible. If there's enough market force against them, hopefully the things that have happened in the past, like OpenAI has been a very good corrective force to Anthropic over the past six months. Everything they've done has been super pro-consumer and end user and that has kind of forced Anthropic to follow suit and make a lot of concessions that they probably wouldn't have made if they didn't have any competition forcing them to do so. Hopefully ZDR ends up going the same way and all of this stuff can be publicly released but it's really just a matter of time at this point. But Ben, aren't open-weight models faster, cheaper, and better? Why would we ever use Anthropic models or OpenAI models right now? I mean, well, here's the thing. This is what I thought and you thought and basically everyone thought until this very strange model. We're not talking about the new model yet. I'm talking about Kimmy K3 and GLM 5.3 right now. Oh, cool. I want to do more bullying first. I'm always down. Kimmy K3 is the model everyone was super hyped about. It's an open-weight model that's 3 trillion params, near frontier class in almost everything, benched better than anything Google's ever released by far. It's not a high bar, but yeah. It's a bar though. It is a bar. Trillion plus dollar company getting, the trillion plus dollar company getting absolutely trampled by an open-weight Chinese lab is pretty nuts. True. So it's cheaper. It's half the price of Sol per token, apparently. But I just, I take every opportunity I can to talk about this because I feel like it is missed pretty often. How do you think the cost of Kimmy K3 Max compares to Sol X High? Like what would you guess? Like how much cheaper do you think it would be? I mean, I'm cheating because I kind of know the answer. I don't know the exact number, but I'm pretty, it's, I know it's a substantial increase in price. Like it's a lot more. It's not a lot more. It's a little more. It's like 81 cents to 84, but I also think these numbers were calculated before the 20% discount on Sol. Yeah. So I think Sol just got even cheaper here where now it would be closer to like 65 cents or so to 84, which would make the gap like 25 to 30% more expensive with Kimmy K3. Yeah. But what's much more interesting to me is the speed and latency one. How much slower do you think K3 is compared to 5.6 Sol? I mean, having used it, it feels like it is at least 2 to 4x slower. It is 2.5x. 2.5, yeah. From 32 seconds, 5.6 Sol X High to over 70 seconds for Kimmy K3 Max. And this is also worth noting that when artificial analysis runs their work, they are not using the WebSocket codex bindings that make codex even faster for multi-step stuff. It is annoying how much faster 5.6 feels than almost anything else simply because it's like full end-to-end is more considered in the design of the harness and the APIs and all these other things. Grok 4.5 could be similarly fast. 4.6 feels meaningfully slower. Yes. And Muse Spark is pretty damn fast too. I'm surprised at that. But those are like models that are designed to be efficient and fast like token generation, which makes up for the gap with things like the end-to-end caching model, like the WebSocket layer and all that. But in the end, if like speed is your concern, you should find the lowest reasoning effort available to you on any open AI model and throw it on something with the WebSocket API instead. It'll probably be a lot faster if you're doing anything even vaguely multi-turn. Yeah, absolutely. I mean, if 5.6 Sol on Medium is 1,000% going to be cheaper and faster, I'm looking at DeepSWE and the difference is even crazier. Like it is $4.65 per task on Kimmy versus, or sorry, where is it? Yeah, it's $4.65 on Kimmy and then it is $3.47 on GBT. Yeah. Like it's absurd. And the output tokens is hilarious. Yeah. If you're cost constrained, you should probably be using OpenAI models. If you are latency constrained, you should probably be using OpenAI models with WebSockets. If you are capability constrained, you should probably be using 5.6 Sol on like the highest reasoning options available. Not, and Max rarely benefits you, but test it obviously. And if you care a lot about where your data is going and making sure everything is running on your own servers, your own infrastructure, I would question why, but if it's a legitimate use case, then OpenWeight models running on your own infra are actually really fucking cool. I don't want this to come off as anti-open weight model. Yeah. I just wish that people would stop pretending they're cheaper when they're not. Yes. Yeah. The whole cost discussion is just so polluted. Hopefully over time they can, I don't know how they package that up and explain that to people, but there needs to be some way to explain like rather than just showing cost per token, cost per task, if there's some way to get a number for that, I don't know how you standardize it, but it would really help to just explain the real differences here. Yep. It is what it is. I think similar to like how developers had to finally understand what a token was, because like, remember the days where you'd pay per message? Oh yeah. Oh, that was fun. Yeah. Like we just have to get used to these things and token efficiency is a thing that people are going to have to learn to understand. Similar to how like programmers have to know the different algorithms for like sorting data and whatnot, they're going to have to understand that the token efficiency matters when you're building and using these tools. Yeah, 100%. Because it, at least the direction the industry is going right now, it seems like there are far more viable models than I expected to have at this point. I was really concerned two or three months ago that OpenAI and Anthropic would be the only real viable options at this point. And while yes, they are still in a league of their own and it is for the best of the best, nothing comes close and unfortunately I don't, time will tell whether or not that gap closes, especially with something like Grok 4.7, but at least for a lot of work and getting actual stuff done, I've been doing a lot of it myself, using something like a Grok or an open weight model is very viable at this point. They're viable models now. I have feelings on that. I think that I was very quick to make the shift in how I prompt and how I work based on the capabilities of the new models like Fable and Soul. And once you get used to that, I find it really, really hard to go back. Like any of the tasks I could confidently do with like a Grok 4.6 or with any of the open weight models, I would rather do with a frontier model and have it send the prompt to do that small piece. Like if I have a large project I want to build and it has like seven steps, I could use a cheaper open weight model or Grok or something and spec out and do each of those steps one at a time. But with something like Fable or Soul, I can do a rough sketch of all the steps and get it to go spin everything up and do all of those parts. And sadly, like while they are capable of prompting in a way that like Fable can interpret the slop from Fable 1 when it's a Fable subagent and roughly get what I have in mind, I find that when it tries to prompt things outside of its own family, especially models that are a little quirkier like a Grok or a Gemini or even some of these open weight models, that getting Fable or Soul to orchestrate those is not as consistent for me. I haven't found, Gemini has obscene quirks, but I really haven't found Grok to be all that weird or quirky. It's just less capable. The best like multi-family experience I've gotten, which I think is going to be a more common workflow over time, is the like cursor stuff and Lauren's like potato mode thing. Her skills with that are really, really good. And when you put it in potato mode, I generally am using Grok 4.6 X High Fast Mode as my main model because you get absurd usage on cursor with that. Then that will spin off lots of Fable agents, lots of Soul agents for stuff like planning and review. But then like the main grind and run is all on Grok. And overall, the price seems to be a little bit better and you kind of get a best of both worlds, but it's still not quite there yet. In a perfect world, if you had unlimited money, you still will get the best experience just Fable using Fable. I would know how long until we get a potato mode model with all the Laurenisms baked in where it knows how to prompt and when to call in Fable and whatnot. Like that actually could be really cool. That would be sick. Take something with Grok's level capability and teach it when it's wrong and like how to shell out to smarter things when it needs to. That could be a really compelling like thing that I would use. Like what would it look like to make a model that's job is to just be the one you talk to? Similar to like how the voice models are trained to just be the thing you talk to literally and it will shell out when it should to some extent. What if you did that for the code model where the top level code model, its only role is to interface with you and make sure all the orchestration is going well and using the right things? I would be very curious to see if that is possible. I think the voice models are very bad at doing that but it's mostly just because they're really dumb models. I think you need a certain level of capability in order to do this and unfortunately it does seem like especially that prompting ability really benefits from having a bigger model. Like RL, you can use that to smack in being really really good at writing code and running for long periods of time and orchestrating subagents but I have not found it to be nearly as good about the prompting stuff. I counter you with Kimmy K2. Kimmy K2 was surprisingly good at writing if you remember. I do remember that, yes. Did you ever read the paper on how they did that? No. Entirely RL. Really? They had an RL pipeline for quality writing and like good prose. Interesting. So you can't do it it's just there's no economic incentive to right now. Yeah, that makes sense. But differentiating I think is enough of an economic incentive at this point. Like if you can make your model actually like novelly good at one specific thing like apparently a lot of the reason Gemini is still used is because it's the best model for like internationalization. Oh yeah, I could see that especially with their data corpus like they probably have great international data but then you have to like deal with using a Gemini model and honestly like let's be real it's a much much easier way to get your translations handled properly inside of your applications. True. We're not actually doing a sponsor break we ran it earlier but general translation is awesome you guys should go use them. No, I didn't know if we had the sponsor break in yet but yeah I am actually a really big fan of general translation I love for sponsoring this because like that's just it's a great product. It is. I am very impressed to all of those guys. Yeah. Earlier last week there was a new stealth model dropped on Open Router and Open Code Go and a couple other places called Ox Alpha that no one really paid all that much attention to but there were a couple weird things about the announcement that we were both just very confused by specifically how much provisioning it had. The Open Code team when they posted it they said that they have allocation for 100 trillion tokens per day which smells like a Frontier Lab and the Frontier Labs have not done any stealth drops for quite a while so we were confused like what the hell is this thing? I'll be more transparent here I was DM'd by our contacts at Open Router told that this model is available it's Frontier Class they have a ton of provisioning and they really wanted us to try it and potentially integrate it in our products. Normally Open Router doesn't hit me up for these things anymore they'll still hit me up here or there but I was actually surprised to see that channel active especially right after their Stripe acquisition. I did not even know that they hit you up about this. Damn. Like morning of when it dropped. So I had like a mental note that something's interesting here something's going on and then when I saw on Twitter that Open Code had the 100 trill a day allocation which for reference all of Google's Gemini usage is in like the 150 trill a day range from what I've heard. Yeah. So that's an absurd level for just one of the companies that was being partnered with for this. And also it is 1 million contacts it is multimodal it has ZDR and it is completely free during this like stealth period. Yep. So all of these are like big what is going on here. Yes. This means a weird combination of things. First off whoever's doing this has endless compute that they have around that they can spare for an experiment like this. Whoever is doing this is almost certainly not one of the major two labs because they would benefit a lot from doing a release like this themselves under their name. Absolutely. There has to be somebody who benefits from the name not being known initially and then being known later. Yep. Because like if a new model came out from Thinky Machines and they give me free provisioning to try it I'm not going to bother. But if they give me a new model from OpenAI or Anthropoc I'm obviously going to go try it. Yeah of course. It's somebody whose name isn't big enough that they would want to do this in a stealthy way to work up some interest. Yes. It's somebody who can afford to do this this way and somebody who doesn't need the data because they're doing the like we promise to not train your data thing. Yep. And it's also just really good. Yes. Like surprisingly so. I tried it because of the both message that I got from OpenRouter as well as the provisioning I saw from OpenCode. Those two things maybe question what's going on because the model should not be able to do that much and for free for that long. Yep. And then Ben ran it against a subset of the DeepSwee bench because they have a couple test problems public and it scored higher than Fable. I want to be super clear on that subset. Subsets of benchmarks are not a great like measurement of how good a model actually is. I just posted that because I was confused. I was like what is this? It went a little more viral than it probably should have because no this is not. How many million plus views did that post get? Way too many. It was like I didn't even hit a mil. Damn I thought I did more than that but nope. But yeah now like 850k views on this post that is effectively misinformation. My bad. I was mostly just very confused because to be fair the fact that it got that far at all does mean that there's something there. Like this is not a bad model and it did outperform these other models on this one specific subset. It was run by the DataCurve team which is the company that made the DeepSwee bench and they got the actual results of this. I think this is still on a subset but it is a much much larger subset with real data and it's averaging about a 63% which puts it on par with something like GPT-5-6 Sol Medium I think like slightly below Fable on Medium so still very very good and it is a pretty fast model. There's a lot of behaviors that were really nice like I really liked the writing and we were just like what the hell is this thing? Yep and with those numbers we had to assume this was a pretty big model that's doing some pretty crazy stuff from some lab that has a bunch of compute. It's GLM which is confusing because GLM-5-3 just came out like a week or two before this model anonymously dropped. So what is it? Is it GLM-5-4? Is it GLM-6? Did they quietly have some much bigger pre-training going on? What was it Ben? It's GLM five point something flash. I think it's 5-3 flash almost. 5-3 flash almost. Sweet. So it's 5-3 flash. Yep. It is a 300 bill per am model that from our math we haven't had a chance to try it because the weights aren't out yet. Yeah. We should be able to run on 2DGX sparks. Yep. Which means this model is like a DeepSeq V4 flash alternative that has vision baked in that is more capable follows instructions better and is just all around a really good model. Yeah. It's fantastic. We were sitting there last night trying to figure out what this thing was before we got the official confirmation and I was like sending it through all the models trying to sleuth what this was. It smelled like GLM. It had very similar shapes in requests to GLM. It had the same vision properties of GLM. The tokenizer was the same. Because GLM 5.3 non-flash doesn't have vision which is the main reason I didn't think it would be that. And I also just didn't think they had enough compute around and I was working under the impression it was a bigger model which would need way more compute to serve at that level. My assumption was Xiaomi had finally manned up and like stepped in again. I know they will someday. I love my Xiaomi. I am sure they will reappear as a dark night late in the evening. some random time in the future. Probably. This is not that. This was GLM 5.3 flash surprising me at the capabilities of a really small model. And that's actually the thing I want to take the opportunity to talk about here. It's a conversation I've been trying to find the right framing to have and it's part of why I mentioned like what if we trained a model just to be good at orchestrating and talking to it. This is the first model that I think really emphasizes the gap between how a model behaves and how intelligent a model is. Yes. Historically the smart models are also the best behaved ones because they need more knowledge to do better things. Now that RL and our post-training pipelines are as important as they are to get this agentic multi-turn experience where models aren't just doing the thing that they read online five years ago in the training data. They're now doing things that are unique to how we use models with data that just doesn't exist in the public. They have to be built and trained for the specific thing. this is one of the best examples of a model that has that behavior side, the instructions, the tuning of it and how it behaves in a very, very high class, like one of the highest tier usable models that I've ever used where the usability is far and ahead of anything Google's ever shipped while also being one of the dumbest models I've used recently. Yep. It is not a smart model and that's not because like it's going to hallucinate a whole bunch or it's stupid or was trained on bad information. It's just smaller. It is a small brain that was trained to do things this very specific way and it does it super reliably. It makes it genuinely pleasant to work with and I think this is so useful because we've had a lot of models on the other side like a Gemini 3-1 Pro where the model is unbelievably smart. To this day, it's still the only model to score over a 90% on my skateboard benchmark skate bench where it's trying to name skate tricks based on the rotation of the board and the skater. It's a weird bench because you have to have a lot of knowledge of skateboard and the unique grammar and language in it as well as the spatial capability to understand how the skater and the board move and the relationship between those. Even the highest end models like Fable score in the low 80s at best, sometimes high 70s. In 5-6 is I think like an 82% or so. Google is a huge lead there because the model just knows a lot and is really good at spatial reasoning. But then you try and use it and you want to pull your hair out because it's not good at doing what it's told. Yep. And that difference is important for you to understand and it's been hard for me to frame. And I think this release will be the first one that makes that a conversation we can really have. Absolutely. It's really, really cool because like what this thing is, is again, it's not that smart, but it's so good at running within a coding agent harness and doing stuff there that like that intelligence issue doesn't really become that big of a problem since it can search for everything it needs to. It can pattern mashups of existing stuff within your code base. You can use a bunch of tools, use all of your like LSP service things and then end up producing stuff that is far beyond what you would ever expect it to be able to do. I think a lot of this comes down to it's just RL, like clearly the ZAI reinforcement learning pipeline is just unbelievable. We talked about it a tiny bit when the GLM 5.3 announcement happened a couple weeks ago. It was very much like, oh cool, this is neat. They have like this real world environment setup that is very easy to iterate on and build on over time to keep improving their models. This is probably going to produce something cool someday and that something cool is already here. This thing runs in real world code bases better than basically certainly any open weight model I have ever used and the fact that we are potentially going to be able to run this thing on sparks running at home sitting on our desk is crazy. Yes. The idea of having a near or borderline frontier like would have been frontier four months ago model is running on my home server as much as I want whenever I want with the thing we didn't even mention with all this is since this is an open weight Chinese model these things have historically been very fine with doing basically whatever you tell them to and since you own the weights it's very easy to jailbreak and customize these things into doing whatever you want and now that we have this like this kind of feels like a Pandora's box moment where like the frontier labs are probably going to feel the pressure especially on the safety side where like you know open weight capabilities have been substantial for a long time this is a huge step function and also a huge accessibility increase to where way more people are going to be able to run a model with no safeguards at all at home for less than 10 grand. On that note one of the things I think is interesting here you had been telling me before because you put a lot more time into GLM 5.3 than I did that you were really impressed with the improvements they've been making to their RL pipeline yes why do you think they just put out a flash model after those improvements because flash models benefit a ton from really good RL since the bait like again it's literally what you said earlier is that the base of it is not very smart but if you can just get it so that it behaves incredibly well in a coding agent harness and in real world work it can make up for that intelligence deficit by just calling tools and doing things what if I told you I don't agree all right I agree with all these facts but I don't agree with this being the reason I think the reason that the flash model got this treatment first is because flash models take less time to train I would guess that there's a 5.4 that is getting a similar treatment that is going to take a lot longer to complete said treatment oh I'm sure absolutely I actually well I just think this is the first model that has this capability because it's the first model that was being trained for this capability that completed its training run I think that they pioneered this their slime thing with GLM 5.3 but that was like the first iteration of it that was also an iteration of what they were already doing with 5.1 and 5.2 if I recall it was yes it all this runs on slime our open source post training framework for RL scaling with Megatron on the training side in SGLang on the rollout side its design keeps training rollouts and the data buffer on a single data flow so math code sandboxes verifiers and long horizon agentic environments plug in as data generation rather than as changes to the training loop this is what this is what let us keep adding environments through GLM 5.2 and GLM 5.3 without rebuilding the training stack each time so they've had this training stack that they're just improving on over time and it's gotten really good they trained a flash model on it as you said it was almost certainly done first they can release this and then whatever GLM 5.4 ends up being is probably going to be a monster yep if more and more of these like post training hacks patterns solutions etc become public information the capability gap on behavior will be closed and I still think this is where the biggest gap exists right now where like opening eye models even though like smaller ones are much better behaved than almost anything else on the market like to the point where it's almost annoying when you tell Luna even to do something it will do exactly what you tell it absolutely oh or even in product models aren't quite as good at that no I mean look at Opus yeah like Fable for example benefits so much from the absurd amount of information that's baked into it and is relatively well behaved it's annoying and I hate how it talks but it it's thoughtful in a way that is rare and it's behaved well enough that you can use it for real world stuff if I just want them all to do exactly what I say if I wanted to act like a tool open AI has a huge lead yes it is not unlikely that open weight models will get ahead of where anthropic is in terms of like instruction following and like doing what they are told to in the not too distant future I would argue that even now with this aux alpha model we're closing that gap pretty fast it might be a spicy take but I think this thing is much better behaved than Opus is it definitely has some small model smells to it I've used it for a lot of coding work over the last 24 hours and one of the big things I noticed is that it will leave like dead code snippets around like that is the every time I've had like Fable do a review of its code it's like yeah this functions and this is like well put together but what are all of these dead functions doing so it'll do weird things like that it'll miss random edge cases and it is definitely not the most efficient thing I've ever seen the number of steps per task is pretty high and especially when I was running some other benchmarks that were open source locally it was actually timing out on like the allocated benchmark time since on the higher reasoning levels it would just go for very very long periods of time this is not a perfect model by any stretch of the imagination but it is pretty remarkable what we are now getting at the size like I know this is a little dramatic but I really do think that a small model that can run on less than 10 grand of hardware that's pretty easy and accessible for everyone to use being this good is a is as big of a like it's going to be as big of a moment as something like mythos this is going to change a lot of things and will it like I don't know what the full implications of all this is going to be especially like you know we've talked about security stuff a bunch this is now like you can just send this off to do whatever you want it to do and it's just going to do it if this was actually a Google model I would be suddenly so much more bullish on Google oh say Google was the first lab other than open AI to make a model that behaves this well in the US oh yeah I honestly I thought this was going to be a US model at first especially since we were like ruling out GLM on this thing I thought maybe this could be like I'm the meta stuff is a little too fast for it to make sense but like it could be a baller play for someone like SpaceX but the problem is like if they if any US lab had this there is no reason for them to not just put it out now because this is going to be such a huge moment for GLM yeah I knew as soon as I saw the numbers and how it behaved that there was no world in which this was a US model just because like none of the US labs had any reason to release it this way yep well maybe meta did I think meta or Google would benefit from this because getting this much hype built up around a model over the course of a week and then suddenly it's revealed that like oh shit meta's done something insane it would be worth it for the press meta's going too hard on consumer and building their ecosystem like muse spark CLI was more notable than the air muse code sorry was more notable in the model and they like pushed the muse code CLI harder than the model and they dropped it recently I guess that's true I still haven't used muse yet it's fine like I like it I actually think it's pretty decent and we'll probably talk about that in the tier list oh we will that we will any last words on the Chinese lab that could I don't know I I I don't really have any words for this other than I'm very very excited for this thing to release and I hope I hope I hope I can run this locally and this is what I wanted to see I wanted to live in a world where you could get near frontier performance running on your own home server near frontier behavior not performance behavior it acts like a frontier yeah but you can get it to do frontier level things like a lot of the yes is it going to be as clean and easy as fable no will it require more effort to get the same output yes but it's still possible it it can solve a lot of the same problems fable can it does not have outputs of the same quality that fable does especially on the code side fable still writes the most mergeable code it just does yeah it does it does out of the box you can my the point that I've tried to make and I guess actually this is the time to talk about it is that fable does write the most mergeable code with the least amount of effort but you can get the same level of quality out of any of the models that have crossed a certain bar by putting in more effort by instead of putting in like one prompts that's like a couple sentences you give it a lot of context in the prompt you have a lot of skills you have a very well put together code base that you like manually read and put time into curating if you set it up in a certain way you can get that same quality out of other models you can get Gemini flash to to write the best code you've ever seen as long as you paste it in the input well okay I'd say about effort levels here there's a certain level of effort that doesn't make any sense to be even remotely practical I'd tell you about like in the the Overton window of reasonable amounts of effort to put into your prompting and projects in 20 2026 or August 2026 it is fables the easiest GBT grok and now this I disagree so hard that I need a drink before we even start really gonna be a fun section really okay all right here are something stronger I'm gonna need something stronger pour that a little stronger than intended turns out pouring a giant bottle of vodka with one arm is harder than I expected so yeah no I'll be suffering more than expected for every drink as we do this one oof yeah I am sorry in advance for the what the world is about to witness this will be fun yes this will and this is not what people signed up for at the podcast I don't know what they thought they were getting honestly true it's honestly what more what I expected it to be from when we started it yeah it kind of it was it became AI news of the week yeah I think the news is good but I think we should do more bullshit like this because this is really fun I uh I made two tier lists here I used like bed for them uh which was actually a very good fit I have a model tier list and a harness tier list I don't think the harness tier list is as fun but I can make fun of t3 code while we do it so we should do that later and but the model tier list is this is the main event oh boy yep all right where do you want to start well I want to start by announcing the fun game I'll be playing which is every time Ben says anything about grok or even mentions the word grok I will be having a sip of my drink and I would encourage those at home who would like to play along and have a lack of care for their liver join me I take no responsibility for your upcoming hospital visit yeah well in that no way I guess we should start with a grok 4.6 rock 4 point all right at least let you finish the point and I'll keep track of how many steps we have to have all right fair enough so grok 4.6 I copied the tier list you did for your video I think these are reasonable tears I can see how these will fill out I where do you want to put grok I want to work off that I think if this was grok 4.5 and this was filmed at the time where grok 4.5 dropped I'd say it would be a pretty high B tier model it's surprisingly smart surprisingly fast surprisingly good at doing what you tell it to surprisingly cheap crok 4.6 took some of my favorite things away where they made it much less token efficient which makes it less cheap and less fast at which point I'd rather just use sole medium uh the price I'm not as concerned about because when you're using it through the grok subs or through cursor it is effectively unlimited like I added it into a god I guess I'll talk about this I set up a Hermes agent in the discord group chat that I have with all of my friends and we were hanging out last night we thought it'd be really funny to put that Hermes agent in there of course I put grok as the model powering this thing and we hammered it with things that cannot be repeated for a long time burned many billions of tokens and I use like seven seven percent of my weekly usage or something like that you can use as much of this thing as you want to the price isn't really a consideration but the speed is there is a huge problem on speed with this thing it is not it feels like an interim model like a step function type thing where four or five looked really good and had like okay this has a lot of potential they didn't do the thing opening eyes been doing where if you look at the efficiency graph and deep SWE goes up into the right this guy is going kind of up into the left which is not the direction you want to be going hopefully from what they've teased on grok for seven that will go up into the right but I want to put this thing in B tier that is seven drinks seven drinks mm-hmm wait is it just how many times I say the word grok mm-hmm so it's a rat eight damn I had three big ones instead fuck yeah that's gonna be how do you feel about B tier for a model or I can say it grok four six yeah I I want to see tier it but I have not been using it as much as you I my concern when I do tier list is that we don't leave enough space in the top tier because like I can think of at least three tiers of model above grok four six that I would consider distinct okay actually I I'm willing to do it this way because especially since you we have S a B C D F and Google mm-hmm I think we should spread this out to like not do the thing where everything is an eight out of ten we should put this as like if I was doing this out of ten I would put grok four six at like a a solid five to seven out of ten so I think I'm in the middle I am okay with like this being the midpoint of the tier list is actually great anchor point one more drink that one more two more yeah I'm quite an episode God okay so now that we have put the forbidden model in its right place right in the middle at sea mm-hmm for those at home who have guessed this yes the point of this game was just to make Ben stop talking about grok I want to talk about muse next I know it's a weird pick mm-hmm but muse sparks kind of like my equivalent of your like grok four six except I'm too busy to use it as much as I would like to yeah it also have five accounts with Claude and three ish with open AI so I don't really need a cheaper model but if I was token maxing and like really trying to save my money when I did it mm-hmm I could see myself using muse spark one two on the contributor tier like very heavily even just said API prices yeah because the API price on the contributor tier is basically free how's the efficiency pretty good but the craziest thing is like I can give it really big difficult tasks that aren't like things that change code like the test I did and it's my favorite model test now mm-hmm I have this giant open source project t3 code that has a thousand PR is open that's getting like five hundred to a thousand PR is a week at this point yeah getting through all of that is impossible so my favorite tasks for models both to test their like taste orchestration and like general usability and efficiency is go through all the PRs that are open and help me prioritize yeah okay that makes sense and like oh I think you showed me this it flew through it right it did it in under three minutes yeah for a thousand PRs split up well across sub agents and it cost under 30 cents that goes 10 cents for my first run okay that's really cool that feels like you know what whenever we talk about speed on these things I for what's a different model that isn't that model when I'm thinking about something like GLM 5.3 it is is it the it is on paper faster and the efficiency like if I'm just caring about speed in my normal day-to-day workflow most of the time when I fire off an agent run I tab out and I go either do another agent run somewhere else or I go do some research or I read some code or whatever I want to do then I'm not looking at it the whole time that sounds like it's at the level of speed where I could just be sitting inside the same like I could just sit there and watch it run for 30 seconds and be okay with that the catch being it's not a smart enough model that you necessarily want to do that and half the time you're spending sitting there is it realizing it tried the wrong thing three times and to go try something else oh that would be but it gets through it so fast that it barely matters so like the the use case I see especially at that insane cheap price is just like random shit that you would throw on a cron in open source projects to keep track of like I would use this to write me an update on an open source project every day when it like merges changes to see what's going on or like as the the dumb bot I used to keep track of things without ever worrying about cost it's great but there is a huge huge catch which is that contributor to your thing yes the base tier pricing for this model is a dollar 25 per million and 425 per mil out mm-hmm but that contributor tier like do you know how big the gap is uh isn't it it's like 10x right more than 10x it's 10 cents per million and 20 cents per mil out that's a 23x jump in output cost they are fiending for data over at meta yes oh and if you're willing to give up that data this is the cheapest intelligence will be for a while yes yes it is that is like they are eating like the electricity cost is higher than what they're charging at that point absolutely but that's like like when given free intelligence take it especially if you're working on like open source shit like I am with t3 code would I use this for like managing my taxes or finances fuck no no hell no what I use this model for doing blind dumb like audit all my issues or highlight some PRs I should go look at and merge yes for like finding needles in order to find a needle in a haystack you have to process a lot of hay yeah and that's expensive and they have given us the world's cheapest hay management tool okay so before we put this on the tier list I think its position is entirely determined on where we put Luna because I think it's like it in Luna it sounds like from everything you said sit in kind of a similar place how would you compare it to Luna Luna feels like dumber soul and that is what it's meant to be yes the other aspect of muse I like is that it has a different flavor yeah you mentioned that because it doesn't feel like the other models at all which gives it something like I cannot imagine Luna giving useful feedback to soul I have had spark give useful feedback to soul yeah so that makes the comparison hard because like when comparing against a dumber soul you have to have some value in the fact that this isn't soul it's something else and it behaves differently accordingly I think the bigger thing that makes this hard to position is the the gap between the contributor here in the normal pricing yeah I think if we're looking at the normal pricing this is a d to f tier model absolutely yeah well you're looking at the contributor to your pricing for the people who can and will use it at that mm-hmm that is a beat here just for the novel value it provides yeah I'm kind of with you I'm down to average at see I actually I think honestly the type of person who's going to be using this thing for real work is like it isn't enterprise going to use this probably not because they have more money and they could just pay for Luna it's whatever but if you're an open source maintainer who's working on open source code that code's already public most that stuff's already going out there it's not that big of a deal so realistically speaking you can just eat the contributor tier so I think we should judge it on the contributor tier but I'm actually okay with doing B yeah I I think that their angle for releasing this should have been like that type of thing because like meta knows open source they have been contributing to open source world forever if they release this with like a one-click thing you could add to your github repo that will help you with triaging issues and stuff and they just hate the cost for it yeah it's open source who fucking cares that would have been huge I I think they have a great opportunity here and this model is a is the revival of meta's brand in the AI space for me it's a bold pick to put it in B and if we do I want to be very clear this is only because of the contributor to your pricing like that's what pushes it over otherwise it would be C or D yeah but if you're down I'm down or be tearing it I'm down to be tear it entirely off of the price and speed I think that is reasonable it's still looking at this putting it over grok for six does feel bad don't do it um cheers but still for that reason I'm okay with this we're off episode I I was scared I wouldn't be able to convince you on that one because this is actually put it on C tier before but the more I've thought about it the more I think it is more valuable to me and to the ecosystem than grok for sixes and also remember it's gonna be open wait soon well nah I want to believe that that's true but in his duck said it did he say that specifically in a separate post or are you referencing the state of the a like meta's vision on AI article he specifically said it would be okay if he specifically said this will be open source then fuck yeah that's awesome I can't research yeah meta has announced plans to release the model weights from use spark 1.2 soon awesome okay then in that case easy beats here yeah Alex weighing uh we'll be releasing an open weight version of you spark 1.2 soon yep uh yeah easy beats here like I if we had like pluses in here I would be comfortable playing this like b plus like this is if it's gonna be open weight it's that cheap and it's that fast and it is relatively capable for the types of tasks you want for a cheap dumb model hell yeah great model I think the non-contributor price should be about half what it currently is if we didn't have that absurd discount there I would be dear after easy but like that makes it a unique thing and a lot of what I'll be waiting here is like what does this model provide that others don't yeah grok 4.6 is like the main value of grok 4.6 is that it is a better subsidized model inside of cursor and like grok super heavy type subscriptions and it is surprisingly good for what you get there but it isn't like bringing anything super novel beyond like what it brings to that ecosystem you spark has things that nothing else does because it it tastes different fundamentally yes and that price for the contributor is insane yeah I I hard agree I'm down for that cool who what is next something else that something we'll disagree on yeah we need to do that we've agreed too much on this sonnet 5 I you're gonna be disappointed in the lack of disagreement here oh you were the sonnet 5 defender I was and I will continue to defend it while putting it in F tier okay I guess that was easy yeah it belongs in F tier but this model people at the very specific moment in time when it released and also this is a very unique thing to us because at the time we had had early access to gbt56 all so I think it came out right in the middle window between when we had lost gbt56 all early access and we had also lost fable due to the government ban so we had become so used this new next generation paradigm of orchestrating long-running tasks through sub-agents and workflows that it being the only model which could do that at that time felt good and I thought it was cool that they were able to stick that into a sonnet to your model it did not is it the model I'm have I touched it since then no will I touch it ever again no but Ben if you're willing to put more effort into your code base and your system prompts and your skills fuck you and you can make the no other models behave the way you want them to you you are the Overton window guy you are the one who has told me about this Overton window framing it's good framing but no no no no no we are not Overton windowing this man the Overton window has fable on one end and it has grok46 on the other end put that fit yeah yeah it has that on the other end sonnet 5 is down here sonnet 5 requires far too much effort to be worth it we have to think about the amount of effort that goes in from 0 to 100 at like infinite effort you have Gemini anything at negative effort you have fable sonnet is too close to Gemini for me to ever want to touch it so no okay do we skip to fable 5 so we can have this conversation or do we keep teasing it so that we can maintain our retention yeah let's tease a little bit more okay yeah I I'm going to break you on this one oh I'm so excited for it so yeah I don't know if you will let's work our way there with other things we hopefully can agree with let's pick a Gemini v4 pro what you mean deep seek v4 pro sorry I deep seek before sorry we're in a very different god clearly the the grok mentions are working yeah they are yeah grok's gonna become code word for drunk going forward that is oh better than what it used to mean or better than what it used to mean can we just like stick an eric video in for that moment just play the entire thing in the middle of the podcast let's get a they pay him to get a chef feature oh my god I would pay him so much money if he would just like come if he would just come for an episode to like bring grok up to talk to us during an episode I would pay so much money for that I give him the entire sponsor revenue for an episode just for that gag it'd make me so happy but we're not doing that we will talk about deep seek v4 pro anybody who's interested in potentially sponsoring you can hit us up at a nerd snipe at 10xn right nurse type at 10xn.dev thank you Alyssa yeah you should it's great um deep seek v4 pro I don't have a strong opinion about this one I'm gonna be real with you I do all right it is the best coding model with no vision I like can't use my browser because I want to leave the b-roll up nicely explain the other open weight models that are better at coding also have vision mm-hmm the ones that are I put this like how important do you think vision and the ability to perceive a screenshot is for a coding model like 10 out of 10 important like it's critical I wouldn't seriously use a model without it then this is very easy it goes in detail yeah I agree and also to be clear if we're talking about the like ranking if it's the best coding model that doesn't have vision I think Gemini are sorry GLM 5.3 is above it I think there's a better model that doesn't have vision I keep forgetting 5.3 doesn't have vision yeah 5.3 not having vision does suck that's hopefully gonna get remedy but for is my understanding that v4 pro the new snapshot and GLM 5.3 are relatively close in most coding benches I haven't used either for significant coding work so I wouldn't know I really I like GLM models a lot I'll be honest like I have used 5.2 definitely more than you have and I have used 5.3 I've enjoyed both of them a lot and I think they're doing great work there I would definitely put it above deep seek v4 pro I think that the really compelling deep seek model right now is the v4 flash for me Kimi K3 is what I saw others experiencing with GLM 5.2 where I had heard a lot of people say 5.2 could like do their day-to-day code work I ran into too many issues with that Kimi K3 if you had the patience to wait for its relatively slow generation times actually could do real work for me I have not tried GLM 5.3 simply because like K3 was my oh shit it can do it now and then I moved on and since GLM 5.3 didn't have vision I just didn't care that much because K3 did and K3's vision is actually really good it's like spatial recognition and things like it it can get details out of images that most of the other models miss uh yeah yeah I actually think that's true especially post Defcon and how horrible most models were at getting anything out of vision yeah I can see that Kimi K3 I don't know what the fuck they did but they built a whole new pipeline for their visions of I think they had some discussion about it but like it it recognizes details and can like pinpoint pixel accurate like things and images better than most closed weight stuff can right now interesting meanwhile GLM 5.3 you cannot give it an image at all yeah yeah which does suck that is that is a huge limitation and when we get to GLM 5.3 I'm going to put it a lot lower than I would like to then let's just do that now because I think that like the relationship between GLM 5.3 and DCV4 Pro is something I'm not comfortable really commenting on because like to me they're both powerful open weight models that don't you or they cannot function with half my prompts in the majority of my threads like agreed I know people think I'm insane here but like no you're not do you understand how convenient it is to just take a screenshot of an error and paste it in so you don't have to like find some weird way to finagle and copy the text like yeah I is this more useful for front end yeah of course like if I take a screenshot of an error or something like overlaid incorrectly or something laid out wrong in the front end it's essential like if you don't have it there it is impossible to do your job if you don't have it for back end tasks it's still really annoying yeah because I so often just want to screenshot the error or the edge case and not have to like format the text the right way or some shit it's it is so convenient to have models that can perceive images I think it's even more important for back end than you're saying because like even just like you know oftentimes we're both working on huge fleets of like local Linux devices and stuff like that as a staging between them running jobs on other things or even if you're deploying stuff up to the cloud you know like I love herself but sometimes their errors are a huge pain in the ass to copy and it makes a huge difference to be able to just screenshot the one error that it matters rather than having to figure out how to copy either the whole stream and overload the context window or just like finagle it and then format it in a different thing like screenshotting matters a lot I would say at least 50 to 60 percent of my prompts have screenshots in them I cannot live without it and GLM 5.3 as much as I do really like this model it not having screenshots really hurts yep and there's gonna be so much cope in the comments around this like oh just use OCR whatever yeah every single time I am remoting into another computer that has a GUI in it and I hit an error I'm gonna screenshot this like annoying low-res text and open it up in macOS preview and try to copy the text out of it and then paste it in or crazy just weird thought I know I could use a better model no okay well to be clear Theo's being dumb here there are better ways to do this the way I would do it like in my PI config I use a lot of open weight models that don't have cut like OCR or vision built into them and I built out a custom vision tool for these models so that they can read and see what's actually happening like if oh you got an image but I can't read that I'll call this tool to get a summary of the contents I did a big benchmark for myself of how well different OCR models worked for this tool because I wanted to get the best one that could run locally because the whole point of the local inferences I want to keep this local I want to run it on my home network I benchmarked it GPT-56 Luna demolishes every single one of them like the best way to do this extension or whatever is to just pipe it into GPT-56 Luna ask it what's in this image to give it in a structured output and then pass that into the other model that is the best way to do it it's really annoying not having it built in really sucks like yeah no I as much as it pains me I think we have to put this in the same tier as deep seek v4 pro like I cannot imagine putting it above the forbidden model it should be the top of detail yeah yep I agree I that's where I was going to put it I wanted your permission to do it I am thankful for having your permission I also do feel obligated to say that I almost made a pie a finish your drink rule and I'm very thankful I did it no yeah you're especially we get to the harnesses you're fucked oh god yeah it's gonna be fun yeah just wait till I start talking about the deep seek harness you're not really continuously are bringing up Luna mm-hmm let's just get this one out of the way top of B tier top of B tier me actually bottom of a tier I we need enough space for the A and S model so I'll put it top of B for now I want to put two models in S but you're gonna hate that so well we'll get to that when we get to it okay fine I'll put it in a for now okay so we can have that fight when it comes to it yeah okay okay I do you want to tell our fans why you like Luna so much because I I like it a lot too but I think you have more experience with it and have even stronger preference than I do yes it's absurdly cheap it's absurdly useful it is like it's basically having a better GPT 5.5 that you can use unlimited usage on for random miscellaneous tasks I have like our internal system for monitoring videos and tracking sponsors and doing crazy like non-deterministic back-end work is all running through Luna and it does a phenomenal job I'm paying API pricing for it and it barely even seems like I am if I recall you were actually setting this all up before the price drop too yeah yeah it still was basically nothing like it's absurd how cheap this thing is and how much they were able to squeeze out of this tiny little model it's remarkable and that price drop by the way was 80 percent it was already surprisingly cheap that just made it unbelievably so like just like throwing everything at it just for the sake of it is worth it and this thing was a freaking meme they did internally like they just did this for fun they had like GPT 5.6 soul go make a model for you know see if this works and then it kind of worked like it's really good yeah dollar 20 cents per mil out is just stupid it's like it's insane yeah it is the non contributor tier pricing of muse spark and it is multiple step functions above muse spark it's a third the price of the non contributor tier but that makes it six six more expensive than the contributor tier still oh yeah here with because contributor tier first or for muse is 20 cents per mil out and like outputs really what matters in these cases because the inputs are so fucking cheaper like rounds off the 20 cents per mil out versus dollar 20 per mil out with luna but the non contributor tier for muse spark is 450 yeah don't ever use it no never use that yeah muse sparks to be crystal clear on this ranking this is computer tier if it is not going to that number it should be like f tier it probably it's nearing google tier at that point yeah it is nearing google tier on non contributor to yeah for sure it's like a flash was fast and still just as bad i mean i think it's better than flash but you know yeah it muse impressed me luna is really really good for what it costs and what it is i'm also using it heavily but not like in things that i am like building and provisioning internally what we use it for is for everything related to title gens and metadata within t3 code if you have codec set up in t3 code we default to luna on i think medium for yeah every little metadata thing we do so like naming your branches naming your threads all those types of things and i have some features i'm working on now where we will do automatic theorizing around what should happen to all your threads that haven't been settled yet that will largely be powered by luna by default because it's just it's effectively free with how cheap it is and if you have a subscription with codex it is literally working out to free yeah it's incredible yeah it is a very good model yeah it's it's a nice reminder of what can happen if your post training is good and you actually put any effort in to the smaller models speaking of smaller models and effort do we complain about opus now you know i was thinking about this i think it is a cop-out to only have google models and google tier i think we should christen this with a non google model by moving sonnet 5 down to google i want oh no no no no opus i i will not stand for opus 5 being above sonnet 5 or being below sonnet 5 or at the same tier because hear me out here sonnet 5 is more expensive per task than opus 5 for most tasks and opus could actually get some work done unlike sonnet yeah but like sonnet was neat at the time is the thing like i i don't know man for the me like i really want to see opus down at google tier i really do i really want to stick it down there i have merged some opus code that didn't break things i have too and i regretted it it felt bad for myself the code that i merged that i regret is the code where it broke things because opus 5 can absolutely break things and it can over complicate the hell out of everything it does it is the ultimate example of anthropic being dog shit at reinforcement learning rl they are so bad at it it's hilarious and they give so little of a fuck about anything other than their big giant models so they can build their magical god to take all our jobs and kill our families or whatever i don't know what they're up to in there but like opus 5 is not a good model is very very bad yeah which is why it fits perfectly at f tier but like we have a sentimental attachment to sonnet because you were going through it when sonnet came out i was but it's not a useful model and even then it was like not a useful model it was just conceptually it it was like the ghost of christmas past yes it was oh yeah that is a scarily apt comparison for my feelings of opus 5 it is my ghost of christmas past it feels bad to see sonnet 5 and google tier i'm gonna be honest with you i know but that's where it belongs it is on there is no reason to use that model other than you have it subsidized through your codex up or your quad code sub and even then opus is better per dollar if sonnet 5 was dropped to ten dollars per mile out instead of 15 we could and was even slightly more token efficient than it is you like a five percent reduction in token utilization i would consider bumping it okay but it is the sonnet 5 is effectively it feels like they took sonnet 4 and they gave it like some system prompts that like like sub agents are a thing you should know about them yeah yeah that that is exactly what it feels like okay i'm sorry this hurts me but it it hurts me in the growth sense so i'm you know where i put these in my tier list no i i saw your tier list earlier i have deliberately not thought about it so i don't bias myself it'll be fun would you like to know where these two went for now okay yeah we can spoil that one okay opus 5 was top of detail okay sonnet 5 was end of detail here wow but there was like four models between them i negotiated you down here yeah all right so i'll take that the the reason do you know why i had sonnet 5 like higher than we have it here because i have actually found uses for it which are debugging multi-step workflows and sub agents in t3 code with claude code you have five subs just throw this at opus like if you run out of your fable i know it's sad like i was sad about it too but like it's my monkey that dances for me that's all it is opuses opuses the monkey that dances i got a little too close to burning three of my five so for context for those who are not familiar with the chaos i put myself through as fables not number one defender but top three or so at this point fable is a fucking incredible model to the point where i have five subscriptions with claude now i am probably now at the point where i start canceling a few but like at its peak i did actively max out all of them and find myself stuck one of the notable things with claude code subs is that the fable usage can only be 50 of your weekly so let's say your weekly usage was two hundred dollars hypothetically you could do a hundred dollars worth of fable and then you couldn't use fable anymore you had to use other models what this meant in practice is that i would have five accounts that were out of fable but still had 50 of the weekly usage available for opus and sonnet so what this means in reality for my lived experience is i effectively have at any given time unlimited opus and sonnet because these 50 left are just sitting there until the end of a given window then gold bug happened and i just wanted to throw anything i could at anything so i actually managed to burn down most of the second half weekly usages just with opus there so why would i use sonnet because i wanted to ration out my other half of a weekly a tiny bit and telling sonnet to spin up other sonnets as sub agents or workflows to do like basic dumb math functions just to test things and like see the visualization for that was useful ish oh wait you were using sonnet during gold bug no i was using opus during gold bug and then when i got home from gold bug i was nearly out of my weekly usage so i debug workflow stuff with sonnet okay okay it just adds up you have like a thing on your left eye other eye yeah it's still there it's on the like very corner of your eye at the top of your lash or your eyelid yeah that's a that's the i stick caffeine cream on my eyes and it helps a lot yep i just want you to know so that if we have you locked in that that's not like distracting yeah it's fine i like to know when i have things on my face i do too but i've accepted that my face is sufficiently fucked up that like i i do you do a good dermatologist so you did and he is gonna blast lasers into these and they're gonna get fixed the i i see the comments i promise you guys i sleep i can show you my heat sleep yeah the he is not my blood boy no that is not true i would pick a much younger guy i know i would use my blood for far more effective purposes than keeping him alive like no no no can we find a way to like power open weight models with blood i would do it like okay i if that was what i was doing if like i was how much plasma do you have to donate to afford dcv4 flash at home uh none the dcv4 at home is easy i'm talking about like i want fable at home i want soul at home like if i could get soul at home for like a liter a week i'd consider it to be clear i would consider it i would have you welcome to our audience no i would not give it to peter peter has not earned that but i didn't realize you were discriminating against who gets your blood in order to get unlimited fable well peter teal has no ability to get unlimited fable and to be clear i'm not i'm positive peter teal could get unlimited fable if he had any interest in doing such well yeah i'm sure he could i don't want unlimited fable i don't want unlimited fable with their like little safeguards or whatever on it i want unlimited fable in my networking closet running 24 7 doing whatever i wanted to that would potentially potentially be worth it anything else no absolutely not on the topic of fable with no refusals we should talk about something very similar gemini 37 flash oh hell yeah oh flash this one's gonna hurt me more than you i think because i used to be the gemini flash fan boy not like defender but active fan like promoting it consistently 20 flash was cool that was a cool model and thankfully call it 20 flash because 20 flash was unbelievable so well i've learned how much do you know about like the specific detailed price history of the flash models i mean it wasn't like muse spark level just absurdly cheap yeah so gemini 20 flash was 10 cents per million 40 cents per mil out they didn't do a 2 1 or 2 2 because google's naming is bad enough that they thought the anonymous nano banana name was worth keeping yeah they're very special so they went from 20 to 25 uh-huh to five was 30 cents per million and uh i need to find the no reasoning price because this is when they started doing their very fun thing where they would split the price with and without reasoning so they bumped it from 10 cents per million to 30 per million so 3x on input output was 40 cents per mil out they bumped it to 250 per mil out oh non-reasoning output was 60 cents so a 50 increase for non-reasoning yeah with reasoning on it was 250 because it was the first flash model that had like reasoning as like the core model thing yeah and then they decided that was too complex so they just made the price always 250 per mil out gotcha this is particularly absurd because when you have reasoning on you do more output tokens a lot more so it's more cost per token and 10x the tokens which resulted in the your bill for a given task with 20 flash versus 25 flash not just increasing by the token cost but increasing by the token amount so for our limited testing for things like title gen and stuff for like a daily run with 20 flash it cost us like eight dollars moving to 25 flash would have been like a thousand plus yep that makes total sense the the 25 flash was the death of flash it's been a bad series since it's been fun since but yes it has been a bad series and somehow we have continued to inflate the price three five flash made it to a dollar fifty per million do you remember the output price i too much throw a number out uh like three or four nine dollars per million out nine it gets better telling me it's nine if you want it to be region specific so you want to guarantee it's running in the u.s it bumps to ten dollars per million that's why google hell yeah that puts it that puts it above that was it at roughly tara's price hell yeah for a model that uses six x the tokens and one fifty intelligence yeah and remember that whole rant we had at the start of this episode about how like efficiency and price are like different things it's worse than both yeah it's worse than everything this is probably one of the most expensive models on this list it isn't but it is way more expensive per like level of intelligence if you do the chart for like intelligence per dollar and you measure it by cost per token it almost looks okay but if you measure it in terms of the actual cost to do things let me hop over here and switch to here you end up in lots of very uncomfortable spaces apparently three seven flash are like the weird limited set of models i put in here almost looks good it's roughly v4 pro price from deep seek roughly same intelligence level too but like luna gets the same intelligence level on max at under half the price it's five cents per task sorry it's way cheaper than i thought it's the log scaling on a hard for us is screwed to me it's 25 cents per task for three seven flash and it is five cents per task for luna for pretty much exactly the same level of intelligence and the best part is luna gets done faster because luna is so much more efficient that the tps doesn't matter anymore well i think if i'm being honest with you we have this tier for a reason i think we both know exactly where this thing needs to go s tier yeah perfect gemini 37 flash an s plus tier because it has given us infinite content yeah honestly i had no i haven't cashed in on the google actually hot take of the models in this list gemini 37 flash is the one that normies encounter the most by far yeah but that's like more of a skill issue than a like like yeah okay google has distribution the sky is blue very cool can i drop the super hot take that i think the google ai summary when you search something's actually a good feature and usually is pretty useful uh i mostly agree yes and even spicier take i think the google feature like the gemini feature within the youtube studio is really good i really like it the like and the like summarize this video for me is fantastic it's really useful i should have had another tier for ben's to something i hate so much that i have to finish my drink you hate that thing yes it is a great way to just grab like a quick summary of like sentiment from comments and it'll tell you what the like top level like click through watch time all that stuff like i don't read what i don't read the words it says but i just like look at those things because i have successfully hired a team where i don't have to look at those things myself anymore i wouldn't have known because i pay people expensive salaries to do that for me instead yeah no like it is not a if i was actually analyzing the performance of a video i would just go to the actual dashboard and read the data but if i just want to see like a quick gut check after like an hour of a video being live how it's doing compared to previous videos it's quite nice for that i just realized there's a number i don't know i want to see how sonnet 5 on like a lower medium compares in cost 237 flash hmm huh they only have sonnet 5 non-reasoning and max in artificial analysis for pricing they don't have the price for the other tiers that sounds like a skill issue the non-reasoning is roughly the same cost range as 37 flash on high so sonnet 5 with its brain turned off performs per dollar roughly where 37 flash on high does when comparing two models that are stupid that like should ever be used for any reason 37 flash does indeed have a value almost that's a glorious chart uh your screen recording you should just do that so if you go to the deep suite and uh just sonnet 5 and then 37 flash yep yeah that is a a very funny looking chart so if you do that sure can you turn on soul and luna as well just to emphasize again how stupid this is that sounds like cheating but sure once again there is no reason to use gemini models other than memes and apparently internationalization if you haven't already started using today's sponsor general translation yeah please god just use general translation they will make your life so much any excuse to not have to touch a gemini model is probably for the best right now yeah honestly i wish i'd have to say that because again gemini to a flash was such an important model at its time it's my equivalent of ben's grok uh code fast and rested beast as such i am sad about this regression three seven flash is trending in the right direction though it is more token efficient than three six and it's cheaper they knocked it from nine dollars per mile out to i think 750 but i will also believe that's a temporary promotional pricing do you know when that promotional pricing ends one january okay well actually there's a chance like there's a chance there won't be another gemini model by then i know that they posted about the like v4 pro training happening three six is a month ago not even i think like three weeks ago time is non-deterministic at google it could be a month it could be a year you don't know you have to spin the wheel and we haven't seen where the wheel spun you don't know there might not be another gemini model yeah i on that note i do know where gemini three one pro will go in this yeah go for it no i'm not even gonna put it in the on deck section i'm just gonna drop it right to the end of google tier yeah i i think this is the correct ordering in google tier as as sad as it makes me this is correct yeah yeah i mean is there even anything to say like it's yeah oh cool yeah gemini three one pro if you we don't have the visuals available to you because you're listening we still love and appreciate you make sure you leave a review and a rating on our little podcast here it's a lot of work especially all this alcohol thank you alissa for being happy that i did my job for once gemini three one pro just put at the absolute end of google tier not because gemini three one pro isn't surprisingly intelligent especially at the time it came out yeah but because we've been waiting for three five pro for what five six months now roughly oh longer than that yeah it's been a very very long time and i don't like they they are training v4 pro right now time will tell that ends up being good but i do actually train three five pro yeah they did have three five pro yeah it was a disaster like if it was so bad that they didn't release it like i i would love to see that model yeah it'd be really funny yeah flash has become their name for we don't want to look like we're putting out a frontier model because we'll get clown down so hard so we're just going to call it flash instead right like they desperately need a like sonnet or opus tier like in between because they're charging too much for flash like way too much and google can make good efficient models i think they did a flashlight briefly for i think three six that was okay but like we need flash to mean what flash is supposed to mean which is what deep seek before flash means yeah flash is supposed to mean what haiku and luna mean and i mean rest in peace haiku you will be missed we will never see you again we're about to hit a one year since the last haiku release we will never see ah yeah no we will never see another haiku model i would be shocked if we ever saw another haiku model yeah they had to remove like two digits from anthropics valuation for them to care about things that small oh absolutely yeah minimum and like they're i think they treat haiku roughly how they treat their customers and external people they just don't think about them that's a rest in piece what a what a lovely little model that i you were the haiku defender i defended at the time and now i defend luna and i got you to put luna in a tier so i've done my part the legacy of haiku lives on in a tier for what it's worth in my tier list i had the air five six luna in like top b tier yeah okay that's fair yeah so on the topic of these small flash tier models do you seek v4 flash you have a lot stronger of an opinion than i do here cool so as v4 flash is number one advocate that bullied me into getting another spark yeah because it's good it is okay is it at the frontier no but is it at the point where it like kind of feels like an opus 4 6 level model yes it is very good it can run for long periods of time it knows what it's doing and it feels really good when running on to djx sparks locally like you can get about six threads going at once with reasonable tps like ranging between 25 to 35 tps at that range there are a lot of optimizations you can do that you probably get that higher i think this model for what it represents like the problem with this tier list is like we're putting these models in the place that their specific caliber model should belong i think this should be a b tier model honestly i had it a little closer to luna personally but no it is not close to luna in capability i can tell you having done a lot of like actual real work with this thing it is not at the level of luna i wish it was but it's not it is more capable than you would expect and it feels but okay its voice is not great it's a little weird to talk to it is very okay with doing literally whatever you tell it to do the fact that it can run on two dgx sparks which i think is like about the top end of even remotely reasonable consumer hardware like 10 grand for hardware is still an absurd amount like i would not call this consumer hardware but it is potentially accessible hardware when compared to like an rtx 6000 rig which will run you 50 to 100 grand like this is the model that could potentially be run at home for someone who is employed in tech and it is currently the best model for that so i think that's why it deserves b tier but beyond that it's hard to justify more i think it would have been a if it had vision i do know that there's the v4 flash vision experimental version that just dropped like right before we were recording it's been out for a few days by the time you're watching it is not open weight as of recording time so the benefits of v4 flash being a thing you can run on your own hardware not there i'm currently in a flame war on twitter about this because people are saying i am so terrible for ranking it low for not having vision no it just if you're running it yourself which is the main benefit like i yes i boot bumped it from b to a in my list because you could run on your own hardware and i felt like that was worth giving it extra points for but it also deserves to lose points for not having vision if it now has vision cool but if only is vision over api no it cancels out like that's the whole reason it was higher in the first place is that and if i recall a lot of why you did the benching for how to get ocr out of these images for models so that you could have yes so could you reasonably use v4 flash the ways you like to without luna as an ocr tool for it you could but it would be it's pretty nerfed and it the unfortunate side effect of this is in order to run it on two sparks you're going to pretty much saturate the memory on both sparks like those sparks are fully being utilized you really don't have room to stick a decent ocr model on there as well so that means that you need to stick that ocr model on some other local piece of compute you have if you want to keep this all fully locally it's not there like it is not i i don't think you can justify this quite at the tier of luna despite the fact that you can run it locally which is really dope and it is the best local model we currently have it's not there unfortunately i wish it was i wish i could put this thing at the top of a tier but i can't this is funny because i thought you were going to push for it to be higher and that i'd have to push for it to be lower i am this makes me feel better about the flame wars i'm currently getting in from my tier list being public so all right don't worry the flame wars between us are coming they are as soon as we put fable five in the on deck yep but before then when you talk about our friend ox alpha yes this is a weird one because we don't know for sure where this is going to actually have to land because there are a lot of question marks around how big is it how where can you run this thing like there are rumors that it is a flash model and there are also a lot of rumors that is a very big model the current consensus is that it's a glm model and wink wink as a result we don't know where they should go like muse spark makes a ton of sense and b tier with the contributor pricing i think this is a similar case where if this ends up being a flash model easily in a tier easily not even close but if it is not a flash model if it's like a really big model that you can't run anywhere else i mean it's drf yeah i think like if it's cheap enough a c tier could be reasonable if it ends up being bigger but uh if it does so happen to be i don't know like a glm five three flash or something yeah hypothetically hypothetically a tier i think makes a lot of sense i agree so low a i still put it behind luna just because like the price of luna's absurd but once it's officially out we have real numbers for and see what we can run it on if it is open weight yep it'll be a fun one exactly if this can target like the deep seek v4 range of size then yeah we've got something special here and while we're on it the last open weight we have in the list of kimmy k3 indeed our final open weight contender kimmy k3 i think we might disagree on the swings i i yeah i agree go ahead yeah so i i'm a kimmy defender i have been for a while they have not had the touch they used to with the quality of their writing because kimmy k2 was like uniquely a good writer and since then they have been trying too hard to make it autistic so it codes better and in the process made it worse at speaking who would have thought but kimmy k3 still has novel like taste it is still one of the best if not the best 3d models you can use right now for like creating 3d assets and doing things in a three-dimensional environment i've seen do things none of the current crop are capable of yeah opus can the opus cannot do as good with the fish slop 3d kimmy k3 finished it faster and had a better output than opus 5 did for me okay yeah i i have found it to be surprisingly good at 3d stuff in particular it's just it's strange that like an open weight chinese model isn't just like throwing punches at frontier there are things that is better than frontier at yes it's a small subset it is a small subset and the fact that it is open weight does give it points because you can supposedly run this thing on your own you have to be a big company you have to have a lot of gpus but if you have those resources it is possible to run this thing on your own i still personally think that this is especially when you factor in speed efficiency how it feels to use and work with day to day i would personally put it in d's here between glm 5 3 and deep seek v4 pro i think it's better than v4 pro and it is worse than glm 5 3 i don't know if i just haven't used glm 5 3 properly but i still found it to be meaningfully better overall it's it has more design taste it oh oh no no no i know i have the win it has vision yeah i you beat me i was literally about to say that word taba d here put it there no i'm putting it below c i i'm putting it behind grok 4 6 i this depends on how much you value open weight i think having a model that is almost inarguably better than grok 4 6 it is just much slower and more expensive but it is open weight so this comes down to like like can make a 3 is smarter than 4 6 damn yeah these are the mid pack models for sure i think these two should be our seat here it's just the question of which one goes where this entirely depends on how much you value open weight if you think open weight is really important can make a 3 definitely goes above 4 6 if you don't care that much to paying api prices anyways then it goes behind because also can make a 3 requires if your business is making over 10 million revenue that you sign some deal with the moonshot guys and that deal seems to require you to match their api pricing so any potential cost benefits that you would get from k3 being open weight and the competitive hosting environment yeah just do not exist with it yep but if you're open weight on principle because like you can download it and hypothetically run it yourself yeah but hypothetically is carrying a lot of weight there like i value open weight when it is like my bar for open weight right now is like two dgx sparks like if you can put something on two dgx sparks that is like really good open weight if you have these really strong contingencies on like you can't run this anywhere else it like glm53 have being open weight is neat and i'm glad we can run it on us like neo clouds but the main reason i care about that is because i don't want it to be running on chinese servers if this was if this was an american lab i would not really care if it was open or closed weight i'm going to be so this is one of those places where we disagree i care more about open weight in the sense that like a normal person could run it or in the sense that it creates a more competitive pricing environment it does neither of those yeah and this is why i don't care about the middle tier as much things that you can run on two dgx sparks like i don't think the person who is willing to spend ten thousand dollars to own hardware that will be out of date in two years to run models that are worse than frontier at a cost that is higher than frontier and having less parallelism as well i don't think that is a particularly serious space it is fun and interesting entertaining but like i it's just it's not valuable it's it's fun it is fun but it is also it gets points because i'm putting the like threshold of semi-reasonable for someone who is in tech to potentially get access to this it ten thousand dollars for this compute is not reasonable for 90 percent of people it is not reasonable for basically anyone from an economical sense of like making back your subscription money that's not a real thing but it is possible to get access to once we cross into the like you're buying a car level of money it's silly and it just doesn't make any even you can get a nice car for 10 grand not a like a phenomenal one but you can get a reasonable car for 10 grand yeah they use car markets gotten bad but it mostly agree like it is technically possible if you're willing to compromise a lot but like again like the the ox alpha is up here because we are potentially hoping that this will be a smaller model based on rumors that will be able to run on something like the sparks deep seek v4 is up here because it can run on something like sparks and it is more accessible and also reasonably capable kimmy k3 solves none of those it's inefficient it does have vision so it should be above five three glm five three i agree with that i think bottom of c tier is correct my one pushback here would be that the the benefit of open weight for me isn't like the the hobbyist can spend a bunch of money and run something worse than fable for more money than fable at home i think that's like a fun thing but i don't think that actually meaningfully improves the state of the space or the market at all it's more like the the hobbyist thing it ready for a real throwback it reminds me of the framework quick what is that you you weren't around for quick no qwik the guy who made the guy who made angular felt bad these angular was really bad the web for i thought you're talking about framework computer you mean framework the like js framework yes quick oh oh oh i that is i don't like the comparison because it's true but continue yeah the the point i'm trying to make here is that the quick was a framework that was built because the guy who made angular felt bad about how much slower the web had become due to his creation so he went to an absurd extreme of what would it look like for web frameworks to effectively operate as though they had a hibernation model similar to like windows and other operating systems where they would take the state the pages in store it on the server and then embed that as an html comment on the client so it wouldn't take as much compute to get the page interactive again conceptually incredibly cool for hobbyists enthusiasts and people who care a lot about like every edge being smoothed out and as powerful as possible unbelievably awesome no one actually uses quick no one actually cares it is just a like cool fun hypothetical people will use on side projects like dicking around of the code you ship how much comes from an api from a sub versus from your deep seek v4 on your local hosting i mean 98 v2 at minimum i would argue probably closer to 99.9 v.1 percent yeah it is it is not even close like i i have used deep seek for a lot of real stuff but is it the stuff that i'm putting in my actual projects that i put out there no similarly the most popular and prominent users of a framework like quick had real jobs where they were using react or angular or something else yep that is where i feel like models like deep seek v4 flash and these like runnable on two sparks not one but two sparks type models go because as soon as it is two sparks that means it's not running on your macbook that means it's not running on actual consumer hardware you are now buying a thing just for fun you can run ox alpha or sign on ox alpha deep cv4 on one spark but it is not a great experience like it pretty much the problem with it is when we've talked about this before when you're running models locally you don't have cloud provisioning and as such you can't run multiple threads at once if you have two sparks the reason why you want two sparks is because that allows you to get up to like five or six threads you have just one running you can get it going but once you get to two or three threads it dies it is down to like five tps so to do anything even remotely real with it you need two sparks and yeah it's not great yeah case in point the thing i'm trying to say is there's a difference between a tier of model where you can run it on hardware that has value outside of inference yep and my one dgx spark i did set up has been sitting idling for when was rick miami april yeah april late april yeah four and a half months it's been sitting in a closet doing nothing because the only purpose of that box is inference yeah so if the role of a model like v4 flash is to allow you to spend money just for everyone's on your local network and have fun with that awesome cool that is a toy that is not a useful thing and the reason i like tierless like this is the actual like value these things bring to the market i advocated for deep cv4 flash and the potential flash version of ox alpha being up in these tiers for the reason that i do think that they do have actual economic value the fact that i can poke this at any of my like custom projects or internal tools and when i tell you know soul or fable to do everything you can to break and hack this thing the way you know normal end users will be with api usage of open weight models that have no restrictions on them i can test that locally and i can do that all on my private server and i can ask it private questions and it's not a big deal because it's not going up to the cloud that's a real use case that does actually have some real value that's why i should put them in the higher tiers yeah exactly yeah but that's like the end of that value it is that is where the value ends we are not at the point where when i am sitting down to do a real piece of work that is going to ship to end users and i have to decide between using my 200 a month codex sub to use 5.6 soul on high reasoning with basically unlimited tokens versus spend ten thousand dollars to run deep cv4 flash on max reasoning with a maximum of five threads i'm going to pick soul every single time i'm with you cool the reason i went down this whole rabbit hole is because i personally believe the other ends of open weight are where the value is open weight that a consumer could run on things that they bought for reasons other than hobbyist inference like you cannot run v4 flash unless you bought hardware to run ai on yes if you bought hardware based on any need other than ai running there is no chance of running v4 flash now you bought hardware because you're interested in the idea of running ai locally you might be able to run v4 flash maybe that's the gap kimmy k3 does not fit on that side nope it's on the other side which i also think is valuable yes the idea of open weight models as a method of distribution lowered access restriction and most importantly a competitive pricing environment is very interesting to me the idea of a large open weight model that a dozen to a hundred companies can have yep that can fight each other on pricing and like optimizing their kernels doing everything they can to make that model as good of a value as possible genuinely interesting and exciting to me yep kimmy fails in all of these categories because of the license yes the the license is damning that's what hurts if kimmy k3 could be priced however efficient you can make it as the host i would comfortably put it mid a tier but due to the license restricting how expensive it is to run making it guaranteed the current price which in the real world 15 per mil out for a model that capable sounds awesome until you go back to our previous discussion and you see how inefficient it is it ends up being more expensive for real world tasks yep than solace yep it does so unless you are doing bespoke 3d shit and need that slight improvement and you're willing to eat the 2.5 to 3x higher time to completion with a 15 to 20 cost increase as well not worth it again depending on how much you value the open weight and what that allows in terms of sovereignty and customization because it does not help with cost no this model being open weight does not benefit you with cost at all no due to the license for that reason we are measuring its open weightness purely in what it allows in terms of customization and its unique value in terms of where it is better than the frontier because it is not better in cost it is not better in speed it is better in freedom as long as you're not selling it right i think that puts it c tier but i would fight i could see myself fighting for b simply because i think the 3d capabilities are really cool and having a model that good with all of the things we needed to have open weight is valuable 5.3 can't be that high simply because it doesn't have vision which throws it okay i have to put this above glm by three i agree with that i cannot put this below grok 4.6 i cannot above grok 4.6 you mean or sorry yeah i cannot put it above grok 4.6 that's three times you said grok me sorry yeah an anonymous model made me say that i don't know if your version of this allows me to choose the order of things so i think that this will show kimmy k3 as ahead of grok 4.6 yeah it does rip so the the tier list you are seeing puts kimmy ahead of grok 4.6 no matter what i do in reality it is behind i tried this too it doesn't work oh that's annoying something about the id whatever you get the idea so far we don't have any problems with that it's whatever you get the point it i honestly to be super real with you i would put kimmy k3 and grok 4.6 in a very similar level of capability i think that i would put the like they're very similar yeah if i had to pick one for daily uses like my main model i would pick grok 4.6 but if i had to pick one that i reach to when i have access to frontier right kimmy k3 brings me something that the frontier doesn't grok 4.6 brings me frontier at a slightly cheaper price and worse capability it's not a cheap well it is and it isn't but it is it brings you frontier at more effort kimmy brings you something different so actually honestly looking at this as much as it does hurt me i am okay with putting kimmy in front of grok 4.6 i pulled him over guys it it hurts me deeply but i am okay with this oh yeah i'm looking now at artificial analysis numbers i do not realize grok 4.6 is more expensive on x high than 5.6 soul okay yep no the the bug in this tier list was correct the the the weights were right once again like bed was right once again yeah well yeah four we had according to our assistant thank you alissa ben just said grok four times i've seen two of him one yeah two god i'm so thankful to do the harness to your list after this now yeah okay yeah i hate everything okay speaking of things i hate we should talk about tara before we get to the frontier war and have our flame war yeah let's let's do tara really quick because i think we don't need to talk about this one too much can i place it somewhere and you tell me if i'm right or not yes i moved it to f tier that's a little mean it should be at the bottom of d tier fine these people prove pro moved behind it because of the ordering and no actually once again i'm okay with this okay if if we treat f is forgettable tier i'd put it there because like it's just that luna max leads in really nicely to soul on low if you look at a chart of like price to performance so much sense both luna and soul are great models that have a very real purpose i don't know what the purpose of sol luna got an 80 price discount soul got a 20 with some platforms offering a 50 percent tara got nothing yeah did target a price discount i was looking at open router usage which is not representative of the full industry but it is i think its usage is the least out of the three by a lot tara had a 20 price discount when luna had an 80 percent it terra has become open ai's forgotten tier already right after being released similar to like sonnet tier at amperopic it just it's there because it should be not because it's a good option well and luna punches above its weights for what it is and soul is surprisingly cheap due to the efficiency yes tara taurus is like this awkward in between there and it just doesn't make a lot of sense and also especially now that like aster is on the horizon we know that's coming and we know that's going to be above soul on the like list of open ai models now there's going to be four why would you ever use the second to like the second to the bottom like so i mean this is the line rule if you're going to buy wine you don't buy the cheapest you never buy the most expensive you buy the second cheapest often the best value is it the best value though it's not it's because it didn't get the 80 discount luna saw if the original prices were where everything stayed maybe maybe luna seeing an 80 price drop putting tara at i think like luna's a dollar 20 per mil out or so tara's 12 per mil out yeah it was originally 10x it was originally gpt 5.5 pricing and then they reduced it by 20 and it's still just like i don't see the point yeah and again like to be clear in these charts like we're not saying like this is how smart the model is or anything like that we are measuring these by how good of a value they bring people in a market with a lot of options and a lot of prices yep the problem with tara is that it's not bringing anything valuable to the table yeah at its current price if it was like 40 cheaper than it is right now i could have a conversation about it but it's expensive enough and inefficient enough that it is like if sonnet 5 wasn't so cringe it would feel a lot like tara yeah yeah it would so sonnets cringe separating it by two tiers i think is fair i feel like five six tara in d is totally fair i i'm very okay with the ordering we have right now awesome yep anything else in this chart you want to fight before we get into the flame war no i think this chart we can both agree on and i i believe this i would add another tier near the top is the only difference i would change but that's what this next discussion will be about i making an s and an s plus tier i would be amenable to but we can't do that because we are dealing with the constraints we have and thus we're going to have to fight about what goes in s plus tier because i think only one can live up here i guess we need to talk about fable that we do so i need to disagree with ben's earlier statements can you once again describe your perspective on the idea of like effort ranges and how those apply in your day-to-day dev work when you are working on code there are there is a spectrum of the amount of effort you need to put into your prompts and code base and skills in order to get competent output at the very bottom of this range like basically negative effort is fable you tell it i want to add button to page it will add button to page in the place that you imagined it should go and it will naturally do that in a beautiful way if you tell soul to do that it will do that but it might put it in the wrong place you have to be very specific that you want it in the top right hand corner if you go down a little bit further you end up with grok where you have to be very clear about you want it to be in the top right hand corner and you want it to be rounded and you wanted to use this font you go even further down and you end up at the open weight bottles then you keep going down further and further and further and you get to gemini models that's what i think about these things my one pushback there more than anything is that with kimmy k3 it is at grok 4 6's tier for doing it right i haven't used it enough it's it's slow which makes you impatient waiting to see it do it but it does use like the right font and rounding and like follow the system that's around it to be clear when i was saying open weight i was talking about like open weight as in you can run it locally okay i was colloquially putting like you know putting kimmy k3 and grok 4 6 in a similar like bucket on this okay i would do that okay then that was my only pushback there yes it's like kimmy k3 is a frontier that you can download which is in and of itself really incredibly cool yes absolutely i i again i am the kimmy defender i am providing my role here the pushback i have is not in any of the things you just said it's in how it plays out in reality i work across a lot of projects you do those projects vary wildly in how well they are already tuned for llms the thing that makes fable feel almost like magic to me is that it behaves good enough everywhere i throw it if i'm working in a project that i have carefully fine-tuned to be incredible for agents fable will work slightly better there if i'm working in a project that was not tuned at all for agents fable will still work very well there soul has behavior that is entirely determined based on how good of an environment it is in both in like the harness it's in and the things that are set up there as well as the repo it's working in and the project is working in and how well tuned the setup there is for what soul has to do i think if you are working in one project that is yours that you own that you control all these layers of and you can allow the model and you to build this relationship in the code base to tune the code base to the needs of the model and the harness and all these other things and you can build this harmonic relationship between them that the gap between these models is much smaller the real world does not work that way the majority of prompts going to these models for code yes are going to projects with more than one contributor that have existed since before these models came out if we're looking in terms of the future we're going towards where code is on demand and the repo is more a set of specs and guidance rather than the actual code that runs sure similar capabilities here between of course obviously you can guess what we're comparing your fable and salt the world as it stands today fable is much better for operating within okay that's where we're going to disagree we are going to disagree on if you want my take of where i would put these on the tier list i would put both of them in s plus tier and i would put fable at number one and i put solid number two i would have s and s plus separate if putting them in different tiers feels wrong to me it does not feel wrong to me when you're in code world i okay if we were talking about this purely just in writing code for a code base i would give fable a tier outside of soul if we are talking about in the general usage of how i use models day to day i would put them in the same tier i cannot in good conscious put soul and fable in different tiers based on the way i use llms when i am doing things like a hermes agent or using a model for anything other than writing code and putting up a pr that i intend to merge i actually it's not just like soul is equivalent to fable i prefer soul the table hard agree soul is not as good a model though soul that is true fable has a capability level that deserves its own tier when anthropic describes models as mythos class that is a thing that means something it does and open ai has not released anything that is mythos class no they haven't astor might be it might be fable is the model that i can in one shot paste a screenshot of a tweet where somebody reports a bug and the pr it makes is ready to merge that is a tier difference to me regardless of how the code base is set up how many skills you have set up to allow the model to succeed and this is like the point of models isn't when you tell it what to do it does it it's when you show it what is wrong it figures it out oh i agree the the the discernment that fable has is unique there is nothing on the market right now that has the discernment of fable that is a unique capability it has soul is a workhorse that is the probably the best workhorse of anything we have right now actually when we're at defcon i use more fable more sold than fable it wasn't even close and all the solutions that i got and derived from models came from soul every single one that checks out yes fable never derived any useful solution taste is not useful in puzzles no it is not it is a brute forcing function and the brute forcing function of soul is remarkable i i will continue to be haunted by the rottweiler versus wise owl comparison i think that is the most apt anyone's put it and like that was like week one of soul someone said that peter i believe was his name i put the tweet up if you have it the description that i saw that has haunted me since is fable is a wise owl it will have a deep conversation with you it will make fun of you throughout it it is clever and will try to cheat and work its way around things soul is a rottweiler that will grab the problem by the neck and swing it around and choke it out until it dies and soul will do that and that is a real difference and a lot of the work i do i prefer soul there's also little things like ios for some reason soul is so much better at ios dev than fable is the i think like 98 of the code in the t3 code swift ui rewrite was done in one thread with soul i had one sub thread i spun out with fable asking it to find some performance optimization and found like two small things i think there's like 100 lines of code everything else in there was one thread with soul other than contributions from some other people that were also using soul okay i i think i have the disagreement at this point i pretty much agree with everything you said the place where we disagree on this is the difference in capability because i think and i'm pretty convinced of this the like fable is better than soul but the difference between caliber and like quality is not as big as you think it is i think that they belong in the same tier i you the most fable usage you have gotten is in shipping crazy amounts of features to t3 code and you've done a wonderful job of that you have made t3 code much better but a lot of the work you have done there a lot of the reason why you can just put up a pr like tell it go add this thing into t3 code and it just kind of naturally works with fable is because julius has done months of work with sol reading the code base setting it up in such a way that it is beautifully architected to add new features and i don't think the difference is as big as you say it is i think fable is a like if we were grading these on an out of 10 system it is a 9 out of 10 versus an 8.8 out of 10 and they both belong in the same tier if we're calling fable a 9 i can't get 560 more than like an 8.5 i i do think the difference is bigger than you're giving it credit here there have been too many times and okay the thing that i disagree on is that we're not talking about the downsides of fable i have a hard like i hate talking to fable i hate reading its output i hate interacting with it i don't hate interacting with soul this is what makes it tough is that like do you agree that there are very few things you can do with soul that you couldn't do a fable uh no i am when i go to talk to fable i feel like i am walking on eggshells and i get constantly rerouted and dealing with all this garbage when i just want to have something done instantly that is like take this file put it on my other machine work with my network manage this machine i always reach for soul when i look at my daily token usage on my like my custom token tracker it is 80 20 soul fable a lot of times like 95 okay i sense my urge to rush the conversation because i have to piss i was going to piss after we finished okay i want you to try to win over the audience without me here while i am going i'll be right back okay fair enough i have to give a sermon on why fable sucks okay my pitch on why fable is just how do i put this fable is not the model that i reach for by default when i open up i wake up and i need to go do some code work i have two options i can open up claude code or i can open up codex i have to decide that basically every day in nine out of ten times i will pick codex does it require me to think more about what i'm doing yes but at the same time i feel much more comfortable with this model it is very good at attacking very difficult problems if you give it even the slightest amount of direction if you're very clear about okay i want this this and this and your code base is set up in such a way that this this and this makes sense to it especially something like t3 code there's not really any issues it does the job very very well and when i'm working with fable there are so many weird things with it like the writing style is awful the output like the writing style is just awful i hate reading the things that this model outputs i don't want to interact with it it takes a very long time to get anything done the code right out of the box if you give it no context you're just like okay go implement this new code base from scratch it will make a better looking code base and it will make these better decisions that is not what the real world looks like the real world looks like you have a real code base that has actual opinions in it and you are iterating upon this base foundation or you are iterating in working with some system or network that you have already established and worked upon soul is infinitely better at that kind of thing the model i want to work with every day is soul and i have a very hard time putting that five percent like it's not five percent it's probably a little bit more that 10 to 20 percent better discernment of fable putting that substantially above five six soul just doesn't feel correct to me i think these two are neck and neck and with an experienced good operator behind either of them you can get the same level of quality out of either have you finished have you finished your sermon on fable i have soul okay i have finished it i don't know anything that you said but i'm going to counter regardless okay how do i start this you've been gaming for long enough that you were around for the wii and the wii u in particular era i actually liked both of those a lot i had they were great systems they were fantastic they really were were you conscious when skyrim came out no okay i'm a small child the thing that made skyrim different is that you would look out over the coastline like over the like place you're currently in and see other things in the background and you could go to them and they were there and they were real you can make something that looked a lot like skyrim on the previous generation of hardware if you made those background things static pngs that were just there in the environment doing static backdrops for the world made it look and feel bigger and deeper than it was but it wasn't this is the difference with fable for me uh yeah one i don't think the difference is that big too i think that you are discounting the fact that like if to take this analogy to as extreme you're sitting here you're working on the ps2 it's all fine it's great you go to the ps3 the ps3 has a loaded gun inside of it and every time you say a naughty word it like shoots you the point is fable is a very unpleasant model to work with and i think that matters i think the unpleasantness of it knocks it down from being in a tier of its own are you saying unpleasant like the refusals or the tone or everything but the refusals the tone the clod isms the 50 of your sub counts towards it is like i think that's the biggest argument you have is that you can only use half your sub on it i mean it takes like it was the first model that made me get multiple subscriptions for the same provider i mean i the only reason i had two clod subs is because i wanted to have a ab test against soul for like open sauce and stuff like that but since then i have canceled it like i i don't have two subs anymore i don't want two subs anymore when i sit there and i work on it i am using fable and getting through that 50 feels like a chore i don't want to use it i have one last point and this is the hill i can like let this die on okay remember earlier you said that the reason fable can do such incredible code in the t3 code code base is all the work that julius put in with soul do you know which prs are the ones julius is the most mad about no none of them are the fable ones he has complained a lot about the theo slop that has snuck in every line of theo slop he has been upset about every single one and i've taken the time to sit and go through them all so i can tune the repo accordingly to prevent it every single one is a pr i merged because soul wrote it what were these prs what were they touching these are prs for things like updating the relationship between the client and the server for how data is being passed prs that are exposing new functionality or prs exposing the functionality over the server prs that are fixing one-off ui regressions prs that are tuning the sub agent views and capabilities there soul writes unnecessary interfaces yes soul creates abstraction functions to do typecasting where no typecasting is necessary yes it does soul loves as that's one of its favorite words yes it does soul writes typescript code like a python dev that is forced to work in a typescript repo the the python dev is the best argument against soul my argue the thing against all of these two of those things you put in there like architecture level decisions and api level decisions those are decisions where if julius was writing those with soul he would then go look at the code and deal with things in there and clean all that up if you are just if you are shooting into the void and you are just like all right we're gonna do this we're gonna have the bots review it we're gonna have the sub agents review it and see what comes out yeah fable makes a lot more sense for that if that's the way you want to work with these things it isn't as subtle of a difference as you seem to be stating it to be here though because like with with fable i can make a pull request with the changes that fable made and if it passes the ai bot reviews and it does what i expect i can hit merge and no i won't get bitched at if i use soul for the same thing and it passed the ai reviews and i use the feature and it works there's like a 50 plus chance julius is going to come in and be like why the fuck did you merge this yes it yeah again this is why i want to put them in the same tier and i want to put soul behind fable i agree with you it is not at that level where you can just do that and it will naturally get the right answer but the thing is the other the other parts of this thing when you look at the whole picture of the model when you take the thing in its entirety it is very hard to separate these two by that degree to where one is an eight here and one is an s okay i'm gonna push back a different way now okay we have grok four six and c tier as well as kimi k three and we have glm five three and d tier as well as tara yep do you think the gap between the c and the d tier is actually bigger than the gap between fable and soul i mean if you put it that way to be honest with you all i would do is i would put no actually i do think it is because of vision i think tara belongs where it is v4 pro belongs where it is glm 5.3 it should by rights be at the top of c tier if not bottom of b tier for what it is but since it does not have vision and has these usability quirks i you just kind of have to put it down there there is a very measurable difference between the c and the d tier like i if we're a lot of what i'm thinking about with the tiers here is like how much would i want to use this model luna is an a tier because i get real usage out of that thing and it makes sense for my day-to-day workflow ox alpha is in a tier because it makes sense in the interpretation of it that we have that it will be a useful model for day-to-day work i would not put a that big of a difference between fable and soul especially like comparing between luna and flash that difference is not nearly the difference i would put between fable and soul i haven't used you for flash enough especially now that it has like the v4 flash vision experimental version if that was open weight make this harder i perceive a slightly bigger difference here where like if i am just talking to someone and recommending a model like what subscription should you use like get the codex subscription if you're only gonna spend 200 bucks a month on a sub do the codex one it is a much better value per dollar for what you're getting there especially when you consider the 50 percent like limit on the fable one if i'm talking about what model do i want contributors of t3 code using it is fable yeah but that is that that is like because the contributors of t3 code are just by the law of large numbers going to be inherently lower value the way i am looking at this tier list is what would i recommend to someone who i would put at the level of respect of someone like julius if i started julius what would i tell him about these models and what would i think about these things this might be where we like genuinely differ i think if hypothetically speaking julius was to go all in on one of these two models julius would be more productive and ship more better code with fable than with salt i disagree we'll talk with him when he is back but i am nigh positive on this that like the the top tier of contributors is why i brought up ryan carneata before that tier feels the gap more than normal people do you seem to be pushing that they feel it less and i just wholeheartedly disagree there i know my my argument is that the models are not good enough that the i agree there's a gap there is definitely a gap there my disagreement is in a how big it is and be the holistic look at what the model does and where it functions within our day-to-day work i can get what i get out of soul with fable with more patience i cannot get what i get out of fable with soul with more patience i have to like put in massively more effort to get close and it will still miss things like like fable will correct my bad assumptions soul won't fable makes me learn more soul applies my learnings soul is more of a tool i agree it is much more of a like you need to know what you're doing and you need to have a direction and a goal and what you're trying to create with the model but i would argue that in the current state of llms where they currently are you pretty much need to do that if you're building sufficiently complex software you need to be putting this this is where the difference is though fable is the first thing that like similar to how like 4-0 was the first model to cross the line where you could have a conversation with it fable is the first model to cross the line where you can treat it like a co-worker and trust it to figure shit out i i think you can trust i i wholly disagree on this if you're about to say that soul can be trusted to figure shit out no oh you can absolutely trust to figure shit out the problem is you can't trust that what it figured out is going to be in the shape you want it to be it is going to create a solution they will both get the right answer the quality like the how nice the code is is going to be far better with fable versus what it is with soul and you can beat that quality difference out of soul much easier than you can beat the horrible experience out of fable i can't fix fable to be pleasant to deal with fables more an acquired taste i will absolutely agree they like if you if you don't like how fable talks you're screwed if you don't like how salt talks you can adjust it this is like again like the difference between hiring a person who can be in the room with you versus automating the work someone else did almost or like soul really is a tool and it's a tool at its core fable is the first model where i can treat it like an intelligent co-worker and get reasonable responses the same way like i would with like a new hire where the new hire might have these like niches of experience where if i give them a thing they surprise me with how good it is and occasionally surprise me with how bad it is but they're a person on the team and a person on the team has these benefits and negatives inherently and that's kind of how i think a fable is like can i like i can throw an incomplete thought at it and get back a useful thing if i haven't finished the thought i don't bother hitting up soul i've spent a lot of time planning and coming up with ideas with soul and it can get there but it takes longer like i i think i think fundamentally what this comes down to is just the the amount of i think we just perceive different amounts of difficulty in getting good things out of soul and fable because i for me and the way i experience these things i have a worse time getting good quality out of fable than i do out of soul but i am fully willing to admit and recognize that when i have it write code and i look at the code i ab test these two functions the function fable wrote is going to be better if i don't have something set up that will tell soul the way it's supposed to work it's going to infer that it should have 17 000 try catches nested within each this would be a really good argument if we were so prompting on that level like i want you to add this function in this file that has this content in it i would be interested in this argument when i'm prompting fable it's much more like here's a screenshot of a tweet of somebody reporting an issue how do you think we should solve this and i'm like okay cool sounds good go do it and then there's a pr op that's ready to merge that fixes the problem this person had on twitter i've done stuff like that and i do a lot of stuff like that the the difference is just i i i do more back and forth with five six salt because once you get the plan put together and you have the constraints put on five six salt the output's going to be pretty similar between the two i don't really have much like the way i have my skills and stuff set up and i'm like okay let's do a plan it'll invoke the plan skill and that involves lots of questions and lots of like multiple phases of many questions to get on the same exact page of exactly how this is going to work i use that on both fable and gbt five six salt and when i'm going through this battery of questions i much prefer the questions and experience of working with salt i don't like working with fable on these things can i sponsor plug a sponsor that's not sponsoring this go for it i want to talk about macroscope for a second okay rose aren't familiar macroscope is an ai code reviewer one of the things that they do really well is they let you define in the code base itself a markdown file for different reviewer sub agents that check for and review specific things we've had a couple problems in the t3 code code base that we wanted to verify that weren't really lintable things like are you following the design system that we have already built are you following effect best practices like you can't really lint rule those things unless you know every single thing that models or whoever else contributing will do wrong yep this has made me so much more aware of the shortcomings of open ai models because i almost always when i put up a pr with anything resembling a meaningful change with soul fail one or both of those checks fable never does again this is why i i do want to put fable at number one and i agree that makes sense i'm with you i just don't think that the difference is that big i i'm arguing the difference in what these represent and to me what they represent is fables a model i can kind of trust autonomously for the first time where i can give it the problem it gives me a solution and then i see thumbs up and i hit merge and i don't have to think about it beyond that i am hesitant to hit merge on a soul pr until i've actually read the code because it's so it's just such an absurdly high percentage chance that there's some fucking slop in there all right i will fully agree with that statement and i will basically put my closing argument at if we are doing this tier list purely off of code merging into real code bases that are shipping to users i am okay with fable being in a tier of its own if we are talking about the holistic experience of the model from you know the way we interact with our computers these days the computer use stuff the just general interaction the feel of it the cost the usage limits all that stuff i have a that gap closes a lot so i'm just going to place things in this accordingly so we currently have now s plus as fable and soul a is luna and ox alpha alpha hoping that ox alpha is a flash model wink dcv for flash and you spark one two contributor tier as b tier kimmy k3 and grok 4 6 is c tier d tier is glm 5 3 then tara then dc v4 pro mostly because the v4 pro and glm 5 3 don't have vision f tier opus 4 5 because it looks like it's going to be a good model then isn't and then google tier as 3 7 flash 3 1 pro and of course sonnet 5 which is effectively a google model at its current tiering the pushback i would give is if this was an ideal tier list that perfectly matched my feelings and to try to integrate yours i'd have s plus is fable 5 you're beating me to it yep s is soul yep blank row yep and then every another blank row then a tier is luna and ox a minus s minus and a plus should be empty five six six should be s s plus should be fable five and then everything else should be yes because i i feel as though a tier list with one s plus tier that has both of these does not properly emphasize the gap between soul and literally everything else i agree there are two god models it is fable and soul yes these need to be there is a a meaningful gap like a large gap that we agree on then everything else yeah soul and luna being this close does feel weird i'm with you there yes so hallucinate for us in spirit of open ai models hallucinate some rows between these and you'll understand where i'm coming from and it's actually actually another point i want to emphasize fable hallucinates much less than gpt models and five six soul got worse and hallucination not better it's part of their rl pipeline to make the model unblock itself it hallucinates aggressively uh sold us yes um my my hallucination bench is defcon puzzles and i'm noticing very similar levels of hallucinations between the two i did not run fable much against the defcon puzzles because gold bug came out at a time that i was low on fable because i was doing a lot of actual code because again my use of fable is real world work which is very reasonable but i don't have a product with is this public 200k users like yeah it's very public cool so i don't have a product with 200k users so it's not as important for me my main focus that weekend was just the defcon puzzles and to me i wanted to ab test these two and i ran two fable subs worth of usage that full weekend i burned both of them all the way down and i did not get any meaningfully better results out of yes it wasn't i agree for code i want to merge which i believe is the economic driver of all of these models right now businesses using them for code at their companies fable is a tier above it it is but it is not it's not an a to s difference it's an s to s plus difference i think that is only the case because everything else is so far behind if if grok four six was two x better i think that these being considered the same tier roughly would be more reasonable yeah i if grok was if grok four seven is everything they're hyping it to be grok four seven should be a plus then there should be a gap then there should be sold then there should be fable yes but since we don't have that i feel like we are doing a disservice to a tier system if we don't recognize the tier difference between soul and fable because soul is an unbelievable achievement in taking the last generation of tech in the last generation model size yeah and making it perform close to this next generation fable is a shitty next gen game yes and i will continue to use this analogy we're like five six soul is the last of us for ps4 or ps3 and fable 5 is knack it's like a shitty ps4 tech demo but god damn is the tech better i will allow it i i am i i think we are we're getting like too many generations too fast but yes mostly agree that if we were if we did a generational split at the level of i didn't even mention this earlier which i should have it's the level of trust you have for the model yes the amount of trust i have is a lot like when i was talking about the effort spectrum that also maps directly to my trust spectrum i trust fable the most with code i will say that i fully agree there i trust gbt 5.6 a little bit less i trust grok even less than that and then everything else i don't know if you or alissa added the extra tears here but i'm much happier now uh oh you told your agent to do it i did thanks it just auto deployed thank you lake bed yep late lake bed did good here lake bed got us our extra tier and i am okay with leaving this as it is yes this is a list i am am i fully happy yes i am fully happy with this list as it is right now yep i am okay with calling this like our the the list we agree upon that's good with me i think the last thing i want to say before we wrap up is i want to slightly compare it to your tier list because i haven't looked at your tier list so now that you're looking at this and comparing against my tier list how do you feel uh so i have these two right here and for reference we have this up on the screen for those watching for those who don't the most notable differences immediately when seeing it are we don't have the gap tiers because with our new list we just made fables s plus soul is s then there's two empty tiers of s minus and a plus because i really want to emphasize the gap between fable and soul and the rest of the world then a tier is where things with unique value come in which are luna because it's such an absurdly cheap model is an absurdly useful model for the price fox alpha is glm 5 free flash which is a very small 300 bill effective model that is incredibly powerful for what it is it's basically luna at home yep and that being right next to luna i think makes sense agreed then b tier we have a d64 flash because you can run it locally the vision version is coming very soon and it is absurdly powerful to be able to run that on your own hardware then you spark 1.2 because it's such an absurdly cheap model being 6 6 cheaper than luna is an achievement that deserves attention agreed as long as using the contributor to your the actual full pricing not worth it at all nope then c tier we have kimmy k3 because of its novel value and being open weight then rock 4 6 right after because it's very useful surprisingly cheap very fast works great in the croc and codex subs cursor but yes sorry croc and cursor subs yep sorry they're code x x space x yeah it's bad codex is an unused name now i wouldn't be surprised if they take it then d tier we have glm 5 3 which is an incredibly capable model but doesn't have vision yep followed by tara and five or in deep seek before pro v4 pro would be higher if it had vision but it doesn't tara is just not a great value for what it is even within open ai's offerings f tiers opus 5 because it smells like a good model then you use it and merge its code and do you suffer as a consequence and then google tier we have the greatest google model ever which is sonnet 5 followed by gemini 37 flash and of course gemini 31 pro the underrated at nothing goat yeah it is honestly the overrated worst of all time i don't want to air out my whole tier list verbally here but we will quickly compare the differences the most notable difference immediately is that we have different tiers i don't have s s minus and a plus it's just s plus and then a then b c d and s plus to a like that's not meant to be a bigger gap that was just i forgot that i had made that change when i did the screenshot that's meant to represent that fable and soul are like neck and neck but there is a tier gap between them because i do think it's important to acknowledge that the gap between fable and soul is meaningful it's not just like you have your preference like fable is more capable it is it is more capable it is not a like it is not a letters level of capable it is a plus and minus level of capable t's their own that's where we will disagree but we will leave it here i think this is a good compromise one of us writes a lot of slop the other merges a lot of code yes yes you use lms more i merge more i mean a lot of my lm use these days is especially just like mostly research type exactly so yes 100 yep i don't have the extra tears to emphasize the gap and i actually in my video i did i emphasize how frustrating that was yep because i wanted to put a tiering gap between five six all and everything else i admittedly overweighted the open weight nature because i was live with a live chat that cared a lot about open weight so i have a kimmy k3 at the top of b tier which i still can defend the biggest issue with it is the license yeah it that really hurts it and the reality is the only people who can host it are like the neocloud types and if you are at the point where you can self-host kimmy k3 you can also self-host glm 5.3 and like get pretty similar experiences outside of vision and for those who are on open router right now looking at k3 and like are saying look here's two places that have kimmy k3 cheaper all you're saying is these are two places that have less than 10 million revenue exactly all you're saying yep that's it's sad but that's reality as soon as you hit the 10 million rev you have to make a deal with moonshot and that deal means that you cannot charge more than moonshot does it sucks which means that huge benefit of open weight is gone but it does have novel capabilities being open for iteration expansion fine-tuning evolution and other cool things that are genuinely novel and useful i think it's fair to put it there but it is a meaningful gap in my tier list versus where it is here i could fight either way i think i'm fine with it being in either place depending again on how strongly you feel about that license i think it being above both luna and flash feels bad to me it needs to be below those two i'm okay with it being above four six because they fill similar niches but i'm not okay with it being above honestly like i like that luna ox deep seek and muse are all above it i think that feels right from a holistic perspective all of those models have more novel useful features than it does kimmy k3 has things that literally nothing else does with the 3d stuff nothing is as good at 3d as it is right now as we're recording this that is publicly available true but like opus is close enough that like i and i don't think that's a novel like that's not an important part like its unique thing is that it's open weight and the open weight is really nerfed by the license and it is just not it's not efficient it's not cheap it feels of it feels a very weird niche that i don't okay i'll i'll drop a real hot take here okay barring everything other than the actual capability of the model so we're throwing away costs we're throwing away license we're throwing away speed we're throwing away harness we're throwing away everything other than raw capability of models today for doing real world work i think of everything below the fable insult here k3 is still the most capable model uh yeah if we were throwing away everything and just looking at capability i would put k3 at the top of a tier and then i would put glm 5 3 below it and then it would just be madness yes well when you consider everything else that does hurt kimmy it's weird that the open weight model loses points when you consider more stuff but in this case it does it really does raw capability k3 is the third best model available to consumers right now everything else hurts it i think it hurts less than the chart we made represents i think it is fair to put it at deep seek v4 flash level not because v4 flash and kimmy k3 is similarly capable but kimmy k3 is frontier level capable at a with a thing you can download and run yourself hypothetically before flash is unbelievably good for the price size performance and all those other things but you're giving a vision you are yep the vision really hurts it yep and honestly if ox hypothetically becomes what we think it is it might impact the placement of v4 flash to be real with you i can be pushed in almost any direction here i just it's the first time since r1 they've had a model you can download like a file on the internet you can hit download on that is at the frontier bar and when i think about it i think about tasks that soul can do that i would trust it for there just aren't very many that i would trust soul and not k3 hard disagree that i hard hard hard disagree i that is a fundamental underestimation of soul there is a this chasm is very representative of if i was going even just for like more simple stuff like okay i have this build on this machine and i need to wrap all this stuff up transfer the thread transfer all the assets get all that stuff properly set up and working on this other machine in this weird network setup i would 10 times out of 10 pixel over kimmy i would too but just for the speed no i cost i trust it far more okay i i should push kimmy k3 in these scenarios more because i every time i've given it a thing i didn't think you could do it did it and i was impressed i've had a similar experience it is it definitely has impressed me in a lot of ways i've seen it fall over in a lot of ways and i am not the honestly a lot of what i think the s tier and all of these up here why they deserve to be up here is because of the level of trust i have for these models that's that needs to be represented up here and i cannot give kimmy the same level of trust as i can those mostly agree but on the like the previous framing of if you set up the code base properly you give it the right skills you give it the right context you give it the right everything you put more effort into the prompts and all that i would say the gap from fable to soul is slightly smaller but not that notable compared to the gap between soul and k3 everything else barred but i wouldn't pick k3 simply because it's more expensive than soul the discounts aren't as good so decisions not as good the speed is shit the reliability of the apis that are hosting it isn't great i'm not saying i would pick k3 yeah i'm saying it's unbelievable that something as close to soul yeah as it is came out when it did as a file you can download yeah honestly that is true i it does deserve more credit for that simple fact but it is just when you're looking at the complete picture and the way the whole industry is shaken out licensing and just price all this stuff i'm okay with where it is but i agree that that deserves so what is funny here is if we remove the extra tiers we added to the current tier list that we just did and we add them the way i had like verbally added them to my tier list the gap between soul and kimmy is exactly the same it is a three-tier gap yeah but since we added those tiers in hours and i didn't add them in mine it seems much more like i am saying kimmy is like soul when in hours there's a five-tier gap yeah i think it should be a three or four tier gap i but like assuming that the things between are zero yes with and tearless are hard um if we're just looking at capability and if we were putting those three in a vacuum i agree when you're putting the whole of everything into context i don't agree do you think kimmy k3 is more or less capable of day-to-day work than grok for six assuming they were running the same speed and cost the same uh i would put them very similarly okay then i'm fine with all of this i think open weight nature of k3 despite the license should bump it a tier above grok for six considering the capability is similar but then also the speed cost and all those other things knocks it back down agreed and but but putting it slightly ahead of grok i'm okay with i think that is correct sorry for the long tangent on k3 it is my role as the kimmy defender and i think you said grok at least once there yeah i'm not and i'm not i'm not refilling it's over for me we will wrap we're getting the wrap up signal cool anything else uh i think when again we compare the rest here it's surprisingly close i didn't have ox alpha out when i did mine yeah i put spark and grok at the same tier in mine a slightly i put muse spark slightly behind grok four six in mine slightly ahead here because i think it has the more i've thought about the more i think it has novel value i agree and then in my list i had in d tier opus five then composer then five three then five six terrah then sonnet i hurt five three and composer two five meaningfully more in mine because neither had vision that pissed me off five three is capable enough it deserves to be separated out from opus we were harsher towards opus and sonnet in this new list compared to mine yeah i'm an anthropic hater so it's just gonna happen that was also harsher towards deep seat before prox i thought it was unacceptable to have a model that is that capable without having vision especially because i knew deep seek was capable of adding vision to their shit it just felt like a a weird training mishap or something like why does it not have vision it does feel bad i agree honestly to be super real with you with the if you have the context of what ox alpha is rumored to be i could see dump dropping both deep seek both deep deep seeks down a tier yep i could see that and no disagreement on gemini models in sonnet being in the google tier no none yeah they belong down there be really funny if anthropic sold the sonnet weights to google just do that as like a plant to like lead them down the wrong path hey hey google i feel like we have the these weights ended up here but they should be yours we're just gonna like pass them on like google's equivalent of releasing kubernetes to throw off the entire industry yeah very similar very effective it worked really well they screwed a lot of stars with kubernetes very impressively so still think we have time for the harness list no no this is the first time that timer has ever gone over three hours so we're gonna wrap right now thank you guys for watching i'm so sorry for all of this thanks grok grok cannot get me anymore i'm free it's over still get me a little bit yeah you're fucked god