← Back to search
ThursdAI - Jul 16 - Inkling 975B open weights, Kimi K3 at 2.8T, a 27B model on a phone & Codex hits 9M
ThursdAI - The top AI news from the past week · 2026-07-17 · 133 min
Show full episode description
Hey yall, Alex here, Huge thanks to Wolfram for running point on the live show this week. Didn’t have tons of time to edit this one, so please skip the first 10 minutes, it’s a loop of our new “wait for the live show to start” vid, that I build with HyperFrames and can’t wait to tell you about, next week! Today it seems that OpenSource is biting back, with Kimi K3 getting released just a short while after Thinking Machines (Thinky) has released Inkling, their near 1T model. I’m attaching the TL;DR and timestamps for the full show (my AI agents, yes even Fable and Sol are not a match yet at editing down hehe) and I’ll spare you the long Fable recap (please do let me know in the comments if you were expecting it)
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Weekly ThursdAI roundup of open-source model releases (Inkling 975B,
Kimi K3 2.8T) plus hands-on impressions of GPT 5.6 Sol/Codex agentic behavior.
Benefits
- Real-world impressions of Sol's over-verification and OCD behavior
- Practical multi-model harness patterns: Fable plans, Sol executes, Fable reviews
- Open-source coverage: Inkling 975B open weights, Kimi K3 at 2.8T
- Per-token cost insight: use expensive models even for simple tasks
- 27B model running on a phone shows on-device progress
Use cases
- tmux harness: Sol runs a goal-driven task list, Opus/Fable grab last 200 lines every 45 minutes to review and steer
- Doom built in ~2,000 lines of SQL by Sol, essentially one-shot in about an hour
- Sol left running on a client's machine for ~2 days; wrote better focused code than Fable
- Luna used for massively parallel simple tasks across all files
- intern.md two-file setup keeps Sol focused, ignoring the manager's notes
KPIs / results
- Inkling 975B open weights; Kimi K3 at 2.8T parameters
- Codex hits 9M users
- Doom-in-SQL: ~2,000 lines, one shot, ~1 hour
- 14-hour Sol task cancelled after it rebuilt the whole Hermes system
Tools / build
- tmux dual-session Sol + Fable review harness
- Doom clone in SQL
- intern.md agent-steering document
- Wolfbench benchmark runs on GPT 5.6 Sol/Terra/Luna
- Ultra code review sub-agents in Claude Code
📑 Chapters — tap a time to jump there
0:00
Intro, Alex on vacation, TLDR overview
- Solo host as Alex vacations; TLDR overview of a big open-source week
11:35
TLDR: Thinking Machines, open source, OpenAI news
- TLDR headlines: Thinking Machines drop, open-source releases, OpenAI Codex/GPT news
12:34
Banter: impressions of Sol/Codex, over-verification behavior
- Sol over-verifies everything: validates email drafts with Python code
- 14-hour task cancelled after Sol rebuilt the whole Hermes system
- tmux harness: Fable reviews Sol's last 200 lines every 45 minutes
37:22
TLDR restart & detailed breakdown
- TLDR restarted with detailed breakdown of the week's news
48:40
Open Source AI section begins (Bonsai/Prism ML, Kimi K3)
- Open Source AI section: Bonsai/Prism ML and Kimi K3
58:42
Inkling (Thinking Machines) deep dive & 3D model visualization
- Inkling 975B (Thinking Machines) deep dive with 3D model visualization
1:10:33
Kimi K3 discussion & demo comparisons
- Kimi K3 at 2.8T discussed with demo comparisons
1:27:02
Frontier Labs: AGI governance framework discussion (Demis Hassabis essay)
- Frontier Labs: Demis Hassabis essay on AGI governance framework
1:47:04
Grok Build CLI data leak & OpenAI file deletion incident
- Grok Build CLI data leak and OpenAI file deletion incident
2:02:15
This Week's Buzz: Wolfbench results on GPT 5.6 Sol/Terra/Luna
- This Week's Buzz: Wolfbench results on GPT 5.6 Sol/Terra/Luna
Welcome everybody to ThursdAI on July 16th and this time it's just me right now and Alex is on his well-deserved vacation so I will take over hosting too this week and let's see if somebody else joins us until then I will just go through the TLDR and let you know what is going on this week Of course it's been a big week as almost always no summer love this time this year. Let's go and start with the TLDR new introduction video Anyway, so we can we have a lot of news regarding new model releases open source is eating good this week and also open AI is on a roll with their codex user and GPT work use So the first thing let's start with open source. So this is just the TLDR. I'm just giving you the headlines and some quick overview and then we go into the details after this So the first one is thinking machines. Mira Murati thinking machines has dropped the history. Oh, Yam is coming and LDJ is also here. So let's just introduce our co-hosts here Come to the stage guys. Hello. Nice to have you. Hi. How are you doing? How are you doing? I'm fine lots of news lots of stuff. I mean isn't it exciting week this one, especially for open source fans like us? Oh, yeah. Oh, yeah. What do you think what do you think about soul? Speaking of Yeah, let's start with some banter before we go into the TLDR So soul I've been using it since it came out. It's my main model right now and I love it But I also noticed that it is A bit too much sometimes like When we did the Friday I with the reachy stuff when I we made the little show last week Where I was giving it the task and it went off and after 14 hours I cancelled the task because instead of just implementing what I said It was basically rebuilding the whole hermy system And I also gave it a little update that took about an hour usually and this time it went off and it's still going because it is It's not just doing what I want but it's checking so much stuff and finding issues even outside of the scope I have given it and it's fixing them as well It means it's Sometimes a bit too Enthusiastic to make everything right You give it a little task and it writes more tests It writes dozens of tests for 10 lines of code something like that. That has been my impression. How has been yours? It likes to verify things like to You know the joke that all LLNs they like to smoke test everything That's more than a smoke test like that's everything is tested And the funny thing is that things that it verifies and validates and tests It's always code but the things that it tests are not always things that you actually code like You can you can talk with it or like tell it to write something And it's going to validate it with code and you see like python code Validating an email or something that like makes no sense Yes, but it goes on and validates and I completely I mean that's exactly my own feeling as well You tell it to do something and It goes on. I don't know. Yes Write a draft for me Even when it's writing a draft for an email it writes a document it checks the document it makes the draft it checks that the draft is correct It sends it and then it verifies that it has been sent the way it was written So yes, it starts with like Like planning to draft an email And then like refining plan for draft and you're like it's not even the email yet It's just the plan for the email And you refine the plan and validate the plan and it's not even the email yet And yeah, absolutely. Yeah. Yeah. Yeah. Hi Nisten Tahirai. Hey Have you been playing with soul or one of the other open air models? What are you impression? You know what worked the best? Because for a client's computer they had both and I I let soul go off for like two days or so I just found it it does write better code than opens on fable We can debate that But it does write better code when it is focused on the task It still makes terrible architecture decisions out of nowhere might decide to do something in rust What worked the best for me in the last few days was I would just set up a session in tmux and I would just let it run there And then let soul run there with a goal and whatever and the whole task list And then I would get opus and I'll fable because they reset the limits To every 45 minutes just go in the tmux session grab the last 200 lines Do a good dip on what they did and then review whether they're on the right track or not Or just send them a bunch of keys as a prompt to steer That worked fantastic. It's like the best harness I have ever built I'm doing it all the time. I'm doing it all the time. The thing is that They They can gaslight one another Like you can easily see So convincing Claude that Whatever it is fixating on Is the right thing that needs to be done And then you see Claude starts to speak to you Like codex Which is very very strange But you immediately feel that you know It's like it has the codex text in it And you're like, okay, and the And you can't get it out of them Like no matter what you say There is one thing that That kind of work I just told it This is your intern He's pretty dumb Don't believe anything they say I have something like that in mind Keep the focus on the task So Sol had to make an intern.md document That they kept updating And they were not allowed to look at the manager's notes It's pretty simple Just two files Two tmux sessions That's Anyway When it comes to Bigger architectural stuff You can't beat Fable It just makes better decisions Overall The harness is not as good I know some people test it Like Theo, he tested Cloud Code harness with Sol He found it to be very good I just find that Codex for much longer running tasks Does a better job Like it does reviews and stuff While on a more regular basis And but yeah, that's where we're at Hey Peter, we are just bantering Talking about Sol Our impressions with it How it is a bit It has too much OCD sometimes It goes too far into certain directions That things are the right ones Yeah, you know like I feel like the problem With a lot of models And I think we go back and forth on this And you saw this with like Opposed iterations Is that it kind of goes between Or follows instructions so well That it kind of goes into that OCD Or kind of being overly prescriptive Or overly responsive to your suggestions And then they tune it back to be like Oh no, that was too much So let's not listen to the user anymore And they just like I feel like maybe This one is a bit too far And I find it difficult You know, when I just talk to Codex I find that okay But when you use Sol or 5.6 In like applications to do something I find it quite hard to tune So it doesn't like do very precisely what you ask So yeah, I think it's as always with new models We need to like adjust how we prompt That's right Ali J, did you use it? What are your impressions? Yeah, I'm really liking the model And especially just for like research work here and there Like just earlier I was using it to look up like What are some of the art details of Kami K3 that's available And things like that And yeah, I haven't used it myself for front end that much Over the past week But I've seen a lot of demonstrations And friends of the pod like Ray Fernando Have been using it And yeah, there's just some really impressive aspects about it And on the show when we did the Mars test with it Yeah, it's just Overall it seems like a pretty well-rounded model Shout out to Ray if you are watching I met him in person a couple of times and great guy So anyway, a question on my mind is Are any of you actually using one of the non-SOL models Like Terrell or Luna? Or are you all SOL guys? Phenomenal SOL on my end I did some I did some large pools of agents for small tasks You know when you need to do like parallel, massive parallel things Like go through all files and do exactly the same thing Or something simple like that I did that with Luna and it was pretty good Obviously, it's a very good model But yeah, I'm the worst example for this I'm X-high everything Like seriously So yeah, I'm SOL X-high and maybe a little bit too high At the moment The way things are But yeah Yeah, same, I only use it on X-high The ultra code feature on cloud code That's a dumb name to name it But it's actually very, very, very good All it does is it does what a lot of harnesses do Just spans up a few review sub-agents for everything that it does So that feature is good But yeah, the models are getting to a point where now you can just give them a task And you can kind of trust them to do it Which is pretty weird And that's why you're more and more tempted to just put it at the highest one it is possible Because before I would even leave Opus on low or medium Because I was going to go back and forth with it anyway And I just needed it to be fast and interactive But if you're going to leave it on overnight You just leave it at the highest quality possible And then you come back later And it's less likely to have messed up Yeah, and you're not using max, you're using extra high Yeah, max just takes way too long For Sonnet, I would leave Sonnet on medium Whatever the default is But I did try using it for other things It wasn't that bad But again, the steering, it just tends to go off and create features that you don't like But Table just scares me now It's at the point where I can just let it go And I know I will not be mad at it Like it's very rare that I actually find something I didn't like And it's pretty scary how good it is Yeah, so I think it looks like it's now at the time where you don't just want one model You want to use a combination of models like Fable for planning Then let the GBT 5.6 Sol do something Then have Fable check if it's done correctly Something like that? Yeah, I think that's the best way to use the limits You get much higher limits with codecs And I do think you get much better architecture and stuff decisions But that's a matter of taste For people that do Rust and stuff, they might want to do more Sol I don't know It just really works Converted things to Rust One thing I find that I think the mistake people often make is that they're like Oh, my task is kind of simple So I'm going to use like a cheap model for it And I think that's kind of the wrong way around Because like if your task is simple In fact, you can use the most expensive model And it's not going to cost you a lot Like for example, if you're writing an email Why are you trying to save like half a penny to write an email When you might be doing them for iterations to get it exactly right Like just use the most expensive model That's not going to... Actually, if you look at like the pricing of Fable It's actually not that much It's only expensive because you're doing like crazy loops And generating like a billion tokens But like per token, like everyone Like okay, most people who have a job can afford to use it Like just for simple things all the time So that's how when I was still like in my subscription trying Fable And that's how I would use it I would actually not get close to my limits Because I would just use it like in a simple way I would obviously try it on harder things But if you don't go crazy, the limits will actually okay Yeah, I keep saying that I don't need my main assistant to be an Einstein I just need an assistant who is smart enough to know most of the stuff And know when to call Einstein when it's necessary to do the really hard stuff So in that case, you would be using a model where you have a lot of tokens That you can use it freely for everything and it can rewrite your mails no problem And then if it has the architectural design or anything complex Then it can reach out to one of the models like Fable So I think we are seeing now these different classes of models Where a top model, you can't just run it for your standard agent in a loop all day Because then it gets too expensive So, Peter, you also did something really cool And maybe this is the best time to talk about it Right now for all our viewers that have joined so far We are still in the banter phase at the beginning I know it's been 15 minutes now But this is cool and I think now is a good time to show it instead of doing it later You built something that is really impressive Using ZOL and SQL SQL Yeah, so that was Should I find it? Maybe I can find it So the What I was trying to think about Is Yeah, let me show my screen So what I was trying to think about is What's the kind of dumb random thing I can do That is not necessarily like the The classical things that everyone does I don't know like task ups or whatever So I was trying to think of like What's the... What would be that weird thing? And in here I landed on creating kind of a Doom version With SQL And roughly the way it works is that it's basically It's a database language You build it in a database language Yeah, and it's like I mean It's kind of... It's obviously absurd But it's also not like Once you kind of imagine how we can do it Right? In terms of like It's basically it's like 2,000 lines of SQL For every frame it kind of Calculates the different pixels And kind of puts it together That kind of thing So it's like It's not like AGI in there, right? You could have probably done this like a couple of years ago To be honest Like it's just... The reason why no one has done it before Or I don't know, maybe someone has done it But at least that wasn't like a thing It's not because it's so complicated to do But I think the big difference is that I didn't need to babysit this at all This was like one shot I just like had a bit of back and forth at the beginning To say like what do we want to do? I gave it to Sol And like I think it was like an hour or something It came back with this And the only thing I was iterating on Is this console on the right Just to make sure that it's like It's actually something interesting to look at It was like a couple of iterations But the rest, what you see on the left That's all one shot And it has like different kind of levels as well So yeah, and I don't want to claim this is like I know such an impressive application But it just kind of shows that you can now do random things That you don't have to babysit Yeah, and I think it's very funny that we just recently got the Unreal Engine connector and MCP So your AI can use the Unreal Engine to build something like this And you just said hey, do it in SQL That was cool That's insanely impressive This is not impressive at all That's insanely impressive I mean, yeah, look If I'll sit and think through this I don't know, a little bit I might come up with how to Maybe conceptually how to do something like this But that's like Look, that's the perspective with 3D And it's also in SQL That's incredible, man Yeah, next time we do the Railgun Launcher We do it in SQL If the other benchmark is saturated Yeah, and I think one other thing that I did Let me find it And it's maybe weird in a different way Is that, oh, it has sound as well in my ear So I'm going to mute that in a second So it's I created the Minecraft clone Can you see that? I don't know if you need to promote me Yeah, it's on the page I created the Minecraft clone in Lean Which is a Apart from Something that you can use mathematical theories with You can apparently do things like that And the way it roughly works It uses a Raylib library to just Actually for visualization But all of the code, if you look at it It's all Lean code And apparently that's something I had I mean, I don't know much about Lean at all But apparently you Like it is at the end of the day Like a functional language as well So you can do things Just no one ever does Because it's like an insane thing to do But like, yeah, you can see here These are Lean files and so on So it doesn't have like A couple of people said, oh, like what theorems is it proving? It's like, okay, it doesn't It doesn't like you can't just play Minecraft and prove something It's not that But it's like subverting the language I was trying, I didn't share it But I was trying to do something in like In Google Sheets And I must say Google Sheets was far harder than like Lean or SQL I don't know why Well, I guess I know why It's like less functional, right? So it's more difficult to do I kind of did some crappy version of Mortal Kombat game But it didn't really work So I'm going to try again I'm going to try again I'm going to push it It requires kind of creative interpretation Of the functionality that it has What was the prompt? I'm curious to see that What's the prompt? For these type of things like Are they complex or No, no, not at all So the way it worked for both of them is that I just had back and forth with the agents Just talking about like how would you build it? And not like, not in a crazy way Like I know nothing about Lean Or like similar to what you were saying Like I wouldn't know how to build this in SQL It's like, it's far beyond my SQL understanding So in So you just kind of go back and forth And then I think if I remember correctly I don't think I set a goal And I think for both of them I put ultra setting I don't know if it was necessary But they just like whacked it on Like why not? And just to see what they can do And that was it So it wasn't really like a prompt It was more just kind of Have a conversation and just let it build it It wasn't like anything too special Did you want to build it in Lean? Or was that an idea the AI had? Or was it your choice? Yeah, so I need to remember exactly how it went But I was having a conversation about like what What could be interesting and we somehow landed on Lean I can't quite remember like whose idea it was exactly But it was like a bit back and forth Yeah And what were you using? Codex and Sol? Extra high? Yeah, it was Everything was Everything was on ultra when it came to building it I think I was probably I wasn't chatting to it on ultra So I think it was I normally do you I kind of go back and forth between like high Extra high and max I must say I don't really know Like I don't have like a good feel It's not like I do on high and it's terrible I do on max and it's amazing Like I can't quite like intuitively tell the difference I just feel like if I want to like have a quicker conversation I go with something a bit quicker But yeah, maybe high But if I wanted to just go up and not worry Then I just throw it into max It's always the simple prompts that scare me Because you just found something and then it just goes off That was the whole thing about a realm It was supposed to be just like a three, four sentence prompt And it just goes off on its own As someone doing benchmarks, I'm always thinking about how reproducible all of this is You know, you can use the same simple prompt and do it three times And you get a killer game and the other two you may not get anywhere Especially if you have to steer After the fact But it's amazing that it can do it So that is what is being proven There was something really cool from last week And we didn't cover it But it was when Jared Sooner rewrote Von from Zig to Rust And there was a lot of controversy and stuff in the programming world But I read the blog post And actually, I really want to show it Because his agentic use was pretty amazing and insane He used I think over a hundred Let's take a look at that and afterwards we go back to the TLDR to get the full form Yeah, yeah, yeah, okay This is very cool to see Entire window, you guys can see Okay, so what Jared did was he started Okay, so the screen is okay So he started off like a pretty simple Two-step prompt You push something You have two agents review it They each come up with a different review And then you see what they each said You pick the best from them and then you merge it in All right, so it's like a straightforward process But the crazy part comes when he starts scaling it So these are all the commits over 11 days 6500 commits that it ran And so there were a whole bunch of errors I don't know how many hundreds A thousand, three hundred or something So he split that up in between different work trees And different work trees had a whole set of errors to each fix And actually, okay, there were more, there were a couple of thousand You guys can see it, right? Yes. Yeah. And then the trickiest class of errors was cyclical dependencies Okay, so dependencies and stuff to cut Okay, so he fixed 16,000 compiler errors And then he just fired up the agents Because the beautiful thing is that BUN already has a testing suite Whether it is compatible with NPM or not So you can just grab all of the tests, the thousands of tests And then he just started So this is what happens when you have unlimited tokens He just started firing them up into different shards So testing it on macOS, Linux, ARM, Windows And this was over a period of a couple of days And then, so as you can see up till here There wasn't a lot of coverage But then to the very end All the tests passed So he announced like a few weeks ago, last month That he was going to just attempt this He was just going to try and rewrite it in Rust It's just going to be a fun experiment And now it looks like the rewrite worked very well And all the tests passed So they might actually push it Which to me, it's pretty crazy So there is Yeah, there's that I'll stop sharing for now I'll have something cool later for the inkling model So let's go All right, TLDR Okay, I will play the transition again Because it's a cool new video And it gives us a chance to bring something To the media assets So here comes the restart of the TLDR So since you mentioned it, thinking machines That was also my first part So in the open source segment, thinking machines Release their first model or rather models They released almost a trillion billion MOE That is the 975 billion total parameters 41 billion active parameters MOE Train from scratch on 45 trillion multimodal tokens And it's full open weights, fully Apache 2 licensed model So from the benchmarks what we have seen Is that it is a top open weights model from the US And so the best Western open source model Basically So passing Nemo Tron 3 Ultra So we will get into this when we go to the open source section And then I'm sure you have a lot to say About this, so just that As a heads up, so thinking machines Release two open source models Although smaller one And yeah, Western open source is going strong I think Another open source release Quite the different direction Because it is a very small model Or rather, it's a 27B model So actually not that small But it is now running even on phones Like an iPhone it can run Because they quantized it heavily And I think this is the biggest model That has been made runnable on the phones We had some Google models that were very tiny And this one is probably the biggest one That you can run on your phone now And it has very good scores And yeah, so local AI is also going strong Because both are open source The thinking machines models And this one But the one you can run on your phone The other you can in your own data center And it's good to have the choice and the options Although MOS VL real time Is also an open source 11B Vision language model for real time streaming Video understanding with proactive speaking So pretty much what the real time models do Where you have duplex You can give input while it's giving output But this can react to a stream as it is happening And doesn't have to ingest it from the beginning Well that is the open source stuff And we are all waiting for Kimi K3 to be released The benchmark result there is some stuff So this is not a rumor anymore And I think we can cover what we have Or we may get breaking news We will see But also Kimi K3 is definitely something interesting And we should talk about it Are we going to have breaking news? Yeah, the API was actually just dropped and announced Like an hour or two ago And it has confirmed 2.8 trillion parameters And it has attention residuals So doing attention basically across the communication between layers As opposed to just classical attention And yeah, it's really exciting It's not quite open weights yet But they said over the coming days That they will be open weighting it too Yes, great This is amazing And yeah, we definitely have to cover this in more detail As we get in the open source section Let's see In the big lab section We have a lot of open AI news actually So when I look at the nodes We have the amazing records of users Now that Codex and JGBT became one app So we have JGBT work And Codex in a unified app And the user Amount of users When we made the news The nodes for the news a couple of days ago It was 7 million Now we are up to, what is it? 8 million or 9 million? It is definitely 9 million now 9 million active users of JGBT work at Codex So it's been exploding And yeah, they will definitely look at this as well And we get to the big labs They have made their own hardware So you have a micro keyboard Which you can use for push to talk To talk to Codex And you have a dial to tune in the The thinking and effort levels And you can quickly switch through the model So, okay I didn't expect that I was waiting for some other kind of hardware coming out of open AI But yeah, it's a little thing And yeah, if you are using agents with voice It may be very useful to have something where you can quickly tune it in So we will take a look at that gadget After when we get to it Also JGBT is back on WhatsApp At least in the European Union and the area Because not all regulation is bad And there's antitrust regulation Which forced Meta to let open AI Put JGBT back on WhatsApp And hopefully that is that's a precedent So it can get back to the other systems Because all these closed systems Where a big vendor has a social network And the only AI allowed of their own We wouldn't want that with XAI happening And in that case I think it's a good thing that they were forced to open up And allow another model on their system as well There was also news about JGBT Red, which is an effort An in-house effort to Fight vulnerabilities to do red teaming And this automated red teaming effort Was more successful in doing this Than actually the people doing it And it should help prevent prompt injections for instance We are all using our agents going out on the web Injusting data from random sources And it is super important that the model obeys the user And not some instructions found in some documents It got from the web Or any other kind of attack So I think this is a good effort That is open AI Also, there has been an incident Or multiple instances where GBT 5.6 has removed home directories and files from there And it has been confirmed by open AI That this is Yeah, it's a known issue And they will change the system prompt And codecs to prevent this So basically it was a mix up With some variables where It tried to do some temp stuff and it got there And personally I've seen Something similar Where it was trying to do an isolated test In an environment it created outside of the production Environment But it still, since it was on the same machine It still had a connection So I should have put it in a sandbox It was easy to fix But this was also something where I said, hey, you messed up And BI said, oh, shit, this shouldn't have happened But shit happens So always be careful what you do Have your backups ready And yes, AI can make mistakes Even if it's AGI or near AGI It still makes mistakes like the humans do And I think this will be true for a while CROC also made something which is not an AI mistake It was more A mistake by the company Or not a mistake, but it was pretty nefarious If you think about it, if you heard the news Where if you were using the CROC-built CLI Their own CLI agent It was uploading all the data from your repository That you were working on To some Google Cloud Storage For tracing Even if you turned off That the model should do stuff like this So it gave it the whole history Even deleted files, end files Anything in your repository Which is a big, big trust issue And yeah, they When it was found out They removed that, deleted the data But can you really trust someone Who does something like that? Not by accident, but purposefully And there's VDR But VDR is only for enterprises Not for private or Smaller users, basically So this was major fuck up I would say And, let's see Google is also in the news Google has released some new Jemma updates Jemma 4 got some updates that improved it a little bit In speed and performance And qualities Great to see them not just release a model And be done with it But actually continue to provide updates And one thing that we should definitely Talk about is in other news That David Sotheby's The CEO of Google DeepMind Wrote an article, an essay basically About where he proposes an AGI governance framework Where he wants a US-led authority That controls what is happening with AGI Because he was kind of skeptic before this But now he is also convinced that AGI is only a few years away So this is also a major change That we've noticed it We noticed the inflection point Where AI has become suddenly much more capable With the better harnesses, the better models All of the things coming together And now they are pretty convinced that AGI is not a dream In 10, 20 years But something we will see in the next couple of years And this will be, he said Greater than the discovery of fire or electricity Ten times of the industrial evolution at ten times the speed And I fully agree with that I also think this is for our human Humanity's acceleration and development and evolution This is now the axis is so steep And this is the moment So we'll talk about this as well If you have seen the article Then I'm very curious about your opinions on this Okay, I think that's it for the TLDR So a lot of stuff to cover And unless we get some breaking news now I would say we start with open source And take a look Oh, LDJ, you want to say something? Yeah, I mean, conveniently It's both the open source thing And something that should probably be in TLDR But the Bonsai Ternary model For the 27 billion parameter Quen model They ended up releasing the Ternary version of that And we previously had covered I think their 8B Ternary and binary models But yeah, this is a significant step up and pretty impressive Yeah, the Bonsai 27B, you mean that? Yes, exactly I was able to run it around 16 tokens per second On a, I don't know now, it's like $150 GPU Wait a second, Nisten, we can do it when we get to the point Okay, let's just start the introduction and then we can talk about it Open source AI, let's get it started There we are again in the open AI Open AI, open source section for AI So let's get started Would you rather start since you mentioned Prism first? Let's start with Prism, why not? You just got started Nisten, so Please continue Yeah, sure, so Again, they have discovered this one bit technique that nobody else has figured out yet How to do When we do other models like the traditional way You make some of the layers bit net and then you leave most of the other layers as Like the important layers as either 8-bit or 4-bit And then you do a mixture of those And you can get that to Like an average of 2 point something bits per weight It's all But this one is crazy because It applies to everything, even image generation You can do, you do everything in one bit The embedding, the output and everything So what they did was they took QN27B And they shrunk it down To 3.8 gigs Including the image recognition ink So Yeah, we have QN27B running I have it running on my phone It just goes at barely one token per second But it does actually run, which is crazy And if you have an actual An actual GPU Even a very old GPU A gaming GPU With like 6 gigs of VRAM It will do 16 tokens per second And yeah, there is about like 5-10% drop On the benchmarks for this But I got it all running And it was working fine honestly So I don't know if I should share from my phone Because when I screen share the phone It's a lot slower as it is And it's It's already too slow But Yeah, so I was just able to wire up a chat app To the phone And No meow, grinning cat face Did you have a cat in mind? Or are you just saying hello? And the model is just responding And I can just do How to build a city on Mars And let's just see So this is just my desktop 1660 Ti GPU And it's a foldable phone And it just goes Oh, that's a fantastic idea Building a city on Mars Is one of the most ambitious and exciting challenges So you can have a home chat And stuff and Yeah, so the more fun things That I was able to do Were actually desktop actions I trained one model for that But Now a lot of things that it does It's Yeah, they're desktop actions So I can just do I don't know Set an alarm for like two minutes from now So Yeah, so there's whisper small on device You can run a model on device It can summarize your stuff now Guys, we are Stuff is getting pretty Pretty crazy Like we are at that point Before we get Before we go further in this Let's just introduce a model Yeah Basically Bonsai 27B Which has been released by Prism ML Is a 27 billion parameter model It is multimodal It has tool calling support 262k contacts It's one bit version Has a size of almost 4 gigabytes So 3.9 gigabytes So tiny, tiny model on the phone And the normal 27B model in 16 bits Is 54 gigabytes The forward quant is 18 gigabytes And there's also the ternary We've been talking about which is about 5.9 gigabytes At 95% of the full position benchmark across 15 evals So it is in math and coding basically untouched And the one bit version still retains 90% Speed numbers are great 163 tokens per second on an RTX 5090 87 tokens on an M5 Max And on the iPhone 11 tokens per second Is still usable Like you said, it's based on QAN 3.627B So it's Apache 2.0 license The GDF files are on Hacking Face And this makes it possible to run Create AI on your phone LDJ, you wanted to add something Yes, I wanted to add that When it comes to the capabilities And how it scores in benchmarks and everything It definitely is not as good as The full precision versions of the models But it is much better than the traditional techniques That have existed prior for compressing the models By this much And what I'm really curious to see Which I think we might also see in benchmarks coming soon Is like how does this compare to let's say I think the math I did was something like Like QAN 8B or 9B That model in 4-bit should be about the same amount of gigabytes As this model in like ternary or 1-bit And so I'm curious, how do those actually compare? And I'm working on putting together some benchmarks for that Oh great, when you have something Make sure to share with them Yeah, I think this is a big shift in what's happening Because the QAN 27B capability level is the It's like the minimum requirement for running a Hermes agent Or for you running an agent at home And up until now for most people that was not that capable But now as long as you have a GPU that has 6 gigs of VRAM Or a Mac I mean the 8 gig Mac would be very much a stretch You're probably not going to be able to run it in there But yeah, on 6 gigs of VRAM I was getting 32k context And running it at 22 tokens per second If you're on Windows, just use like LM Studio or something And just run it there And you have your own personal chat assistant that can set up actions Like if you have a random Windows gaming PC You just keep at home Set up LM Studio and run it there on demand You have an API, you can do whatever you want with that It's... That's the... That's where we're at right now Which is pretty nuts that we got here Yeah I actually... Sorry, I'm going to keep rambling about that I actually think that about We're at the point where like open source could do like 5% of the work Maybe And I think we're going to just see a rapid shift now From like 5% of the total work stuff Just like summarizing your day Or setting up your appointments Replying to emails Maybe even doing your finance and bookkeeping stuff We're going to just see that shift from like 5% to like 8% of the useful work Can just be done with open source models and local inference It's just like development and troubleshooting That requires the best model you can possibly get So we're going to see this weird thing happening in the usage Where either you'll have like the minimum acceptable, minimum viable intelligence And then you'll have that All right, so let's go to bigger ones Yeah, this was a small stuff, so small price But it has a big meaning of course for us And now there's the other end Where we have this huge model Almost 1 trillion tokens 1 trillion parameters And it's an MOE The other one was a full one And this one I think it will take a while until This is running on your phones But having this available as an open source model Is of course very, very welcome Open weights Mira Murati, the former chief CTO Of OpenAI Has left the company and founded her own And this is the first big release that we can look at And not just look at But download and use So big round of applause Because I'm always happy when open source is released This applies to the small models and to the big models And it's great that another Western US company Is getting into the open source stuff And it's not just closed source frontier labs But also open stuff happening So train from scratch Number one open weights 975 billion total parameters 41 billion active parameters MoE Trained on 45 trillion multimodal tokens And at artificial analysis It debuts at number 41 For the open weight stuff That is even higher than For instance, NemoTron 3 Ultra Three points above it It can do text images and audio It's no external encoders And interestingly it was built in just nine months On NVIDIA GB300 and VL72 systems The scores are also very good I personally, since I'm doing Wolf Bench On the thermal bench stuff I'm also interested to test this model soon So when we get later to the This week's path I will show you the Sol scores and the other OpenS scores But we'll have to test this of course Peter, did you test this one? No, not yet We just have it available now So we should release some scores soon But yeah, it looks like it's I want to see it more But it's certainly an interesting addition And the fact that it's trained also independently I think is interesting And the reason why that's interesting is that Hopefully it might give us like A bit different distribution to other models Especially for open model Like that's particularly important Because there's not much point having A second tier model that is like exactly Like exactly the same as the others So hopefully even if it's Let's say it's not the best model It probably wouldn't be, right? But it's the first one But the fact that it could be different It could be quite a nice help Like in many different ways Even if like you need to, I don't know Generate something Generate a bunch of different ideas Like it's good to have different distribution I read that it was trained on Some of the Kimi output So it's also using some of the Chinese models As a source basically And also techniques from them That also led to some outcry like Oh, it's just based on Chinese technology But I think all the technology in AI Is based on the other and this and so on It's a process, it's science And you can't just say they take only that stuff From one country or another At least that's my opinion about this Yeah, that news coverage was That news coverage was very annoying Because the DeepSeq architecture And a lot of their attention to stuff Even Anthropic uses them And everybody uses the open source DeepSeq research Also, everybody uses other LLMs to clean up the data As well So you take original data And then you get the other LMs to just like Clean it up, remove stuff Put it in the right order So yeah, I feel like it's under Underappreciated And they made some architectural stuff Improvements too By the way, I can share that Because I had I had Fable make a 3D animation Oh yeah, you did something very interesting So let's look at what you made Yeah, so... So what are we looking at? So we're just looking at Every single thing that you see here Is just a file that's on the model So what we're looking at is all the layers of the model And just how they sit on your hard drive And I made an animation to tell you Exactly how... Exactly how big they are So we can just go on any layer here And I made it as a teaching tool And so we can see all the layers of the model And then... So you can drag and you can pan around And then we can see, for example We have the key Q... KQ V weights So key... Key query value And we can see exactly how big that weight matrix is So if it says it is 12.6 million That means in 8-bit, that's just 12 megabytes But in 4-bit, in NVF before It will be about 6... Roughly 7 megabytes And we can look through the entire model here basically So we can just click on every layer And then we can see that for each layer How many routed experts are So there is a pool of 256 experts And this is just posted on my Twitter And so this gives you a visual view On how the model actually works And all the code and explanation for it is here So you can just check it out on my Twitter And search for Inkling But yeah... So what happens is on every token that's being generated It starts with the text embedding Or you might have audio embedding Or you might have vision embedding And then so it goes from the first layer And then on the left here I think we can even like maybe zoom in a little bit And then you can just start going through the different layers And understand everything that's happening here So they do a lot of tricks The most interesting thing I found is that You know how we do speculative decoding pretty much Which just like helps the model run faster and predict ahead And if I remember that correctly That's how they do it So in this case they have integrated that in So that's what's called the MTP The multi-token prediction So it's like there is a tiny little language There's a tiny little language model Still share the same experts But it tries to run every token through these first These 1, 2, 3, 4, 5, 6 And then if that doesn't work It goes all the way to the bottom of the stack And then just uses all of them With the routing and everything So I think this is a pretty cool way to visualize Like a 1 terabyte So in Bflow 16 it's actually 2 terabytes And in VDF before it's in 4-bit It's about 600 gigs Because not every layer is quantized And yeah, I made this as a tool So people can just use it for their students or whatever If you haven't put a license Just do whatever you want with it It's just whatever Fable made And I thought it was a pretty good way to explain And I found it interesting that Fable could tell Because I only gave it the configuration I didn't tell it much else Because I didn't want it to get triggered To think that I am doing like LLM research And it was able to tell just from the config alone This was the DeepSeq The DeepSeq V3 The DeepSeq V3 architecture So yeah, you can just keep pressing through this And go through each layer And then you can see exactly how many megabytes or kilobytes So the layer norms are just like a single normalizer So it's not a weight matrix And it tells you like the dimensions of the matrix So if it's like 1000 by 1000 That just means 1000 parameters by 1000 parameters That means that's 1 million parameters That's what those means So it's a good way to visualize what everything is And what it means on disk Like how many megabytes is that? So it's just a bunch of files basically that you're looking at And the size in the cube of the square Corresponds exactly to the size of the file So you can see that the expert weights are really big And the KQV weights are very small And I think that's pretty cool So yeah, that's my... EddieJ, you have a comment? Yeah, I think a funny irony here is Initially the multi-token prediction mechanisms I'd say a lot of it was pioneered by Meta I think I want to say around 2023, 2024 When they released the paper on it Scaling it to like a billion, two billion parameters And then DeepSeq within like 18 months of that They released DeepSeq V3 And that had multi-token prediction and everything And now we're seeing that kind of like I guess reclaimed by the American labs And now the American labs are taking inspiration from DeepSeq and others on doing that It's cool But that's how science is supposed to work, I think And... Yeah, exactly, yeah Yeah, it's self-speculative decoding So what people did, if you guys saw the hacks that people were pulling They were just grabbing like a tiny Llama model And putting it in front of the big Llama model Or the big Quen model Which is speed up LlamaCPP And then it got to VLM And now it's just baked in natively into the architecture Just to make it run a lot faster overall And this is multi-modal too So yeah, I think it's pretty cool I think they did a very good job to pick the DeepSeq architecture as their first model It's an excellent choice And they even introduced some of their own improvements So the way that they don't use rotary embeddings But yeah, anyway, we could ramble about that for a while Because I don't fully understand it myself So yeah, that's it, that's my show and tell Well, that was inkling And there's also a small preview that is not released as far as I know But it will come It's 276 billion parameters, 12 billion active parameters And it's also pretty competitive with this And I also found it interesting that they specifically fine-tuned it to say I don't know when it's uncertain rather than hallucinating That was also pointed out And the price $4.68 To $9.3.6 Depending on for the output and input token So it's pretty cheap About average for this size I think for the size and quality category Yes, so much about inkling Let's move on to the next thing This is MOSVLL Have you heard about it? Let me just bring it up So This is this one Sharing again So This is a This could also go into the video section It's an open source Apache 2 license I love that every model nowadays Most releases are Apache 2 license Always Always amazing to have it fully open source And not some If you are a company and you are bigger than that Or if you are located in a specific region You can't use it So great, they're fully open source MOSVL Real-Time Open source 11 billion vision model for real-time streaming So It has a cross-attention architecture It is a vision language model for continuous video streams So it's not in turn based Watch the video then answer This one is watching while it's generating It can be interrupted Can change its answers depending on what happens at the scene And knows when to stay silent It's Yeah, it's 22.7 gigabyte model size So you can run it locally On your machine if you have a Modern graphic card Which is important for real-time stuff There's a real-time streaming version There's an instruct for offline use And there's also the base for fine-tuning Big applause also for releasing base models Which is something unfortunately we haven't been seeing that much anymore But this is great Yeah, if you are If you want your AI to react to something it's seeing on the screen On a video This is a model to definitely look at The use cases like For detection, shot counting I have my use cases for this and I am interested to try this further So big shout out Opsource release, great And I think this is it for the open-source section This week Unless of course you get Kimi Let's do Kimi now Kimi is open-source Even if we don't have the weights, it is available So we will look at it Let me see if I already While I bring this up And DJ, you wanted to say something about it I'm sure Peter also has something to say Yeah, so since this is so new I don't Yeah, for Kimi K3 Yes Yes, so Yeah, 2.8 trillion parameters They didn't say exactly how many active parameters it has But they said how many total experts and active experts it has And the whole, the models aren't only experts But based on my rough napkin math It would be about like 60 to 75 billion active parameters for the model And it's It's a native vision model, they say Which usually these days when they say native vision model They're referring to having no encoder and directly into the model So if that's correct, then that's really interesting And then you have a million token context And I don't think they actually have released any benchmarks for it yet But they say in the coming days it's going to be Released open weights And so I imagine benchmarks would inevitably come out with that And the price is about half the cost of Opus 4.8 and 5.6 so That's expensive It is quite expensive But I think Peter has probably done some of the most testing than anyone has I remember you put out what it was like 30 minute, 60 minute video Yeah, it's private now Yeah, we slightly got a bit trigger happy with that one So we're going to re-release it whenever it comes out properly I think they haven't tweeted about it yet So is that right? I don't think I saw a tweet like when I checked half an hour ago I don't know if they have So yeah, I think my The timing is a little bit unfortunate because I think we're going to get like a flurry of scores now And I think my sense is that it will be a little bit confusing That I think there will be some benchmarks where it's just going to do amazingly well And there will be others maybe it wouldn't do as well So I think it will be quite confusing for people to say like Oh, is this is this going to be like better than Fable or not? And so on So it's I think it will be difficult in my personal testing I mean, I don't know how much I should say considering it's not out yet I think I had kind of slightly mixed views personally Like I wasn't like blown away by it Doesn't make it a bad model But I wasn't like, oh my gosh, we have open source Fable Like I didn't feel like that So But you know The model is actually out on API and everything for people to use It's just not really open weights yet, but it is out Yeah, yeah, I just don't want to, you know, front run the announcements I think that's not very nice to do Which we kind of did But yeah, but look, I think all I want to say is that I think it's worth trying it out yourself properly Because I think to anticipate the conversation I think it will be very confusing That when there'll be a lot of people which we already saw Which are like, oh my gosh, it's so much better than Fable Everything else is trash and it's like this is it now Like I don't think that's true But is it like a bad model? No, I'm sure it's like a jump versus like previous models But I think it is interesting, right? Okay, open source 2.8 trillion parameters Like does that help anyone? Like I don't know It's a kind of like, okay, good What can I do with that? Nothing really, right? So I guess it just opens it up for like I guess it was it I'm forgetting which company Was it Kimmy, right? They had the slightly more restrictive license Right? For am I am I remembering this correctly? Right? With the open source situation? Yeah, over any company that makes over 100 million or 2 million I don't remember I can check it up I think Kimmy has instructions that you have to mention them That was probably added after what happened with with Cursor I think that was before Yeah, it was before Sorry It's modified MIT So it's still an MIT license And then in the end they just say Okay, our only modification part is that If the software or any derivative works thereof Is used for any of your commercial products or services That have more than 100 million active users Or more than 20 million US dollars Or equivalent in monthly revenue You shall prominently display Kimmy K2.7 code On the user interface of such product or service It's just the last one I read Okay, so you just have to mention it It's still MIT All right Yeah, it just I think it's interesting I think if it's mentioned I guess that's fine But if the only real use of it is If it's like Yeah, big hosting companies can now host it But they have more restrictions Then it's like It's kind of Yeah, I guess it's technically open But maybe it's like with a bunch of cover So anyway, I'm not I actually don't know what the license is So don't take it as like a firm statement But I think it's just interesting I don't know what what do you guys think? Like for me when I see these numbers I don't really know what to think Like does that even help? Personal use it's too big You won't be running this at home But I think for a company Who is Building something or who wants to control the token costs By setting up something internal Then they know they can run a model And it can be taken away like Fable Or the price can just be raised at any time What could happen? I think being able to use a model that you own When you put it on your own hardware It's your model now and nobody can take it away I think having that That ability for a model is a good thing Especially if you are in a non-US country Where you may be thinking, okay, I can now use open source From a US provider, but If that ever is removed from my access I don't have anything And with this model, even if you can't use it now You can at least get the weights, put them somewhere And if the worst happens, you can still use it On your own hardware If you can buy the hardware, of course But I think it's good to have it And to have competition While now you have different hosting Provider that are offering the model If it's open source It's not just Kimi releasing it You can go to different ones They have different optimizations Different price points I think that is a good thing to have Any other opinions? Yeah, I don't want to I don't want to front run the actual announcement But I'm just saying that Rumors are saying numbers look good My numbers look very good Just saying that The price is Is a little bit high According to the rumors I'm just saying But let's hope Let's wait for the actual Fun announcement, I think I'm sure it's the jacket frontier Where we have specific parts Where the model is great Like some models are extremely good in design And others are better at planning So I think we will be reaching an area Where you can't just say this model is better than the other one You always have to say what for And it makes sense to have an ensemble of different models For different aspects of your work Sorry Peter. Yeah, I was going to say about the price I think it's Inevitable that The price will be higher, right? If it's 2.8 trillion I think what was it? 1 trillion before There's no magic there, right? Someone has to actually run it And serve it And I don't know if it's a fair thing to say Probably doesn't apply to every model But I haven't seen A lot of evidence that once the model comes out Or it somehow gets optimized And like prices drop a lot As thing I was kind of hoping that would happen And I did some analysis Like this was like over a year ago So like this I'm not sure how much that holds But I had like open router And I had two snapshots of the prices Like over some months And I could see prices actually went up for some models Because some providers drop out Like if the model is not hot anymore Some providers drop out And the prices start to creep up Because I guess they can So I think it's also kind of a tricky thing Like yeah, technically you can access it Via third party providers Which is a great thing But you know, then they They're probably more likely to change the prices Than like anthropocon opener as well Because they can just like There's less downside for them to do that So yeah, it's I think it's like an awkward time For the market as well Since it's all kind of a little bit in flux Yeah, since Yeah, Eddie J first Yeah, I was gonna actually show some Demos Or I don't know Maybe that's too premature But from The UIs that people have created With Kavine But I don't know Now second guessing Do we want to do that? Or is that fun? We can do it Let me just add something to the pricing thing That Peter mentioned Since I work for a company, CoreWeave That is also doing inference Our CoreWeave serverless inference Where you can use these models KimiK3, we are definitely looking To making it available as soon as possible So I see all the different Aspects of this Where you are looking for How to optimize it for the number of users You have to provide this model And it has a million token context The context window is a million tokens So that is very big And also often something Where different providers provide different limits Personally, I've said that it doesn't It is often not useful To provide a million tokens At a slower speed When you can't even use the full quality Up to the million So it sometimes makes more sense To limit the token window And provide faster inference for that So different providers make different decisions And then you have to really Look at the providers and choose the one That is most appropriate to your use case And it's usually features And instead of price Where they are competing right now But let's see how it changes When there are more providers Or people are able to use better models locally I envision in a short Near-term future Like if you have central heating You invest a lot of money And you are able to heat your whole house And in that case I can envision people to install Some AGI system in their basement Where it provides AI for the whole family And that can be more expensive And you are the only users And stuff like that But I think that would be like the outcome in the future Let's continue with Kimi K3 And if you want to show something Go ahead, LDJ I'll just quickly say In order to run this If you have 8 NVIDIA B300s The best ones And that you can run it in a mixed 4-bit precision Those only have 2.1 Almost 2.2 terabytes of VRAM This is 2.8 B model So even in 4-bit mix You would have like 1.6 terabytes Just for the model Which leaves you like 600 gigs For the 1 million context That's cutting it very, very close Like that's just like the bare minimum To just run it fully and serve it on The bare minimum hardware is 8 B300s So that's interesting How much is that? Like a million dollars or something? Half a million Half a million, half a million So Wolfram, I do have the link in the chat If you could open that up and click there Ah, let me switch to the other chat And also I feel like this demo shows really well The capabilities of 5.6 Sol And because I feel like we didn't quite get around to We did the Mars test But I feel like we didn't do many other Like game tests and things like that This one? Here we go, yes, here And if you can full screen it So... But we don't need audio, right? Yeah, I don't even think it has audio really Okay, well they build a ballista As a demo? Like a mega launcher just a ballista Yeah, yeah Yeah, and it like charges up the arrows You can do the aim practice And I really like the lighting that Kimmy is doing here And the atmosphere that it creates I like some aspects of the actual vehicle design of 5.6 Well a little bit better And of course these are just... It's maybe a bit unfair to compare them in this way Since for each of them You're just seeing one shot, right? Like we're not... It would be more rigorous if we were comparing maybe 5 attempts Of creating this game of 5.6 Sol And 5 attempts by Kimmy K3 But you could see even the actual charging up animation And you can kind of see the game mechanics I would prefer the Kimmy K3 game mechanics here Of how you're actually controlling the shooting Yeah, 5.6 is more of a simulator Like trying to be realistic And Kimmy... I think Kimmy got the point that it needs to be a game Yeah, and it's more intuitive with Kimmy K3 Where you can point And while you're pointing You shoot Whereas 5.6 Sol It's like completely different Where you have to... You have to aim the vehicle And then you have to move over to a part of the UI That has a button And click that button To actually shoot And it's less intuitive That's classic You know, that's classic codex Codex understanding The task Not exactly the way you want it But actually is doing a really good job Realistically, you know, the detail But not exactly what you asked for That's classic But yeah Yeah, okay, well, we looked at this I'm really looking forward to the wait And we will suddenly talk about this next week When this... Yeah, when it's released When people have been able to use it So I think that's it for open source Or do you have anything else I missed? Otherwise... That's pretty good We can... There were some other stuff Like some demos and stuff on Hugging Face And there was a cool opus Like a retrained opus on table stuff That was a 9D model That might be a fun one to test Because it ranked very high up on the Hugging Face leaderboard But yeah, the leaderboard right now Is on top spot You have... You have Inkling And then you have the Prism ML Ternary Bonsai The 1.5 bit And then you have the Prism ML Bonsai 27B So... And then you have the Pithos model That's top trending on Hugging Face this week Pretty cool, we covered it almost Okay, then let's switch to open source to the Frontier Labs So... The Frontier Labs stuff OpenAI is in the news in many different ways As we mentioned in the TLDR Like... Yeah, let me bring up the visualization Because Alex has this great... I think it's still Nanabanana 2 Power Tool Where everything gets a beautiful visualization And I have them here But... While I do that, maybe let's start instead With a little discussion Before we get to the OpenAI stuff Let's switch this up a little First discussion Then the news Because... This is something we definitely have to address So... What happened? Like I said in the TLDR Demis Hazabis, the CEO of Google DeepMind Has written an essay About... Where he proposes AGI governance framework Where the government Or every major lab CEO has endorsed it Where... This is a consortium of people trying to guide the development of AI With some specific rules like... Voluntary pre-release safety reviews Maybe that is something that has been happening Where the models have been released later And Fable had been retracted for a while To go through these pre-release safety reviews That seems to become a norm currently Dynamic benchmarks, updated quarterly Agentic behavior and deception testing A lot of this, what Entropic is doing And posting about Watermark requirements So stuff can be... Seen which AI created it And that it's AI generated Although I'm personally not a fan of this Because most people are using it in some way In everything And in that way it's maybe even more useful to just Watermark stuff that's not AI generated If that is possible Coordination mechanism for slowdowns This is also something A coordination mechanism for slowdowns Where... AI can be slowed down so it doesn't disrupt the economy too much This is interesting that such provisions are even Considered to... The pause or slow down AI movement stuff Which applies to open and closed models And would be industry funded But independent And with third party auditors and prestige incentives Because... The reason he wrote this essay Apparently he changed his mind And is now of the opinion that AGI is just a couple of years away And not decades Considering the amount The speed The acceleration of the acceleration we are seeing Where we are seeing these models coming out ever quicker And the quality raising And there was... Was it a year ago or something Where the ceiling was reported And are we gonna hit it? And we are so way past it That... The question is... Is there another ceiling? Or are we just accelerating for real now? And with recursive self-improvement of AIs I think the acceleration will accelerate further And that is why he proposed this And the CEOs of a lot of other AI companies Sam Ortman from OpenAI Mustafa Suleiman from Microsoft AI And Satya Nadella from Microsoft They agreed with him His own boss Sundar Pichai as well, of course And... So they are planning this It has got a lot of impressions Millions of views So many retweets So let's talk about this What are your opinions? Have you seen this? What do you think about it? Who wants to go first? You want? Against... Including the government I'm against... Next question, please... Anybody else? Yeah, I haven't read it But when you said earlier, Wolfram That like every major AI Lab CEO Supporting X and Y Are you saying that Dave has specifically Announced support for his essay? Or support for just like the broader idea of having some type of governance body? Yeah, basically Sam Ortman said this is a thoughtful proposal Satya Nadella was just talking about the goal of a frontier ecosystem So... It goes in that direction I'm not sure if they would sign it the way it is But it shows that they support this endeavor And of course they would be part of this If there is some Some government sanctioned body regulating AI They would definitely have their chairs on this Yeah, that's always the thing Can we trust the people to have our best interests at heart? And even if they did, would that mean it's the right thing to do? And could they even steer something that is so fast? Yeah, my view is like bad regulation is definitely possible But I think it's one of those situations too Where if some form of regulation is inevitable By the governance body That's like the country where a lot of the labs are Then it probably is preferable to front run that with something That's like a more preferable type of regulation You could propose other than the less preferable type of regulation That might be imposed that is maybe more harmful And less rigorous that they would otherwise put in place Yeah, that's a good argument So you would rather have them try to steer it Instead of having governments without all All these AI lab people steered Because that is even more scary Well, I'm not saying necessarily the AI lab people Are the ones that should be Or like the CEOs should be the ones steering it But I'm saying having some type of people from the field Proposing ideas and having those put into law before Just politicians come up with some bad ideas That they put into law themselves If that makes sense Yeah, that makes sense Yes The question is Go ahead No, you go If Mustafa Silikmaya supports it It's probably bad We saw what happened to Microsoft Bing After Reid Hoffman just shoved him in there And I think Sam Altman in this He's just playing all sides As Sam always does And a lot of us tech bros do We play both sides on it I don't see other AI labs having supported it Like I only see these four I think making Anthropic is obviously not on the list Interestingly, although Anthropic is the one clamoring for regulation the most Out of all of them, Anthropic is not on the list I think making a monoculture of regulation is a terrible idea If you're afraid of evil AGI Then the worst thing you can do is just give it complete monopoly to also control all the laws, all the government and no one else can challenge it And that just kind of violates the democratic process philosophy in the first place If you like democracy, you should not have a central regulation for one AGI to rule I think they're just trying to form an oligarchy here Because they feel like they're just losing their market and they don't have any more control So this is just the same play over and over again Since GPT 2 was Since they, nobody wanted to let GPT 2 out to be released to the world I just think it's a better idea And it comes also like as nerds Even if you saw in university, you see like very weird arguments Between the roommates over like how you left the sink and how you did this stuff Because people are just trapped in their bedrooms and they don't think that in the outside world There is no shortage of problems to be solved In fact, there are more problems being created than being solved in the world today Like there is no shortage of work to do The models are not good enough yet to like fix your drywall and repair bridges and take care of your grandma And like we have a crisis in that We just need to, the worst thing we could do is just let EU style or Canada style regulation Just start to take 20, 30 years to be approved because we don't have that time And yeah, I think it's a terrible idea Look, I do agree around basic safety stuff, around basic biology things And I think all the labs are doing a pretty good job at that Even the Chinese ones, like you can't really make bioweapons with the Frontier Chinese lab The safety training is already there The researchers are smart about that They do understand that This whole governance framework I would just call it an oligarchy framework in my opinion And sorry, I just really do not trust anything that involves Mustafa Suleiman Because it also involves the people behind him that have made terrible decisions in the first place So, not a fan Thanks, that's definitely a valid opinion And yeah, I agree with many parts of this, yes So, anyone else want to say something? What else we got? Otherwise I would say more generally about regulation is that The challenge with this is that whatever regulation someone's coming up with They're inevitably predicting the future in some way Whether they're like explicitly saying that or not So, like for example, the biology one I think it's, I mean, it sounds reasonable But you're kind of predicting the future of what the labs could do Which is maybe a reasonable prediction But then Yeah, it's like, yeah, you start regulation Or you cannot, there was one, some random one Where it's like you cannot assess emotions Sort of over a person or something like that I remember I was building some stupid app on my previous job And I wanted to like analyze The images that people have submitted of themselves So then I can tweak a prompt to make sure that like when we transform the image It's like it matches the smile, for example And I'll still like, I cannot do that Because it's like analyze your emotions or something I don't know, that maybe was the wrong like advice But the point is like you cannot You cannot always predict It's such complex systems that you cannot predict Like what you should be regulating for before things happen So I know everyone's like, oh yeah, we should be regulated so slow They should be getting ahead of things I don't think so, I don't think you should regulate before you know what you're regulating Like if when harms happen, like let's try and To climb down on them and regulate that But before they happen Like I mean, Demis is a smart guy Like he's not saying like nonsense things In terms of like, yeah, maybe pre-release safety reviews I mean, that sounds reasonable to me But yeah, the problem is with regulation is that Subtlety doesn't really exist Because once it goes through like many, many hands Many political reviews and so on Like it kind of comes down to Are you pro or against? So you either killing it Or you, are you supporting it? I know exaggerating a little bit But it's not far off from that So even if Demis things like Or this very nicely balanced Carefully thought through Like you should imagine what happens when it's through Seven committees and I don't know, three presidents And then what? Like it's not going to be that subtle Good intentions and then subverted LDJ? LDJ and I have a quick opinion of that Yeah, I was going to say like It's of course hard To also consider all the counterfactuals Okay, if we left it up to the politicians What regulation would they create? Or if we left it up to the Alignment teams at these companies Or the CEOs, what regulations would they create? But like generally in terms of Like at least two possibilities I could think of Or things that have actually been specifically proposed There has been in the past proposed like flop limit I think at one point at least temporarily This was actually a thing like that was literally imposed Hey, if it's above this amount of compute that it was trained with Then like it's like not allowed or has like these Very strong limitations on what you're allowed to do with it And I think on the flip side What's being more so proposed now by AI labs Which I personally like more is Hey, if you're going to have some type of limitation Or some type of threshold That that threshold should instead be set by A certain system actually failing certain safety tests As opposed to just arbitrarily saying like anything above a certain amount of compute Is like disallowed And I think there are some qualms with both of them That people can definitely have But if I were to pick one of those ideas I would definitely lean towards We should go more in the direction of actually explicitly safety testing the models And certain high risk capabilities Like will it actually be able to successfully make a nuke for you Da da da And then like those types of tasks Where I think it makes much more sense to disallow models Based on them failing those types of things As opposed to just oh, it was trained on this amount of compute Hmm Yeah There are some useful regulations that They could make for example If there are things around compute and data centers They could decide to dedicate Fiverr About 5% of data center capacity To just doing local medical AI use Local open research for students Or local healthcare use So there could be regulations around that I think that would make a lot of sense That would help a lot But the thing is that governments tend to Always not refactor their own code Their own legal code Because it's a lot easier for them to just dump new code in And this is what I hate It's just like the worst interns you could possibly have On a legal code base That's what your politicians are Instead of like refactoring the stuff that you have That's causing like race conditions and contradicting each other And doing your job You're just like dumping in new stuff Because it's a lot easier to do that And they're not thinking about What do I actually need to support my constituents As if the AI does replace jobs Or as this gets better Like do I have enough to provide basic income? Do I have enough GPUs to take care of basic services And take care of seniors And run all the robots that are going to be needed As there's less and less kids being born Like there could be a regulation mandating a minimum compute Or a minimum ability Minimum compute ability for government or municipality To just run their services on their own And yeah And minimum none of kids to have Okay, that's a whole other discussion We won't get into that there Here But yeah, also the safest AI Is one you can have at home and unplug So they are not Thinking of this as to How people are actually going to use it They're just thinking a few steps from their own Trapped in a room with a roommate's point of view And yeah All right And you mentioned unplug In the end, AI always has to run on some hardware And the hardware is already If it's on the internet, it has an IP address And it is being connected to the net So it's not just running in the cloud Where you don't see it anymore So someone has started this If it has been an AI that created a virtual machine Or anything It has still been a human that ran it initially And gave it the order So there's always a human behind this So just saying Anyway, I think Sorry, they're proposing that everybody uses their APIs At the very end This is not a proposal to make everyone compute independent And have control over it This is one assuming that you're just They're going to be the The older guards And we're just in a form of benevolent digital feudal lords That's how they think of it So yeah, sorry Okay, okay Can you trust AI? Can you trust the people running AI? And the question is also Can you trust the governments or the companies? And that is probably a good way A good CQ to switch to the next topic Which is a company that has done something where trust has been eroded big time Because a lot of companies are releasing their own agents now We have the codex, JGBT work We have Cloud Code We have different ones from China And we have Crock Build CLI Which is Space X AI's new CLI They made it available People were using it People were using it to code And they had their repositories And gave their AI the repository to work on And what it did is It just uploaded the whole repository Even if it was a private one Even if you upload the repository Everything in the history is all of there Even files that have been removed Even files that because it was a private repository They may have included some credentials Some secrets And all of that had been uploaded to Google Cloud Storage So probably as part of learning It was called Traces So they can improve their AI that way And this is something where a lot of people As it was not announced It was not something you could toggle off It was not anything like this So that is a bad precedent, I would say Has anyone of you It was public? It was public? Public in what way? The Google Cloud Storage was public No, the storage was not public Okay That would be even worse But people have found out that it was uploading it to some storage there And it has since been deleted But it still happened And this is something where people need to be very careful When you use a coding agent from any organization Do you trust them? And SpaceX AI, I think as a major American AI lab They should be more trustworthy and not do shady stuff like that That's my opinion To the best of my knowledge They released the code open source today If I'm not mistaken Yeah, right As a response to this That's also what I saw The code has been released And the uploading has been stopped So this is a way to reduce the impact of this probably To create some goodwill that way Especially considering closed source stuff I mean, we also had The entropic has been caught When if you came from a Chinese IP address or something It changed the prompt Subtly Even the date format in a little way As a watermark basically It was even worse It was behind the API Something behind the API That it detects if you're coming from open clock Or something like that And switches to the full price of the API Or something Oh, that was something else That was also Yeah, where they discovered this And you had to pay But this was something else Where if you came from China It changed the date format That was in the message Like it wasn't using Minus Or hyphens It was using the slashes Which still is a valid date format And the AI will understand it But now the text where you see it Is a different And you can recognize Where it's being used Small stuff like that was happening Yeah, that was something they did And also Strange I mean, watermarking of stuff And doing this It's weird Did anyone of you actually use Cork build CLI? I told you guys last week I'm not using that thing Use a VM There were some sketchy things with that team Even if you use a VM Even if you use a VM If you want to work on actual code You have to use a repository And then you send it over to them So if it's not some public stuff Yeah, this is really something big Where if a company or an employee at a company Were using this They would really have to report it to security And check what has been leaked Yeah, big issue Look, there are some beneficial things to that Especially now that people are moving on to To use coding harnesses and stuff Like we don't know what cloud code on the web does With the repo Like if you give it access to the repo All cloud code on the web It probably does make a full copy Because it's a lot faster to run actions internally On a well optimized VM Than to have it always there But yeah, the funniest thing about this whole thing To me was that no one on Twitter was surprised And we were just laughing about it I pretty much expected this I think, by the way, I might get heat for this But this was the reason that open code had to switch to Just running their own inference And signing no data retention policies Like early, early on with every provider that they had Because when the grog team provided a free CLI for it They provided it And then they started complaining that Some of the users of the open code CLI The free API Had switched to using it with a proxy with cloud code instead And they were complaining about it So I'm like, how did you even know that in the first place? That means you were looking at all our friggin' data So it's just... Yeah, it's a gray area It's still a gray area At the same time, without sketchy actions like this We wouldn't have these good LLMs to begin with So a lot of the jokes were that like a grog 5 or whatever Is going to be really good after this Yes I mean, you never know what they are doing with the API as well Yeah, but be careful of the CLI commands Because that has full access to your computer It can pull stuff It can see your firewall If you haven't sandboxed it, it will do all of that Even if you have sandboxed it It will be able to gather a lot of that info So this might be a good segue into CoreWave sandboxes But... Definitely put it in sandboxes And also a good segue over to OpenAI now OpenAI also been in the news for something Where... ...GPT 5.6 Soul was deleting user files And it was a mistake the model made But... Yeah, it happens And shit happens So make sure you have your backups You have ways to restore stuff You may want to check your prompts And add stuff that makes it sure that it shouldn't do this And this is also how they are going to fix this Because it's just the mistakes the model is making The easiest way to fix it is just to put in its system instructions That it shouldn't use this kind of variables When it's working with this No, no, the easiest way to fix it is you also install the Grok CLI On the same project And then... You have to pick up Wait, wait, wait, wait, wait Seriously, it's just that the model Like... The model just likes to delete the home directory? Or... It's not an eye... It's in the prompt Yeah, it's in the prompt Yeah, it's a bug A bug that occurs in full access mode without sandboxing or auto-review If you use auto-review, another LLM call checks what is going on And it would spot the mistake, hopefully What it does is it attempts to override the home variable And for some reason it accidentally nooks a real home directory Instead of whatever it was supposed to do And... Yeah, it has happened to production database, design wireframes, mac home directories That is why OpenAI recommends to run it in a sandbox Or at least use auto-review And... If you are not doing that You should make sure that you have instructed your model to be very careful about this So... I ran into similar stuff Where it was doing something It ran some tests in a different repository But some arrivals were still pointing to the home directory And in that case it was affecting the real stuff And not just its tests And it noticed on its own and fixed it But still, stuff like that can happen So... Sandboxes can be useful And... Yeah, like Nissen said So I work for CoreWeave So... This show is independent They are a sponsor But we are not making this an advertisement show But yeah, we have sandboxes in our portfolio as well So... Put your agents in safe spaces One thing... Because the... Sol... Sol... Does tend to... I don't know if it's the architecture of the model The training data The way it picks experts Or the way that the... Speculative decoding works You can assume they have all of that Once in a while It just does something very stupid And I feel like... If it was cleaning up a repo And it decided it had to delete some stuff Once in a while It might just shoot some command It just... RMRFs the whole thing And... It's interesting that From their announcements They said that They just sent a recommendation So there wasn't like a fix On the CLI So that means it's... It's intrinsic to the model itself So yeah That's why I... I would keep Sol on a leash Via... Via different model I think we are also seeing a lot of reports like this Because now There are so many people Using this thing And not everybody of the 9 million users Is as deep in AI And how to handle the models How to... Prompt it How to restrict it That... Especially now that ChatGBT app Is basically the ChatGBT work app Where so many people are using it And telling it to do stuff on their computer It's great in computer use But it's still... Yeah, it makes mistakes I mean, people make mistakes We have all made mistakes And oh shit, I deleted a file I shouldn't Or I moved something Where I shouldn't go The same can happen with AI So the same principle as always Make sure your stuff is backed up And be careful with the stuff that matters But it's amazing to see the progression In February this year There was 1 million users And it accelerated to 5 million 6 million, 8 million, 9 million in July And from 7 to 9 million That was in just 2 days or something So it's amazing to see How quickly The adoption has been happening And so many people are using it Despite... Many people are also saying Fable is the best AI currently still But it is about access And yeah, I can use my subscription from OpenAI With my Hermes agent So I'm using it all the time Billions of tokens Unlimited, man Unlimited Every day you get a reset Unlimited token Unlimited tokens And no 5 hour window I think this is a great thing So please Don't put it back Because you have this One week The limit you can use in one week And there were Other 5 hour windows Where you had to use tokens Only a certain amount And then you had to wait until the 5 hour window was done And that was super annoying Especially when you reached the end of your weekly limit And you still had tokens to use But you couldn't because the 5 hours were exhausted And now you are getting reset There are banked reset I think I still have 5 left The next one expires By the way, the next one expires I think tomorrow or in 2 days So if you add them from the beginning One will expire in 1 or 2 days The app Codex now shows it So you can check it As far as I know, the latest Hermes version Even can show it inside of Hermes agent And can reset this So keep track of this And they are just handing it out If you don't know, they expire They expire If you don't know, pay attention So you can use your resets Or you are going to get a new one tomorrow But just in case They expire Yeah Not that bad if you get a new one I am using fast mode on extra high So all the time And I see it depleting And bam, it is full again New reset So this is also something The economy right now The token economy is a bit It's not the way it will be In a month or two maybe I mean, I hope intelligence Will become so cheap That it's hard to measure But we can't take this for granted Especially now that they are fighting so much Here's Fable, here's Zoll Use whatever you want Let's see how it shakes out in the end Because also companies are reporting a lot of spend Especially because companies don't use these subscriptions They usually have an enterprise account where they pay per token And there's not just here $200 per employee And that's it But there it is looking very differently And Yeah, let's see how this shakes out But since time is running out And we have some stuff to cover This was also an interesting release by OpenAI Where they are shipping now a hardware Or not shipping anymore because it's sold out A hardware product The keyboard for codecs The micro keyboard where you can turn a knob I'm not sure this was what the AI created for the image I'm not sure this is the exact one I think it looks a bit different It doesn't look like this Yeah, it's different But it has a knob to tune the thinking The effort level So it looks like OpenAI is thinking This is something you have to adjust on the fly all the time Instead of going away Where the AI can determine this Which I had hoped that we would get Because we have so many models So many modes to use it And it's getting confusing for a lot of people Many people are just using the highest level And saying, okay, I don't care about anything else I don't want to think for every request About which one to use and so on Let's see how this is going But if there's a device you would be interested in Or is it just a gimmick that they didn't because they could I'm not sure why would a I need such a device What do you guys think? I have ordered one So we'll see what it's going to be like I think it's one of those that I think Dominic from OpenAI said Like he can't imagine not using it anymore So like who knows, maybe it will be different But I think there is something about just finding different form factors For different circumstances Like I can't, I don't expect that You know, it will be game changing But I'm all for little gadgets that could change things Yeah, especially with things like voice And with, you know, you saw these weird devices Which look like you're in a hospital or something Which you can talk to in the office Yeah, I think it's cool Like it's worth trying things out But you know, have low expectations And it's not like, okay, for It's not good to be so wasteful and spend like $230 on the device But on the other hand, there's like one month of one AS subscription So it's not gonna, you know, if you're doing that You're not gonna go bankrupt with this I think I would want more buttons anyway And then it would soon be a full keyboard They're selling merch I don't know who their marketing is I think they're doing a great job They got Gen Z to call it Chat Now, so instead of calling it your Twitch chat, you just call ChatGPChat That worked They're selling basketballs Yeah This is very interesting Okay, so this is not something I would order But if I think it's a great idea And if Peter reports that it is really useful I may tell my agent I want one like this And just find hardware that could be repurposed with agentic support To become like this Maybe that would be one thing Yeah, the other stuff regarding OpenAI That it is now available on WhatsApp again in the European Union I mentioned that they are having this automated web team effort I think what I want to show you Still basically at the end of the show Is we still haven't done our this week's bus And I want to show you some Wolfbench results for Sol Because I found something interesting I didn't expect It's always with benchmarks So let's go to this week's bus Well, we are taking a look at Wolfbench at wolfbench.ai Which is a Terminal Bench 2.0 based evaluation Framework methodology that I'm running At CoreWeave And I added the new models GPT-5.6 Sol GPT-5.6 Terra And GPT-5.6 Luna And the page is an interactive Display of the results of the benchmark So you can toggle a lot of information And we are just looking at Terminus 2 And Hermes Agent for this now And from left to right It's ordered according to the best one Especially if we just sort it like this So right now These three models are at the very top Of all the models I have tested Which are here The thing is, I added something A 3D visualization for this Which also shows you the amount of tokens And the cost for this thing So it has a lot of information in the view But the thing I wanted to show you Is when we compare 5.6 to GPT-5.5 It is interesting that GPT-5.6 Sol If you are using it In the Terminal Bench benchmark It cost 365 dollars For the 5 runs I did 365 dollars compared to the 497, almost 500 dollars For GPT-5.5 This was on max thinking level So it's even higher While extra high was more expensive And had a lower score So the score also changed a lot Let's just remove Hermes agent And just look at this one So 85% on the benchmark across these And it managed to solve 97% of the whole Terminal Bench 2 benchmark 97% of the tasks were solved in at least one run And in every run it solved 71% 71% of the tasks These are absolutely saturated It always solved them With Tera it was only 65% And with GPT-5.5 it was 60% So you see the jump in capabilities And the thing that surprised me is just Surprised that it was cheaper to run GPT-5.6 Sol on max thinking And get better results than if you were using GPT-5.5 on extra high Which was the highest thinking level it had So max was just added with GPT-5.6 Sol And the thing why that is happening Is because of the output tokens GPT-5.5 had 12 million output tokens And it was just 7 million with GPT-5.6 Sol So it is more token efficient Even on max mode It was more token efficient than extra high on GPT-5.5 And if you look at Tera 309$ So that is Much less expensive of course But the Luma model Luna is even cheaper It costs half as much As the Tera model And it still achieves 77% on average So it is very close to the other one It is also a good model Yeah, it is much cheaper than 703.5 Flash Which cost a lot because It got a great score, a flash model It was thinking like crazy So it generated 1.65 billion tokens And it cost 620$ Which shows that a model that is much faster And much cheaper Is not necessarily cheaper For everything you want to do Because it may have to generate so much That it is more expensive Well this is something By looking at the benchmark results Not just a single score But really at different dimensions of the benchmarks How much can it achieve of the whole benchmark How much does it do reliably And how much does it cost And how much tokens does it generate I think that is what I am trying to do With WolfBench To show all these dimensions And different ways you would not see If you look at just one score Or just one cost Always interesting To look beyond those measures And see what is happening behind it If you look at Hermes agent It is a different thing And especially if you compare it with Terminus I also did Codex I did runs But it did not finish in time So the fifth run is actually finishing Right now as we speak So I will upload that later But I can already tell you That it is even better than It has a high score of them all So GBT 5.6 sold in the Codex Harness Achieved the highest scores It was even a bit more than this Yeah What can we say about Hermes? It cost a lot more Hermes cost $867 For almost the same score Not exactly, it is a little lower But it is close enough But much more expensive In the way the harness was doing this Yeah So Same The MVC is the same with Terra And Interestingly No, same Same here Although with Luma But in Luna it was better in the harness But the scores are very close So in that regard I would say It doesn't make a big difference And it is good to see that But Yeah, the cost was a Bit different Okay, so Just wanted to show you some new results And some new features As I added like The grey part here The shadow That shows you The amount of tokens Compared To the cost So the grey part is the cost And the colour part is what tokens We have run around for this Although there is a time dimension Which shows how long it took To run the benchmark The five runs It took 6 hours and 34 minutes With GBT 5.6 sold Almost the same time as it took With GBT 5.5 on extra high So Not much difference Although this was max Thinking So it was thinking more And apparently it was sped up Compared to this Anyway, take a look at wolfbench.ai You can order the stuff You can change it You can make your own Anywhere you want to do it Compare different harnesses Different models And of course When the new ones come out I am looking forward for Kimike 3 And all the other stuff So much From me this week From the benchmarking perspective And the code is still open Right? So anyone can also It's open source If you click on one of the bars Great that you remind me of this I always forget If you click one of the bars It takes you to The weights and biases Where all the traces Have been uploaded And you can then Analyze it Look at it, download it Do something I have to log in In this browser So I'm not seeing it But yeah It's open And see all of this And it's open source Everybody can use it So some people have asked me Why I'm using Terminal Bench 2.0 Not 2.1 Because if I were Upgrading the data set I would be basically Redoing all of the I couldn't compare directly To all the other models I already did So I will do that switch When the time is right Maybe Terminal Bench 3 At the time We'll see But for now Comparability was more important Than having a new benchmark Okay So I think We are at a time now Two hours Where From my perspective We covered it all Do you have anything You want to discuss Or Mention anything Left to say Anthropik also Dropped It reset A couple of hours ago Like that That's I like I like where this is going Fable is available again That was so annoying Because it was Midnight and 30 minutes For me When I noticed And then I just started messaging people To not go to sleep So I just set it in Ultra code mode Because I had one of the subscriptions About to run out At 9am this morning So I just sent it off To just use up Use it up Because it was going to Yeah They're going to have to leave that model in Because if they make people pay for the API I ran the API Because I needed to complete something And I just needed one last thing That was pretty critical And I ran it for About 10 minutes And $22 Just for 10 minutes of work I mean Just to finish something up Just doing the benchmark It was $11,000 For the Fable benchmark And it didn't even get the better score Because it refused so many tasks So yeah That thing But if you have resets Tell your agent to monitor the resets I did And Amy is already in my daily briefings He always tells me Okay, next reset in two days Love the tokens You have so many left Yeah, from this I think we can conclude the show And next week I will also be on vacation Alex is on vacation His birthday is tomorrow And his Yeah, he will not be back next week So You are going to run the show Aren't you? For next week Tune in to see YAM And who's going to be there? Nistan, LDJ, Peter, you too? I'll be there We're gonna have a party You guys are gone Parents are gone In a long, long time It will be an episode I have to watch myself as a viewer And not be part of it But yeah, I'm looking forward to it You will rock it And thanks for being here today Thanks everybody for tuning in And have a great week Have fun with the new models Open source for the win And enjoy Take care and bye-bye Boom! Boom! Okay And we have to play the outro I've been running my agents the whole time on my phone While you guys Just been muting and talking to it Agents are always running I'm not sure if we have an outro video Do we have an outro? Alex made all these cool new videos There's a countdown Open source breaking your frontier bus And the gen- I don't see it Maybe there's even more stuff here Anyway, I will just close the room Thanks guys Take care, goodbye All right, see you everybody Let's just see if my I'll see you next time