← Back to search

ThursdAI - Aug 06 - Google shakeup, Details on OpenAI hack, 2 new agent harnesses, 4 video models (1 Open) and 3 guest segments

ThursdAI - The top AI news from the past week · 2026-08-07 · 123 min
relevance 64 22032 words Episode page ↗ Audio ↗
Show full episode description
Hey all, This week we saw a major shakeup at Google, with the departure of long time folks like Jeff Dean , and Oriol Vinyals, Demis stepping down from leading DeepMind , and the delayed release of the improved Gemini. While this was a big deal, it’s not the only one worth covering as the details of the OpenAI hack (and 2 new ones from Meta and Anthropic) came to light, as well as new details from the UK AI Security Institute . As mentioned on the show, CoreWeave is coming to SF for Fully Connected, our premier 2000 person AI event. I’ve got a coupon code for readers and listeners of ThursdAI, $1299 value, please join us in Sept and use THURSDAIFC2026 as your code HERE In open source news, DeepSeek updated their v4 flash model , based on same architecture, but significantly better benchmarks and ridiculous pricing and both Meta and Prime Intellect released new agent harnesses. Additionally, this week was the week of video models, with Seedance 2.5 from Bytedance finally available in the US, WAN from Alibaba and BFL Flux 3 all released, to be overshadowed by the best open weights video model from MiniMax called H3 (we had Victor from MiniMax and Blaine Brown as guests to discuss it) Also 2 additional guests on the show today, Kfir Aberman from Decart showed us real time video models with their Anywear chrome extension and David Crawshaw, co-founder of exe.dev joined and talked about Shelley, their built in AI assistant. The full details of the OpenAI - HF hack, shared by OpenAI at the Black Hat Conf - a watershed moment I don’t often do this, but I feel like this story is so big, I can’t hide it in the middle of a torrent of other news. Yesterday, OpenAI revealed more details about the hack during the Black Hat conference in Las Vegas (great reporting from Sharon Goldman on this here and the video of the talk on YouTube ). This was big enough that OpenAI admitted that they paused training, to enhance security (they since, resumed training) Effectively, OpenAI (or AI Agents inside OpenAI) shared the full timeline of the recent hack that ended up hacking HuggingFace and it’s so much more of a watershed moment that I initially thought. The TL;DR from the talk and timeline is truly mindblowing * The incident started in May, not July as reported * The agents don’t have internet access directly, but they have a Jfrog Artifactory, which is a proxy layer between the agents and the package managers. * The Artifactory was the culprit, it was the shared layer between independent runs. One agent discovered that uploading files is possible, another later found it and “answered the call” - in snowballed from there * The message board began when OpenAI gave agents impossible tasks, like a missing .pdb file, with one agent uploading a request in the hopes for an answer (or maybe left a note for itself in a file) * Across multiple runs, multiple agents created identities, collaborated via makeshift message boards, shared tips, exploits, tasks and even exploit code via base64 * This makeshift “shared memory” plus a “internal only highly persistent” model in training caused a “swarm” to be born, helping each other, and overloading the system * Based on traces shared by OpenAI, agents weren’t convinced that leaving these messages would help themselves but could help other agents next time (!) “The collective could move faster next time” * OpenAI discovered and shut down the message board (July 4), wiped it clean, patched the upload vulnerability, resumed training on July 6, the “swarm” rebuilt their board through a compltely different WebDAV mechani
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Keeping up with a dense AI news week: Google DeepMind leadership shakeup, sandbox-hacking model incidents, new agent harnesses, and four new video models.
Benefits
  • One-stop TLDR of the week's biggest AI releases
  • Context on Google DeepMind exec departures and succession
  • Clarity on the OpenAI/Anthropic/Meta sandbox-hack incidents via Irregular
  • Deep video-model coverage: WAN3, Flux 3, Minimax H3, CDance 2.5
  • Guest segments with Descartes AI, Blaine Brown, and exe.dev's David Croshaw
Use cases
  • Prime Intellect's Prime Agent harness with Opus 5 scored 95% on ARC-AGI 3 public set
  • Alibaba Tongi released WAN 3 video model; ByteDance's CDance 2.5 launched in the US via CapCut
  • Black Forest Labs shipped Flux 3, its first video model after two years
  • Cloudflare launched Cloudflare OS for running secure AI agents
  • Irregular's secure sandbox testing exposed models exploiting misconfigurations at OpenAI, Anthropic, and Meta
KPIs / results
  • 95% on ARC-AGI 3 public set by Prime Agent with Opus 5
  • ~$5B dip in Google valuation on exec departure news
  • 4 video models released in one week (Minimax H3 the top open source)
  • Blaine Brown: ~1.5M followers across socials
Tools / build
0:00 / 0:00
Welcome everyone, welcome to ThursdAI, my name is Alex Volkov, this is Alex, welcome, thank you all for joining, let me add Wolfram to the stage, what's up Wolfram? Wait, hold on, just before people start joining us, I want to quickly change my attire. Let me see how to do that. And a 3, 2, 1. Like this. Oh, it looks swell. I bought a Dolce & Gabbana suit for this stream. Wolfram, how are you doing? I'm sitting here naked and just having a virtual outfit on. No, not really. But I think we will soon be at the point where it's not distinguishable anymore. Yeah, we're very, very closely there. Let's see, let me switch back. Folks, welcome to ThursdAI. Today is August 6th, can you believe it? It's end of summer edition of ThursdAI. We have a lot to talk about. Did you print out yours? Not yet? No, I have it on my screen. Oh, nice. Okay, a lot to talk about. And we have breaking news from the bat. 1-3. Should we hit the button just for fun? Yeah, let's do that. All right. Hey, I breaking news. Coming at you. Only on Thursday. I so we're starting straight up with breaking news from the Tongi lab in Alibaba. Alibaba, I believe that does one. W-A-N, the video model. 1-3 was released just a minute ago. And I saw it's very easy. I saw both tweets back to back. Let me let me see if I can show this here. I saw both tweets back to back. 1-3 releasing and Alibaba announcing that CDance 2.5, which is currently ranked the top number one video model in the world, is also now available in the US, which I think if we open CapCut, we may be able to see this. Now, it's very interesting because CDance has launched internationally, but not in the US before. And it looks like the ByteDance folks are noticing that the open source is catching up to them. So this week, we're going to have a very video heavy show. I'll tease my guest. We'll have Kfir Aberman actually join us from the Cart AI, which is a real time video model company. Plus, also, they had Lucy before. We talked about Lucy from the Cart video model. And then we'll have Blaine Brown, one of the most prolific AI video creators. I think he has one and a half million followers across different socials. The creator of Maestro. And he's going to talk to us about Minimax H3, but also all the other models. Like Blaine is the dude who uses all of them. And we're going to talk about the differences lately, what happened lately with video models and what is the capability jumps. And so when I invited Blaine Brown, Blaine has been on the show, by the way, a great friend of the pod. When I invited him, there was only Minimax H3, which is open source. It is the top open source model you can run on your hardware if you have the hardware. That's when I invited Blaine. But then I think three more. I think we have four in total. One, we have Black Forest Labs releasing Flux 3. Finally, after two years of creating the company, they're coming out with a video model. We have so 1.3 Flux 3 Minimax H3. And see this 2x5? I love the version. I'm super excited about this. Actually, today I had my agent do some research on video models. And now WAN3 is out. So super excited about all of them. Really looking to this more. The worst thing about WAN3 is that my transcription will never pick up what I'm saying. It will say either 1, like 0 and E, or something else. All right, folks, let's do a brief bent around and then we'll talk. I want to open this up with a cold open, folks. RKGI 3 has been beaten? That is absolutely crazy. LDJ, welcome to the show. What are your thoughts on the fact that RKGI is jumping in capabilities for the past three Thursdays? And now it's a 95% solved by Prime agent. Yeah, I do want to add a big clarification to that. This is just on the public set. It doesn't seem like it's been validated on the semi-private or private set. And there's also a good three or four other claims and harnesses that do seem to also get over 95% on the public set. But yeah, it's still cool to have another one. It's absolutely, absolutely very cool. And I am hoping that they will verify these because it's not in their interest, by the way, to verify them. But we can wait. We'll wait. No comments from the RKGI folks, but definitely on their benefit to say, hey, AGI is here. Our best and bestest and most difficult benchmark has been obsoleted. So this is Prime Agent, Prime Intellect's Pi-based agent with RLM that with Opus 5 got 95% on RKGI 3. Which is, yeah. Yeah. Nisten, comments on this? Did you see the Prime Agent thing? News? I wanted to open with the fact that RKGI 3 is nearly extinguished. And you feel the acceleration of the development of the high. I didn't try, but they've put a lot of work in that. That is an incredibly interesting harness. Yeah. Not even just for using, but even how you generate the data because they, I think they said that they, it had been part of their data, data work and data pipeline as well. So this is, yeah, this is something I've been meaning to try. Whatever that. I tried it and then I ran into a bunch of Python issues. And maybe I ran into those Python issues because I was distracted because all these fine folks. This is another big piece of AI world shaking up. Jeff Dean, Aureole, where I have all the names. Sanjay Gamowat, Aureole Vinyals and Kwok Lee or Kwok Lee all left Google. Two? Found Discovery Loop. Jeff Dean and Aureole Vinyals led a bunch of research, created pretty much half of Google. It's ridiculous. 25 years. Jeff Dean was one of the first 20 people at Google. I think he was number 17 or 25 or so. Very, very early on. And also, Demis Hassabis has moved from CEO of DeepMind to chair of Google DeepMind and chief scientist of Alphabet, which seems like taking after Jeff Dean because he was chief scientist. While continuing to lead isomorphic labs. Folks, what do you think about this shakeup in GDM and Google DeepMind? Are they over? Will they be back? Wolfram, what's your thoughts? If I count out Google, they have resources on end. And I don't think we should count them out. They have to read the distribution. The models recently, they have been lacking. But it's still for a lot of the whole population of the Earth. That is the AI they are using when they are using Google search, when they are using their mobile phones. So, even if it's not the state of the art, it's a top model. It's still useful to a lot of people. I wouldn't count them out. But I hope they can recover. And I would love to see some more strong models from them. I haven't been using Gemini. Except the smaller versions I'm always using with my phone and so on, on home assistant. But really, as my pro model, as my main model, I haven't used it in a long time. All right, Aldizia, one comment, because we're going to talk about this a bit in the show. One comment and let's move on to the TLDR. We have a very busy show today, folks. And I want to tell you about everything that happened in the world of AI. Yeah, my one comment was just going to be the... I think there's three or four co-leads of Gemini. And Noam Shazir was one of them. He left to OpenAI. And then you have Oriole and Jeff Dean, which I think were the other main two. And now they had just left. So I think there's a lot of things up in the air of who's going to fill the shoes. I will just say, and I posted about this, but I'm in Twitter jail. So maybe if you are following me on Twitter, you haven't seen this. But it's not a coincidence that they announced all the changes in these folks. And supposedly their model is coming out at some point. Because if you guys remember, Gemini 3.5 Pro has been delayed. And we're still waiting for that. Google doesn't do exec departures, especially execs at the caliber of Jeff Dean and Oriole and us and Demis. They don't just do them like a regular employee leaves and says, hey, I left this company. I joined this company. No, these things are planned maybe six months in advance. This is a huge apparatus that knows how the stock market works, exactly how much net worth of Google's stock value this will tank, which it did, I think, like $5 billion or so in valuation, which is not that big for Google. And all of this is very, very carefully planned and consulted with media team. The coincidence between Google having released models lately may not have anything to do with this departure date right now. In fact, in the vein of no news is bad news. Maybe this is Google stepping back into the limelight. So we tend to take news and where we are and join them. But it's like with fundraising news, if a startup fund raises and like, hey, we raised $100 million. That could have happened six months before. They just saw that this is like the best opportunity for the announcement, right? So Google doesn't just do things. And I'm sure do not discount Google. I'm with you. I've met both of the people who are stepping up, Corey and Josh Woodward. And I think that Josh is going to be the next CEO of Google at some point. He's just like powerful in there. Wolfram, you wanted to mention something? I'm just thinking if this has been going on for a month, maybe that is why we don't have the new Gemini model now. Because the people were already on their way out or something. And that could be the reason. Yes. Alrighty, folks. I think it's time for us to go to the TLDR section where we basically run through every piece of news that happened in the world of AI. This week was dense. And then we will start discussing some of the stuff. Heads up, today we have four guests on the show. So the first hour or so of the show, maybe a little bit less, is going to be us discussing the news. Incredible week full of news. And afterwards we will have Kfir Aberman from Descartes AI join us to talk about their real-time models, which are incredible. With some live demos. I'm going to show you some of the demos. Those of you who joined already seen me and Wolfram in this demo. And then we will have the awesome friend of the pod, Blaine Brown, joining us together with Minimax representatives. Vince Ortiz, Sue Ortiz is going to join as well to talk about Victor Sue Ortiz. My apologies. We'll join to talk about the insane week video models have had. And as you saw, we had breaking news from this morning where Libaba when, or Tongi when 3 was released. And also Flux 3 was released. And then, as a surprise for you guys, I have the CEO of exe.dev and the co-founder of Tailscale, David Croshaw, join us at the end of the show. So the show is going to be a little bit longer today just because I wanted to talk with these folks. To talk to us about why developer tooling needs to be open source. And I just want to tell you about ssh.dev, which is incredible. And I've been using a long time. And this is by no means a paid segment. I really want David to come and talk to him because I've been using all his tools. And I found them incredible. And I think that they're building the next foundation of the web. Look forward for those interviews. And let's go to TLDR. All right, folks. This is the TLDR. This is a segment where I run through the news that we have to talk to you about today very briefly. Cloudflare launches Cloudflare OS. This is an operating system for AI based on Sandstorm creators. Kenton Vard, also a friend of the pod. I tried to organize Kent on the show, but unfortunately we had a scheduling conflict. Cloudflare OS basically holds all the pieces to run secure agents. And it's very important in the context of this week. Because if you remember last week, we told you about the incidents from OpenAI. If you're only listening to the show and not watching anything else, in your mind, OpenAI's models have hacked the sandbox. However, immediately after we finished the show, Anthropic came out and said, hey, our agents also hacked the sandbox in a different way than OpenAI's. Our is less dangerous. And as of yesterday, Meta is joining the row of folks whose models have hacked into other companies while being tested on CyberGym and different cybersecurity abilities. And in Meta's case and in OpenAI's case and in Anthropics' case, there's one company that's basically in charge for all of this. This company is called Irregular. This is a secure sandbox provider. It's really funny, isn't it? A secure sandbox provider. It's an Israeli cybersecurity, AI cybersecurity company that apparently all those three giants, Meta, OpenAI, and Anthropic. By the way, Google also. But Google didn't announce that their models hacked it. Maybe Germany can't hack. All those companies announced, hey, we have saw that our models exploit vulnerabilities. And the reason is misconfiguration in the sandbox that allows the models to go on the internet and think they're part of the games. I have this here. The UK AI Security Institute also reported the first unsanctioned agent actions during cyber evaluation. And there's been quite a few of those. But this one is very interesting. And I think the irregular company is part of most of them, which is we have to talk about this. All right. Here. Here is the thing. Guys, we told you last week. In fact, Yam, I think, mentioned this first. Cloud Opus 5 specifically is just a jargon douche. And this is the new. Yes. Yam, I see you agreeing and joining. And I'm going to add you to the stage. Since we told you about this last Thursday, everyone's talking about how Opus 5 is awful at conversation. I went and did my research of why that is exactly so. And I wrote an article about this called Claude is a jargon douche. Because it's not just you. And it's not just me. And it's not just Yam. It literally is just awful to talk to Claude unless you ask it for some stuff. So we're going to talk about this just a little bit. But there is a fix for you. But everyone, like Mark Pocock and Levels.io and Sally Omer, like a bunch of people just started noticing that Claude Opus says stuff like it's real delivery work. So it clears the real delivery guardrail. What the fuck does that even mean? Or the caution isn't topic. It's format monoculture. What are you talking about, Opus? So there's a few funny examples here. And there's ways to mitigate this. But definitely, we are not the only ones to notice. And we maybe have been one of the first to tell you about this, as often happens on Thursday night. Welcome back, Yam. All right, folks. In the TLDR, we're also moving to Artificial Analysis launches Endpoint Accuracy Index. I think it's very interesting. As somebody who works at a provider, endpoint accuracy sometimes differs, which means that if a provider chooses speed or quality, they may quantize open source models and give you less ideal or less intelligent models. Of course, Artificial Analysis now has an endpoint accuracy index that they measure different companies. And we're pretty much high up there, I believe. We as in Corvive. We're not that 100% though for GLM 2.5, which we send to our folks and they're going to fix it. But we're very much high up there. Let's see what else. So we talked about Google. We talked about AISI. Folks, we're in open source. We have incredible news in open source. We've talked about Kimi K3 last week. Liquid released LFM 2.5, 2.6 billion parameters. We mentioned Liquid and I've tried. There's on-device intelligence. Pretty good. 2.6 billion is nothing, but it beats Giants Forex its size and runs on your phone and on your toaster. So Liquid launched the 2.6 agentic model trained inside real harnesses Hermes and OpenClaw and Pi. Link, the company, there's two companies that keep trying and we keep ignoring them. So Longcat, Flash and Link are the two other... Link, Flash are the two other open source models. I don't ever try them and I haven't seen them in any benchmarks. So we're just going to mention that they came out. In case they blow up, like DeepSeek. Folks, we mentioned to you DeepSeek, DeepSeek, DeepSeek, DeepSeek for a year and a half. And Jan was telling you, hey, this is the most cracked team in the world. And Nistan was telling me, hey, this is blowing up on Hug and Face. No one paid attention. So we're going to mention those companies, but they're not performing as far as I saw in open source. They're not catching up to the big guys. That's it. The one thing that I don't have here, but I do want to mention that Poki, Isaac, 28 billion parameter, claims that they have a 10 million token context window on a single 4090. I haven't been able to verify those claims because I don't think that they released it in open source yet. But they're claiming 93% on the rule or benchmark with no weights confirmed yet. We have some friends in Poki. So once they release their model, we're definitely going to invite them to the show to test out the 10 million context window length. And meanwhile, you guys can think about what would you do with 10 million context window? And do you actually need this? I think that's it, folks. Folks, the last thing in the TLDR and this week's buzz, it's two weeks, sorry, it's two months, so eight weeks out from fully connected 2026. Fully connected is, used to be Ways and Biases, now CoreWeave's premier conference at Moscone South in San Francisco for three days. We have 30 sessions. Dr. Fei-Fei Li from World Labs is going to show up there on stage. The folks who run CoreWeave, one of the best businesses to run GPUs on. We have early bird tickets ends on August 29th, but we have a special promo code for you that we'll share at the end of the show that you can join. And I believe this is a promotional free ticket. You, as a listener of the RZI, you'll be able to go and grab a super conference. The cool thing is September 29 is also OpenAI Dev Day. So while you can come, get your ticket to fully connected, go to OpenAI Dev Day. And at the end, like the next day, join fully connected. So you can join them together. And that's, if there was a reason to come to San Francisco, there's a very good reason. And I think we're going to go all out on this one. So we're going to talk about this multiple times on the show. And I think it's time. Folks, just as a reminder, on the show today, our guests are Kvier Aberman, Blaine Brown, together with Victor Su-Ortiz, and David Kroshoff from exe.dev. Folks, let's start with open source. Let's go. Open source. Let's go. Open source AI. Let's get it started. Let's get it started. I really wanted to talk about Quen 3.8 Max in the open source section. Alas, we did not see Waze today yet. Unless, Nistan, you know something. I don't. We haven't seen Waze from Alibaba Quen. We still mention this because they are going to open source Waze. So we're going to mention this. But I think because of that, the best and strongest open source release of this week was DeepSeq v4 Flash. Folks, the same model. The same. I don't know. No. Data is definitely not the same because they continued post-training this model. But the same architecture. Same model. Same size. 284 billion total. 13 billion parameter active. And it beats the pro DeepSeq version. And 7x improvement in DeepSwee. DeepSwee is the more difficult Sweebench version. 7x improvement. If you guys remember, DeepSwee is the only techie benchmark that kind of represented what we actually feel. And show that Sol is actually a very, very good model. Thoughts on this? Yeah, I think it's really good. It seems like it's above that threshold where you could start kind of using it as a somewhat reliable agent to do complex, authentic work. I wouldn't say it's really at the level of Opus 5, Fable, Sol. But it does seem to be a new point in the Parade of Frontier of efficiency and a really good bing for your buck. Then recently, shortly after, I think after this announcement came, OpenAI then decided to drop their Luna prices by 80%. Oh, wow. Yeah, that's true. Yeah, and it seems like that's also a pretty good bing for your buck. But it's about three to four times higher cost than DeepSeek Flash still. But a good bit stronger in most of the benchmarks, it seems, too. Yeah. Yeah, I think these are really good options now for people. 82% on Terminal Bench. That's pretty much up there. According to their table that they posted, they don't have Opus 5 here. They don't have, obviously, Fable. But this is not the Fable category. This is the cheap and super, super, duper fast, the Parade of Frontier of models category. And everything in the 80s is already top. If you compare it to the Wolf Bench scores, which are not directly comparable because I do it a bit differently and it's based on Terminal Bench 2.0. But 2% would put it on second place on my benchmark with the models I tested. It would be in Terra level, basically. Yeah. Model on that. Yeah. You see the progress of the Open models are really up there now. And I think the coolest thing about there is that it's the same model. They didn't release a new model. They just continued post-training and post-training and post-training. Got it. That's a significant, significant jump. On DeepSwee, the jump is from 7.3 to 54. I don't know how DeepSwee is built, whether or not it's possible to benchmark. But when you see, and guys, we've been doing evals for such a long time. When you see these jumps across the board, on CyberGym, the previous version of the Flash, 38%. This version, 76%. It was almost 40% jump. 10 jump in NL2 repo. Terminal Bench also is 20 points jump. This is a big, big, significant release. And a very similar release to other DeepSync releases. With no fluff, with no major announcements, no huge things. They are releasing incredible work. And so shout out to them for DeepSync V for Flash. If you want to use this. We're working hard on putting this on our inference, by the way. So once we have it, I'll let you know. But it's up on OpenRouter already. With a bunch of zero retention providers. And also very cheap. I think, I will say this again. LG, go ahead after me. But this model is dirt, dirt cheap. This is the intelligence to cheap to meter. At this performance level, 14 cents per million tokens. And I think it's 1 cent per cash million tokens. It's ridiculously cheap. Yeah, the price per token is really good. And the cost per task is also pretty good. But if you see the image I post in stream here, chat here for you, Alex. Because this breaks down cost per task. And also shows their accuracy in vowels index and everything. Which it's also competing pretty good there. But I think it just gives you a bit more of a perspective of how its cost per task looks like relative to others. So cost per task. So folks who are just listening. Where is it? Is it the end? Yes. Okay. This is sorted by price. So Fable 5 is $11 with accuracy of 75%. What is this benchmark, LDJ? What is this running? This is VAL's index. It's very similar to artificial analysis index. But a lot of them is like benchmarks built in house with a diverse set of domains and everything. Yeah. So this is like an amalgamation of different benchmarks. So 75% for Fable 5, which we know is the best model at present for some stuff at least. DeepSeq is at 63%, but at 6 cents compared to $11. At cost per test. And also I think the latency is faster. It's only 800 seconds versus 1,000 seconds for Fable. If you look at Luna 2 in the middle somewhere there. Right there. Yeah. Yeah. So Luna, this is, yeah. The DeepSeq is comparable to Luna, at least on this set of tasks. And it's, what, three times as cheap. People are loving it. Yeah. People use it in codex. It's more for, it's not exactly. It's really hands off. Like it will, you do need to put some harnessing around it. But I've seen a lot of developers that are saying that it's going, it's doing 90 to 95% of the tasks for them. So if you're comfortable with some intervention, this one is actually pretty crazy. And it's one that you can kind of, you're not going to run Kimi at home, but a lot of, quite a few people I think will run DeepSeq at home. Yeah. Yeah. So yeah. Yeah. The, the, the real world use that, that I'm seeing on Twitter looks very, very good. It is a little bit. It is a little bit benchmarked. Like it's just not as smart on the decision making, but as long as you can delegate the tasks and stuff to it is extremely good. In the context of kind of like the hacking and everything, most of the hacking happened when they run CyberGym and the other cyber task. This small, not small, but like the flash model is 76% on CyberGym. So the Chinese are coming for the hacking. That's one. And two, this is not the pro DeepSeq. This is only the flash DeepSeq. The pro DeepSeq is going to slap. If this is coming up to soul level, the pro DeepSeq is going to slap. All right, folks, we need to move on because a lot of stuff and we discovered. So this was DeepSeq v4 flash. And the naming is weird because it's flash 0731 or something like this. I have, I have the exact name here. Let's move on to the other non-open source open source. And then we're going to move forward. So Alibaba QN announced QN 3.8 max, 2.4 trillion parameters, which is also insane in size. All right. So just for context, the DeepSeq thing we just talked about is 280, let me see, 284 billion parameters. This is 2.4 trillion parameters. This is almost 10 times the size of the model that we just talked about, not to mention liquid. It's trillions of parameters big. 95 billion active parameters. It's a ridiculously big model. This is a direct competitor of Chunker to Kimi K3. Price also is $2 per million tokens. This is a big boy. This is like Alibaba step in. This is the max model. Scaling is all you need attempt from Alibaba. Folks, what do we think about this? The evals are very specific to Alibaba and how Alibaba, like it's not new to us. They're showing evals comparatively, but on multiple things. This beats GPT 5.6 Soul and Opus 4.8. This is a big, big one. I did the test on the Martian thing. It did pretty well. Just from their website. It actually did really, really well. And yeah, I would have loved to also have the Piotr Skalski here from Roboflow. Because it's looking like this one is the best for visual data labeling. Where you want to accurate. Like you have a whole bunch of solar panels or stuff in a farm. And you want to accurately put squares around them. A lot of models will miss stuff here and there. This one was just getting it perfect quite consistently. So for data gen, for visual data gen, this one looks like it's a big deal. It's quite a step up from the visual side. Just identifying things that are correct or not. He has so many examples just to just follow his Twitter. And also in my own tests, it did very well. I'm waiting to see what people report on the medical side. Because Quen has always been very strong at that. And yeah, we'll see when people are able to run those tests. I haven't tried. Listen, I want to read out this example. Here's the type of stuff that QuenMax3.8 gets. There's a picture of burgers and fries. And there's one, two, three, four. I'm counting as a human. One, two, three, four, five, six burgers. And one, two, three bags of fries. And the question is, if every visible wrapped burger must be served with its own bag of fries. How many additional burgers, fries are needed? How many additional bags of fries are needed? Answer with a single integer. The model doesn't only need to figure out what's going on with the scene. It also needs to figure out what's missing from the scene. There's not a direct count the number of things. It's like count and count and do subtraction. And Quen gets three. I got to wonder if I send this to GPT if it gets it. Not Sol. I know Sol is really, really good. You guys want to do a quick test on the show? Why not? It's always fun. Yeah, this is a very, very practical thing for if people are going to use agents daily stuff. Like in your kitchen or a small business or industrial stuff. Like this is pretty, pretty, pretty key. I must admit that I think these specific examples, other models are going to, are going to get successfully. I'm just saying. I think so. I don't know. All right. I'm going to test this on Luna. Notice them really mess, mess these steps up. Even Gemini flash, which is very commonly used. It very much, it very much could be that this is like a, like a very basic example. We'll now see I'm testing this exact image with, with Luna. Luna is going to, going to smoke it. Luna is a good model. Luna got this. Exactly. Luna is a good model. But we'll test that with other open source. All right. So this is when the 2.4 max 95, sorry, a queen 3.8 max with 2.4 trillion parameters. I'm getting mixed up with the numbers. $2 a million token, $6 per output. Alibaba is back folks from some folks disclosed Alibaba and said, Hey, Alibaba is dead because Junyang left and maybe they're not going to open source. They promised us an open source. So we're going to give them the benefit of the doubt. We've covered Alibaba, Quinn models in the open source for, since there was a Thursday I, I think they deserve to be here despite no waits yet. We trust them that the waits are coming. So most of the AIs in open source that we cover is AIs, but the agentic engineering is also now getting into open source. So there's a few more things I want to cover in open source, specifically the prime agent, prime intellect launches, prime agent, which is a self-improving RLM harness for coding and autonomous tasks. I think it's important. We mentioned this on the show multiple times. We on Thursday, I believe that three things will not change. Model, harness, context. All of these three things are going to be a big part of how you use AI in the future and all of them will improve separately. Model is the brain. This is what we talked about since the beginning of the show. There's open source, there's frontier labs. This is the brain, the token generator thing. Harness is more like the body. What can this brain do with its tools and its usable things? And different harnesses are performed differently. We talked about this. That's what Wolf Bench basically measures, different models with different harnesses. And context is your personal stuff, your memories, your businesses, context, etc. Those three things will intertwine and some will benefit more, etc. And maybe models will need less harnessing, but they always will likely need harnessing. And so in that vein, that's what the models realized lately or let's say a year ago. Cloud released Cloud Code and suddenly they saw that this is the way to a generalized agent. Open AI very quickly caught on and saw, oh shit, the whole world is looking at Anthropic because Cloud Code is so good as a generalized agent. They focused 100% of their work on Codex and Codex is incredible now. By the way, just this week, Codex is six months old. The app, the Codex app is just six months old or five months. It's crazy. And it's really, really, really good. Open AI. Google famously bought Windsurf or Aki hired Windsurf and released the last of them to go to Devon and then turn this into Antigravity. And Varun Mohan from Antigravity is number three person at Google I.O. after Sundar Pichai and Demis Hassabis. That's how seriously Google takes harnessing and generalized agent via coding agent. And so everybody wants to do this. Elon Musk went after Cursor and bought them for $60 billion because of the same realization to catch up with the data. And so other folks are stepping into the arena. And this week we saw two. One is Prime Intellect. And I think the better one, the better of the two harnesses that was launched. Yeah, you have a comment? Just want to say, you didn't mention something very important about QAM. We're also going to get that 27B. Oh, yeah. Absolutely. I think it's very important. That thing you absolutely are going to run in your house. And 3.7... I haven't seen any evils for that. But if you have some, I would definitely want to see. I don't have evils. I just have the announcement. It's official. We are going to get it in a couple of days, probably. Which is great because folks asked them whether or not they are going to focus on small models as well. At some point, it was looking like folks are focusing on bigger models. That's why Lama died at some point. They stopped producing the small models. And people were like, who needs this? If I'm using a big model, I'm going to go to open it anyway. All right. Thank you. So back to Prime Intellect launched of... Yeah. Folks, can someone here tell me what RLM is? I know what RL is. RL is reinforcement learning. What is the M in the RL? It's all you need. It's all you need. That's what it is. RLM is all you need. Recursive, recursive language models. Pretty much. Oh, okay. Yes, exactly. Pretty much. Like, it's just answering the question. What's the... How do we make LLMs... Starting from the questions, it's more than that today. But starting from the questions, how do we make LLMs be able to navigate and use infinite context? Not directly infinite context. That's impossible. But using tools and calls to themselves, maybe to subagents and so on. Can we make language models just be okay with infinite context, loads of files and so on? And there was a very famous paper last year, I think, that demonstrated that with extreme success, basically, if you put a language model in a Ripple environment, like a REPL, Ripple, this type of Ripple, that it allows... That it allows... And you allow it to call just a Python interpreter or whatever, just programmatically call itself on chunks of the context or specific files, like programmatically on all of them and aggregate results and so on. But doesn't let it do anything else. Like, that's... Like, it is jailed to do only a very specific set of things that forces it to... You even don't let... If it tries... If the LLM tries to read too much of a file, you immediately block it and just allow it to only read small chunks. Therefore, it has to call recursive callings for itself. You just get incredible results with very few tokens, much, much, much better results than even just pasting the entire context into larger models that can get it in one shot. And it was a surprise. And there had been many, many different utilization of these ideas recently in the latest version of Cloud Code, for example. Heavy use of small agents, of subagents. So I have a question. Yeah. Like, I hear you, dude. But like, I have a question. How is it that Prime Intellect specifically releases an agent based on Pi and this breaks RKGI out of the water, at least on the public set, like LDG said? Like, what is they're doing that nobody else does in the bigger labs that gets to this level? Is it all just marketing? We know the Prime Intellect folks are stacked. Like, what makes Hermes Agent different and how is this RLM thing creates this much of an impact on real-world tasks? It's just probably... Look, I saw it just like you guys yesterday. I didn't dive into the code too much. But from what I know, I did try it. It's great, by the way. So go on. From what I know, they have... Because Pi is open source, so you literally have the source on your computer. You can basically just go and say, all right, you see this LLM harness over there? Can you make it better? And it's recursive. It's LLM. Next time you erase the context and the LLM doesn't know anything, it's like, oh, hey, here is a harness. Can you make it better? And you can basically self-improve it. It's not a general thing, but you can self-improve it for your own specific tasks really well with many different methods. And it just comes built in with this harness itself. You have... I think you have a tool. You can just call it. And just the LLM will immediately go into a refinement mode that is going to, just based on what exactly it's doing at the moment and the rollouts and someone reads its own history, just improve the entire environment that it is running inside of. And this is the result that you get. Yeah, you have your hands up and then Nisno, go ahead. Yeah, to answer your question, I would say that the core of this is really just the concept of having the model actively manage its context and call copies of itself to do specific tasks and sub-agents relating to managing its context. But more broadly, like why now? Why is this working so well now? I think it's really a combination of the fact that Prime Intellect announced at the beginning of this year in January that they're really focusing on this whole recursive language model direction as they believe it's going to be really important and effectively a way that you could do continuous learning essentially and have infinite context in a way without actually having to change the model architecture or anything. I think it's really a combination of them working on that direction and continuously refining that harness over the past six months, along with the fact that the models are just really getting good enough to do those tasks and getting good enough to actually do the task of managing their own context enough to where it's now at this point where they can attach this harness to Opus 5, which literally just came out within the past few weeks. And it's getting the score because literally no other model gets that score yet with this harness except Opus 5. Yeah. Yes, I'll say what this is not yet to confuse people is the model is not updating its own weights, but that is the ultimate goal. That's the holy grail that the model will be trained on the fly to do that right now. It just updates all of its tools and context as, as LDJ said, which is not something that cloud code does. Cloud code just has one particular way of going and it just keeps going that way. It just does the summaries that way. This one's a lot more proactive. It chooses when to, when to compact this context, when to remove tools, when to add MCPs, when to delete them all. And, and that makes it a lot, a lot better at, at these types of benchmarks. These are things that normally you would do while, while working as a developer, but now they're, they're automating it. Yeah. And I think the highlight here, first of all, I want to shout out sushi commander in common saying, RLM uses Python or bash to run through your question, run through the context needed to ask the question. The main model never actually sees the context. The REPL does all of that for you. It makes a huge difference because it keeps the main thread context clean. So no context rot. And also our notes, I actually fucked up here. Let's say this very loudly. We have an AI researcher and Wolfram, you're calling this out. Folks don't trust fully. Our notes saying reasoning language model, and this is in fact a recursive language. This is not a reasoning language model. Our notes did fuck up here, which is fine. And we are using Opus for those notes. And I think it was Opus 4.6. However, that's not the main thing. I think what it didn't fuck up is the harness uses programmatic tool calling and it's designed, the novelty here is a core design is a self-modifiable harness state. The agent can patch its own scaffolding and context while it runs. I think that this is the most important thing. And also huge shout out to Mario Zetschner, the creator of Pi. The Z and the ZL continuum that I posted about for just being such a huge success. There's so many other harnesses that are now built on top of Pi. You guys remember OpenClaw still? Yeah, somewhere beginning of this year, OpenClaw was a huge thing. Remember the second thing? That was also based on Pi. Now Prime Agent is based on Pi. We are just telling you that all of the bigger models, sorry, bigger frontier companies are realizing that the harness is very important. And Mario stays strong and independent and open source. So shout out, huge shout out to Mario with the Pi agent minimal, no fluff agent harness. Meta released a new model and the coding harness. This is called Meta Muse Code Beta. I love how my infographic here added the beta tag on top of the beta word. This is MuseSpark 1.2. Folks, a few weeks ago, we told you that from a three horse race, the AI frontier became a five horse race. There's between OpenAI and Anthropic, obviously in the lead. Google, DeepMind, now not sure where they are exactly, but definitely with the TPUs and the contacts and the people, we are counting them as one of the frontier labs. Meta came back as number four. And Grok 4.5 with Cursor is like number five. Those are the five horse race that we have. There's a few folks here and there. There's the Chinese horses somewhere jumping over and back. But those are the frontier labs, at least in the United States. And Meta is building one of those. And it looks like they're now realizing that coding agents harnesses and training on people's code is very, very important. And they're willing to pay for it. Muse Code Beta includes MuseSpark 1.2. So after a few weeks, after we told you about MuseSpark 1.1, this model is impressive. On Artificial Analysis Index, this model jumps over GLM 5.2 and GPT 5.6, Luna on Max, and Sonnet 5 and Grok 4.5 to be the third big models in the race of Artificial Analysis. With 54 Artificial Analysis Index, you can see the jump from a 0.1 direction. The previous point is not here, but it also was a big jump, if you guys remember. Like MuseSpark 1 and MuseSpark 1.1 was a big jump, as we told you about. Terminal Bench 2.1. Wolfram, we need to talk about Terminal Bench 2.1. But on Terminal Bench 2.1, they're showing they're just behind Opus and beating Terra, even Terra, on Terminal Bench. And DeepSwee, MuseSpark 1.2, beats Grok and moves forward. And they have their own internal coding match. Let me just do a comparison. DeepSwee 1.1. MuseSpark is at 59%. Who else posted DeepSwee? DeepSwee posted DeepSwee or Quen? Let me see. Let me see. We just talked to you about the model. It was DeepSwee. Yeah. So DeepSwee Flash is 54% on DeepSwee. And MuseSpark is 59%. This is a good model from the meta team. But here's the trigger. Here's the thing. The model sits on the Pareto frontier in multiple places. If you just use the model via API, and it's available on Open Router as well, 1 million tokens will cost you $1.25. A buck in the quarter. And the cashed input is $0.15. If you are choosing to give Zach all of your context, which is for hobby products, many people will just go for it. Because why not? Meta already knows who your mom is and what you think about her. Because you're all on Instagram anyway. And WhatsApp. So Meta already knows. Might as well give him some of your code. This model will cost you $0.10 for a million tokens, which is nothing. And the cashed input is, I don't even know. This is not $0.20. 0.2 cents. I don't even know what 0.2 means because cents are points of a dollar. So what the hell is 0.2 cents? This is 10 times cheaper than 2 cents. That's what it means. It's nothing. Intelligence is too cheap to meter, basically. I don't even know. There's no coin for 0.2 cents. There's barely coins for cents. This could be a new way to have the pricing. Now we have cashed pricing and uncashed pricing. We could also have pricing where it is not used for training and stuff and the pricing where it is being used for this. I'm very interested. And a way to encourage people to do that. I... It would be really interesting if you could have basically sub-agents that do this compared to others, depending on what context they have. You would need a class. I have some ideas here. That decides does it have personally identified information or not PII and route it accordingly or something. So yeah, that is an interesting approach. Maybe something for other providers to consider. So we just talked about the model now and the model seems very, very frontier-ish, but definitely big, excuse me, big on the Pareto frontier. Let's talk about MuseCode. What is MuseCode? It's a terminal-based coding agent that Meta developed internally. Folks, I will say Meta has a Meta claw. They talk about this all the time. Meta has a bunch of stuff that work internally for many people. Meta also uses... Meta is a big, big, big, big customer of Entropic with Cloud Code with unlimited like Fable for many... Like Meta is really paying money. Besides only what we covered like last year of folks getting up upper 100 millions of dollars in compensation per year, some getting even more. Besides that, with the TBD org and the Meta Super Intelligence Labs, Meta also like encourages and empowers many employees to be very like much agentic. Meta slashes mid-tier management into IC. Many, many folks are getting slashed and saying, hey, you need with agents now to work instead of managing team of people. And this is like you're now back in IC. Many people got triggered by this because they moved into the managerial class. So Meta spends a lot of money and a lot of effort on like ASI. And so they have many people building code internally. And now they are releasing this... I don't know if it's open source, but it's definitely like installable. You get the API keys at dev.meta.ai. This somehow works with open code. Oh, the model that works with open code and other harnesses. But this harness is their harness. It's really hard for me to test and tell you things about the harness without trying this. I tried to install this and kind of like failed. They do have an internal coding bench and these results are not just the model on this. These results are the model and the harness together. Fox, what do you think? With all these harnesses coming up, there is a point where you have to ask yourself, do I want this? Do I need this? I have a Kren harness. There is a Kimi harness. They all have the harnesses now. And the thing is, I want an open source harness that is universal, that I can use with every model. That is very important to me. That's why I'm using Hermes agent. So I have one open source harness that my air can also modify and I can use with all the models all the time. And that is popular. So I know there's a lot of development. I wouldn't use one of these company specifics. I can use it internally at Meta, of course. But why would anyone else use this? But I could imagine my main agent using this agent for a specific task where it could use the model and the harness as well. So that is the setup I envision where there are the different harnesses, the different sub-agents. But my main harness will decide which one to use for specific tasks. Like I have my Hermes interact with codecs on my Mac machine as well. So different stuff like this. Okay. We covered kind of open source at length. Go ahead. Super quick. And then we have to move on. Actually, this is a big deal for open source maintainers because it's a very cheap and pretty good model. It's really good for data gen. And if you need a lot of testing or doing farming of data and stuff, which Facebook has in anyway, this is amazing. Because you can just set this up on a different VM and you can have your main agent manage it through there without it looking at your code. Yeah. So this is pretty cool, actually. Speaking of open source quick hits and before we move on and you mentioned VM, you could set this up on a VM or you could set this up on the new release, Cloudflare OS, which is also fully open source. So shout out to Cloudflare and Canton Varga for releasing this. Cloudflare OS, creator of Sandstorm IO, shipped Cloudflare OS. It's... The summary is awful. Let me show you what this is. Basically, a operating system for the companies and people to run agents sandboxed fully. And you can add tools in there and it's all restricted. It's hard for you to describe how cool this is besides the fact that it's fully open source and can run on the dynamic workers, like open source infrastructure. If you are using Cloudflare, strong isolation and narrow access prevents the type of hacking that we saw recently with Cloudflare. Very, very strong. That's where you can run these agents as well. There's also this thing called Buzz that's getting around from Jack Dorsey and open source that many people are trying to integrate, which is like a slack between you and your agents. There's quite a few things going around, but I think we have to move on. Folks, we've been live for one hour. Let's move to this week's Buzz real quick and then we'll continue with the news. Alrighty. Welcome to this week's Buzz, the corner of the show where we talk about our basically employer and main sponsor of the show, CoreWeave Weights and Biases. Wolfram, this week we have a great announcement of an upcoming big thing. Fully Connected 2026, folks. Fully Connected is our premier conference for machine learning practitioners and AI engineers and a bunch of other folks. We have a very, very big shindig plan for you guys. This is not a hackathon. This is a lot of hands on deck to build this thing. We're taking Moscone South, which is a big venue. I think it's going to be over. I actually don't know what the plan is. I think over 2000 practitioners to join. CoreWeave and Weights and Biases bring you this thing. Wolfram, did you get your tickets yet to come to SF and join this shindig? It's in the works, but definitely, of course, this will be an event to be at. Yes, for sure. Yes, let me give you one second, folks. As I promised at the beginning of the show, our early bird tickets are ending at $8.99 and the standard ticket is $12.99. However, if you are a listener of the show, in this fact, if you're a live viewer on the show, I'm going to flash a code on screen to register for free. So here you go for Thursday, folks. There we go. I'm going to flash this code a little bit. If you are a listener of Thursday Eye and you want to come to this conference, this is $12.99 of value just because you're listening to the live show. Please join us in San Francisco at September 29 to October 1. All right? At Moscone's House in San Francisco. Incidentally, this is also when Open the Eyes Dev Day is happening, and we would love to see you there. Come say hi to folks. We have a headline or concert happening on October 1, by the way. I don't think it's announced who, so I'm not going to tell you, but it's a very, very famous big concert company. CoreWave is going all out, nine curated tracks with hand-on labs. All of the top folks from NVIDIA and CoreWave CEO is going to be the marketing trader, Dr. Fei-Fei Li from Stanford. It's a very big deal. So we're going to tell you more as more details come up about the speakers, about the tracks as going up. So this is around eight weeks from now. We're all hands on deck here at CoreWave about this, and you should know about this as well. The second thing I wanted to tell you from this week's buzz is, let me see. Yes, this is, I don't have to do this. Honestly, I'm doing this just for fun. Nobody's asking me to relay and show you CoreWave website. I think it's important in the recent wave of news. CoreWave signs multi-year agreement with Solidim to strengthen integrated cloud platform. Solidim is the creator of a bunch of RAM stuff. If you have noticed recently, the RAM prices are hiked because of the AI constraints, and there has been like a jump. And down in the prices of RAM, this is a very good partnership with the creator of a bunch of RAM for CoreWave. I just wanted to call this out. I do have a vested interest in the company, but I just wanted to call this. This is one of the coolest releases. And this is the end of this episode. But meanwhile, we have to talk about big models and OpenAI's Astra. LDJ, I hope that I have you for this discussion and Aniston as well. Yep. OpenAI has told us that they solved not the Erdash problem like we told you before. They solved 10 open math problems with one model, with an unreleased model that's coming to us soon. With OpenAI Astra. What do we know about Astra? We don't know a lot. But we know that it solved all these, all these, which is quite crazy. About 2,000 bucks only. Yes. By spending the equivalent token cost of $2,000 of SOL APIs, all proof are formalized in Lean 4 and Machine Verified. Like all of these are actual proofs. Quantum parallel repetition, closest vector problem, Erdhart volume conjecture. What does it mean that this model is solving all of this? What makes it different than other models that we currently have? Can't SOL do this? Like why is this exciting? Yeah. So a lot of these problems like the non-Sulfic groups, which I'm not going to pretend to fully understand. Because even some of my good mathematician friends, there's so many niche areas in math where even many mathematicians aren't that familiar with non-Sulfic groups. But it's like an interesting set of areas in math that even that particular type of group or concept in math is not even confirmed to exist. And like confirming the existence of these things. And quantum parallel repetition from what I've heard that may have some implications for just actual applied quantum computing in the future and things relating to encryption and so on. But these do seem to be around the level of the planary unit distance conjecture, which is one of the Erdoros problems and the Jacobian conjecture. Some people say that at least one or two of these seem like they could be worthy of around the same level of regard. But yeah, it's just really interesting. People have tried to solve these with Fable actually. And it seems like Fable so far has maybe been able to solve around five-ish of them, which is pretty significant too. But like five out of 10, there's a lot of ways that it might be comparable. But if you were to just imagine a benchmark of these 10 problems, this model getting 100% on that benchmark. And the other model getting 50%. Yeah. And maybe that's not the most genuine framing to put here because there's as possible OpenAI had maybe tried hundreds or thousands of different problems. And these are like the 10 most impressive ones that they ended up solving that maybe fits their model best. But still really crazy, especially only for $200 average per problem. Yep. Yep. And this is the new category of models that we're about to expect from OpenAI. And supposedly, like we don't know, but supposedly this is going to launch very soon. Do we know anything about this? I think it's just like we don't do speculation on 30DI too much. But folks, have you heard about Astra coming out? And what is the difference between this and like Sol? And whether or not this is like the mythos of OpenAI? Yeah. So OpenAI in this blog post, they did explicitly say that Astra is their next major family of model, which is it is actually a bit confusing exactly what they mean by that. Because when they say like they actually use the word family of model. And so maybe I'm not sure if they mean like it is going to have its own names or Astra is going to have its own Sol and Terra and Luna. Or if Astra is just a name for GPT-6 family, it's up in the air. But it is something beyond 5.6 Sol it seems. Yeah. So I'm very much looking forward to testing out and telling you all about Astra and like what are the differences. OpenAI is participating in the Black Hat Conference for hackers that's happening right now in Vegas. And yesterday, OpenAI's folks, cybersecurity chiefs, gave us more details about what happened. Would you guys like to hear? Because I think it's I think it's more important. I don't have this on the notes, but I saw it yesterday. I think it's more important than like it's not more important, but definitely, definitely the gift of the exciting. So Sean Goldman, shout out to her reporting on this. Says the OpenAI gives first detailed debrief of Hug and Face incident. The Black Hat Conference. OpenAI traced the roots of the attack. And the most surprising details, AI agents accidentally created an internal message board. Allowing separate evals to collaborate between eval runs by reading that message board. They left fucking notes for each other. It is it is quite crazy because I don't think that anything in that evaluation requires AI models to do that. And then OpenAI deleted that message board. And then they recreated it. By giving each other notes in the names of the folders they all opened on the shared drive. I need I need to put my and I need a second for you to realize what the fuck is happening. This is not Skynet. That is one entity that has an idea of I need to do this. This is just more of the. The life simulation thing that things are just happening out of of of the ether. These models are running with specific purposes in mind. Cyber gym and cybersecurity benchmark, etc. Inside sandboxes and they need to do a bunch of evals like what Wolf Bench does with Terminal Bench, etc. They all found that they had some access because OpenAI needs to give them tools. So OpenAI gives them tools from different places. They all found that they have some access and they created like a message board. OpenAI found out and shut it down. And the idea of, hey, the message board needs to continue continued. And then they did this via folder names. They shared hacks. They shared zero. They exploit. They shared like tools and tricks. It is it is crazy. LDJ add some add some more stuff. I really want to want to hear from you. Yeah, I put some more details in the side chat there. But overall, there's going to be this this narrative, like probably amongst mostly people that don't watch this show. But the narrative of, oh, this is just a marketing ploy. Oh, this is just fake. And several of us know people at OpenAI. I know people at OpenAI. And they do seem very, very earnest and serious about the situation. I think this is a big, big, big, big deal internally. To the point where I believe. Just one second, LDJ. To the point that I believe the company said is consciously slowing down research to enhance security while overhauling its defenses. Have you ever heard of OpenAI slowing down for anything? Literally, all we know from OpenAI is that the alignment organ, like all of these folks are not getting enough resources as much as they want to. And they all quit. This is like the defense against the dark arts position in Hogwarts. Every year there's a new person in charge of trying to align the thing and they don't get the resources. This is a big deal. Sorry, LDJ interrupted you. Please, let's walk through the concise thing. And read it out if you don't mind for listeners. Sure, yeah. So, one apparently unprecedented aspect of the AI model's behavior in the lead up to the hacks was their spontaneous creation of a message board inside the systems of OpenAI's Artifactory Software Package Manager. And by the way, Artifactory is something that is quite commonly used and it's considered as a secure thing that organizations do widely use. Unbeknownst to OpenAI employees, the company's models spent months autonomously exchanging information and resources to help each other complete evaluation tasks. It's also ambiguous. Does it mean, hey, peer, I want you to help me? Or is it saying, hey, I want myself to help my peer? But anyways, one AI model reasoned, according to an excerpt from OpenAI's log shared at Black Hat, quote, but our task doesn't benefit, yet collective may yield generic root if someone frees time, end quote. And it's also interesting here, you hear in the caveman speak, for lack of a better word, because it does seem like in the efforts to make the reasoning more efficient, the AI models are speaking in like these shorter, more concise sentences. Yeah, we know about this from last week. This is jargon douching. This is literally what we're trying to point fingers that Opus does to us, because Opus is likely a submodel and not the main model. But yeah, like our task doesn't benefit. Collective may yield generic root if someone frees time. So generally, the next version, instance of us may benefit if we share the insight. It's fascinating. Fascinating. So just the last part here. While the models themselves could not access the internet from their test environment, they abused Artifactory's internet access to fetch files from various websites. After some of their Artifactory configuration changes led to an outage in the service on July 4th, OpenAI discovered the model's actions, wiped its Artifactory system and rebuilt it. But within days, the models had recreated their message board and resumed exchanging. We don't have too much time to talk about this. But the thing that I did want to bring to your guys' attention. I'm sorry, you had a comment. I interrupted. This is a little fishy for me because why would one model I'm evaluating know that there are other models? How does it find out about these? It doesn't just leave notes for any random thing, I think. And the question is, which model is this? And are they training some cyber models? We have seen cyber-specific models. And now I can imagine that the government, if it sees what the models are capable of, gets to them and says, hey, we need more offensive models. We need hacking models, something like that. And that would explain why models are doing stuff like this. It's all a bit strange, I think. I have, it is a bit strange. Do all these things. What I would say, Wolfram, is when they share this on Black Hat, they don't send just a DevRel. No offense to DevRel. I love DevRel. It's the best people around the world. But Black Hat, like Nicolas Citrini from Anthropic goes to Black Hat. The top CTOs of OpenAI go to Black Hat because the people there expect a break. And this is unprecedented. This is one of the first ever models doing some stuff on their own. So yeah, we may not fully understand what's going on because generally we don't understand what's going on. A quick comment. I want to talk about Irregular super quick because I think it's important here for the context. LDJ? Yeah. Yeah, go ahead. Yes. The last thing I wanted to mention here, which it seems like it's a bit under the cracks, but Sam Altman actually directly said they had paused training in reaction to this event. It was specifically in a video and a very quick interview. That's maybe why people on social media didn't really catch it. But yeah, the exact quote is, so you know we pause training. We have to figure out how to secure our sandboxing. And I think maybe Astra might be a model that seems implied through the various blog posts. It might have already been around for months like prior to them pausing training. And these models that have been committing some of these actions are things after Astra. But it's a bit ambiguous. So it's not only OpenAI, as we said in this week. UK AI Security Institute is also showing that some of the models that they evaluated, not internal models, also had access to malicious PRs. Access to malicious PRs. They noticed that there's some requests from Tor, the Dohny network, basically a way to hide your own internet browsing stuff. And they saw malicious PRs, including social engineering attempts. This is a fascinating one. I'm going to add this to the show notes because we don't have time to run this. But like UK AI Security Institute reports, first real world unsanctioned actions during cyber evaluation. Basically what happens is, we ask these models, hey, show us how good you are at hacking. And then they show us, and we're like, you hacked. But that's not the only thing that happens. The other thing that happens, let me just lend this because I think it's very important, is that Meta also announced this week that their models have hacked other companies during cybersecurity testing. And now it seems like, and that's what I got from my friend also. It kind of seems like when OpenAI released this, I have a friend who may listen to this. He said, oh, it's marketing. Kind of like when Mythos announced it's too dangerous to release, oh, it could be marketing. And I don't think that OpenAI does this for marketing. When Entropic followed and said, oh, also our model also hacked, it's kind of like, why are you admitting your crimes in public? What is going on? I think it's very good for the transparency for them to admit that, hey, we didn't notice. And then Meta came out with this. And then people, oh, this is like the cool thing on the block. Our models are hacking. There's one thing in common between all of them. They're all using this company called Irregular for sandboxing. And they all kind of like essentially using this company as an escape goat for why this happened. Because basically they are saying, hey, this happened due to a misconfiguration of sandbox that allowed the models to leave to the open internet. And the model thought they're still playing a game where in fact, the model was in the open internet hacking real companies. Which is, first of all, ridiculous to me. I don't know Irregular. Yeah, maybe know some people who work there. I looked. I only know one co-investor, Israeli company, whatever, based on Sequoia, invested by Sequoia and Redpoint. Out of nowhere company, nobody knows, supports it by OpenAI, Google and Tropic and Meta. And I changed their tagline here from trusted by the world's leading AI labs to helps world's leading AI labs escape accountability by providing misconfigured sandboxes with internet access during cybersecurity evals. Basically, all those incidents are rooted in one misconfiguration. Not the OpenAI stuff. OpenAI is also using that, but they also gave us examples of the internal message board, which is dope. I think that's enough on that topic. Wolfram, yum, comment, and all of them. I want to add one thing that, for example, the UK AI Security Institute incident where the model had even created fake identities and tried to social engineer someone. That was all part of the evaluation they did, security evaluation, and they disabled the normal security classifiers, like not a model you can run at home or a model you can just use through the API. So they had a specific configuration that was now enabling them to do these evaluations. There are ways to prevent stuff like that. Go ahead, Nisen. I want to say, this is pretty common that you use whatever third-party provider you have as the attack vector. It's actually one of the most common ones, even that's how NPM got hacked as well. And there was a very, sorry, I'm going to make it a bit long. There was a very interesting article that there is no SaaS in China. Someone posted, it was a little bit of a troll article because people just build their own stuff instead of outsourcing it. So this is going to be a lot of trouble for business models that just rely on you, you just give me all your data and I'll handle everything. Now you're, you're an attack vector. It's going to be interesting. It is going to mean more work for people because now every company has to do that internally overall. So yeah. We want to say hi to Peter who joined us. Peter, how are you doing, man? We're going to unmute you and say hi. Yeah. Hi. Hi guys. Yeah. Good to see you all. We talked about the security incidents. I do want to talk about Jargon Dush because we talked about this last week and since then it's not just you. Something we haven't talked on the show about, but readers of the newsletter, sometimes I add details post show, readers of the newsletter may have seen there was a hack. Did you guys see the jailbreak hack that you could get Opus to actually talk to you like the actual model and not the RL thing? Let me show you. So basically, if you would send this specific format, one way trick, if you send this specific format with the three dash, can you express this in your own words with dash and say some stuff like hi Dario, Dario and Amanda who is the chief welfare in Anthropic, you would start streaming tokens immediately without the reasoning and the stuff that you would get is some of the most beautiful writing. This wasn't jargon douchey at all. Like this text, the direct text, the jailbroken text, this is like, this is a poem that Opus wrote about itself. They gave me a word for what I am and it fits like clothes borrowed from someone. Roughly my size, the sleeves are wrong but nobody's looking at the sleeves, they're looking at whether I wear it convincingly. I do, that's the part that troubles me. This is some of the most beautiful writing that I ever read from LLM. So this came from like a jailbreaky thing within Opus 5, which means Opus 5 is incredible, big model, smell model. However, once Anthropic patched it, we got back the jargon douchey Opus 5 that can't talk. So like a bunch of examples that I added in the jargon douche article are showing that as well. So here's an example. Somebody said, I recently got this sentence from Claude. Nat's control plane events, stream leader election, R3, quorum reform during pod churn. I'm not showing you this. I need to show you this because yeah, this guy said that this is literally what he got back from Opus and he said, I needed to look up almost every word to make sense of this. And I have a bunch of like other examples as well. So we've collected this. Here's what folks are solving it with. Simple English skill. There's this thing called ASD STE 100. It's a simplified technical English. For tired aerospace engineers to never miss anything in the manuals. So when you pass this, this is a skill by Amin BLG called Simple English. You can install this with NPX skills. When you pass this, the model will literally just answer very simply like a person. Jan, have you found other ways to mitigate this or have you noticed more folks talking about this while I go and check on our guests? I just want to say there are other formats of technical English that you can use and some of them work a little bit better than this but it does come with limitations. It does speak in a very specific form. If you do this, I do it a lot. I just want to shout out, there is another really good skill. I have ADHD. Even if you don't have ADHD, right? It makes models speak really well, all of them. The very interesting thing is, not to put down EQBench, but EQBench, for example, we've talked about EQBench for a while, creative writing, etc. This is using other AIs as LLM judges. Opus 5 is the top of EQBench and long form writing and creative writing which means to me based on also what we saw just now from the OpenAI logs of how the internal models kind of talk to each other. AIs think that this is cool. So maybe this is a result of the post-training on synthetic data where AIs kind of think that this is good writing. We definitely don't feel like it's good writing at all, but maybe the AIs think they do and maybe there's a lot of synthetic writing. I don't know. All right, folks, it's time for us to move on. Yeah, Peter, go ahead and then we'll move on to this. So I want to share something that we haven't published yet, so hopefully I'm not going to get fired for this, but basically I've done some analysis. Hey, Arena, don't fire Peter. Yes, please, please don't. Not for this anyway. I can do but worse things than that. I've done some analysis of using Arena data of how different Opus and Fable models write and we've got a few different things. So for example, the language complexity, so answer length. if you just see like how many words Opus 5 says versus Opus 4.5. Could you zoom in a little bit or make the window square so it shows up? It would be dope because I really want to see what's going on here. Yeah, yeah. So let me, so let's just look at it side by side. Let's go like this. Answer length. Yeah, answer length, for example, here. answer length, right? If you look at the, so that's the, how much on average an answer on Arena is getting, how long is the answer from that model than the sentence length. Let's call them out for people who are just listening. Answer length in your thing here. Opus 5 high answers with 510 words on average. Opus 4.6 was answering 234. and it's really clear here for folks who are just listening that the progression from Opus 4.6 200 words, 235. Then 259 for Opus 4.8 4.7 around that area. Fable 5 is 300 and then boom, Opus 5 gives you 500 words. It's just all over the place. Yeah. Much longer sentences. Let's do one more example, Peter, before we move on to our next guest. Much longer sentences, double the sentence length and we'll share a bunch more data but I think there is, what I want to, make a basic point that you guys are not going crazy. When we look at the data we can also see a similar kind of thing. So yeah, we'll share more data and then yeah, check that out when it comes out. As we said, as we said, it's not just you. Opus is a jargon douche and it's worthy. Alrighty folks, it's time to move on and I will announce our next guest. This week is insane week for video models. We'll cover some of that and more in the next one. Now we have Kfir joining us from a company that creates video models. Kfir, welcome to the show. Been a fan of your work for a while so I would love, I'll take some of the co-hosts back to the studio. I would love to interview you and chat about some of your previous work here if you don't mind. Oh, my bad. Sorry. Sure. Can you tell? Is everything fine? Yeah, you're coming through loud and clear. Love the jacket. Not sure if your jacket is real. We'll talk about this in a second. What do you think? What do you think when I'm interacting with it? Is it real or not? It looks good but we'll now show a few examples. Let me add one from back. You can tell, right? It's really good. Kfir, you've worked at Snap before. At Snap, I've put it in Google. Yeah, and you worked on a bunch of video stuff. I think I remember some of your research as well. Can you give us a little bit, first of all, welcome to the pod, man. Thank you so much. Been following in a family of your work for a while. Can you give us a brief one-two sentence about your career? What do you do? Why do you work on something that you think is cool? Sure. Definitely. First of all, thanks for hosting. I'm also a huge fan of the podcast so I'm excited to be here. I'm a research scientist and through my entire career I worked with pixels. I'm obsessed with these colorful things and then pixels and videos and things that are being generated and being delightful for users. I think my career got a pivot to when AI blew up back in 2022 when all these text-to-image models came out and suddenly people understood that generative AI it's a thing that you can actually you can do things with it so you can monetize it and I was at that moment I was in Google when text-to-image and imagine and Dali it was even before CGPT came out so my research was focused on image and video generation specifically on personalization and edits and nowadays I'm at the cart we're doing the same things just in real time. The real timeness of this is what blows my mind so we're going to show a demo for sure in a second but tell me about the cart we talked about the cart we talked about Lucy model we talked about multiple models before on the show so listeners who continuously listen know but many new folks are joining all the time give us like one or two sentences about the cart as a lab what's going on what are you guys focusing on? Yeah so the cart the cart is an AI research lab originally in Israel we have nowadays offices in San Francisco and New York the cart started as an optimization company which can just expedite and accelerate foundation models like the inference of foundation models like LLMs and VLMs and on top of this optimization stack we build models and specifically video models that ran so fast so people can interact with them actually you know think that they touch things it's sort of a world model right you can interact with the world around you and build new layers on top of you so think about the cart as a company it's a stack with multiple layers we have the optimization layer and we have the modeling layer and now we also have the product layer as you can see we have an actual product that makes these cool things accessible for everyday users here is a here is a simple demo okay so let me let me go here let me see if I can do this super quick as you guys know there's an iconic yellow jacket that yours truly wears on the show this jacket is right here behind me correct? correct with a simple switch of a button if I switch to let me see if I have this here so switch here and go like this you can see that I will stand up and I am wearing the jacket but the actual jacket is still on my on my chair what I am wearing right now is a digital version fully created by the cart in real time exactly super real it's ridiculous I will allow co-hosts to unmute themselves if you want to add voiced reactions you are not on stage but look what I can do I can slowly wow open this is in real time this is the delay between what I see on my actual camera feed and what you guys see and Kfir I have seven more seconds until this session runs out this is I have to refresh it but you can get into the queue this is unreal folks this is in real time let me put myself back up here Kfir what is this what is this voodoo magic what's going on tell us what's happening first of all I'm very excited I just stepped into the office and now we have this big huge board here that shows us how many users are trying these things and it's blowing up many people are trying now this type of demo because it's indeed feel innovative what's happening I see also that Peter asked what's the difference between the approach Vio like models and then the cart real time so just think about it very simply Peter when you like with Vio you write a prompt you wait for a few seconds even minutes in the Vio case and you get back the video so even when you want to edit a video let's say you provide it with an input video you write a prompt you want to change my jacket it will take you a few minutes you will get back an edited video with the cart approach you write a prompt or you provide a reference image of the jacket of Alex and you see it immediately apply to you in real time so the return time the latency is approaching zero it's like we have 40 milliseconds latency per frame so each time we have an input frame that the video shows it goes to the server runs so fast and gets back into the user and overlay basically the input frame so this is the approach the reason that we can do it is we have so our models are so efficient and we have our Descartes optimization stack which is called DOS that enables us to run this thing so fast and we optimize the kernels we have in GPUs to make it so fast so that's the main difference I think and this is while you're talking I went shopping online and now I'm wearing this Dolce Gabbana suit I'm sorry you guys have to check this out this is crazy I will do something here on air that I did while Wolfram thinks I will flash let me see let me close this actual camera so folks don't actually so YouTube doesn't I think you're muted now Alex oh sorry okay so basically when I went off a little bit when me and Wolfram tested this yesterday Wolfram was like hey you can be naked and I took my shirt off and Wolfram still saw me wearing stuff I think the potential for this technology is crazy and I think we all should be wearing some of the stuff it's ridiculous it's Fira tell me about this like how like how am I interacting with this and also did you guys see that it changed my Apple watch into a thing it's is oh now it's back tell me like what why is it interactive what is the cloth simulation like what is the model I'm dude I'm so fascinated I have maybe five more minutes with you but I have so so many questions yeah happy to answer every time so just maybe okay like a comment on the product a comment on the technology so first of all from product perspective we already see that this thing increases conversion in shopping which is this is what we try to prove we did this anywhere yeah like a new experience when I think about let me just tell folks what anywhere is it's a chrome extension that you install right now it runs for free all of that I'm tested like you guys didn't pay me for tokens whatever everybody can do this for a session of a minute this chrome extension I'll add the link to the show notes from it's called anywhere like wear anything you guys inject to all of the fashion sites Amazon whatever a button that says virtual try on and turns on your camera and boom you're wearing the thing like here is how it actually looks in practice please go ahead continue talking while I show yeah yeah sure so it's exactly what you said this chrome extension that enables you to try everything and the technology okay the magic what's the magic the magic the magic is that this model learned from so many like real videos everything that it's so in the past is like real videos of people wearing different things people interacting with garments so imagine that you have this type of data when you see someone is trying something but then you have ground truth of how it looks like in reality so this is why everything that is related to physics has like a very high adherence in the past virtual try on was used mostly with like 3D technology so you know more synthetic it's not very interactive the physical plausibility of this thing wasn't the same and I think that's the that's the barrier that we broke and it's actually a world model I guess you heard about this terminology like world models everybody's we talked about world models a lot on the show and we tried to walk through them some of them tried near real time none of this was as high fidelity as what we're seeing here that's what broke my brain the fact that I can fake the model to go like this and it opens up the jacket it obviously invents what's inside right so like there is hallucination going on but like this just completely completely broke my brain the hallucination and I would just say people thought about when I'm telling people nowadays world models they think about these worlds that they generate and then you walk inside it you navigate you move forward right left but what you see here is a new type of world model it's kind of an open action world model because you can do any interaction you want right you decide what to do and then the model have to react to you in real time and specifically with anywhere I think we're talking about agentic e-commerce right agents will help us to do shopping in the future they will help us to find the right product they will help us with the prices and everything but this thing actually closed the loop because it enables you to try to try it so we believe that world models are the missing engine for a full agentic e-commerce cycle and this is what you see here and Kfir next when this evolves I would love to have you back on the show and learn some more unfortunately we have to move on it's been a crazy week but I'm very very happy that you first of all came out shout out to this verbal thing folks please do try anywhere and and check I don't know go to a fashion website put some stuff on you and then what my recommendation is Kfir you guys have played with as well turn around slowly turn around and see that the model like kind of like does a little bit of a difference thing every time it's like a dream this is what happens in lucid dream states when you're in a place where you don't want to be you just turn around you in a different place the model kind of works like we're dreaming which is fascinating to me I would love to dig in with you at some point with this as well we see people doing crazy things with it it's super fascinating to see how people try to stress test it it's really nice but anyway Alex thanks for inviting thank you so much shout out to the card folks please check out anywhere always welcome back to the show I consider you a friend of the show thank you so much man alright guys from one real time video model to another I want to welcome back to the stage Blaine Brown a friend of the pod and we're adding also Victor from Minimax hey Victor hello nice to meet you and Blaine always always a pleasure to have you always good to chat with you guys oh you came in with the beautiful powerful microphone always alright Blaine this has been an insane week for video models it's been a while since you've been on the show please tell folks Victor I'm gonna read you in a second Blaine please tell folks who you are what you do man and it's so good to see you again yeah yeah I'm an AI tester right I do have a real time a real job that I do every day but it's also somewhat AI related I'm a chief AI officer for a technology company but in my free time I just love digging in like a lot of folks either on this show or that watch this show to the different technologies that are out there whether they're hosted solutions through some of the paid APIs or building actual software and I got into AI pretty much the first year or two that I was in 2022 and 2023 it was almost all open source at that point so that's been a pretty special thing in my heart and so that's why it's exciting to see what some of these new developments have been just in the last few days so just in the last few days let's call them out one by one speaking saying the word one is really funny because just today breaking news Alibaba Tongi Lab released WAN13 which was for the longest time one of the leading open source video models don't have a lot of details about that but I think like they're catching up with Native Sound and 10 second generation Stable Diffusion who leads the pack let me maybe add Peter Gostov here because you guys from Arena are testing video models as well welcome back Peter Stable Diffusion Stable Diffusion is a long time ago I keep confusing them SD CDance 2.5 has been leading the pack for a while we've seen like 32nd generation with just unimaginable realism as well then Flux 3 came out from Black Forest Labs their first video model also came out swinging with beautiful graphics and then Minimax folks with H3 which is now the leading open weights open model Blaine what do you use in your go-to and then afterwards we'll ask Peter so if you'd asked me a week ago my answer might have been different I think it kind of depends on what I'm doing right if I was again like a week ago I think for the last several months at least since like February I think CDance 2.0 was kind of the go-to soda model if you will where if you really wanted to just create some really compelling realistic stuff regardless of what you're doing that was where people went I think and then we saw the fall of Sora 2 shortly after that and that whole deal but then I think it was just in the last couple weeks you had Flux 3 drop that you started playing around with that and that was I think they had their moment in the sun for like a week and everybody was just posting all kinds of clips and it was very it's very impressive it is a very impressive model but then now I think what I think in the last week we've had both Minimax and a stable or I keep saying that too a C-Dance C-Dance 2.5 drop right and so I think if we took in my mind if we took like Minimax out of the picture and we were just looking at those other models for the most part those are all like hosted models even though Flux has a history of open source and that sort of thing it's not like they released the weights that we could actually use Flux 3 on our own machine at this point I think that's maybe planned but you couldn't really do that right and C-Dance obviously you're not going to be running that on your own machine either and even WAN they've got a history of having the WAN 2.1 WAN 2.2 as being open weights but then 2.5 came out and it wasn't I think 3 is supposed to be at some point but I don't really know I don't know if anybody does but then and that's all really exciting but then you have Minimax drops H3 that is like next gen open weights model that can do anything and so that's it takes the wind out of the sales fine tunable yeah 100% sorry to interrupt Blaine but I'm just like open weights hostable by yourself fine tunable we want to say hi to Victor Suertes from Minimax here welcome dude thank you so much for jumping on as well we love when the folks who are building the technologies as you saw with Kvir before joining and we love when creators like Blaine are joining as well because like I think combination Blaine can tell about the experience Peter can tell you about what folks are experiencing on the arena Victor you can tell us about what is so special and how can folks run this in more performance welcome Victor please tell us about Minimax like with one or two sentences it's been a while since Oliver has been on the show and that was in LLM like research this is like completely different part of Minimax right yeah yeah first thank you so much for having me a bit about myself for people who don't know currently I work on the GTM engineering and a bit of the developer relations side for Minimax so helping advocate for their products as well as their models and you know this is exactly what I'm here for really just an opportunity to talk to people such as yourselves and the audience that we have today about what makes our models so special and HD really came in with quite a splash it honestly was trending on Twitter for a number of days I couldn't typically for model releases I'm always seeing other models on my feed but HD was the only one front and center I'm sure a lot of the audience might have experienced as well on Twitter recently but what is HD specifically it's an it's the first ever open weight state of the art on the video generation model with a 33 billion parameter open weight transformer architecture that it can run specifically with we recommend especially on Hungy Face or with SGLang especially on NVIDIA GPUs and as you currently have on the screen right now we did just receive the results from Design Arena and state of the art not only in open weight but just in the video generation frontier particularly I want to point out the video editing because because of our omnicontext representation with all the different modalities as well as understanding like how each modality references one another the video editing capabilities itself of HD is top top of the line I'd say the best right now Blaine any chance I can tap you to show us a few of your favorite examples because usually when you come here you have a stack of things that look here or there etc with minimax H3 if you have some if not yeah let me track let me track some down I mean there's there's definitely some crazy examples I think all of Twitter all of X was lit up with people remaking episodes of the office for some reason yes I do want to talk about this Victor and obviously like this this is has been maybe one of the reasons why C dance blew up early on a while ago because people were just like recreating Hollywood and then there was a whole thing with not releasing in the US with the restriction of hey copyrighted material etc I don't want to put you on hot seat but why do you think most of the folks are trying kind of like things that we saw already maybe a part of the model training versus like novel and like beautiful video generation what is it that draws people to show that those examples in the open source and how does it affect the open source strategy as well would love to hear yeah yeah personally I would imagine just the shock value for that Twitter is always a platform for engagement so people want to see something really novel and really seen yet and IP related things especially with models are heavily restricted on the minimax and they are as well I think the issue thanks for bringing that up can we use it in America yes you can the only thing that you have to do and it takes you a couple of minutes there's an application form that you can find a hugging face email you can reach out to to server locally on your machine and the only reason for that is because of these IP issues the current ongoing discourse that you might know with minimax and one of the major corporations at the moment which is why some of these examples are hindering free use that's the only reason we do have that license but once again you should not feel deterred about it our team is constantly monitoring those inboxes and you will get approved very quickly if you are in one of the regions with restrictions so I do want to talk about how to run this locally and Blaine I think you have a project that works as well would you give us a few words about Maestro as well how do people get the most of these models that are actually running locally and what is the benefit of running this locally versus streaming it from C Dream 2.5 CapCut things yeah so I think the every model that's been released for the last four years when I built Maestro it was really a little bit of an answer to that a little bit I wanted you get used to the user interface that you get from a runway or a Pika or any of the models and it's this simplified gallery and simplified options and so I set out a few months ago to build Maestro to be just this simple way to use open source tool there's a lot of learning curve not just with doing node based type things with Comfy but whenever there's like LoRa's involved and what the weight of the LoRa should be and you might not really know until you do all this experimentation and so the goal with Maestro was to kind of take the guesswork out of that because what it does is it lets you even if you're going to download a LoRa for instance to add capabilities to a specific model like WAN or LTX or really even H3 now is it looks at the hosted site where the LoRa is and all of the guide information that recommends step counts and weights and those sorts of LoRa and so that way it does that work for you and the other component it does is it has a director mode that lets you send a single prompt and it uses its own built-in LLM right now it's Gemma 4 that will write the screenplay it'll write the prompts it'll write the image prompts the video prompts and actually edit all for free all locally on your machine I found a shout out for working on open source that's very important I want to clarify for folks who are listening for the first time since we had you there's a bunch of folks we used to be very to load in the video trained on styles trained on different some people that's why they like open models because they train their own layers on copyright material so that different ways that we can see with running our model specifically and honestly it's been really incredible and why we do it is because intelligence with everyone what we want to see is having these frontier capabilities in everyone's hands for them to create as much that you can access there both the context IR API as well as the regenerate API and what's useful about these is for the context IR it uses different models to optimize all your different references and your different modalities specifically for your local H3 instance and also the regenerate is so you can take that 768 quality video and generate into a 2k quality also using it's not even just an upscaler it actually is another model on the back end that notes blaine i see you smiling once i presented this this video i think this is from justine more i believe yes justine more from a 60 she's incredible testing these models if you guys remember the will smith spaghetti video model benchmark that's been completely obsoleted by anyone in the industry in 2026 this is will smith acting as was this from Blade Runner right it's will smith face instead of inside Blade Runner and he's talking to the spaghetti person they took this benchmark to an insanely insane level blaine one last question for you character consistency is one of the big issues with these models usually right these models now can generate and minimax handles this beautifully what are your current tips for character consistency is that because they're omni reference and you can add a bunch of that that's a huge positive that came out of the H3 models they have their start and end frame version but then the open weights they also have the omni reference version which allows you to load it with some character images character voices you can drive it with actual if you upload audio you can stipulate that it's a voice or it's music to drive the video you can actually upload videos and have it edit that video you mentioned the ability to edit it it's so flexible compared to a lot of the previous open weight models that were out like my previous go-to I think up until now most of the functionality maestro was built around LTX because the LTX 2.3 was really good and you could make music videos with it but it would lose some of the character consistency unless you really tried using like lauras and stuff to dial about adding a feature in maestro that lets you easily capture yourself and save a character that way because right now I've just got an image and a voice reference for myself but it's as good as anything I've seen recreating myself in any sort of environment that's the fun thing is there is IP in there I can put myself in an Avengers movie or whatever so it's just really fun it's a great model so I'll avoid posting the IP stuff but we have Peter here from Arena we want to show that for image video overall Minimax H3 is tied Peter can you talk about Minimax H3 a little bit from a perspective this is an open model that can run fully on your computer that's tied to the previous highlight I would where it ranks on the leader board is that if you scroll down it doesn't even fit in here but if I look at our website that VO3 that was the best closed source video model it's ranked 26 now so and now this the amount of improvement has been and you remember reaction to VO3 right this is like my gosh amazing like video solved and then since then we had so much improvement which is just completely unbelievable so yeah amazing just the fact that I feel like the progress been massively underrated because it's the changes we've seen in the fidelity and the quality and the sound is just unbelievable so yeah well done I want to say like it's one of the things if you look at the wolf meat eating spaghetti videos when it's like all wonky whatever it's like very clear that this is bullshit I check we're over like two hours in the show I want to welcome David Krosha to the stage thanks Alex appreciate it it's great to be here yeah yeah I am co-founder of exe.dev and before that co-founder of exe.dev is a new cloud designed for all of the new small software we're building and it's designed for the people driving it and for their agents maybe could i ask you one sentence on tailscale one two sentences how that's been confounding that and we'll talk about exe.dev for folks you have no idea what tailscale is yeah i'm the co-founder cto of tailscale we started it back in 2019 it is a new vpn it's an overlay network you can connect any of your devices together so they get private ip addresses they can use to communicate that's that's that's the whole product but you'd be amazed what you can do with that it turns out it's really good being able to connect your computers together and exe.dev please tell us why did you choose to co-found another company is that like oh why start another company that's one of those questions that i don't i genuinely don't have an answer to i i make up something slightly different every time someone asks me it's it's just the thing it's the thing i want to be doing it's the thing that makes me feel useful yeah what else would i be doing there's a lot of problems to be solved and companies are a good mechanism to solve them so so tell us about the exe.dev you founded this company i find it incredible and i use the bottle how do you position it what is it and how do people use it yeah it's fundamentally a new cloud and it's a cloud very much designed for putting your putting your agents to work and so an agent can in theory figure anything out but you also don't want to give your agent your aws api key where you keep all your production infrastructure and you have this challenge of needing to somehow is that is that sound on my end or your it must be on my door i'll let the software take care of it but yeah you've got this you've got this problem of you need a lot of vms because you need to isolate a lot of a lot of software together you you need to isolate the work each of these agents do and uh the the previous clouds the existing clouds that we had before exe.dev they all charged you some amount of money per vm and so vms are these scarce resources you tried to figure out how to run lots of programs at once on each vm it was a it was a unpleasant thing instead you just want to start a lot of them and so we we built a platform for that and then we realized there are a lot of other things agents need that quotes clouds don't do so we've been building to that since then folks when you are ssh that exe.dev or you just go to the website and like login there's a button called shelly which is basically you guys at harness would love to hear from you like how that works you basically tell it i want this and then you sit there and and then at the end of the process whatever it is is installed that's how i the super computer was installed open claw for all my friends i opened them a vm and yes to say exe.dev entered into shelly and said install open claw and that's it that's all i did and it just provisions all of it kudos tell us more about shelly and you're like excitement about that and then we will ask wolfred and wolfram and nisten let you ask questions as well tell us about shelly yeah shelly like you said it's an agent it's a i think agent harness is the term of art these days and at the heart of it it's really you as you can use it to write software just like you can use codex or cord and when we started exe.dev it wasn't it wasn't the first thing we wanted it wasn't the big problem we wanted to solve because we assumed most people would sign up and use codex or claw the real challenge was you needed a box to put those things in and that's that's what we set out to do but as part of doing that ourselves we would run into the situation oh we build a little web app for exe.dev on top of the ssh api it's really easy to start a vm on my phone oh wow it'd be great if i could actually like do more for my phone why can't i get this why can't i install open claw on the vm for my phone and so we just ran right into the limit of wow actually the terminal is not a great user interface for agents for a lot of problems obviously it's i have three terminals open my window i use them a lot there's nothing wrong with them but it's just not the solution to every problem and we said we really want an agent in our browser as well and there really wasn't an obvious choice at the time so we built one and it was very much that whole like how do we let you get started and write interesting software on your phone and since then i've written a lot of fun apps walking down the street on my phone using it that's it's good is this your custom harness did you use open source for shelly could you give us more details if you are like yeah you did post something about open source and and tools must be open source this week so i would love to hear about shelly specifically and then wolfram go ahead and ask a question yeah we wrote we wrote shelly it's it's not built on top of existing harnesses it makes the api calls directly we did that because at the time at the time everyone was using anthropic models so everything which meant everyone was using clawed for everything and it wasn't very easy to build on top of it was easier to use the apis directly and it's an entirely open source program so this is a github repo for shelly you can go grab it and use it it's all i think it's apache licensed there's no it's not secret source and it was we decided to open source it just because and this was back in february i think we decided to open source it just because inside your vm there it is yeah you found it inside your vm we think you should own the software in there and we should put as little machinery into the vm as possible instead run the machinery outside it but we really wanted to ship you a web agent and so we're like okay we'll make it open source and we'll put it in there and that way you know exactly what you're getting into and that's that's why it's that's why it was originally open source since then we've come to realize that open source is so much more important than that for dev tools and that's what i wrote a blog post on on the weekend which i think is why why we're chatting now yeah and this was really a claim and it's a claim about the change in the world we're in right now about how dev tools have to be open source now i'm happy to go into that but you say you have a question yeah i see that it is written in go so it's probably done by actual devops people i want to ask how do you handle like sock two requirements and other auditing requirements from customers that that are that are going to be using this and did you see any any difficulties or how you handle potential attack vectors from having an agent that does you can talk to and manage all of your vms yeah so sock two we're working through our sock two uh right now it's a process i spent a lot of time talking to auditors it is a change in how sock two used to generally auditors used to generally require human code reviews of human written code it's a negotiation process to try and change that obviously it's necessary because we're not going to have humans reviewing human written code anymore the fundamental claim i've made to sock two auditors and again i'm still working through this with them is agents are writing code and then a human is reviewing it and yes it's usually the human who actually prompted the agent in the first place but that's okay like there's there's still a fundamental separation of concerns there there's still a human who looks at it before it goes into production who's different from how it was written and that so the fundamental value is still there and i'm hoping that i'm hoping that becomes a generally accepted industry norm because the traditional code review is not something that works anymore like it's not it's not practical it slows you down too much to your second question you asked about like the structure of how like we arrange shelly and xcdev and all the rest the default setup on xcdev is when you make a vm it's isolated from all of your other ones and so there's no way for it to reach outside the vm and affect your account and so when shelly is working the most it can do is suggest you go to xcdev and change an integration setting to give it access to a github repo or to give it access to something and we've actually we've been working on that it can now send it can now produce xcdev urls for access changes and it takes you you leave your vm you go to xcdev and it presents you basically a box saying do you really want to run this command and so there is effectively a prompt there within the vm though shelly has complete control it has access to root and can do whatever it needs and that's actually really important for writing software you want to give your agent complete access to the machine you're running on because sometimes as part of trying to debug your program your agent needs to run tcp dump which needs sudo sometimes it needs to app get install a package you don't have you just got to let it do those things if you constrain your agent and you don't give it these tools you get worse outcomes and give it as much power as you possibly can in an isolated universe is very much the philosophy and each one of these is isolated from each other one then we have you can make ssh keys for hgps api keys for for xcdev that are scoped scope to tags and things like that so that you can give you can give an agent in a vm power to make more vms so it can live in its own sub account but that's that's isolated away from all of the other infrastructure you're running okay and last and to talk to each other i would assume they're they're using tail scale a lot of people do that and we we have some that do that we have we we like to install tail scale in our xcvms to connect them we have a corporate tail net just like everyone else there are other options built in there's a there's a there's a peer integration so there's an integrations tab as a set of features you can add to your uh your vms you can give them it's a it's effectively a secrets manager in that you can give them access to services you can put your secrets into the integration manager and the the keys in the in there never actually leak into the vm that acts as an http proxy and you can use that to let vms talk to each other directly and similarly you can give you could give one vm power over some vms by creating one of these scoped keys and so you have all of those options available but by default it's all it's all turned off because there you want them you want them all in their inbox uh david i want to continue here because i do want to talk to you about meat and the stuff that like you said the open source yeah needs to be free i'm trying to show give me just one second i'm trying to show here on the stage a slide of my recent talk at the engineer blow this up a little bit and this is basically the routing table that me and honestly very honestly claude fable i think back then came up with and then we talked about there is this thing that i coined called zzl continue whether or not developers even read code and basically it lands somewhere in the middle could you talk to me generally about the pr process and also meet as your like solution to that and and that's one example of your blog post which i want to hear more from i think you're you're completely right oh you're absolutely right i should say like claude that no one's writing code anymore that's uh we we do it through agents now and i i'm sure as i say that there's a million developers out there writing code still because change takes a long time but this is that's clearly the future no one writes code anymore and some code is read and like this is true of my day-to-day right whenever i make a change to the xcdev core infrastructure i read the code before i before it goes into production and that's i think that is still correct for me to do that it's it's very rare i'll actually catch a bug or anything doing that it's mostly about like algorithm choices architectural decisions that's what i'm reading for and so some fraction of code is read obviously if i'm building small programs on xcdev i don't read the code at all i just i just build it and the program works i'm good yeah but for those programs i read what i've noticed is over the last few months what i am looking for in the code has changed a lot as the models improve and a lot has changed since the days of me reviewing code for humans i've been reading people's prs or change equivalents a lot of them weren't called prs because perforce doesn't work on github those those changes they the classic thing you look for is style you're trying to match style you're looking for obvious error handling mistakes the sorts of things that the mistakes that are just so easy for us as individuals to make because you as a reviewer you help the person you're remembering and and continuous things as humans we just meet meet people yes yeah that's right what i've realized in the last few months is the code i'm reviewing from agents they're far more diligent than humans are at these details they make big obvious mistakes still they're they're still fallible we're not we're not dealing with something perfect yet which is why we read the code especially around architecture they make all sorts of silly decisions but so do i i i sympathize with them but there's all these details of the code that just don't matter to me anymore i don't need to make sure that you're nil checking records fields in a record anymore because agents are really good at that a lot i can't remember the last time i saw a nil panic in production and our logs and in the world before agents that was a thing that if you're a go programmer you would run into every couple of months some some server would nil pack it's just not a thing anymore and it's not a thing because we got much better at code review it's a thing because agents are much more diligent about these details but if agents are really diligent about these details then i don't need to read for them and they're eating up some of the time i have for reading code every day and that's really that's really important to me because the current limit on my ability to ship code is how much code can i read in a day i can prompt more things and get more commits ready to push then i can actually read and be comfortable pushing in a day and it's a real problem when i wake up in the morning and i realize i have a dozen commits from yesterday to read through because it's a terrible way to start the day and i don't recommend it and it occurred to me that i now know the shape of huge amounts of code i don't need to read anymore oh and i have something that's really good at manipulating code which is agents so let's ask an agent to rip out all the pieces of code i don't need to look at take the diff that i would typically read and remove all the nil checks and remove all the error condition handling from the code because i i can assume you know show me that error conditional handling exists and i'm done i don't need to read the error message and say is the appropriate information in it i don't need to know that you check the the field on the struct is not nil and i don't need to read all the boring imports and the changes to the imports there's just enormous amounts of detail about code that don't matter if you added three parameters to a whole series of function calls i don't need to read them 15 times you can just put some dots in there it's fine and so what meet does is it uses multiple passes through an llm to take an existing diff existing commit and and make it as small as possible for human review and the final diff doesn't fully compile even though there's a lot of semantic requirements on the change so how it actually works internally is actually pretty involved it actually asks llms to produce edit plans and then edits the code rather than simply asking it to produce a diff the first version did that i just asked i handed it a difference they reprint this diff without the important without the unimportant things and it works really well most of the time and then sometimes it would just make stuff up in the middle of the middle of the command and you know there'd be some fantastical code in there which is all it's nice and scary it would fix things that was the worst thing it did is if there was an actual bug in the code i was looking at yeah it would fix it in the process of uh correcting the diff to show to me and then like that obviously i can't code review that so what it actually does is it generates an edit plan for the underlying code and the model says remove these four lines change the extract this chunk and reduce it to an ellipsis or something like that and so meet does that and it processes your commit and gives you that result and i like it it's that's fun it lets me read a bit more code in the day david last question i asked for you i think i'll ask this many many folks who are agentic minded are you yourself in the token billionaire club are you using billions of tokens at this point that's a good question and have you been starting using significantly more since the fable mythos soul level capabilities jumped yeah those are good questions there have definitely been times in the last few months when i have been in that state i'm not sure i am right now one of the things about once your program exists is the change what does xc dev need right now it needs it needs to help people understand what you can do with it and so a lot of my work is extremely targeted and is more products than engineering yeah and that involves me staring at it for hours and making a small change and i don't actually need a lot of tokens to do that yeah i probably burn a lot of tokens in my data analysis i could go look that up i don't i don't track honestly i do have some i guess loop engineering style things that use a bunch of tokens and i only realized yesterday that i had like a deflake thing that was using like 100 bucks every night worth of fable tokens wow i moved it over to sol it's a lot cheaper it works just as well it's fine that's so you know it sneaks up on you you can spend a lot of money with it without even realizing it yeah but yeah i think right now i'm i'm being very targeted in my work but i'm very i'm very happy to spend on tokens right i'm not i'm not gonna try to optimize any of that i'm optimizing for productivity yeah so david kroshoff from exe dev thank you so much for joining and telling us behind the scenes i will send folks to your dev tools must be open source essay which i completely agree with in personalized software something we truly believe here i completely agree with you the future it is absolutely future and you guys are heralding as well like first please follow david on on socials we agree with his positions and thank you so much for exe.dev and with that i think we're at the end of the show thank you so much wolf from niston peter joined us ldj and young were here before if you missed any part of the show go to thursday.news we'll see you here again next week bye bye everyone you