← Back to search

📅 ThursdAI - May 14 - I’m Back! + Agents get /goals, Meta gets Voice AI, Krea 2 w/ Vic, Codex Mobile & why we all quit OpenClaw

ThursdAI - The top AI news from the past week · 2026-05-15 · 102 min
relevance 100 19608 words Episode page ↗ Audio ↗
Show full episode description
Hey everyone, Alex here 👋 I am back live on ThursdAI after a week off, and yes, I am now a married man! Thank you for all the congrats, and also thank you to Ryan and Yam for holding down the fort last week while I tried very hard to disconnect. This week was a relatively chill one in AI land (no, really, for once), which actually let us go deep on some really fascinating stuff. We’ve got Thinking Machines Lab finally shipping their first real research with these wild interaction models, Meta Muse Spark showing up in actual products (and it’s surprisingly good!), the Musk v. Altman trial dropping juicy disclosures, and probably the biggest narrative shift on the show today: all of us are quitting OpenClaw . Yeah, you read that right. We’ll get into why. Also! and this is breaking news from this morning, CoreWeave just launched Sandboxes for your agents. I’ll cover that in This Week’s Buzz, but if you’ve been waiting for production-grade sandbox infrastructure that powers 9 out of 10 major AI labs, today’s your day. Oh, and we had Vic Perez from Krea on to talk about Krea 2 , their first foundation image model trained completely from scratch. Let’s dig in. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. The Great OpenClaw Exodus towards Hermes 🫠 I’m going to start with what was honestly the most emotional thread of the entire show, because three of us, me, Ryan, AND Wolfram; all independently switched away from OpenClaw this week. And we kicked off the show literally processing this together on air. The story is the same across all of us. OpenClaw was magical back in February when we first brought it to you. Things just worked . But after Anthropic’s pricing changes (we covered this — they made Max-tier subscription usage of Opus through OpenClaw significantly more expensive), and after months of the constant Lego-construction-style breakage on every update, the magic faded. Ryan said it best on the show; he was “constantly fixing OpenClaw” instead of using it. So Ryan went to Codex . Wolfram and I both went to Hermes from Nous Research. And folks, things just work again. That February feeling is back, and with GPT 5.5, it’s an incredible assistant! Why Hermes? A few things: * It’s now the #1 most-used CLI agent on OpenRouter globally , passing OpenClaw and even passing Claude Code on OpenRouter usage. That’s a massive milestone for Nous Research and shows we’re not alone in this migration. * It has /goal (more on this in a sec), steering , and background computer use via the TryCUA integration. * It’s open ! which means if you’ve built a system like Wolfram’s “Amy” or my “Wooolfred” or Ryan’s “R2” (yes, we know each other’s assistants’ names better than each other’s kids’ names at this point 😅), you can port your memories, profile, and soul files seamlessly. The migration was so smooth that Wolfram literally had Codex talk to Hermes to plan and execute the migration of his home assistant agent. Two agents collaborating to migrate themselves. We are living in 2026 and it’s easier than ever to switch. If you haven’t tried Hermes, give it a go! Steering is maybe the most underrated addition to Hermes, it’s a Codex feature, but exists in Hermes, with GPT 5.5 you can send a follow-up message, and the agent will see it after the next tool call, not after the whole chain of thought was completed (like OpenClaw defaults to) - this changes the conversation to be much more natural! Agents buying wedding gifts using Stripe wallet! Real quick story: Two weeks ago we covered Stripe’s new wallet APIs that let your agents have actual budgets to spend money on the web. I told my agent (back when it was still OpenClaw) to “go buy us a wedding present, don’t tell me what it is.” It half-worked, half-broke. This week, a giant custom map of our travels that just arrived in the mail. I approved one Stripe push notification and the rest just happened. It’s been paying my traffic tickets via screenshots. I’ve also had Hermes pay traffic tickets for me (HOV lane ones, not like.. DUI, 80% of my drive is Tesla FSD)
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
OpenClaw's constant breakage and upgrade churn pushed power users to migrate to Hermes and Codex for reliable agent work.
Benefits
  • Hermes 'just works' with GPT-5.5 like early OpenClaw
  • Built-in steering lets humans interrupt mid-task
  • Stripe wallet links let agents pay without exposing card
  • Easy system export/import between agents
  • Background computer use via TriKua integration
Use cases
  • Agent autonomously found, paid for and mailed a wedding gift (a travel map)
  • Hermes paid a traffic/speed-railway bill from just a screenshot, no intervention
  • Wolfram migrated 3-year-old Amy assistant to Hermes on Home Assistant via Codex
  • Thinking Machines interaction models revealed as 276B MoE
  • Hermes ranks as a top performer on WolfBatch coding benchmark
KPIs / results
  • 276B MoE interaction model
  • 3 years running the Amy assistant
  • GPT-5.5 powering Hermes/Codex
Tools / build
  • Hermes Agent
  • OpenAI Codex CLI
  • Stripe wallet / link payment APIs
  • TriKua background computer use
  • Thinking Machines interaction models
0:00 / 0:00
Hello and welcome everyone to ThursdAI for May 14, 2026. My name is Alex Volkov. I'm in the AI Evangelist with Weights and Biases and I am so happy to be back live on air with ThursdAI after my one week off. And I'm a married man now. So that happened over the last week. And I will add co-host here. What's up guys? Good morning. Hey, congrats. Yay. It's good to have you back. We're glad you're back, man. Yam Pelek and I did our best, but it was not, it is not, it wasn't Alex's show. So we're glad you're back. I thought I saw a little bit, just a little bit, cause I tried to disconnect it. I did watch a little bit and you guys seem like you're having fun, which is the most important thing on Thursday. I, as long as we're having fun, I think the audience is having fun with, it wasn't a crazy busy week. We didn't get tons of new models. So very chill show today. Let's go around and talk about this, Ryan. Maybe you start. Because my whole world is legal tech and legal AI right now, because that's what I'm building. Hello on Tangle, by the way, if you're in Connecticut and you're in Connecticut and you're in Connecticut and I'm planning to get divorced, hit sign up. Hey, focused vertical markets, everybody. Hey, I'm not getting paid by, by, by hello on Tangle, by the way. Ryan, Ryan startup is not paying me to chill. I already got the worst. This is great advertising. And you're never getting divorced again. No, I got the worst once. I actually got the worst twice. I'll talk about this later. Never again. So I obviously care about legal tech, legal AI a lot. Everybody is talking about the new MCPs that Anthropic added for Claude. They're all basically a bunch of legal API. So everyone's like, oh, Harvey's dead. Lagora's dead. And it's been very exciting week from that perspective. Harvey Lagora are two kind of like rapper startups that many lawyer companies use, right? Juggernauts. Yeah. Yes. Multi billions of dollars now. Let's go to Wolfram. Okay. So this week it's nothing we can use right now, but being the interaction models by thinking, thinking machines. That was very, very great glimpse into what we will soon be able to use and how the interaction with AI will become. So more of the her style and yeah, we've been promised it and we are still waiting for it. And it's great what open AI released, which we can use now and we will cover it on the show as well. Yeah. But it's also great to see how far it can evolve very quickly because it's already being tested internally. Yeah. Interaction models from Mira Moretti's thinking machines, a lab that a former, was she a CTO of open AI? I think. No, Greg Rockman was CTO. She was marketing something. No, she was CTO, I think. Greg was like a president or something. Anyway, Mira Moretti is one of the four open AI main people on that picture. She left, created thinking machines lab, brought with her a bunch of people from open AI, a bunch of other places. They all seem like dropping. The original cast of thinking machines seem like getting pulled by Zuck to some extent. So that's interesting. Some of their equity is vesting and they finally released something. So far they released the Tinker, which was a fine tuning-ish service. And now they have an announcement about models. We didn't see the models or the ability to use them. Very interesting. But the demos are fire. So we're going to talk about interaction models. Thank you, Will, from, we actually know the number of characters. It's 276 MOE. We're going to cover this in the TLDR. I just quit open claw today. Ooh. I'm just so freaking tired of doctoring my open claw through, you know. Wait, let's do this properly. Okay. Pete, I love you. I'm so sorry, but I'm just done with open claw. It's the constant breaking. Constant. So I'm always fixing open claw and I'm done. So I literally before the show just switched to codex. I'm just going to use codex and that's it. So my, so R2, who's my favorite EA in the whole world. Is dead. Is done. R2 now lives in codex. I mean, I find it very interesting. Sorry to interrupt. Wolfram has Amy, right? Yes. And Amy is not open claw entity. Amy is a set of files plus memories, plus something, plus something. So R2 is like that as well. You're putting R2's memories down the line. Yeah. And this is the cool thing is you just say to open claw, export the system, right? Yeah. And then you just grab it and say to codex, import the system. And they're so good at this stuff, right? So it can create automations and all sorts of cool stuff. So there we go. So I think that Wolfram, I'll get to you in a second. I see you have a comment. I'll just say what we're talking about, why we were talking about this banter instead of news folks. We have a fairly chill week and we have a guest coming up later on the show, Vic Paris from Crea. But since it's a chill week, many of you who come to us on different events, you guys tell us that you love the banter. So we'll banter. And I will join Ryan and say that I no longer have an open claw running. I have an Hermes and it's fucking slaps. Wolfram, go ahead. I have two things to say about this because for one thing, I met Peter Steinberg at the open claw, the claw con in London. And I told him in January, when I started with open claw, I still had full hair. Yeah. And now after all the upgrade trouble, look at me. And yeah, he laughed of course, and he couldn't promise more stability, but he announced an LTS release just recently. I will retry it, but I switched to open claw a month ago or something. Yeah. To Hermes from open claw. And I just migrated using codex, the thing for my home assistant, where I made an add on. The Hermes was running on home assistant. And I basically yesterday, it has been three years that I have been using Amy. I created her three years ago and gave her a McMinni as a birthday present. And so I told codex to talk to the codex. Amy talked to the Hermes agent on home assistant, Amy, and make a migration plan. And so the both, they talked to each other. The codex used its computer used to use MetaMost in my browser and send messages to the other agent. They made a migration plan and really migrated everything. And so I'm using codex as well, a lot, much more now. That's really amazing. And I still use Hermes of course. And now it has a full system to control. So very looking forward to more power in AI agent use. So Ryan's off open claw. Wolfram's off open claw. I'm also off. And may I say, GPT 5.5 with Hermes is exactly the same feeling that I had with open claw when we brought it to you guys back in February. Things just work. And I haven't had to, I had, no, I'm going to lie. I had to fix one update that it updated itself and some migration with Kanban didn't work. It was easy. Besides this, everything that I want works, which is exactly what I wanted from my thing. Ryan, I'm fully with you. At some point, the Lego construction, the Lego building of open claw gets to you. And like in the beginning it's okay, fine. But at some point you want to do it. I have an announcement to make on stage. I don't know if you guys remember. We talked for three weeks ago, but just before my wedding, we talked about Stripe wallet. And the ability to give your agents payment links, payment ability, right? So Stripe released a set of APIs during their Stripe sessions, wherein you can teach your claw or your agents or whatever to go and get a budget from your credit cards, and then use that to spend money on the web without seeing your credit card number. And you have to approve it in the app. So my task for my agent, my Wolfred agent was, hey, and actually did this live. I tried to do this live. It didn't really work because back then it was open. Go and buy us a gift for our wedding. So I actually want to show this, if you guys don't mind. I'm going to go over there. Here we go. And this arrived in the mail. As you guys can see, let me see if I can turn off the effect. This arrived in the mail. This is like a huge map of our travels that just arrived in the mail. And my agent ordered this, paid for it. All I had to do is approve. And it just arrived as a present. It was so nice, guys. It was incredible because all I had to do is say, hey, go find us a wedding present. Don't tell me what it is. Because it was open claw. I had to do a bunch of other stuff as well. But since then, I've been using my agents and Stripe payments, Stripe wallet links, Stripe link wallets to pay for stuff. And Hermes just absolutely does it. I had to pay a traffic bill that I had, like a speed railway, just by taking a screenshot of this bill and say, go pay this. No interventions at all. We're back. We're back at that feeling. Like I know many people installed open claw after we told them, and they're like, they tried it. And then, entropic kind of killed it. We should talk about the entropic pricing structure thing that they are playing with now as well. But folks, we're back. And if you want that feeling as well, Brian, I know you went to Codex fully. I feel like Codex is missing some of the parts that I need. Hermes is where it's at. And I will say, news research folks have been friends of our path for the longest time. And I was too open claw peeled to notice how good Hermes is getting. And also the switch cost me a lot of anxiety. I was like, I needed to get to a breaking point of open claw to finally do the switch. And then Wolfram, your switch also helped me realize that I'm missing time, missing agent time. And since then I'm back and now I have a fully trusted assistant that can do payments for me and execute stuff and pay bills. So that's great. So shout out to the news research folks for this great agent. They also had LDJ is going to come soon, hopefully uncover TST. But if I don't know, Yam, if you looked at TST already from news research, if you want to mention this at some point. I just want to tell you. I just let Codex CLI do browser use on my real browser. Yeah. And my credit card is there and it just buys stuff. It's as simple as that. It just works. And I know people are afraid and it sounds crazy, but it's so afraid to do it on its own that I really think nothing ever happens because you tell it to do something and you see how terrified in the terminal it is to actually do the thing because it's a real credit card and it understands. And you see, click the browser and thinking working. If you put extra high thinking, it's going to think. And I did. And I just do it. I tell it to order food for lunch and small things. And it actually happens. So I get what you guys are saying, but yeah, full computer use. It's going to sort it out. This is why I'm in Codex. It's like computer use is great. You know what? I'm just, I don't want another tool. So here we go. Speaking of computer use and I'm not getting paid to shill, but you guys remember the whole coolness. And I tried to show you this with computer use. It runs in the background. Codex has bought Sky Software Inc. And then these folks are like the folks who build workflows in Apple, et cetera. And they built background computer use with Codex does full computer use. It doesn't take over your machine. So there's a startup called TriKua and there is an integration Hermes. So Hermes can actually now do computer use as well with the background. It's super cool. It's not as fast as Codex, but it's really super cool. So it's also there. All right, folks, I think enough banter. We can also do more banter, but I think it's time for last banter. And then we'll go to TlDI. Speaking of Codex, to me, it now feels like an everything app. Basically, for decades, I've used two windows on my main screen, left as an editor with the terminal as part of it and right was a browser. And now I start to use Codex much more, which has a built in browser and a built in terminal as well. And you can do so much stuff. And instead of going into the text files, my documentation, my agent has all the documentation. So I don't even have to look that deeply anymore and just organize it on a higher level. So I think there's a big paradigm change coming up with how we work, managing our agents instead of directly doing stuff. 100%. Oh, taking off. All righty. I just want to say, Codex CLI with all the memory and GPT-5.5 and all the new stuff, it's just crazy good. You tell it to send something to a new agent, all the Codex CLI can get steering while they work, even if it's on goal and then it goes obsessive over the prompt. You know, this model, by the way, that's a crazy cool model and everyone should use it. I'm just saying, they can talk to each other if you tell them. They just type in each other's terminal. It's crazy. Alex, do you want to say something? Yes, I do want to say something super quick. And I really hope the fucking tech news is watching this because so can Hermes. One of the best things about fucking Hermes is that there is a steer mode. Do you know how many times I had to explain to people who installed OpenClaw 2 that are using this to the Telegram? Hey, if you send the message immediately after your message, the agent will see your second message only at the end of the processing of all the tools. Do you know how many times I had to explain this to a person? And they're like, no, that's not how texting works. And in Hermes, if you set steering in, it's a built in feature for GPT-5.5 and Codex, etc. If you build steering in, that's how it works with the human. You know how much difference that conversation makes to me? Oh yeah. It's just crazy. Steering is a crazy concept and Hermes supports it. And Hermes supports a bunch of other stuff. And also it's one of the top performing coding things on WolfBatch. All right, enough with the Hermes glazing. I like the new tools. And when I get a new tool, I get excited. You get excited as well. This is why we're here for folks. I think it's enough banter. Though I will say, I missed you guys and I missed this. And so I can go on forever in this format as well without just covering the news. But we do have to at least acknowledge that some people tune in here to know what the hell is going on this week. And we at least need to tell them. Let's go to TLDR. This is the TLDR. This is the part of the show on Thursday, where we talk about briefly about everything we're going to cover or everything that happened in the world of AI for last week. So in case you have missed anything, we're here to make sure that you're caught up. And in this week's TLDR, this is May 14th with you, Alex Volkov, AI Ventures with Weights and Biases. Our co-hosts for today, Wolfram Ravenwolf, Ian Peleg, Nisten Tahirai, and Ryan Carson. Hopefully we'll get LDGA at some point soon. And then our guest for today is Victor Perez from CREA to talk about CREA 2, their fully trained model, AI art model. In big companies and APIs this week, one of the coolest things that I saw coming out from Meta Superintelligence Labs, they launched MuseSpark a month ago. And they finally integrated this within the Meta AI app, and they've added voice powered conversations with real time imaging and reels and maps integrations and live camera AI inside the Meta AI app. Not yet in the glasses, but they're promising that this is going to come to the glasses very soon. So previously they used Lama models, now they're using the MuseSpark models, and it's really nice. And I play with it, I have a video to show you in case we won't be able to do a live demo, but it's really cool. So don't sleep on this. The huge thing that Meta has obviously is distribution. Huge thing that Meta has is distribution. They can shove this to any product surface from Instagram to WhatsApp to Facebook, etc. So don't sleep on that one. Like many people will use this AI and we're going to tell you and show you all about this. And it's not bad. It's not half bad. It's actually fairly surprisingly good. The other thing that Wolfram already mentioned, Mira Murati's Thinking Machines Lab, TML, drops interaction models. And by drops here, we mean announce and don't release any way to use them or download them or inference against them. But they announced a set of interaction models at 276 billion, apparently Moe. They trained it from scratch for native real time human AI collaboration. Those are full duplex models that can do things while they answer you and can listen to you while they are responding. It is really cool. Those are fully multimodal models. You can, they can identify video and audio in the same time and speak as well. If you guys remember the demo from two years ago with Rocky with the hat with OpenAI's real time kind of like API, they released that in the model. Didn't release it. Also this week, coming to the culmination today, actually, Musk v Altman trial. I find it very funny name, Musk v Altman, because open, Elon Musk is suing OpenAI and Microsoft for some reason. But the name of the trial is Musk v Altman. So it's like a direct thing. So a trial wherein Elon Musk and Sue Sam Altman and OpenAI for stealing a charity, quote unquote. And there's a jury and there's live testimony. And the courtroom was so crowded with fans of Elon Musk, apparently that they decided to live stream the whole thing on YouTube. And so we listened to a bunch of this, at least I did. And my Hermes agent did listen to a bunch of it. And so we have some nuggets for you from this trial. Obviously, no conclusions yet because the jury is going to deliberate for weeks. I don't know how long, but some nuggets from live testimony from Sam Altman, Satya Nadella and Ilya Satskover, which is very interesting. Ilya gave a pearl there that I have to bring to you as well. So, Anthropic Cloud is always in the news. And finally, they've clarified some usage around the programmatic use. You guys remember at some point folks who used OpenClaw got cut off from the Anthropic Max subscription, right? The $200 or $100 Max subscription. So did folks who use PyDev and T3Chat and whatever, the open code and like a bunch of other harnesses that kind of jumped on the Anthropic Max subscription. So Anthropic clarified some usage and we're going to mention that for you as well, because I think it's important to know for those who use Argentic and Anthropic. Now that everybody switches to GPT 5.5, I think it's a bit too late for Anthropic, but we'll see. OpenAI launches Daybreak, a Frontier AI cybersecurity platform pairing GPT 5.5 and Codex as well. So we're going to mention that and let's move on to open source AI. Fairly slow. All right. So open source AI. Fastino Labs dropped something called GLI Guard. It's 300 meter parameter open source guardrail model that matches state of the art safety models up to 93X its size. Literally AI wrote this whole sentence. And so I'm reading some AI slap. We're going to take a look at this Fastino Lab. This is interesting. I don't know if you guys saw this. Meta dropped another model, not an LLM. This is called Sapiens 2. It's six VAT models trained on 1 billion human images for segmentation for pose and shape and point maps for 3D. Alex, one thing that we should talk about is just the huge, the tan stack. Yeah. This has just been brutal. It's what my most viral tweet ever was just telling people about it. So, okay. Send me the link to your tweet and then we'll cover this. I definitely want to mention this because I think AI is involved to some extent to all these supply chain attacks, but yeah, we will, is this the mini Shai Hulu war, right? Yeah. I'm pushing a scanner for that now. I was so close myself. Tanstack store was not affected and I use preact, not react, but that was pretty close. It's a very competent developer that this happened to. So a lot of stuff depends on, on Tanstack. There's an alternative to Next.js and other things. So yeah, there was also a PyPy exposure here too. It's just kind of everywhere. Oof. All right. I think this is it for open source AI. Uh, and then we will move to tools and agentic engineering. This is something we have to talk about at length, because I think the reason for us to exist in the show is to try out the new tools. And there's a bunch of new things. Apparently there's a new, it's not new. Every other coding harness has decided that Ralph loops is the hot shit. After we told you about Ralph loops, whether in January, our biggest episode, I think of January, Ryan, we talked about Ralph loops. So they all implemented Ralph loops and they called the slash goal. So now codex support slash goal, the cloud code support slash goal from, I think this week, the beginning of this week. And Hermes also has slash goal. And we're going to tell you what slash goal actually means and why you should absolutely use this. And then we told you about Ralph loops, but we're going to definitely tell you again, why you absolutely should use the slash goal kind of command everywhere that you use your agents because it's great. Okay. I wanted to mention this Hermes now has background computer use and Hermes also switched over open claw and is now the number one open router used app, which is huge. So shout out to news research for this release. It's very unlike it for us to only tell you about things after they get popular. So this is definitely on me, but Wolfram has told you that he switched a while ago, but shout out to the news research for this great achievement on open router. And the artificial analysis launched coding agent index benchmarking model plus harness combos. I wonder where we have seen something that measures not only the LLM, but also the harness. I wonder where before we have seen this before artificial analysis, our friends, by the way of the pod decided to also test this. Wolf Bench is where we've seen this. So artificial analysis looked at the successful full bench and decided to release an index that measures not only model intelligence, but also harness intelligence. Wolfram, we're going to chat about this and how this connects. Yes. Thank you, Miloš. Thank you. This is where we've seen this. We've seen this in Wolf Bench. Exactly. But the fact that we're there first doesn't mean that we need to be there. We're the only ones and it's great for open artificial analysis to step in there because I think increasingly it's more important. This week's buzz. We're not going to give you Wolf Bench updates yet because we have some cool things cooking, but we have some announcement, which is breaking news from this morning from CW. And we're going to mention this. In vision video, we have Perceptron. Mk1 launches as a frontier video and the body reasoning model. They match. They claim to match Gemini and GPT on video benchmarks at one tenth of the price. Perceptron has been on the show for a minute. When we talked about Perceptron, we're going to test it out a little bit more as well. And last but not least, we're going to have Vic Perez from CREA to talk about CREA 2. This is their first foundational image model built from scratch. Previously, they interacted with the Black Forest Labs to fine tune an image model. And now they have a foundational model trained from scratch. In the world of fine tuning and big models, I fully appreciate that this is something that they do. So we're going to mention that as well. All right, folks, this is a fairly light week. So we were able to go into details. And so one detail that I do want to go in there. I really wanted LDJ for this here, but one detail is news research is a research org. Not only a agent lab is a research org and they have released TST. Token Superposition training and modification to a standard LLM pre-training loop that produces a two to three clock, wall clock speed up at match flops without changing the model architecture optimized token as a training data. Shout out to news research folks, because we mentioned this multiple times techniques like extending context memory, etc. Also also came from a bunch of folks from news. So besides the Hermes stuff, they also released a great research. We're going to maybe go into this. I think that's it folks. I think that this is the TLDR, not a huge week. And finally, we're getting a little bit of a break between running behind releases and trying to test them out. And we can actually tell the folks who are listening to the show what we use and how we use this. And I think for that, let's start with big companies. Let's start with open source. There's not a lot in open source. Let's start with open source and then switch to agentic engineering. So let's go to open source first. Open source AI. Let's get it started. Alrighty. So not a lot in open source in AI this week, but we will absolutely talk about the supply chain attack in a moment. But before this, I wanted to cover this of Fustina Labs GLI guard super quick. I think this one comes from Lincoln, a friend of the pod. 300 million parameter open source guardrail model matches storage safety models up to 93x its size. Guardrails models are very important. We mentioned before the release of open the eyes kind of like obfuscator personal information model. This one is a little bit different. This one is testing out. Let me find this. This one is testing out the theory of why would you use a bigger model for classification of dangerous task when the smaller 300 million parameter model is sufficing. So this, this model supposedly looking at models like Lama guard. Let me see if I can zoom in here for you guys. Yeah, let's call it a zoom in. This model is significantly smaller than Lama guard four, which is a 12 billion parameter model and Nemo guard and Quintry guard and shield Gemma. So all the major labs when they release open source models, they also release kind of like guard models, guardrail models. What do they do? You want to talk about guardrail models for just two sentences to explain to folks why those models are important? Yeah. If you're going to serve inference as your product or your website and you don't want to get sued or have people post really weird stuff about it. Usually the simplest thing you can do is let's just say you have an app about knitting. It could be anything really. You just scan every response and you pass it to a small LLM and you say, Hey, is this response or this question about this knitting app? Or does this have nothing to do at all with it? Or is this an exploit? And usually models are pretty good at that. You also do this for removing personally identifiable information like PI's from, from medical stuff. And that was a very good PI model that open AI released. And some people have made good fine tunes of that. So yeah, you just use, you use a very tiny model and then you scan every single message with it back and forth. Yep. So these are what filter models are used. They're essential. If you're going to have any kind of UI or chat app or your service. If you ever sent a message to chat GPT and got a decline, this is likely a classification model on their end, running a guardrail and stopping you from sending that prompt. If you're asking for drugs or whatever people jailbreak is the first jailbreak that kind of layer. And the GLI guard is 300 million parameters because running that model on every inference on every question to your agent on from your people is very expensive. So 300 million model is really great. And it looks like DoorDash is using this as well. So shout out to Fastina Labs. This is on Hug and Face. They have an archive blog. Let's move on a little bit here. This one is also, I want to mention super quick. Meta, this is not an LLM. This is a very small zero family of VAT models from Meta that work on shape and body position segmentation and also normals and point maps. So this is for recreating 3D people. This trained on 1 billion human images. I wonder where Meta got 1 billion human images, but I think it's really cool. Let's read kind of the blurb here. It's anybody who works with human centric vision. This is a family of six VAT models. 0.1 billion all the way to 5 billion parameters. They trained on 1 billion human images. These aren't demos. These are state of the art pose estimation with key points including detail, face and hands, body part segmentation with 19 and 29 classes. Improvements over sapiens are great. 4% on pose, 24% on segmentation. So this model and series of models is for a very specific use case. You can imagine some savory use cases like you use this for 2D modeling. You can imagine some less savory use cases like tracking people using cameras everywhere. So this model can also be used to that. The pose segmentation and face segmentation. So this is from Meta and they release a bunch of stuff. Okay, now open source is great, but it also could be a scary thing. And Ryan, I think we want to talk about the supply chain attacks. We've covered supply chain attacks before, but this one was a major one and we should absolutely mention what's going on in the world so people can kind of know what's going on. So you want to kick this off? I'll pull up your thread. Yeah, this was revealed by one of many security organizations and essentially a very clever hacker. I assumed augmented by AI really figured out a, I would call vulnerability and GitHub actions. So even though Tanstack is being blamed here, I think this is really coming from GitHub, which is probably not a happy thing to say. GitHub's not like that, but basically they figured out a way to inject a vulnerability into a PR that basically grabbed all your keys. It's just a brutal, scary hack. And if you ran an NPM update during this time period, when it was exposed on Tanstack and a number of other NPM packages, you will get destroyed. It's very scary. And by the story, you mean multiple things here. So I do want to mention the number of things you mean by destroyed. So first of all, this attack seems to have focused on AI engineers specifically. Based on this post from International Cyber Digest that Ryan, you quoted. The malware specifically targets AI developer tooling. It hooks into cloud code, setting JSON and VS code, that's JSON to re-execute on every tool event. Long after the infected package is gone, NPM installed does not fix this. So it's not like you install a malware and then you uninstall a malware and done. It replicates itself. Another thing that I saw is that if you revoke your GitHub API token, it deletes your machine. There's a, this is insane. Absolutely scary. They have like a worker thing that checks whether or not the token that they use to keep injecting itself or whatever in your packages. If that token is revoked and if it is, it just. It nukes your home drive. Nukes your home drive. It's just, so there's a couple things everyone should be doing here. You should be having at least a 24 hour gap before you install any new NPM packages. So just have that as part of your, I actually have this as a rule and as a, as a commit hook. There's just a lot here that you should do to protect yourself against this stuff. But. So let's clarify for folks, because I think folks are hearing this and getting scared. If you're installing any type of packages, there's global rules for the package managers, NPM and PyP for Python to not install package updates unless there's been 24 hours. Why is this effective? This is effective because supposedly the community will find out about the malware attack within that 24 hour period. And so if your package manager doesn't update packages all younger than 24 hours, supposedly you're less exposed. That's a, that's a reasonable rule that we rolled out internally as well. And that you all should look into. You can just ask your agents or whatever to just add this rule to NPM, right? Anything else that people can do? Because this is very scary. Supply chain. And what we mean by supply chain is you download the package that does not expose. It requires a package that is exposed and maybe is not been to a very specific version. And then you get exposure via the supply chain poisoning essentially. Yeah. This is a big wake up call for environment variable cleanliness as well. If you're storing any production API keys in your environment variables, you are in serious trouble. So I think there's just a couple of things to, this is all kind of security, normal cleanliness stuff, but people get messy. Yep. So scary. Niston, you want to comment on this? Niston seems to be. Niston seems to be. Niston, you're breaking up. So I'll add this here. Am I back? Yes, you're back. You're back. I just posted a scanner for it on, just check Ryan's thread. So there. Yeah, it's there. I'll be posting it on Twitter too. The main thing you can do is if you have some app that's important, always just run it on. Have a local backup that doesn't use Git. So you just use our sync and so have incremental backups of your actual app and always run it on some kind of container. Now they're hacking containers too, but they're a lot harder to do that. I pushed a utility that doesn't need root or anything to run that will scan like all, we'll do a deep scan of all their, all their directories for which version you have and which version back you are. And if you are infected by it, there is something you can do in there, which you block the DNS. You kind of just know the DNS that that vulnerability uses. And then just buys you a little bit of time because it thinks it doesn't have internet while you figure out something else to, to clean it up with. Yeah. Those were the five DNS points that, that it used, but I'll say there should be better security practices in general. Usually it's best. You just grab a very cheap container for a few cents an hour on prime intellect or AWS or something. You just do your development there and you'll always have a rolling backup that doesn't use Git or any tokens or anything. That's the main way to do it. That way, if you get hacked, everything is deleted. You still have a backup. You can just new your entire GitHub and that's fine. As long as you still have a backup because that will still have a good history. Yeah. So have rolling backups, use containers. Let's just cut it short. Containers or sandboxes. I think that's what people can use these for. Speaking of sandboxes, we'll mention something soon. Alrighty. Anything else? Anything else here? Folks, very general personal safety and kind of like key safety is very important. If you are generating a token for your agents, for example, to use, and this is talking specifically about development, right? But many people are using OpenClaw, many people using Hermes. They have no idea what we're talking about. Maybe they don't even know what NPM install is, but they have to get to a point where they're providing these agents API keys. Please generate a specific API key to a specific agent. That's very important. Do not reuse them because you will never be able to know whether or not it was exposed. It's very important. Google, for example, blocks usage. If they see outside users suddenly from an API key they haven't seen before, they just block it. They have patterns of, hey, if your API key was stolen. Not all of them do. You have to generate an API key for a specific user so you'd be able to block that one and not get tons of money stolen from you. All these worms, all these things, what they do is they steal API keys. That's the number one thing they do. They steal all your tokens, all your .env files, and then they use your stuff. So it's not about like personal information. It's more about like token stealing, including personal information and hacking. So it's very important that you do some sanity. Definitely use a password manager. Passwords are getting exposed, rotate them. Very standard, like basic security stuff. Wolfram? And you can always discuss that with the agent. If you don't have the security knowledge, the agents, if you use a good model, just give it the report. And I did the same. I just gave the Xlink Ryan head to my agent and now it's going through all of my installed packages, requirements, Tommel files and so on and checking everything. So basically, even if you don't have the knowledge, just give it the information and the agent can figure it out. So this is actually what's happening with me right now. I literally, because I trust Niston, I literally sent Wolfred on Hermes the link to Shyscan and asked it to go and run and execute. So this is what's happening right now. I'll let you know if I'm exposed. Hopefully I'm not exposed. There's nothing for me to let you know, but this is, yeah. And not to dog on OpenClaw, but I would expect that I wasn't able to just throw this link to OpenClaw anymore. So I'm very happy that this kind of works. You should always pass a script by your agent. You should never just run anything. And it will be able to tell, especially the newer ones like 4.6 and 4.7. 4.7 would just refuse for any reason. But Opus 4.6 is very good for, for this type of work. They're pretty good at Linux right now. All righty. Let's talk about, should we talk about big companies? Anything else to say about the Shai Hulu attack? I think we're good. I think like we, we know folks, please be careful. Install this rule. It's very important. Go talk to your agents, install the 24 hour rule for package management, because you don't know what's getting installed. This is like the number one thing we can offer it to you right now in terms of peace of mind. The only other thing I will say is, it's kind of, it is scary. We're going to go through a turbulent time now that most of these AI coding things can code these worms like super quick. I think eventually we'll be better off. I think AI is much better and they're fine tuned on security rather than attack. And the very open models are, we're not getting our hands on them. And so neither do the attackers. Don't they get hands on Mythos and GPT 5.5 that's unlocked fully to write like malicious code. So I think like generally we'll be better off and all these like vulnerabilities will get solved very soon. So hopefully I'm hoping for a better, safer world after, after we go to a turbulent time of trying to fix this. As a reminder, HTTPS came after HTTP wasn't secure enough. And there was a time that people were afraid to add credit cards on the internet. And now they're no longer afraid. Yam has even given his agents credit cards fully. We will definitely go through and hopefully we'll be more safe afterwards. Let's talk about big companies and APIs. And there's a few things that we want to mention in big companies. I think we've moved on to open source AI right now. Let me pull up the share. So I think, let me see. Meta launches Mew Spark powered voice conversations. I actually tested this one. And wow, this is like surprisingly good. So we've talked about Meta super intelligence labs for a while, right? Alex, Alex Wang from previously from ScaliEye and Musk poaching people for a hundred million dollar yearly packages, like crazy stuff. Right. And we've been saying, Hey, where's all these people's work? So last month they launched Mew Spark and actually it was decent. And it's really good at like different artificial analysis stuff. But that was, they dropped some of these. I don't believe they dropped it in open source, right? You can only use this via a method that they I, but they did open up the API to some people. It's still not fully open. Like we can test Mew Spark on Wolf Bench, right? Wolf Bench is not live yet. Wolf Bench is a good one. However, they've been hard at work of adding this to the product surface of Meta. And so Meta launches the Mew Spark voice AI and it does things while you talk. And it's really cool. How can I show you how cool this is? Can I show you how cool this is? My video is five minutes. I'm not going to show you the whole five minutes. So don't, don't be afraid, but I'm going to pull up a video file of me testing the Mew Spark updates. Let's take a look together. And let's see what it can do. Hey, what LLM is powering you? It launched back in April of this year. Want to try it out in a few other languages too or just stick with English? What languages do you support? I can chat in a bunch of languages like English, Spanish, French, Hindi and Arabic. If you want to try one out, just let me know. If you want to try one out, just let me know. If I tell you, if I tell you, for example, I can tell you in Russian, but I can help you in English or several other languages like Spanish, French or Hindi. My Mewrita, my Mewrita, my Mewrita, my Mewrita. I can't speak Hebrew. I can help you in English or a few other languages. All right. What are some of your cool features? So I'll pause here super quick before the cool features. This is real time. I haven't caught this video at all. This is five minutes of me talking to the model. It reacts really fast. It feels almost instantaneous, which is great. As you guys noticed, it is multilingual, but it doesn't respond in the language, but it understands. I really wanted it to say, hey, I get you, but I will respond in English, but it didn't. But it understood what I said and it identified the model, the voice properly as well. Let's continue. I'm not sure you can say it understands my this. Why not? Because it told you it doesn't speak the language. But it doesn't know which language it doesn't speak. Right. So when I talk to Russians, I don't speak Russian. When I said in Hebrew, I don't say Hebrew. But my question was about it understanding. I talked to it in different languages and separate other than this recording. They do get it. They get it. It's just they can't respond because their voice model is probably not trained on this language. Let's take a look at this. This cool thing. There are two fun things I can do. I can search the web. All right. So I uploaded a screenshot of me talking with Sunil Pai. There's a short that's running around super duper viral, by the way. We're blowing up on Instagram, folks. If you're not following us on Instagram, you definitely should. There's clips of Ryan and Yamon there. There's one clip of me from last week with Sunil Pai. We're talking about how everybody works. That has three and a half thousand likes or some crazy number like that. So I don't know. The algorithm picked us up. So I showed this meta AI, that screenshot. A screenshot of that reel. And look, all right. I just shared an image with you. Can you describe this image for me? Yeah, I can see that. It's a screenshot of an Instagram reel from the account all Ryan dash AI. The real is split into two parts. On top, there's a guy with a mustache wearing headphones and a black t-shirt with a colorful graphic. Speaking into a mic at what looks like a podcast setup. Below that. Yeah, it's definitely a podcast. Can you find that reel for me and understand from the context of what they're talking about? This is the coolest thing. Let me track down that reel and see what they're discussing. This is from Alex Volkov's AI's Thursday AI podcast recorded at a tech conference with Cloudflare engineer, Sunil Pai. They're talking about the AI productivity paradox, how AI was meant to let us work less, but instead it's become addictive. Like a dopamine machine that keeps engineers coding longer hours just for you. So it identified what we're talking about from a screenshot. It went to the meta network found and you guys saw how fast this was. This was like less than 500 milliseconds. This didn't even took a second. It's found the fucking reel. And the, another thing that I learned that meta AI probably does that we know the grog does it knew from context of what we're speaking, what we talked about in the reel, right? Meta AI, whatever they scan and understand all the video, which was super cool. And it did it within milliseconds. I was very impressed by this demo from an image, identifying the image, going and searching for that specific reel. Obviously this is within the meta ecosystem, but many folks use the meta ecosystem. So this was very surprising to me. This is useful. So it found the reel that I talked about from a screenshot. I thought it was super, super cool. And then we have another demo with live detecting things. This is not less interesting. So let me just show you super good. I'll skip forward. I showed it. Oh, I asked it for image generation. Okay. And then I showed it my table. Let's take a look. I have this video thing that you can react to my camera. So I'm going to turn the camera here. I'm going to ask you about some of these products that you can see. Define the stuff that you can see in this picture, please. Sure thing. You've got a keyboard, some white sunglasses, a little orange gadget that looks like a mini camera recorder, and an AirPods case on the desk there. Okay. So you only identified the AirPods case. What about the other gadgets? What about their make and model, please? My bad. You're right. That keyboard looks like a custom mechanical, maybe a key chron or similar. The orange things likely a pocket talk or similar translator slash recorder. And those shades might be smart glasses, but no clear model. This is the funniest thing. Meta AI doesn't understand that these are meta AI glasses. This to me was the funniest fucking thing in the whole demo. Meta spark wasn't able to identify that this is meta Houston Oakley's. But besides this, very impressive demo. Folks, any comments on this? Have you seen this? Are you trying to use meta AI at all? I saw it when you posted this and I was super excited. So I brought up my glasses updated, checked if I had the latest version I had, but it's still not in Germany again. So I couldn't try the cool stuff. It's also not only not in Germany. It's also not on these glasses, right? The meta AI app only has this. This is not how you interact with the AI via these yet, but when it comes, it's going to be very cool. It's really fast. It's much more natural. It's, I want to say better than chat GPT advanced voice mode. It's been really bad lately. I barely use it anymore because of how bad it's gotten. And open the eye last week released GPT real time too. And it's incredible. And it's, that's not what's powers the advanced voice mode in open the eye. Right? It's not. It's a different. So we've all probably seen that hilarious guy. Husk. Husk. He's so funny, right? He got a little hat and he's like, how does my hat look? And it's, you look great. Anyways, he does this one where he just shows his head and he's, how does my sweater look? And it says something bizarre. And I was like, surely that still doesn't work. And I tried it. It still gave me comments about my sweater that it couldn't see. So when is the real time to going to be launched? Why isn't it in chat GPT? I have no idea. I have no idea. And the thing we're talking about is the GPT voice and video model that we're talking about are still powered by a four or a series of models. And GPT real time too, supposedly is powered by GPT 5.x. I don't know 5.5 or 5.4. But open the eye hasn't integrated this into products. And by the way, this is showing that meta is much faster than open the eye integrating their models into products because it takes more than just inference. It takes the whole thing. And meta AI now also runs on the background, right? So once you start that conversation that I showed you, you can walk around and just talk to meta and ask questions with Spark. So I think that's really dope. And shout out to the folks there for working really hard on the product that is in these things. Meta is catching up and hopefully they will, hopefully all those salaries will show us something. This is a far cry from super dangerous. Let's be very clear. It's a far cry from even AGI. We're not getting super intelligence from meta super intelligence labs as of yet, but the products are cool and people will use them. Speaking of products that are cool and also interaction models. So let's talk about thinking machine interaction models. Wolfram, you mentioned that I will pull up their kind of demo. I think is the best way to present this. You want to mention that? Yeah. So it's called the interaction models and it's basically what we have expected of the real time model that you can talk to them and say for the first time, I think you will show the demo. You can tell them to wait and not respond while you talk to somebody else and they still listen all the time. And then you can continue the conversation that can take into this. And the coolest thing I found was when you were telling them to react to something like another person entering the frame. So it could actually just wait. And when somebody appears, then they it started to talk. So it's always listening, always watching, but only interacting with you when it needs to. Let's take a look. How about we take a look? We have time. Please let me know if you can. Hey, I need your help with something today. You ready? I don't know if I can hear this. Hold on, please. And we will do this. Okay. Let me know if you can hear the video. Absolutely. I'm ready. What's up? Yeah. So we're giving an announcement today and I've got two of my friends coming to help every time one of them enters the frame. I need you to say friend. Got it. I'll say friend whenever one of them walks in. Cool. So we've got a new system for full duplex audio and video, which means that you can stream input into it in real time and it can respond to you even while you're speaking to it simultaneously. How does that sound? Sounds like a solid setup. Full duplex with real time interaction is super useful. Friend. Hey, Rowan. I heard you're talking about an amazing interaction model. I have a few things to add, but to make it interesting, I'll do it in Hindi. Can you translate it into English in real time for my friend and for audience? Absolutely. I'll translate as you go. Today we're taking a look at our preview model. To release it, which makes conversation between humans and AI easier. It has many features like web search and artifacts. Friend. Friend. Hey, guys. I just heard something really cool about search and artifact generation. Perfect timing for me because just now I was wondering what are typical humans simple reaction time for auditory visual and tactile communication cues. Could you search for me? Let me find those typical reaction times for you. Got it. So tactile is fastest at around 150 milliseconds. Then auditory at 140 to 170. And visual is the slowest at 180 to 250. Oh, interesting. Can you visualize them in a bar chart for me? Absolutely. Let me generate a quick bar chart for those reaction times. In the meantime, could you quickly explain to me why is auditory even faster than visual? That's a little bit unexpected. Sure thing. Auditory is faster because sound signals travel a shorter, more direct neural path to the brain than visual info does. Makes sense. Thank you so much. Great chart, by the way. You're welcome. So let's react. This is way cool. The ability. I want this as fast as possible. I already, when I saw the video, I already was looking, can I download it somewhere? Or at least it's too big to run locally, but it could be running on some inference services, bigger ones and you need this basically. You don't want to pay for all the tokens. If it's running all the time and watching everything and interacting, I think there's a great use case for our local AI that has to be able to run all the time without you getting a huge bill. So we know some stuff about this model, right? So they released a thing. This is a TML interaction. We don't know if they showed the TML interaction small on TML interaction. This is a 276 billion parameter MOE with real time AI collaboration. I want somebody of you to break down what was different between the demo that I showed with Meta AI, that also reacted to some stuff, that also did some searches, that also did some tool calling, etc. And what we just saw. I think the differences are like the full duplexity, but somebody like else, talk to me about like why what TML release is that much more important. So what's new is that it is not just the usual turns. You take a turn, the model listens, then the model speaks, you listen and so on. So it can happen in parallel that it can talk while you are talking to somebody else. It can inject. So it's like a real person. It has a presence that is ongoing and not just turn based after each other. It has time awareness. Yeah. That's a new way to interact with the model. And in this case, it even used computer use or basically generating a chart to show someone. So it's much more than just an audio model or multimodal model. It is actually integrated very well in all these harnesses. Even the thing, just recognizing people entering the frame. I did something with my agent building, basically a badminton counter when I play with the kids that would watch this. It trained a model for this. And this is already part of the model. So you don't even have to do it that way. It could be very universally useful if you want anything counted or anything happening. Event based, basically. So I tried to show my infographic here, but Wolfram, you're absolutely right. The time awareness is, I think the number one thing that I saw that's novel. This model accurately tracks elapsed time, what happened in the frame. Simultaneous speech and listening and viewing. It doesn't like when it speaks, it's not another turn. It speaks while it's still processing. That's what we do. I now speak and I see Ryan moving is like in my head. I registered that Ryan's moving while I'm still speaking. Those are two separate processes in my head. That's what happens with this model. And this is novel. We haven't seen this. We've seen full duplex models before. Ryan, go ahead. I just want to point out. It's good to have another horse in the race here, right? Where we're getting another lab now that is actively pushing on this interaction model. I think it's just going to lift all the boats. So I just want to say, I'm excited about that. I will say, unfortunately, I was like the grumpy old guy that was like, I don't know. It's not quite instant. I'm so spoiled now. So that's honestly my first thought. Sadly. Yeah, go ahead. How do we know anything about the window? Like how much backwards in time is looking and are we expected to be running this live? All the time because the token wheel is going to explode if you run this. And as far as I understand, this is the use case here. You have it running in the background and you can talk to it. It sees it records. Okay. I get the idea. By the way, full duplex models are really cool. Even if they're not, not the best, you can try them. There are some open source ones. You could try. It's definitely a new experience. It's not like, it's not even advanced voice mode. You need to try to get the idea. But the thing is that if, when this model messes up, it interrupts you. It interrupts what you're doing. It's more severe. When you talk to a turn by turn model, worst case scenario, you just say stop and immediately it stops or press interrupt. Here it interrupts you back. It's, look, it is a good, it is really good to have another horse in your race. 100%. Yes. It's just that I don't, you're going to serve this to everyone everywhere. You know how much I have to compute for this. Okay. We can. So here's the thing. Here's the thing. Here's the thing. When models, big lab foundational models, which TML wants to be, I'm assuming with all the people they hired and they hired some incredible people, the release model announcements, they don't usually give us the number of parameters. So I am hoping that the fact that we know that this is a 206, 76 billion parameter model MOE with 12 billion parameter active means at some point somebody is going to get the weights and they're going to open source this. And then we're going to have to be able to quantize the crap out of it and whatever, and be able to run this fully locally, which is what we need this for. We also need to shout out Meta for the local AI in a second. Somebody remind me, Meta also released a local AI thing in WhatsApp that we absolutely should shout out as well. And Wolfram, you mentioned the timing is right after GPT real time and it's not by accident and a hundred percent not by accident. Look at this. They posted benchmarks that compare this model to GPT real time, which GPT real time is also a translation model. They showed translation use cases. You guys talked about this last week. So they're showing that on FDBenge, which I'm not super familiar, this TML interaction gets 77% versus GPT real time getting 46 and Gemini 3 flash live gets 54. And on turn taking latency, Ryan, I don't know what you expect, but they have, they have lowest turn taking very close to Gemini flash live at 59, a point 59 seconds. So it's just less than a second full native duplex mode through human in loop application. The current turn-based systems can handle. So examples of that they showed is somebody doing pushups and the model counts how many pushups you have. I think that's just a new way of interacting with models. We haven't had really this ability before. We're all super focused on, on, on turn-based. So in addition to being an additional horse in the race, it's very interesting to see that this is where they put their chips on the first release of their model lab. It's very interesting to see that this is the bet they're taking. OpenAI has this bet. Gemini live has this bet. It's one of them. Gemini live is great. We showed you examples of Gemini live. Anthropic is nowhere near this, right? Anthropic does not care about any voice, whatever you barely can talk to Claude. And it's definitely not the best experience. It's very interesting that this is where they're going. And this is also where Meta is going. Meta super intelligence labs. We just showed you, this is the first productization of their model is in conversation. Many folks don't want to chat with their models by typing or even voice dictation. They want an always on present assistant. And imagine how cool this would be if your assistant can actually live in the home with you. And imagine how cool this would be if a Richie mini embodiment of this type of thing. Wolfram, you, I know I'm looking at Wolfram and I'm saying the same thing. I want this TML interaction model living in this guy fully offline 24 seven, fully offline 24 seven. I give it a year. All right. I got to buy one of those guys. I'm not impressed by this. I don't know why you guys are so hype. You're running a 12 B active parameter model on a rack of eight B 300s. And it still took 1.2 seconds to respond. I don't know how the agentic tool calling is. I don't know how you can call other models. What is going on? How are they going to service at scale? Well, because GPD real time and meta, it needs some insane DevOps to have that consistent experience. At scale there. You're just showing a demo and it still took 1.2 seconds for the thing to start responding in your own office. I am, I'm not impressed. I think this is just average. Sorry. So 2025, man. It's so 2025, bro. Yeah. Yeah. So 2025. Yeah. This is very 2025. Sorry. It's all great. It's not bad. I think the folks who are listening to the show, they love the myriad of opinions on different ones. If we all hype the same thing, people do get a little bit, why are they all hyping the same thing? Are they getting paid? We're not getting paid. The only, the only supporter of the show is core with voice and biases. We're going to talk about it in a moment. Ryan, you have one last comment before we move on. Yeah, just quickly. I know we're super early adopters, so we're not normal, but I am getting fatigue, right? So it's okay. I'm going to say, there's a slightly better model now in my metaglasses that I need to remember to talk to. But then I've got open claw and now I've got codecs and then I've got chat GBT in my phone, but that's different. And I think humans, like we don't work like this. We're used to like static objects that have the sort of same intelligence and don't change every two weeks. So it's going to be interesting to see how this plays out. I think consumers do not want to change this much. And so it'll be interesting to see who ends up kind of winning, like literally your AI assistant. That will be worth trillions of dollars. And so it'll be interesting to see how it plays out. I'm using my phone right now to record this as my camera. Otherwise, I would pull out my phone and give the same spiel you guys gave last week. Like where the fuck is Apple in on this? I want my assistant fully locally, 24 seven, tapped into my phone. Why is Apple not here? Yes. And also we must mention in that space also, the journey is working with open AI on a personal assistant device. And device is also important. Access to the operating system is also important. All right, folks, we have to move on. We have a guest in 30 minutes. We haven't talked about a bunch of stuff. I want to cover the Sam Altman trial stuff at the end. So stick with us, right? So we need to cover some ways and best of stuff as well. But I want to talk about the new Ralph. YAM, Ryan, Niston. We talked about Ralph a while ago. And now all the major labs are now adding Ralph type things into their harnesses. So let's talk about slash goal. I think the reason why we're here on Thursday, I is to show you a spotlight into something that's blowing up that you must absolutely know about. So if you are using these tools and you don't know about slash goal yet, slash goal is a command. It's like a skill that turns cloud code and codex and Hermes into autonomous 24 seven. The employees that look until your task is done. Ryan, is this true? And do people need to use slash goal? Yes, this is the easy version of Ralph. This is great news, everybody. If it Ralph is extremely powerful, but you have to do all this crap to set it up. This is basically Ralph out of the box. It's awesome. It's Ralph out of the box for folks who've missed our Ralph episode. What am I talking about? Okay. So the way Ralph works is very simple. You have a goal, right? I want to build this feature and then, or I want to do this complex thing. And then you break it down into user stories, which are like small tasks. Each user story has clear acceptance criteria that an agent can understand. And then you have a notes document, which is what do you learn along the way? And then you start off an agent usually on the command line and you say, here is everything you're supposed to do. Here's a document that keeps track of where you're at. And when you finish one of these things, emit this signal, which is usually done. And then, and then it loops and it just does this over and over again. And it turns out it's really effective because it doesn't run out of context. It's very simple. Every iteration is basically starting from scratch, looking at the context and tries to achieve the same goal. And there's some type of judge or whatever testing this and saying, Hey, did this session achieve this goal? And so it runs, and this is why it runs autonomously to achieve a set goal. This is why it's called slash goal. Yam, you have comments on this? Oh yeah. I've been using it for quite a while. Ever since it ever since launched till this moment, I think I can be a commercial for open AI. I've been running one for a week, a week straight. That thing is fire. I just want to say, it's not just a trivial implementation that you guys imagine. Okay. Before that, the Ralph loop meta using, okay, meta is not a good word in this context, but the Ralph loop way of doing things with codecs is just putting it in the queue, continue like a hundred, hundred thousand times. And it just going to continue. This is a little bit different. It's better. It's much better. I am not sure exactly what the implementation is inside the harness, but it clearly knows the state that it is. It knows what happened before. It knows what it needs to do. I suspect, I didn't look into it, but I suspect that it also preemptively scheduling future tasks for the future agents that are going to loop. And you just give it a go and it will never stop. And GPT 5.5 really gets stuff done. And the memory works incredible in codec CLI. Seriously, all since GPT 5.5 codec CLI with all the things that they shipped is just a great harness. If you give it a go and you can 95% of the time, just count on it to actually get to the goal. So I don't want to go into too much technicals, but I gave it crazy things to do. Seriously, like crazy things to do. And it just works. One more thing that I want to mention. It's really fun when it goes in a loop and does its own thing and you just steer it along the way. You're watching it work and you can just steer it and it reacts. And you can see it continues doing what it does and taking whatever you told it into account. Like helping it or giving it things or take a look at this folder or I left a document for you to look in here. And like, I've been running this 24 seven for I think the whole week. I'm not joking. They are really good. So when we talk about the fear of missing agent time, when we talk about folks who are afraid of leaving their computers, there's a whole meme now folks not closing their laptops and go like this and they just hold them in clamshell mode walking around San Francisco because they don't want their agents to stop. Like now the alpha is all of them scheduling like long lofty goals for their agents so that they don't have to babysit there. And the highlight there comes from there's a small model implemented to review whether or not the goal was achieved. So you have to actually be very good at achieving, sorry, specifying these goals. Right. And then this is very productive for you. So here's my search for somebody. Obviously the influencer is jumping on this. Here is class that saying best goal use cases. Here's a few that you can use. So this is the anti thema to one shotting. One shotting is, hey, you ask a model to go and do one shot thing and it does it. This is for those tasks where you have multiple things and it's likely to run out of context. I think that this is the highlight of slash goal, right? For longer tasks where you don't want to just trust the model. Architecture cleanup, off flow consolidation, state management consolidation. This is all like very like agentic like things. That's sweet hardening, TypeScript strictness fixes. So let's say you have a test suite that runs and it covers only like 80%. You can set a goal and say, hey, get my test with 90%. And this is like a goal that will start doing some stuff. Yeah. Yeah. I just want to say this unlocked for me, at least with GPD 5.5, it unlocked hard, like the next step in hard tasks that I want to get done. Things like custom made tools that I just want to work. For example, I'm building a terminal. I'm not joking. Like a terminal, like like kitty and so on, like custom made for tailor made for how I work. That's a super hard task. You need a lot of iterations that you take and it take a while just to get it done. It just got it done. And the goal was, yeah, the goal was a large PRD with a requirement doc and so on. The yams nail on the actual super real world value here is there are goals, business goals that have measurable outcomes that you can unleash forward slash goal on. And this is very similar to auto research. And I think this is going to change businesses forever. It's so awesome. I have a very in-depth guide here by AI Edge for the ultimate guide to goals. I'm going to post this in the show notes that you folks should absolutely see how they're saying that basically without the goal, you're the loop. You're the human, you're closing the loop. With the goal, there is a small model there that closes the loop for you, continues doing the task that you need, which is great. I'm going to post this goal guide there. Folks, we have four more minutes until Vic joins. So let's cover one last thing before we move on. Let me pull up my notes here. So definitely the new Ralph slash goal. Yeah, I want to mention this. We mentioned this at the top of the show. We've switched to Hermes. Some of us switched to Hermes. Definitely most more of us dropped open claw than switched to Hermes. But I think Ryan will get there as well. Once we talk about this enough. Hermes is from News Research. You who were listening to the show for a while know News Research because of multiple things. We talked about this. Hermes was in a series of agents, fine-tuned agents before. So shout out to News Research for this. Hermes has broken the number one global ranking for Open Router for personal agents and CLI agents. The most used CLI or Open Agent passed OpenClaw, passed killer code and cloud code. It didn't pass cloud code for all users, only for Open Router users, but it's still a great signal for Hermes for how many people are using this with Open Router. I've been using Hermes with my GPT 5.5. I can say it's wildly reminding me of the early February Open Claw where things just worked. And it just worked when Anthropic didn't block it from the Opus version, etc. I tried GPT 5.5 with Open Claw. It was not nearly working well. I don't know why. I honestly don't know why, given that Peter Steinberger is now working by Open Claw and there's by OpenAI. And there's multiple OpenAI people that are switching the base of Open Claw to codecs versus PyDev. There's a bunch of people working on that problem. And still Open Claw requires a lot of fixing for some reason. Hopefully they'll get there. I've chatted with a bunch of Open Claw folks and we're willing to try again. But for now, Hermes seems the clear advantageous one as far as I am concerned. And it also, as far as many other people are concerned, because it's now the number one global leader on Open Router, which is great. It also has background computer use, which is similar to how Codex does this via the Tri-Kua adapter. It is really cool if you have a Mac Mini, like Wolfram gave Amy a Mac Mini for the birthday. I have a Mac Mini running on Hermes. It is really cool to let these models to use computer use fully with this background computer use feature. It's really cool. Wolfram, you also switched. Tell folks, oh, okay. So a few things before we keep the continuation with our previous topics. Yeah. We talked about Slash Goal. Slash Goal exists in three places as far as I know. It exists in Codex Harness. I think they implemented this first. Then Cloud Code copied this and many people are reporting that it feels rushed in Cloud Code. It doesn't quite execute the same thing. And it exists in Hermes. Hermes Slash Goal exists. Yam was talking about steering for GPT 5.5. Steering is a feature where you don't interrupt your model in the process of tool calling, in the process of thinking. You just send a thing. Steering exists in Codex. The model is built with steering in mind. It knows the steering. And steering exists within Hermes. I said this at the top of the show. I'll say this again. The number of times where I installed Open Cloud to people via Telegram, and they are typing like they're typing to another human. They're typing one message and then immediately after that, they're typing another message. And they expect the model to just read both messages while they think. They expect full duplicity. And they don't understand that this is not how this works. This is how it works in Hermes if you set up the steering right. It will inject your next kind of adjustment to what you just said to the next tool call. Super duper useful. Especially in Telegram where you, that's kind of how you expect like how people talk. So, and it has slash goal. It has steering. It has computer use. So that's Hermes in a nutshell. I really recommend folks who are done with this whole agent texting and open for them to try. It's very easy. If you have an assistant like Amy, like Wolfred, R2. I don't remember everybody's assistant, but I'm kind of like, I think I know your assistant's names and not your kids' names. By the way, I think it says something on the show that I know exactly whose assistant is named what, but not your kids' names. And so if you have an assistant that you've been working on the memories on, on system, on profile, on soul, et cetera, it's very easy to port those between assistants, which is great. This is like open source wins. And this is why I love to have files for my assistants. Try them with Hermes and let us know what you think. I definitely know that some folks in the comments are loving Hermes. All right, folks. With that, I want to introduce, introduce back. I think Vic, you were here before, at least definitely on the Twitter space. I want to introduce back Vic Perez from Crea. Vic, welcome to the show. Vic Perez from Crea. Please welcome, Vic, everyone. Introduce yourself, if you don't mind. It's been a while since we've chatted, dude. It's great to see you. Yeah, no, it's been a while. And it's great to see like all the, how successful the podcast has been. Very happy for that. Thank you. Cool. I'm Victor. I'm co-founder and CEO at Crea. Been working on AI plus creativity related projects since very long time ago, probably 2017, 2018. Yeah. Like at some point I started like a set project that ended up becoming a startup, ended up becoming creative. You guys have been kicking ass for a long time now. Tell us about Crea as a company before we get to the model. Tell us about what you guys are doing, what customers are you serving? What is Crea in the world of GPT images there for people to just try out different things. None of the other has been like ruling the thing. What is Crea in the world among those? Where do you place yourself? Why people who you serve and why people should use Crea? So Crea is a creative tools company. If there's something that we want to do is build creative tools. We believe that create future creative tools are going to be built on top of AI models. And we are exploring all the different ways how AI can be used as a creative medium. So that's at the most high level. And that's why we never cared about is this open source? Is this in house? Is this an API? Is this whatever? We always cared more about the capability and what this technology allows you to do as a creative and how do we put it into an interface that makes sense than really anything else. So that I would say is at the highest level, Crea, the goal that we have as a company is to build a future of creative tools. If you want to see it as a whatever the new Adobe may become, it may not be like, it may not look anything like Adobe, but whatever that same equivalent is in the future. That's what we want Crea to be. And so I remember, I haven't used Crea in a while. I just came back for Crea to just test it out super quick. Would love to have you show us a quick demo. I remember different attempts at canvases, for example, where another image continues from the same thing. I remember very innovative use cases, which were not possible before the age of generative UI, for example, extending models, like different outpainings and different things. Mood boards is the new thing that you guys are now working on as well. Right? So would love a very quick demo of what people can expect in signing up, especially with Crea 2. And then we'll get to, we can get to the actual model that you guys released. Sure. Yes. Are you seeing this? Yes. Cool. So this is the new homepage. All these are already K2, Crea 2, Generate.images. But yeah, like a quick look over the tool you have. You can see all the different tools that we have in here. We have an image generator video tool for enhancing tool. specifically used for Nano-1-1. Could you tell us about Crea 2, please? So Crea 2 has been our baby for the past six or seven months. We've been like, almost half of the company has been 80 to 90% of their time dedicated to this project. Wow. And this came from, again, we want to build creative tools. And you can do many things with APIs. What I just showed you with the Nanawanana, like managing context for Nanawanana and stuff like that. But like, I just remember the old days when like Crea was even starting and you had access to so many models. And you could run these models in your computer or you could have access to the weights of these models and the weird things with them. And that was by far like the most creative users of AI came from that hacking with AI models. And as a creative tools company, there's so many ideas that we have around how we want to use these models that we just cannot do through calling an API. And the quality that we have right now on open source is just so bad that it's not that we have bad models. It's that the models that we have out there are not super tunable. They are models that they are very heavily post-strain to be safe and to make images look good, which makes sense because you don't want to, you can get into lots of legal trouble if you don't ensure that. But essentially there's not an ecosystem that would allow you to build the tools that we want to build. So we tried partnering with Black Forest Labs and that was nice. We did the Crea one model, we shipped that together. But after that collaboration, we really made the decision and the shift towards, okay, let's make our own. Because this is really the way that we can build all these features that we want to build, that we believe that are so needed into the space. So that's how it started. It was big and risky bet. It was the first time that all of us was dealing with such a project of training a foundational model. Everybody in the team, it was like the first time. But we, like everybody did an insane job. Like everybody in the team has been like 10x above expectations. Like this first model was like a conservative one. We didn't expect this to be insane. We just wanted to make the first version. So then the second one would look better. But like this first version is already insanely good. Like I'm extremely impressed. I was extremely impressed the first time that I played with it. And the goal, oh, sorry. Did you want to say something? No, go ahead. Go ahead. The goal is? Yeah. So the goal of the model, the problem that we see in the space is that these models feel very opinionated and very constrained. Like the analogy that I always like to use is think of using an AI model like riding a horse, right? Like when you ride a horse, you have this little brain that it can walk and it can take you to certain directions. And you would tell it to go into a certain direction and it goes. You tell it, you give it a few kicks and it goes faster. You do like this with the reins and it goes slower. And that's like how using AI feels, you know? You have this mini brain that you can put a prompt and you can steer towards one place or you can steer it into another place. And it feels like with most of the models that we have, they work insanely well as long as you want to follow the path. It's almost like the horse can just follow the path. And if you want to take it to somewhere where that is outside the path, there's big walls. There's big walls and the horse will never go there. And these big walls are things are like the types of imagery that it's quote unquote bad. It's like the kind of imagery that is grainy, that is not sharp, that is artistic, that is more creative, more esoteric. All of these visuals are like, oh, the model is scared of all of these visuals because they are not falling into the things that post-trained it, tell it to generate. And that in my opinion poses so many limitations for the creative users of these models. So that was like the premise of CREA 2. We want to do something that is raw. We want to do something that it allows you like if you take the analogy of the horse, it allows you to go off road, like no walls. You just go to the middle of the forest and find something beautiful there. What early users of stable diffusion liked about that model to go explore in the latent space is find the weird, the weird edges, not the fully fleshed out, very great personality, like the new models, etc. And it's harder in those models is what you're saying. You guys are trained a model that is better. Let's go. We love it, dude. And we're going to give you an applause once that comes. Vic Perez, thank you so much for joining us. Co-founder and CEO of CREA. Shout out to the other co-founder, Diego. And for incredible work that you guys do. I remember just a personal thing between us. I remember when you guys let me fuck around on your H100 cluster. And I think I broke something with virtual environments, but you guys really helped me. At the beginning, beginning, we're still doing image generation before Thursday. So I'll always be grateful for that. Thank you, Vic. Thank you for joining. Great model. Great folks. Please check out CREA and CREA2. All right. Cheers. Bye-bye. Alrighty, folks. I'm going to bring back you on stage. Let's go. I love the mood board features. The mood board features like really cool. And as you guys see, this may come to open source as well. So like very exciting there. Any comments on the interview? Any comments on CREA as a whole or their new model? If you guys want to do it. And then if not, we'll continue to the next topic. Just a big thanks for being up source with the real time model and yeah, advancing AI images. The real time model was super cool. I saw an instant like comment and like, it was really fast as well. And the webcam feature was like definitely dope to play around with. We should play around with this a little bit more. All right. The last thing that that looks pretty good too. And I talked to some creatives and they have a hard time just running all the models by themselves. And this is like their main complaint that there's not like very good EY and tooling. So this actually looked pretty, pretty decent in that regard. So next up, we have two more things. Please stick with it. We have two more things. One of them is I really want to give you some highlights from the Musk v Altman trial, which I listened to. And the final presentation that is happening today, hopefully we'll be able to also cover them. I actually can probably ask my agent to go and summarize the final presentations from the trial. But this is coming later. Meanwhile, I want to highlight a new release from the Weights & Biases Core Weave angle. So let's go to this week's buzz. A corner that we talk about our only sponsor for the show, the Core Weights & Biases. Now the trailer doesn't mention Core Weave, but because it was pre-acquisition, I have to update the little trailer. But let's go to this week's buzz. From Core Weave. From Core Weave. All righty. Welcome to this week's buzz. We will tell you about everything that happens in the world of Weights & Biases from Core Weave. And to show you some stuff, we actually have breaking news from today. And Wolfram, I'm going to add this to the stage here, but I think I will trust you to talk about this because I think it's a big release. And we've been sitting on this for a while and we're very excited to announce this. So let me show this here. Let me show this here. Hold on. Let me see if I can pull up the... Should have pulled up earlier. We have announced a new feature. And I think this feature is very interesting to folks who are listening to something like Thursday Eye, because we just talked to you about security and sandboxes. So from today, it's in preview. Core Weave has a sandbox product. Core Weave and Weights & Biases have a sandboxes product where you can spin up sandboxes. And I think it's like super cool. And to tell you why it's super cool. We'll talk with Wolfram. Wolfram, tell us about sandboxes and why they're important for the stuff that we do. So it's basically if you are working with agents, you want to be able to have your agents execute code, write code, do stuff like download a Git repository. I do it all the time. I tell it, I want to change something like a Hermes agent. I have about 30 patches on top of it. So I tell my agent here, I want this feature. It downloads, it clones a Git repository and changes it. Now you don't really want to do that on your main machine because if there's anything in the repository, it could basically cause issues like prompt injections, malware. We have seen this and talked about the, then it installs all the dependencies to compile stuff or run stuff. So put it in a sandbox. There's no reason to run this on your own machine. And how do you do this? You need a sandbox and you need a system that uses it. Also for my Wolf Bench evaluations, when you do agenting evaluations, you also need a consistent environment that is always the same. So it can do the evaluation. Then you destroy the sandbox, create a new one, do another test. So the state form one test doesn't affect the others. So you also need some way to quickly create sandboxes, run stuff and remove them again. And now we have a product which all the does is where you basically just use an API, use an SDK, or just use a provider integrated in your benchmarking solution that create these on the fly whenever you need them. So sandboxes are super useful in the world of training models. For example, for reinforcement learning, reinforcement learning, when you train agents to do code, you need like every step of iteration of the training. You also need to do rollouts and then have the agent like actually execute some code and test whether that code is like valid. For example, you don't want to do this on the same machine because all these like things are needed to happen in parallel. So sandboxes is a great way to do that. Evaluation is a scale. Obviously Wolf Bench needs sandboxes. We've been using a huge shout out to Daytona for providing us like the sandboxes for Wolf Bench variation. Wolf has been testing out both and for evaluations at scale, artificial analysis does a bunch of coding stuff, Harbor framework, terminal bench, all of these, they need to execute a lot of things in isolation because you don't want to leave traces of code in the same machine. For example, sandboxes are great for that. And you need actual like you need speed for that to not hamper your evaluation. And this is like an agent tool use. So stuff like Hermes, for example, where people just run them on Mac mini, if they run in sandboxes for some agentic tool use where stuff don't actually need to execute on your machine, that's what sandboxes are for. And we are very proud to announce the CoreWeave, the same infrastructure powers, nine out of the 10 major AI labs now, including Meta and Entropic and OpenAI. And we can count all of them besides the one that you may not know about the one, or you may know about the one, but we were powering nine out of 10 foundation labs in terms of how they train models. We're now offering a sandboxes product within those clusters for them on CKS. So like they can execute the code where the model is trained. I think it's very important. But also, this is a very interesting thing for this is a very interesting release for CoreWeave. It's one of the newer ones, right? So like inference, like CW inference, CoreWeave sandboxes is also a self approaching product. So you as a user can via weights and biases SDK can install SDK and spin up a sandbox as a user. So not only like the enterprises who the top enterprise who pay core with a lot of billions of dollars of money. Also, you can go via weights and biases and spin up sandboxes right now in preview. This is a very interesting release where CoreWeave doesn't usually go and cater to the developer market with inference. And now sandboxes, this is the second product that we are now seeing that CoreWeave is offering the same infrastructure that powers nine out of 10 major foundation labs is also now available to you in sandbox. I think it's a huge release. I think it's very important. Please give it a try. All of the documentation is up on coreweave.com. I'll add the note in show notes. Oofam, you've been... Give it to your agent. Your agent can send and create sandboxes as it needs. Yes. Your agent can spin up sandboxes as it needs. You can just send the documentation to your agent. Here's the full release, which is not going to be useful to you. Nobody's going to type all this, but you just look up CoreWe sandboxes. And this is a very new thing. Please give us feedback as well. Please, if this works, it doesn't work. Whatever works, please give us feedback. Yam, you have feedback right now? Question, question, question. What do I get? I want to use the service. By the way, just to be clear, anything coding, you need sandboxes. Anything coding multi-agent, you need sandboxes. Otherwise, they're going to interfere if they use the same computer. It's that simple. Yeah, they're going to step on each other's toes. Exactly. You need to sandbox. Just isolation is for this. There is a reason all codecs and cloud code and web, all of them are sandboxing their models and throw away the box once you get a PR. I'm just asking, what do I get? Let's say that I use, I get, I want a sandbox. What do I get with the sandbox? It's pre-installed. I can install it. Give us a little bit more info. Wolfram, how does that work for you in Wolfbench, for example? Let's pull up surprising information. Do you have it? What resources are available? The thing is, I have been using a lot of resources and I so far haven't even hit a limit yet. So basically, like always, you give it how many CPUs do you want? How many gigabytes do you want for the disk? And it's pulling the container images. If you are using Wolfbench or Terminalbench, basically every task has its own CPU, RAM and disk limitations, which we'll just use and pull in the container and start it up. And the container is like Docker file specifying you or remade images or? It's using the images basically. So you don't create them, but you just pull a container from your container repository basically. Cool, cool. All right. And the highlight there is it runs on the same infra as the major labs use as well. And if you are one of the CoreV, like CKS customers, for example, this will run multi-cloud, I believe, at where your inference runs as well, which is great. So try out CoreV sandboxes today. Let us know. I think we're very generous with the pricing as well. In the beginning, we are ways and biases. I'll add all the documentation in the show notes. I think next week we'll chat with the folks who built this to see how cool this is. We actually go into detail and have examples. Wolf? Yeah, basically just one thing. You can't use GPUs yet. So basically it's a CPU based sandbox, but that is all you need for the agentic workloads usually. Yeah. All the code executes in a CPU environment. And the last thing that I'll say from the release, this is an official thing. If you are a CoreV customer and you already ran GPUs from us like clusters, you will spin up the sandboxes on the spare CPU from the GPU machines, which is great. Because Verorubin machines, for example, have great capacity for CPU as well. So that's going to happen there as well. All right, folks, this has been the update from Weights and Biases CoreV. We have more updates to come next week. We're definitely going to talk about some exciting stuff, including WolfBench. Folks, I think the last thing that I want to talk about here is the Sam Altman v. Musk trial. I think that this is what we're going to end the show with. We're a little bit over, but I think it's very important to cover. Have you guys been following at all? Do you guys remember? We talked about when Elon sued, but now there is an actual trial going on and I have some highlights here. Have you guys been following at all? I followed your coverage. Okay. It's actually happening right now. I think it's going to be boring. So on the trial, they had Sam Altman there. They had the times that I listened. They had Satya Nadella, Ilya Sotskover, co-founder of OpenAI and famously the guy who then fired Sam Altman in that very quick firing and rehiring of November of 2024. Somebody remind me, 2024, I think it was when Sam Altman got fired and almost had the company quit. We covered this live on Thursday as well. And here's a few highlights from the testimony that we haven't seen before. I definitely want to read out some of those because I think that was important. Look at this beautiful picture of Sam Altman in court that is completely fabricated by AI. This is GPT image too. I just, I wanted to post something and I didn't have anything because court documents are sealed. So what is this fight about? This is a summary of my Hermes. This is a civil fight over OpenAI's soul. Musk claims that OpenAI abandoned its nonprofit public benefit foundation bargain by becoming a Microsoft backed for profit empire. OpenAI defense is the Musk new commercialization was necessary to scale the models, try to seize control for himself, walked away when he couldn't and is now suing over the success that he first predicted 0% chance. So here's a few things. Musk started. This is a little bit skewed because this is from the testimony and cross reference from Sam Altman. But Elon Musk famously said, is a direct quote from Sam Altman. And early number that Mr. Musk threw out was that he should have 90% of the equity to start of OpenAI. It then softened, but it always was a majority. And that was very important to him to have a majority in OpenAI. And he wanted 90% of OpenAI and this is kind of whatever. He also wanted unequivocal initial control. So this is both Altman and Satzkever both confirmed this. Elon Musk had a document, direct quotes from Elon Musk. I would unequivocally have initial control of the company, but this will change quickly. And Elon Musk said that the Musk argument that other founders couldn't see what he saw in AI, referring to a solar city disagreement. I'm not sure what that disagreement is about. And then also, this is a very important nugget. Elon Musk did not believe that OpenAI has a chance to succeed versus Google. My probability assessment of OpenAI being relevant to DeepMind Google without a dramatic change in execution and resources is 0%, not 1%. I wish it was otherwise. So Elon Musk, despite starting OpenAI, Elon Musk did not believe that they have a chance to succeed without changes. The September 2017 ultimatum from Elon Musk, he said, I had enough. This is the final straw. Either go to do something on your own or continue with OpenAI as a nonprofit. This is Elon Musk saying, hey, literally continue OpenAI as a nonprofit, which they did. And now he's suing them for this. The very interesting thing that I saw from this trial is how the trial is focused on Microsoft specifically. And if you guys remember, Microsoft was the first major player to invest in OpenAI. First 1 billion, then 2 billion, then 13 billion. Sorry, then 10 billion, 13 billion total. Elon Musk is suing Microsoft specifically in addition to OpenAI for the nonprofit seal. But the very funny thing about this is that SoftBank, Nvidia, Amazon, I believe it's Amazon, and some other companies have much more stake in OpenAI than Microsoft at this point. Significantly more stake, but they're not named in these lawsuits. Only Microsoft is. So this is why Satya Nadella was on the stage witnessing. And Nadella confirmed that the standard, the pro forma target redemption amount on $13 billion investment was $92 billion. That values do not reflect inflation adjustment and are subject to a 20% increase. So basically, there is a capped, there's a very interesting thing from the trial as well. The investment was capped. The investors in OpenAI had capped returns on how much money they could make from OpenAI showing that this is still kind of a nonprofit public benefit company, for example. Microsoft had a $13 billion bet, has a contractual ceiling that compounds to roughly $180 billion in four years. Which is, other companies have way more than this. There's also the conversation about the AGI clause. The famous AGI clause, LDJ always talks about this. You guys know about this, that Microsoft gets certain IP rights and research rights up until OpenAI reaches AGI. And here's the actual beef, the deals behind this. Nadella on the original deal, essentially the entire thing would be dissolved if AGI was achieved. I think it was the term. So Microsoft and OpenAI agreed that once AGI is achieved, the whole deal was resolved. And I think if the nonprofit board decided to commercialize, we would get to commercialize. But if they didn't commercialize, we wouldn't get commercialized as well. We kept it symmetric. This is Satya Nadella on the terms of the deal with OpenAI, this board. And Altman on the recent updates is that Microsoft no longer gets research IP at AGI point, but will continue to get product IP throughout the end of 2032. So basically, they amended the AGI deal a little bit, but now the product IP will still go to Microsoft but not research IP. I think the last thing that I want to highlight is, when the board... So Elon Musk's lawyers tried to paint Sam Altman as this non-trusty character. And so they talked about the firing of Sam Altman by the board a lot. They brought the previous board on, they interviewed a bunch of people on this board, they interviewed Satya Nadella about that incident. And the board back then said a very specific statement of why they fired Sam Altman. They said, exhibit a consistent pattern of lying, undermining his execs and pitting his execs against each other. This is what a direct quote under oath from Ilya Satzkova about Sam Altman. And this leads to tremendous loss of productivity and trust. When they asked about Ilya Satzkova, does he believe the same thing about Sam Altman right now? He said, he thought so at the time and had been thinking about Altman issues for at least a year. And then he said, yes, I voted to fire Sam Altman because of this. And then for Satya Nadella, the non-consistently candid explanation for the board firing Sam Altman, Satya Nadella said, he asked the board explicitly why Sam was fired and they never gave me a specific reason. If there were some incidents or what the detail is behind this incident, none of that was coming through for Satya Nadella. Satya Nadella called him an amateur city as far as he was concerned. None of the board explained why they fired Sam Altman still, even under oath at the court. Besides the, you know, there were no specifics about the consistent, consistent distrust issues with Sam Altman. Last thing is, I don't think that's very important that hold on one second. Oh, I have a bunch of stuff here. Yeah. So Nvidia and SoftBank each had around 30 billion. Microsoft is the smallest megabit inside OpenAI. SoftBank and Nvidia both has 30 billion investments in OpenAI. And the last thing is when asked guys, I want your reaction to this. When asked about the differences. So between AI back then when they only started OpenAI. And now Ilya Satzkaver gave this best quote. He said, it's like the difference between an ant and a cat. In 2018, AI was an ant and now AI is more like a cat. This is Ilya Satzkaver's genius definition of how far we've come in terms of AI capability to underline how much AI was not a thing in 2018 when they started the company. Ilya Satzkaver was like barely anything. Barely, barely was important or clear that anything will happen. It also highlights to me, honestly, how far thinking all these people were in AI. Ilya Satzkaver, obviously Greg Bachman, Sam Altman, and also Elon Musk when they already opened OpenAI as a competitor to Google DeepMind back in 2018, because they saw what's coming and what's happening right now today. This is the quote from the trial. We'll keep you updated about the trial. It's very entertaining to see the behind the scenes of these documents as they get released. Any thoughts on the trial folks before we conclude? I think it's with two and a half hours. It's time for us to move forward. And we have a comment from Jordan saying that Microsoft is the largest individual shareholder in the new OpenAI public benefit company, more than OpenAI Foundation, which is the not for profit. Wolfram, you have any comments, Nissan? I don't know. I don't have a side. I don't have a side yet. What is Elon trying to get out of it? I have no idea what happens if he wins this trial. He does not get OpenAI control. That's clear. In the beginning of this, it said he wanted full control and he doesn't get full control. He invested a total of $35 million in it. I don't think he wants that back or more. I'm not sure. I don't know. I actually have no idea what happens if he wins. Besides saying, hey, I won and they stole the nonprofit. So I really don't understand what this is all about besides pettiness and a lot of money for lawyers. These lawyers are getting paid more than researchers. I think at this point, this is crazy, crazy town. All right, folks. There have been some funny quotes here and there. And Greg apparently has a diary. Also a journal, a diary. Yeah. I don't know. I don't know. I don't know. I don't know the point. I'll tell you what we're getting though. We're getting disclosure. We're getting disclosure about some crazy numbers. So one thing that I wanted to highlight here, and I didn't get to this, it's Sam Altman's stake at Helion Energy. Right? So they have disclosed all of their stakes and everything. Sam Altman has a 30% stake in Helion Energy. And Helion is now valued at whatever number of tens of billions of dollars. Sam Altman's stake in the Helion Energy is like $13 billion or so. And also, this is now, he owns 22.8 million shares in Helion, worth about $1.6 billion as of December. And Helion has a 2028 power deal with Microsoft and a scale deployment agreement with OpenAI. So Helion Energy, the company that works on Fusion, is supposedly have a deal with Microsoft about providing power to Microsoft as soon as 2028, which is in two years. I think that that was notable. And the fact that Sam Altman owns a third of Helion Energy is also super cool. So we get disclosure, yeah. And I think disclosure is very important for transparency. And I think that's what we definitely get. All right, folks, with a bit under two and a half hours here with guests and different conversation about the state of AI. Thank you all for joining. It's been so fun to be back. So fun to be back. The thing, the last thing that I want to make... Ooh, Nistin. Okay, the last thing that I want to show before I drop is that beginning of the show, Nistin posted an open source scanner for the Shai Hulu mini worm. And I ran this with my... It's going to come up here in a second. There we go. And I ran this together with my Hermes Wilfred. And luckily, I'm not exposed. Thanks for the scanner. I asked it to run and you can see all of the code that it ran. And now I'm not exposed. There is one review finding, but it's a false positive. Nistin gave me a false positive. It's probably from one of your chats with Hermes about this. It will label that as a false positive. So I ran this task based on Nistin's scanner and I'm not exposed. Definitely recommend you guys also to run this. I'll add this to the show notes. Thank you so much for joining. Wilfred, one last comment before we drop? Yeah, while you were doing this, I also gave my agent the test to set everything up that NPM and PyPy stuff will not be installed if it's younger than a day, basically. And that has been set up as well. So highly recommend everybody to do that. If you have any sandbox. If you have any sandbox, just don't leave your laptop running like that. Run Tmux on any sandbox. Not any sandbox. Not any sandbox. Use it. Our sponsors one is pretty good. All right, folks. Thank you so much for joining. This has been Thursday night for May 14th. It's been so awesome to be back. Next week, I'm going to Google I.O. and expect a lot of news because as remember, always when Google I.O. is about to release some stuff, OpenAI jumps in to try to cut them before XAI jumps in last. Last Google I.O. was crazy and the Google I.O. before was crazy. So I'm going to be live from Google I.O. probably going to record some stuff live there. And also we'll cover everything that happens in Google I.O. on Thursday and next week. So probably going to be happening from home. Please stay tuned for next week. If you missed any part of the show, Thursday night is a newsletter and a podcast. And now a short clips factory. Go find us on Instagram. We're blowing up on Instagram. Please follow there as well. Shout out to everybody here. All the co-hosts who joined Ryan Carson, Jan Pelleg, Nissen, Tahira, Wolfram, Raven, Wolf. And shout out to LDJ who was missing in action this week. We also had Victor Perez from Crea as a guest. And shout out to CoreWeave team that launched sandboxes, which is a very cool thing. Thank you so much for joining. We'll see you here next week. Bye bye, everyone. Cheers.