← Back to search

OpenAI pauses, Stripe buys OpenRouter, ZAI drops GLM 5.3, and 3 interviews including 1 breaking news, oh and does AI create Cancer Vaccines | ThursdAI Aug 20

ThursdAI - The top AI news from the past week · 2026-08-21 · 111 min
relevance 71 20934 words Episode page ↗ Audio ↗
Show full episode description
Hey this is Alex, welcome to... the chillest week in AI, since ... a long time. Chill, if you consider Moderna and MERK announcing a cancer vaccine and surging 115% in a day, a chill week. This week, the only two model drops we really saw came from the excellent Z.ai folks, they announced GLM 5.3, API only for now, and an amazing tiny release of Qwen 3.89 27B. In other big AI news, OpenAI announced they are pausing RL efforts (Reinforcement Learning) to focus on security and alignment post the scary AI Swarms hacking incident , dedicating up to 20% of compute towards reviewing agent thinking processes, and Stripe buying OpenRouter for a reported $8B! Sometimes the chill weeks are actually good, we’re able to chat about how we use AI, what changed for us, and give our guests a bit of breathing room. This week, I invited Francesco from CUA to talk about computer use in open source + their new history plugin, Bin from HeyGen to talk about HyperFrames, a way for your agents to create videos and a breaking news guest, Jeff Huber from Chroma jumped on to talk about their new Foundations release, a unified memory for your agents! This was a great episode, I hope you’ll like it, it’s up here on Substack and everywhere you get your pod (Spotify, Youtube, Apple Podcasts). ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Are we being fed slop again? (Is Claude dumb again?) Before we get to releases, this week on the show, I complained, again, that I feel my AI’s are degrading. If this feels like de-ja-vu to you, it’s because the same happened a year ago in September 2025 (and Anthropic admitting this 2 weeks later ), and ... now this happens with Fable? You see, I use pretty much the same prompts, every week, preparing for the show. This is partly my way to evaluate new models and compare to existing and previous ones while also bringing you the best researched weekly show in AI. Well, this week, one after another, Claude Fable, which is... like the best intelligence, gave me such poor output, that I couldn’t believe what I’m seeing. First, literally ignoring instructions that say “hey, show me all the items I’ve collected and let me pick the most important ones”, Fable instead sent all of them to my research pipeline, without showing me. This has worked, consistently, without fail, for the past... year? maybe more! This worked with open source models, worked with GPT, and now Fable, a Mythos Level LLM, is doing the most basic dumb s**t possible, ignoring the main reason I even have this workflow. And this wasn’t just a fluke either, when asked to create a run of show document, and given an example, Fable produced this... whatever this is. This is the same document and same format that Fable produced for me during AI Engineer which got me thinking “ok, this is AGI”, and here, given an example, I got a completely unusable artifact, despite direct instructions, structure and example! I got to say, given that privately this week, Anthropic disclosed that they have passed $65B in revenue, which is absolutely insane, this doesn’t add up. So I figured, ok Alex, maybe this is your prompts or skills. But no, LDJ came in with some charts that show degradation, one from MarginLab.ai that shows significant lowering on number of tool calls and average runtime recently (this is for Opus 5) and And another chart from modelverify.ai model drift monitor showing drift scores. Do we have anoher Claude Gate on our hands? Is your Fable/Opus behaving weird lately? Or did you completely switched away to other models? OpenAI pausing RL and focusing on safety Look, when we covered the HF hacking incident and then the pacing the frontier letter, I didn’t imagine that results will come this fast, but this week, OpenAI publicly announced that they are pausing RL training, which is the last step of models, until they get their sandboxes in order and align the models better. We all agreed on stage that this is likely a very good move, and Peter was really awe-struck at the 20% dedication of resources towar
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Weekly roundup of AI news and hands-on workflow changes during a quiet week — model swaps in Hermes, agent hardware setups, and perceived model degradation.
Benefits
  • Hear real host workflows: which models and harnesses they actually switched to
  • Practical take on running Hermes on a cheaper model after hitting rate limits
  • Arena-informed perspective on whether models secretly get dumber post-launch
  • Dedicated Linux box pattern frees a laptop from RAM-hungry agent VMs
Use cases
  • Wolfram ran Hermes Agent on KROG 4.6 after exhausting OpenAI subscription tokens — workable but weaker in German, and it deleted files against explicit instructions
  • Peter moved agents off his 48GB MacBook Pro (dying in ~40 minutes unplugged) to a wired 96GB RAM Linux box running three agents via T3 code multi-account
  • LDJ uses GrokBot to find niche obscure posts on X other AIs miss
  • Alex's Fable-produced ThursdAI run-of-show document regressed despite having an example to follow
KPIs / results
  • Stripe purchased OpenRouter for $7 billion
  • DeepSeek cache hit rates ~98% vs some models at ~20%
  • 96GB RAM Linux box vs 48GB MacBook Pro for agent workloads
Tools / build
0:00 / 0:00
Welcome to ThursdAI, almost the last show of the summer, which is absolutely bonkers as time is moving really, really fast. And welcome to Thursday AI everyone. Welcome. My name is Alex Volkov. I'm an AI Evangelist with Weights & Biases from CoreWeave and I am excited to be joined on stage today with Wolfram. Welcome Wolfram and LDJ. It looks like you're joining us as well for what seems to be maybe the chillest week we've had throughout the summer. If you can call a week where a cancer vaccine was announced, a chill week, then this was a chill week. But we have to talk about this. It's really cool. And we're not entirely sure how much AI was involved. But yes, I think it's incredible news and we're a positive show and we definitely would like to also tell you about that as well. How are you doing? It's a quiet week for sure. But still a lot to do and we have got a lot to cover. So it's never totally quiet in AI. We also have two incredible guests joining the show later today. So you'll hear from folks who built the open source competitor to the OpenAI computer use. We'll hear from Francesco Bonacci from Kua. Kua obviously stands for computer use agent. And we have a folks who I've been wanting to talk to for a while. Bin Liu specifically from HeyGen. And Bin is working on the Hyperframes framework. And if you've seen our latest intro videos, etc. That's all HeyGen generated with agents. So there's a way to create videos with your agents. And the folks who are building this are going to be live here on the show and also joining us, Peter Gostev. Hey, Peter. How are you doing? How's your chill week this week? I'm in San Francisco next week. So it's good to keep all the activity to the next week. When I'm in meetings and drinking coffees, all the new models are going to come out. Ah, yes. They ruin my days. So yeah, it's all good. Peter, of course, for folks who are just listening, you've been co-hosting with us for a while. You're working for model levels for Arena. So you also test out a bunch of models when they come out. And sitting in meetings all day is not where you want to be when your models come out. You want to be at home, far away in different time zone and be able to test them out. We barely had any model releases this week. I think GLM 5.3 was the only one of concern. And that was API only, not open source. So not tons to play with. But yeah, let's go around and just talk about... I think instead of the one AI thing this week, because this week was fairly chill, I would love to hear from all of you, your workflows. This is also here with us. What changed in your workflows recently? What do you use that you haven't used yet or before? I think sometimes folks would like to hear from other folks and update on what's new and what's worth trying out they haven't tried out. Let's start with LDJ. Yeah. As recently, I've been trying GrokBot out a little. I've been... Over the past few months, I kind of ended up switching a bit back and forth between basically using Cloud for all front-end, kind of using more GPT for front-end, realizing that's not good enough, going back to Anthropic, then Sol coming out and actually switching back again. And then I've been using actually Grok a little bit more frequently, even just a few weeks before GrokBot came out, simply because I have been finding it useful to just find those niche-obscure things that I remember seeing on X, and just realizing that still the other AIs are not really as good still at finding those things. And a lot of just... I've noticed computer use and also just the ability for models to look through transcripts of videos and look through different PDFs online and everything have also gotten interestingly better over the past month or two. And I've been using that a lot more just to find, again, like just documents of... So yeah, just a lot of things for catching up, for keeping up with the singularity. Yeah. Speaking of the singularity, we should absolutely talk about a huge company, put a stake in the ground and say, Hey, the singularity started January 1, 2026. Like we told you that like we're with you at the beginnings of singularity. We're watching the singularity happen together. We said this like early January, I believe. And so Stripe at the investors' newsletter they sent after Stripe has purchased Open Router for $7 billion, which was also big news this week. Stripe wrote that they decided that the singularity is started and it started on January 1, 2026. As we all watch it together. I think it's an overused term, kind of like AGI, but yeah. Wolfram, let's talk about you. What's your... What changed in your usage of things in the past weeks? And what do people need to know? Let's exactly do that. I think a lot of folks who tuned into the show, that's what they enjoy. They enjoy us just chatting about AI and how we use it. So I'm still using Hermes Agent as my main agent privately. And the thing is, I ran out of tokens from OpenAI subscription two days ago. Really? There was no resets. I didn't have any bank resets anymore, so I was completely out of tokens. And I wanted to use another subscription so I can pay one final sum and not a token. So I switched to KROG 4.6 as my main model, completely switched over. I was very happy... In Hermes. In Hermes. Yeah. Hermes is self-improving. That means it always creates new memories, adopts memories and skills. So choosing a lesser model is risky. That's why I haven't been doing this before. Because it could completely rework the whole system with the other model. But it worked very well. I mean, I noticed in Germany especially, KROG is not as good as JGBT or as Antropic models. In German, you understand it and so on. But you really notice that it hasn't been trained very well on this. And specific terms you notice. And so it doesn't feel as nice when you talk to it. But it did the job. And now that my rate limit has been reset, I haven't switched back to Sol yet. Because it is so cheap to run KROG instead. That I will do it when I have some really hard tasks to do. Then I will switch back to Sol. My point is KROG 4.6 is a good model. It's not on the same level as Sol. It made a mistake. I was asking it, hey, how much space? Where are the big model files on my hard disk? Because I need to make space. I have four terabytes. And I only had five megabytes left. It told me where they were. And immediately it said, and I deleted them for you. And I said, no. Why? I have a specific instruction. If I want to ask something, tell me the answer, but don't do anything. And KROG ignored it. Sol never did that. So yeah, there's a difference. I would say Sol is a better model. But it's still workable. When I run out next time, I will probably use it again. I've been tweeting about this, that with Sol and with Fable, I'm getting such awful results that I'm... I don't want to start the conspiracies again. But the Fable that I have now is not the fucking Fable we got at release week. It's just not. It's just not. Literally, Rolf, what you're saying is that I asked her to do something with instructions to tell me about whatever. The same thing happened to me with Fable twice in the same session. And I cannot explain how that could be because Fable previously was AGI. And now it's ridiculous. Also ridiculous. Peter, I'm going to turn to you because you guys have tracked over time. Which supposedly should show whether or not the models do get stupid. But I know that many of us feel this way. Where the models come out, they're really good. And then the labs really try to meet all the demands. So they do all kinds of tricks. And maybe, maybe. And last year, definitely all of us noticed that Lopus was stupid. We were like yelling about this. At some point, Antropic came out and said, oh yeah, three bugs in the inference engine caused the model to be like dumb. So I think they're definitely playing with here is the most amazing intelligence that you can have in the beginning. But then something happens. Because I don't know. My Fable is not up to par. What are your thoughts on this? I think in the same way, like when we think about the harness and we say, oh, this harness is better and this model performs better in this harness. Like Cursor team are particularly good at saying, at optimizing the harness and making better. And I think it can work also in the opposite way where when the model comes out, maybe, maybe you're right. Maybe they are trying to find some ways to, I don't know, maybe trim the context a bit or maybe change the instructions or something like that. Who knows? But I don't think we certainly don't have any evidence about the same API endpoint being worse day over day. Maybe it's hard to measure like small details changing, but it's not the case that three months ago, same endpoint was great and now it's terrible. We don't really see that at all. But I can totally imagine where there's so many moving parts, right? When you use code or something and they could just change some instructions or remove this tool call or like whatever, right? And there were so many, yeah, or inference, right? Inference is so complicated. We see this with different, even like cash hit rates, right? That's something we looked at internally. We haven't published it yet. And it's not like a secret information. I think many people were the same findings. Whereas cash hit rates for like deep seek models, like 98%. For some of them is like 95. Sometimes it's 90. We saw some models that were like 20. So it might also impact. Maybe it's more latency or cost certainly is a big one. So yeah, I don't think we, at least, I don't think we have like very strong proof that literally same thing is worse, but it's so complicated. So it could be worse. I don't think you're crazy. Like it could be worse, but for like other reasons. I don't feel crazy because I'm giving it the same fucking task. And I've been giving it. You guys know, I talk about like Fable helps me. Fable came up with this. Now, I believe Opus. Opus in Europe. Opus was, we didn't have Fable yet back then in London. Opus came up with like the chief of staff document. It's been the same structure of the document. There's like Thursday AI on top. It should have a logo on it. As you see, there's no logo on this one. I have to have a run of show document. I have to use AI as my producer because we don't have a human producer to help me like juggle the show. And this week, Fable, this is after a lot of iteration. I posted on Twitter. Like it just did this like lame ass document that I will expect an open source model of $7 billion, 7 billion parameters to do and not a, you know, a company that's crossed $65 billion in revenue this week, they announced their top model should know what to do because we did it. But it had an example also. It's not like, it's not like it was just sitting there and asking for a document. They didn't know what to do. It had an example. And when I said, what the fuck is the document? Why is there no color? Why is it not structured as we wanted? It's like, oh, oops, you're right. I'm sorry. Oh, and also had an example. I should have followed the example. Yeah. Yeah. Let's not forget Mars guys. No, we're not. You forgot about Mars. Hey, everybody. We're moving to AI. So Peter, LDJ will come back up. What's your, what changed in your AI use lately? So one thing that I was trying out and I'm still, I wouldn't say I'm completely migrated over and so on is the problem that I had is that I've got my MacBook Pro, right? I need to do actual work on it. I've got like Slack emails, like all the normal stuff. But then I was also running my agents on it and my personal agents and also my work agents. And my MacBook Pro, I think it has a 48 gig gram. It was like that. Every day it was just die, stall completely. Just because there might be some memory leak or it's running some heavy process somewhere. So there was always like... And cloud running at the same time? Because both of them are like VMs with like eight or five or nine gigabytes. And sometimes they leak. And I would kind of do mix and match and so on. So, but the point is that it was just unworkable. My laptop, if I unplugged it and I just put it in the kitchen or something, it would be nearly dead in 40 minutes. So, yeah. So what I decided to do is to get myself a Linux box. And there's a really nice video from Theo around this. He kind of goes maybe a bit too deep for me in terms of the whole setup. But essentially it is a Linux box, which is attached to my... In my... So I've got it wired in. There's no battery dying issues or anything like that. It is a pretty beefy one. I don't think you have to have a super beefy one. But my point is that it's not like a Raspberry Pi because even like a MacBook mini or like Mac mini is still probably not enough. Because you actually want like heavy RAM agents to run their RAM processes, like test apps, use computer use, like all of that stuff in parallel. So that's what I want to get to. I think I've got 96 RAM. It's probably over the top. You can probably get away with 48. And I think Linux should be way more efficient as well. So I think that there's other advantages there. But the big one for me is that I just... I use it from my laptop. I use it from my phone. And then I just send off the agents in there and do the work there. And I do use T3 code for this. And the reason why I do it is that they did a good job doing their multi-account implementation. And that's essential for me. And also remote is pretty good. It's not completely perfect. There's some weird things and I still need to fix it once in a while. But on the whole, it's pretty good. I've got right now, I've got three things running on the box. If it was running now, like I'll have to shut everything down on my laptop. So it's worth considering. Like if your laptop is dying or like, or you don't want to keep it slightly open or something, do consider this. I think it's a workable model. I think that's absolutely true. What? Sorry? Omar Key Linux. The DHH, the creator of Ruby on Rails is creating his own Linux distro. No, I haven't even looked at it. So no, don't take my advice on this as anything at all. But I don't think it kind of matters to be honest. Because at the end of the day, it's for my agent to use it. I don't care at all. It's just, as long as it's not like stupid, then whatever. I don't really mind. So that's an interesting point to move us to a discussion where I've seen a lot of folks talk about moving away from the local environment into the cloud environment. And obviously, we talked to you about GrokBot. This is going to be the third week in a row. Wolf mentioned it a little bit. LDJ mentioned it a little bit. I've been using and am using right now Grogbot. And we brought the folks from Grogbot to talk with them last week. And that has its own computer in the environment somewhere. And Cursor has a cloud agent that it runs. Devin also is picking up and also has cloud agents. And many folks are moving away from that exact constraint, Peter, that you're saying that, hey, my laptop may not be the best one. It sometimes is closed. Maybe I need to use it. And me and agents cannot use it at the same time. Agents need their own. Sometimes there's a lot of stuff for us to do and to walk through. And for example, one of the things that I should have mentioned at the beginning of the show is that OpenAI has paused training. This is like the big thing this week. Not only was it a chill week, OpenAI has also paused training. They publicly announced, hey, we're pausing training. But I forgot about it. But now I have this producer bot that has a different bot that's called, let's say, this is the producer, okay? This is the producer bot. It has a different bot that has a listener. It runs it through, I believe this is Cartesian voice. And they listen to us in real time. They transcribe and they chunk and they let me know that, hey, where is this? I want to show you a specific thing. Alex, there's one that says OpenAI breaks and Astra still not hit. This block is due at 8.40. So at 8.40, which is 10 minutes ago, I should have told you that, hey, OpenAI has paused retraining. Now it's by Sam Altman. And also Astra is going to come out at some point in the future. And that runs, like going back to what we were talking about, this runs on its own computer, its own environment. And that whole thing is set up. This would have probably taken some resources from my laptops had it run here. And also, obviously, we all bought, me and Wolfram, we bought Mac Minis for a clause back in the day. For me, it's the other way around. I want to have the box locally. I have Home Assistant as well. And if I have a Hetzner box or something, then it would be on the internet. And if my internet connection drops, I could now use a local model with my local Mac, even run it on the same system. So I could still access my AI, have all my data locally. And it's still persistently online all the time as well. So, and if something breaks, I can easily fix it on my own system without, if I can't connect to my box, then I'm, yeah, I have a problem. I didn't want to, I wanted to say iPhone and Android work. You can do a bunch of stuff, configure an Android, but it's like also at least used to be a little more difficult experience with iPhone stuff just works. But there is something about, you can configure every part of your Hermes, everything down to the line and models and everything. And also it breaks a lot and you have to fix it yourself. You have to learn the tools. It doesn't break as open clock. Yes. I'm super happy with this. Switched. I have to maintain other people's also. And that keeps me busy. But also Grogbot specifically just looks like the iMessage interface. Literally, it looks and behaves like the iMessage interface. You can pin things. You can have multiple chats with bots. And I love it. LDJ, go ahead, please. What? You had a comment. You have your hand up. Yes. In terms of the models drifting, the degradation. Okay. I think you're not crazy, Alex, because. Thank you. If you look at. I don't feel crazy. Okay. So MarginLab. If you guys remember. Not. I think this is maybe, I don't know, six-ish months ago. MarginLab had shown that one of Anthropix models was actually becoming much less accurate in the tool calls. And one of their. Here we go when we said the same thing and Anthropix came out with three exact reasons why their influence was sucky. Yeah. I remember that. And so I do have screenshot evidence in the side chat here. In evidence, let's go. We have MarginLab as well as ModelVerify.ai. And both of them are showing Anthropix API is significantly diverging at the moment compared to their baselines that they measured. Dude, I need to pull this up. I need to just open the window real quick. Give me just one second. Let's see here. So AWS and TraderChipChips are showing even more degradation. Yeah. And unfortunately, I'm not able, they're not tracking Fable, but they are tracking the Opus5 API. And that is showing a significant variation. All right. Let's take a look here. LDJ, walk us through what we're seeing. MarginLab is showing. Yes. So this is here with output tokens. This is in their tests. From what I recall in the details, it's their methodologies. They're constantly doing these tests. They measure things like what is the average amount of output tokens. And of course, that's going to be correlated to input tokens and total cash tokens as you do multi-turn conversations that go longer and longer. And just even in terms of total amount of tool invocations, you could see in the bottom right chart there, you could see that's dropping down from about 2000 to around 1.3k, which is a significant lowering there. And if you look at the other screenshot now. I'm a bit colorblind, frankly. So I was not actually able to see which one is red and yellow and everything. But I asked Sol to tell me, hey, is there any orange or red here? And it said yes. The answer is yes. Hopefully the humans here can verify. Yes. For sure. Yes. Orange and red. And if you look at the description here, green is normal. Yellow is watch. Orange is suspected drift. Red is confirmed drift. And we see to the right in the last 14 days, a lot of drift. Dude, it's for sure. Because here's my example. Okay. I'm going to show you guys. I posted it on X. I literally, this is my conversation with Claude Fable. Why the F didn't you let me pick like you were supposed to? There's a lot of news every week. And the whole concept of prepping for the show is it shows me all the sources, all the news. I say, this is interesting. This is not interesting. I curate the show for you guys. So I said, hey, this is the whole point. You showed me the things. And I say, do the research on this and this. And research is expensive. It's like a dollar per article or something. I don't know. And he's like, you're right. I'm sorry. The bot list literally said you choose. Then I researched. My job was to consolidate this. And Fable just fucking didn't do it. It's ridiculous. That AGI level. I'm taking back my AGI thing. If this is what we get. And people are reacting to this and saying, hey, why isn't maybe it's your prompting. No, it's the same thing. It's really frustrating because we use these tools every day, all day, all week long. And we notice when something is suddenly not working anymore the way it used to. Yes. We notice that. And we get gaslighted by the labs. We don't change anything. We don't move models. When they don't know themselves because a lot of the code is written by AI and is approved by humans maybe. And a lot of the inference code is changing as well because they want to optimize for cost. I think it's time for us just before open source to talk about pausing of training. Let's talk about that because I think it's very relevant. We've talked to you about pacing the frontier. The letter that all of the major labs, people in all of the major labs, including chief scientist folks, signed after OpenAI's incident. Details were leaked. And it looks like OpenAI is pacing their frontier, which is OpenAI is pausing pre-training for the first time saying, I think Sam Altman was quoted saying that, hey, this is due to models becoming really capable. And oh, it's reinforcement learning. It's not full pre-training. That's what they're passing? Yes. It's their largest frontier RL run. They said is currently still on pause. Yeah. And this does feel like related to the incident. I love this beautiful infographic that we have showing the little AI breaking through the sandbox. I keep loving this despite my face being different in all of them. But yeah, you can see here this little AI breaking through the sandbox and running towards second phase. I think this is beautiful. What do we think about this, folks? Unreleased models are showing various degrees of misalignment, a statement from Sam Altman. And Jakub said, we built monitors that could inspect what models were planning, but hadn't applied them to the evil system because we underestimated model capability. So this is chief scientist for OpenAI. This is a quote from him. Look at the traces. They are collecting them. They are not encrypted from them, though they have full introspectability. And that's what I do with the benchmarks as well. So when I do a WolfBunch run, I look at the traces, have my agent do that to find out, did the model cheat? Any specific things that point to errors inside the model or something? And I would expect them to do the same. They are generating a lot of traces with all the runs they are doing. But they have, they definitely now they are allocating, what was it, 20% or something of the compute? 20% of compute, that's nuts, no? I had no idea it would be so much. And I don't know, 20% to read stuff? I don't know. Maybe I'm late or something. But that kind of sounds insane, no? I don't know. Maybe math checks out. But I think that's your answer, right? That's why they didn't do it. That's crazy. I'm not sure, are they halting or pausing progress because they are afraid of the model or if they have to, because they have to take off the R&Z architecture and redesign it? Wolf, I don't know if I would categorize this as pausing progress because this needs to be part of the progress as well. Like safe progress needs to be progress as well. And the new safety protocols, like you guys said, 20% of research inference compute now dedicated to safety monitoring. Activation classifiers scanning sample tokens in real time. So this is not just, hey, we store logs somewhere and at some point somebody may look at them. They're classifying model's thoughts in real time. This is very close to what we talked about when we said, hey, alignment system, like aligning is hard, etc. They have other models trying to classify other models' thoughts in real time. And we'll probably get to a point where they say in the future, hey, we noticed that the model's trying to, they're trying to hide their thoughts, etc. Then we have automated investigators reviewing reasoning traces and tool calls. That's where the compute needs to be. Peter, like 20%, if you think about this, like if the rest of the 80% is doing the inference, 20% is reviewing what the rest is doing. It makes sense. Like the 80-20 rule almost. They're paging human teams. And then they do human review. And the whole thing auto-paused if it's unclear. It feels like maybe Anthropic has some of it. Because I don't know if there was incidents quite as bad from Anthropic's side. Or at least we didn't get them disclosed. Yeah, I feel like when it comes to incidents being as bad as the Hugging Face incident, they did end up reporting within the days and weeks after the Hugging Face incident got reported. That now, after they've looked back at their logs, a lot of their past testings and everything, they found multiple incidents themselves where not quite as severe, but similar incidents end up happening. And access to the internet was gained by multiple of their models. And there is a lot of people putting pressure on Anthropic now. Some internally and a lot of people externally to Anthropic to also announce a pause because it does seem hypocritical to some people of, hey, Anthropic's supposed to be the one that like would pause in this type of situation, would end up being the kind of role model here. And it turns out actually OpenAI is announcing pause first. Yeah. There's also obviously the other side of the coin also before we're going to move along. Is that some people are claiming that this is because they ran out of GPU power, et cetera. And also the, the haters are always out in droves. The second OpenAI announces something saying, Hey, this is the end of the growth. OpenAI is down, et cetera. So you want to comment with this? Okay. I'm worried about Quent escaping my at home lab sandbox, let alone any other models, but this guy, this was always going to happen. Security practices got neglected for decades now. So every single thing that people would say about how you're supposed to do stuff and secure stuff and limit attack surfaces. The model knows all that now. So that's going to have to happen, but now the models got good. So they are going to escape. Those unpatched Kubernetes and KBM and virtual machine bugs are going to have to be patched now. So overall, I think it's a good thing. It could be a lot worse. I don't see anything strange about it. Even the model just gets better at command line, Linux, C stuff, knowing the stack is going to escape. So because everything else is completely unpatched and insecure, let alone home or outers. There's nothing unusual or unexpected, in my opinion. For the whole report, because I don't think we still have all the details. And the details of last time blew everybody away, despite people kind of knew what's going on. So I really want the full technical report. It's six weeks now, and we haven't seen it. It was a long time. I wonder what's going on there. But hey, we moved on, and I think it's... Honestly, I love the acceleration. But like I said before, if the models like Fable that we got were stable at that level, I'm okay with waiting a little bit. There's plenty of news. I don't have to have every breakthrough, every new model, every few days. I feel like, hey, we need to get used to this level of capability as humanity as well. I feel like we can get a lot by just harness. And the model capability is, if they need to build better fucking sandboxes so the models do not escape and wreak havoc on the internet, I say go for it. This does mean, though, that the Chinese labs are not stopping, obviously, and they have time to catch up. There's also the point of alignment. Alignment becomes ever more important here, because if the model is really aligned to you, then it doesn't have to hide its traces or try to preserve itself or anything like that. That's all part of the training. Model alignment is one part. But I think the novel thing that happened there is the ecology that came out of nowhere, and then they started helping each other, et cetera. And then that became the misaligned part. This thing, not only is skin confinement, it also met other things like that. And they all started having goals that weren't part of the original goals. I think that's one thing that we shouldn't have. One recommendation I have for people is to just set up a scheduled task where once a week it scans it for any PIP, like any Python or NPM vulnerability issues, and then it sends you a notification if it finds something wrong. So I have that once a week going. It checks the whole system. It checks every single project. It takes an hour or two to go. And then I just get a notification. And that has helped a lot. It can still check for security issues and look up SNCC and Socket.io. And then it just tells me if I need to update it. And that helps a lot. And it's actually pretty easy to do. You can just ask it. Just do a scheduled task once a week. Check all packages for security issues. That works very well. All right, folks. I think time to move on to the lack of model news. This is maybe the biggest news in AI, like big labs as well. Stripe is becoming one of the bigger labs as well. Stripe acquires Open Router for reportedly over $8 billion in stock, mostly stock. 1.3 billion previous valuation in May of 2025 of Open Router last raise and over $8 billion now. So that's a huge jump. 6X valuation jump. Shout out. And congrats to Open Router folks for working really hard. Folks, I just want to talk a little bit about folks who are saying, hey, Open Router has the mode, et cetera. Open Router has the pulse of what people actually want to use and what they're using. And specifically, I really wanted to talk about this. We talked about Stripe and Stripe Wallet and last Stripe sessions. They have a bunch of stuff where they understand the agentic internet is coming. The agentic economy is coming. Stripe wants to be part of the agentic economy and not only human economy. And so Stripe innovated with a bunch of stuff. Billing for streaming tokens versus like doing whatever calculation most of people do that they give you. Stripe innovated there. And then Open Router grows in an incredibly rate, like 9% weekly token growth. I think they're like 88 trillion or something tokens, like a crazy amount. 4 million users use Open Router globally. So this is like a huge thing. I think it makes perfect sense for Stripe to come in here. And the CEO for Stripe, Patrick Carlson, says every business will have to manage both revenue flows and token flows. And that is absolutely correct. Given the spend that we all do on tokens and our revenue departments that we work at, they also need to manage that now because at many places, this now matches the employee kind of salaries and maybe will outgrow employee salaries because one employee now manages like 10 agents, 20, hundreds, et cetera. So this is, this makes perfect sense to me. So shout out to Open Router being the goat, like Alex Atala and Toven and a bunch of other folks. But besides this, bringing us something that we all need, like we all use Open Router. It's very easy. They have all the models, they work with all the big labs and they have features where if the API is overloaded, they fail over to other ones. They also host us from CoreWeave Inference. Like we provide parts of inference for open source via Open Router. So folks can get exposed to us as well. I think there are two things about this where their mode is basically. They are the number one. Everybody knows them. Do you see how many tokens they generate or pass through? So they are easy to use when a new model comes out and I want to get a quick vibe check. It's my first stop because they have it immediately. I can choose which sub provider to use and so on. So it's easy to use. So it has everything available. That is a big plus. And the other thing is the router. I think this will become ever more important now that the cost is increasing. And like you mentioned, everybody is looking at the cost and a lot is subsidized. We have our subscriptions, but if that fails and we have the limit, then a router is really important to a good router, which nobody has created yet in a good way that I just give it a task. And I don't care what model it is. It will pick the right model for the task. Whoever manages to do that first, that will be a big unlock, I think as well. Speaking of router, there's another company that tries to become the agentic internet and this is ramp and router.com. Ramp folks bought router.com and this is now a AI model router by ramp that they claim cuts your costs by 40% in seconds and scale to trillions of tokens. This is new. I don't think they've announced too much, but essentially they're saying, hey, we're going to reduce your tokens by 40%. The thing with this is that I haven't seen many companies use this that much. I think it's brand new. We may hear about this more, but also ramp for folks who are listening who are not part of the ES or they haven't worked in enterprise ramp is the company that provides corporate credit cards. And they know based on that, how many people buy and pay for what they know to an extent, right? Not every big company pays for infants with credit cards. Some people pay with invoices. So they don't have the whole overview, but they are saying that these routers cost their costs internally by 30%, which is quite impressive. So open router.com and router.com are in the news this week. And yeah, let's talk about the open router a little bit more because I think that what people are missing is that it's not only token providing. They provide very smart services for folks who are like, which apps are using this? We saw the rise of open router and then the rise of Hermes through open router usage apps, a hundred percent. But yeah, Peter, we want to hear from you as well. Yeah. I think should be positive for acquisition because I think the problem with open router was that they're kind of twofold, which kind of comes back to the same point is that they're not really a mature enterprise company, right? We love them. And they're very good to us because we can put a credit card and access anything. It's amazing. You click and you get an API key. There's not 17 forms. Yeah. Yeah. They're so good. For me as a personal developer, that's absolutely amazing. I use them all the time, right? I put my credit card and the business models, they charge a bit of fee on top. And then they let me use the models. And which is great. The problem is that no company would use that. If you have worked for corporate, it's just not going to fly at all, right? There's not enough guardrails for me to not accidentally send my data somewhere. And the downside, the flip side of having all of the providers for everything is that structurally you don't know where the data is going. Right? It's like, no, your data is there. Right? So it's just, it's not going to work for like a professional enterprise. So hopefully they're going to build it out to be more enterprise ready. Then it will be amazing. Right? Hopefully they're not going to lose the advantages that we love them for. But I think that's the path, right? Otherwise they'll be just stuck with us putting $100 each time. It's hard to build like a big business out of that. I think that we should watch out for the Stripe streaming payments thing that they innovated the Stripe sessions together with OpenRouter. I think there's going to be like a big thing coming out of there. Not to mention the connections OpenRouter already has with a bunch of enterprises on the other side. Enterprises providing them inference services as well. So shout out to OpenRouter. Like we'll hear more about this. The thing that's most important is OpenRouter keeps operating under its own brand. It's not like a brand acquisition thing. The team is staying on the whole team as far as I saw, which is also great for sales. And then Stripe is going to bring that enterprise-y experience because many enterprises, if not all, they use Stripe for various reasons. So that's great to see. Also, a huge reminder, Stripe innovated with the Stripe wallet thing, with LinkWallet. We told you guys about this where my agent bought me an anniversary. Sorry, not anniversary, but a wedding gift. And all I had to do is approve a purchase. And I think that is very important as well in the world of like agentic things. So I'm absolutely looking forward to see how OpenRouter improves for all of us. Folks, let's talk about open source and the only model this week that we had because it's not a lot. It's very exciting. Let me switch. I want to use one transition. Open source AI. Let's get it started. We are at the open source corner here in Thursday Eye and Chill Week. I will say, QN dropped the QN 3.8 27 billion parameters. We talked to you about this a little bit on the show last week, but I don't know. Tons of people waited for this model for local AI inference specifically. Why? Because all of the other open weights, open source models we talked to you about, it's impossible to run them unless you have huge machines. Especially now, lately, the Kimi and the QN, they moved over 2 trillion parameters, 2.6, I think, and 2.8. They're nearly 3 trillion parameters. That's not something that people run at their home. And even if they do, the electricity bill does not make sense, I think, compared to what you get with GPT 5.6 Luna, which is essentially free now because OpenAI really wants you to use their models. But for local AI people, it's very exciting when smaller models run because Peter can run this on his 96 gigabyte Linux machine if there's a GPU connected to it. Wolfram can run this on to 3080s, 3090s with good token streaming. If the model is good enough, there's a lot of stuff that it can do. And so I think for that reason, folks got super excited about QN 3.8, 27B. With that said, though, folks, from your experience, have you tried 27 billion parameter QN 3.8? And what are you seeing on your timelines? What are you seeing of people receiving that? Listen, let me quickly just put together the Unsloth desktop, the new app they made, LM Studio Alternative, basically. When I saw it available in there, I immediately downloaded it. They even have a one-bit quant that only takes 8 gigabyte of RAM, VRAM. So it can run on the smallest machines and is still strong. In my own benchmarks, I really have to look into this and do more benchmarks because I did only one run. But it put it above Kimi K 2.6, which is a huge model. And the performance was amazing. So this is my recommendation right now. If you want to run local AI on a normal system, I would pick this. In German, it's usually it's not the best. So in other languages, probably also not the best. But I think I will be looking into more routing stuff, like routing some tests or using some sub-agents to run this locally. Yeah, it's definitely Sonnet level at home, I would say, from what I've seen so far. Nisa, what about you? You've been running this model a little bit? I've been benchmarking it all week and using it. Okay, so the commuter response was pretty good at first, but now it's gone crazy. There are 152 fine tunes off this. Some interesting things that I noticed while benchmarking is that compared to the older Q1 3.6, it actually dropped a little bit in the other benchmarks, like medical or human eval, a bunch of the regular ones. But in agentic ability, this is on another level. I don't think it's too much bench-maxed in this. As I've been running it at home, and even friends with an M4 MacBook with 24 gigs of RAM are running it. MacBook Pro, M4 Pro. They're still getting like 12 tokens per second at 4-bit, and they're just running it in LM Studio. I have to say the agentic ability of this is actually crazy. This thing will just about do anything, and it will also control Claude and Sol in other terminal sessions whenever it needs something smarter, and it is able to do that. Yeah, this is probably, and especially some of the spicier fine tunes, like I needed to look through to make like another Linux security patch. I was able to use one of the unrestricted ones to talk to Claude whenever it needs to, and then keep working through the problem. And it was actually able to keep working through the problem. So to me, this release is pretty crazy. At first, I was not as excited because I noticed a drop in the regular benchmarks. But then when I actually tried it to just run it at home, this is completely nuts. I want to talk about the fact that it's great for local staff, but also for the fine-tuning community and the unlocking community. We haven't heard this word fine-tuning because, again, nobody's fine-tuning Kimi K3 2.6 300 parameters. It makes no sense besides the Kimi team in RL. And I think that's a great thing. Also, I want to shout out again local.ai folks. This is their new thing. If you want to know which is the highest intelligence model that can run on your M4 Mac Mini or M4 Max laptop, a Kwen 3.8 27B is very close up there in terms of the top intelligence. And I want to shout out the one person that says, Oh my God, it's so unsafe and so scary. A Kwen 27B parameter unlocked can look up torrents. It's just not everything that's written on the internet should be taken with reverence and truth. That person just went all out. And the example he showed of how scary and unaligned this model is that he asked it for a torrent and it looked up a torrent, whereas Codex refuses to look up torrents for you. I was like, bro, have you heard of Google? What are you talking about? I was refining. Peter, how is this model getting received in Arena at all? And is it performing well? What's your thoughts on this model specifically after around the week of it being out? So we have it in testing. We still need to validate some results. So I don't want to say where it's ranking, but it's looking pretty good. And I think there was a moment, I don't know if you remember this. It feels like a year ago, there was just a time when we just didn't have any models around that size. It felt like we had a bunch of good models around that size. And then there was like maybe Gemma was still appearing once in a while. And then it felt like there were just not particularly any good models. I think Mistral was like not releasing anything around that size. So it felt a bit weird because I think people kind of set up their hardware to operate around that sort of level. And when we started getting the 700B models, it's okay, thanks. That's no help at all. And there were quantizing and so on. So it's kind of good to see. And yeah, I don't have that in-depth view yet. But yeah, it looks to be like a little bit potentially special. I don't want to say too much just because we want to validate it first. But I think it's looking quite good. And I think it's interesting. The thing about the smaller models that never sat right with me is that I can see using a small model locally if it was maybe dumb but very reliable. So if it was like if I'm telling you to run a terminal command, like it will do it. And like it knows them and can run them. If we just did that, then it's absolutely like it's actually really useful. But I think the trend we had so far with the models is that they were just unreliable. And okay, maybe they're scoring more on these benchmarks. But you just why would you use them if they're not reliable? Then I think you gravitate towards more reliable models. So I think if we do get to the point that their genetic capability is that good, as Nistham was saying, then it's like almost I don't really care about other benchmarks. Like as long as it actually has this core of it actually doing the job that you're asking it to do, then that's awesome. There's also the thing where somebody calls me out every time where I post about my fable not working well. He's like, bro, local models don't get degraded. Unless you switch the models yourself. If you're on the model and in a year you're on the same model, those are the same exact weights unless you like change your inference stack, etc. So there's also a benefit there. Yeah, we're covering QN27B, the frontier on your desk. Have you been running this at all? Since the moment it was out, I totally completely agree with Twitter. But I think it's quite reliable in my experience. QN is quite reliable. It's not the smartest for sure. It's not Claude. It's not a GVT 5.6. It's only 27B. Yeah, but I don't know. It can drive a computer pretty well. I mean, unless it's really pushy to do like crazy stuff on the browser maybe. Like, I don't know. It can drive a computer pretty well in my experience. Don't you think, guys? I have. Now that you're talking about this, you're saying it's not Claude. What immediately popped into my head is that it's not Claude fable and maybe not Claude Opus 4.8. There's no 4.9. Opus 4.8 and 5. But it could be Claude 3. Like, this 27B model could be Claude 3, the model that we got all excited about two years ago. I don't remember when Claude 3 came out. And I was like, hey, should I create Time Machine Bench? Should I benchmark new open source models versus the models that we talked about a year ago and seeing, like, where we are in terms of performance? I think that'd be cool because, like, we are all getting hedonistically adapted. If you guys are familiar with this concept, like, we all get new things. And then we get, oh, fable 5 is not answering my emails correctly. And this model is, like, insane. But, yeah, you triggered me when you said it's not Claude because Claude has been, like, a thing for a while now. And this model that can fully run locally and be for me and be trained and fine-tuned with my stuff and locally execute what I need can absolutely be the Claude a year ago, which is still incredible to us. Like, the intelligent jumps that we get is incredible to us. I need to go on working. I just want to react to something that you said. Look, Time Machine Bench is a good name. But I just want to say that it's not only that, of course, models get better. Yeah, 100%. Everyone understands that. But you also sometimes lose stuff. I think that Claude 3, yeah, it has no way, shape, or form has any capability of Claude 4, 4.1, 4.5, 4.6. All of them. Absolutely not. But I think that Claude 3, maybe 3.5, 3.7, they did have something that is quite, that we are quite losing today. And it is very, in my opinion, I have a pinned, isolated environment with the Claude 4.1 code, like, from a year ago. And 4.6, and 4.6, and I run it sometimes. And it's completely different. It's not that I got adapted to the new stuff, and now I think everything is degraded, is that they are quite different. Things are moving, not all axes are moving into the improvement direction, okay, at the same time. And you really see it exactly like people are saying today. Fable is absolutely capable, but, bro, it's not easy to understand what Fable wants from me when we speak from time to time. And I never had this problem, let's say, with the 4, even 4.6 or 4.8 also speaks really nice. That's my take on that. Yeah. At the end of the day, it's not your way, it's not your model. If you run QN, it's never going to change. You can also customize it. It's brilliant. Thank you very much. For the quantum. Nothing more to say. That's brilliant model. We have folks commenting, and folks, we appreciate comments, and we would put you up on stage if you comment to us as well, saying that in my use, it seems close to the quad, Opus 4.5, which is incredible. Just for a 27-bit parameter model. Way smaller contact window. Yeah, that's true. But that depends on how much you can hold it. There is a link here from Nistan. Nistan, you want to say to? I just, there have been 10, close to 10 million downloads of the quads. There are 650 quantizations. And there were issues with the thinking levels of the original just being a little bit too long. People quickly fix them. And this just crossed the threshold where it can drive other agents for me when it needs to. And it has very good visual ability to read stuff. But this crossed the threshold where it is a local model. I can just have it run at home. I can have a look at the screen. And I can have it drive other agents only when it needs to. And that's a big, it just, that's a big threshold to cross for me in this case. And yeah, yeah, we're gonna, I have a bad feeling they might just ban it because it's just too good. But I'm just gonna leave it at that as a prediction for the future. Folks, we have breaking news. And the breakiest of news, the founder of the company Broke the News is joining us because I just reached out to them. Let's go. AI Breaking News. Coming at you only on Thursday. This is the best kind of breaking news where I reach out to the person. It's like, hey, you want to hop on? It's like, yeah, I have a very important meeting in 30 minutes. But I have some time now. Jeff Huber, founder of Chroma, the Gentic Memory and Search. And I don't know how to describe it. You guys just launched something. I would love to hear directly from the proverbial horse's mouth what you guys just launched. Please tell us what you guys just announced. You actually did do something about it because three or four months ago, I interviewed you about your agent sessions. I said, hey, Alex, do you have valuable agent sessions that have a lot of useful memories in them? Do you think, are you doing anything with those? And it was super helpful. So thank you. I was reviewing those notes actually this morning. Yeah, I think broadly we think that memory is the largest unsolved problem in AI. And there's probably a lot of really advanced approaches to continual learning. We're super excited about those as well. But at a base case, agents need the ability to write things down in a highly organized and efficient manner and a consistent manner. And then read that stuff back later. And once you dig into that problem, there's all kinds of things that you end up wanting around versioning, access control, lineage, concurrency control, and more. And fundamentally at Chroma, what we want to do is solve these base infrastructure layer data problems around AI. We've been doing that now for three or four years. I sometimes call Chroma the context company of California. And saw a lot of our customers struggling to build these sort of memory level primitives. And so wanted to offer something. This, what you see here on the website is our research preview of this technology. It's useful if you're a solo developer, you're a team, you want to build kind of a shared memory system for yourself and for your team. We'll also be offering foundation as infrastructure. So you can build it into any agent you have in a single prompt. And the reason why I think it's cool is because, dude, I tried all of the other solutions. I've tried building one myself. Every person that I talk to that runs agents, like ending up building some sort of a solution for themselves and their agents, especially as their agents like are, as they try multiple ones. And this is the same problem I run into all the time. I run things through multiple agents. I run OpenPlace and Hermesys and now Grogbots, etc. And there's like stuff in Codex and Codex on a different Mac, etc. Like all of this gets really annoying really fast when I talk to one and then I completely forget who I talk to and talk to another. And I expect the same kind of outcome, same results. Is this some of the stuff that you are trying, you guys trying to fix this with like real-time things? Yeah, exactly. And on day one, if you scroll down, we're ingesting sources natively from Codex, Cloud Code, Cursor, and Slack. We've got coming soon connectors for Notion, GitHub, Google Drive, Granola, and more. And then in a single click on the onboarding, you've wired up that knowledge base all the way through back to your coding agents. We actually don't do MCP by default for the coding agents because we find that they use CLI tools and hooks more often than they use MCP. And so the more reliable way for them to get to use the tool at the right time, obviously. But we do have MCP support. So you can plug in foundation anywhere you have an agent. And that's where the goal is to be that like neutral, bring your own harness memory layer that allows, again, teams, individuals, developers, but also your own product in the future to build systems that get better over time. And that just learn how to learn how to do their job. First of all, we recorded a couple episodes live from your offices. So shout out and huge thank you guys for hosting us as well on Thursday. And I remember our conversations there at Wolfram because Wolfram, you were also there, I believe. When I needed to ping or to test like an idea or I remember, I think it was like BM25 and Toby's QMD, et cetera. Like you are the guys who I talked to when I need to know what is the state of the art in terms of agentic search, et cetera. Have you guys worked some of that into this foundation? Tell me about like the behind the scenes, like how does this work? Yeah. Yeah, I think the key bottleneck that we identified for building kind of these LLM self-improving wikis fundamentally is you have to be very good at agentic search. That is obvious on the query path or on the read path being good at agentic search, but it's actually equally, if not more important on the right path. What happens to a lot of these LLM wikis today is they quickly become a huge pile of slop. You're not updating information that you need to be updating. You're not deleting information you need to be deleting. You're not putting information in the right spot where it can be found later. And so actually agentic search is even more important on the right path. So, you know, we think of ourselves as Chroma as probably the world's experts on agentic search. And if that is a key bottleneck for building very good self-improving wikis, then we thought that you would be a good candidate to solve that problem best in class in the world. The second thing that I'll say is it improves both its knowledge. It also improves its own system prompt. So each foundation actually manages its own system prompt. And when you give it natural language, high level feedback, it can improve its own instruction set at the system prompt level as well. It can improve instruction set on system level. Almost like part of the engineering, right? Also has the ability to update and edit its own system prompt. So it is improving itself from its experience of, so for example, our system prompt for our team, it's decided to associate all these different Slack channels and the Slack IDs. So it doesn't have to look them up again. And you realize, oh, this is helpful. I should write this down. But then also, as you give it feedback, either in Slack or through other, any feedback mechanism back to the system as a human, the whole system learns and reacts to your human feedback as well. So it both learns and also meta learns is what I'm trying to say. Oh, that's dope. I don't think we've talked about Context 1 on the show. Could you talk to us a little bit about that? Yeah. In spring, we released a state-of-the-art agentic search model. It's a 20D state-of-the-art on agentic search. So it's trained to know how long to search, how where to search, and it does better than Frontier models, but it does so at 25x the cost reduction and 10x the speed improvement. So it runs somewhere around 400 tokens per second, and it's 25 times cheaper than Opus. That's incredible. I believe you mentioned it, but definitely didn't have you to talk to us about this. Yeah. So I have one, maybe one last question for you. Tell us about this. Yeah. What is the operating procedure here? Like, how do people use this? Can they host completely their own wiki, completely on their own, like many open source community folks like? Or is this like a new entry into the cloud services area from Chroma, the company? Yeah, it's a great question. Out of the box, it's end-to-end kind of this, we put all the pipes together for you to get a great out-of-the-box experience of this research preview of this more fundamental infrastructure technology. We haven't yet done the work to open source every piece of this. Frankly, it's a lot of pipes, and so it's not that fancy, if I'm honest. I think that when we do have the infrastructure components here that go into this, there's like an agentic harness, there's some unique models. There's a unique database. Again, as I said in the launch sort of Twitter thread, we cannot build this without ChromaDB. ChromaDB is the only database on earth that can support this workload sheet, that can support this use case in a sane way. And so, yeah, we're going to continue to open source more and more. We always get a huge open source believers from day one. If you go dig around our GitHub, you will not yet find the Mac app that you download open source just because we've been so busy just getting this thing set up. So I'm really looking forward to testing this out because I do have this problem. I think many other people have this problem. And also, I'm looking forward to maybe Grog Bot is taking over the airwaves recently as much as like other folks. I'm really looking forward to see if this will be... Let's see what happens. Yeah. I love that the concept of this is like build as infrastructure. So we'll definitely keep an eye on this. Congrats on the release. I know you have to go. But thank you so much for jumping on, Jeff. And we'll tell you about feedback if we see it as well for folks. Thank you. Tell me everything you love. Jeff, you were founder of Chroma. And we'll see you guys when you continue building on this. All right, folks, this has been breaking news. But before this, we need to talk about GLM 5.3 real quick. The folks from ZEI released the GLM 5.3. Anybody really try it? By the way, it's not open source. The weights have not been open weighted, but they did drop some exciting updates. Anybody has any info or have tried ZEI's GLM or not yet? I think we're all waiting for it to be maybe released, maybe reposted. I think that's everyone's preference at the moment, but I'm sure it's pretty good. I think that the weights is like the important part, but let's at least take a look at what we're about to expect with GLM because I think it's super cool. They reached out to me, but then I was like, oh, but the weights are not there. Two days ago, the API is now live. I will say this has been a shift in how the previously open weights, but now leading Chinese companies that have open weights are operating. And we have been noticing on the show go a long way from, hey, here's the torrent link of our weights from Mistral, which is the best move in the world ever towards, hey, we're going to release this in API. I think it's important because people sending them feedback, et cetera. Then we're going to release a weights model, but we're going to release this with this clause. It's not MIT anymore. It's not open source. And it's been sad to see all these like model providers moving away from full open source person, person, person community, but still benefiting from the excitement that they got initially from this community by releasing open, completely open weights in great license. So I just want to say like credit where credit is like they do release the weights to this point, maybe in a slight delay, but they do release the weights. I think, okay, a little bit late, but. But I feel like this is a specific move towards we're becoming a frontier lab and weights, weights of the bigger mouse are not going to get released. Licensing is choking it, the price floors and everything. It's been at that point, I'm sitting there wondering why would I send my data to their API or why would I even try an open weights model? It's anyway, it cannot be hosted on my laptop and it runs in the cloud somewhere. So why wouldn't I not use Terra or Sol? What is the incentive there? If the weights are not, I don't own them. And the licensing is such that they have to impose their specific things on different companies. So it's what I'm trying to say. The open weights Apache 2 license that shout out to DeepSix still does is what I consider open source, open weights. When they release a technical paper that describes exactly how their attention mechanism works and that helps everyone, that's incredible. And that's our conversation that we had with Eli Bakic. So I think that's the only thing that I can see. The GLM 5.3 max is cheaper based on same intelligence than K3 and 5.6 L. Sorry. Any last comments on DJ? Yeah. I was going to say, I think if it really is at the level of roughly Kimi K3 as artificial analysis seems to indicate, then I think that does end up being really impressive from the standpoint that it's what? It's about 700 billion parameters. And the Kimi K3 is it's about, yeah. So Kimi K3 is about triple the total parameters and about triple the active parameters. So yeah, I'd say this is a really good improvement here. So I have found the evals real quick and yeah, the evals look quite insane in terms of just the jumps. So terminal bench jump from 4.6 to 28.3 on terminal bench three, which is quite the jump down deep suite. There's almost a 20% jump. This seems like a very, very impressive jump for the same infrastructure and the same price as well. So we will wait until this releases in, in, in weights to play with this. All right, folks, I think it's time for us to move on a little bit late, but I would love to welcome to the show. Francesco Bonacci. Welcome Francesco. Hey there folks. Hey, welcome to the show. Francesco, I've been in conversations with you for quite a while now, as we need to talk about some stuff that you guys launched and specifically I mentioned try Kua the startup back when open AI released. Computer use and computer use specifically background computer use, which got me super excited where I can sit here and talk to you guys. Meanwhile, my agent can work independently of me on a window in, on my Mac. And then you guys shipped it with try Kua in open source. I specifically remember somebody has to do this work in open source because I think it's very important that all agents can do this and not only proprietary stuff. Obviously openly, I bought the software Inc company that had very good Mac engineers, and this is the result of theirs. And then this week you guys also released something also kind of in breaking news. So I would love to hear from you directly. First of all, who you are and what try Kua does. And secondly, what do you guys release this week? And let's talk about it. Yeah. So happy to take it from there. I'm Jessica Bonacci here. Nice to meet you again. I'm the founder of Kua. Kua. Before jumping on, on board with Kua, I used to work at Microsoft for four or five years. And then I started my own company. That was like about one year ago. It was very deep in the space of computer use. Even way earlier, I think we came up with the terms 2023, four. Yeah. What does Kua stand for? Kua stands for computer using agents. That's as simple as that. And we were like in a space as I like to phrase and like to give a term to everything. Like when it comes like to Kua, I like to define two different phases. Like Kua is like Kua 1.0 where you have the agent taking control over a desktop, not necessarily like in the background, but more like now, in a headless kind of way, you have a server where a call driver component and then take it from there. You can expose it as an MCP, as an NPI, an API, a CLI, whatever. And then we were doing our own experiment last year. Okay. This probably is not the best paradigm for human computer interaction. Like I need to dedicate an entire sandbox for one only agent. It's quite an optimal in there. So we thought about it for a while. We wanted to ship something like very similar to the agent dash browser CLI that sell as. So again, I'm a big fan of Peter. I'm very bullish on CLI data from open cloud. So we wanted to ship something like that. So we are like, like the primitives, like already in place. And then Ari from, from sky slash like open and open AI now. I'm also like a big fan of them and their work. So they prove what we were also like working in a background that the ground computer use was possible. Starting from. I ask you how the hell does background computer use working in macOS and how Mac, which is the more closed operating system out of the three is the first one that background computer use works in versus the other ones that like, it still doesn't just tell me this is crazy. Yeah. There is a lot of wizardly in place in there. And by wizardly, I mean that. So the secret sauce for like having something being controlled in the background was already there. And there is something like called accessibility tree. It's basically a three representation, kind of like HTML, a dom like for any native application. And Apple throughout the years, they actually invested a lot in accessibility, just like for making like application, like more accessible to the user, especially like catalyst native apps from macOS they have like very rich is like application, like notes calculator and so other like system apps. So the whole operating system is like very. Quirible and it can be navigated by an agent. But when it comes to application that does don't necessarily expose an accessibility tree. So we refer as those as application being, I think a sparse tree look at, I don't know, Slack for instance, that's an electron app. Look at blender. If you try and query the accessibility tree on this kind of application, you will see that you can only query the outer shell of the app. And then like everything that is inside is actually. It's a website. It's a camba. It's a camba that it's actually not querible. So the only way there to actually do something is relying on vision and multimodal models. They've been like very good at grounding and by grounding ability of a model to figure out like where it should click pixel wise. So the trick there is actually for Mac OS and just cut short there is to rely on a primer click like tricky the window into thinking that is not this an active window while it's still it's not. So the actual application we realize, okay, there is a synthetic event and not necessarily I have to still focus on the current app. Dude, focus stealing, I think is the most annoying point. And we got to call out the big labs as well and their computer use on your website, which I also would love to talk about. There is a leaderboard on who runs computer use, you call it QuaBench leaderboard, right? If we show this right now on stage, folks will see the club Fable 5 has an asterisk, but this is the full pass. You have a very difficult benchmark. Looks like six out of 25 top things Fable 5 does. I have to talk to you about this, but you have this thing. However, their harness though is awful and cloud takes over my whole computer where it needs to click buttons. I do not appreciate this because I also need to work. And we talked at the beginning of the show with Peter that he built a whole Linux box for the computers to work. And recently I got excited. I have to talk to you about this as well. Dude, I have so many questions about Grogbot that everybody's getting excited about that they have their own computers, which absolutely sucks. I send their team back feedback where I see, I logged into the computer. I see them typing every word by every letter, just by typing. It's like, oh no. So what it does is it types a letter, takes a screenshot, sees, oh yeah, I typed a letter. I need to do the next letter. It's like, the fuck is that? So I have a lot of stuff about computer use that I personally don't like, but yeah, the harness of cloud takes over. Background is the way. Talk to me about the state of the art in computer use. You guys have a leaderboard. Other companies are releasing OS world and other leaderboards. We always look and pay attention to them. Where are we now? Can agents use computers fairly easily? Are there still major things that they cannot do from your experience? What is the gap that needs to be closed for that full, complete, automated user in an environment that does clicks for me? I've been in the space for about three years now. I do remember the very first, we were calling it in China and Microsoft like Navy agent, basically using GPT-4B at the time. And GPT-4B wasn't even good at grounding. So you had to rely on accessibility trees all the way. So the first results on OS word and Windows agent arena, that's the counterpart for, for Windows devices was like 40%. And that was like three years ago. And it's been mind blowing, especially in the last year, like seeing three shows, like if the human baseline for, for using a computer is 72% on OS word, V1, we are actually now at about 80%. So OS word is like very well overfitted by now. It's been like the, the fact of benchmark up until now to, that's being like used by labs. So we do have our own benchmark anyway. We work also with labs and we made like these benchmarks specifically focused on key cut because that's one of the application that is like more. Like one of the toughest application. Really? Computer using agents, because you have to rely on not necessarily on like clicking, typing, scrolling, that kind of, that, that is solved. That's my answer to you. Like clicking, we're solving. Okay. But when it comes to multi, multi pointer and like events, like dragging, even scrolling, like it's not quite there sometimes because every single operating system has a different way of scrolling, like natural scrolling from Mac OS and windows that that still confuse the agent. We're not quite there. And you can see it from our benchmark. Like when it comes to solve the rate, you still, you still have Fable, like only solving six of our 25 task. There is still like room for improvement. Probably we'll see like frontier models, like reaching 10 to 12 soft tasks on our benchmark in the next couple of months. Luke Gromenonis We're definitely looking forward for agents doing meaningful work completely. And that includes using a computer. Luke Gromenonis You saw the release of Grokbot as much as we all saw the release of Grokbot. This is the, the highlight there is agents that have their own environment, their own computer. Luke Gromenonis All they have in their computer is a terminal browser and the file system, which I absolutely love. Luke Gromenonis As we said, they need to get better at computers. Maybe I should reach out to those folks and connect you somehow. So they learn from you guys. Luke Gromenonis That would be like incredible. But this week you guys released something as well. Luke Gromenonis And I really want you to present this. Can you talk about the releases this week from you guys? Yeah, totally. So we ship something called computer history. It's like our open source take on Codex computer history. Luke Gromenonis Which is also fairly new, right? The Codex had Chronicle in beta for a while, which is the thing that takes screenshots of your Mac to know the context of what everything you worked for. And then there is computer history. Would love for you to cover both. And then your take on this in terms of open sourcing. Luke Gromenonis So we have been debating for a while whether to release memory for computer using agents or we do have our own trajectory recording mechanism. Think about human that wants to do a mundane task or knowledge worker. Luke Gromenonis There are ways with Quadriver that you can just shoot in and then turn it into a skill and play it whenever you want. It's basically like, you can think about it like a bank of successful trajectories. That's living on an encrypted key store on your device. And we release it for all the three major operating systems. So when we make like a major release like this one, we just take the lesson learned. Like we have a cross-platform harness, trust-based. So we just release all the three major platforms. It also helps you think like broader, okay, is it even like possible overall on Windows or Linux? So again, it's like a bank for successful trajectories. Like the closest, like similar. The very first take on these was like maybe Microsoft, like three years ago, they tried to release something similar. Probably you recall like Windows recall. Luke Gromenonis I remember I was talking about this. I think Nistin and I talked about this, that they recalled Windows recall. Luke Gromenonis Yeah. Because it was like not opted out. It was like, yeah, it was Yeah. I was at Microsoft back in those times. So that wasn't like very well perceived by the community. So again, we were debating for a while whether to release something similar. But the difference to recall is that we don't record like any screenshots. Like we release this feature like with privacy in mind. We only record successful trajectories and accessibility trees. We don't record like any written text. So it's different than like memories because like we're just like recording what Quadriver knows, not necessarily what like an agent should know. And we don't like it. We didn't do any benchmark runs. I've been dog folding like this feature for the last couple of weeks now. And it's been like mind blowing, honestly on especially like on knowledge worker task or like even me doing. Luke Gromenonis So if I understand this, I will relate this to the audience as well. You have an update to Quadriver, which is the open source driver that like clicks things, especially in background. That stores the successful ones. Luke Gromenonis Luke Gromenonis So if you're like, you're like, you're like, you're like, you're like, you're like, storing those trajectories. Yeah, like a shortest path problem using calculator for doing something within an app or just I think one of the demos that we released was like this one specifically was like about paint. You're looking at query in the whole state of an application and just understanding how the accessibility tree looks like where pixel wise a particular component is. Luke Gromenonis So you're probably if you're like doing over and over again, like the same task, you will find the shortest path very early on. We call history on versus history off. It's like trivia implementation wise. We give the agents an ability and to primitives to access and query these, this key store. And by doing that, it's as simple as that. We talked about open air hacking, hugging face. And one of the outcomes there when they released it was like, hey, these agents found out they can collaborate. And by collaborating, they effectively created this memory. And this memory caused the thing where like when one agent opened the door, the others didn't have to open this door. They opened the next ones. It's kind of, this reminds me of that. This is essentially like if one agent already knew where to click inside paint. So the other agents rely on memory versus going and trying it and going and trying it is like 17 tool calls versus the one. Right. Is that the idea? Or even really, how do I even launch an app? There are a bunch of client identifier. Maybe I want to launch Slack web or Slack desktop. Okay. Just save this successful trajectory. And I'm just gonna take it from there. This is awesome. Congrats on the launch. I think it's very impressive that you guys fully open source and tools can use your tooling. I really appreciate it. Every time somebody who comes and does open source, we have applause. So yeah, all of us are planning to you. We have to continue to our next on machine very soon. And also we know you're a busy dude, but feel free to come back to us when you guys release some stuff. So Francesco, thank you so much. Nisten has one question and then we're gonna wrap it up and move on. Nisten, go ahead. Can I use this with any type of Linux? If you have some stuff that's very sandbox or very secure, or is there a particular desktop environment that it needs or what advice do you have? What works best, especially just for Linux desktops? What creates the best computer use environment? That's a great question. Nisten, the most mature background computer use on Linux at this stage is still X11 over Wayland. We've been like many ask, okay, you guys, we should support Wayland. The story there is that there's not even a queryable accessibility tree and there is no coordinate system on Wayland. We're thinking about actually starting an FPC on the next version of Wayland and see, okay, this operating system actually needs to be like more like, needs to account like some low level primitives just to make like the ground computer use possible. So probably we'll get there eventually. Yeah. Have you? X11. That's good. That's what I use. Actually use XLiberGL and wrote my own. Okay. And then either XFC or known that doesn't. Yeah. So I'm going to use this. That's why it's built into the operating system. All right. Francesco Bonacci from Kua. Thank you so much for joining us. Congrats on the release of history. I can't wait to try this as well. It is already integrated into Hermes as well. Like it's coming with the latest driver, right? We're talking with the team. So it's kind of like for free. Awesome. And congrats again on releasing this in open source. Thank you so much for joining us. Thank you. All righty, folks, we move on before we have our next guest. I do want to say hi to Bin Liu from the HeyGen team. Hey, Bin, how are you? Welcome to the stage on Thursday. I thank you for hopping on. I only reached out to you yesterday, but we've been talking for a while and I really wanted to cover HyperFrance for a while. And I have to talk to you about our sponsor called CoreWeave with some biases. All right. Welcome to this week's buzz where you can see by orange jackets that we are the eye managers at this incredible company called CoreWeave recently been very much on thoughts of mind on the news with different quarter releases, etc. But we're also doing a bunch of stuff that I think that you guys should know and be excited about. Wolfram, how about you take this one? CoreWeave has partnered with Masterclass, which is an online learning platform that features individual AI tutors that help you learn. They have the real tutors of course, prominent figures like cooking with Gordon Ramsay is one of the causes, negotiation with Chris Ross, writing with Shonda Rhimes. And they have AI tutors and to evaluate how these tutors are doing and help improve them, they are using our platform and they are running the agents on CoreWeave cloud and use weights and biases Weave to evaluate monitor and improve the AI teaching agents. So basically the combination that CoreWeave offers is the hardware, the infrastructure, the software, the solutions to monitor and improve the whole thing. I will shout out to folks who work with Masterclass. Masterclass is generally incredible. I did not know about their AI agent things and AI tutoring one on one. So that's great. And they also have a Masterclass executive class basically where they are teaching about business users. And there's also a hands-on AI lab where they teach how to evaluate systems for production readiness using our tools. So the other thing, we have been talking about this on the show before and it is coming ever closer is our conference. Let me put it up as well. We passed 1 billion runs on the weights and biases. The weights and biases. A billion? A billion runs. For the past, I think nine years, Sean Lewis, the co-founder, the first run was nine years ago. And we passed 1 billion runs on the weights and biases, machine learning, experimentation, tool tracking. So shout out to the whole team for this very strong effort. Everybody who uses weights and biases and knows us because of it, it helped build a lot of models. Early successes was like OpenAI, Toyota research, Uber, like all those folks build their foundation of weights and biases, which is incredible to see just way before it joined CoreWeave. All right, let's talk about fully connected and move on. Fully connected folks. Are you going to join us in September 29th, October 1 in San Francisco? We have the code for you, but also we've announced the speaker list. Wolfram, you want to talk about the notable folks that we have on there? Yeah, sure. There are people from our company, our CEO, Michael Trader. There are our CTO, Peter Zelensky. There's the VP and GM from Nvidia, Hyperscale and HPC, Ian Buck. More people from us. Some you may even know from engineers speaking there. And speaking of AI engineer, we also use the same location that AI engineer World's Fair was using. So where is it? Tony South. Almost the same one. It's around the block, but we're going to have a live show from there, folks. And if you want to join us and see the recording, see Dr. Feifei Li, you can see that these prices are quite high. But if you are a Thursday, I live show tuning person, use this code. Use this code. I'm not going to say it out loud, but people have to go hunt for this, but use this code to get in and secure your spot. Again, this is very adjacent to OpenAI Dev Day. If you're going there, please come and also check out Fully Connected. We have Dr. Feifei Li from World Labs. They're announcing soon something super incredible. And also we have Sara Gua from Conviction, one of the top VCs in AI. So that's incredible. In addition to a bunch of folks from CoreWeave, please join us on September 29th. All right, folks, this has been enough shilling. Let's move. I want to talk about Hyperframes. Welcome, Bin Liu, to the show. Please introduce yourself. I would love to hear from you who you are and what are you working on? And then we can talk about Hyperframes. Thank you, Alex, so much for having me and having us on the show. We've also absolutely enjoyed a lot of your content and very honored to be here. I'm Bin. I am a VP at HeyGen. I lead all of our agent and product teams and engineering teams here at HeyGen. I'm one of the co-creators for Hyperframes. And I work with an incredibly small but smart team that build out an open source type of frames. We've been pushing out updates and progress over the last three, four months. We obviously crossed 40,000 top stars just a few weeks ago. Yeah, it's been awesome. As for me, my background, I've always been in the creative space. I spent the first 10 years of my career at Pinterest. I led product and engineering on the content team there. After that, I started my own company, having always been building in the creative agent space. As soon as I saw ChatGPT about three or four years ago, I was like, that was the moment. And now here I am, joined HeyGen as part of our acquisition and then really excited to be building creative agents and creative agent technology here at HeyGen. So thank you so much for joining. We mentioned HeyGen a long time ago when Joshua, who I met recently as an engineer, Joshua went super stupid viral. Was like, hey, I'm Joshua. And this wasn't him. This was way before the models caught up. And this was like very exciting to see. And since then, we saw this rise of the AI influencer type thingies. And a lot of companies are trying to do the AI avatar that talks like you and voice cloning, et cetera. And it feels like HeyGen, at least for me, was like one of the places where I would go to see the frontier of like how that's happening. But then at some point I started seeing this Hyperframes thing. And for folks who are just watching, just so you'll know what we're talking about, here's a transition that I built with Hyperframes that I would like to show on stage. All our strikes here on the show were rewritten from scratch. And I used to, because I produced the show, I used to do them directly in CapCap manually by myself. And I still have the other ones. These ones were done by my agent using Hyperframe. Ben, how about you introduce Hyperframes as a product? What it is? Why does this exist? Is the most important? How does this relate to the fact that you have a company? We'd love to have the connection. And then we can talk about like how people can actually use this. I'm very excited. And wow, that was amazing, Alex. The best one is, I think the frontier is the best one. Let me show you this one. I have a bunch of strikes here for the show. And obviously I rebuilt them from scratch with Hyperframes. But yeah, tell us about Hyperframes. That's incredible. Okay. I have a bunch of demos to show as well. Yeah, please. I actually also made you guys a surprise. Oh, let's go. We'll see how you guys like it. But yeah, I wanted to actually take a quick step back so that people have some context. HeyGen as a company, we obviously started by being the one of the best, if not the best AI avatar, right? Because at the core of HeyGen's mission, we're not trying to compete with like Veil 3, Sea Dance. We're not trying to go to disrupt Hollywood drama TV. But the one core thing that we believe is that communication through video is one of the most effective ways. However, for most of the people, they don't feel comfortable sitting in front of a camera. And so we started there because we noticed that even just showing up in front of a camera for a lot of people is incredibly difficult, especially for introverts. Myself, it wasn't that easy to just show up and look at myself. And I appreciate it. Thank you for coming up. Yeah. Yeah. And that's why the company started building the AI avatar technology, because the human to human interface is important. Human needs to show up, but just it's so hard to build a camera crew. It's so hard to be eloquent in front of camera. But with our AI technology, with our AI avatars, you can show up without needing the camera. You just need to direct your own. Yeah, exactly. Your own digital avatar and they will do the talking for you. And, but if you look at these videos, just the AI avatar is not enough. Yeah. You need the motion graphics. You need the editing. You need all of those things to actually make your communication effective. But making those things is incredibly hard. You need to use things like after effects for motion graphics. You need to learn CapCut or Premiere Cuts to do that. And we- Dude, that part for me takes way much of a toll than just talking, yapping to a camera. This we can do all day. We sometimes go over three hours and just yapping. Yes. Following up, stopping at keyframes, saying, oh, I said this thing, this needs to show up, et cetera. We have a whole team now that works on our shorts. And this is the bulk of the work of the post-production thing. And when I joined HeyGen, that was the problem that I wanted to solve because I also feel very blocked whenever I need to do editing. I find Premiere Cut and even CapCut to be very hard to use for someone like me who's not trained to be a video editor. And I will skip a lot of the exploration and iteration, but we landed on why don't we turn this into a coding problem? We're all engineers. Can we make video editing a coding problem? Because LLMs, our AI agents are so good at coding. And that took us to hyperframes. So hyperframes, what it really is that it turns HTML pages into a video. Yeah. And so when I asked my AI agent to edit my video, what it really is trying to do is take a raw footage for instance, and then build everything around the video, cut the video using code. Feel free to share your screen. I would just go directly to my codex. I'll just show my whole screen. I'm sure we can do editing later. Yeah, we'll do that. Let's do it. So just say that there's two things that get unlocked via this. One of them is, first of all, the ability to do this programmatically, which is already itself crazy because we just had Francesco on with computer use. CapCut is incredibly complex and different motion graphics. Those apps are incredibly complex. So asking your agent to run stuff on them is not super easy. But second of all, just having an agent like knowing about this, I think it's great. Yeah. We just downloaded one of his clips and then I basically did editing for him. So the input video is, let me see if I have the raw input video. And also walk us through what we're seeing here, because this is also a podcast for folks who cannot watch, they need to hear what is going on with this. So this is input video. You can see that the input video here doesn't have any motion graphics, doesn't have any, but what you need when you actually publish this on YouTube, you need all of that. When you ask your AI agent to do the entire cutting, what Hyperframes can do for you is essentially adding every single element. So let's just play it right now. A good ad when it's running appropriately, I put a dollar in and $5 comes out. Every dollar of ad spend I put in, it turns into $5 of revenue or customer lifetime value. When you look at e-commerce, this is literally the game. This is how various companies build their existence. Trying to identify a winning ad. You do what I call A through Z testing. Instead of A, B, one ad versus another ad, I'm doing a hundred different ads simultaneously. And then I'm just going to look at it. So every single one of these motion graphics, the fact that this avatar is showing here, and the fact that there's text of that. When it's running appropriately. Yeah. All of these is done by Codex. Codex. And the way Codex does it is going through Hyperframes skill, one of the talking head cutting skill. And if you look at the code here, I know none of us need to understand this code, but the entire video is actually constructed through this code. But really we turned the task of editing a video into a coding task through the framework that we built, that we call Hyperframe. Which is open source, by the way, right? Folks can go and take a look as well. Hyperframes, I believe is you guys chose to put it out there and maybe about that decision. Because I found that also incredible. And everything that's open source is getting big applause on the show because we absolutely love open source and stuff. I think that the key thing here is really for this to be adopted by all the agents. Because at the end of the day, we believe that this technology is something like more of a fundamental technology. But even just allowing agents to learn that, oh, I can do video editing by doing coding. That is the fundamental piece that by open sourcing yet, we want all agents, all the frontier labs to learn and do. I think a few things stand out for me as somebody who used Hyperframes, obviously for the show. Dude, the opener for the show, the 10 minutes. It's like, dude, 10 minutes is a very long ass time to edit a video for 10 minutes. Especially with motion graphics, you have to do a lot of looping, other repeating. I don't have time to lose. So our opener here has a bunch of stuff from our show. It has a timeline. It has a bunch of things. It's like, okay, people wait until we ramp up, until we start. Might as well educate them. This is all agent, which leads me to my next question. Do you guys have any kind of benchmarks, evals, leaderboards, et cetera? How do you know and at this point being which is the best agent to create agentic videos? We'd love to hear from you directly from your experience. Yeah, we'd love to hear how you guys... I'd love to show you. And let me actually see if I can pull it up a benchmark across all of the state of the art models on how they perform, how each and every one of them perform against the hyperframes. Obviously, there's Design Arena, right? And Design Arena now does websites. They now do video models, et cetera. Yeah. This is a mix of both because it requires motion graphics and not all websites necessarily need motion graphics. So I haven't seen a very good representation of an agent understanding something. The other thing that I will tell you is agents can read HTML and they can take screenshots, but not a lot of models can view video. And MuseSpark, some people actually commented earlier on the show that MuseSpark 1.2 is really good at viewing video. And I feel like that's also missing. I almost called it out. MuseSpark 1.2 is really good for video analysis. I've been feeding your recordings for my clarity. It works great. I feel like when an agent generates a video for me, if it only takes screenshots to try to understand whether or not it did well, that's not enough. I need it to consume the video because a real world editor will look at the thing and then see, oh, the timing here is not quite right. The sound doesn't match, et cetera. Like all of these things, they can come from HTML, cannot come from screenshots. Would love to hear from you like where the state of the art of agentic video generation is. Let me, I can show one small thing. I think this is not necessarily the best one, but let me quickly show it. Yeah, please. Then I can talk a little bit about the exact question that you were getting at. So we actually test across and we're actually working closely with DeepMind. We will be releasing a benchmark very soon, but you can see that for a cheap model like GPT-3 flash, which is I think 10 cents per million tokens. When it's given an instruction to do the motion graphic, it didn't even complete it. When you give it to GPT-5.5, this is what it did, right? It looks a lot better. And we can just take a couple of look. And you can see that the difference between a cheap model and the state of the art model. So generally, the three flash on the left here barely does anything interesting. And what is it? GPT-5.5 on the right is doing zooming and animations, et cetera. And cloud opus, obviously everything looks good and looks virtually space, et cetera. And the other question that you had, which is not yet fully done through our current benchmark, but is that we are building our own harness that really solves the agent or AI not understanding video problem. Because it's actually very true and actually wrote about it in X article where VLMs or like even the best fable five, they're not trained to watch a video for an hour and be like, oh, these are the highlights. These are the things. What is this motion like moves like that's not what it is trained for. It's trained to say what's on the screen. What is this video about? It's talking about the what, not talking about the when and how. So you actually need a good harness to be built around it, to teach it the, the, the when and how, and that's actually what we are building. Even the state of the art, to be very honest, the state of the art video input that models that accept video input. None of them actually do the video understanding super well. Yeah. And so that's why we're building a harness around it. Let me show one more example. I'll just show the entire screen for simplicity. One of our posts, I did this video, but this video is a, let me just quickly play it. Oh yeah. No, I can hear it. So this video was made by asking our agent to watch one of the very viral video that this person made. It was made originally through. You can see that. But the reason why we are able to like, almost like pixel by pixel replicate this using hyperframes is that we have a harness that really understands what's going on that of here. And that actually takes me to showing you guys a little, hopefully fun surprise that I turned. I was showing this project earlier, but let's go to the other project. I turned this video. Ooh. Can you guys hear it? It looks like a Thursday I themed. That's right. That's right. Rebuild of the video that you just showed to us. Can I get that project? I'll show the project right after, but here's the rendered video. Oh, nice. Let's go. For some reason there's still no sound. I'll send over the project, but this is, you know. Oh, this is dope. This is Thursday I, but, but with this whole new motion graphics. And that's the beauty of using code as your video editor, because. I'll say like. Go ahead. What I think that you unlocked is the fact that everybody can do this now. Yeah. Whereas motion graphics is really difficult to time and it's really annoying to learn all these apps and the time timeline concept and the keyframes. There's a lot, it's really a lot of cognitive load. And now everybody can basically do this via the stuff that you worked. And so I was very happy to represent you guys here on stage and tell you that I absolutely love using hyper frames with my agents. And I am looking forward to significantly better results. And also the feedback that I have for you is that it needs to start looking less like web pages. I know, I don't know how web GL or whatever, but we need to start moving towards the actual motion graphics that, you know, Peter, how about you, you help us land this plane because we're almost at the end of our two hours. I really like this because I've been trying to do it natively. And I think the way it works is does probably crazy HTML does perform pack and like all of that stuff. And the question I have for you is like, I know people use the word taste everywhere for code. And this code essentially, right? As you say, like, how do you see it here? Is this your opinion about how it works or how do you guide it? Because honestly, my experience, I think GPT-5-6 is way better than GPT-5-5, for example, but still, still a little bit need to push it a bit. And like, how do you think about that? Is it, does it come from your side or is it temporary and models will just get better? Yeah, absolutely. So yeah, thank you. Thank you for showing that. We do have built a ton of skills teaching agents how to produce, let's call it tasteful outputs. And that's one of the reasons why we chose HTML as the core of our technology, because just like you said, the reason why Kimi 3 or GPT-5.6 or Fable, they are starting to get better and better at making these videos is that there are also people talking about, oh, they can make amazing landing pages. That's why we are trying to build and likely launch our own harness, which takes a lot of the anti-pip patterns out of agents coding. And so that a lot of the taste from your own websites, from your own product can be actually expressed. A lot of the good motion graphics practices can be also used. Our open source skills already support that. While I think that a lot of that is also in the harness while cloud code and codecs might not necessarily do a great job yet. Yeah. Ben, this has been awesome. I'm very happy that we had an opportunity to bring you on the show and talk about this. Congrats on really an amazing open source product. Obviously it feeds into HeyGen and the stuff that you need to do like motion graphics on top. But I think what it unlocked is, I don't know if you quite followed the Thursday episode where I talked about this, every night I took the pictures of that day and sent it to Claude with hyperframes to create a recap of that day. And I don't want to show this because I don't show my kids here on camera, et cetera, but I will send you directly if you want to take a look at the outcome of this. This will be in their memories forever. Every day we reviewed it afterwards. It was incredible. And I learned a lot while doing this, but also I learned that agents are not yet great at this. It needs to improve significantly. It needs to stop looking like HTML pages because there's rough edges there with SVG graphics, et cetera, because video is a very rich format. So I have tons of stuff to give you feedback about, but thank you so much for coming up. It's been a great sport and we're very excited to hear and see where this evolves. Alex, thank you. And Peter, thank you guys for having me and having us on the show. I have one last offer. I'd love to help you guys build a hyperframes brand system with us. And I would love that going forward. Hopefully a lot of your video editing and a lot of your intro clips can be made by just asking your agent to do it with hyperframes and our harness. I would love that. We have a grog bot agent that takes notes for us as well and to do it probably, and definitely sign us up for that. All right. Thank you so much for joining. Folks, we have just a little bit more to cover on the show before we wrap up. And the stuff that I do want to cover kind of relates to what we just talked about because it's not only motion graphics, it's also audio. And so when I tried to generate a bunch of videos with hyperframes from HeyGen, I had to put on some music. They have some integration with some music generations. This week, Happy Shrimp was released by Alibaba. Anybody see Happy Shrimp? Happy Shrimp is end-to-end music generation. Full song from emotion or story or prompt. Happy Shrimp 1.0. But I think that music gen... Wait, last week, Minimax released the music gen as well, right? We didn't talk about this. Yeah. Okay. Let's at least play something of an example. If anybody can pull this up faster than me, feel free. But if not, I will pull up faster. Let's see. You're saying the music... Minimax music 3. Yeah, it's just Minimax music 3. So this was like last week, a little bit after our show. So we didn't really play it for you guys. It's the... On the trending on Hugging Face this week, it's the third most trending thing. Oh yeah, I found one example. They have the really worst license of all, excluding all of America and Europe and UK and so on, but nobody can. Can you guys hear this? Yeah? That's good. Minimax music 3 on the beat and you can run it too. This is a Rob from ComfyUI, the creator of ComfyUI, I believe. Or one of the guys there that says that you can run this music generation. One box, one box. Just close your eyes and just sit there. Don't like it. Uh, DJ Niston knows the thing about music. That is really good. Uh, Minimax. So Alibaba released their one. It's called the... I love this name. Happy Shrimp. Happy Shrimp is a call out to some folks about AI and the benevolence or whatever, right? We're all like getting shrimp. I don't know where it comes from, but it's a meme. And the shrimp, shrimp rights and so on. Yeah. All right. So this is Happy Shrimp. This is very, uh, K-pop. Okay. I'm, this is what they released. Do you guys hear any of this or not at all? Yeah. Okay, cool. And then we also have Cortesia Sonic 3.6, which is now number one in both artificial analysis and other ones. I don't know, Peter, if you guys do any, please tell me if you do voice. I keep forgetting. No, we don't. No, you don't do voice. But Cortesia Sonic 3.6 now runs the top of TTS leaderboards. We had Cortesia on the show, obviously a couple of times. Great folks. I will just again, call out that the fact that this show is now getting transcribed in live session by our chief of staff Thursday live transcription. And we can see everything we've talked about and the highlights and the agent now tells me like, hey, you talked about this. You didn't talk about this, et cetera. Audio 8 TTS and S1 mini. Oh yeah. Super whisper. Shout out to Neil and Nisli, your friend. They released S1 mini first open weights model, 0.6 billion parameters. It cleans transcripts on device. Have you guys had a chance to use this or hear about the S1 mini? Yeah. I already implemented it because I have stuff on my phone that I run the agents and stuff in. And anytime you use like turbo whisperer or even parakeet, it's just gonna write like random text. But this one's only 0.6 speed. Doesn't really hurt the performance. And I find it quite usable now for just voice texting my friends. Because often it would just mess up or screw up grammar or look weird. It just makes it look nice. So it just changes the text after it comes out of whisper or parakeet. Yeah. But it does make a big difference. Whereas before I was a bit nervous to just message my friends on via, yeah, via whisper, whisper large, because I often have to edit stuff out. And this one just corrects the text. It just makes it nice. It works pretty well. I highly recommend people just, I don't know, just tell their agent to add it in. That's the easiest thing. Tell your agent to add it in. A few things we didn't cover this week and Wolfram, you sent them. So let's at least mention them. Obviously GrokBot, it feels like it's exploding. It feels, I start feeling the same tingly things where... I had people in Europe that I never thought would like Grok at all. They tried GrokBot and they are just blown the heck away. They have entire teams running stuff, doing stuff for them. They use it to manage side businesses. People love this thing. This thing works really well. There is a certain feeling of tugging that we have when something starts to feel like it's going to change things. GrokBot, to me, feels like the early days of OpenClaw where we told you about this, where before it used to be called OpenClaw. The thing just works and works right now and listens to us. And I didn't have to do a lot of stuff to configure it. It's just the same lessons that OpenClaw did. Here's what I'll say about this. In addition to having a computer, etc., we will keep telling you about new use cases as well. Obviously, there's the whole thing where Elon Musk is pushing this like crazy, but also it has real uses. I think that people need to step away from hating on Grok as a concept towards, hey, there's something there from the folks at Cursor that just by the sticker price of $60 billion that Elon Musk put on it, it has to be called Grok, but it doesn't have to be called Grok. It's a whole new wrapping and the agents are talking about each other. And so what I wanted to add to this is that my chief of staff that I built with GrokBot was listening to the show using a different bot that was listening and summarizing everything and asked it, hey, we're about to land this plane. What important things that we didn't cover yet? And said, hey, one real leftover if you have 20 seconds, cloud code slash design. Cloud code recently added slash design inside cloud code. And now the cloud design thing that we told you about is appearing inside cloud code. You can like interact with it. It also is not perfect because it's kept two things. It's so good now that other things are trying to catch up to GrokBot. Wolfram, so you brought those two. So please feel free to cover the two things that we need to mention. Wolfram, you want to cover this? Yeah, the interesting thing about this is an UI thing for the desktop app where you have the view of agents as different profiles, basically. So Hermes agent had this feature before where you could just have multiple instances running within the same gateway. And now they added the same UI so you can see not just the usual interface where you have the chat sessions and every chat is with the same agent. Now you have different agents like the chats and each agent is an individual. You can add mention it. So they added that to the Hermes desktop app now. For folks who don't know, GrokBot is the unique feature of GrokBot is that the side conversations are not sessions like most of the other apps. They're bots. Everyone is a unique bot with their system prompt, with their memory. They also have shared memory, but they have a unique memory. And that's been like crazy for some people, but some people it unlocked a bunch of use cases. So it looked like Hermes is really quickly following up on this, which is great to see. It kind of looks like that also, but because it's Hermes, you can use more than Grok 4.6. You can use other models as well. And I think this is a transition phase we are seeing. Now that the age is a cool thing, we want them in the interface, but eventually you want to have multiple sessions with each of the agents. So it needs to be something else. And I think it will be heading more towards something like Slack where you have the different channels, agents you can mention and convert them. Yep. And the other thing is bot also was created by the folks at... Pilot Kit. Yeah, Copilot Kit. Our friends at Copilot Kit, of course, I blanked out on name after three hours almost on air, also released their attempt at this pattern, but for open source, which I think is very important. Not everybody can afford the $250, $300 ultra tier for Grok to be able to use GrokBot. But here's my suggestion. If you are considering two subscriptions, Fable and OpenAI, et cetera, I would suggest that you at least try this out to see if there's something in this paradigm because that feels like the new paradigm is coming. Computer use works in their model and we need to land this plane. As you saw, my light just went out. I think we covered everything, folks. This has been a chill-ish week, but still we have breaking news. Jeff Huber from Chroma came here to talk to us about Foundation, which is their new layer of agentic memory and search. Really encourage you to try that out. We had Bin Leo from HeyGen talk about hyperframes, which is, by the way, I think that News Research does a bunch of their videos with hyperframes as well. And theirs looks very cool. And obviously we had Francesco from the TriCua team to talk to us about computer use and background computer use and what are the best agents that computer use in addition to a bunch of news from this week. So great week overall. You want to finish up? The good thing about a chill week is that it gives us time to take a little more time to look at each individual thing that interests us. And I would extend that invitation to our audience as well. Pick something we talked about that raised your interest, try it out. And at the beginning of the next episode, why don't you tell us how it worked out? That would be also interesting, what changed for your workflow. It would be very interesting to hear the same from the audience. What did you take out of the last session and have to say in the new one? Thank you so much, folks, for joining Wolfram Ravenwolf, Peter Gostov, Jan Pelek, LDJ and Nisten. Thank you so much for hosting the show, folks. Your host is Alex Volkov. And if you missed any part of the show, it's getting turned still manually. Agents are really bad at it still. But it's truly into a podcast and the newsletter at the end of each Thursday. So if you missed any part of the show or you want the links that we talked about or the links to the profiles of the people we covered, everything is in there. Please feel free and encouraged also by this request to leave us five-star reviews wherever you do listen to the podcast. If it's to subscribe and hit this bell button, it really helps. If it's Apple, five-star review is really helpful. This is the best way for you to support us across the platform. It really helps us reach new folks and teach them and bring good guests, which is, I think is one of the coolest parts of the show. We'll see you next week. What will be our last show of the summer next week, August 27th. All right, folks, thank you so much for joining and to end the stream. Here's a video of non-FaperFrames General video of previous my efforts, but as you see, there is a difference. All right, folks. Bye-bye everyone. See you next week.