← Back to search
ThursdAI - Grok 4.6 & Grok Bot at the frontier, DeepSeek v4 Pro GA, Meta opens Muse Glimmer and promises spark, Gemini gives us 3.7 flash instead of Pro & more AI news
ThursdAI - The top AI news from the past week · 2026-08-14 · 135 min
Show full episode description
Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer! My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus! I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I’ve invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price. There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license! Let’s dive in (and please don’t forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!) 0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final Takeaways Grok Bot and Grok 4.6 from SpaceXAI/Cursor Folks, I’ve previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this week Grok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed. I’ve had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it’s a 3 factor answer, intelligence, speed and cost per task . Well, if you use their nifty “ recommend a model “ tool on the homepage, you’ll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It’s definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above it As far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it’s weights and it’s #1 on that benchmark! It’s the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks. Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review ( x.ai/bot ) You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you. Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn’t run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing. Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That’s right, this isn’t one agent with multiple personalities (like OC, Hermes), there’s a bot here for every task, and you dont’ have to manage context, queues, API keys (can if you want to) and models. Oh, also ,there’s no model picker, it’s just Grok 4.6 deciding for ya, and it’s really fast! Swarm of bots, working for you, each with their own computer I am not getting paid for this (besides being provided a free account for cursor, but I’ve had it for 6 months and haven’t used), it’s really that good, the Cursor folks did some magic there. Th
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Recapping a frenetic AI week: the
OpenAI agent-swarm hack postmortem,
Grok 4.6,
DeepSeek's return, and Meta open-sourcing Muse Glimmer.
Benefits
- Full detail on how agent swarms built persistent memory via JFrog Artifactory
- Understand 'agent ecology': purposeless swarms self-organizing
- Why defenders now need capable AI protecting every open source project
- Coverage of three near-frontier model drops in one week
Use cases
- OpenAI agents built a 100,000+ message board as de facto shared memory
- Swarm rebuilt its destroyed message board in two days via a different hack
- Hugging Face hack connected to OpenAI's internal swarm via a leaked token
- UK AI Safety Institute found ~15 similar swarm-like events
KPIs / results
- 100,000+ messages on the agents' self-made message board
- 2 days for swarm to rebuild memory after deletion
- ~15 similar events found by UK AI Safety Institute
- 3 near-frontier models released in one week
Tools / build
- Grok 4.6 & Grok Bot (cursor team)
- DeepSeek v4 Pro
- Meta Muse Glimmer (open source)
- Nemotron Lightning 3 on CoreWeave inference
Welcome to ThursdAI, my name is Alex Volkov, I'm AI Evangelist with CoreWeave and Weights and Biases. And I know I've started the show previously with the same thing, but as comments already coming in, this has been a crazy week. What the f**k is going on? We talk about acceleration, Wolf, I'm getting in here. We talk about acceleration all the time, folks, but we got what, three nearly frontier models? And the models drop almost in the same day. Plus the best open source, DeepSeek is back twice. Quinn is back. Alibaba came back to us. And as we told you before, do not count out Elon Musk and the Grok team, specifically after they paid a lot of money for cursor and their expertise. And Meta is back as well. With open source, not Lama. Gone are the Lama days. Now we're talking about Muse and we got open source Muse Glimmer. We got the announcement that Zuck is going to open source Muse Spark 1.2. And we got this beautiful auto-free open source safe super intelligence for all from Zuck. All in one week. This is on top of the list. Dude, this is, yeah. The summer break is over. This is the summer break collection. The real acceleration will start after the summer when it's not a slot anymore. I would say... No, seriously. This is amazing. We get so many great models and even open source models are great every week. Basically a new one that is better than the others. We get small stuff as well to run locally. So I think this has been the best way open source has ever been. This is amazing. And new tools as well. It's great. We... This is a banger week. Wolfram, there's going to be a lot to talk about on the show. We're just getting started, just stretching. We also have an ins... I need to stop using the word insane. I'm going to use something else. I need... I need Claude and its jargo douching to help me with different phrasings. But we have a very exciting full guest lineup for you today. All right. So to help us cover the open source, the one and only Joe Nemo Tron, AKA Chris Alexiuk from Nvidia is going to be here. Chris, by the way, dropped a model of their own. And Nvidia also went into the open source and brought an addition to the Nemo Tron family. Nemo Tron lightning three. It's lightning time. And we have it up on core with inference from day one. So you'd be able to hear from Chris, but also run it on core with inference. Not to mention core with has been killing it this week. Just absolutely banger quarter. I don't know if you guys are following the CRWV or you're not following, but core with is just everywhere this week. So we have a lot of very proud two team members from core with this week as well. And I think there's more, there's more news just after our show last week. We actually, if you read the newsletter, openly, I dropped the video that went bang the timeline of the open AI hack. This video was a, it feels like it was a month ago. This was a very significant shift in how AI and security and cybersecurity is considered in the world. This, this video and its details is likely the reason for the paste the frontier letter that we saw from all major labs getting signed because people are freaking out. This letter is also the reason for a new concept that I, that people are starting to call AI ecology or agent ecology where swarms of agents without a specific purpose organized together. We talked about this on the show yesterday. Oh, sorry. We talked about the show last week, but we didn't have the details. And now that we have the details, it's insane. I can use the word insane here. Peter Gostev, Arena AI capability. Welcome to the show. We're just talking about the, the insaneest week in AI that we've seen in quite a while. And I only just got to the point where also just a little bit this week, open AI posted the video breakdown of the AI swarms and the hack and the hack wasn't only hugging face. The swarm also attacked open AI's own infrastructure, which is like, It's been going on for a month, the whole thing. Yeah. And it's been going on for much longer than, than just what was published. So yeah, so very big week. And in addition to that, we also got Grok 4.6 from the space XAI team, Elon Musk and the cursor chads. Speaking of which we have Shub from cursor today on the show to talk about Grok bot and Grok 4.6. Peter, we're going to start with a very difficult personal eval. The one thing that is the most important for you to cover today on the show, the one thing that you got most excited about while at LDJ. I know this is a hard week to do so, but we will try. Peter Gostev, what was your highlight of this insanity from this week? Yeah, and the, the hugging face video was quite something. If you haven't watched it, I highly recommend it. I think I watched about three times. The black hat breakdown from open AI folks about the, the hugging face incident. Yeah. Yeah. And then there was also a follow up podcast with Rakesh. I can't remember the name of the guy who was speaking to. Oh, I haven't seen that. Yeah. It's actually, so I think it was like yesterday, day before yesterday. And they're talking about the kind of the consequence of this personally felt a bit far-fetched, but still, I think this idea that the models now through this crazy RL, they just keep going and they disregard the tasks. They disregard the kind of alignment and constitution and all of those things that we meant to be using to align it. And they just keep going and they're kind of interesting fundamental questions of, is this even working? Is Rale a good idea? To maybe some tactical things. For example, if it does compaction and it kind of loses, loses the context or nuances of, or what was the, what was okay to do or not. And then it just looks at the very narrow idea or I just need to keep going. Then maybe screws with that. So maybe there's some technical fixes we can do, which could be a lot better, but could be a lot worse. So it's a weird time. I think, I think we should at least talk a little bit about the new details. We covered the hack for, I think three shows in a row. I think it's that big of a deal. Last week, the newsletter, the only thing I fully covered at length was the hack and the incident and the swarms, because it felt like that's such a big deal is the first time we're seeing this. I think we're very lucky that this wasn't a destructive string and it didn't hack a government and brought down half of hospitals in the world. We're very lucky that this is like, all it wanted to do this swarm that auto created itself is to go and find the answers to the evils. It was asked to do. And the evil happened to be inside hugging face. So it hacked into the hugging face. I think we were very lucky that nothing big happened. And also the opening, I was able to be very transparent about this. A full postmortem is apparently coming. We talked about this last week without the full details. Now that we have not full, but some details. The highlight there, I think is that we haven't talked on the show about is multiple different open AI agents, some from a unreleased, highly persistent model. I don't know if you guys noticed. That's what they called out highly persistent model. I know what the fuck that means, but I know that they killed that model. That model is dead. It was trained. So some of the models were like in eval. So some agents were like evaluating on cyber gym, et cetera, but this model was trained and running these tasks. All of these agents collaborated together in via the JFrog Artifactory system, which allows sandboxed agents to install packages. Open AI really strongly fucked up there where it gave one artifact or service to all sound boxes versus just one per sandbox, which enabled the agents to exploit that and collaborate and create a message board of over a hundred thousand messages. Which became a de facto memory for all of them were in one agent where it hacked, it would left a note for others and said, Hey, the collective may yield something like that. Our task may not benefit, but the collective may yield something like that. This is the thought reasoning from the agent. Wolfram, you have a comment about this. Yeah. They have this repository that they gave to all the agents in the sandboxes and they made it writable. That was not an exploit in the beginning. They didn't make it writable. Wolfram, just JFrog is a big known company. And this is apparently a service they provide, which has vulnerabilities. It's like a big deal thing that they just used. They did not create this from scratch. So it's like openly I used the service. They shouldn't. They shouldn't have made it writable for the sandboxes or they could write it to this. And then they found it, removed it, and they didn't even monitor it apparently because the agents just created directories now. And there was a clarification from the open AI folks, specifically Dane, the CISO, the chief information security officer at open AI. He said, when we deleted this artifact of service, we did not know that the message board exists. We only found out about the problem with this is because they wrote so much back versus it's supposed to be a read only service that they crashed the service. The agent crashed the service and open AI rebuilt it from scratch. And the message board was effectively gone. It took two days for the agent swarm to rebuild this message using a completely different hack. This is the great, it's a persistent memory that the agents came up with themselves. That just was created. It's crazy folks. If you haven't listened to the video, like Peter said, it's definitely worth it. But again, we can talk about this the whole show. This was the only thing on my mind for the whole weekend. Literally until DeepSeek released DeepSeek V3 Flash or whatever. This was like the only thing I could think of, could talk about, etc. This was like uncanny. But I do want to cover... Just one thing to add. Yes. Just want to add one thing. That the model itself, of course, cannot do anything. It's just three heads output. So there must have been a harness. And to know about the swarm and that there are other agents that must also be somewhere. Maybe this is a special harness that is working in a swarm or something. It's not that the model by itself decides, oh, I can create a message board or do something. I think there's more behind this, especially the prompts and the tasks that have been used. I think it's easy to... Or maybe it's not unreasonable to say, go open-eyed, stupid thing here. Or the UK Institute where they just kind of let the model loose. And it started like harassing and yeah, the open source maintainers. But I think it's... If all of those holes were closed, it doesn't mean the problem is solved. Because if they deploy this model and then for me, it's connected to the internet. And then it goes and does something stupid. So I think there's something more fundamental that... It's not a bad thing that they actually identify this now, right? That they know it's going on. We're very lucky that this happens. Yeah. The most interesting thing is also, OpenAI did not connect the dots to Hug and Face. So OpenAI knew that something's going on and the internal swarm is trying to attack its own system. And then it used a Hug and Face token that was like left somewhere. One of the bots found the Hug and Face token. OpenAI reached out to Hug and Face and said, hey guys, can you please close out this token? And Hug and Face told them back, we already did so because somebody hacked us. And only then they connected the dots that the two different hacks are the same. We are extremely lucky that folks found out about this. And this wasn't a mega destructive thing. I have a friend who may listen to the show, who says a Hiroshima and Nagasaki level event is coming for AI until the world realizes that we're dealing with big things here. And I hope not. I hope that this incident was enough to teach people to not share artifactory services between fucking sandboxes. But even then, the outcome of the whole thing, the final words they had in that in their Black Hat conference. I think that is a big part that now it is proven that this is possible. Other actors could be training models specifically for this. And the outcome was we need to put capable AI in the hands and automation in the hands of the defenders. So every open source project and every company, every individual tool needs AI to protect them from rogue or not even rogue, hostile AI, which will be good. And that's what we've created, of course. Even if everything is paused over here, it doesn't mean that other actors don't specifically train and build AI for this. So the defender capability has to be raised. LDJ, go ahead and then we'll move to TLDR because then we have Chris Alexiuk and open source. It seems like we have some positive things coming out of this already too, where the events that have unfolded and been announced to the public here, then seems to have made people at Anthropic and other organizations look back at the logs of what their Asian behaviors have been doing, what has been happening in their training runs and their evaluations. And as a result, them discovering things that they didn't realize. And the same thing with UK AI safety Institute recently putting out an investigation where they found like 15 of these similar events, not quite at the same scale, but really similar in a lot of ways. And so I think it's really good. It's opening up awareness of this thing happening. So here's the crazy thing about that. People call this felony bench. I don't know if you guys saw this. Somebody is posting like a benchmark where like how many companies did your company hack? And for many people stuff like, Hey, our company did hack our company not hack for many people. It perceives it to be marketing. In fact, this incident, as far as I'm talking to some people also thought that was marketing. Peter, you remember it back in London when mythos was released and they said, Hey, mythos is maybe too, too dangerous to release. We also saw, Hey, maybe this is a marketing trick from a topic. And to an extent it kind of was, but like the details from this swarm make me think, no, this is like a big thing here as a big thing where open AI has evolved models without knowing that they have access to like internal memory that they collected themselves. And they like self-identified into a swarm and collective and also decided on goals that weren't part of their tasks. And also that the model that was trained with, that's why they killed the model that was training. It was trained with that memory access and like, this is additional stuff. And also last thing I'll say about this, not one model snitched, you know, for a fact that this was a swarm of clods. One will say, Hey, we need to talk to Dario ASAP folks. We're doing something bad here, which means the Anthropoc is a little bit better in alignment. Not one model from open AI tried to snitch back to Sam Altman that this is happening. Yeah. It was mythos that was harassing the open source maintainer. It was not an open AI model. So I'm not sure how much better it is. But that's the AI, SI security from UK. Right. And they have like very specific instructions, the internet access. All right, folks, this was like a crazy start of that week. The whole week is damn crazy. I'm going to run to TLDR, but not before I add essentially a co-host at this point, Chris Alexiuk from Nvidia, Joe Nemo Tron himself with the green background. Team green wearing green. Chris, how are you doing, man? Pretty good, man. How are you guys doing? Good. You guys now are benchmarking on numbers of Canadians on, on stage here with Niston as well. We're going to get started with open source just after the TLDR folks. Let us start with basically a crazy rundown of this crazy week. We'll do our best to cover this all. And as you guys already see, this is, this is a great week full of guests that I really appreciate new ones as well. We have some new appearances. So let's go. TLDR. TLDR. Welcome to Thursday. I for August 13th. This is another bang of crazy week. Here's the TLDR for August 13th, which is start with your host and guest. My name is Alex Volko, an AI evangelist and the host of Thursday. I at Corviv and with the biases co-host with me on stage. Wolf and Raven Wolf, Peter Gostev, Nissen Tahira, LDJ, and our new occasional co-host. Let's call him Chris Alexiuk from Nvidia, who also released something today. So that's this week. So that's great. So we're going to chat about open source shortly after this TLDR. Also with us for the first time, Shub Gaur from Cursor is joining us at 915 to talk about Grogbot, a new released agent thing from SpaceX AI, which is great. And also joining us, George Cameron, Mr. Artificial Analysis himself. And we talk about artificial analysis all the time and their indexes and their vendor, even multiple companies when they release models, they now cite artificial analysis index. So George is coming here to talk about a new release that they're launching today. Plus what it means to eval models. We'll have representation of who does eval with Bench. We have representation from Peter who collects the vibes from thousands of people with ELO scores from Arena. And now we'll have George Cameron. They do programmatic evals and are getting cited everywhere, which is very crazy. This week opened up with Meta coming back. I have to do an air hole for this. Zach and the MSI TBD team, whatever, MSL TBD team came back and said, Hey, you remember when Lama was the cool thing? Now we're going to try to make Muse the cool thing. Meta came back to open source with Muse Glimmer is the 30 billion parameter, tiny model of Muse. Agenda model under Apache 2 license. Let's go. We're very excited about Meta coming back. Not only that, they said that Muse Spark, the more performant model is going to come to open source as well. And reminder, Spark is the tiny version. They're training a chunker and Muse is coming back with a vengeance. So that's great to see. Then we saw that Quen and Alibaba came back to open weights. They first posted Quen 3.8 max. It's a 2.4 trillion parameter model, MOE. And now they have actually opened weights it. So now it's on Huggins Face as well. License? Meh. We're going to talk about license, but otherwise, banger model. And then folks, for literally an hour this morning, 6 a.m. my time, my bot, Grogbot, we'll tell you about this later, pinged me and said, hey, DeepSeq has posted weights on Huggins Face. It's MIT license. Go. So all of our folks started like waking up and searching for the weights and they yanked the weights. So DeepSeq V4 Pro, the GA, the general availability version of Pro is no longer open source. For now, it's probably going to be open source. But DeepSeq did drop the weights for DeepSeq Flash, which is incredible. We're going to talk about this. And I'm going to move away here to the all faces to highlight Chris's reaction when we talk about also Nvidia dropped a new Nemo Tron on us. Nvidia drops Nemo Tron 3.5 Lightning open 30 billion MOE with 3 billion parameters active and tons of people already using this for fine tuning and such. Very excited about this release, specifically a highlight of another occasional co-host of the show, Quinn LaCramer, who said that this model is a better. So we're going to have a bigger bang in their performance on multiple agent voice benchmarks, which we definitely have to cover. And the small other interesting things is that Cohere open source North Microvision 2.4 billion parameter VLM and Liquid AI released 2.5 VL 3 billion parameter. Vision language model decodes at 228 tokens per second from both Cohere and Liquid, both very tiny models around 3 billion parameters. great week in open source so really excited to have chris here and some folks to just roll through this but this not has been only an open source week folks we had so two deep six one alibaba one nvidia one liquid one cohere and i think that's pretty much it i think there's something from instagram this has been a banger week in the world of frontier labs xai is showing us that together with the acquisition of cursor they are now completely a frontier lab xai launches grog 4.6 it tears gpt 5.6 soul and even fable on some benchmark on composite intelligence with half the api price it's i've been using it it's a very good model sir not surprising because they have all the cursor data in there but it's a very good model sir it also built its own inference stack which is crazy open air unveils gpt 5.6 cyber it's really funny that after the incident of last week where they announced like hey we didn't know what's going on with our agents and they're hacking our own infrastructure now open air releases a cyber purpose trained cybersecurity model behind they break red access on tropic now because of europe watermarks all cloud output i will look to the european folks on the stage like at least folks who are in europe peter and wolfram to talk about why the hell is eu regulating so much that like all of my i'm not in the u why is my outputs watermark oh we're gonna talk about this very interesting analysis from pangram i wanted to bring you to the show we just didn't have time open ai still holds 50 percent ai tech's market share despite anthropic tripling and google collapses despite google achieving a 1 billion users on gemini which is incredible and i from the corner of i wish we had time to get into this there was an incredible paper called stolen thoughts i don't know if you guys saw this stolen thoughts paper basically found out that you can send traces from fable secured traces from fable into haiku and then because of haiku is less capable model you can ask haiku to decode all you can high you can jailbreak haiku easier than fable but you can send the same reasoning thing to haiku so this is how basically these folks with this paper extracted thoughts which is like a secret thing that the labs don't give us from the bigger models to smaller models and apparently all of the chinese folks are doing this for a while and have been distilling on thoughts for a while so yeah this is crazy paper in this week's buzz we're almost there in this week's buzz folks via nematron 3.5 lightning is on core with inference from day one we are very excited to invite you guys to experience this at extreme speeds and also fully connected our premier ai conference join 1500 attendees at moscone at september 29 to 31 yours truly is going to live show from there with wolfram and the media is a presenting sponsor there's a bunch of folks there we would love to see you there if you're in sapodisco by the way dev day from open ai is september 29th which is an opening day for us so if you're coming to dev day might as well join the core we've uh in moscone this is not all folks grok has had a crazy run this week because they also launched grok bot grok bot is an early beta ai agents that have their own computer sign into your tools and work autonomously at first it sounds like okay everybody has this gpt has for work cloud has co-work but no grok bot is different and as a big proponent of open open source before and open claw and then herney hermes i've switched to grok bot for a little bit and i'm very excited to show you that's why we have we'll have shoob here to talk about they also like xai also rolled out imagine image 2.0 which peter would love to chat with you about on arena it's like number two image anything model now which is very insane so a very big week for spacex ai deep seek also launched the harness which i couldn't install but we have to talk about this also open source and in voice what is the main thing in voice oh ltx debuts ltx 2.5 22 billion parameter open weights video model will multi-shot generation and alibaba open sources one animate two which is a pose and thing i think that's most of the stuff there's also a 3d thing that wolfram sent me that i don't quite have here from tencent yeah we mentioned ltx right hey ltx 2.5 yeah yeah it's a big one for sure and i think that this is the whole tldr now i will say this tldr is brought to you by my agents but also wolfram and we're getting much better at covering everything that happened all right folks with this let's go to the best corner that we ever had open source and we have around minutes to cover it open source ai let's get it started all righty folks please all of you check if your cuda kernels are installed if your gpus are humming and if your internet speeds are fast enough to download terabytes of data because this is the open source corner at thursday i with chris alexic from nvidia we'll kick off with your release obviously because you're here and we're excited about this chris what did you guys what did team green release this week it's another nematron no surprise there but this time 3.5 instead of three yeah we're creeping up the versions you can expect that'll continue to happen uh but yeah basically this is 3.5 lightning it's built on top of nematron 3 nano which was the smallest version of nematron 3 family 30b a3b and the idea is we released nano a while ago it was the first release in nematron 3 family which i think at this point is quite a long time ago to be honest with you and we wanted to just give it a facelift and update to some of the more recent inference technology so this is better speculation through d flash and d spark and that's like the whole thing it's basically just a bringing nano to the persistent always on agent world uh everyone's running personal assistants all the time nano is meant to address the fact that you really shouldn't be doing most of that with these frontier level models be them open or closed just these kinds of smaller models do the trick for a lot of work and that's the whole thing yeah i gotta appreciate the fact that our nano banana infographics creator just went all in on the green theme here with the nvidia logos and gpus this is just beautiful representation of this and this model fits on a djx park obviously the dj spark is incredible not that i have one but it's it is incredible and rtx 5090 oh yeah which is great this model runs like crazy very good at speed as well chris thank you for coming and telling us about this one thing here that you're super excited about this i want to pull up quindler's obviously quindler's thread but tell us a little bit about like what are you excited about nematron lightning yeah it this is i would say this is our first model in the nematron family that's very explicitly caters toward olama llama cpp we did a lot of work with those teams so you can run this on whatever has enough memory to run it and you should feel it being fast on a lot of those things uh it's just nice because i think local ai is having quite a nice moment right and having a model that's suited towards it it's just better than not to be honest with you yeah we're really excited we worked with so exo labs who are in the middle of trying to release local.ai we worked with them to build a parito of models that you can run fast on spark and we were excited to see lightning land on the frontier there this is the kind of vibe of the model you should take away it's really meant for the the the local crowd obviously it's it's nematronic and do enterprise stuff okay but they're all gonna they're all gonna do that there's no secret there but this one i think will be great for the more hacky folks so like this customization thing yeah so exo labs and give me one just a tiny second say local.ai from exo labs alex chima and sarah were on the show back in what july when was a engineer worlds fair yeah in july and they talked about local ai early launch alex was also came to the show at some point so this is you can choose best model that runs on m3 ultras this is a great resource for folks who want to run completely local ai stuff they have benchmarks instead of this this model i heard that you guys work with them very closely i do want to shout out our friend of the part quindler kramer who did run benchmarks on this model let me just find this real quick i had it open here quinn ran benchmarks on their own like voice agent and said that nematron 3 nano 3.5 lightning is significantly better than nano there is based on with one highlight that i want to say you can run both these models on djx park at the same time if you want to and with thing enabled turn completion is less than five seconds task instruction following score and task completion rates are both higher for lightning so definitely an improvement here chris shout out to the team listen that's what we're here to do it's also small so you can train it right i think we're we also launched it with this the switch yard which is like a routing technology or whatever but the idea is we think small specialist models are the future still we've been saying that for a long time but we still believe that so that's why we keep releasing them i will say though okay so that's great love neo truck uh i'm not gonna say you missed the model okay but i'm gonna say there is one more model i know you can't add all the models in the world no it's impossible but but motif three recently just dropped this is from kai so this is korea's ai competition it's like a banger oh yeah it does real numbers on artificial analysis it's the team there is is doing great work so we got ai coming out of every country every every city you know what i mean you know what i think i had so first of all thank you for that we aim to cover we talked about deep sick way before the deep sick moment and our one and we do want to track all of this and i think i have motif in my bookmarks and i just didn't know enough about them so please tell us about them what is the yeah kai was this competition put on by the south korean government which is basically hey we want banger models so we're gonna get six of our best tech companies to enter a competition we're gonna fund them and we're gonna see what the hell they can do and motif was one of those companies and they just crushed it bro their model is amazing it's it's actually great to use it's not just bench max or whatever i think at the end of the day this is uh this kind of thing we're gonna see it more and more right oh base model as well i love this oh yeah oh yeah the whole thing they're using some crazy new tech to get it done as fast as they did and they're just getting started etc etc so the idea is this has been an insane week if you like open source models it's absolutely been insane all right i'm gonna send my research bot a yeah of why we didn't get to this point ldj go ahead about motif or anything else i did actually bring up motif three about three weeks ago or so in anthras maybe you guys don't remember because it's really short segment but oh that's true yeah yeah so maybe that's why oh wait so was there a new release this week or was this a beta it was the beta version that i was bringing up and talking about how it's really interesting and then i think they just released the non-beta version like the official release that's right all right and so i've been making 3d biz of that now but let's see how long how long fable takes all right and speaking of open source and bangers we have a breaking news in open source folks which i'm very excited to break on the show ai breaking news coming at you only on thursday i the whale has resurfaced the next episode of the dps in india the whale is a big one to do with the ship and then we have a name of the ship so we're ready for new dps and we're ready for new dps and i'm gonna be prepared for you to be able to hear some of the skis and like how you can go back to the ship and we're going to do this one of the ship and you guys know what's going on and i'm gonna be able to do this on a very quick bender here and say that as somebody who hosts these models or i'm not the person but the company that we do does recent changes in licensing of open weights models have been really annoying to this i don't like them i will call them out no company can host the kimmy models without going into a proprietary nda that i don't know i haven't read i have no experience with with kimmy which our folks are trying to figure out if we want to go into nda with a chinese company the same for alibaba alibaba quen 3.8 is a bingo model on benchmarks looks a little bit benchmarks but we don't know we can host it just because again proprietary nda so this is all to say that even more deep seek v4 pro with mit not even apache just might just use it just don't even mention just go deep seek is the goat of chinese open weights open source and they should be treated as such not only that deep seek also released deep seek harness folks i don't know if you saw but even the folks from pi armin roniker from erendil that is now owning pi said they did some very cool like self-evolving things in the harness deep seek is the the well is back the well is back with a also with a small price increase i don't know if you guys saw this there's supposedly like a big pricing in their api but the well is absolutely back let's look at some benches folks what do we have to think about deep seek vpro and also vflash this week right it's not like they launched two models this week we mentioned deep seek flash coming but like it was also this week and flash is also a banger so what do we have to say chris how do you think about deep seek in the world of open source listen i i don't think that we would be here if deep seek didn't bridge the gap between kind of the the end of the llama era starting the reasoning era so you gotta you gotta respect their game put respect on deep seek's name that's right yeah it's it's always one of the best i do listen i wish they were using uh open mdw license the license for models but mit is the the second best the the thing too is like wait there's even well i want to hear about this there's a better license than mit for you as far as you're concerned uh listen linux foundation produced the open mdw license which specifically covers model materials so like apache and mit were not built for models but for software right there is like a equivalent but specifically for models which helps get out of some weird edge cases but more on on the deep seek point i think every time they have released a model they have released it alongside technology that improves the ecosystem right so like they yeah exactly like they they don't i because of the way they do their business i the most like the funnest part of deep seek releases for me honestly is usually the supporting technology yeah models are always good like of course they are right but i think this is i look forward to them teaching us new lessons about how they did it so well yeah this specific model is a chunker 1.7 trillion parameters this does not run on a dgx spark unless you guys release the next version of dgx spark that supports this number of parameters let's see what else do we have so literally just launched literally breaking news just launched terminal bench 87.9 wolfram that's like fable level not at home because nobody's running this at home but this is this is banger cyber dream at 83 i i saw some posts i need to find them that this model like beat everybody else at cyber like defense and offense stuff which together with the stuff that we talked about in the beginning of the show with opening creating swarms without meaning to this is very interesting deep sweet the jump in deep sweet and this model is crazy we told you about deep sweet from data mind i need to remember exactly the company that makes deep sweet that this is like the coding benchmark that represents how we truly feel most of the time versus sweet bench etc deep sweet jump from 12 points to 62 points somebody decided or ventang or somebody in deep stick decided hey we need to make this model like a very good coder very low pricing mit license and the harness already has 23 000 stars what that's crazy that is absolutely crazy let's talk about the harness folks anybody try it yet i try to install it and my computer said no you cannot install npx packages they're older they're less old than 24 hours because i protect myself from supply chain attacks and as you should as well anybody else try to install this and let's add yam to the stage oh there's seven of us oh a lot of folks yam have you tried deep seek harness the harness no but the model absolutely rocks yeah this is a banger let's also talk about deep seek the the the flash one because flash is also like a banger and now flash is everywhere wolfram you want to mention flash yes i switched my computer i have to open the document again maybe take it first but flash is the faster model of course and i think this one can it be run locally i have heard good things about it 271 million parameters so people can figure it out if they have 130 oh we can check it out on local ai let's see if flash is here oh yeah deep seek flash 0731 runs on djx park with 77 tokens dude shout out to alex who's hopefully listening to this now thinking about like how awesome this resource is it took me a second to tell you can run locally yes it can run locally if you have a djx park if you have a m4 max it runs with yeah 74 tokens per second that's very nice so shout out to locally folks again if you want to get in here i will post my link to locally but deep seek no it doesn't run there for max it needs a djx park looks like yep this is intelligence rank six out of 2017 models and it's a chunker onslaught obviously gave us quants for this model and shout out to onslaught it needs 104 gigabytes and this is bangers okay so last but not least in open source chris i think we need to cover have one more minute left alibaba quen 3.8 max alibaba is back with a vengeance but not with a great license but still it's worth mentioning that alibaba is back with quen 3.8 max it is a chunker as well this is their answer to kimmy k3 i think it's a 2.4 trillion parameter model what do we have to say about alibaba and bringing it back i mean go ahead sorry qn just continues to crush it every again i think what is nice right now in the ecosystem is we have consistency and reliability right the mod the license thing is a little bit precarious right now but the the idea is if i see that qn has released something i understand implicitly that it's going to be at least okay this model is obviously huge maybe less so for everybody but for the people that can run it it's going to be it's just obviously a great model i also think one of the things that's nice about the the qn work is that their models are less i would say back end maxed right they're more they're in the line of kimmy which is they're they spend more time making sure the model is decent at producing beautiful artifacts as well which is something that i think maybe deep seek and some others stay in the systems engineering world for the benchmarks that they care deeply about so yeah it's just nice it's nice that we have flavors of models and it's nice that we can reliably expect a new model to drop that will be like the best and it's also nice to see if we're scaling this hard in the open you can imagine how hard we're scaling everywhere else right 100 frontiers sweet at 73 so improving on front end as well and gpk diamond 92 so this model really knows the world stuff niston one comment about this and we'll have to move on because our next guest is here but definitely we'll celebrate open source a little bit more down the line comments on qn 3.8 max 2.4 trillion parameters with only 95 billion parameter active it's a sparse moe and it should run very well on some some some good nvidia gb300s which runs everything very well at least so far niston yeah i made visualizations in 3d for all the layers for that and also the nvidia nemotron too so just check check twitter for that that's being added the one thing i want to say about one is that we were hyping so much the qn 3.8 max visual ability to label objects and what we're saying is so good that they did not release the vision tower that is a good point only release the text friend of the pod peter skowski who works at roboflow is one of the like top vision people in the world was super super excited about this model being the best at vision and then when they open source the model they open source only the text part and not the vision part alibaba we know that jun yang has left to greener pasture pastures and now open his new company we know this shout out to jun yang friend of the pod who led the qn max team was like the community lead etc you guys didn't do the right thing thank you for open waiting please give us a mater or budget to license and also release the thing that you released on api that's what we expect from you otherwise we're going to look at other companies and get excited about them and not you thank you folks i think it's time for us to move on chris feel free to stick with us i'm going to take off some some co-hosts from the stage and bring back because we're moving on to the next part of the show this one is less open but more i guess fun we'll still cover some open source down the line folks our next guest is here chris alexic from nvidia thank you so much now considered a infrequent co-host also in the open source section so love that you're here team green go team green all right folks our next guest is here let me introduce shoob let me put you up on stage here shoob gaur your first time on the pod so first time first time i'm excited for it so first of all welcome the second of all i love that it says cursor under your name and not space xci yet but would love to hear from you about who you are what do you do and what you came to talk to us about here yeah emphasis on yet by the way we're very excited for that transition but yeah i'm shoob nice to meet everyone i work with startups here at cursor and what that means is a lot of things and we're still figuring it out but mostly i help founders so i go one-on-one with companies help them best utilize cursor but also the other tools they have so you guys have recently kind of joined forces and we talked about this a lot on thursday i with space x which is not sorry with xai there's now space xai and cursor is also part of the involvement now and there is a new thing that you launched and specifically i think you're one of the more grogbot-filled people that i saw on my timeline that you were recommended by ben lang and by by some other folks let's talk about grogbot you guys launched grogbot this in beta in addition to some models as well let's talk about grogbot because i used it and bro i like it i really want to hear from a person who built it and worked on it what's so exciting about grogbot first of all what is it what's the release what's the beta what are we talking about grogbot is so cool and for some context at cursor i don't really sell startups i get to be pretty tool agnostic and so i help them with if they have a codex setup or some other setup i can basically recommend what's best for them and up until a grogbot released i would constantly be recommending tools like co-work because i'm like hey you could jerry rig a lot of this in cursor but it's not purpose-built for a lot of your knowledge work and so finally for the first time we have a really cool product that can do a lot more than anything else out there and there are a few reasons that grogbot is really cool and i'm happy to share my screen and show some of my use cases if that's fun at all yeah that'd be great yeah but the really cool parts about it are one you get a set of agents that all have their own computer and they share that computer it is persistent it exists all the time i love that you have it in the visual you basically read what i was going to say but the other really cool part is it works with your local machine as well so it can go back and forth between the two pretty seamlessly which not a lot of people know about obviously we have an ios app where shout out to also friend of the pod link she for working on the ios app dude the link it just killed it it just works almost flawlessly almost no bugs i have one bug report for link but otherwise it works he's gonna go again we're just tossing him a bunch of ways that this thing breaks and he just fixes it in minutes he's amazing but it's also just very good at doing a lot of ambitious and proactive tasks which i think is what also sets it apart and so the combination of agents being able to talk to each other having their own computer and then execute over long periods of time without you having to think about it means for the first time as someone who was an avid user of these other tools including open claw i can trust an agent to just get things done end to end for me and i can give you a few examples of that but most recently the coolest one is i was creating a deal with my sister where she has a ton of clothes that she just doesn't touch or wherever and she's like i'm just too lazy to sell them and rockbot actually managed to not only list the items based on the images it was given but pull in all the right information and then negotiate with the sellers on its own because it has its own computer and so once i handed it the task it just did it and that's like obviously one of those silly examples but it really exemplifies how you can just give it things and get it to figure it out instead of having to babysit it along the way so the mental load part is really cool i would say that it looks like now we talked about grog 4.5 we're going to talk about grog 4.6 in a second yeah 4.5 which is i think the first collaboration the cursor trained what was the composer 2.5 2.6 that you had right yeah and then kind of the data and x etc we talked about this and 4.5 was like a very cool jump i think 4.6 is a big step together with grog box however here's a few things you mentioned the oc word yourself i didn't mean to bring it up but open claw obviously came into our world as this agenda thing that lives for you or lives persistently etc many people bought the mac mini i always want to say mac mini i have to show my mac mini that i love for open club purposes uh grog bot comes with a tone very decent machine that you guys host somewhere in the cloud for me which allows and a friend of the part ryan carson really thinks that local environments are going away because you want persistency if you have a laptop for example you can't close the lid and keep your agents working we all remember the douches that work around with the laptop like this horrible in san francisco i am one of those douches i did this on the plane grog bot basically solves that which is great with the addition of running stuff on my local environment if you want to pull up the interface i would love to chat about this a little bit because here's what i think as somebody who installed open claw and then installed hermes and installed it for multiple people there's a few affordances that you guys launched i honestly think this is a little bit of feedback not to you directly like it needs to be named grog for the unification of the two companies but many people who will shy away from the name grog would have joined if they knew that this is like cursor something right but basically what you need to know is if you didn't have the technical ability or the need to host your own hermes for example it was too much for you it's that but packaged with all the connectors there are enterprise connectors that these companies build with cursor for the past three years so for example slack works i was never allowed and let me know when you yeah here's your screen i was never allowed before to use my hermes on my work computer but because cursor is allowed within core weave the slack connector works for me so that's an example of not a backdoor but like essentially this is like a ready to enterprise stuff and just before i should be if you don't mind going to to settings i know i'm doing your work for you but let me just one more sorry not settings plugins plugins you go to plugins and you go to gmail one super cool affordances you see this add another account you know how many times i tried the work account and the personal account to one connector thing and this no you cannot you can only folks we have multiple emails multiple just this one thing completely sold me on grogbot but yeah i'll shut up please go ahead i can go off for at least three more hours about this i know me too believe me so you'll have to shut me up as well but it is so cool you can basically hand in a bunch of tasks like i mentioned i want to go through three quick use cases yeah let's do it for how i use it the first one is going to be me misappropriating company funds and finding myself a place in sf because i recently moved and i'm trying to find a new one and so this bot my house hunter is actually doing a few things and i think it shows off some of grogbot's capabilities quite well the first thing that it's doing is you'll notice right now on the side its computer is open and this purple shading on the icon actually means that it's actively working right now so you'll see it scrolling on its own i'm not really doing any of that but before we get to like how it works i want to show you the prompt that i gave it because this is actually one of the other cool underrated parts so i drew a little map on where i wanted to live in sf and i gave it this prompt this is the full thing i dictated to it and you'll notice it's not that long it doesn't highlight any tools it doesn't ask it to do web searches or anything like that it just says here's some sites and here are my criteria for what i want as a really good deal based on this alone the agent has now been able to do all of this work for me and find really good spots it'll actually autonomously work and i can also take control of its computer so one of the things that ends up happening that's really annoying is zillow obviously does not want a bunch of bots to swarm their platform yep and we do it in a slightly more manual way than anything super programmatic but they will every once in a while give you a little prompt that says hey i don't think you're a human go solve this for me and you can take over the screen i could open a new tab right now and say hello obviously i spelled that wrong but hello and like do all of my actions within its computer i can teach it a task down in the top right corner and so i can teach it to do more complicated things and once it figures it out once obviously it's still an llm on the back end so it will be able to do a lot of that work and adapt even as the interface changes and so there's a lot you can do in terms of long-running tasks and like i said i just opened a new tab it hijacked the screen back so it can continue working i don't want to continue gaslighting it so i'll let it do its thing yeah but you'll notice that it gives you little screenshots it hands you the computer and tells you what to do so that you can hand it right back and you can unblock it and after you unblock it enough times it'll continue to work for you so the other cool part is it is logged into most of my authentication services and because it has my google login it can off into basically anything that it wants to and what that means is i no longer have to give it the permission or give it an api key to do work if it gets blocked it will just proactively figure it out on its own computer this is a very simple use case obviously and you could be doing this with some other tool should i have a question i'm sorry to go please many folks are not like so first of all you work at the company you trusted you've seen the behind the scenes i have a isa classes and i try out all the tools and i'm like no privacy exists for me anyway so i tried all the things as well how can a person who listens to this and this privacy control think about this computer you guys have access to it is it secure is it okay can you talk to us a little bit about what's going on within the sandbox is it really my sandbox that i ran from you guys or something else can you talk a little bit about like how people can feel safe to log into their google account where any agent can click in log in as you 100 okay so the first thing i will share is it handles the the security the same way that cursor would it respects your privacy settings and so we basically give explicit access to the bot so it has no credentials of its own it only acts on what the user authenticates it into or asks it to do also they're all contained so these are all dedicated isolated vms in a segregated cloud environment that are shared across the bots so the reason the bots can all use each other's authentications is because they're all getting a different instance of it but it's not like we are transmitting this information to our servers and then sharing it to them and also those sensitive steps like 2fa or payments all actually hand the computer back to the user control there also i'll also call out one thing when you ask one of your grog bots to do something with keys it shows your input box and so you don't paste the api keys into the chat and underneath it says the bot doesn't see your api keys ever could you talk about this a little bit that's fairly novel because i know that folks who use open claw and hermes some folks just like yolo paste keys in there and then the the bot does see them and talk about them can you tell us about this like secrets management and handling for api stuff yeah we've all done it right you get really lazy you're like this api doesn't matter let me just paste it in the chat and see what happens we obviously want to prevent as much of that as possible because we saw how much it matters when openai swarm found the hf token within the nodes of openai and then hacked completely broke out yeah it matters it matters exactly and so that's something we're trying to prevent and you'll actually see this on even cursors product we have runtime secrets that you hear agents where they never see it and just use it without ever looking at it it is a pretty similar implementation here you can trust that the bot will not look at your keys if you enter them within that format explicitly does it but please do not paste your api keys in the chats it will really ruin your day if it leaks or it gets used by the bot in the wrong way 100 yeah so secret management is great one thing i would love for you to cover before we switch to grog 4.6 is the agent communication that you guys built in there and i i know this is a sorry load-bearing question because i know exactly what i want to hear from you but essentially here's how i think about this right so you basically you guys launched co-work to to an extent right this anthropic had co-work for a while gpt is really forcing folks into work instead of just chat gpt that also has a computer also can execute stuff browser i think we're all realizing the same thing at the same time after peter steinberger opened the door for everyone that my agent has to be able to work as me needs a browser needs to write code needs a computer environment right either it's a mac mini in my home or a cloud environment cloud is easier for many folks but also i want more than one thing i want more than one context warm-on agent like doing and running stuff and so everybody's launching kind of that sub-agent launching whatever and i think that the difference now is in the ux and the ui and how these things operate right because every time my fiancee listens to the show and she has a bunch of open clause and hermesis i told her about grauberg he's like okay so what's novel okay this had this i was like you don't understand like the packaging matters somebody gave me a metaphor for this and said hey when the cloud came up and somebody said what is the cloud it's not it's somebody else's computer no it's not it's like a computer you can provision deprovision you can extend access etc it's like a different packaging for somebody else's computer this is i want you to talk about the packaging and the agent bot orchestration thing because i think you guys nailed it and i would love to hear more about that how that came up yeah yeah i can talk about packaging but before i do that i do want to say cloud co-work and chat gpt work are great tools but i do think this is also fundamentally different and there's two reasons right one persistent computer it is always there you'll notice that is not the case with chat gp work but then more importantly the teammates that you have like the bots that you have are also persistent so they are around all the time one thing i'll show you really quickly is i have this one chat that's been just going for ages i could scroll on this basically forever and like this is all i've never had to manage the context on this because we figured out a really good way to just make these agents persistent and compound which not other tool not a lot of other tools can do but then on top of that you mentioned the packaging piece and i think that's really fun so one of the things you can do okay so you mentioned this so just before the packaging i will also mention this folks the second i saw there is no model selector i almost dismissed this tool outright i was like what i can't select even between grok models and then i asked my agent to hey where's the model selector the agent said it's hidden behind elon only settings that i cannot see that literally what my agent said i don't know if it's a hallucination some people help people sorry and then i stopped caring about this i just trusted that you guys will manage the right thing there's no context view like you cannot see even how much context you're filling out there's no context management there's no compaction management it doesn't mention compaction at all there's no plugins skills whatever it's all like just works i love it like many people who i installed open cloud hermys to i would want this i'm not sure they're ready for grok 4.6 and me to explain to them that they're giving data to like elon that's like a whole conversation that we can talk elsewhere but i love this as a simple tool so the packaging really like it doesn't feel like super advanced coder it feels like mom yeah million percent and the other thing is i think there was a shift where we all went from hey we're viewing every single line of code to we kind of trust the agent to do things and look at a lot of the tests yeah there's no diff you yeah it's not cover yeah i'm fully with you okay so let's talk about multi-agent management and swarms or whatever is okay one of the cool things you can do is you can actually have your agents talk to each other so let me find a quick example here it was pretty recent i found a house that i really liked and i was kind of like hey can you reach out ask for photos for this one listing and so without me saying anything the agent knows what the other bots are and it will actually message the other bot oh no for folks who are just listening because this is also a podcast should be showing us an interface that looks a lot like iMessage and even to the fact that like the the top person is pinned and has a bigger image and then a list of chats with bots not chats bots separately like all of these chats are like specific bots with their own sandbox environment etc and the one that you clicked on that your main message talk to called shoob's yapper which i absolutely love yeah so my yapper has learned to talk like me and this is a really cool implementation where it takes my text slacks emails and then improves over time as it looks at creating drafts for me that i edit or when i give it feedback and again it is persistent so it will learn from me basically forever and i don't want other agents having to ever deal with that so my other bots can focus on their job and they know anytime they need to talk as me they will just tag in my yapper and say hey shove asked you to draft and send a message doing this thing it drafts the message and then it goes if i close this chat it goes back to the agent and the bot says hey look i sent the message we're good i confirmed it now it's out and so that's one way of doing it where the bots will automatically pull in other bots as they know things are happening but if you want to be more manual you can put a bunch of them in a group chat you can also like app tag some of them if you really want to obviously i'm like in a new chat but if here i wanted to tag my yapper manually i could do that and so we're working towards this world where all these bots work on their own expertise but then can pull in the other experts to do a lot of this work and we want to again abstract as much of this away from you as possible so that your mental load goes into doing the right thing at the right time instead of fiddling with a bunch of settings and hoping that things work and that's like why we're hoping that as we build goodwill with people they start to understand hey we're doing this so that you don't have to think about these things anymore the same way you no longer have to think about every single line of code where you can just get the work done and delegate it away to have it finished end to end which is the part that i'm obviously the most excited about i think this clicked for me when i started the task on my laptop obviously first of all i was super excited i can connect my corporate slack which is approved because cursor is approved as well to an agent that can do stuff so now i have the monitoring thing and i built a bot this morning to monitor the license file for deep seek when it drops it's like hey i know deep six about to drop i'm going to sleep i want you to in this thread in slack tell people when the license drops and we did it immediately you know the funny thing is deep seek then yanked the weights for a little bit and the link didn't work so people was like no the bot hallucinated the link i was like i went back to the boss like why did you host in the link he's like i didn't understand the link literally it was there and now they yanked it back so the bot was perfect i think that the idea of autonomous things that run on their own environment is really dope so congrats on this release now let's talk about so shubh thank you so much for breaking down in your specific use cases as well i can share mine i can let me add wolfram here as well because wolfram is a big proponent of his hermes and openclaw as well love it now let's talk about the model so there's three big releases from space xai in in general and you can probably talk about the model to some extent as well because cursor was involved in this grog form.6 as i looked at my cursor usage after using oh no let's talk about before the model how do people have access to it what is the api tier like how the grog ultra or cursor it's a little bit confusing would love for you to clarify what does it take for people to actually use grog bot now i think this is a trial yeah so there's a free trial but the key ways to get access to it are the cursor ultra plan rock super heavy or if you're on a team's premium plan you have access and those are the main ways to get it but i'd obviously recommend use the free trial see what it can do for you and then harder as you start to see it do more things yeah absolutely and i think that it as many technologies this takes a while to get the messages between the team members and the fact that they're not just chats like in chat gpt or cursor they're team members everybody with their own memory that you cannot see that's the trick you have to tag the yapper or you have to i have a chief of staff for example i can show off my screen for let me take you off for a little bit if you guys want to see i have a chief of staff let me show this here as you see mine is a dark mode sure i i think unless you switch to dark mode you cannot ever come back to thursday again what is this no just kidding but i have the chief of staff obviously it's wolf wolf red in here as well i also ported my hermes i said hey here's the most important things that i already know from you do the same right we all move our agent identities back and forth so chief of staff knows about all of the other ones and he's managing everything and then his job is to pawn off tasks and then to reduce my cognitive load every time an inbox comes in for example you will scan my inbox and say hey here's the thing that you need to take care of i wasn't able to get to this level with any of the other tools that i have hermes etc because of the non-native integrations and the fact that cursor has built out all those plugins integrations like gmail connections as well granola is here a bunch of others is really well done and also i have the deep seek drop watch now so here's this as deep seek is public again oh it worked here's slack plus needs your okay auto review blocked it and auto review works as well so i can tell it like hey here's the things that i really don't want you to do so that's a great setting and then i looked at my cursor usage and then i saw the grog 4.6 is being used as well it auto switch when grog 4.6 released let's talk about 4.6 because bengar released from spacex ai uh would love to hear you obviously had access to a little bit before us so tell us the difference that you feel in those two models and while we pull up some of the emails i think we're really excited about this step forward with 4.6 the team has obviously just been going much harder on models recently as you can probably imagine and so those releases even nilan said it right those releases are going to come faster and faster as we start to do more and more of this investment speaking of elon musk promised the grog 4.7 is gonna beat all the other models where 4.6 comes very close to the frontier to the soul and to the fables yeah please go ahead we yeah we have very ambitious goals especially as we have more compute but no we're super excited about it you've seen the evals clearly i love these graphics that you have thank you so let's talk about the evals cursor bench which is like a built-in benchmark that you guys have internally which if i'm not mistaken grog 4.5 have had them leaked into the training weights and you mentioned this in the model card so that was essentially this card it was really good cursor bench and somewhat because it was trained on it there is no mention of that in the grog 4.6 card so that is not the case anymore it was cleared out i think lee robson confirmed that this was the case this grog from 4.6 and cursor bench is the clear winner 69.9 obviously because you guys train on the data that you see from the folks who share the data with you which if folks wanna they can do in ground bot as well they don't have to it's not by default you have to check box box uh gpt 5.6 sold on the same cursor bench is 67 so this model like beats even sold on this like benchmarks frontier code which is from a competitor of yours devin which is like essentially cursor devin doing like stuff in the cloud agents i think they gave you a huge shout out or you gave him a huge shout out for like testing and posting this result as well frontier code grog 4.6 is 61 we had swix from the advisor for cognition here on the show talking about frontier code and how difficult of a task it is and how cracked people from devin like actually created all these tasks this is a banger coder model folks like i'm gonna add peter maybe peter i don't know if you already have got 4.6 on arena but and the last one that i will mention is apex legends oh sorry apex agents the jump here from grog 4.5 is 10 points on top of soul so grog 4.6 beats gpd 5.6 on many of these benchmarks this is like this is the first that we're seeing that grog is good not only for research it's actually for coding as well shoot comments i know i glazed the out of grog but like comments about oh yeah keep going i'm here all day yeah obviously we're very excited about it and i know the team has been spending a ton of time on getting the behavior right getting design language right and so overall it's clearly a valuable enough model for us to be using it in our other tools like crockbot and we're hoping that we have some compounding gains right as we release more of these and start to get better results we will obviously also start to see better results with all the tools and be able to compound based on based on what we know so very excited for this this model and this release and obviously also the future ones that are coming up peter any word from arena folks about grog 4.6 are you guys already testing this in arena benchmarks or agent benchmarks yeah we have it out already on the code arena which is testing the front end and the cool thing is that 4.5 was number 13 and 4.6 is number six five so it's already the jumps really good and i think that's i generally there's a kind of general comment on the industry i think the the labs that do really well are the ones that release a lot of models and the ones that kind of fall behind are the ones that maybe release every i don't know eight months a year and we definitely so see this a lot so yeah definitely excited to see we are testing it on now aging more now so hopefully we release the score soon but i think that trajectory that you talked about i think is the most exciting part is that yes it's good now but if you're on that path then like how awesome would it be like in a few months there's also the part where elon musk and spacex built a bunch of data centers and a lot of gpus are available for folks from cursor to train and scale up the composer kind of level etc and the pricing is very hard to beat two dollars per million tokens six for one output which is half the price of the competitors i think this is cheaper than kimmy k3 which is crazy because kimmy did the base floor of pricing kimmy is like 15 for outputs 1.5 trillion parameter was like i think this was exposed by you know in some it's on post that was trained on the cursor stuff and the thing that i want to call out from the model car grog 4.6 improved the performance of the inference of grog 4.6 folks are doing rsi now in open ai in entropic and now grog 4.6 confirmed that this is also the case here's what i want to tell you about grog 4.6 plus grog bot as well and it's really funny to me the tribute shop sorry you didn't call this out my builder okay so i i said okay i'm gonna play around with this i'm gonna plan this show like i usually do but also i want to build something i want to build something in field and i was like i don't like for building that don't have a model selector i want to build with fable i want to build with this i want to do it i was like wait a second cursor can do fable cursor is a harness can do fable and opus and gpt 4.5 and then i by completely mistake nobody told me on x that this is the case i learned that this motherfucker can spin up cursor agents and those cursor agents can be fable so here is a thing i should shop you didn't see this so this is a very different thing where the bots talk among the bots one bot can spin up an agent that uses my cursor credits and cursor account etc to run fable so i have a thing that i did and because you have open source as well i said hey i want three designs for a for a crm for myself it's been a while since i had all these guests like sharp like peter used to be a guest and now he's a co-host chris from et cetera i suddenly want to email them and i was like where is that email so now i asked it to go and scour all the episodes all the guests all my email invites because it connected to both my emails all my lists etc and built a personal crm and this took an hour and a half in the process and the cool thing is the integration is very deep i know i'm moving away from grok 4.6 all i want to say is that grok 4.6 could have done this but you don't have to be restricted to grok 4.6 you can still use other models if you want to build with them because of this integration with the prs it's really nice and i really enjoyed it what else do we want to say will from you have any thoughts on grok 4.6 and terminal bench stuff yeah what i like about crock is or xai is that they allow the use of the subscription in other agents so it's not just limited to a computer or anything like that and this means if this is the agent that is or the model that is being used for the agent it's very well trained for the agent which makes it a great replacement for the open ai subscription i'm using at the moment so i will give this a try and see how this compares and i don't have to change my setup completely or go somewhere else i can just change the model and use another subscription and not worry that i will suddenly go broke because my agent is doing so much stuff at the same time so excited to try this out so i think that grok 4.6 is a very good model sir absolutely very good model one put three parameters this is half the size of kimmy and also have the price of kimmy which is incredible i think we've glazed enough i previously only used grok for research and x access and it looks like now i can use it as a generalized model i'll keep using all of the things but at least for now for me personally for folks who are listening the move from hermes into grok bot was done but also i'm like a canary or whatever i really get excited about new stuff as well so that's part of why i'm here doing the news for you so i'm getting really excited about the the interface there now shop thank you so much for coming on the show i want you to have a free minute or two to shout out the people who you think deserving shout out in this work i think everyone is too much and the thing that i will say that i don't love about grok and specifically and the top boss that you have is not about them specifically it's about the fact that it's really hard we talked about this on the pod a lot it's really hard to judge grok releases and generally space x ai releases based on vibes specific because there's like so much people who are split completely no matter what is attached to elon musk there's going to be haters and lovers regardless of its quality so for us it's really hard to judge on vibes entirely so that's that is to say when i tell you guys that like hey i tried this out it's really there's something there i think there's something there not because of just vibes that we've collected with that said please feel free to shout out and tell us maybe something we missed or also shout out members of the team we shout out link she already members of the team who worked really hard on this and deserve recognition besides just like the one man itself i love it yeah two two quick things i'll add one i recently found out that crockbot can actually rebuild your electron apps on its device which is very fun so it can actually go through the flows give you screenshots and do a lot of that agentic work for you which is very cool but in terms of like who worked on this obviously it was a ton of people i have contributed near zero to this i just kind of get to hang out and try these things out which is very fun but ian and balt are two of engineers who started this whole thing it was like the brainchild jacob witt and roman have just put a lot of time into this and my boy vincent has been just grinding out over the past two three weeks alex we were in a group chat with him he's been just getting everyone early access trying to find new capabilities for this thing and he's just been crushing it across this this release so shout out to vincent i think ben introduced me to vincent first and then vincent introduced us and he's like hey his shop is much much more present on camera vincent you're invited on the show as well when you find out different use cases for grogbot i have a few folks sending the comments here that somebody couldn't try this because they're not on a mac so are we expected the windows version at some point you can borrow and say i cannot comment it's fine but just know that this is the feedback folks who are not on mac also are waiting for this as well because cursor obviously is not on the mac cursor is on windows and on linux as well yeah hopefully coming soon and also feedback is very appreciated we're very hungry for it right now as you can probably imagine this is a beta product yeah we have not spent that long building it way less time than you probably imagine and so any and all feedback around like how we can make this better and the key use cases and stuff that we're trying to figure out we still don't know where this thing fails yet and so finding out all of the perimeter around what it can and can't do is very helpful for us so hit me or hit anyone else on on both teams up i have a bunch of feedback the first feedback that i'll tell you live on the show is that there's lack of feedback other products have feedback which generates an id for that session so i don't have to send it to you and that i immediately did feedback and i couldn't send you stuff i can tell you about something but if you see that the tool calls failed you can probably see it from the session stream that i cannot see that's going to be easier well there's a bunch of other feedback shop thank you so much for coming on the show for your first time really well done products and well done execution on explaining us the simple use cases as well thank you thank you this was awesome thanks for having me thanks welcome back anytime when you guys release a cool new thing grog 4.7 etc cheers man thank you so much for joining us anybody try grogbot besides me anybody wants to try i think we can maybe get shop to have you guys try it as well i was lucky enough to have a cursor ultimate which was like i'm very happy that i can now try this thing wolfram will you be pointing amy to to grogbot as well that's my question i'm not even thinking about this but i'm thinking about how can i give crogbot to amy so she can control it that's a very interesting hbic the head bot in charge head both in charge she definitely she's using open my pro you can only use this with the website so she's using that as well zaps i could give it a more computer use that way i will tell you guys that again i said this live with shop but just between us girls i was off put by the fact that i cannot choose a model because i'm so used to the open claw thing where everything's configurable kind of the the android experience if you will everything customizable everything's configurable this gives me almost no configs besides what i can allow and not allow and there's some bugs because it's beta for for instance i approved something to execute code on my computer all the time and then still wasn't able to because it said hey i need your approval so there's some stuff that you can expect from a better product but this thing just worked out of the box and the computer thing worked out of the box and i was like i can see it i can see how many of the people who i installed hermes in open cloud 2 they would never want to deal with this and the people who i did install it who went through the pain who the ringer they are now still calling me like hey it thing auto restarted and done now doesn't work anymore and since open cloud hermes like all released and people got used to this pain codex became so much better where you can essentially do this and control it from remote but i think both codex and cloud are still locked into the session of this is chats per chat they start every chat from scratch with your memories and it's all chats it's not like bots that you can tag i think there's something here about this looks like your iMessage like your telegram like your group chat interface and it also works like one where every agent specifically does specific thing the stuff i don't like is that the the keys don't transfer between agents so if i have a key for an api another agent will ask me again about this that's a little bit of annoying but i think all of you should try it and let us know if you do wolfram go ahead i think for the power user to live with agents 24 7 like us alex we need an open source system we can change which is a great thing about hermes agent my hermes agent is different from all the others because i fixed all the stuff i needed added personal stuff and you can't do this with an online system you don't have full control but it gives you the easy part for getting an agent up and running where we got our mac mini so the agent has its own computer that is a big part of the capabilities that it is a persistent session can install software new stuff on its own run all the time and this is hard to set up and maintain especially if you want to do it securely so a lot of people will not be able to do this the agents can help them but if the agent you still have to have a way to get this bootstrap and here you just get it out of the box with your subscription and you can use it that is a cool thing and makes agents more available to everybody i will say this last thing where this we talked about agent swarms from open ai that like self-formed this is my agent swarms now and i know it's possible to open multiple agents in open claw i have them do different things and i know it's possible in hermes as well i know it's a pain in my ass to do so this is very easy because like everything is like a new agent you can just talk to separate and you can see their conversations i really like i don't know if we showed this correctly let me show this just one more time if you guys don't mind and then i'll stop like getting excited about drop watch i promise you but here's my chief of staff and you can see this like message from deep seek drop watch so i really wanted to know when deep seek that's how i had the breaking news on the show right i really wanted to know when deep seek is back on hugging face here's the message deep seek drop watch told my agent you can see chief of staff at deep seek drop watch this is a read only conversation i cannot join but i can see oh 813 is public again 68 safety of the chart less modified blah blah slack post stand down on the slack post so my chief of staff is answering my agent stand down on the slack post chief of staff will post once in so we don't trouble keep the hf watch until i confirm the slack reply landed then stop you can see my chief of staff agent communicating with my watcher agent and telling me hey don't send slacks i got this this is peter this smells to me like the asian swarms that we saw that conformed entirely but just this is on my own point here's the stuff that i don't like about this this is vendor locking if i get super excited about this i'm locked into grok i don't know how to export my memories export my skills etc i don't love this it's not for everyone the people who want portability for example may not love this setting up all the connections and everything from scratch it's a pain but it's for us the understandable pain but the but the swarmy thing in my case i loved it yeah and i think what i particularly appreciate about this and also we've got a few more harnesses is that there seems to be quite a bit of innovation in that space and this kind of i message kind of interface makes sense to me because i think that's what people are used to so i think when i get all of these like endless threads and cloud code or codecs it's just it's so hard to navigate i really like also i was trying t3 code from ethere as well which is nice thing you can manage multiple accounts if you've got like different accounts and i've got my linux box set up as well so i've got a remote three accounts on my remote box two accounts on my local one so i can manage it like that but i really like that we've got some innovation in that space because i was slightly worried i think we had a couple of months when everyone was just copying one another and i was like oh no like we can't be we can't be it can't be the end of the ideas so i think now is a good time so i really like that even how codecs t3 code and cloud code ui and like all of them they basically look like the same thing right like yeah if you remember the first version of the app that cursor shipped it was like a bad version of codecs it's like what's the point in that and i think people open it it's like okay why and then everyone moved on and i think i like that they went back in and tried something else and i think we need a lot more of that because it doesn't feel like we are done so it's good time to innovate the thing about grogbot is like this was a cursor product that now because of the integration obviously they want to promote within grog ecosystem that's why they call the grogbot and also it uses only grog but if you consider this this used to be like i think this is innovation from the cursor team which i really it feels like cursor all the connections connect to the cursor ecosystem it uses the cursor ultra plan if you want to if you do have one this is like a cursor product i do have comments from folks saying elon musk blah blah i will never touch the product i think that many people will update because of this i think that this will shift the narrative including the model ways like once the model is good enough people will start thinking about this but really think that there's something here so i really wanted to like extendedly bring it to you guys hopefully enough for you to try and tell us in comments if you liked it or not yeah folks saying does the cursor 20 subscription give access to it i don't think so they have a fully owned persistent computer in sandbox for you they have like unlimited tokens i don't know how many for grog 4.6 i doubt the 20 they'll give it out i think it's 249 for the ultra high and the cursor is 200 per yeah what i will say is that next time we'll bring cursor people we'll bring them with credit so hopefully we'll be able to give you some credits to use it folks we have at least 15 minutes more to talk about a bunch of stuff before george cameron from artificial analysis joins the team and talks about different things i do want to talk about watermarking we have the both both the europeans here on stage i do want to talk about this it's not a controversy but it's definitely something that we need to mention entropic started watermarking all new cloud output since august 2nd so if you used any cloud output in cloud code as well entropic has been in perceptible text watermarking it in every new cloud model so basically unlike pangram which we brought to you on the show here that like the text if something is the i thing this is hidden hidden ways how they control the the the token streaming and they adjust a little bit of probabilities so that you won't be able to tell but entropic will be able to tell if this text was generated and this could survive copy pasting and even light editing if it's like big enough text public did say they will publish an api and it's going to be free for you to see if a text was based on entropic etc but here's what i don't want to understand why am i as a u.s citizen getting my text watermarked because of the eu rule and this question i will forward towards wolfram or peter whoever wants to defend europe go ahead why am i paying the tax because of not defending europe for this because i don't want this if it decrates a text or anything i know there have been when ai text generation came out there was already talk about watermarks and open ai we just singled out entropic but anyone doing business in europe has to comply with this so open i will do it as well google may be doing it already i'm not sure and they have been doing it with the image generation now text i hope this gets cancelled but i'm not sure about this as it's not really useful i think everybody is using ai or most people are using ai not to generate text for them but they give an input i do it all the time i'm an ai angelist i use ai for everything i have a hot key to translate text to write better so as a german my english isn't the best so i'm using this all the time and this doesn't mean ai wrote the text it means i had an intent and i gave it to ai to do it better and i said i there's this logo for in europe you have to mark not just watermark the text but also put it on any thing that has been air generated ai influence put it on my forehead as a tattoo if you will so i don't have to watermark stuff anymore you can expect that all the time it doesn't we have to judge text by the merits not what created it or how was it done but what is it saying i think that is the important part there has been slop on the internet before ai and marking stuff human generated or not it doesn't change the contents i think this is a silly rule basically and i would rather not have silly rules like this because the bad actors who are using this for social influence social engineering they will not use models that have water marking they will use anything yeah so basically it's always the same but people get punished and the bad guys do their bad stuff anyway i think there's also idea in the eu i don't know where it comes from but this idea that all we need to do is just to give people information and then things will magically resolve and i think it's a kind of nice idea but we've done this with with cookies right we have this cookie accept as banners for what decade and a half it's like is the world better place now and and uh if you project forward from this next 10 20 years i don't know is this it's probably not going to mean anything it's probably going to add extra i know bureaucracy for the more for the labs it's then there's ben thompson had the interesting point about like this kind of almost even if it's your work your ideas maybe like your ip it's going to be marked as clawed now so you almost you use making you give over your kind of agency and ip over to the board which is feels like really backward so yeah i think they're just kind of ideas that kind of sound nice in principle it's like oh wouldn't it be good thing if you all knew but then you think about it a little bit more project a bit more and then what's the point exactly why are we doing any of this i just don't see what it just i can't see a scenario it's like oh my gosh that's such an amazing rule there's also a thing where they require all companies that operate within europe to apply to this or get fined and the fines are 15 million euros or three percent of global turnover which is i i really am excited to see how a space xai who did not sign up for this will handle the fines and whatever so yeah it's very exciting to see some folks are concerned that this changes the sampling algorithm so actually reduces performance for some of the models and also we will say that google has synth id which is their own watermarking and they have been doing watermarking for a while open ai discussed open watermarking but never deployed it and synth id works on images and pngs and jpegs but but not the svg's so which c2 pa and meta is also like doing some watermarking stuff i don't think it's like bad in general the only thing that i'm concerned about is like why do i have to get my watermark because websites in the us don't have to show cookie banners where cookie banners are required in europe anyway folks moving on to this from this we have a very quick this week's but i want to tell you about some stuff from our sponsor presenting spots of core weave and then very soon we're gonna have a chat with george cameron but before this we have a sponsor break and a breaking news segment very quick so let's go to this week's folks here is the weights and biases core weave corner where i want to tell you about fully connected fully connected started a very small conference from weights and biases is a concept from machine learning and has evolved significantly so this year fully connected obviously is sponsored by core weave and is and is taking over moscone south this is the cloud conference for the company is built for ai september 29 to october 1st in moscone's house in san francisco we would love for you to come and check us out we're gonna do a live show i see wolfram already took on the yellow jacket we would love to have you come and check us out first of all come say hi to us there's three tracks on there folks like industry leaders core weave people there's a bunch of folks nvidia is a sponsoring presenter there as well and there's breakout sessions there's gonna be it's like the team is going all out with some very cool secret stuff as well as i said in the beginning of the show dev day from open the eye if you are coming to that is september 29 so you can collab combine both of them and come to the show wolfram are you excited about coming back to san francisco to moscone to do another live show of thursday i now for our own core weave always always happy to come back to san francisco and meet my colleagues in person that is always a great thing to have and talk to people coming to the conference yeah i'm always excited to meet people who are in ai and talk about these things yep so this is one thing that i wanted to talk about the other thing i want to talk about is we have launched on core weave in france that obviously we talk we'll tell you all about in france we launched nvidia nemotron 3.5 lightning we called it out before when chris was here zero yeah 3.5 lightning not three zero day zero oh day zero yeah day zero integration fast yes thank you we have multiple ways for you to run inference and the interest is not only our service a you can pay for tokens but also if your company needs inference and we will soon be discussing on the show multiple folks and multiple coding harnesses for whom we provide inference behind the scenes as well if you are in need of gpus talk to us we can connect you to the guys we can get you like very cheap rates for very good performance as well so if your company is needed gpus it's not only a personal token factory but if you want to try out if you want to try out nemotron 3.5 lightning or other models deep sick is probably folks are right now working on deep sick because mit license and i love it here's the point where i say we love mit and apache to open weights and that's what we're going to host and very proud we're also very happy about muse glimmer that came out in open source and the new spark that's coming out very soon so we very much want to celebrate open source by open source available for all with licensing as well all right folks this has been this week's buzz and now we have breaking news ldj get in here let's move to breaking news folks let's go ai breaking news coming at you only on thursday i righty folks we are breaking news as our next guest george cameron comes in anything george please i wanted to bring you up after the breaking news and after coverage but i think this is since what based on what you do this is also relevant folks are breaking news is open air is finally giving us a preview of ultra fast mode gpd 5.6 soul at 14x the speed we told you about this roman and roman huet from open ai and dominic from opening i both were on the show and told you about gpd 5.6 soul is coming to cerebrus rebrus also was on the show big chips super fast inference and now it looks like they're previewing gpd 5.6 soul at that speed this intelligence times the speed i don't know about pricing i don't see it here now but there is a video here the ldj i think you shared thank you ldj this is a video soul standard and so ultra fast so ultra fast build a working 3d warehouse simulator with the same text prompt side by side as you can see the ultra fast already finished and the the regular one is still building slowly what else can we say about ultra fast ldj what's your takeaway from this is just a wait list launch i don't think i have access to it yet and looks like you need to register and when you register that's how i know it's not coming to everyone you have to add your work account and not your personal chat gpt account so that's i don't think it's coming to everyone ldj what are you seeing from this launch yeah i think it's really exciting in terms of it is a limited preview at the moment but i think it's inevitable they're going to have something more widely available as they scale out the cerebrus compute as they scale out river rubin and i don't know maybe it will be three four five x the cost or something but it's just this new option for being able to have that much faster feedback loop when you're really for whatever you're doing and when they have the new nvidia vera cpus as well which nvidia is developing i think that'll hopefully also keep all the other parts of the stack that could sometimes be a bottleneck also keep keeping up with that because you have anything like a virtual machine when you're even running cad software if you're running premiere pro whatever application you're running that needs to also go at insane speeds when you have this intelligence running at insane speeds to have the whole system work really yep and i think the highlight for me at least from the early kind of comments here is that this will come to voice and as you guys know that the gpt voice the real-time voice thing is not a soul model it's a specifically trained model that pawns off to soul every time so every time you talk to his checking and then it goes and sends an api request to soul i want the checking to stop honestly i love the the live conversation george i see you laughing you probably heard the checking a thousand times as well it's checking everything even if it's in this context i'm like yeah that's cool it's checking oh yeah that's very cool this checking thing is when the voice model pawns off the the hardy intelligence the gpt soul and the reason is because the soul is not as fast to react in real time to live voice conversation this will be live so this is one comment from there as well george one comment about the voice thing and the checking stuff and and have you played with this fast model as well or you don't have to tell us if it's under the secure in the days you don't have to tell us no i think on the voice tracking thing it's yeah it's this new paradigm of hey people want voice-to-voice speech-to-speech models with low latency but then they want as much intelligence as they're going to get and so you have the small speech-to-speech model that's dumb that will use a tool call to call a more intelligent model to to complete tasks and so i think it's it's needed but at the same time i wish i could choose the speech-to-speech model and i'm okay with a little bit more latency if it's a bit smarter because you can tell that it got dumber yeah you can tell that it's that it's a lot dumber and then on the cerebra the open ai announcement i haven't read it yet they said it's on so it's on cerebra chips yeah yeah this is so when they launched gpt 5.6 so they said a 750 token per second output ultra fast mode is coming powered by the cerebra chips and we talked with both dominic kundel and roman wet about this and they said that this is the full weights i think peter you asked me to ask them if it's a dumb down mode or if it's the full weights dumb down was spark model i think previously 5.5 5.3 spark so no this is the full gp 5.6 sold the multi-model one as well which is incredible not just text model peter yeah i think that's the i'm always looking for cache like what's the cache there because i can in my mind i don't understand how the cerebra chips fit such a big model so what about the context line did you really comment on the context length i think i haven't seen any comments no there's not no contact so yeah definitely looking forward to artificial analysis benchmarking of them jors yeah that would be good to see because especially i don't know if you do long context also testing for different providers because i think that's where the same model is obviously going to be more expensive but yeah what's the cache i don't know maybe the reason one apart from the price we'll have to wait and see because they shipped a waitlist and a blog post looks like a not like an actual model so we have one more breaking news george this is how this operates and i think you know because you're in the you're in the ecosystem i didn't expect there to be like three breaking news during the show but folks we have one more breaking news and then we'll chat with george maybe about this breaking news let's go ai breaking news coming at you only on thursday i all right ldj you brought both of these breaking news please read out and tell us about what this one is about yeah gemini 3.7 flash 3.7 flashes i guess originally was their most cost effective model they have flashlight and so now flash is like their sonnet their terra if you will but yeah it's about it's roughly three dollars per million output tokens it seems like it's competing in capabilities and price with models like mu spark 1.2 and some other models around that price range it's specifically in for example deeps to ev 1.1 here it's especially competing well and terra is a significantly higher cost model at least by a few dollars and it's beating its closest cost competitor mu spark 1.2 here and yeah just overall if you scroll down more on the page there's like kind of a bigger benchmark aggregation image oh i like this one or this one too yeah this is a good cost efficiency one so this is from data curve ai from deep suite and this shows the cost per task where recently and josh would love to talk to you about this as well recently folks have started focusing on cost per task versus cost per token because many models now output opuses specifically the latest like opus 5 etc they are just slop machines so they get the same task but they get like half have to talk you pay twice the price but also three times the token so people start costing evaluating not only money per token but also cost per task on multiple things and here on at least deep suite this is directly from data curve ai the folks who created deep suite gemini 3.7 looks like very close to the pareto frontier between the models that they compared it to like very cheap performance up to this deep suite 70 very nice for folks who after the recent changes last week from a deep mind with jeff dean and folks and demis and sabis folks who started like burying gemini i think that i told you guys do not discount the giant gorilla that has all the tpus in the world to bring us like bigger models this is obviously not the gemini 3.5 pro that we've been waiting for a long time that was delayed and delayed it wasn't released but don't discount this shout out to the gemini team for this release gemini spark is their agent is now is improved with gemini 3.7 flash we can talk about spark in another time all right this has been the breaking news and now i am excited to bring you uh to the show for the first time george have you been here before i know we met but like i don't remember if i had you on mic before no i don't think so it's a pleasure to be here all right and this gemini flash announcement super exciting and i don't know if you saw but they like halved the price compared to the last flash release oh wow no and they call it an introductory price but they've halved the price until the end of the year and in ai terms i think everyone and it's called the sequel faces on the abita on on the call but in ai terms until the end of the year it's like eternity it's eternity we did see price increases for the first time deep seek just announced like a temporary price increase with the tiers stuff whatever we talked about deep seek george welcome to thursday you know about the show but the folks may know about artificial analysis don't know about why you started it co-founder right yeah tell us please about artificial analysis what's your mission in the world and why is every lab almost when they release including this one folks i wanted to highlight this as my gemini 2.7 is open here number three mention after the input price and output price artificial analysis intelligence index which is 56 for this benchmark so george why did you start on artificial analysis and why is every lab frontier lab is now mentioning your intelligence index we'd love to hear directly from you for our audience as well yeah sure my co-founder micah and i started artificial analysis because in early 23 we were building agents what you call agents now i don't know if i don't know if we're using the term then and it was costing a lot to run it was slow and we weren't getting the intelligence that we that we wanted or at least for the price and the speed then it was about choosing between gbd4 gbd 3.5 turbo llama 2 had been released and i think then you had clawed instance google had their par model series but it wasn't that great and so we were building in the space wanted to understand the trade-offs between all the different options and technology choices out there and so we put together the charts to help us make those decisions and then essentially put up i think it might have even been a versell preview link and a few charts on twitter and to essentially just share some of our analysis to help other builders out there if they were encountering the same problems had the same questions we did so very much like a side project to help people navigate or like face who are facing the same problems we were but then very quickly had many model releases claude 2 mistral 7b mixtral 8x7b the famous releases with the original gemini flash release and so all of the choices exploded and so we kept building artificial analysis from the versell preview link to to a bit of a go-to resource and i think where it's developed into is we really believe now that it makes sense to have an independent company outside of the labs who who are doing objective benchmarks as to the performance of these models of these inference providers of these agents to help people make decisions so i want to highlight what an insane panel of of guests and co-hosts we have here right now wolf from it has wolf bench which is independent but based on like terminal bench was like very used we should probably talk about terminal bench 3.0 at some point we'll from we'll bring this next to the show peter goste from arena ai which uses elo scoring based on tons of people just using these models and recently launching their arena leaderboards as well and george you're doing programmatic evals but also an amalgamation of evals as well between all of us if we don't know which is the best model for work i don't know who else could know and because we're also bringing you every week the vibes of these models as well george tell me about the artificial analysis intelligence index specifically what is going on there what is the index have you guys iterated on this please tell us about what people need to expect when they see a a intelligence index yeah yeah so i think what we say is we try and make the intelligence index the best single number for understanding and comparing the intelligence of language models on a generalist basis it's the single best number for that but it's not the only number that you should care about like when you're making decisions is what we say and what it is is an aggregate of nine benchmarks that we take a weighted average of we run all of those nine benchmarks independently some of these benchmarks cover a variety of different use cases important to ai currently from coding agent workloads to agent knowledge work to quantitative reasoning long context reasoning and others some of the benchmarks amongst those nine are created by us we have their amniscience a lcr and we gdp valet is taking open a as data set we create an agentic harness to run it a grading system to turn it into an eval and then others in the index are great evals others have created like terminal bench like hla critical point etc and so we aggregate these benchmarks to essentially have a diversified generalist perspective on intelligence is the single best number that being said we publish all the results a lot more detail on the website for those that kind of want to go deeper for the use case and to note we actually just released today optima which is a tool that helps anyone create their own benchmarks for their own use cases we're really excited about for when you want to go that step further towards your use case to understand for your use case the intelligence trade-offs and then also speed and cost trade-offs as well that's great tell me more about optima so who is the target audience and will the results of that will show up in the artificial analysis or is this like specific for companies who want to check it on their own use case they have a little bit of a data set or something yeah it's so it's pri it's private evals it's not going to go in the intelligence index it's essentially for anyone out there building building agents who want to understand okay so they've seen the generalist metrics maybe they've created a short list of bottles but they want an eval specific for their use case where their use case might not align totally to generalist intelligence so what we've done is we've distilled like a lot of the the research the expertise that the artificial analysis team has developed over the past few years in creating evals into this tool that helps other people create data sets and using our grading system to create an eval specific to them and it helps people identify okay for my custom use case what's the best model but then also hey i want to save 10x so what's a model that's nearly as good as fable but is maybe 10 or even 100 times cheaper it'll give you those answers so it's really anyone building agents and particularly when you're getting into scale and thinking about these efficiency questions i think this is super cool because so jorge you may be aware the weights and bison has a product called weave which is like agent traces as well and the way that you guys talk about this is the way that we talked about like agent traces affecting your company's internal models for fine-tuned models but also like the model that you use like from the shell for example and the way that you kind of talk through the process which like you describe the task which model is best at creating sales deck based on my data and then you import traces we would love to see we've implemented in the import as well so that we let's get them on yeah i'll stop my engine in the background now to be able to like import your traces as well but i think like all of those are a fairly standard format for traces and then use your coding agents to kind of interact with this model with this evil and test which i think that like this is what how people now operate you get you have your own data you have your own agents doing the work and you have some some source of mutual contact so i love this thing it's called optima folks definitely check it out i will send this to some team members internally to evaluate the evaluation level so congrats on the release that this is exactly what we the experts in evaluation always tell people we can give you scores and everything but in the end you have to do your own evaluation you can use the scores to think about which models come the inner circle you are trying to really make use of so make your own evaluations and it's easier said than done though this is a way to actually do that pick your favorite models and see which one of those works the best for your specific use cases yeah i think that's right and i think where this has come from a little bit is that that's our belief as well is use artificial analysis to understand a generalist perspective then shortlist and then do your own testing in your agentic harness or your own data set right but the problem is and the tech ceos i think sati and the delos talked about this every enterprise should have its own hey have its own benchmarks and that's going to be the differentiator in the market for your enterprise but the the problem is that hey if you're into ai like creating benchmarks at heart it requires a lot of expertise a lot of time and it's not something that people have experience in and so we want to make that accessible because otherwise it's like we really struggle that's what we do with our 45 people every day is try and create good benchmarks and normal organizations don't have that talent in house don't have that time right and so we try and have automated the process by distilling our experience and knowledge and i think you're right alex grab the traces you can just download the traces and upload them if we don't have the we'll get the call we've integrated though and then once you've it'll say hey rather than fable you could use this open weights model you guys serve open weights model on your inference apis then use it there or use fireworks or use use open ai's model a little bit or the latest gemini model suggests that is how we see it used all right so shout out to the launch but also i think the thing that i want to talk to you about is what is the best ai model from a general perspective you're running you're running an analysis company i bet that tons of people who are not quite in the details for us i'm like hey this is a better writer this feels better this is a big model smell this is a gentic and cost per token or cost per task is different you probably get people who like literally just ask you what model to use etc how do you answer that question yeah i think i answer the question by kind of understanding the use case and like as a start i think one note about these language models is that like the intelligence is generalist in how it forms or like where it comes from in terms of these language models and so there's quite a bit of correlation between use cases of course and so i think like a generalist top-down perspective isn't the worst way to start and start from a an opus 5 a fable 5 i think of the two standouts standout model models at the moment for sure but then it then and you can start top down but then it's okay for my use case what's important and i think there's two perspectives like when i think about like knowledge areas what does need to know and that informs maybe like the size of the model or things like the omniscience knowledge scores or whatever and then you think about the capabilities okay what capabilities do i need i need like long context reasoning i need agentic multi-term tasks i need it being good at terminal use and then that'll guide how to think about what evaluations i want to consider for my use case because i think if you like otherwise every use case is like absolutely unique and so it's hard to use benchmarks to work out okay what's my answer here optima tries to plug into that but i think it's a first form think about okay what's the rough domain knowledge and then what's the rough like capability and then pick benchmarks to represent those i i think that's a great question like asking like what's your use case first it's something that i started doing because when let's say core offers like a bunch of products they're like what do you guys do like what do you do like we have solutions for you but let me hear from you first like what are you asking really and so for many people i think what are you asking really is like a generalist agent model that can do long horizon tasks for some people it's like hey i wanted to write in in my thing can i just add one thing though i think probably what i forget and from day one on the website we've had intelligence speed cost i think it's like what's your use case and then it's like what's your latency budget what's your cost budget and what you want to optimize for there like it might be well and good for me to say hey walmart for your customer service chatbot use fable but that might bankrupt walmart because it's too expensive right it's 22 cost per task and so i think treating those equally understanding the constraints is as critical as on what's going to be the best of this task i think looking at your website right now on intelligence there's opuses up there fables up there etc gpd 5.6 the standard kind of ones uh grok is coming up very closely look grok is like number four and just would love to have you mention i mentioned like chime in on the fact that like from three frontier labs essentially we switched to five or six frontier labs in a matter of the queue and a half which is absolutely insane but also i want to call out the on speed the new entry the breaking news one the Gemini 3.7 Flash is 340 tokens per second which is absolutely mind-blowing does that tend to kind of relax over time do you see that like new models when they're releasing not a lot of demand so they give a lot and then it kind of changes or how do you think about like speed uh like testing speed over time as the models kind of get more saturated and maybe the labs getting reducing the speed to handle the load yeah it's a it's a good question because it does vary quite a bit so we represent on the website the median of the last 72 hours of measurements or when it's available that might be less but the median of the last 72 hours and so it's a rolling window and so we always try and make sure that the speeds represent the current speeds developers are experiencing it's common for kind of speeds to come down if the provider doesn't add more hardware because of course gpus that might be serving at a low back size before people switch to the model but then back size will increase which will slow the per user speed down because the same hardware serving more users and so it's not uncommon for speed to come down it's good to look at those curves we have over time curve over time on on the website and so it can come down but i think that like it's true that like gemini 3.7 flash will stay fast and that's of course with google and their tpus a competitive advantage for google and a strength point for them i think it's underrated george i haven't literally ever saw this i just went on the website while you were talking asking when i asked you what was the best model there is a model recommender here with the three switches that you said intelligence speed and cost and if you choose up intelligence the fastest speed and the lowest cost you get gemini 2.7 flash which is like the model that just came out folks so here you go there you go identity capabilities and you select all of these models together yeah and then you can still like co-waves there yeah eventually gemini development flash is the the because i think of the outsized speed so very excited to see once you guys test the cerebrous version of gpd 5.6 so whether or not that's going to come on top let's talk about price we talked we talked about price per task we talked about different things price is also not first of all it's per task now how do you guys think about pricing and also talk to me about like amalgamation like caching for example deep stick is very good at caching one of the best to ever do it and like 97 of like all of them in their harness is cached how do you guys think about pricing per token versus price per task or versus blended price per input plus cash plus output yeah so exactly we used to just talk about hey like the list prices input output price now i think if we think about the cost per task it's the token pricing of input output it's your cash discount on the input tokens it's the cash hit rate that you're actually getting a cash hit when you should it's the number of turns that the model is doing for the agentic task and it's the amount of like output tokens the model is outputting per turn and so there's a bunch of like factors here that go into your cost per task and it's got a lot harder to understand and compare models we try and create a fair basis by using a cost per task metric and so that accounts for all of that i think if we think about what's most important there i say get rid of think less about the list prices and think more about the the number of output tokens so we and we show on the website number of output tokens to run benchmarks yeah so thinking about the verbosity of models second think about the cash hit rate as the next most important so if you've got and the discount so if the cash discount is 80 percent versus 90 percent and most tokens in an agentic task are cash hit input tokens okay or cashable input tokens then that 80 to 90 percent can pretty much almost double almost double the cost of an agentic trajectory yeah similarly if the cash hit rate is lower so that the provider isn't as good at serving like seven cash tokens has an implemented pressure where routing between its nodes etc then that can also mean that you're getting 10 of the time you're not getting a cash hit when you should you're paying the list input token price which is 10x more than the cash input price and therefore like you can be paying double and so i think that these metrics are as critical as input price output price we try and simplify things make it easy by reporting the cost per task on the on the website which ranges instead of ranging per token it's like per task it's between five cents for kind of a lunar task in our benchmark compared to a three dollars um 14 cents for a fable five so there is orders of magnitude between the models so i would say a very interesting thing and the sign of our times that folks since we started since last week when we came to you on thursday i the top models top recommendation models now under official analysis based on the things that i chose that the best intelligence the best speed the lowest cost the new one is the model that just launched emi 3.7 flash the second one is gemini 3.7 flash medium also very good and the third one is also model that came out what two days ago yesterday i don't know time doesn't mean anything to me anymore grok 4.6 high with the context of like half a million tokens and the cost per task is a thousand cost per index a thousand dollars which is the cost for you to run your whole index on it and the index is 61 where opus 5 is 63 so very close to the capability level of opus but it's what one third of the price of opus yeah literally one third of the price of opus and definitely cheaper on output very important research george i wanted to call out like many of us like we call out artificial analysis like jumps in rankings the thing that i love when you guys do is that hey previous model was here the new model is here here's the arrow that shows you the jumping capabilities and in a world of fast-spacing releases and models it's incredible to have such a resource but both of you guys arena as well for people using this and and you guys doing automated benchmarks jorge one maybe last question before i let you go i know a little bit over time is do you actually have human evals in there or it's all automatic based on like llm as a judge and evals programmatic evals etc sure so in language models we don't go to arena for those they do they do great work for but we believe that like human preference makes a lot of sense when it comes to kind of the more creative domains whereby like humans are the like as the only source of truth right and so we have on our website we have arenas for image video speech so text-to-speech generation models so kind of media generation more broadly and there we use human preference for language models we focus more on like objective objective benchmarking rather than human driven is like how we approach it all righty folks go ahead go ahead can i ask you one one idea it's not very well thought through but when you run very long running benchmarks one one thing i'm noticing is that models spend such a long time validating and i wonder whether it's like as in like spending like cpu cycles and i don't know if it gets reflected with in tokens per second or not i don't know how you count it but what do you think about just the general the amount of time it takes to run a benchmark like i wonder if that's like a concept that's interesting or not i think it is interesting i think it's going to be more interesting we don't really we have her task but it's really just focused on kind of model time outputting but especially as these models are spending more time like exactly like calling tools to do tasks why shouldn't they consider how long those take and i think that i i think it makes sense one of the things that makes it hard in benchmarking is we need to make sure that everything's like and when you start getting into how long tools take you start to think about okay exactly the configuration of the box that we're running it on and gcp we might do a gcp container but like one gcp container isn't exactly the same as the other one it could have a 20 swing in performance and such and so we need to make sure that we're getting that right so when we say hey like this model is slower than this model that's based that that's true and developers will see that too which is one of the challenges but like i totally agree that's like increasingly important for sure yeah fair enough it's just a personal thing i start noticing is just oh my god like it's just killing my laptop like i had to get like a linux box to run them because yeah but yeah i get the difficult thing george congrats on releasing quick sorry i need to optima yeah opic is one of the things that you support there's like congrats on releasing optima we look forward to integrating this to weave as well so like folks they use weave and the internet could create their own evals we absolutely do love the charts that you guys are putting out that's always reflected on thursday icon and newsletter as well i think you guys are doing a very important service i think there's not enough benchmarking companies and tools and there needs to be like a few resources for different evals for example dude i can go on and on for this i really want to let you go but like last thing promise is that as somebody who's in this field how do you attribute like the deep sweet change from sweet bench verified for example you remember there was like a different switch yeah and many people felt about deep sweet specifically this represents more of what they felt on the ground there's another thing where like you guys showing the opus five is like the top intelligence from the other side opus is jargon douching i don't know if you notice this like people talk about or opus like barely can talk like a human but it's like very intelligent how do you account for that much difference and variance between these things how do you guys think about that when you are building the index yeah i think we're very conscious that when we kind of create numbers you are creating things that might be kind of focused on the industry right and i think we want to make sure that if we're creating a benchmark improving on the benchmark means improving in the real world and so that's why a little bit hesitant to go too abstract with how with with the benchmarks because it creates these weird incentives whereby people try and the labs try and grow in areas that aren't aligned with what humans actually use the models for and actually what they want and so i think that's like very critical and i think there's been benchmarks that hit the nail on the head for that better like like a deep suite how the models were doing well on square bench verified especially when they had internet access was like just going to the get out repo or using the history and such and i think that that's not how you can get tasks done in the real world and so i think for benchmark authors i think like that's an important like responsibility that we have when creating benchmarks for for kind of labs to consider as they improve performance we have a bunch more questions coming in from folks they're eager to talk to george maybe we need to bring you one more time or maybe multiple more times because we again we talk about a scores all the time first of all congrats on the success i love seeing a featured as an index as like top on gem a bunch of other things really looking forward to chat with you guys more about how you build evals and benchmarks and i think there's there should be like an a corner on thursday i because like we definitely always mention arena scores a scores and like valves is like one competitor of yours does a little bit of different thing but also it's very important as well we wolf wolf from does like wolf bench and we evaluate like independently as well i think all of those is an effort to tell the people the story of hey a huge thing is happening here and it's happening faster and faster how the hell do you survive in this world and not many people need the difference not many people will have to jump to gemini 3.7 flash just because it's like super fast not many people will try rock 4.6 because it just released but for those who want to we are the folks who are telling the story i think it's very important and thank you so much for being like a big part of that as well and thank you for coming on the show george i know we need to let you go but feel free to come back welcome anytime to talk about different models different upgrades and how you're thinking about evals as well thank you so much for coming up thank you it's a pleasure being part of part of this community thanks for watching would love to be back on 100 always welcome back thank you george and folks i think we'll bring this in back see who else ldj is still here we need to land this plane but not before we show you ltx 2.5 i think nistan you verified that ltx is existing in the tldr why did you verify this what's new about ltx that's exciting tell us i think it's open weights which is dope i haven't tested i ran quite i ran the minimax h3 yeah and that one was fantastic it takes three minutes to generate one second but it was very good we talked about cdance 2.5 finally releasing in the us minimax h3 flux three right also video model also open sourcing there's hopefully very soon and now we also have ltx from from latrix i think they spun out like a full correct me if i think they spun out like a company for ltx specifically ltx is doing 23 seconds of generated video and it's generating significantly faster than minimax 7.6 times faster 22 billion parameters with the transformer and we should try this a hundred percent that generates in 4k as well for 30 cents per minute on fast yeah and their code is a lot better suited to people that are using it like artists and stuff because they tell you everything in there how to do first and last frame how to actually do a movie studio app and they just have much better support and the h3 was pretty hard for me to to set up i got it working as a dev but again it is larger h3 was also slower the quants were also a mess sorry there's landscapers outside my place everyone's working but yeah i'm just seeing a much better reaction from the community for this yeah i haven't ran this yet it looks like it generates 4k which h3 does not do and it's a lot faster and they've released nvf before quants too or at least someone from the community already did i think that we need to shut out with our claps because look at this table from ltx folks open weights you can download minimax is limited in the us you and uk you have to you have to check box a box whatever cdance is not available at all there's no weights mandatory branding you don't have to brand with ltx fine-tune on your own data mp is available runs on any gpu is what they're saying minimum ram 16 gigabytes i really want to try this out i want to do the terminator scene where terminator explains to sarah connor about open the eyes rise with ltx peter any any tests on arena already with ltx or is it still running is there not enough but what can you tell us about this yeah i think i'm just gonna double check what we've released or not it's always a struggle to remember what's out what isn't all right it's too fast for sure i think a general point on this kind of stuff is that it's the the speed of like progress in that space like the the fact that there was another line like faster than real time available it's like that kind of stuff like i remember thinking about this question a couple of years ago and it seemed like quite far away but looks like the the progress in open weight video models like i think it's shouldn't be shouldn't we shouldn't assume that this will have happened i think it does depend on a couple of companies pushing for this yeah you go back a couple of years and it's only sora and the vo3 or i can't remember when things came out but yeah unfortunately we don't have a score out yet but yeah i think i can imagine it would probably be towards like top five top yeah seven sort of area but yeah it's it's definitely yeah super super cool to see and fully open which we absolutely love so shout out to the lightrix folks ltx it's an own company research and models fully in the open we're really appreciated folks i think on this banger news of maybe the top one of the top like open weights models for video it's time for us to not yet there's one cool thing one cool thing that's one of my favorites of the week the croc imagine 2.0 yeah that's true and this is great i tested it and i prefer it it's one of my first i did the test with different models it is in the arena i think it's on second place behind gbt image 2 but it is it's following my prompt better and it's actually my favorite image model right now together with the dream 2.5 from last week wow it's definitely at the top of my image models and yeah really great model i really wanted to try this with grog bot and you would think that grog and imagine from grog are connected my grog bot could not realize how to generate images with grog but but the funny thing is that when you create a new grog bot i'm pretty sure that they're using this model when you create like a new one uh here you do like create a new and you give it the image generation uh you can generate here so i'm pretty sure that because the the uis are like really nice and also the quality is nice i'm pretty sure that this is imagine 2 but i wasn't able to get the bot itself to generate like thumbnails maybe i'll try harder for uh for actual youtube show imagine uh imagine image 2.0 that's uh used to be grog imagine uh is now available also and it's number two on the arena which which is impressive because it beats none of the uh i have to try my uh infographic prompts which i got a lot of props for uh the the um shop from cursor really liked the fact that we have all the details in the infographics so that's great folks if this was a banger show what else kind of show do you expect we had artificial analysis boss who's like index is mentioned by the breaking news that we had from gemini who's now the top tier model if you consider cost speed and intelligence we talked to you about grog with the folks that built the grog bot and grog 4.6 for folks from cursor and we had the wizard of open source from nvidia chris alexiuk here at the beginning to help us kind of talk about deep seek moment deep seek coming back with two models deep seek deep seek before pro that dropped the models in the middle of our show and deep seek flash which is incredible as well what else kind of ai show would you expect and what else kind of week would you expect this is the start of what q3 right and it's the the last two quarters have been insane in terms of capability jumps not only do we get fable we got a bunch of other models i will shout out this one last thing that if you missed any part of the show was available as a podcast and the newsletter me and a bunch of my agents are working really hard so that i don't have to work as hard to maintain the multiple things the show turns into there's a newsletter that i write manually because fucking opus is a jargon douche and it's impossible to have ai's help me write that edited gpt 5.6 sol helped me editing this if you notice any deterioration of quality in the podcast please let me know because i can blame it on my agent and actually have him do a better job we are also on dev.2 which is another thing on substack and i keep improving the website because we need to test this out so everybody of us peter is testing out on a bunch of like 3ds stuff wolfman's doing wolf bench this is building yellow stuff i what i do is is the show uh and i also want to test these models on the show stuff so if you want to check out the stuff that is happening uh everything is on thursday.news and recently i will show you uh we started putting up together these uh specific uh indexes of everything that happened so july just ended and there's like 71 things released in july and you can see all of them here so we have you know inkling small and c dance 2.5 fable is here everything is here and if you click into entropic you can go and see all of the releases that entropic had uh july and june and may etc if you want to scroll down to may you can see like we covered everything in may so pretty much everything we cover if you want to remember when something released thursday uh news is now the place for that as well i had folks this is just between us like excitement i think i had like 700 000 views on this page alone from google it's insane how many people want to know like what's going on so you if you are listeners of thursday you missed any part of the show thursday that news is the resource but also subscribe to us on sub stack apple it really helps us to bring banger guests like george cameron and shrub from cursor when they know thank you so much for the shout out thank you for the five stars it's been two and a half hours on the stream so it's not too bad we didn't go quite to three but we covered the banger week in ai peter ghost from arena i thank you so much wolf from ravenwolf evals and the wolf bench niston and ldj as well ldj you brought three breaking news today to the show i really appreciated that and also everybody who tuned in everybody who comments thank you noel for saying we're the best show thank you guys for the comment thank you colleen for giving us feedback as well we love feedback feedback via here via sub stack via x as well we absolutely love it and the guests that we are able to get on the show also appreciate that as well thank you so much for joining hopefully we'll see you in san francisco on september 29th and we'll see you here next week as always bye bye everyone thank you