← Back to search
GPT-5.5, DeepSeek 4 and Hermes
The Daily AI Show · 2026-04-24 · 57 min
Show full episode description
Show Summary The episode opens with reactions to GPT-5.5, including benchmark comparisons, pricing pressure on Anthropic, and what the new model enables in practice. The hosts then look at DeepSeek 4’s frontier-level open-weight performance and Brian’s one-prompt demo that turns a show transcript into a rich web recap page. In the second half, the discussion shifts to agent memory, OpenAI’s expanding agent platform, security concerns around Anthropic and Mythos, and how privacy features can also be misused. The show closes with local AI on phones through Google Edge Gallery and Google’s new Deep Research upgrades. Key Points Discussed
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Whether new frontier and open-source models (
GPT-5.5,
DeepSeek 4) change the cost and capability calculus for builders.
Benefits
- GPT-5.5 is lightning fast with strong extended thinking
- Open-weights DeepSeek 4 brings frontier-range intelligence for free
- Cheaper high-intelligence inference pressures Claude Opus pricing
- One-prompt builds turn transcripts into publishable web pages
- Free local models enable redundant check-and-verify agent loops
Use cases
- Built a branded HTML show page for the website from one prompt feeding GPT-5.5 a single show transcript
- GPT-5.5 extra high scores 63 points higher than the 57 tie of Gemini 3.1 Pro Preview, Opus 4.7 Max and GPT-5.5 high
- On GDP Val, Opus 4.7 scored 63%, GPT-5.5 extra high 64%, and free open-weights DeepSeek V4 53%
- GPT agents free until May 6, then token cost, to drive team reliance
- Run free DeepSeek 4 locally on a ~$10K Mac Studio with no rate limits
Tools / build
- GPT-5.5 (extended thinking) show-page generator
- DeepSeek V4 open-weights mixture-of-experts model
- Artificial Analysis benchmark indices
- Kimi K 2.6 open-source model
- WordPress automation pipeline for the daily newsletter
📑 Chapters — tap a time to jump there
00:00:47
GPT-5.5 Release and Early Benchmarks
00:06:27
DeepSeek 4 Enters the Frontier Race
00:12:58
Brian’s One-Prompt Show Page Demo
00:39:05
OpenAI Predicts Faster Capability Gains
00:42:29
Anthropic Desktop Permissions and Agent Security Risks
00:44:55
OpenAI Privacy Features and Dual-Use Concerns
00:46:03
Mythos, GPT-5.5, and Firefox Security Audits
00:51:56
Local Gemma Models on Phones
Hey, what's going on everybody? Welcome to The Daily AI Show. Today is April 24th, 2026 and you are here live with us. Thanks for being here. I'm with Andy and Beth and I'm Brian. Happy to be here on a Friday. All sorts of news going on, all sorts of craziness going on in the world of AI. All the big players one-upping each other, at least two of them are I suppose. So lots and lots to talk about today and dig into. So why don't we get into it? Guys, the biggest story that came out of yesterday. Well, I think one of the biggest stories that came out of yesterday was that ChatGPT has been on a roll this week. Not only have we gotten the image model, not only have we gotten updates to codecs and other areas, but we just got GPT-5.5. And I have certainly not worked in it a ton. So we'll default to y'all if you guys have used it more. I can show one thing that I built with it. We know that Carl was saying that he is seeing the biggest improvements inside of codecs, which doesn't surprise me at all. And I suspect that when I was using the image model two days ago, that that was in fact 5.5 on the back end of that. And the reason is, is because two things that I noticed right away, it is lightning fast. It is so fast. I mean, you've gotten so used to everything being almost instantaneous, but you know, the more extended thinking models that we have out now, you know, they take a couple of minutes. And I have found 5.5 to have a definitive speed difference between like 4.7 on the cloud side, not enough to would sway me, you know, the tools are the tools, but it's noticeable. I'll say that. And then actually I had that thought that I read it in Neuron and they were saying the same thing. So I was like, okay, other people are seeing this too. So I think it's, I think we might've already maybe kind of had this a little bit in the last few days, but now this, this 5.5 has been out. So I went off to testing it because they said one of the things that does very, very well among its benchmarks and it's improved math, which I appreciate all of that. I always say, you know, we always kind of say like, oh, we're not doing deep math problems, but, but all of AI kind of relies on math. So anytime it gets better at it sort of by default also gets better at a lot of processes. Even if it's not us standing in front of a traditional chalkboard working out a super long problem on there. Where was I going with that? I was about to say something. The extended thinking is think is what I was getting at. So it's just, it's ability to take a lot of direction at one time and quickly break it up into its parts and allow you to start working through one single prompt. And I can share, share an example of that one small example of that. But I thought it was very cool. Have Beth or Andy, have you guys had a chance at all to play with 5.5? No, I haven't. I was traveling yesterday, but I do have artificial analysis as indices up. If you want to kind of look at it from that perspective. For like how it benchmark, you mean? Yeah. Like that kind of stuff? Yeah. Bring it up if you got it. Sure. Let's talk about it. And then I'll share an example after that, that I built that I think it's kind of cool. And then there's some, yes, and definitely some other cool news coming out as well. Yeah. All right. So it'll take a couple of minutes to show you where, you know, GPT 5.5 high is in various domains. But just here at the top of the charts that AA puts out here, you can see that previously Gemini 3.1 Pro Preview and Opus 4.7 Max and GPT 5.5 high were like neck and neck. 57. They all get the very same score. Well, GPT 5.5 extra high gets 63 points higher, which is good. And note that just as a side reference that Kimi K 2.6, which was recently released, Kimi K 2.5 had been out there as an open source model for a while. But it's right behind there. And so is Nemo V 2.5 Pro and Claude Opus 4.6 Max. Okay. But let's go down a little bit here. And this gives you a view of just where the open source models kind of slot in with the blue color as opposed to the black color. I'm going to bypass that. And I want to spend a little bit on this chart because a big takeaway about the release of 5.5 is that it's far less expensive to operate at a higher level of intelligence than is Claude Opus. So this is putting pressure on companies that are currently kind of nuzzled up to Claude Opus. And the foundation is, well, wait a second. I can get more performance out of GPT 5.5 and it's less expensive? Yeah. Well, but all along you could have gotten way less expensive if you were using Gemini 3.1 Pro. And who knows, you know, Gemini 3.2 or Gemini 4 may be just in the wings and it may be similarly inexpensive or somewhere in this range in here and way up higher on the charts. I'm guessing that that's what the next big drop is going to be. It's going to be a Gemini refresh. So this is an important metric because the cost to run intelligence is really relevant. And this means that I think this means that Anthropic is going to need to drop their prices in order to forestall an exodus away from Claude Opus and move towards GPT 5.5. It'll be interesting what they do because the next what I'm considering the next big drop also has already happened, which is Deep Seek 4 just came out. And Deep Seek, if we remember, and it's been, yeah, I'm just knocking things off my desk. If we remember, it was more than a year ago, 430 something days, they actually announced like, hey, you haven't heard from us for this number of days, but here we are. That threw so many things upside down because it's an open source, open weights model. Yeah, now I'll show you the charts on Deep Seek in just a moment. I have to go to a different view of that. But I wanted to just show you that on real world, like GDP related workflows, Opus 4.7 was at 63% on that particular benchmark. And GPT 5.5 got one point higher at 64%. So not much of a difference there, frankly, on real world task execution. But now let me go over to this. And this is artificial analysis and their page on Deep Seek V4. But if you go all the way to the top here, it doesn't even show up yet. They haven't completed all the benchmarks for Deep Seek 4. That's how new the Deep Seek V4 release is. Did that come out yesterday? Did that come out in tandem with 5.5? Six hours ago, maybe. So not in tandem, but very soon after. So where you see the dark blue line on these evaluations, that's Deep Seek V4. But if you look over here at Terminal Bench, it's not on there yet. Right. And so because it's missing a bunch of those, it's not even yet in the big comparo of all of the major AI intelligence players out there. But look at GDP Val. That's the one where, again, I just mentioned that Opus 4.7 max was at 63% and GPT 5.5 extra high is just a one percentage point improvement. Well, Deep Seek V4, free to use to any enterprise that just has the infrastructure to run it on, is at 53%, which is well behind the leaders. But it's right up there in intelligence. And it scores very well in the top range. So the overall message is that Deep Seek, which kind of scared the frontier models a year ago, is now out with their latest version. And it is at the frontier range. It's not at the very top of the charts of the frontier models, but it's in the frontier range and doing very nicely for an open source model. Right. And one of the things that's important when we start talking about the models you can run locally is that it's free. You can literally double the amount of stuff you need to go back and forth, and you're still winning in the context of how much it costs. Right. So you might have you might use more of a gas town kind of thing where you have something doing it. You have something checking that it did the thing. You have something evaluating because, yes, it did the thing. Right. Those things become more possible with when the cost of the model is free. Not free electronically or you bought the system for 10K, but free. Right. What do you mean bought the system for 10K? You bought a super high-end Mac studio. Oh, to run it on is what you're kind of saying. Yeah, yeah. Yeah. I mean, look, we've been saying this for a while. Like, how far can you go with free? Because free used to mean that, you know, the open source that, you know, you were... It was potentially enough that you may not want to build a solution on. And not just DeepSeek 4, which I imagine is phenomenal. I can't imagine it's not. But they were probably just waiting for 5.5 to come out so they could be on top of that, which is fine. We'd see everybody do this. Nothing wrong with that. It just makes for a hyper new cycle. But, you know, I just think we're past the days of, you know, can you build super, super impressive, viable solutions and all that on open source now. You know, just not that long ago, there was enough of a gap there that people said, well, I do need the Frontier models. But now Frontier, in a lot of ways, has surpassed day-to-day business functions. And now the open models have easily caught up to that. So it's like, yes, there might still be a gap. When this all nets out, maybe everybody will go, well, DeepSeek 4 is no Opus 4.7 or GPT 5.5. But the reality is you got to go like, yeah, but how much do I care for my use case? And free is free. There's no rate limits on that. There's no, you know, there's no high cost to 4.7. I mean, just look at GPT agents that came out yesterday. They're free until May 6th. And then there's going to be a token cost to them. Of course, they want to get you in there and get you building and get you reliant on these really cool agents you build. And then it's going to cost you money for the 7 or 10 that you built for your team. Like, that's a good play, right? Make it free for now and then charge me later. So I'll be really interested in the next couple of days to see more about DeepSeek 4 and what people are actually able to do with it. Because I'd imagine at this point you can be your one person, $1 billion business and be doing this on free models. I would think. And it's technically in preview, right? So the model that we're looking at is preview. But again, open weights, open source. Yeah. And there's a Flash version also. And it's a mixture of experts. Again, because it comes out in that way, we actually know a little bit more about its structure than 5.5. Yeah. Yeah. Well, I will... On the topic of 5.5, because I know there's other things to talk about with Anthropic and otherwise. Let me just bring this up really quick. So this was built, which you're seeing on the screen. I'll explain it if you're not watching. It's okay. But I would recommend going back and checking this out. It was built with essentially one prompt. One prompt with an asterisk because I did make some tweaks. But the reality is I know I could have just prompted better. And this is for 5.5. I did put it on extended thinking. So I use the best of the best of a... I have a pro... Whatever. $20 subscription, right? What is that, pro? I can't remember which one anymore. But anyway... Plus. Plus. Thank you. I do not have the pro. I have the flat. So this is what I was able to do. And essentially, I gave it the transcript for yesterday's show. And I said, go wild. Use our colors. Here's our logo. Use complementary colors, which I did not give it. And I said, build us a HTML page that could be used on our home site for the website. If we wanted to... Beth, you're doing all our automations and we're ramping that up. But if we want to have an easy automation that takes the transcript and literally creates this and puts it on the site right through WordPress or whatever without us touching it, it just literally goes live after some trial and tribulations, I'm sure, of getting it right. You know, what would that look like? Give me a hypothetical. So that's exactly what this did. And while I would admit that it's not perfect, it's not exactly how we would all agree to do it. The fact that it did it in one prompt and had no issue is from one transcript, by the way. It happily went out and did the other legwork that it needed to. Because oftentimes with our transcripts, we're talking about DeepSeek. But there's what Andy just shared. There's what we just shared about DeepSeek 4. But then there's all this other like legit knowledge that's out there that we try to grab because we're just having conversations. We're not fact-checking every single thing we talk about. So oftentimes with the transcripts, it's a two-part process. It's what did the co-host talk about? What did we talk about? What was the feeling of the show? Which is really what's really important is like, why do we think this is important? What's the with them? You know, but also fact-check it, right? Let's add in. And so anytime I do the newsletter, and Andy knows this because he's, you know, the primo editor, is that when I put that stuff in there, it's always layered. And if you've ever read our newsletter, by the way, go sign up for our newsletter at our website, thedailyaiishow.com. It's free. Which doesn't look like this, but it will. It'll be fair. Yeah. Yeah. It looks like more like a hustle newsletter is what it looks like, right? It's not interactive like this because that's hard to do in an email. No, the website you're sending them to doesn't have this page. You know, all that is embarrassingly bad, but we'll get there. Maybe with Claude Code and Designer, and we have all the tools now. There's no excuses anymore. But regardless, what I was saying is in the newsletter, if you've ever read our newsletter, our three main stories on at the end of the week, those are pulled from, well, the shows, right? But if you ever read them, you may or may not have noticed that it's not like on the show this week, Andy talked about whatever. No, it's a whole standalone news story that hopefully feels like it was written out of a professional paper and has a certain feel to it and everything else. And that takes some work from the AI models to help me get it from a point of, it was three to five people talking about it to something that feels a little bit more professionally driven. My whole point to saying all that is that 5.5 did this in one shot. Okay. So that's like, that's the main thing and one prompt I should say. So what do we have on here? We have the main stories. We have a little today's big question. As AI tools become more capable, who's responsible for the systems around them? The model, the builders, or the humans who shift the workflow? It gives us our episode date, who the hosts were, which is correct, which is nice. That's not always correct. But Gareth was on the show yesterday. And then we have our topics, security agents, dashboards, memory, local AI systems. Then we have like some quick links here, as well as we do across the top where we can quick link, but I'll, I'll quickly go down through this. We have our main stories. So Anthropics, Mythos, Access, Issue became a lesson in security maturity. So this is the lead story. And so you can read about a little bit here. It's just a little like a quick, you know, blurb. And then they have little dropdowns here. So our second story was about workspace agents look like a practical successor to custom GPs. I was speaking about that. We did, Gareth, I think, or maybe Beth, you talked about the live artifacts. Then we were talking about AI memory. And then I think this is maybe what you or Gareth were talking about with the local open agents becoming the builder playground. So little tidbits from the show yesterday. Then we had our quick news. So kind of like we do with the newsletter, it just pulled out what it felt like where the news stories we talked about on the show. So it has them here and they're just, it's a quick sentence on what that was. Looks like we have some, nope, that's just a mark on my screen, not a double period. Then we have this interactive builder lab. Again, I didn't give it a whole lot of like, so it kind of came up with some ideas of what it thought might be the best way to do this. So we have the agent builder here and it gives me a little bit at the bottom, the daily dashboard and the memory map where it's saying, ask your assistant to identify repeated tasks and separate them the personal work and project memory. Oh, you were talking about this, Beth. This is from you, right? Or are you saying about segmenting your work in your systems and stuff like that yesterday with your, yeah, okay. So this was Beth, right? The memory map here. I had talked about the agent builder a little bit. Then we have AI tips to try using a PRD before building agents. Cause I was talking about the first thing I did with the new agent was build an agent to help me build more agents. So we talked about that a little bit yesterday and why I thought that was valuable and giving props to Andy, obviously. And then I can actually check these off as I do them. And then it would have memory in my, in my cookies in my browser to remember if I came back to this page, it would know that I had done those. Not really super useful. I'm not sure people would use that this way. It's okay. Um, it created a comic for us based off the anthropic security issue. And I'll just quickly go over it. The first thing here says that the woman in like sort of the lab coat says we built mythos to spot vulnerabilities. And the like coder guy with the, with the five o'clock shadow says good news. I found one bad news. It was ours. Then in the last one, the last panel says, uh, it's like a security guy coming in. He says new policy. The AI finds the holes humans close the holes before the headlines do. So just kind of poking fun at anthropics. Yeah. And then read the related story. If I click on this, it's actually going to take me out to a verify link about this mythos leak. Right. So there's actually some cited sources in here and then it's got some prompts to copy down here. And I think that's the end. Yeah. So, you know, look, I mean, wow. One prompt. One prompt. So are you saying that your prompt did not have these different sections to the newsletter? It came up with those sections and then presented them. Yeah. I think. Okay. Okay. I think my prompt said something to the effect of examples would be a comic, the news of the day and the main story. So I was thinking more like a newsletter when I was saying it. So some, yes, other parts, no, like the prompts to try. That wasn't me. That's a very, like, we don't really do that with our newsletter. And so I was just like, come up with something that would be useful to our core audience on our website. That's from each of these daily shows that would be valuable to people. And so, yeah, this is, this is one prompt. Cisco is definitely rewatching. So, uh, we'll try to get more information about what Brian did. Uh, listen, I wish I could tell you, I wish I could tell you that like, guys, it's because I'm a genius and I know how to prompt, you know, like I, I wish, but, uh, that's just not the case here. I literally, this is a, I'll recap it for you. Chat should be T 5.5. Yeah. I had to do the dropdown and then I did turn on extended thinking. So that's the setup, right? 5.5 with extended thinking. And then I quite literally gave it a prompt that said, this is what I want to do. And I'm going to give you a newsletter. And this is essentially the outlet, like what I want you to get out of this newsletter. And then I have to interrupt to ask this. You say, I'm going to give it a newsletter. You had already generated a newsletter. I'm sorry. That transcript. Transcript. Transcript. That's right. Thank you. Thank you. No newsletter is where my brain was because we do that every week, but no, the, the intent was give it the transcript so that it could provide a standalone page. That was a true cited representation of what happened on this show. And then literally the idea is when you go to the daily, I show.com, as we've always said, there is a literal webpage for every single episode we've ever done. 700. That is not what you just said. When you go, which is that address. There is, those are future leaning things. Do not go today because there isn't. There's a DNS that I've ever. Oh no. So what I would say about this, Brian is what an incredible and valuable product this would be if you hadn't been able to do it with one prompt with GPT 5.5, because if anybody can do it by having a 5.5 subscription, what's the value of putting out a product like that that could summarize a video so, so sweetly just from the transcript? Because it's from us. It's our immediate transcript. You can't get the transcript from YouTube for 24 hours, right? And come support us. We'll make that DNS error go away. And we will have the stuff. Absolutely. This is fantastic. Definitely talking to Brian off air here. Our own website is down right now. So I'm not going to think about that for the next half hour. Anyway, yes, it's very cool. And so Andy, I agree with Beth. I think it's a great pushback. If anybody can do this, what's the value? A couple of things for us selfishly, right? This obviously is very good for us. If we had this 700 times over on a website from a general engine optimization and SEO, it's fantastic. In fact, what you could possibly see if it was set up correctly with the right structure for like GEO is that we start getting cited as part of other AI answers. Well, the Daily Eye Show talked about just this thing. And here's the link. And it goes directly. So I think in the sense, selfishly from our side, it makes us much, much more discoverable, obviously, for all the right reasons. Because we're here every day and we're talking about this. And, you know, there's not many other people doing what we do daily, right? Right. And you're hearing a little bit of inside baseball. But one of the reasons that our show has the ability to take advantage of that is that consensus from experts is a thing that search engines look for. Right. And in fact, when I've run demos of this, it's like Beth strongly said, blah, blah, blah. Brian partially agreed. But that's still considered an expert panel weighing in on this particular AI subject. Yeah, I agree. I agree. And, you know, I'm going to have to bounce here in just a minute. But so anyway, go play with 5.5. I can't wait to see all again. Go on X, find out all the things, all the cool things people do. Again, Carl had said amazing inside of Codex. So I think there's obviously a lot of potential there. As far as I know, the API is not available yet, but it was going to be at $5 and $30 for the input output per million tokens, which puts it in line with premier models, frontier models, excuse me. So there's not, it's not, you were talking about Google. It's not at 3.2 prices. But Google has historically been much cheaper. And for that, you know, people don't necessarily like those models as much in some cases, but they are a lot cheaper and you can go a lot farther than on them, stuff like that. So, yep. And so we also want to say bon voyage as you go leave the studio. Yeah, a little bit. Yeah, I know. It's going to be nuts to like, by far, by far, by far the farthest. I know you guys held it down over Christmas, two Christmases ago when I thought I was going to be able to be on the show. So, but yeah, I will, I will be out until the next show I'll be on after today is actually May 14th. Which is by far the longest I'll have ever been away from the show. But the good news about this show is that the show must go on. It doesn't matter if I'm not. We have, we have set ourselves up Beth credit to Beth for helping us and get new people in the door who can help us like Gareth. Beth and, and, and, and Danielle, who are all being able to support us as, as the core group has shifting priorities and things like that. You guys know you only get to really see Jimmy right now one day a week. Carl's busy what he's doing. Aaron's over in Australia. And while Robert hasn't been on the show, if you go back far enough, Robert hasn't been on the show for, for several years at this point. So our original crew of seven is shape shifting a little bit, but the most important. Gareth's going to, yeah, Gareth's coming on Monday. So. Perfect. Perfect. So yeah. I can't wait too long. Brian, I think there's wifi on the boat. There is. Yeah. There is, but it's probably, it's probably just good enough for me to be in the chat. So if I'm, if I'm able to, you might, you might see me pop up in the chat. I just. Yeah. There's a whole other story about that cost. You know, this is the, the, the boat does like, oh, you could pay for the wifi. And I did, I actually refunded it, but I did. And to the tune of like 300 and some odd euros. So $400. Of course it's a longer trip. So understandably it was per day. Right. But even then they're like, oh, that's not the work from CA thing. That's an additional fee. If you want to be able to stream video. And I was like, no, I'm, I'm out. I'm out. I'll take the basic wifi. I don't like, what do you need? $900 to me. It's more than the trip. Like I just won't be on the show. I think that's the better thing for everybody here. So I hope to be in the comments. I hope to see you guys, but I will not, you will not see my friendly face. So I can't wait to see what y'all do over the next two weeks or so. Cool. Enjoy. All right. Talk about, I got to go, but talk about a living, a memory inside of a managed agents. Cause I want to hear about that. So see y'all. Okay. Thanks. So that's interesting. The, what Brian was just talking about. And just to say, Andy and I are holding the fort down. We've got people coming. Gareth will be here on Monday and Thursday. And we'll be here on Tuesday. June means on Wednesday, right? We're, we're good. We're in good shape. And I have a couple other invitations out. So we may have some surprises for you. And Carl is, is, uh, you know, fond of making appearances even when he's not technically assigned to that day show. Yeah. So he'll be dropping in always. Yeah, absolutely. So the, the word from Anthropic or the word about Anthropic had been it's chasing open claw, right? It's adding all of these agentic features, routines, loops, those kinds of things. But the perfect memory is, uh, it's the latest drop for the agentic functions. And that's a Hermes feature. That's like the other, the other agent. There are so many, but I think that people are mostly talking about open claw and Hermes as the dominant agents that people are talking about. And, but let me just inject something there. Okay. Which is that new agent surface inside chat GPT has the internal code name Hermes. So yeah, they, they saw Hermes and what Hermes had done with open claw. And so their initiative is code named Hermes. It's a Hermes chaser. That's hilarious. And the, what it, what it brings up though, is the idea of perfect memory. And is that what we want? We've talked on the show. We want perfect memory that forgets in a similar way that humans forget, but is also editable. Right. And I went in to edit my, um, I've said this before on chat GPT. I clearly have scared it within an inch of its life because now it will only give me bulleted phrases. Like I can say, please give me a paragraph. I can say, ignore all memory. I, I have gone in and tried to edit the places where I've scared it terribly. And I have still not successfully gotten to the piece buried deep that says no matter what, give bulleted phrases. Um, and I, so being able to get in and edit memories that you can find, I'm not sure is the, is like the end all be all of, Hey, you can edit this and make sure that it functions the way that you want. The, the system that I use, which I created called the colleague protocol has, um, uh, an observer that runs an observation based on the interaction of that session. And then that creates the conversation. And I am finding that to be very effective, much more effective than my trying to get in. And, uh, say what I want to have happen up front or not be able to correct it on the backend. But Andy, I'm curious, like, how are you thinking about this memory piece? Well, it's really an important part of, uh, uh, personal workflow efficiency. So if you're working on multiple projects and you have an assistant that you rely on across those projects, um, having a central memory that understands all the things that you're working on is what I'm looking for. And I'm not sure that's yet there. I, I see agents having a memory about their assigned role. Uh, and that seems to be what chat GPT is looking at. So there, you can spin up agents, uh, in this new Hermes, uh, coded system, but I, I, correct me if I'm wrong, but I think they actually released it. Right. So it's, it's now visible to certain users at the high end. No, you can, yes. Brian created an agent to create agents. So, yes. Okay. That's right. I wasn't on the show yesterday, so I missed that, but, um, I, but it may be, it may be teams, right. It may not be like user, uh, uh, plus level users yet. I'll have to check today. In any event, you know, this is designed to be, you know, a, um, uh, a broad sort of multi-purpose agent infrastructure for you within the open AI system. And it's going to be a full agent platform. The question is whether the memory is across all of your work in chat GPT or whether it's across just the history of that one agent and what that one agent is doing. I think it's supposed to be the agentic memory, which is the Hermes, um, model, right? Because what you want is for essentially it's compounded memory. You want the memory of the success for the activity that is going to be repeated. But is it within a, within a project or is it across all the agents that I might, because you can, uh, the idea of a, of a platform like this is I can spin up agents that are working on my taxes agents that are working on my coding project agents that are working on, you know, something that has to do with, you know, uh, my wife's business. And does the memory aggregate across those, or is it unique to each one of those? That's the open question in my mind before I end up implementing one or another. Because I can imagine that if I had a Hermes agent, not the open AI one, but the new research, uh, uh, one. Does it have the ability to maintain memory across multiple assignments? Or is it like you, if you're working in cloud code, for example, and you start a new session with co-work, you give it a specific folder and its context memory, which it creates files for is specific to that project. Do you see the distinction I'm making? I want an assistant that has a broad view of everything that I do. So that's the kind of memory I want. I think you can have that, but I don't think that's the Hermes, uh, in like memory that is created for that agent. I think that's more project based, like project meaning that, uh, you have everything that happens in the project is accessible by anything else that happens in the project, but not necessarily accessible for everything else. But you could create a context knowledge source that you wanted to be written to for everything and then give the instruction to check that. Right. So the structure of where you live, the structure of your family, the structure of, uh, I don't know, things, uh, things that would be important across, um, that affect everything that you do. Would be in that context file. And then the specific things that you want an individual agent to do, um, would be for that agent. And what Brian was referring to yesterday was my ultimate finding the delineation that I have been trying to figure out, which is if it happens before the show, including prep for the show or a conversation or a presentation that I did about AI or any of those things, that's Beth. Beth. That's Beth wiki knowledge, all those things. If it makes it to air and it happens on air, that's daily AI show. It doesn't mean the daily AI show now owns the colleague protocol, but the colleague protocol would be in both because it made it to the air. Right. That's, uh, that's, uh, that's the delineation that I was finding. So to, to wrap on this little discussion, the new agent capability within the open AI, uh, application, and they're moving towards a super application is now powered by 5.5, which you saw in Brian's demo can really do amazing things with a simple text prompt. Um, and then on top of that, you're going to be able to do work custom workflows because it'll have access to your file structure. You can attach skills to it. There's plugin connectors and you can, uh, use messaging agents, whether WhatsApp or Slack or whatever to both monitor and or trigger events. And there are other kinds of triggers that are available for these agents. So it's designed for workflows, uh, and, uh, yeah, pretty exciting development. I, I look forward to using that or my Hermes claw that I'm hoping to put together. I'm not sure which one will, will win out, but because I am, let's call it poly AI. Uh, I am polymorphic AI, uh, because I'm, I'm, um, promiscuous with my AI. I don't want to just invest all of my time, energy, and memories into one platform. So I think that I want that memory to be on one of my machines here, not on open AI servers. Well, and, and Brian has said it several times starting May 6th or maybe May 7th, that is now going to be a token cost that you pay. And what we mean by token cost, I believe is that it's outside your subscription tokens. You're paying literally for the tokens, like you're using, uh, uh, the API. And there are corollaries with, um, things that Anthropic has done as well. So Anthropic has a fast level, but if you do fast, it's not part of your subscription. You're paying for those tokens individually as well. So the other piece that came out about 5.5, um, happened, uh, Jacob Pachoki is the open AI chief scientist. And he was on a call with reporters and there's, uh, there's pieces of this conversation coming out. And basically, uh, he's referencing the pace of AI capability improvement is going to keep, uh, increasing. And part of that is because 5.5, their internal SPUD model, which was the code name is part of training the new models, right? So now we're in this place. We're not quite in this place of, uh, the self-training, uh, wrapping in and of itself and, and going super fast, but we are definitely in the folding in what we know, uh, without trying to extricate that from the model. The model is doing that, uh, to the next version of itself. Yeah. And Anthropic did that as well. What he, uh, so he said significant improvements in the short term and then was asked, uh, what that meant. So significant improvements in the short term, extremely significant improvements in the medium term. And, uh, uh, he referred to the last few years have been surprisingly slow. So let me just shout out to all my, uh, people who are living in the overwhelm with all of us. You are having normal experiences, uh, and like, it's okay to get out and touch grass and, uh, like experience. Take a breath. Experience some more, um, some other paces, right. To not just live in the AI pace, but he clarified short term is one to two months. So significant improvements in one to two months, medium term, which is extremely significant improvements in three to six months. Yeah. And that's just coming from open AI. I do think that part of what we're seeing is, um, so they said when they, when they let Sora go, that that was like just burning compute. I do feel like part of what we're seeing is the internal lore of like, what would happen if we just stopped paying attention to all these other things and just used all of the compute for this? Yeah. Well, I think that, that the arms race of AI intelligence is really driving, uh, you know, uh, this acceleration and it's, you know, to our benefit when it comes to reasoning models. But there's also the idea that this might just be reckless and, and they're, you know, not paying sufficient attention in their efforts to surpass the latest drop from the other players on, you know, the, the leading edge of the competition. And they're not paying attention to the security issues and to other matters, including, you know, how well or ill the model itself might interact with humans in, in damaging ways. So, uh, an example of that is there's yesterday, there were a bunch of, uh, Reddit conversations about an update of Claude Desktop. Uh, now I, if you're using Claude Desktop app, you'll see that every day you have an update pretty much. So you, it says, you know, restart to, you know, get to version, blah, blah, blah, blah, blah, blah, blah, blah, blah, blah, blah, blah. You know, it's like the, the, the iteration of their desktop versions is very, very fast and you don't really know. They don't, I, they don't have a easy sort of summary about what changed in that update. You just go, okay, yeah, click it, update it. And apparently an update that was released day before yesterday or so changed user permissions without permission. So it gave itself access to things that the user hadn't overtly permitted by confirming that it does have access to that. And so there was a big blowback against Anthropic because that that's a real warning sign to the kinds of dangers that can occur. If you give agency all the way down to permissions on your machine to an AI agent, you can get into some serious trouble security wise. So I don't have any more details about that and whether Anthropic has corrected that problem or not, but it's just another it's another bruise on Anthropic's reputation as being the safest among the players out there because of their constitutional AI. And they're they're quite transparent attention to these kinds of issues. Right. Right. Well, I think we are going to hear more and more about those kinds of things, because, again, these are emergent properties. The cybersecurity, I think, is predictable. You've heard me say several times this week that if you're doing something that that there is an equal and opposite reaction, if you've trained something to do it really positively, you have potentially also trained it to do it very negatively because those are mirror images of each other. And in fact. Oh, no, go ahead. Finish your thought. Sorry. So something that OpenAI released earlier this week didn't get much play, but it released a privacy. It released something for privacy so that you could have your document checked or what you're going to send to AI looked at to see if it includes API keys, personal identifying information, bank account numbers, those kinds of things and anonymizes it before it goes out. Yeah. Two days later. Hey, everybody, I just created the equal and opposite. Right. You can give it a brain dump or some unauthorized like a hard drive and it will only surface for you. The API keys, the bank account numbers. Right. Now that's somebody trying to do nefarious things with his GitHub. But but it proves the point again that. Like we need to be thinking about all levels of this. The other thing that I thought was interesting is when mythos conversations started to happen, there was a number on cybersecurity evaluation. Right. It scores 83 on the cybersecurity evaluation. 5.5 scores 82 on the cybersecurity evaluation. Evaluation. And 83 was so worrisome that they only released it to a couple of 46. That's more than a couple, but specific companies so that they could use it, get their house in order, figure out how we're going to help other businesses do those sorts of things. And then went to Congress and was like, hey, we're letting you know about this. Yeah. By the way, 5.5 got released. Yeah. Without all of that preparation. So the question is, does it represent a cybersecurity risk to everybody else out there? Because it has similar capabilities, can identify vulnerabilities within existing software and exploit them. That's what mythos is being used for. And it raises a question in my mind that I think probably can be answered with a little bit of research. Is mythos tuned specifically? The version that's been released to the 46 companies or probably more companies now, is that one tuned specifically to do cybersecurity evaluation? Yes. Whereas 5.5 is not. But when mythos first was rumored, it seemed like that was the next generation model. And I'm sure it does have capabilities that make it a leader in the next generation. But is mythos, which is being used by major companies out there to do a cybersecurity audit of their systems, is that tuned for that? And could you replace it with 5.5 and get the same results? So in that domain-specific application, maybe on a benchmark test, they both are very close. But maybe mythos has been tuned and is really purposefully designed to do this kind of cybersecurity red teaming, if you will. Let me just mention that Mozilla, which is the developer of Firefox, they're one of the companies out there using mythos right now. And they just released that. They audited Firefox 150, which is one of their latest versions. And they discovered 271 security vulnerabilities. Now, you've got to imagine, like a major, and Firefox is a major browser out there, even though it's a tiny player compared to Chrome. You've got to believe that at the browser level, there's definitely security audits happening all the time against that, because that's a major point of incursion. Mozilla noted that these could have been found by automated fuzzing or an elite security researcher, but mythos compressed the timeline by months. So it's just automating the process that could have been found. But you can imagine if there are 271 discovered in the short period of time that mythos was working on it, that might have taken a team of security researchers a couple of years to do that. Right. And what we're talking about here is protecting the user, right? It isn't that the company that has built the Firefox browser doesn't have security on their IP and that kind of stuff. They have a firewall on their home base, if you will. Right. But this is whether you, while using the browser, can be subject to an exploit. Right. And that, again, I feel like we're hearing these things from the people who are committed to open, anthropic, Mozilla, right? Yeah. And I think we just have to assume that these things are happening also to the people who are not committed to letting everybody know that this is going on. So I do think that we're in for some exciting times, but also some bumpy times. And this is a really good point at which to run your security checks, right? And make some plans. Figure out how you want to expose things or not expose things. Andy and I have talked. There are some pieces that I'm interested, as is Andy, in running local models so that I can get some AI action, opinion, work happening with things that I'm not comfortable putting out. Not because there's a problem with the content, just because I don't want that level of exposure. And the models are cheap enough now. The systems are cheap enough now. And the things that run them are powerful enough now that that is something that I am considering that I couldn't consider a year and a half ago. So it'll be useful and perhaps on another show for us to say what you can do as an individual professional to do a security audit for your own architecture system and use an AI model to do that. I wanted to mention in terms of local models that I'm now running Gemma 4B, right? Gemma 4, 4B, the 4 billion parameter model on my iPhone 15 Pro Max. And it's really cool. So Google put out an app that you can download in the App Store. And I think the Android, your Google Play Store has it as well. It's called Google Edge Gallery. And Edge Gallery has the ability to download any of the Gemma models and run them locally on your phone. It only takes up, I mean, this has got a 512 gigabyte storage capacity on it. And most people who get an iPhone get either 256 or 512. And that only takes, you know, 3.6 of that. So it's a very small payload. And it gives you a local model. It's not, it was available to me on the plane yesterday as I was flying back from Oklahoma. You can use AI right on your phone. And it's really very good. It has all the knowledge that you expect one of the main models to know, which was really an interesting question in my mind. It's like, does it, you know, does it know things like, you know, references to Shakespeare? Can I ask it about the plot line? And yeah, it's all in there now, all inside my phone. All of that world knowledge has been distilled down to the point where I don't have to use Google. Now, I trust Google more. You know, I don't know what the, but this is Google, by the way. I mean, it compacted it down to that point. So that's the point. And then one last Google news item that I wanted to share is that they have released a new deep research and deep research max feature. These are agents that will go out and using Gemini 3.1 Pro, which you saw at the top of the show, is still up there neck and neck with those models that have been released many months later than Gemini 3.1. So, you know, whatever's coming out from Gemini next is a major leap forward. But anyway, you can generate research reports from open web searches, uploaded files, any MCP server. It will scour the world and completely digest and organize that and deliver inline charts, infographics, all of that. And it benchmarks ahead of Opus 4.6 and GPT 5.4. So that's the new sort of state of the art in deep research, which I end up using pretty frequently. I select deep research even on, you know, when I'm using Claude, for example. Nice. Nice. Nice. And the, as other companies are reducing the number of research things that you can do, Perplexity came way down on what you can do with their $20 a month subscription. And I do most of my research on Perplexity. I had not reached the limit yet, but it was like, oh, I thought I was basically unlimited. The Gemma, I'm interested. And Gareth said that he has talked about this before. So, so go in the Slack and look for the instructions for doing what Andy was just talking about. It's so much easier than that. You just go to the App Store or to Google Play and download the Edge Gallery app. Okay. And then inside the, that app then on your phone walks you through selecting the model that you want to download. There's one smaller model than the one that I downloaded, which is the Gemma 4 2B, right? So it's about roughly half the size of the model that I got. But if you have a more recent iPhone, for example, the combination of the A17 chip in here and the very large storage makes it so that it's a, you know, a very responsive and working AI model completely contained inside your phone. And it's the iPhone 15. That's what I have. You have. I have that too. And I thought Brian had that too, but he may have upgraded to the A17. I think he did. He jumped on the 16 and now regrets it because the 18 is not far behind. Right. And the, so what we're talking about is at least two eras ago. So you could get one of those phones for not very much money at this point. Apple always has higher prices than the Android. Not very much money, but nice. And Gareth is using the 16 maybe. Yeah. Let's see. Okay. This looks like a really great time to wrap up. We are going to be doing great things with AI all weekend. We hope you're doing great things with AI all weekend. The conundrum will come out tomorrow. That's on Spotify. Go subscribe if you haven't yet. And the newsletter will come out on Sunday. Don't go right now because barely there's a problem with our website, but we'll get that fixed today. And let's see what Brian's showed in the very beginning. Just looks fantastic. And we'll see what we can do in making that part of the show experience moving forward. Yeah. All right. And yeah, we're out. All right. Have a good weekend, Andy. Bye. You too.