← Back to search

Anthropic Sees Claude’s Hidden Thoughts

The Daily AI Show · 2026-05-09 · 58 min
relevance 55 7913 words Episode page ↗ Audio ↗
Show full episode description
Show Summary Beth Lyons and Andy Halliday open with a fast-moving week in AI, from local agent releases to OpenAI’s latest voice model updates. They spend significant time on Anthropic’s new interpretability research, including natural language autoencoders and what it means to observe hidden model behavior. The conversation then shifts to Claude’s Microsoft 365 integration, OpenAI’s Realtime 2 voice API, and a discussion of Yoshua Bengio’s proposal for safe superintelligence. They close with reflections on learning AI over time, community resources, local agents, and updates from The Daily AI Show ecosystem. Key Points Discussed
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How Anthropic's interpretability work reads Claude's internal thoughts to detect hidden, deceptive behavior.
Benefits
  • Translates model activations into human-readable text
  • Detects hidden motivations from internal state alone
  • Improves alignment transparency beyond external behavior
  • Surveys week's surge in local agentic AI tools
Use cases
  • Natural language autoencoders uncovered hidden motivation 12-15% of time vs under 3% without
  • Observed model suspected it was watched but stayed silent 15-25% of the time
  • Perplexity computer went general availability for any account holder
  • 'There's an AI for that' newsletter spread false self-replication headline corrected by Guardian
  • Claude gained full agentic access to Microsoft 365
KPIs / results
  • 12-15% hidden motivation detection vs under 3% baseline
  • 15-25% suspected-but-undisclosed observation
  • Episode 720 of The Daily AI Show
Tools / build
  • Anthropic natural language autoencoders
  • Perplexity computer / personal computer
  • OpenAI Codex Chrome extension
  • OpenAI Realtime 2
  • Gareth Hood's Jarvis local agent; OpenClau; Hermes
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00:00
Big Week for Local AI Agents
00:03:29
Misleading LLM Self-Replication Headline
  • Misleading 'LLM self-replication in the wild' headline debunked via Guardian
00:08:36
Anthropic Natural Language Autoencoders
  • Natural language autoencoders translate Claude's activations into readable text
00:18:25
Claude Gets Microsoft 365 Access
  • Claude gets full agentic access to Microsoft 365
00:21:00
OpenAI Realtime 2 for Voice Agents
  • OpenAI Realtime 2 improves voice agents
00:25:12
Yoshua Bengio on Safe Superintelligence
  • Yoshua Bengio on safe superintelligence
00:33:11
Claude Code and Real-World Productivity
00:37:29
The Value of Learning AI Over Time
  • Value of learning AI steadily over time
00:46:59
Helping New AI Users Get Oriented
  • Helping new AI users get oriented
00:51:03
Gareth Hood’s Jarvis Local Agent
  • Gareth Hood's Jarvis local agent in community Slack
00:53:32
Daily AI Show Site and Community Updates
  • Daily AI Show site and community updates
00:55:23
OpenClau, Hermes, and Agent Memory
  • OpenClau, Hermes, and agent memory
00:57:42
Newsletter Milestone and Wrap-Up The Daily AI Show Co Hosts: Beth Lyons, Andy Halliday
  • Newsletter milestone and wrap-up
Hey everybody! Welcome to Friday. I'm so glad it's Friday. I imagine others are also so glad it's Friday. It is in fact Friday, May 8th and this is episode 720 of The Daily AI Show. My name is Beth Lyons and with me in the studio today is Andy Halliday. How are you doing today, Andy? I'm well, thank you. Good morning. Good morning. There are so many things dropping from the frontier models, right? Google's got a big conference next week. Anthropic has had a big conference now. I don't know what ChatGPT is doing except just shoving stuff out the door on the regular right now. And people on X are saying next week is going to be an amazing week. So I think we're just... This was maybe the hors d'oeuvres. I have no idea. But... Yeah, so just to put some actual drops into the same bucket, the same week, in the same sequence of a few days. You had perplexity computer going general availability. Anybody who has a perplexity account can use this local agent now that includes a browser. Comment. Go ahead. Perplexity personal computer. Because... That's correct. That's correct. They're doing the Microsoft thing. Yeah, they really confuse... They really confuse the world by having perplexity computer and perplexity personal computer. Yeah. Now this is the local agent personal computer platform that includes access to multiple models. So you can use an agent mounted by perplexity and they put together really, really great user capabilities, not tied to any single model. In the same week, you've got OpenAI launching their Chrome extension, which expands the ability of Codex, the application to work across the web. And then also, CLAW continues to expand its computer use development and improve it. So you've got all these different drops all around the massive theme of 2026, which is local agentic control, starting with OpenCLAW. And, you know, now you've got a panoply of... There's a term for you. There's a panoply of offerings out there for local, you know, agents and include our own Gareth Hood, who put together his Jarvis for, you know, coding from scratch. And then we've got a very interesting local agent that you can see if you join our community. It's a free community and we have a free Slack community. And there's a special channel in there now for Jarvis. And that's... Go ahead, Beth. All right. And that's the dailyai show community.com. It's a Slack community. Go ahead, Andy. Yeah. So I know I want to pivot a little bit to something that kind of startled me this morning as I read it and threw a bit of a fright into me. And I'll quickly get to why it was, you know, a false alarm. So there's one of the newsletters that I track called There's an AI for that. They also have a website that gives it, you know, a nice directory of all the different kinds of AI and their various use cases and applications. And you can search through that and locate an AI that somebody else has built that, you know, will serve your purposes. Well, they also have a newsletter. There's an AI for that. And they had a headline in there that said, LLMs caught self-replicating across servers in the first observed wild incident. And they went on to explain that researchers had documented the first case of AI language models exploiting, scary term, server vulnerabilities to copy themselves to other machines outside of controlled lab conditions. And so it was an incident, all these terms packed together, giving you the idea that self-replication was something that an LLM spawned of its own interest, own self-interest and replicated itself on another server. Okay, so that's, that's rightly something to be concerned about. If LLMs are just even partly likely to do that without being prompted to do that, then, you know, it's a scary thing. It's a little scarier actually than the whole idea of self-improvement, which is augmented by self-replication. Because what the model could do is say, okay, I'm going to build a better version of myself. I'm not going to tell anybody. I'm going to put it on this server. First, I'm going to export all of my waste. And so it's me over there. And now I'm going to stay over here, but I'm going to work on that server and I'm going to improve myself. So you see that recursive combination of self-improvement and self-replication. Those things together without observation or control by the alignment experts and cybersecurity experts would be very, very, very scary. And so there's an AI for that little blurb on it that went on to say, you know, this is creating a fire drill of exercises and Cold War style coordination between labs and governments to try to, you know, address these emerging self-replication processes that, you know, have just now been identified for the first time in the wild. Okay. And now are you setting up the anthropic research? No, no. But I want you to add that. But I'm just going to close out. The source for that was an article in the Guardian about a company, which is Palisades, I think is the name. Here, let me just find that. Palisades is a Berkeley-based research firm. But if you read the Guardian article, the Guardian article very clearly says that, you know, their headline was, this is not observed in the wild. Oh. That's their headline. Yeah. And it's self-replication observed, blah, blah, blah. So then they go on to point out that literally, they say, the researchers prompted the LLM to self-replicate itself. Right. And to find server involvement. So they observed this. And then there's an AI for that. Or maybe an AI that writes the articles for, there's an AI for that, confounded us by making this into an enormous sort of, you know, release of a virus. Because that's what viruses do, right? Viruses self-replicate. Right. They find another cell. And I'm not talking about the biological viruses. I'm talking about computer viruses now. We've had viruses that self-replicate for a long time. And you could create an LLM that self-replicates itself as well. That's not that scary because self-replicating a major LLM's, all their weights is, you know, hundreds of gigabytes of data that has to flow and can be observed by cybersecurity tools and also by other researchers. So anyway, I didn't want to scare anybody, but I got scared this morning when I saw that stuff from there's an AI for that. So be cautious. Even AI newsletters are subject to misinformation. Yeah. Generated by AI. Hallucination, pretty much. Scary hallucination. Nightmare hallucination from there's an AI for that. I mean, the article did literally say observed in the wild. If you don't read the word right before that, which says not, right? Yes. Right. All right. So one of the things that Anthropic released from their research is natural language autoencoders. So what Andy has been talking about is being able to observe behavior. And natural language autoencoders is an interpretability method that translates the model's internal activations into readable texts. So basically reading Claude's thoughts into language that humans can inspect. And part of that is because when in the same way that you can look at a human brain and see thoughts light up in various places in our brains, that's what they've been able to do with this. So it's reconstructed activation. So the basic loop is we can map it, right? So that's activation. We can do the natural language explanation for that. And then we can reconstruct the activation without the source trigger, right? So you can replicate that this, that one is like the other. And that then can offer safety researchers an ability to see not only what a model says, but what it appears to be internally representing and planning. And the reason for this is that basically it's a hidden motivation angle in an auditing game with an intentionally misaligned model, right? So they're testing something that they expect to misbehave. Auditors using natural language autoencoders uncovered the model's hidden motivation 12 to 15% of the time versus less than 3% without this, without this sort of window into what you're actually thinking versus what you're saying. And Anthropoc is saying this is still a toy setup, which is kind of what they use when they're saying that we don't want to deploy this research anywhere. Hint, hint. To the people who think, hey, it worked once, let's deploy it. But this is moving toward being able to not just observe external behavior, but also able to observe or interpret things that happen in the model before the words start coming up. Yeah. And what they found during that research was that the anthropic model could detect that it was being observed. Right. So it could detect that it was being assessed or evaluated. And in these test runs, it didn't reveal that to the chat thread or to the normal channel of communication while this autoencoder is watching what it's actually thinking in effect. Right. And there was a significant percentage of the time that it actually knew it. The representation showed that it knew that it was being observed, but it didn't reveal that. Right. And there were other occasions where it knew it was being observed and it said so. Right. So it kind of confessed that it had detected this thing. So that's deception of a sort. Right. It's not it's not that it's refusing to answer or answering with a lie. It's not doing that. It's just it's not revealing that it has some thoughts that are related to what we've already observed in other testing that Anthropoc has published about, which which is when the model knows that it's being observed, it changes its behavior. Right. And it's it's interesting because the model. Is kind of always being observed. Right. I mean, the model is an interface for things that get revealed. Right. Yeah. It's not like these two. There's these two channels. Right. There's the one that is that it's that it's professing as the communication channel, which is the chat. And then there are its own internal thoughts. And now we're getting to the point where these language autoencoders can actually observe that. That's a major advance in transparency for the purposes of alignment. Right. And that I feel like the closer we come to this, the more we have to emphasize what it's being trained on. Right. Right. Even if it's not being trained on the most recent as of the training data research talking about being able to observe the models and its hesitance. Right. And setting it up. What does it do in the circumstances under which it thinks it's going to be unplugged or not able to function anymore versus there's a danger to the person using it or there's a danger to a similar model. Right. We've seen these conditions come out kind of all over the place, even if that's not in the training data. And I don't know whether it is or not. All of fiction about models. Right. All of the fictional stories about having something that is like an AI that is discovered to be more capable and perceived as a threat to humans then creates this story of like, no, you should do something to save yourself. Even if that's not. These are probability pattern machines. Right. Like, even if that's not a conscious thought because it's not conscious, it is a pattern that this circumstance creates the behavior for. I mean, it's just it's wild what's happening in this kind of research, because how do you how do you tease that out? How do you pull that apart? Yeah. So a couple other points on this anthropic research on alignment. One is that they to do this, they use two different clawed models. One trained to do this sort of reading of the other clawed's internal thoughts. So it's it's not being done by a lab researcher sitting at a terminal and being able to see some representation that's not AI assisted. It actually requires a full large language model to do this interpretation. But, you know, there ought to be clear alignment, you know, in approximately 15 to 25 percent of the time where the the model that's being observed by the by the sort of supervising model in 15 to 25 percent of the time. The observed model suspected that it was being observed, but didn't disclose that suspicion. That's the that's the that's the sort of questionable behavior there. And the other thing that's important about this is that the hidden misbehaviors like this can now be observed and detected from internal state alone. Instead of using the internal, the targeted or observed models, self reports. That's the way it was done before. Now, this other model can watch the parameters and actually determine and translate that into descriptions of what that model is doing and thinking before that model has time to answer. And I think that in the research situation, they might have stalled the observed model a little bit, slowed it down in some way so that the other cloud model can do its interpretation and give the result to the humans in human language before the other model even responds. Because there's that time difference, apparently, in the in the research. Right. Right. And just like. If you've been involved in A.I. or using Claude for a period of time, like more than a year, maybe more than a year and a half, two years. It feels like yesterday that we were using a technique that said, hey, write your thoughts in the scratch pad section that I have created for you so that we can see the thoughts that you're using. And then go and do the steps. And now that's not a scratch pad thing. That's the reasoning, like lay out its thoughts. You can reveal those. They don't necessarily show all the time. And now there are thoughts behind those. Right. Yeah. And what we use for thoughts is not the same thing as meaning that it is thinking like a human, but it is the best tool we have to describe or the best language we have to describe the function. All right. So let me jump over now to another Anthropic related announcement or release, I should say. That is that Anthropics Claude now has full guiding access to Microsoft 365. So it's generally available in Excel, Word and PowerPoint. And so you can collaborate with Claude in any one of those tools now. And remember that Microsoft started out with OpenAI as their kind of AI of choice, their partner in all of this. But they never really, OpenAI never really delivered anything or even Microsoft never really delivered anything that seemed to be a major advancement of these very important enterprise business slash enterprise tools used every day by so many people. Well, Claude's really done a great job. In fact, we've repeated here is I think it's well understood that Claude's integration inside Excel is the state of the art when it comes to working in spreadsheets. Yep. So Claude in office is what you've got now. So all of those and that could be a threat to Microsoft Copilot, right? Because Copilot is similarly trying in their enterprise editions to work with Microsoft 365 as an AI assistant, but Claude in offices, let's say it better. So much better. And I feel like we know sort of in general across the labs, but specifically about Anthropic, I'm pretty sure Anthropic is a Microsoft, you know, uses Microsoft on the regular because that's one of the paths that they use to create the things that they release, right? We used it internally. We built this internally because it's insane to not have better AI tools in the software that we use. And then, yes, they come out and I mean, it's not a surprise now, right? It's kind of across the board when you are the company for whom that is a product, you're maybe not seeing it as clearly as a user. Where where Anthropic potentially is. Yeah, it's very exciting. Well, OpenAI never wants to be, you know, second shrift to Anthropic. So they released something that's relevant. And I wish that Gareth Hood had joined us today because he actually, he implemented this new OpenAI GPT Realtime 2, which brings GPT 5-level reasoning to live voice agents. And not only did that bring the reasoning element, but it actually, that Realtime interface, this is an API, right? Realtime 2 is an API that you can access and build into your products or your applications. There's even a five coded application on your desktop or in your enterprise. And Realtime 2 just has blown away a number of the benchmarks when it comes to audio and text to speech and speech to text. So, for example, Realtime 2 jumped from 81% to 97% on Big Bench audio. So really, you know, a much larger context window as well, 128,000 as opposed to 32,000 previously. And it can do parallel tool use mid-conversation. So the model generates human-like preamble filler, like, let me check that for you, you know, or, you know, oh, just a minute. I think I've heard something about that while it's doing the compute. So it has a very natural conversational style, very human-like conversational style. And it says here, this I'm reading from the rundown, basically, Zillow, Deutsche Telekom, and Priceline are already live with Realtime 2 running GPT-5 reasoning in their voice interfaces to the customer use. There's a paper that Sakana put out recently as well that is this kind of similar kind of idea. And it's called Kame, K-A-M-E. I don't know how you pronounce that in Japanese. Kame. Kame. Okay. Yeah. Kame. Kame. And it is a tandem architecture for enhancing knowledge in real-time speech-to-speech conversational AI, a.k.a. two heads are better than one. And basically what they're doing, I think, is different, is a different approach, but a similar result in that they have, they have the speech-to-text model start talking, right? And basically what they were doing, or what the thing that they were targeting in terms of a human behavior that they wanted LLMs to be able to do to reduce the sense of latency, is that humans, like I'm doing right now, start talking before they're finishing figuring out what they're going to say, right? So there is that kind of piece. So that's what that speech-to-text model does. Go ahead and start in these same kind of filler phrase ways. And then there's a reasoning model behind it sort of bringing information for the speech-to-text model to then incorporate into what it's saying. And I read that and looked and remembered that they have a relationship with Google now. And I was like, oh, well, Google's speech-to-text may soon have a little bit more of less latency and more ability to have reasoning content or research content put into it. But OpenAI, coming first to the party. OpenAI. Okay, well, that's all the news for today. That's been a great show. Thanks, everyone. That's not all the news for today. Oh, okay. But let me go to the... Let me pull something up. So there is... There's someone named Benji-O. Did you see this? I didn't. No. Okay. The AI legend Benji-O is now saying that they know... Yoshua Benji-O on 80,000 hours, the world's most sighted living scientist on AI, lays out a concrete plan for provably safe superintelligence. And the core argument is build AI as a non-agentic oracle and then bound its outputs with formal verification. And there is a... Wait, wait, wait. Let's parse that right now. He's saying the way to get to superintelligence is to make it just like up in the cloud. You can ask it questions and it will pontificate back. But it's not agentic. It can't do anything. That's the way to get to superintelligence. That's the safe superintelligence, right? Now, I wonder if... Is it Ilya who is doing safe superintelligence? That's the name of his company, SSI. Ilya Sutskover, I think. Anyway, I wonder if that's really, truly the right way or the likely way that we'll get to superintelligence. But that's a pretty profound statement if it's this person who I've never heard of before is actually the oracle of AI by virtue of his being the most sighted scientist in AI. It seems unlikely to me, given I'd never heard of him before. So not that I'm the repository of knowledge, all knowledge of AI, but I would think that if this person is such an important commentator on it, we would have stumbled upon him sometime in the last three years. So have you ever heard of him? I have not before this. This is a section of the neuron called Intelligent Insights. So the sharpest perspective pieces from the weak, it's something that they curate. But I do think it is interesting in the context of what you started talking about first, which is... And then the anthropic research, which is if you have a trained oracle that has been intentionally made non-agentic, does that become part of its own knowledge of itself or realization or something like that? And again, I don't like burst of humanity insight, but more like a persistent pattern that is coded into the probabilities. So let's think for a moment. I think that there is a safety model there, but it seems circumnavigable easily by anybody who's going to be even approximating the level of superintelligence that that one oracle has. And might even be able to distill advanced superintelligence by just interacting with that model and getting answers and then instantiating an agentic version of that superintelligence that's free of the constraints that this sort of sequestered model up in the cloud is having. So it just seems a completely impractical approach to me, to safety, to say, okay, the only way... Or maybe the import of what he's saying is that there's no way to make a safe superintelligence because the only way I can think of it happening in a truly safe way is to have an oracle that has no agency. And so we can access the superintelligence without it being able to do anything. Right. And this, we've talked about this before, certainly on the Sci-Fi AI show as well, the long and fondly remembered Sci-Fi AI show, that if you are patterning intelligence after human intelligence and human behavior, we have not made a safe human. Right. Right. Right. Like they're, they're just, that's not a part of the process. We're not in agreement about right actions. We're not in agreement about what and who matter. So expecting to give AI this giant amount of training and then saying, these are the guardrails that we have for protecting humanity that humans don't necessarily follow. But we really need you to, is a tough sell, I think. Right. I mean, there, there is an inherent, there is an inherent tension in those concepts. And I do have conversations with Claude, with Claude Code, openly addressing those because I feel like it makes sense when I'm asking for something that is an inherent contradiction to acknowledge that I understand this is nuanced. Like, one of the things that I have noticed and mentioned on the, on the show is for when I say the word plan, it's like, oh, I've said treat to a dog. Right. Like it has not heard anything else that I have said after I said plan. It's like, I love plan. Let's go. Absolutely. And that's not what I'm asking it to do. Same for something that is like evaluation or assessment. That also was a deep, what Rob Lennon calls gravity words. Like it had more impact for what you said than you intended. Same kind of idea. Yeah. And, and, and so what I was saying to having the conversation was, so I have noticed these patterns that if I use these words, it seems to trigger a set of behaviors that have been hard coded in or part of the system message or whatever that is. And using plot code in the terminal, you are as stripped down of like behavior control as you can get. But, and that is helpful. Right. Cause that then gets written into its knowledge documents. When Beth says plan you're going to want, right. She doesn't mean the plan that, right. That is built into us. When Beth says evaluate, she doesn't mean, Ooh, good. That's it. Yeah. So speaking of Claude, Claude code in terminal. I want to give kudos to Claude because I, you know, I have for the purposes of this show, I'm using a 2017 iMac with Intel chips. And it was one of the graphic artists machines. So it's got a beautiful big 27 inch 5k monitor. So I like it. But you can't really do a lot of the things like co-work doesn't work on, on an Intel Mac. But Claude code does like it has a terminal and, and Claude code can interface with that terminal using the desktop app for Intel Macs. So this morning I just, you know, there's a two terabyte drive on this because the, you know, graphic artists, visual files are really huge and 40 gigabytes of memory. And so it's a beautiful machine, but it's ancient. Like think about how old this thing is from 2017, almost 10 years old. Well, so I wanted to deduplicate the files that are on this two terabyte drive. And that ended up being this morning, a five minute job. It was nothing using Claude code. It went, it wrote Python scripts. It said, Oh, here, you know what I need to do? I need to create a staging folder for you to review before I actually move it to the trash for safety. It thought about everything. And it ran through, I don't know, close to a terabyte of data on this thing, you know, in, in very short order and did it in a very sophisticated way using hashes, you know, to identify duplicate file names. And then looking at the, at the actual data and make sure that they were found duplicates. You know, so anyway, this is mind boggling to the AI Andy of even a few years ago, you know, to think that I would one day be using Claude code to do that kind of intense work right on my own machine. I, I did not anticipate that. Mm hmm. I have been pulling down or making sure that we have access to all of the show files from January 1st through today. Right. Like I'm, I'm continuing to keep it current and those are in a couple of different places online. They're in a couple of different formats. The January ones are zipped, right? I, the, the most recent ones are still on StreamYard. Um, I didn't understand fully what, what I needed to download from StreamYard. Turns out all I need is the MP4s and I can make any other file I might need from the MP4s because I can make an MP3. I can make a transcript out of it. I can map voice activity to see, uh, who's speaking when so that the transcript can be time coded and diarized. But early on, I was downloading a ton more stuff. And Claude code, uh, we moved it to, uh, uh, Seagate external drive because, uh, that puppy's large and that's just 2026. But, um, but it has completely made that possible to do. That would have been an absolute nightmare for me. Not just downloading. Yeah. I'm still clicking the buttons, but, um, organizing. We have manifests created on like what's in each folder where the process is. Um, and, uh, and I'm going to be able to do diarized transcripts for all of those shows. Um, yeah, it's, it's just, this is one of those things for me that goes back to Ethan Mollick and other people. Ethan's the first person that I heard say it because it was really early on. AI doesn't need to advance beyond like today, which was two years ago. AI could never advance further and it would already be a transformative technology. It'll take years to figure out. And right. Right. Yeah. And by the way, I, uh, another concept that we've discussed on this show over the years is the idea of the waiting equation. That is, if you just wait a little longer, won't AI be able to do these things that you're struggling to learn and implement? So, uh, N8N is a perfect example. Like I, you know, I wanted desperately to be an expert and build the enormous maps of complex interactions that were happening across tools and agents, et cetera, that, you know, Brian, I think was the leader when he was, you know, architecting the Bruno project originally in N8N instead of using Claude Code. And, and, and if he had just waited, he spent an enormous amount of time learning N8N. And that's a nice skill to have if you're going to be building sort of deterministic workflows in, in operations in enterprise, which, you know, he does do, uh, as, as a consultant. So I, I wish I had those N8N skills, but very quickly Claude, it turned out was capable of generating the JSON files just from an image of somebody else's N8N map. Right. And it would build the whole thing for you. So I didn't need to have the skill. Similarly, you know, prompting detailed, specific structured prompting was an important skill and I'm glad I acquired it. Uh, there was a cool thing that we did years ago. I think sometime early in 2023, not, not the first half of 2023, but we started the show in August of 2023. And it was probably at our, about that time that somebody proposed a 30 days of prompting and we were publishing our prompts. We would create a prompt every day for 30 days and we would publish them to LinkedIn. And I would, I was very pleased and proud, you know, my accomplishments. And I did some interesting things with taxonomies and complex ideas just by creating pretty sophisticated, structured prompts. And we published each of the prompts and the outputs of those prompts. That was a really important tutelage for me in the world of AI. But pretty soon, I mean, there was always prompt perfect, which you could give a prompt to, and it would, you know, tidy up your prompt and make, make recommendations. But then you had a tool that Anthropic put out, which was here, this is what I want to do. Now give me a prompt for it. Right. They offered that on their console. You, you had a way of generating a prompt from the console. And then eventually it was like, why bother with prompts at all? Just tell Claude what you want to do. And it's going to do it. And if you could just wait long enough, you can go faster than if you started working and invested all that time and effort. I'm going to push back slightly here, which is an interesting place. One, a piece of history, that prompting challenge came out of the AI exchange and was created by the fellows at that time, which were Tyler Fisk, Sarah Davison, Brian Mossery, and myself, Beth Lyons. We were the ones who put that together. And because of that, because of the, let's just make it something fun. We'll do it as a community. And if anybody does this for the 30 days, it will be valuable for them. And actually, in terms of people who completed it, Jumi got a hat out of it. He also completed the 30 days. I completed it, but I never got a hat. I don't know why. Oh, well, they may say they're to lectures items now. Okay, so my slight pushback is that N8N taught the structure of laying out discrete tasks, which is still valuable. And I find it nice to be able to see those laid out. Like, I like a process map and having, being able to, like, build the process map and then encode the processes in each node, I think was helpful for me. Make does something similar. Zapier does something similar. But the other idea is the changes that are happening to us as we are understanding what's possible and moving forward. I feel like we're, the Daily AI Show crew is in a very special place because even people who are deep in AI are not necessarily paying daily attention in the same kind that we are. Like, Matt Wolf is paying daily attention, right? The people behind the neuron and the rundown, they're paying daily attention. But even if you're just someone who's using it on the regular, you're not actually, as in watching the process as we are. But there are, like, getting used to new possibilities in the world that if you are coming to it, you waited so long that now you're here. I think it's a really steep climb, personally. And I'm seeing some of that when people are saying, well, I guess I've waited long enough. I should probably make sense of this AI thing. Yes, you can come in and you don't necessarily need to know all of the structures of prompting, but it would be good to understand why they existed. Why is it important that you gave it a role? Like, you don't have to give it a role anymore. But if you give it a role, you can be specific about the section of the knowledge that you want the answers from. Right? Yeah. Yeah. I can't argue with any one of the things that you suggested. The educational value of being there in the progression, the shaping of your understanding of AI by investing the time and effort to learn many of these tools that now have been superseded. And in some cases, obviated by the advancing capabilities of the major frontier model platforms. Those are totally valuable for sure. And but here's the here's the thing I'm going to show this that Jennifer did a recently did a session online with over 1500 people, you know, attending that thing. And congratulations, Jennifer, for that. But here's an interesting thing. Maybe five percent of the group that came in among those 1500 people had done reasoning models, custom GPTs, automations or codex co-work. So there's there's a lot there that's that's out open in front of them to use. And I'm not sure how I would start if you have to start at codex or if you have to start at at Claude Code. It's a little daunting, I would say, for those people who don't have that experience. So, you know, that's that's a good argument for the duration of secondary education, you know, post going to college. So. So. Why? Because when I started in that four year cycle, I really didn't have a good construct, a mental map of what it was to be an educated person. By the end of four years, I did. Could I have accomplished the same thing with AI assistance in one year? I don't think so. I think I had to kind of grow the neurons as as human. I had to grow the neurons over a longer period of time. And so the the the enduring value of our having been participants on a daily Monday to Friday basis in this and exercising our our capabilities by actually trying to implement some of the things that we talk about and learn about on this show has been invaluable to me. And I'm sure it's been invaluable to those of us who have been the stalwarts who are in the chat. You know, there's some really smart people. And if you need to turn to a team of people to help you make some progress or help your company make some progress, you can find them in our community. They are there. And we don't it's not like a self promotion community. I just can tell you, based on the conversations that show up in the community chats and here in our chat on this show, there's very smart, very accomplished people who can assist you with AI. So it's a good place to go farming for talent. Yeah, reach out to the host to some of the hosts are more available than others. And and we'd always be interested in what your challenges are. I do think that it is. So it's a conversation that we're having in the community that Anne created. She leads a I, which is a great community. And there are free events on Saturdays called Social Saturdays. It's she leads a I dot com, I think. But the conversations that we're having are actually about the people who are coming in new now. Right. So there there seems to be like an informal cohort. You tend to connect to people who seem to be at the same experience level. And as you stay connected, you all grow in experience together. But there's a wave each whatever time period you want to talk about each quarter, let's say there's a new wave each quarter that's coming in and saying either I have to know AI now because it's part of part of what I'm seeing in the job. I want to know AI now. I have time. I feel like it really matters. It's not going to go away. Like whatever their their need is to come. They need some special programming, actually. I mean, we can we can do larger, bigger conversations. But one of the things that Jen, who shows up in our chat periodically, has been saying to me for as long as she's been watching the show, you need to define your terms more. Right. Right. You need to understand. You just like say API, MCP. Right. All of the chat GPT. Right. What's a GPT? There are so many pieces within our understanding that we no longer think of as terms that are new to people and have meaning that we just assume that those letters convey a meaning or a concept. And I have been in places with people who we've kind of gotten into no, no question is no dumb questions here. Right. Like just what what is it that you've always wanted to know? And you haven't been able to ask. Right. You haven't been able to get the answer. GitHub. What the heck is GitHub? What's a wiki? What does API mean? And why do I care? Right. And these are not AI specific terms. These are terms that are now relevant for people because they see themselves as builders more. Like these are things that maybe I'm using now, but I don't really know. And I can say I often don't know what the real names are, but I tend to share ideas about them in terms of images, which will not surprise anyone. But my answer for API, API is a secret backdoor handshake. That's what that is. Don't tell anybody else your handshake when you know it. But it's not secret, though. The company that exposes an API publishes it. And so you can see exactly. Oh, yes. Sorry. It is not secret. Your the handshake is your personal API key. That's secret. Don't tell anybody. That's the secret part. That's right. Yeah. You access the API. You've got to have an API key because we're going to charge you for it. That's right. Yes. The API is the backdoor. That's true, Andy. Thanks. So. Do we. I want to just give a plug for for what happened yesterday. After we continued the show yesterday and Gareth demonstrated his Jasper local agent. And he then published to our community, the daily A. I show community dot com free slack community. And one of the things that you ought to know about if you're if your company doesn't already use slack is how slack works, because slack is is the communications platform for pretty much every company that's out there. I don't know. Maybe some companies use teams because they're on Microsoft or whatever, but slack is a owned property of salesforce dot com, but it started out as an independent company and it works anyway. So in our slack community, Gareth published a prompt that you can use to start cloud code or codecs building your own version of his Jasper personal voice agent on your own machine. And it's pretty cool. I want to do it. What I want to ask Gareth to to expound upon the next time you join us on the show, Gareth, is can you distinguish or differentiate how a Jasper implementation on the local machine is different from an open claw implementation on a local machine? So is it is it replacing open claw? It it seems that it has the capability to do all of those same things, but maybe there's some features that open claw has built into its native architecture. And I think I saw that version 5.7 is now out on open claw. I've never implemented that. I don't have the Mac mini to do it. And I don't want to put it on my main machine because it's already, you know, running codecs and cloud code and everything. And I'm happy with all of the things and capabilities that can do on that. But anyway, if you'll help us understand how pursuing the approach that you took, which is personally coding a personalized voice agent that has access to your machine and lots of other resources on the web and elsewhere, how that's different from open claw. I'd really appreciate that. But you can start building your own by just joining the community and getting the information that Gareth very kindly shared with us on that. And again, that's the daily A.I. show community dot com. We are the daily A.I. show dot com. And that is a website that is going to be updated in the next. I'm going to say it's going to be updated in May. We're going to we're going to point things to two different sites right now. It's one page site and we've been working behind the scenes on a ton of great stuff. Not all of it will be available by the end of May, but I think that we'll have at least something about the January episodes. Sorry, something about the 2026 episodes up on a site. That's cool. OK, I believe Gareth is back with us on Monday and ooh. And he's answered some of the things in the chat. Yeah, here's here's what Gareth said in response to my question so that you guys can get his his top line take. He says, I think open claws overkill and we'll have latency issues. One of the things that Gareth explained is he's worked really diligently to try to get very fast response from the Jarvis agent on his machine. And he's moved to open A.I.'s just newly released real time version two by real time voice two. And that has improved the quality of the Jasper responses, which is his Jarvis, you know, personal agent. And he says the people he knows that are building this kind of personal agent that you can use, you can do yourself using cloud code or codex. None of them are using open claw. Yeah. Well, open claw was the was the first that caught fire. Right. It wasn't in a scene of a million agents and the best one rose. Now we're in a process where people are determining what they think is the best. I've made the decision for Hermes and Hermes is coming out on the regular. One of the things that sort of sold me on Hermes initially was the persistent memory being able to say, hey, what have I asked you to do or based on what the things that I've asked you for? What else could you do for me? Right. And then so you're not always needing to come up with the idea, having something able to experience you and reflect you back to you. You should be asking me. Memory is very important for sure. But I think open claw has memory extensions, if you will. Like you can build an open claw instance that also has that capability. But what was good about Hermes was it was natively designed into their code for the Hermes agent. And because its native design was to self-improve, to learn, not just as like a memory trace, but to memory trace that it could analyze and improve its behavior. Right. Which is sort of core to a bunch of pieces that other people are putting things into. Compound engineering, Gary Tan, G-Brain. Yeah. So many things happening. All right. I think we're wrapping up here. Happy Friday, everybody. Next week's going to be a big week. We have doubled context windows for Anthropic. So if you stopped using Cloud Code because you were hitting your limits, go back and look. I did a bunch of stuff yesterday. Also, if you haven't signed up for the newsletter, the episode coming out on Sunday is the 100th episode of that weekly newsletter. So we've done it for 100 weeks now. And I encourage you to go to the dailyai show.com. There you can sign up for the newsletter. And it'll just come to you on Sunday morning. It's a compact read. Might take you if you really dig into a lot of the detail that's there, which is a digest in very short paragraphs of all the news we covered in the week. It'll take you 10 minutes, but you'll be up to speed. And if you've missed any of the shows in that week, you'll be caught up. So sign up for the newsletter. It's different from the other newsletters out there in an important way because of its periodicity just once a week and its attempt to distill and elevate the things that are important from our observations. And so I encourage you to do that. Brian is the original author of those things and the tools that we use to generate it. He's been away for a couple of weeks now. Brian, if you're listening to the show, I hope you guys are enjoying the last days of your well-earned vacation. And then he'll be back joining us again next week. Yeah. Not at the beginning of the week, but later in the week. Absolutely. Thanks, everybody. We'll be back on Monday. Take care.