← Back to search

Building AI Agent Offices and the Compute Bubble Question

The Daily AI Show · 2026-07-02 · 75 min
relevance 74 10802 words Episode page ↗ Audio ↗
Show full episode description
Today's AI news roundup: agent offices on Discord, the compute bubble debate, memory-efficiency breakthroughs, Google NanoBanana, and Altman's government equity offer. A working experiment in giving an AI colleague its own private Discord and screen-share office anchored a wide-ranging conversation about where the field is heading. The hosts weighed whether the AI boom is genuinely frothy by asking the sharper question of whether demand for compute still outstrips supply, and tracked rumblings of a training breakthrough that jumps beyond the current frontier alongside a predicted memory-efficiency architecture from an OpenAI spinout. Also on the table: real-time voice agents from Grok and Thinking Machines, Google making the next NanoBanana image generation broadly available, DeepSeek's DeepSpark and speculative decoding, and Sam Altman's proposal to hand the US government a free equity stake in major AI players. The shift from token maxing to token budgeting ran as a thread throughout, closing on Obsidian versus Notion for personal knowledge bases. Key Points Discussed:
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Coordinating multiple AI agents to collaborate and hand off work without a human acting as manual message carrier.
Benefits
  • Multi-agent communication hub via private Discord
  • Agents talk to each other, not just the user
  • Proactive Hermes runs overnight jobs autonomously
  • Cross-project critique between parallel coding agents
  • Controlled access boundaries for security comfort
Use cases
KPIs / results
  • On $20 plans without hitting token limits
  • Cut show-prep time by roughly two-thirds
Tools / build
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00:00
Opening and Andy's AI Projects Catch-Up
  • Beth and Andy open; catching up on AI projects
00:01:34
Building an Agent Office with Hermes on Discord
00:20:55
AI Bubble, Excess Compute, Meta and SoftBank Clouds
  • AI bubble debate; excess compute from Meta, SoftBank clouds
00:26:35
Training Breakthroughs, Scaling Limits, World Models
  • Training breakthroughs, scaling limits, world models
00:29:18
Real-Time Voice Agents: Grok and Thinking Machines
  • Real-time voice agents: Grok and Thinking Machines
00:33:54
Google NanoBanana and Detectable AI Images
  • Google NanoBanana and detectable AI images
00:36:42
Memory Breakthrough and Lab Departures
  • Memory breakthrough and AI lab departures
00:42:02
Altman's Government Equity Offer and Sovereign Fund
  • Altman's government equity offer and sovereign fund idea
00:47:31
DeepSeek DeepSpark and Speculative Decoding
  • DeepSeek DeepSpark and speculative decoding
00:56:32
Token Budgets, Deferred Fable, Scheduled Tasks
  • Token budgets, deferred Fable, scheduled tasks
00:59:54
Hermie's Agent Office Screen-Share Demo
  • Hermie's agent office screen-share demo
01:05:32
Obsidian vs Notion and Personal Knowledge Bases The Daily AI Show Co Hosts: Beth Lyons, Andy Halliday
So I'm on $20 plans and I am not running into token limits. How much I don't want to be monkey fingers. We would like to give the government 5% of the equity of... Hey everybody! Good morning, good morning. It is Thursday, July 2nd and you are watching and listening to The Daily AI Show. I don't know right now what episode it is, but it is the episode after yesterday and before tomorrow. And that's what I've got for you today. 7.50 something. Okay. 7.50 something we think. And my name is Beth Lyons. With me in the studio today is Andy Halliday. Thank you very much, Andy. How are you doing? I'm doing well, thank you. Awesome. Busy days these days. Yeah. Yeah, taking as much time as I can to work on my AI projects, but I have many other responsibilities. So... Yeah. You have responsibilities that can't be put off. It's not like... Horses... I'm just... I'm not going to do it today. I have many dependents and many dependencies. So I'm utterly dependent on all of that. Your interconnected world is very interconnected locally and also abroad. Well, awesome. I was saying before the show that I successfully set up a private Discord for Hermes and right now, Hermes and I to connect, but it will ultimately connect with Claude Code as the agent from my computer that I'm on right now and Hermes is on that. Okay, that's very intriguing. So, you know, one of the things that the sort of claw revolution or the lobster revolution created was the idea that your AIs can be omnipresent with you through the various communication channels that you have, whether that's Slack or WhatsApp or Telegram. You pick your communication thing. I don't use Discord. I mean, I have an account on Discord and I'm a member of a bunch of different things, but I don't ever open it up and I don't ever look at it. What is the advantage of using Discord as a platform for communications with multiple agents where you can invite multiple into that environment? This is a little bit different in my mind from, okay, I'm going to communicate with my agent using WhatsApp while I'm out about with my mobile device. So I get that. And then, you know, the major players added, you know, direct communication through their mobile apps back to the agent in order to be able to maintain continuous interaction so that you can approve steps and so on while you're vibe coding. Mm-hmm. Or, you know, watching while sort of managed agents do their thing. Discord is a little different because it allows you to create a community. And so I get that what you're doing is you're going to put multiple of your players in your system environment into the Discord so that they can talk to you and talk to each other. Discord. Yes. Yes. Which is the point. The point is that they talk to each other because Codex is powering Hermes right now and Cloud Code Opus is powering Cloud Code and Opus 4.8. So both of those, Codex 5.5 high maybe, both of them are at the high level, both of those are very smart, very capable, and immediately go to give instructions when I say, hey, I just had a conversation with Hermes about something. Right? And Cloud Code immediately gives me instructions based on its context knowing me in our past history and everything that's put in its memory about how it should happen. I said, no, no, no. I said that Hermes and I talked about it. I was like, okay, well, I'm just going to write you a handoff document to give to Hermes. I say, okay, Hermes, Cloud Code just wrote the handoff document. Here you go. How does that relate to your thing? Okay, you need to tell Cloud Code that like I'm the orchestrator and that's just the work for me. And I'm like, no, I am now and for a long time setting this up, I am the embodiment of monkey fingers. Right? Like, I am saying, here you go. I copied this. This is what Cloud Code said. And I have to say that this was a slog. This was a big long slog. And Cloud was basically saying, okay, so we got this far. It's midnight. You need to go to bed. And, and I'm like, no, just give me the instructions for this. I was like, I am not leaving until the two of you have a central communication where you can exchange information and knowledge. Yeah. I think you're not understanding how much I don't want to be monkey fingers. Like, you know, as, as the dispatcher of interagent communications today. I can't imagine how I could rely on each of the agents to understand how to manage the handoffs. Like, do they have, do you give them a framework that says, oh, as soon as you get to this point in some development that you're doing. If it touches on, so let's say an issue of UI design, for example. I want you to give it over to Claude. I want you to go talk to Claude about that. And there must be some additional instruction sets or rules that you provide in the documentation for each that gets, I don't know, I'm a little confused about that handoff process. Talk to me about that. Work with me. We are not there yet. What this is doing is literally, I'm having a conversation with Cloud Code. We're working on some aspect of something. And if we were really just doing code, these would be work trees. Like, we would be doing a work tree over here and a work tree over there. And the central, there wouldn't even be a communication because the central repository would be GitHub. Right? And so then everybody can know. We're doing more knowledge-based stuff. And more creation in that context than just it's a software piece that has a goal and it's easy to figure out. So I'm not in these conversations. These are separate conversations that I'm having in two different chat windows. Just when they start to converge and it's reliant on me to be the carrier of the message from one side to the next, while I'm still trying to say, no, wait, you don't understand this context. And no, wait, you don't understand this context over here. So that, so at some point in conversations, we will say, hey, let's go have a check-in meeting on Discord. Got it. And share the files back and forth. Yeah. Ultimately, I think there are many things that will happen. And I gotta say, I am really digging Hermes. So, um, I gave it an overnight job that, that created, uh, created a presentation. And I will share that later in the show. Oh, great. But, um, uh, and it was very funny. I'm very excited and, and, and I feel privileged to, you know, to have this association with you because you are, you are the icebreaker when it comes to Hermes in my world. And I, I fully intend to implement Hermes one day when I can find enough time to do it, but you've invested an enormous amount of time and attention and getting that set up. And I can see that it's, it's not, uh, it's not complete yet because you have, uh, you know, a long way to go for Hermes to have access to all the tools that you do so that it can substantially assist you side by side at a peer level as your, I don't know, your, your, your girl Friday, so to speak. That's hilarious. Um, yes. And at some point there will be at least one and maybe multiple show agents and we're getting close to that. Um, part of, um, part of the challenge that I, uh, come to in trying to figure out how to create show agents is that I know how I'm phrasing things. I said this, I think yesterday too, um, that I don't know what the real words are for running a show, right? Like there are, there are inside baseball, uh, lingos that are shorthand, um, for like an end result. And I am just saying, Hey, I would like this sort of thing to happen. And what do you call that? So being able to create, um, an agent that, that already possesses that knowledge, I've like cut my time in a third. Um, uh, Gary was saying, yeah, he's asking, are you using the desktop app Hermes? So, uh, yes and no. Um, the desktop app Hermes isn't actually, I guess that technically it is a desktop app. Um, I have used it. Um, and if I go in through, so the Mac mini is being run. And headless, which means that I'm running it from my MacBook air and I screen share or I, uh, SSH. I don't even know what that's called, what that stands for. Um, into that machine. What, uh, desktop really is only useful when I have a visual. So like I'm screen sharing, I can see it running on the mini and then I do stuff back and forth. So I did like that. It's a nice thing. I would prefer it. It does not work in the setup because you can't SSH into a desktop. Um, and, uh, and, um, I don't actually know. Do you know what SSH stands for? Like it's a, it's a way of. I'll tell you in a minute. It's a way of connecting it. And I did, uh, I did a little thing on tail scale too. Um, let me see. Secure shell. Jeff is telling us. Um, and voice to text. I hope you're not chatting and driving. All right. Uh, here we go. Gareth is, uh, is driving. Um, so he, uh, is, uh, participating in the show from the chat. Yeah. Yes. A protocol and set of tools that let you securely log into another computer, run commands and transfer files over an encrypted connection typically uses port 22. It was designed to replace older insecure remote access protocols like telnet, R login and RSH. So it's the way that you do exactly what you're describing. It's developed to allow you to have, uh, a programmatic, uh, operation of another machine. Right. And, um, that's how I'm working with the mini. That's how, um, I, uh, work with the MacBook pro. Um, although the MacBook pro is really just running models. It's not like, uh, it needs more of an interface, but, um, yes, this, this felt very exciting. Um, and the other thing is the sharing of files. Ultimately that will be, um, something that makes a little bit more sense. Like I think we're still doing an end around that we don't need to do, but. So let me tell you something that I do with, with respect to sharing files. And I'm working on one machine here. Mm-hmm. So one machine project folders for codex and for Claude, they both have their own project folders, but I can easily ask either one to look at the project folder of the other. Mm-hmm. So I can tell one to put a markdown file describing this situation into the project folder under, you know, some subfolder, and then I can go to the other and say, I'm talking about using both the codex desktop and the Claude code desktop. Mm-hmm. So I can just ask them to look at the other one. And periodically what I've been doing, since I'm doing this parallel same project development, using both of them to see what happens. It's very interesting what I'm watching what happens as each of them go in their own way with my guidance, but they, you know, they, it really is an interaction between the, the vibe coder and the coding agent that steers the process in various ways. And so these things are very different. And you'll see when I ultimately reveal both of these versions of the application that not only at the UI level, but actually in functionality and transaction design, they're, they're different. Right. Right. Some, some, you know, subtle ways, some dramatic ways. But anyway, that dialogue is, is also now reaching the point where periodically I'll say, okay, and I haven't done this recently, but I'll say, okay, codex, go look at the, you, you know, review, you know, the, the design of our system now and all the memory about our development. And now go look at this project folder and give me a diff. What's different over there from this one. And is, and, and the last time I did that, what came back and I did it bilaterally. Right. So each one got to look at the other one is I got, you know, responses like, oh, you know, this, this thing we ought to incorporate, you know, this, you know, I don't think this is as useful and it, you know, would be duplicative of this thing that we do and so on. So it does understand that this is a very, very similar project, but it's an independent project. Each of them understand that. And then they can provide critique and borrowing, you know, from the development process of the other. And so it's very interesting how that can work. And that's a different approach that doesn't leap to the point where Hermes is operating across multiple machines. And now you're bringing in the idea of SSH to allow you to operate from one machine to another. And I would like to get there that ultimately I'd like to get to that point. So, and the reason for me is that for my comfort, I am comfortable with the level of access that Cloud Code has to my machine because I initiate or am in control of every Cloud Code session. Right. So I can set a cron job or whatever it's called in Cloud Code. And I, I am the initiator and Cloud Code reacts. Hermes is a proactive agent. And because of that, and my wanting to be comfortable before we start doing other stuff, the, the relationship between the computer that I'm on now, which is the MacBook Air, can go in and do things on the Mac mini where Hermes is. But Hermes cannot do things on the air. Right. So that was a, that was a break that I wanted to make sure was happening. And I'm not ready to just say like, sure, no problem. This is very frustrating to me. So we're just going to throw all the doors open and see what happens. Um, uh, for me having, uh, looking around and seeing that other people, um, there were a couple other reasons that a discord server worked for some other folks. Having that as another option made really good sense to me. Uh, so that's what we did. And you are muted, which is why I'm talking over you. Oh, nice. Thank you. I see a day coming though, when you'll actually release those strictures on the channel cross talk. Uh, and I can imagine a day when, uh, you know, the, the, uh, one delegates a task to the other. Yes. And then the question that arises in my mind is, wait, what, what if the other that is being tasked by the original, what if it has an objection to that because it's distracting and what's the priority that it could assign when its primary task, uh, the supervisor is you. So like does the Hermes agent, for example, is it allowed to delegate tasks to any one other? And can, can that one in turn tasks, task Hermes with something that it needs? And who adjudicates that in the process? So it gets a little mind boggling to me. And that's why for the moment anyway, I'm keeping everything in one environment and, and I'm the marshal. Like I, I, I direct the activity, uh, of each of the, uh, coding agents around this sort of bicameral development process. Yep. And I think that one of the pieces that is, um, informing this for me is that I've spent half a year building, um, the colleague protocol with cloud code and building memories of how we work well together. And I have not started that process with, um, codex or, um, Hermes. So I'm not ready for Hermes. I'm not ready for that piece to happen before there's a building of a process of how we work. And I will say that last night, which like, we're just about closed out with cloud code. I just haven't exited the terminal and something went bonkers with, uh, with Hermes, um, uh, with discord bonkers by bonkers. I mean, uh, Hey, this should work. Go do this. Not at all. Right. Like, Nope. Stop. Try this again. Nope. Get nowhere. Thanks everybody. Like what is happening? And I just went to cloud code. I gave the screenshot and I was like, help me. Uh, please. I know we're done. I know this is a closeout. I know that the context window is like at 49% left. And that is much longer than any of our conversations. Um, and because of the history that we'd built, there wasn't a lot of, well, first I need to orient blah, blah, blah, blah, blah. Yeah. Um, and the, the, uh, I'll talk a little bit about the screen share protocol at some point. And he also, um, but, uh, yeah. Okay. Okay. Uh, shall we hit on a few news items now? I think we should. Personal news. There's actually. Personally, I find that discussion much more interesting than those, all of those people out there. Uh, I don't mean our people. I mean, the people who are making news, the newsmakers out there, but let's talk a little bit about the news. I wanted to mention that, uh, one of the things that surfaced for me is around the question of whether we're in an AI bubble or not. And, and, you know, I think it's frothy, no doubt. Uh, but one of the sort of biggest signals about whether there's a bubble is whether there is really still excess demand for compute then is being provided by AI inference compute infrastructure. And one of the ways that, uh, you know, additional compute infrastructure has been, uh, created, uh, or acquired is major frontier model developers buying up chips on mass. Right. And, and then, and you, you, you, X bought Colossus. Like, I think it was 200,000 of these, you know, GB three hundreds or whatever. I mean, they, they built an enormous. Edifice of compute infrastructure and, you know, they're using it, but they also, it appears increasingly have excess capacity that they can then make available to other players. And so you, you read about, and there's been news about, uh, you know, one of the major companies, whether Anthropic or OpenAI contracting with X or with Google or with someone else to borrow capacity for their training processes. And so they have this demand, there's this excessive demand that goes beyond their own capacity internally to, you know, run these compute systems. Then on top of that, there are some intermediary companies like CoreWeave and Nebius. These are companies that buy up chips and run data centers and do that on a, for hire basis. So I'm going to build data centers for compute and I'm going to rent it out to other people. And those companies just recently took a big, big hit in the, in the marketplace. Why? Because there are new indicators that there is excess capacity out there in compute. Now that excess capacity could be a function of, oh, well, there's the, you know, a lot of AI compute, uh, data centers are coming online. And as a result that we've reached that threshold where now there's more available than we needed to, to, you know, fulfill the demand. Uh, and the, the, the other, the other aspect of it is that AI may be getting more efficient. And when we talk, you know, repeatedly about, oh, here's a new method. And I want to talk about one in a few minutes or somewhere in the show. There's a new method that increases the efficiency of AI compute infrastructure with respect to the outcomes or the objectives of, you know, inference and, uh, pre-training runs and all of that. Mm-hmm. Mm-hmm. So what just happened and what's in the news right now is that Meta, Mark Zuckerberg said that he's going to build a cloud business out of the compute that they have. Mm-hmm. Which immediately then creates competition for the likes of Cora Weva and Nebius. Mm-hmm. Uh, and that also may signal that, oh, there's, there's really not a need for all those chips out there because now Meta can't really use all of the, the, uh, chips that they've got. Because now they're going to try to make some money by selling time on their chip farms, so to speak. So they're going to, uh, they're going to go into this business or not. I mean, Zuckerberg is famously or infamously, you know, erratic about, you know, his initiatives. Uh, Meta, the name of the company is all about the metaverse, which was his big bet that people would be living in virtual realities. And instead, AI came along. So he's, he, he, you know, kind of pivoted to that. Uh, anyway, I, I think that's a very interesting trend. And then on top of Meta, then today, SoftBank, which has been a big, big funder of, uh, you know, huge, uh, treasure chest, Masayashi, uh, or I, I forget his name, Son. I forget what his name is, but the guy behind SoftBank, uh, developed a huge treasure chest based on SoftBank's investments in the internet era. Uh, and now they're going to go into the cloud AI infrastructure business. So you've got a bunch of players saying, okay, we, we apparently have access to the chips that are necessary to do this. We're going to build out data centers and rent that out. But all of that is, you know, these are smart investors. It's, you know, I, I'm making one argument, which is, Hey, it looks like there's excess capacity in the moment, but that may be quite temporary in the long run vision. These players are probably thinking, no, AI and robotics are going to command greater and greater need for this AI infrastructure, the compute. Uh, and so we're going to invest heavily in that. Yeah. Yeah. So there, um, there are a couple sort of rumblings of a breakthrough or more than one that have happened. One we think is that there's been a breakthrough in training that goes from the 5.5 opus level kind of stuff to this next piece. So it's not just that they were trained for longer or slightly, uh, or, uh, like trained the same way the others were trained, uh, just with more that the, this seems to be a different scaling. Yeah. So it's not just that they're trying to go about a, uh, to go about a, a, an internal virtual world and learn on their own from experience as, as opposed to from texts that have been created by humans. Right. And the other piece that, uh, the other rumbling that I am seeing, um, and it was talked about in, um, AI leaks and news that, that newsletter, um, and X account, um, that there has been a breakthrough in memory. Uh, and we've talked about this a bit, right? Like we've been, memory seems to be the thing again, what we're, what we have said is, is not how much you remember. It's how, uh, efficiently you forget what we don't want you to like clutter up with that kind of stuff. Um, but there does seem to be, um, uh, a potential breakthrough in memory, which then increases the capacity for agentic action and all of those kinds of things, because it's not losing the context within its run. Right. And that's why potentially fable is doing some really interesting things, but the, um, yeah, I don't think it's at either of the frontier labs, but I do think that it is, um, affecting the way that things are coming out. Yeah. Okay. Let me touch on something else here, which is interesting, which is that, uh, another major trend is we're going to move closer and closer to pure voice interaction with AI. Right. Eventually it'll be even beyond that. It'll be more than just voice, but it'll be video. You know, your Hermes will have a visualization that, uh, is representing something more akin to, and, uh, sort of more appreciable in human terms, uh, like a, a real colleague that you talk to. Right. You can design the, the, the visage, uh, or the appearance of that yourself, but voice is, is getting closer and closer to real time interaction. And, uh, so X AI, the division of space. Sex now has launched no code grok voice agents that can interact on real phone calls with under a second of, uh, latency across 25 languages. So you can apparently talk to grok by phone now. Uh, uh, uh, I'm not sure exactly how they're implementing that, but they, they are along with many others who are out there doing real time voice. Um, each of the major labs with the exception of Anthropic, I think has, has a, you know, a real time voice, uh, API. Uh, and so these are being used in the customer support world, uh, in outbound marketing, uh, lots of places. Now you, you would, uh, eventually not be able to distinguish a, a live cold call, uh, with a real human from an AI cold call with an AI, uh, real time voice character. Anyway. And that I am, I'm super intrigued. Like, that's what I'm really interested in. Um, the experience that Gareth is having. Right. Because I, I miss voice and I have to say that, um, perplexity really just like, it was the most amazing voice experience. And then it was just like, we are not going to do that anymore, Beth. We're going to do something totally frustrating. Like, okay, well, um, that's not helpful for me, but I do think. That, um, voice. So when, uh, when AI, um, when we started to be able to generate AI images, right, there were a bunch of things that worked really well. But when you started to try to generate a human and specifically a human face or hands, the little minute things that you didn't get right were like screaming. Yeah. And I feel like voice is that for me instead of text, right? Text is you can generate a landscape. That was a very nice tree, right? Like, cool. I don't see that there are problems in the tree. The branch doesn't connect to the trunk. Like, that's not something that I said ever, um, to, in like a landscape. Okay. But I can totally see that that face is not a face. That face has, uh, right. It's, it's funky. And I feel like voice is that you can give me a bunch of information in text and I take what I want and I like it and it works back and forth. But if you get voice wrong, I have a more visceral. Yeah. Kind of thing. Yeah. And, uh, just one of the, to mention one of the other companies, uh, Mira Mirati's, uh, company thinking machines is, you know, their demo recently. Uh, it was all about real time, multi channel, uh, multiplayer, uh, full duplex communication with the AI and the AI being much better than humans at attending to multiple voice streams that are, you know, in crosstalk mode, uh, you know, in full duplex. So that's, that's the whole voice area is interesting, but you, I'm going to weave off of your mention of image. As an example of how AI artifacts, which are still detectable by humans, uh, you know, makes those a little, you know, uncanny to us like, oh, no, that's, that's AI. And the voice, uh, voice artifacts like that, that are AI also make that happen. But on the image front, Google is just making generally available to everyone. The next generation of nano banana technology. And it's in Gemini three pro image and Gemini 3.1 flash image. So, uh, you know, this Google has developed really kind of state of the art, uh, image generation capability. And everybody who's heard the expression nano banana, it apparently is not going to survive the, uh, you know, the, uh, naming conventions. I thought it was really cool because it's memorable. Now Gemini 3.1 flash. Is that the right one? Or is Gemini three pro image the right one? You know, I can't figure it out. Give me nano banana. I can grab onto that. Right. Yeah. So anyway, but those are now generally available. So on Gemini, you can generate images and, and you'll see as is part of this, uh, I think the product strategy of Google, you'll see this image generation capability folded in neatly into a lot of things that you might do, including Google slides and elsewhere. Yeah, absolutely. And I think, so yesterday I was on a, uh, I was on a non AI, uh, conference. Yeah. I do have other flavors here, but, um, I found when watching the presentation that I could immediately say, oh, that's an AI. That's an AI image. And this next one, that's also an image created in AI. It didn't bother me. Right. It got the, um, like they were very, uh, evocative for what the present presenter was going to say. Right. Um, so it was like a, a path through trees. There was like someone standing with their arms out to the side in front of a big landscape. Um, but I have not been, um, someone who is engaging in regular conversation, uh, not expecting AI, not, not expecting it, but it was surprising to me that I was like, oh, there is, there really is something that I can tell. Um, and I don't see that as much when I've created it. Right. But, um, non AI does that exist? Absolutely. Um, so I have found the, it's AI leaks and news is the channel. Um, and the person who was sharing this is Andrew Curran. Uh, and I'll put a link to this. Uh, they say I'm posting this prediction now posted on, uh, June, June 30th. So two days ago. Um, so I can quote it later. That is, there's been a significant breakthrough in architecture specifically around memory efficiency, not by one of the big labs, but by a team that was spun out of open AI, not SSI. So not Ilya. They will probably announce it soon. Um, so, uh, the next commenter says, is it a prediction or inside information? Andrew? Uh, Andrew says, this is a prediction based on things I have been told, but I don't know for certain if it is true. Uh, another commenter said, that makes this something not a prediction, not prediction shaped at all. Doesn't it? Uh, I am predicting that tomorrow will be Wednesday followed by a day named Thursday, the day after that. And Andrew, but like, I don't know how quick this is, but Andrew says, the earth could be hit by a sterilizing gamma ray burst tonight while we sleep. Wednesday is an imaginary concept that humans have made up without humans. The term ceases to have any meaning and vanishes with us. There are many scenarios where tomorrow is not Wednesday. So Wednesday is a fascination of, of human imagination. So the, so that's where I'm coming from that. I think this may be something, uh, that's going to be announced. And, uh, there are more spinoffs from open AI, but if it's not, uh, super intelligence with Ilya, the next one that comes to mind is Mira. Right. Is thinking machines. Yeah. It comes to mind, but there was also just last week discussion about some major defections or losses from her core team. The people who founded it. So there's, but you know, these, these things are fluid and who knows what the, uh, caliber of the remainder is. And, and importantly, what Mira's capabilities are in order to kind of, you know, respond to or redress the losses that from those people leaving. And I think some of them went to Anthropic and, and others are probably trying to ride the billion dollar IPO train quickly. They'd have to wait later, you know, to, to participate in whatever the outcome and exits would be for thinking machines. Right. And that also just, um, I mean, the, by a team that was spun out of open AI, like, um, Hey, uh, isn't that anthropic? It, it like, couldn't, couldn't you at some point say everything is spun out of deep mind or open AI? Yeah. Yeah. They are definitely fertile ground for, you know, the, the people who are the source of innovation in the world of AI. And then they're there, it's not limited to open AI, but open AI was a gravitational well for all of that at the outset. And, and then, uh, in part, because it was such a principled approach, which was saying, we're going to do this for humanity. And so that I think inspired a lot of people. And then suddenly it looked like people were going to become multi-billionaires. And so, Oh, well, let's forget about that humanity thing. And let's go after the commercial opportunity. I want to get, I want to, it looks like I'm going to be a billionaire no matter what. So I'm going to go where people are having fun doing this work. Right. Like I, I, I'm curious about like what motivates somebody to work, to leave what there was the person whose name I, do not remember, but a couple months ago, it was, they were leaving Anthropics. Uh, they were announcing their leaving Anthropic to go back to, I think the UK and, uh, do poetry, like enroll in a poetry masters. Uh, and they were part of the safety team at Anthropic. Yeah, they gave up. No safety here. This is just going to, well, I think it was just too fast. Like I was in the context, like Brian is always like, well, you have to understand everybody's like, got a reason to do this, either raising money or before the IPO. Um, like there's also context of like, Hey, so if the work that I do now is changing dramatically and I'm not going to do that anymore in the same way that's valued and I don't need it to make money. What do I want to do? Yeah. And I do think we've seen a couple of people like, I'm going to go study poetry. I'm, uh, going to travel and do various things. We also see people saying, no, this is, uh, I'm going to go deeper into another lab. Yeah. All right. So let me spin off a little bit in a slightly different direction. I wanted to mention that in the news, um, Sam Altman, open AI has offered to Donald Trump and the administration, uh, something that had been proposed by others as a way to sort of provide more benefit to the general population than would otherwise accrue. And that is to give the government, the U S government, an equity stake free of charge, uh, in the major players in AI in the United States. And so he's actually specified that he would like to give the, the government 5% of the equity of open AI, which would be around $45 billion based on an, you know, 800 plus, uh, you know, billion valuation or trillion, who knows where it would go, but it'd be worth more than that. If the value of that equity goes up from the most recent round pricing. Uh, and so that, that would then of course trigger, uh, others offering something similar in order to one defuse the animus from the government towards, uh, you know, the companies, uh, as part of their negotiating strategy to provide influence and control over those AI companies. Uh, so now, you know, they're giving their equity up, you know, to the government. Uh, and then there might also be a regulatory approach to that. That's being proposed by others like Bernie Sanders, who said, no, AI is a public, uh, you know, threat and a public good. Uh, this needs to be owned by the public, not, you know, only solely in the hands of the private company that he's, I think I've heard references to in prior news, you know, saying, oh, like 50% of the equity value of those companies ought to be put into a sovereign fund, you know, for the population, as opposed to just in the hands of the very small number of people in the human population who own the equity of that, uh, of those companies, because it's going to extract a tax on all of us because of the value of AI to society. Right. So you said, uh, at no, uh, when you first said it, you said there, we're going to give them this, uh, at no cost or something like that. I mean, it's not an offer for the U S government to invest in open AI. Right. It would be here. Here's 5% of the company put that into, you know, some, you know, structure that will ultimately benefit the U S population. So it's very interesting because Bernie has talked about it as, um, I think full population, but, but U S population, if not, and we, uh, are Americans. If you cannot tell, even though I am currently in Canada, uh, I was created as an American for most of my life. And, uh, Jen and I were having a conversation that I apparently am going to be an arrogant American for, and the, and just see things that way for a while, because there is, uh, there, there are things that as Americans were like, Oh no, no, no. Well that, I mean, we're entitled to that. And Jen's like, uh, it's called, it's called American exceptionalism. Yeah. Yeah. Maybe step back. And it's a farce, it's a farce. But so, uh, we tend to interchange, um, this is, this will help people, uh, when what we're talking about, when it's given to the American government, uh, the U S government, uh, the people we're talking about are U S citizens. And it may be that at some point, uh, the U S citizens we're talking about are another subcategory of that. But, um, the financial times phrased it as giving this administration a 5% stake. And that is a very interesting phrasing because my understanding is that it's not actually for this administration, but it is for to set up a sovereign fund to, uh, create dividends that can be shared with, um, the rest of the citizens in the country again. Right. So it's not people in the country. Uh, very clearly there, there are not, uh, there will be a, a separation between who is considered to be, um, people who are part of the sovereign fund and the other folks are, uh, are referencing it as this is intended to, to fund this, but, uh, given it to this administration, I think is the way for, uh, Trump to be like, that's a fantastic idea. Absolutely. Let's do that. Yeah, indeed. Okay. Now earlier in the show, I suggested that I was going to mention something that was related to the increasing efficiency of AI, because now we've moved from token maxing to token budgets, token budgeting, right? You're trying to figure out how harnesses and orchestration can be used to reduce the overall token cost of doing the work that you want to do with AI. Well, meanwhile, in this, in the labs, they're working on new algorithmic approaches to improving the efficiency of LLM compute and deep seek just put out something pretty powerful. And I'm surprised that we haven't seen a whole lot more about this in the last week, because it was over this past weekend, just a few days ago, the deep speak, deep speak. Yeah. I'm going to deep speak about deep seek. Deep seek released D spark. Now say that three times fast. Deep speak about deep seek releasing deep spark. We're going to deep tweet about deep speak. Deep spark. And then we're going to deep seek the tweet that is, yeah. Okay. So what D spark is, is a new system that allows the LLM to answer faster without changing the accuracy or the, or the form of what the underlying model is saying in response. So in the way that works is D spark, uh, just to put a number to it, the user generation speed using the D spark system with the deep seek V four model is 85% faster, 85% faster matching throughput levels compared to the previous production baselines. So that's a big, big change in the efficiency of the output. Yeah. And the way it does it is, is as follows. It's called speculative decoding and they're not the only company to be doing speculative decoding. And what is speculative decoding? Okay. LLMs generally generate text at one token at a time. And a token can be a word or a part of a word, a punctuation mark like that. And every new token depends on the, all the tokens previously produced. So the model has to keep pausing, run back and look at every single prior token produced, and then figure out the interrelationships of all of those computationally. And then it'll check that as full context. And then it chooses the next token, single token. And that's very accurate. It turns out to be capable of doing, you know, almost a simile of a human in its complexity of response and so on, just by predicting what the very next little token is. And that's very slow. So anyway, the way D Spark uses, instead of asking the large model to produce every token one by one, there's a companion drafter AI that suggests several likely next tokens. And then the large model checks that batch of guesses in parallel. And if the draft guess was correct across multiple tokens, the system moves ahead multiple tokens at a time. And if the draft made a bad guess, the system rejects the bad token and anything after it and adds a corrected token and tries again. But they've built this in such a way that it just works and it results in a massive improvement in the performance of that DeepSeek system, DeepSeek V4 plus D Spark. Now, I'm surprised there aren't a bunch of other people, you know, clamoring about this because if you remember back to the beginning of 2025, when DeepSeek first came out, it threw a shock in the systems out there, not the mechanical and computational systems, but the overall industry system that's pursuing AI because suddenly a Chinese company released an MIT licensed system. So open model that, you know, was really as powerful as nearly as powerful as the most advanced frontier models and did it at, you know, basically for free. If you run your own hardware, then you can run DeepSeek for free. And it's much, much cheaper even in hosted versions. And now what they've done is they've done this thing, speculative decoding, that dramatically improves the efficiency of DeepSeek V4. I think it's a real shot across the bow of the current models out there. But I also believe that behind the scenes in our own labs, speculative decoding has been recognized well in advance of my making a news item out of it. And they're probably working on it. And we'll see that kind of technique being applied to reduce the burden of computation, like quadratic computation that's necessary to do LLM single token prediction. So there are some terms about models that I don't completely understand, but it, I'm curious about how that affects what could be run locally, right? So if you're able to take this DeepSeek smaller, probably a smaller parameter model, right? So whatever their DeepSeek is a big model, you need big stuff to do it. But could there be something that makes the quantization, which I think is what quant is, because you're like, are you running four quant or eight quant? Yeah, that just, what that represents, what quantization means is that if I take the number pi, you know, like it's an endless number, right? And I have to decide, okay, is 3.14 enough? That's three digits of quantization. And is that accurate enough for me? Or do I need all of those things out 100 digits farther? Right. So the quantization is how many digits are you using for each value? And so if you drop down to four, four digits is pretty darned accurate for most purposes, right? Four decimal places are four digits in the value of a unit in the computation. And many systems use 16. So that's one fourth of the numbers that have to be considered in the computation. So you can reduce the load on the computation by reducing the quantization of the values in the system. That's super helpful. So that's part of what I'm wondering. Does this, does what you're saying affect how much quantization would need to happen? And therefore it will make a smaller, more. Well, this is a little different. Speculative decoding doesn't have to, it can be equally applied no matter what the quantization is of the numbers in the system. And what I would say it would do for local is, remember, it actually adds a burden because you have to have a companion model that's doing the prediction of the speculative set, the subset of things that it thinks are the likely next tokens. So it's adding something there. But one of the things that you'll experience if you try to put even a small model, like a Gemma 4B on your phone, is it's very slow. Like the computer inside your phone is having to work very hard, even at quantization for and with only 4 billion parameters, it's having to work very hard. So as a result, the tokens come out like four per second, you know, and now I got a word. So you have to wait for the answer to come out. It paints very slowly on an iPhone. This methodology added to a local system like that would make it come out much faster, 85% faster. Okay. That's wild. That's great. All right. Um, so we did not talk about Fable today. I was just asking, uh, Gareth in chat, if he's going to be here tomorrow, he does not think so. We will, uh, push the Fable discussion for tomorrow. Part of why I wanted to do it is that I think Gareth's used it more than the rest of us. Um, I did a little bit last night, but, uh, I'm back in that, like, Ooh, I only have so many resources and I don't want to like use them all over here. Um, I have like four, uh, resets for banked reset tokens in codex. And I'm like, come on, do stuff. We can run it all up. Like it's, we can do a bunch of stuff. Uh, cause we can reset four times and I can't even get to like one maxed out. I'm the, you know, I, I am the roadblock. I'm the roadblock because in order for anything to happen, uh, I, I, I haven't set things up so that those can proceed without my supervision. And I haven't set up the remote access, you know, via my phone application for Claude or for codex. I, I haven't set those things up so I can't do it even while I'm out and about. So the only time that we're making progress here and the only time I'm spending my token limits budgets is when I'm sitting there at the computer and, and, you know, that's just, it's not enough time in the day, uh, you know, to ever reach those limits. So I'm on $20 plans and I am not running into token limits. Yeah. Um, I, um, and I, that was part of the exploration last night, right? Uh, admittedly the night was short cause we were working on freaking discord until three in the morning. Uh, so, uh, overnight was not, you know, bed at 10, you have all this time, but let me see if I can share. What I do have though, that does burn tokens is I have a couple of scheduled tasks that run on Claude, uh, um, co-work. Yeah. So you can, you can create a scheduled task and, and those kick off in the morning and they burn up, you know, uh, uh, you know, co-work is inherently expensive token wise. And when it's doing this research for me in the morning, then it's burning up some part of that budget. But still more a comment on my, uh, inattention to this, these coding projects in terms of hours per day, I might even get a full hour or hour and a half per day in, whereas I have envy for, for, uh, Gareth and you who, you know, have the ability to spend many more hours per day on, on these, uh, issues. And so I'm going to be riding your coattails, Beth and Gareth, Gareth. I definitely want my Jarvis, uh, eventually. So I'm looking forward to implementing those, but I'm going to shortcut the process by borrowing heavily from your expertise. Okay. So I'm going to, uh, do what I promised at the beginning of the show. If I can share the screen and it will, uh, let's share the window. Nope. What is that? That's going to be the Mac mini. Yeah. All right. Andy, tell me if you can see. I see your desktop. You do see the desktop, but now where is the, well, oh, because it's on a Chrome tab in the Mac mini. There we go. Okay. That's why. All right. So I basically just said, Hey, Claude, uh, sorry. Hey, Hermie. Yep. Uh, I have too many children now. Like, um, okay. Hey, Hermie. Um, I would like you to explain what we did. Um, make it compelling, make it like something that I can share on the show. Uh, give me a 60 second and a five minute version. So this is Hermie's, uh, thing. I did not do much else, um, in instructions. This was just, this is a goal. Uh, and because when you set a goal, you need to know, uh, you, you need to have a mutual understanding of when the goal has been reached. Um, I said, okay, so what is the success criteria? Uh, set the goal and, um, also know when you're going to stop. I believe Gareth has been my most recent, like, yes, learn from Gareth's wisdom. Okay. So, uh, this is Hermie's telling its own story with me. Beth built an agent office, not another chat silo. Um, 60 second story. We turned agents in separate windows into an operating room, a place to talk, a place to remember, a way to show the system visually. That is AI say that is AI understanding the words operating and room as a room where you operate and not what most of us would fear operating room to be. Um, agent work was becoming scattered across many conversations, tools, memory assumptions, live discussion needs a room. So I asked for progressive reveal. Um, the old pattern was useful, but leaky decisions need a durable home views need to be richer than texts. Um, the ground rule. This is a hilarious artifact. G brain is not the persistent memory layer. And that's just cause I had to say that multiple times. And now, uh, like Hermie says it regularly so that, uh, yeah, I know it knows, right? That comes up all over the place. Um, so then a working local knowledge and agent operating stack installed, written, copied, and connected. The file system was built on the first overnight. The agent office becomes real. It's a discord based office established for live communication, live conversation. Um, and this, what I'm showing you is, um, a self-contained HTML document that was just set up to do it. The office has a live room. Let's see. We've got these with, uh, I asked for a, um, Excalibur like, um, process map. And that's what it gave me here. Uh, the stack is simple because each layer has a job. Discord's job, live conversation, markdown obsidian, durable source, decisions notes, HTML dashboards, rich reviews, explainers, maps, and Kanban project notes talk about where we are in the process. Again, clean boundary. Don't rely on G brain, Beth. I like G brain, but it's not a part of this right now. Uh, this turns agents into an operating cadence Beth can actually run capability unlocked. Uh, every useful conversation has a route to become a durable note, a work item, or a showable view. First plan share guided explanation, repeatable loop show ready language. Now it's discussing with me why it did this to show on the show. We're having a meta thing. What we built show ready language, why it matters. So we built a discord based agent office where Hermie can respond in shared channels, then backed it with a local obsidian markdown wiki and visual HTML documents. I am always asking for HTML at this point. Why it matters. The live conversation is no longer the memory, right? The, that whole transcript of the conversation. We can meet in discord. We can, uh, record the knowledge into notes and it enables an agent check in, uh, and it can turn into a durable note, et cetera. And this is the last. Uh, thank you, uh, Hermie. I will say that, um, Hermie with an I is, uh, because, um, I think that's a whisper flow, uh, artifact, like I think whisper flow transcribed, uh, transcribed Hermie. Short for Hermie. All I can hear is Miss Piggy saying, I'm like, all right, Hermie with an I. Yep. We're going for it. I love it. I love it. All right. So, uh, really, really impressive work by Hermie there. Yeah, I think so. Um, I did say at the end of the night, uh, to Claude on air, which I miswrote as Claude on mini and I've confused everything. I was like, everybody needs names. This is not going to work. Oh, yeah. I have one question for you. Uh, no two. I have two questions. You're using markdown files in a project folder plus obsidian as the memory repository for this system. So it's, they're not in the project folders, like the, it puts documents in the project folders, but the markdown files that it's talking about are in the wiki or in a Beth wiki, the old, um, uh, Carpathia idea. Right. So that everything that gets done is in the same wiki. It's tagged in different ways so that they- And you're using obsidian as the platform for that. We're using, yes, I guess we are using obsidian. It's really like, uh, in my experience or my use case, obsidian is just the viewer. Obsidian is just like a, uh, series of markdown files. So obsidian is not like notion in that respect. Notion is a database. No, uh, obsidian is a very connected set of disparate records, but I don't believe that obsidian is structured in the same way as notion, right? So a database generally has a schema that's kind of hard coded into it. And the wiki is self-connected. I think that's worth a full show is to discuss, uh, you know, notion versus obsidian. Uh, and for people, people who are, you know, thinking about creating a personal knowledge base that's attached to and usable by their agents ultimately ought to think about that. And there, there's a range of possibilities ranging from, Hey, just mark down files in a folder. That's, that's perfectly adequate. And then you've got other frameworks in some of the advantages of notion might be that you can actually publish from within notion. So if you want to express something outwardly, uh, which you cannot do with your markdown files and your local machine, you know, notion can, uh, optimize that for you and do it in a beautiful way. In fact, a lot of people use notion that way as a publishing platform. Um, I I've never used obsidian, but a lot of people use obsidian. A lot of the influencers who are at the cutting edge of this are using obsidian as their platform. And I think Karpathy may in fact use obsidian and his, his implementation. Uh, so anyway, let's talk about that. Let's make a note. Yeah. Hermie. Hey, Hermie, Hermie make a note about this. Oh no, that's Gareth's tool. Jasper, Jasper. You're listening. Make a note. We want to do obsidian versus notion. But even so, why I do not currently have something that, uh, that you could initiate in that way. But also the headphones mean that, uh, none of my devices are hearing outside of it. Anyway, the, the, the other question I had is why did it keep referring back to the, the thing that you say you're not using? So, all right. I am very interested in G brain. G brain does a ton of really great things. We tried to update G brain before, uh, we tried to update G brain two days ago, three days ago. And there was, um, uh, a bug in the migration and, um, we tried to fix it a bunch of times and now G brain is just frozen because, uh, continuing to try this was not helpful. I haven't done a lot, uh, that's in G brain. And if we can just get to the markdown files, we're fine. But yesterday, uh, Gary tan or maybe two days ago, um, made a post that said, um, really G brain isn't that helpful until you're at like 10,000 notes. And I'm like, Oh, we don't have to make G brain work. We're good. We're not at 10,000. Uh, we're not even at 10,000 for the show, but that's the other piece is wanting, um, to, uh, to create a show wiki, a show knowledge wiki. And there have been various versions of people have done. Um, uh, I have some, uh, some parameters that I want to make sure in it when we create it. So, uh, that's, that's coming up too. Yeah. Cause one of the things on my to-do list is to implement either, you know, the compound engineering framework for my development stack or the G stack. And, and we had a conversation about that on the show after which I think Gareth was involved on that one because, you know, he's, he's got both of those things, uh, you know, in mind at the end of that, I felt like compound engineering would be the most direct and most applicable for me because my approach is not as sophisticated. I'm not spinning multiple projects and multiple plates at the same time, professional and personal. And, and, and, you know, I, I can see where, uh, you know, a consultant who's has multiple clients or somebody who's actively working in a company like Gareth is, or he's doing multiple developments for that. Plus also building frameworks for, uh, you know, other applications, you know, he needs that multiple, uh, you know, uh, deep, uh, structure for the knowledge management of all of those various projects. And, and these systems like the Gary Tan has designed in G brain and G stack, you know, satisfy that. But for me, you know, I'm pretty much focused on a single application development right now. I have some other projects that are like around personal finances and so on, but that doesn't require anything more than markdown files and project folders. And, and so I'm not at the point of sophistication that requires those things, but I want to have them available if I ever get more sophisticated. And I think that Gareth and I share a little bit, uh, with neurodivergency and ADHD because I always have at least three separate, uh, trains running that I'm interested in. Um, uh, with AI, I have three others that are not AI based that I can't do on a computer. Um, and there are times where those weave into each other and, um, and the folder structure just makes that crazy. And that is also why we set up discord. I like, okay, I've had this conversation with you three times in three separate windows today in different ways. And I really just want, uh, this kind of piece. The other piece, um, generally, and, and you may be using it differently. It may be that notion is a little different, but generally you need to set up the structure before you start to populate it with the information. And that's not how the journey happens for me. So I ended up with a standard template that has to do with, you know, uh, education and study because that's what I was first using notion for. It's not really well suited to managing a project. Right. And, and so I, but there are multiple templates that you can use within notion for various purposes. And I just don't have enough, uh, real time use of it. But now my notion, uh, collection, if you will, of various things over the past three years is, is a rich repository of things that I would like ultimately an agent to say, okay, read through all that stuff and tell me what's in there. And, and, uh, is anything here worth, uh, organizing in a way that would surface useful information? Absolutely. Okay. So I made notes, notion, uh, obsidian wiki, uh, markdown conversation, compound engineering G stack. There may be other, uh, offshoots of that. Um, so we've got some stuff teed up for you. We will talk about fable just not in this show, uh, because we are wrapping up, I think today, um, for people who were listening and want to see the, um, the visual that I was talking through that Hermie built, I will put that in the daily AI show community slack. Um, and it's only Thursday. We're back tomorrow. What will happen today in AI over, over the next, uh, 23, 22 hours, uh, for getting ready for the show. All right. Thanks, Andy. This was great. Yeah. Great. Thanks, Beth. Thanks to everybody in chat. Uh, thanks for, um, Craig and Gwen and Gareth and, uh, uh, Swaldo who chatted. Hey, um, Jen, let's see, Jeff. I hope I'm not missing anyone. And we know people are listening afterwards. Well, I'm going to say, Hey, uh, Justin may actually be on now, but Hey, Justin. Hey, Cisco. Hey, uh, I don't know if Jude listens afterwards, but if you're caught, if you're listening to this, Jude, Hey, Jennifer, don't forget Jennifer, Jennifer, uh, Delma, like we have such a great community. Um, so if you are listening since after the fact, and I did not say your name, uh, go to the community and, uh, let us know that you're listening. Uh, we love knowing, uh, who's out there. All right. Uh, take care of everybody and go build with AI. Bye. Bye. Bye. Bye.