← Back to search

Grok Buys Cursor, MidJourney Goes Hardware, Hermes Agent & Evaluation-Driven Development

Artificial Developer Intelligence · 2026-06-26 · 55 min
relevance 83 9571 words Episode page ↗ Audio ↗
Show full episode description
MidJourney — the AI image company — just quit image generation to build 50,000 spas that scan your body slice by slice. Then the week got weirder. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ SpaceX buys Cursor — Elon's SpaceX (xAI/"Grok Cursor") is acquiring Cursor for $60B in Class A common stock — a ~60x multiple on ~$1B revenue, largely to buy an enterprise foothold. (Shimin: the first real sign of an AI-tool consolidation phase.) ▸ MidJourney goes hardware — the image-gen pioneer is licensing micro-ultrasound chips to build 50,000 body-scan spas (first one: SF, 2027), aiming for a billion scans a month. Fully private, no VC backers, a self-described "community research lab." Terabytes/second — ~500 hours of HD video per single second of scan. ▸ Tool Shed: Hermes Agent (Nous Research) — the plugin-maximalist opposite of a minimal harness like Pi: built-in memory, a self-learning skill loop, cron scheduling, swappable memory providers, and ~20 chat channels out of the box. Dan: "parachuting in with sixteen crates of supplies and a film crew." ▸ Is AI ruining our skills? (Nature) — physicians' precancerous-lesion detection fell from 28.4% to 22.4% once the AI tool was removed; 52 engineers scored 50% on understanding their own code with AI vs 67% without. Cognitive debt is showing up in the data. ▸ Claude Code is a video game (Provi.me) — the "one more prompt" loop that keeps you up three hours past bedtime, and why AI finally made B2B SaaS addictive. Plus the "agent dice" repo: roll a natural 20 and a stop hook makes the agent reflect and write itself a skill. ▸ Evaluation-Driven Development (Decoding AI) — treat every AI feature as a hypothesis and gate the PR on an offline eval pipeline (built on Opik) instead of unit tests. Gold-standard vs synthetic datasets, code-metric vs LLM-as-judge evaluators, and an "aggression" dial for how big a jerk your reviewer is. (Shimin: Newtonian physics → quantum mechanics.) ▸ Two Minutes to Midnight — ChatGPT slips under 50% share (46.4%; Gemini 27.7%, Claude 10.3%), Nvidia raises $25B in its first bond deal since 2021, and Ed Zitron walks OpenAI's FT-verified financials ($38.5B loss in 2025). ~2B users — one in four people on Earth; no 10x left. Clock moved up to 5:00. ⏱ Chapters
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Weekly review of AI dev news including the Hermes Agent tool and evaluation-driven development.
Benefits
  • Built-in learning loop and agent memory
  • Supports many comm channels out of the box
  • Pre-baked tools with boot menu setup
  • Curates hundreds of links each week
Use cases
  • SpaceX/Grok buys Cursor for $60B in class A stock
  • MidJourney pivots to 50,000 body-scan spas by 2031
  • Hermes Agent runs across CLI, Telegram, Slack, WhatsApp, Signal
  • Cursor reported $1B annualized revenue
KPIs / results
  • Cursor acquired for $60B (60x revenue)
  • Cursor ~$1B annualized revenue
  • MidJourney targets 50,000 spas, billions of scans by 2031
Tools / build
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00
Cold Open & Welcome
  • Hosts intro AI dev news show
  • AI lengthening their middle names joke
01:50
News: SpaceX Buys Cursor for $60B
  • SpaceX buys Cursor for $60B stock
  • 60x multiple, enterprise inroads play
04:46
News: MidJourney Pivots to Body-Scan Spas
  • MidJourney pivots to ultrasonic body-scan spas
  • 50,000 locations, terabytes per second
11:45
Tool Shed: Hermes Agent (Nous Research)
19:54
Post-Processing: Is AI Ruining Our Skills? (Nature)
  • Debate whether AI erodes developer skills
27:13
Post-Processing: Claude Code Is a Video Game
35:23
Post-Processing: Evaluation-Driven Development (EDD)
  • Evaluation-driven development explained
41:44
Two Minutes to Midnight: ChatGPT Under 50%, Nvidia Debt, OpenAI's Numbers
  • AI bubble check: ChatGPT share, Nvidia debt
  • OpenAI's numbers reviewed
Hello and welcome back to Artificial Developer Intelligence, a weekly conversation show and study session about AI and software development. We go through hundreds of links and dozens of newsletters each week so you don't have to. My name is Shimin Zhang and with me today are my co-hosts. Dan, he is standing in the frontier looking at the foundation of the human experience, Lasky. And Rahul, Gemini makes you command an army, Yadav. Hello guys, how are we doing today? Hi. Hello. Is it just me or are our middle names getting longer week over week? It's AI is outputting more text. As AI writing gradually takes over the entire internet, it's like flowing into the articles we cover and therefore like impacting our middle names. AI already has an impact on our podcast. Yeah, even aside from the content. All right. On this week's pod, we will first have the news thread mill as always. We're going to talk about the cursor acquisition and what mid-journey has been up to. Also, Dan, apologize for calling it the pod. I meant the shell, of course. I didn't even catch it. You could have gotten away with it too. Next up, we're going to have the tool shed where we'll be talking a little about Hermes agent. Yeah. And then we'll have post-processing. We're going to talk about whether AI is ruining our skills, cloud code as a video game and evaluation driven development. And what is that? And then finally, we will have our, I hope it's a fan favorite. I don't know. No one ever writes me emails. They only send them to shipment. So someone write me an email please and tell me if you like two minutes to midnight. But yeah, our maybe fan favorite section where we talk about where we're at in the AI bubble and see where we're going to set the clock to this week. Alrighty, Dan, let's get started with our first news item brought to us by Dan. Yeah. So what is that? I don't know. I always say a couple of weeks ago, but it feels like everything was a couple of weeks ago. This is like a while ago, right? Like almost a month ago, Elon had made a deal with Cursor. And in that deal, they were going to provide or like, I don't know, exchange a billion dollars worth of services, if I recall, something like that. And then it also gave them the option to potentially buy the company for 60 billion at some time after their IPO. And sure enough, this past Tuesday, which is the 16th of June, they have announced that they will be doing that. So I guess it is going to be a 60 billion in class A common stock deal. So Cursor is only going to be getting stock out of it. And apparently as of like last November, Cursor had reported 1 billion in annualized revenue, which is pretty insane when you think about like how small that company is. Right. I mean, they've been growing, but it's still like not that big. There's not that many engineers to have a billion dollar footprint. So pretty wild. And I, for one, am I excited about this? I don't know. It sounds fun to say it. So I'll just say it. I'm excited to see what Grok Cursor is going to look like. So they have 1 billion in revenue and SpaceX is paying 60 billion for it. That's a 60 X multiplier. That seems high. I think they're basically paying to get into the enterprise space, which is kind of what the CNBC article that we got this from is hinted at a little bit is like Grok really has their X AI slash SpaceX really doesn't have Frankly, they don't have that big of a consumer presence either. But like they don't really have any inroads in the enterprise space. And Cursor is actually pretty popular in it. It was sort of like an early tool and a lot of folks have stuck with it. Yeah. And if I recall correctly, Cursor was for a while there touted as one of the fastest growing, you know, SaaS of all time when it comes to growing to $1 billion. Right. Like it was like the hottest of the possibly hot unicorns. And to see them being sold relatively quickly and a little unceremoniously, I do wonder if we're starting to see like a consolidation phase from that initial explosion of AI tools. Yeah. It could be an early, early shot in that direction for sure. But we'll have more data there in two minutes once we get there. So yeah, I think that's really about it unless y'all had anything else, but figured it was worth covering because it's, you know, some big news in the AI space. Grokursor is going to be great. What could go wrong? Grokursor? Grokursor. We need a better name. No, it's perfect. Grok. Grok. All right. Grok with a K. Yeah. The second item I have, speaking of a crazy fast growing AI startup, the little known player in the AI image generation space of MidJourney came out this week that they are going to start. Little known? That was sarcasm. Okay, gotcha. Yeah. They are one of the first, I don't know, consumer grade AI image generation companies. Yeah. And what I always appreciated about MidJourney is that, as a little bit of background, they don't have a nice front end UI. Everything is done via credits that you then post your prompts into a Discord channel in order to generate an image. And this is when this is like the newest cutting edge image generation. Like I always find that UX to be super, super crazy. But they came out this week announcing that they are going to, I don't know if pivot is the right word, but they're going to expand into creating these ultrasonic body scans. And they are planning to build 50,000 of these body scanner spas all around the world, generating billions of body scans by 2031. This pivot was so hard. It like, I found myself on the floor with like a black eye afterwards when I read this. Like, what the heck? What is going on? So this is where, Dan, your middle name came from this week. MidJourney, in their announcement blog, said, you know, they were looking there. They felt an obligation as people standing on the frontier to look at the foundation of the human experience and ask, what do we want to be different? And how do we want to be different? And what do we want to become? And they decided to do hardware tech, become a service company by creating spas where people can go to a sauna, go take the jacuzzi, and then also get this ultrasonic scan. But it's going to be cheap, right? Is what they're anticipating. Like, it'll be like 100 bucks or something. Yeah, I think that's the idea there. Yeah. I mean, a billion scan per month is like, that's like one eighth of the entire world's population, right? Like, that's nuts. It's pretty cool. Like, the actual, so they didn't even build the tech, like the hardware tech behind it necessarily either. They licensed the scanner technology from another company that makes these like micro ultrasound chips. So each chip is like capable of basically being an ultrasound. And like, you know, ultrasound is almost as useful as an MRI for some stuff. Like, you know, usually it's like much more localized. But like, if I learned anything from watching, what was that crazy ER show that everyone was talking about like a couple months ago? The Pit? Yeah, there you go. Like, they didn't have time to get to an MRI machine, remember? So he like just uses an ultrasound and like spots the thing. Yeah. So, I mean, you know, maybe something to it. But I guess the way it works is pretty interesting. So you like, there's a tub of hot water and you stand in it and there's like a ring of these tiny like ultrasound chips all the way around the tub. And then it slowly warm, like lowers you into the warm water, which sounds a little creepy, honestly, that part. But as you're going in, the like ring is basically like scanning every, you know, slice of your body as you go in. So it can like see what's up, I guess. So, I don't know, it's not the worst idea in my mind. It's just kind of wild that like it's the same company. But I guess firstly at all words, now we've got my journey. So, you know, what's next? I think this is a fantastic idea. I agree. The news was surprising. You know, they call out that they are taking in terabytes of data each second. And they had this reference of like if you converted that data into HD internet video, you'd need to watch 500 hours of footage for every one second of scan data. So it's like this insane amount of data that they're taking in per second. And it just wasn't possible before, but now it is possible. And so this is one of those like, yes, AI and all that. But given that we have so much more compute now and we will have so much more compute, what are some other crazy things we can do that just weren't possible before? So this was awesome to see. Yeah, and they do have domain expertise. Yeah, well, they do have domain expertise in analyzing AI images, right? Like I could see how there is some, and I hate to use this word synergy in this project. Yeah. Well, plus networking too was going to be my point, right? It's like in order to do like large scale deployments of stuff like, you know, a big image generator. It's like you need to have like really beastly interconnects and stuff like that too, right? So at least for some of it, I don't know. So when I read this, my, maybe this is because I was traveling. My mind went to TSA scanners. You don't have to do the whole stand like an A, make an A. Yeah, there's millimeter wave radar for those. Yeah. Yeah. But like you could also very, depending on like how fast they can do it, you could figure out a way in the future of applying the same technology for all sorts of, you know, security. And then there's going to be some other crazy applications. The other really interesting thing is MidJourney is completely private. They have no VC backers. They are a grassroots organization. I kind of think of them as like, you know, maybe they are actually controlled by a cabal of evil people. But there's this other possibility, which they are like anonymous, but doing good things and building. They claim to be building a community-based research lab. And that is a new model that I've never seen before, honestly. Sort of like a sci-fi trope come to life, basically. Hacker Collective got to do good in the world by building a weird med spa thing. And then scan all of your body data. Like what could go wrong? So I guess to summarize our news items for this week, like we have two super early players in the consumer AI space, Cursor and MidJourney. And one decided to sell to deeper pocket. And another one decided to pivot so hard that I'm still, I'm considering calling a personal injury lawyer from the web lash I suffered. I'm still spinning. So this was a full pivot? That wasn't clear from the website. What they're saying is we're done with the whole image generation because it's been commoditized? I don't think they're pivoting to completely the spot model. But they're clearly investing a ton of their capital into this project. And listeners, if you want to go visit, you know, I think they're planning for their very first spot to open in San Francisco in 2027. I may book a trip. That sounds kind of like kind of a cool thing to do. So, yeah. So I'm wondering if we are going to see a shakeout coming up. Like kind of reminds me of a old Winston Churchill quote. Like this is not the end. This is not even the beginning of the end. But perhaps this is the end of the beginning. Like there will maybe be some consolidation coming up. All right. On to our tool shed where Dan this week, you've been reading The Dark Side of Hermes. Yeah. So I actually didn't even know about this thing. So maybe other people did. I don't know how long it's been out for. Seems like a bit. But there was kind of like a big chunk of news that came out around it in the past like five to six days, it seems like. So maybe it's a newish thing. I don't know. Anyway. I first heard about Hermes maybe a month ago. And I remember going to a local and this is Seattle area AI hackathon where the organizer told me that. Yeah. It was all the rage in Seattle or in San Francisco a month ago. Yeah. Yeah. Well, I guess that means it's dead now. So why are we even covering it? But yeah. Yeah. So the thing that immediately caught. So how did I find it? I was. I guess it got posted to Hacker News just kind of with no context. And I was like, what is this thing? And like I know of the underlying company like News Research because they've done some cool stuff with like their own models. And, you know, just like a couple other, you know, what else do they do? There's something else they did that was interesting to me. I don't know. Brain is all over the place today. But anyway. So I'm like, cool, I'll give it a read and just see what's up. And the first thing that immediately stood out to me was that it at least proclaims to have a built in learning loop. So not only does it have like, you know, sort of agent memory, which like some of the agent stuff that we've talked about here previously, like Pi doesn't actually have built in memory. Now, granted, it's pretty trivial to like make it right and make it exactly how you want versus, you know, some sort of built in thing. But it is kind of interesting that this has it in. And then it also comes with like pretty much a smorgasbord of every possible comm channel you'd ever want to talk to this thing on. It supports CLI, Telegram, Discord, Slack, WhatsApp, Signal, Matrix, Mattermost. I don't even know what that is. Email, SMS, Ding Talk. Don't know that one either. Faixu, Wecom, Weijin, QQBot, Yanbao, Blue Bubbles, Home Assistant. That's a wild one. Microsoft Teams, just in case you hate yourself. And Google Chat, which is pretty wild. You know, out of the box. So the notable things that kind of like I would argue like separate it from Pi is that Pi is like super stripped down, minimal. And that's kind of the neat thing about it. This is very much like a every single tool is included in the toolbox. It comes with this like boot menu thing that pops up when you first run it that asks you like what out of its pre-baked tools you want to install. And it has like huge lists of them. But it comes with memory. It comes with a bunch of skills out of the box that like, you know, they've sort of already worked on. And then the other sort of killer feature is that it has a cron capability, right? So it can schedule tasks, check in on them and do the things that you'd want from like sort of a long running agent. So I haven't played with it personally, but I just found the idea of like the self-learning skills. So it has like not only because I have memory, but it also has this idea of like a built in learning loop. So it can create skills from the things you've asked it to do, which I thought was kind of novel and worth chatting about. So there you have it. Yes. You know, one of the nice things about AI is you can ask AI to take anything you've been working with the agent on and turn it into a skill. Like that is practically free. But this is the only one where I've seen that it is a first class object that does it automatically. I don't actually know how frequently it does it or, you know, how, when does it know that, you know, the slash learn command should be used. But it is pretty cool. And I do want to also point out that they have a memory provider feature. And this is something that I think we haven't really covered at all on the show. There is a whole class of external memory providers on the market now for agents. And Hermes supports Hongchil, Mim0, Hindsight, Holographic, RetainDB, BiteOver, SuperMemory. And of course, the default is OpenBiking. I've not heard of any of these. I don't personally trust an external provider with my memory, especially if I was to use my agent for personal sensitive tasks. But maybe I'm just a little old fashioned like that. I have asked Claude Cole to take a look at my existing PyAgent usage and also go through the Hermes agent documentation and see if I should switch and like what benefits it will give me. And the biggest one is the memory. And I feel like, and Claude actually recommended me against using Hermes just because I have to transfer all of my existing workflows over from Py to Hermes. And also, I already have a lot of these features built in on my local version of PyAgent. So I didn't make the jump. That's the tradeoff. It's like you can have it, either have it build it yourself or you can build it yourself, you know, with Py. But this is very much like, I would argue kind of the opposite approach where like it just has tons of tools. Yeah, it's got memory tools. It has like web search stuff already baked in. And there's like, I don't know, 15 or 20 search providers that you can use. Most of them are paid, but like if you already have a subscription to like one of them, it's pretty easy to just like plug it in. And then, of course, like the thing that almost made me a little hesitant to bring this up is that then they also have their new research like platform. And so if you run it on that, then and you already subscribe to their subscription, then it gives you everything like memory, search, all this stuff all in one place, including the model runner and all that, you know. So it's like, cool, I get it. But, you know, at least I appreciate the fact that they did make it very pluggable, too, in addition to their stuff. And they're recommended ways to get the news portal subscription through like that's how you get the models or that's what they recommend. Wait, what is this news portal subscription? I've not heard of this. I did this and not come up in my research. It does everything is my understanding. So I think it has some sort of memory product. I believe it'll actually host the agent for you. And then it'll also sort of like open router where you can connect it up to lots of different frontier providers or their own models, too, I think. Got it. So if PyAgent is like the bare bones version of an agent harness and OpenClaw is the open source kind of jankly vibe coded version. If you're a fan of OpenClaw, I'm sorry. That's just a vibe I got. I mean, who's to say that this isn't also jankly vibe coded? Yeah. This is what you say, except for all. It's the cat title version. It's definitely like the plug-in maximalist version, right? If PyA is like the you're just going to go out into the wilderness with a K-bar and make it all happen yourself. This one is you've gone out there parachuting in with 16 crates of supplies and film crew. And it's a SaaS, right? It already hooks up to a subscription that you can buy. So it's battery included. You just got to pay them a little bit every month or so. Yeah. That makes sense. Or don't. Or you can go through their pretty easy wizard and don't pay them anything, which is, I think, admirable considering. Yeah. Listeners, if you've had a firsthand experience with Hermes, write to us. Let us know what you think of it. And especially me specifically. Shemin gets all the emails. I'm tired of it. Yes. I will include Dan's personal email along with his work email at the show notes. Just write to him directly. We have his address, too. If you would like to send him his name mail. We're happy to share it. The second part is, like, my email starts with Shemin at, like, I don't know how. All right. Do you have anything else to add on Hermes? Come back in a couple of weeks and I promise I'll have tried it out on something. Or if you send me an email about it, maybe I won't have to. All right. Let's move on to post-processing, Dan. First article is brought to you by Dan again. I know. I'm busy this week, Dan. I'm on a roll. All the links. So this is an article in Nature, and it is called, Is AI Ruining Our Skills? Early results are in. And I'm going to add this part myself. Spoiler alert. They're not good. Didn't actually have spoiler alert in the real title, but missed opportunity. So they had a couple different data points in here. I mean, this is really nothing super new if you've been listening to the podcast, but just seeing more and more actual studies come out that are looking at this phenomenon, I think it's worth discussing. So the first study was they had a whole bunch of physicians that were, like, essentially experts at doing colonoscopies. So each physician chosen for the study had done over 2,000 colonoscopies. And they typically, I guess, are trained to spot something called an adenoma, which is like a precancerous lesion that can be pretty hard to spot, it seems like, based on what the article said. So in the, like, sort of control group that wasn't using AI, they were, or before they were using AI, they found an adenoma in about 28.4% of all colonoscopies. So then they, of course, gave everyone an AI-based image tool that, like, runs while the procedure is happening and, like, helps them out. Their rate stayed about the same with the AI tool. And unfortunately, and I don't know if this was intentionally part of the study or not, but the tool was intermittently unavailable. And so when it was unavailable, folks that had been using it, folks, physicians who had been using it, their detection rate without AI dropped to 22.4%. So that's almost a 6% drop in skills. So that was the first one, right? So it's like your optical scanning ability as a human is not immune to this as well, right? Like, and we talk a lot about software, but like, there's many other ways that it can impact too. Yeah, I would hate to be one of the 22% or the people who happen to have my images read when the service is down in quotations. That's the only way the IRB passed it. Yeah. Or you used up your cloud credits for that day. Right. Cloud doesn't save all that more day. Yeah. I mean, token maxing is over. So it applies to doctors too, I guess. Well, yeah. So in the second study, they, this one's a little closer to home. So they took 52 software engineers, basic coding task. Half of them were allowed to use AI. They were given a quiz afterwards. The average score was 50% in the AI group versus 67% in the non-AI group. And the thing that stood out a lot was that like the parts of the quiz that they scored worse on the AI group was on conceptual understanding of the produced code. So I think how many more data points do we need to bring up about cognitive debt? But it's here. It's here and it's real. So, and there was a nice little pull quote in the bottom that I thought was worth bringing to the, to everyone, which is people need to manage the competing dynamics of relying on generative AI and staying mindfully vigilant. It's by Rinter Kalyaf, which I think is really true. It's like, you can use these tools, but you need to be aware of how you're using them and what the impacts they could be having on you both professionally and personally. So not saying don't use them, but just be mindful. Yeah. And there was this additional article where they talked about a group of accountants who have been using an automated non-AI accounting system continuously for more than a decade. And then when the tool was taken away, the accountants have forgotten how to do several routine work tasks. And that part got me thinking like how much of these tasks we should be able to hand off to AI entirely, right? Like there are certain things we've forgotten how to do, like writing scripts and that's not important anymore. Oh, I mean cursive? Sure. Yeah. Cursive. I've forgotten the name already. Yeah. That's how long ago I've been asked to write cursive. Well, I mean, it's like cursive versus like calligraphy, right? But I guess either one, they've both kind of gone away. Yeah. So. Or just handwriting at all. Right. Yeah. Like you should see my handwriting. It's notoriously bad. And I guess it's especially confusing right now because we can't quite tell what is still needed and what isn't. Like we don't know if the skills that's being lost is a really important one that we should never be able to source to AI and which ones we need to be in control just yet. Yeah. That's fair. Well. I'm reading this book about dementia and Alzheimer's and stuff. And one of the things, and we know this, it's not news from the book, is if you don't use it, you lose it. And people who suffer dementia are people usually who after retirement don't challenge their brains as much or in general don't challenge their brains as much. And then if you don't use the brain as much, then you end up losing it because there's not much to pull from. And this reminded me of that where over time, if we give AI more and more agency, it will impact our jobs. But also a lot of times it might accelerate how quickly people get dementia and old age. There are things like exercising and stuff you can do. But if you're not using your brain as much at the end of the day, you're going to get dementia sooner. So that's something definitely concerning. And then, sorry, go ahead. Well, I was going to say, luckily this episode was brought to you by Leaco problems. You spend 45 minutes doing Leaco problems every morning and you will not get dementia. Leaco problems. Four out of five grandmas recommend Leaco problems. The fifth one is busy playing video games, which is how we actually prevent it in old age. And then the other, you know, unlike the actual AI at work note, one of the arguments that we continuously have for when to use AI and when not is, and where do you draw the boundary, is anywhere where you need human judgment. You shouldn't put AI there because it doesn't have judgment and you need to pair it with human judgment, not replace it. But that human judgment, the underlying assumption there is that human has that judgment and it is continuously being practiced because otherwise similar to, you know, anything else. Again, if you don't use it, you're going to lose it. Over time, if AI is relying on a human, but the humans are not really able to help, we might, you know, almost like make it more obvious that it doesn't really matter where human judgment is needed or not, because it's just not there even when you need it. So if you play it out long term, that stuff will be a big concern too. Yeah. Well, speaking of playing video games like the grandma to stay sharp, we have a article about that by Rahul. I do not know this person's name. They're a product engineer and an indie builder. The website is Provi.me. The article is Quad Code is a Video Game. And we've seen instances of this in the news. I think even Steve Yegge had it in one of the news articles recently where he was like, you know, I'm thinking about this late at night and about what else I can do with my agents and all that. And it's interrupting his sleep and everything. And that's the same thing the author is talking about. It's three hours past bedtime, a quick bug fix turned into a refactor, and then that refactor turned into a brainstorm. Then a feature branch is running parallel, and a lot of time passed. Their reason for that is because using Quad Code feels like playing a video game where you have this like one more turn kind of feel where you give it a prompt, you get something back, and then you have to respond to it. And it's continuously this back and forth. But the end result being you're able to accomplish something in the back makes the world change, the real world change instead of similar to playing a video game where you change the world, but it's in the video game environment. And then, you know, that's just with one agent. If you're running a bunch of agents in parallel, that really makes it, you know, you have to spread your attention across different things. And if one agent is keeping you up so long, then imagine like running a whole fleet of agents and everything. And so it gives you this feeling of continuously accomplishing the same thing. Something related to this that I realized while reading this was, if you look at B2B SaaS, I don't know about the consumer SaaS stuff, but B2B SaaS was not addictive until AI came along. The closest thing you got to like addictive B2B SaaS products where maybe people are just like checking their email all the time to get that hit of there something new or Slack messages. Blackberry much? Yeah, like no one logged into their Salesforce or HubSpot or pick whatever to be like, oh, I'm going to get a hit from this. No one cared for that. But all of a sudden now I see this real sort of an addiction to workplace SaaS products, which is all AI driven because it's making people feel like they're accomplishing more. And it's very much because these are LL models. We've talked about that, how they're optimized for engagement. And so now your workplace SaaS products have engagement built into them, which makes them feel like playing video games like this article is saying, which makes it addictive. And I don't know where that will take a similar long term, but I hadn't seen addiction to SaaS products before this. Yeah. Do you guys agree that using an encoding agent feels like playing a video game? I think so. I can see where the like, just one more round kind of thing would translate to like, oh, just one more prompt feature or whatever. Yeah. Turn, whatever. So I can see that aspect of it translating. So I guess I agree with the core premise. The thing that I think is unique to the usage of agentic engineering to me is it's a little bit flipped. Right. So when I'm playing the video game, I kind of feel like I'm like wasting time sometimes. Like, oh, I could be doing something real instead of just whatever. But then I'm also like, sometimes I just need to turn my brain off and, you know, play some Slayer or whatever. But the part that I have with LLMs, it's reversed and it's kind of funny is like when I'm a step away from the computer for like just a minute, sometimes even, right? It's like, oh, I got to go like, you know, grab a water or something. And then in the meantime, the LLM is finished, whatever it was doing, has fired, done the next thing. And then it's firing up like, you know, a slightly controversial tool prompt where I have to review it and say, OK, approve. Like, I feel like I've wasted time by like stepping away for a minute. You know what I mean? So it's kind of like this, like, I don't know, antithesis in that respect. But I do think the one, just one more turn, man, just one more. I'll just beat this level, you know, and then I'll go to bed. Yeah. And then setting things up to be done while you're asleep or you're running chores and stuff. Yeah. Like if you, I don't know if you guys had experience of setting up like gold farming bots or automated scripts for your video game characters growing up. I'm showing my old World of Warcraft OG version player disciplines here. But that very much reminds me of like leaving an agent on and give it a go and wake up in the morning to see how far it's gone. Like it's very much reminiscent of like waking up at five before school starts to see how much gold I've made in the World of Warcraft auction house. And I also do think running multiple agents does remind me of StarCraft, like the article I've mentioned. Like this idea of you're paying maximum attention, you're always clicking through the tabs. That's true. And sometimes there's a fire. Yeah. Who knew my APMs? Yeah. Well, come back. Those of you that haven't played StarCraft extensively, like they're really like pro players and, you know, even bad players, I guess will eventually learn this like me. You basically set up a bunch of hotkeys that you're cycling through, like your base, your army's forward position, maybe a defensive position or two just to make sure the other guy isn't doing anything wacky. Like, you know, secondary base, like production, you know, every mining facility you've got. And, you know, like so it's just like click, click, click, click, click. And there's always a task each time you go through the loop, essentially. So, yeah, it is funny. It is pretty similar. And what was this like six, seven years ago when AlphaStar competed against the top StarCraft player? So I'm sure they learned a lot of just not just like can we be a human at StarCraft lessons from them that got applied to these products. Yeah, even the problems are similar. Like the reason why I stopped playing StarCraft was the games were too intense, even though they were like 20 to 40 minutes long. I like am pumped in adrenaline, like my hands are shaking after a match. Sometimes I'm very upset because I lost and I shouldn't have. And it kind of reminds me of being overwhelmed by having too many sessions concurrently happening. That's why you should just play big game hunters against the computer in Brood War. Just saying. It's nice and relaxing, takes a couple hours, like, you know, just kick their ass every time. You wind up with whatever equivalent of a Protoss carrier fleet is. Anyway, sorry, we digress. Before we move on, the author created a repo called Agent Dice, which you can, I think it's just skills that you can plug in. Where, you know, you just lean into the whole, yeah, this thing is like playing a video game. And then it rolls a die each time a turn ends and the chance grows the longer the conversation is running. And then when you land a natural 20, the stop hook is going to prompt these. And I'm like, reflect on what happened. When you extract the patterns that you saw in the conversation and it creates an artifact based on that, that, like, improves the system going forward. So they've leaned into gamifying it further and threw a dice into it. I thought you were going to say when you, when you roll a 20, it basically like fires a kill switch and you can't use the agent anymore for like five hours or something to like force you to step away from it. Oh, that would be healthy. I'll be too healthy in the world, in the age of AI. All right. Okay. Let us move on to my post of the week. This is a post from Decoding AI Magazine. They are sub-stack magazine about AI. It's titled How Evaluation Driven Development, EDD Works by Paul Easton and Alejandro Aboy. This is another, you know, new methodology about software development in the age of AI, similar to spec-driven development and VS verification and spec-driven development that we spoke about a couple of weeks ago. The problem is fundamentally when you have an AI feature system, it is very easy to break what has already worked. So if you change a prompt, your factor tool, an old feature may regress and you may not know it. And this is especially hard when the output of the system is non-deterministic, right? Because you won't be able to catch that regression every time. And the second problem is when you're developing a new AI-based feature, it is brand new. So we have no way of testing and see how good the feature is. So their solution and this blog post does have a little bit, relies pretty heavily on OPIC, an open-sourced AI feature observability telemetry app. The idea is you first treat every single feature like a hypothesis. And so based on this hypothesis, you make code changes. And after a code change is done, you then run through an offline pipeline where you have call-code agents come up with imaginary scenarios using OPIC. Of course, you can swap out OPIC with any other kind of scenario generation tool. It could just be another AI that's trained on your existing data. And you run a set of evaluations on your new code base. Only when your feature does its hypothesis and does not cause new regressions do you actually open the pull request. That is the, I think, heart of the methodology. It makes a lot of sense to me because I've been thinking a lot about, you know, what to do with all these AI-powered features that you have no evaluation harness for, right? Then you have no idea how good they are. You have no customer feedback. So are you even building the right thing? And this is one way to solve that problem. Kind of similar to how we had the age before tests where we just wrote spaghetti code. And now we have testing and regression tests and sometimes even test-driven development. This is kind of the eval-driven development part of our next age. The article also goes into different types of data sets. Persistent hand-built evaluation sets that, you know, are the gold standards that you use to catch regressions as well as on-demand synthetic data sets used to evaluate a feature that you're currently working on. I find that dichotomy interesting. And since we're in the land of code, they also have two different evaluation types. The evaluator could be either for code metrics that, you know, look for deterministic, linters, kind of code metrics, or you can use AIS judge, LMS judge to score for subjective things like completeness, accuracy, ranking quality. Also a helpful dichotomy to have. I think altogether, this seems like a really powerful new way of doing software development that is worth keeping an eye on and maybe try out in your own workflow. At least I try to always build in evals as a foundational part of any AI features that I'm building whenever possible. You mean AI-driven or AI-assisted development of? AI-driven. So if the feature uses AI in some way, if the feature is non-deterministic in some way, yeah. Because if it is deterministic, then the old unit test regression test framework still serves us well. You know what this reminds me of, actually? Sidebar. Sure. This reminds me of when we went from Newtonian physics to quantum mechanics. We went from a nice deterministic world to something that is non-deterministic and fuzzy. And we need to have new techniques and new harnesses to work around it. I like the aggression setting. Set it to max aggression. Bring in... Explain to us, what is the aggression setting? So you have these two modes where you can do a manual quick check versus automated experiments. Manual quick check is similar to like, you want to just do a lightweight check. It's not going to have any database. It's not going to set up experiments or anything. And then you have two, which is automating judgments. So when you're shipping your functionality, it creates a whole data set. It judges... The judges are going to score different items and then create experiments to make sure that they can catch differences in everything. You can pair all of that with an aggression mode. It picks how much of a jerk you want your reviewer to be and like how adversarial you want it to be. And you can go from happy path. I click the things and it works. Looks good to me. Ship it to like, no, why does it have this like, you know, one millionth of an edge case that is really not going to maybe do much. And so you can tweak it across that spectrum to get some really adversarial reviews as well. So that was, it's almost like simulating real life because you see reviewers in real life, PR reviewers at both spectrums as well, where you have the people who are like, yeah, I looked at it. It looks fine. Just shut up. Whereas it's someone who just like really goes through everything, nitpicks even the small things. And you can pick which one you want to go for. Yeah, except a nitpicker will burn more tokens. So it's even worse than a real life nitpicker. Potentially lead to over-editing. I don't know. Yeah. And I worry if, if it's too aggressive, it causes the whole system to drift a little bit into, yeah, some unexpected or unnecessary specs. Yeah. And they do call out like, you don't want to go full adversarial necessarily because you don't want to like. Full adversarial. Yeah. Yeah. Like you don't want to prematurely optimize and also just have, you know, waste time. So you kind of have to anchor it into premistries. That makes a lot of sense. All right. Shall we all move on to our favorite segment of the week? Two minutes to midnight. Where we take a look at the financial side of AI and make our best guess on how far we are into the AI bubble. The clock is based on the Armageddon clock by the Bulletin of Atomic Scientists. And we are currently at 5 minutes and 30 seconds. Closer to midnight, the worse it is. So Dan, why don't you kick us off? All right. So we're going to lead things off with ChatGPT's market share has dipped below 50% for the first time. So this is coming to us courtesy of TechCrunch. Up until around January, ChatGPT had commanded over 50% market share. But by the end of May, it has fallen to 46.4%. And the article attributes that to the rise of Gemini, which is at 27.7%, and Claude, which is at 10.3%. And then everything else, Grok, Perplexity, DeepSeek, and Meta.ai have less than 5% market share. The thing that was interesting about the graphs of this show, if you want to scroll down a little bit for folks on YouTube, is that in terms of the actual recent growth leading up to May, and granted these are lagging the current month by a little bit, but Gemini had some uptick. But if you look at visually, Claude seems like it's growing quite a bit. Granted, Gemini is still at 30%, so Claude growing is relatively smaller. But it's interesting, and I wonder how much of that had to do with all of the sort of recent back and forth with administration, both around like the Fable stuff, because that's brought even more attention to it. And then previously, the run-in they had with like the Department of Defense or whatever they're calling it these days. So the other thing that I thought was worth calling out in this article is like less about two minutes to midnight, but just something we've talked about before, which is that users are increasingly willing to switch between assistants. And that data was courtesy of, I think, Sensor Tower was the provider for it. And then the other thing that's kind of wild is growth rates have decelerated in terms of both spend and overall growth. So they're claiming that that is likely a sign that the market is maturing, even as the absolute numbers continue to climb. And then the other sort of really telling thing, which is kind of comes on the back of what we talked about, was it like a month or two ago, right, where you I think you brought the synchium and like developing markets, particularly using AI a lot more than the United States. There's like a lot more haters in the U.S. than other countries. So Asia recorded the first download decline of 3.3% in Q1 of 2026. And then the last stat that I'll just drop just because I thought it was interesting was that 13% of Anthropix users are paying for a subscription plan, which is both less and more than I thought. It's kind of fun. I don't know how it manages to be both, but yeah. Are these only consumers or like businesses? Yeah, I believe this was like not enterprise, but I could be wrong. Still though, overall, like looking at their numbers with chat GPT having 1.1 billion monthly users, Gemini with 662 and Cloud with 245 million. Like we're talking 2 billion, assuming they're non-overlapping, which is definitely not the case, but it's a good idea of the scale of the whole thing. We're talking like one in every four persons on Earth. The TAM doesn't get much bigger than this. There's no 10x from here on out. Just read the SpaceX IP address. The TAM numbers for their IP address. The TAM can get bigger if you just keep expanding your ambitions. Yeah. All right. My news item for this week is NVIDIA is in the process to try and raise over 25 billion in bond issuance, the first since 2021. To put this in perspective, NVIDIA currently has something like $8 billion in bond outstanding. So they're raising, they're going to more than double their total outstanding debt after this issuance. And this comes at an interesting time because NVIDIA also recently dropped something like 5% over the last five days. So it seems like the market's demand for NVIDIA stock may be hitting a plateau or even a slight dip. But the monster must still be fed with more tokens. So they're now moving towards debt issuance. And that's not a good sign, in my opinion. Although I'm still pretty, I always forget which one. Bullish is like you like them, right? Yeah. Yeah, I always get bullish. I don't know why I get bull and bear confused. Think about bears are always sad. That's how you know. Bears, sad. Sad bear. Bulls are angry. I don't know. Anyway, I'm still bullish though because I actually think that their move into ARM stuff with Microsoft is pretty smart. And it will be pretty interesting to see how that goes. Because like, you know, we've sort of hinted at this last week too with the article that Rahul brought in where it was like, the writing might be on the wall for like having a personal agent that's running on your hardware and is just calling out to these other APIs. And I think if that's the case, having a machine that can run a halfway decent LLM on board is kind of smart, you know? Yeah. Especially if it's relatively like all things considered power efficient. And NVIDIA has like half of the revenue from AI when we say, you know, the industry has done half a trillion, like 250 billion or so that is purely NVIDIA making money. So if you look at it from that perspective, they're the only one who are making the margins and the profit and raising is not as concerning as let's say. But they're raising debt though, right? If they're printing so much money, why are they raising debt? And then I'm sure there's probably some like corporate-y reason for it. But like, yeah, I don't know. It does. It does concern me a little bit. It's the same question I have of why is Alphabet raising debt? Your money printing machine. Why is Facebook raising debt? You know, like it's a little concerning. All right. Until our last article brought to you by Rahul. Ed. Ed is back. Where's Ed? Where's your Ed at? Ed Zitron got exclusive access to OpenAI's financial documents that they submitted. And this has been independently verified by Financial Times as well. So there's not too much, any like he's making any numbers up. And he doesn't even, didn't even comment much on the numbers. He just like stated them as plainly as he could see them and leave, save his commentary for the next week. Things to note, there was a, in 2024, they had $5 billion in loss. And that shot up to $38.5 billion in 2025. 2025 is when they did the whole conversion from non-profit to for-profit. So there were probably like one-time losses that came as part of that. Their operational costs have been exploding. They went from $12 billion and changed to $34 billion in 2025. And they spent close to $200 billion on R&D. The revenue growth has increased, obviously. In 2024, they did close to $4 billion. In 2025, they did $13 billion. And there's this like tight interdependency that when Microsoft had published one of their quarterly reports, they had like one customer was the largest like consumer of their computer and stuff. And then later it was like, yeah, OpenAI is basically, you know, I will not fully draw the circle, but you know what I'm referring to. And they've gotten some money from SoftBank and Microsoft, but that was not even a billion. So what are we even talking about here? So yeah, this kind of like, they're planning to go to IPO. Financials obviously don't look neat here, but we just saw the SpaceX going to IPO on Vibes. Elon, to his credit, does have a great retail following, which is what the IPO was successful off of. So we'll see how people feel about OpenAI and can they also write the retail name or not. Is Sam Altman as big of a GigaChad as Elon Musk? I don't know. One other thing here that I did ask Gemini about is in both 2024 and 2025, they have this line item that says net loss attributable to non-controlling members capital. Like what the, what is that? And this goes from like, I think it was 4 billion or something in one year. And then, or no, it was 4 billion in the second year, close to 4 billion in the first year as well. I do not have enough financial acumen to be able to tell what that is. So I will read what Gemini said. It represents the portion of a subsidiary company's losses that belongs to minority shareholders rather than the parent organization. So 4 billion goes to that as it lost. Interesting. Ed also pointed out that OpenAI has just over 50 billion in assets with almost half of that in cash. So if you compare the 50 billion in assets to the, I don't know, 50, 70 billion that they spent in expenses last year, does not paint a great picture of their cash flow. Or the around right. And they didn't raise, I guess they had gotten some like compute commitments and they had the whole like, oh yeah, it's close to 900 billion or whatever in valuation. But the actual money they got was not that much compared to that. And a lot of it was just like, yeah, we'll give you compute. And if you hit some certain targets, we'll give you some more. But they didn't get all the like cash on hand or something. Just totally way outside my wheelhouse. But I immediately want to start like modeling that. Where will these lines intersect? Come back then for some fireworks, probably. This probably also explains why they're going IPO pretty soon. Because they probably need that cash injection. The cash. Yeah, I know. Otherwise. Yeah. And unfortunately, the only signal we have right now is really the SpaceX IPO where like they were up quite a bit. Now it's kind of like tapered off a little bit. I think they're still over the initial list, right? By a significant amount. Yeah. Yeah. But it's definitely some of the initial fire has chilled. No, Shimina came back. They started at 160 and they're currently at 156. So they're compared to the initial price. Yeah, but it was like over two at one point. I guess is what I'm referencing. It was kind of like. Oh, yeah, yeah. Yep. It did. And then this last week, it just went down into the right. Yeah. So likes it. Yeah. Down into the. Oh, no. All that said, how do we feel about the clock this week? I know it's probably super boring to. I don't know what's worse that we flip flop every week or that it's boring to be like, we need to wait and see because it's about to get super exciting. But that's really where my head's at is. Where my head's at. That's right. Come back next week for more puns you can shake a stick at. I think we got three kind of bad news in a row this week, you know, between NVIDIA needing to raise bond. OpenAI is losing more money than we thought. And at the same time is losing market share. Like that is not comforting for me. Yeah. It's fair. I would actually push to move it forward to like 510, 515. Okay. How about just five? Okay. Is that too big of a jump? I do. I do feel. We've been much closer. Yeah. Yeah. But like originally I could have, I could think like OpenAI maybe has it till like November or something. They go IPO in like October. But now it feels like they're actually burning a lot of money. So my gut says five sounds good. Okay. Let's do five. All right. Five minutes of this. We'll find out all this within the next four months. That's true. All the IPOs are likely going to happen before the midterms. Yeah. Cool. That's exciting. And people, you know, if they don't do that, they'll spend all their money on Black Friday deals. Who has money to buy stocks? I'm excited to use for our November 15th episode where we digest all the happenings from early November of this year. We got to get Nathan back for one of those too. That would be great. But as things stand, when they put the clock at five, and as always with the clock being set, it comes for the end of the show. So thank you listeners for joining us for our study session this week. If you like the show, if you learned something new, please share the show with a friend. You can also leave us a review on Apple Podcasts or Spotify. It helps the people to discover the show, and we really appreciate it. A-D-I.