← Back to search
Google is Not a Serious Company
Nerd Snipe with Theo and Ben · 2026-05-28 · 82 min
Show full episode description
Not only did Google accidentally ban Railway's account, but their new flagship model Gemini 3.5 Flash is absurdly bad. Oh and apparently Theo's building his own cloud... Thank you Macroscope and GT for sponsoring today's episode! Macroscope: nerdsnipe.link/macroscope General Translation: nerdsnipe.link/gt SOURCES https://x.com/theo/status/2057359424378097823 https://x.com/JustJake/status/2056881510939283776 https://x.com/KorduGG/status/2059141337895604626 https://x.com/unboringtech/status/2059144145273491610 TIMESTAMPS
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Challenges the narrative that
Google is an unbeatable AI juggernaut, citing weak models and unreliable cloud.
Benefits
- Honest critique of Google's AI execution
- Clear take on where image/video gen actually has value
- Reliability lens on cloud infrastructure choices
- Context on China AI lock-in and IPO turmoil
Use cases
- Midjourney dropped from ~45% to ~8% of survey respondents
- GPT image two used to generate dashboard/UI designs for code models, not human viewing
- Google accidentally deleted Railway's entire account via auto-ban algorithm
- OpenAI's Peter burned $1.3 mil on random OpenClaw pull requests
- Omni model's one real use: video-to-video (add fire, day-to-night)
KPIs / results
- Midjourney ~45% to ~8% of survey respondents
- $1.3 mil spent on OpenClaw pull requests
- ~4 years warning about Google Cloud reliability
Tools / build
- Gemini 3.5 / Gemini Omni
- GPT image two
- Google Antigravity
- Google Cloud / Firebase
- Personal cloud platform (host)
📑 Chapters — tap a time to jump there
00:00
Gemini fallout
- Crashing out over the Gemini 3.5 launch psychosis
05:21
Video gen
- Video gen value debated; video-to-video the one real use
- Most LLM/image output never consumed by humans
10:10
Google Cloud
- Google Cloud reliability frustrations, clown thumbnail
- Railway account deleted by auto-ban algorithm
14:04
Windsurf/Cursor
- Windsurf and Cursor editor landscape
25:10
Manus/Meta
- Manus and Meta breakup and what it signals
30:00
China lock-in
- China AI lock-in and open-weight ecosystem risk
45:00
Cloud workflows
- Cloud workflows and self-hosted services
Can we crash out about Google? Please? Please? Fine. I feel like you're going to crash out about a lot of things now that you have how many threads that you've done in the last 24 hours with your Hermes agent inside of your new Discord server? Oh, too many. I mean, we'll get to that maybe at the end if we have time. Yeah, we need to not do that one first because if we do, it'll take way too long. We need to start with the stuff that I was having at Surface, which is a full list of all the tweets and things that happened with the Gemini 3.5 launch. Very full psychosis. Yeah, I am in total psychosis. I feel somewhat guilty for this. It is what it is, but that's not the only thing we need to talk about today. We will crash out at Google as they very much deserve it. They have been better than I expected at responding to me crashing out at them, which has been a very good sign. And they didn't demonetize me for it this time, which was good too. We also have other economically scary things though, like the Manus and Meta breakup, which I know Manus might not be the hottest topic right now, but what it represents is actually really scary. How it kind of shows a widening rift between AI development and research in China versus the rest of the world. It also might signify the end of the open weight model, like wonderful ecosystem that we've had thanks to the development and research going on in China. We also have the IPO chaos going on right now between the SpaceX IPO, sorry, SpaceX AI IPO, the anthropic potential IPO and the raise they just did, as well as open AI's filing. And I think you found some interesting stuff about some rules that got changed. Yeah, there's a lot of weird shit going on there. I mean, well, well, it is weird shit. It's probably corruption, but we'll get to that. Probably. Joy. Anyway, yeah. If we're going to be surrounding ourselves with incompetence, we should have one additional incompetent topic, which is the cloud that I've been building, as well as other services I've been throwing out myself. I wrote a lot more code last week than I have for like months and the results are chaos, but fun. And it seems like you're actually liking it. And we've deployed or afforded apps on my cloud since I started it. So it's in a cool spot. We'll have a lot to share there. But first, let's get started. Indeed. Also, reminder, check out all the platforms that you like listening to and watching podcasts on. There's a bunch of different ones. I know some people like Spotify. Others like Apple Podcasts. Give us a rating if you can. Apparently, that does help a lot. Really pumped to see the awesome ratings you guys have been giving so far. So thank you all for that. Sorry, I have to do this. If I don't, then Alyssa will continue to hold my cats hostage. So I'm stuck doing it. Now let's crash out of Google. Yes, sir. All right. So I've been looking forward to this one because, look, I don't have a deep vendetta against Google. I have always just disliked them and memed on them and dunked on them for the love of the game because I think it's funny. I always have got like the anti-gravity thing last fall when they released their very special cursor competitor, which was a very awful platform. It's supposedly gotten a bit better, but there's still a lot of weird catches with it. And I don't know anyone serious who is using it for anything serious. Linus Torvalds. Oh, fuck. You're right. Okay. But like, is he doing anything like, uh, no, like I feel like he's using it for like playing around with it. He's not developing the Linux kernel in anti-gravity. That's not happening. There's no chance in hell. But, um, other than Linus Torvalds screwing around, I don't think anyone is actually using anti-gravity for anything for a good reason, but I don't actually have any malice against Google. I want to preface all of this with that. I'm more just tired of this narrative that I keep seeing over and over and over again about how Google is this unstoppable, unbeatable juggernaut. Just, just you wait for the sleeping giant to awake guys. They have all of the things. Look, man, I get it. I too have gone to artificial analysis.ai and looked at that little chart where it shows Google has all the boxes checked and no one else has all the boxes checked. It's a very pretty chart. That doesn't mean Google is competent. That doesn't mean that just because they have all of the things lined up, they can actually do anything useful in time and time again for a year plus now, really since like what two, five flash, I think was the last really good one. If I thought it was okay to our flash was the really good. Yeah. Since then there has not been anything notable or useful on the frontier from Google. And I'm just tired of listening to this thing over and over again, where anytime you talk about any lab, you just have to deal with the stupid Google comments. And I hope we can finally start putting those away. I know we can't. The people are going to be in this video, but Gemini three, five flash was a disaster. That was not a good model at all. But, but, but Ben Omni Gemini Omni. Oh, this is a fun one where, uh, this has happened with Google stuff before. Cool idea. Poor execution. Someone else is going to do it better because the Omni model for my understanding is it is the idea of anything in anything out. You can send it a video and it'll give you some text. You can send it some text to give you a video, whatever you want to have it do. That's a very cool idea. I haven't seen anyone actually use this for anything real. I don't even fully understand how it works. There's one actual thing that this is good for and only one. And what is that? Video to video. Ah, so like it's not actually useful. Yeah. Like you can take a clip of us talking here and say, put fire in the background and now there's fire in the background. Yeah. Yeah. No, that's true. I actually, I was listening to, I think it was like Wancho. They were talking about Google IO and they were talking about like the demo they did of like taking a shot from daytime to nighttime. It actually did that pretty well, which like, cool. That is actually a very good use case of video gen. I'm trying to find as many useful things for image or video gen just because like, I still am not fully convinced that it matters. You need to do something with your GPUs. No, I have. I don't want to do anything with it. I'm just curious. Like what is the actual use for this? Because I'm very curious right now. What are the economic values of all the different AI things we're doing? Text is incredibly obvious. That's insanely valuable. Image gen I've come around on, especially with GPT image to video gen. And still haven't fully come around on for anything other than maybe that. Like that's pretty cool. B-roll. Yeah, B-roll. I've said this for a while and that's one of my like favorite theoisms that like, like the majority of LLM generated text is never read. Most of it is reasoning traces that we don't even get at this point. We just get summaries at best. And then it generates things like outputs like code and whatnot that we barely read half the time. The vast majority of text generated by LLMs is not read by humans, but that doesn't make it not useful because that text can be used for anything from executing code to reasoning to make the model responses smarter to solving things in the background and deciding what the model should or shouldn't do in different scenarios, the tool calls and whatnot. What good is an image that's never seen by a person? What good is a video that's never seen by a person? It makes no sense to be investing similarly heavily into those things. And when everything gets trumped as quickly as it does, it makes even less sense. Like I just saw a survey of like what image gen tools people are using in mid journey dropped from like 45% of survey respondents to like eight. Holy hell. Damn. And they were like the frontier. They were so far ahead and now it just doesn't matter. Are they just not caught up at all? I haven't paid much attention. Nobody's seeing a tool that is just image gen now. So like if you already have a chat GPT subscription or anything else, it just works like that. Unless you use a quad subscription where you can't really do anything but generate shitty code. Yeah. I think the one thing that's changed on this is that with GPT image two for the very first time, I am actually having the image gen generate images for the models, not for me, where before it was entirely just human consumption. I'm doing a lot of image gen, not for me to look at, but to take a design spec that is nicely written out and put together a better design than what the text models would do because the way they train this thing, it's incredible at dashboards and UI and stuff like that. And that is a much, it's really useful for guiding something like five, five into making more common stuff. But you do look at the image it generates before sending it to the model. Yeah. Almost always. Thereby, by statement stands. It's still in the loop, but it is also still like not, I'm not consuming it from an artistic standpoint. I'm consuming it from a business standpoint of like, I need to make this thing exist within the code gen. Informational like chart generation, I would say is similar. You're not consuming it for the artistic value. The point is that there is no value at all without a human looking at it at some point in the chain. Yeah. Yeah. That's fair. Yeah. And it just doesn't really click for me. How, like, why would you make the image models better at doing that? So you can do it without a human in the loop when you could just make the LM better instead. Like it's just, there's no incentive to continue iterating on these things. I'm surprised we've gone as far as we have. And I'm also like, like, I think GBT image two not being a huge moment is kind of the end of the heavy investment in the space. Like probably. Yeah. Sora's dead. Image two is good, but not getting very invested in and people don't seem to care that much. We haven't heard a thing about nano banana in a while, which is still like the most embarrassing Google ism that like they had that code name and it was better than Google's normal naming. So now it's just called nano banana forever. It's hilariously Google. They walked right into that one. Yeah. The one thing I do think image, the image two could be used for is helping make better front end data for opening eye because they need it desperately. Like the fact that their image gen is better at front end than their model is definitely a temporary thing. That's not going to last forever. But for right now, that's the case. And if they can use this to help pick up those capabilities in future iterations of GPT five, I don't know if I still agree with the whole front end is better in other models than GPT five five at this point. I think all of them are so done to death that I would argue none of them are really good at it at this point because a lot of the demos everyone was doing of front end stuff was like, oh, here's like the landing page it makes and all that stuff. And we are so desensitized and so used to all little quirks of LLM generated UIs that even the front end design skill from Claude, which was really cool three months ago, is now a smell because we all know what it looks like. I am proud that I've managed to wrangle decent designs out of all of these things. You just have to like beat out the Claude and codecisms with a stick. You just have to like work at it. You have to go hard on it because like the whole point of design is it needs to feel novel and different because there are a lot of designs that I will see these models output that four years ago I would have been like, wow, that's really nice. That looks really good. But since I know where it came from and users will know where it came from, it feels low effort and therefore it's not good. That's it. That's fair. I want to crash up more at Google specifically on the Google Cloud side because this is where I am most frustrated. I care a lot about reliability of the things that I built. That's just always how I have been. And one of the scarier parts of betting infrastructure on top of things other people built is that you have to rely on them to make things reliable because if they don't do their job, they don't keep the thing working. You and your users are screwed. And in the end, you're the one holding the bag. So I take these things really seriously. And I've been trying to raise alarm bells about Google's bullshit on the cloud side for like four ish years now. I'm pretty sure it was when I did my first video on this. No, I know the thumbnail you're talking about, the clown one. Yeah. Yeah. That was a horrifying thumbnail, by the way. Thank you. Yeah. I made a video where I had two clowns and I replaced their faces with the Google Cloud logo, the Firebase logo, because fuck it. Just none of it works. It's well deserved. Yeah. I, there are outages that haven't even been publicly confirmed to be GCP related that I wish I could talk about from really big companies that like took the fall themselves and shouldn't have because they don't want to burn more bridges with Google. But the recent one was Railway because Google just accidentally in quotes deleted Railway's account entirely because they set up some new algorithm for detecting accounts they could possibly close to free up computer and accidentally banned theirs in the process. It's insane. Unbelievable. Like a, how is there no human oversight? But even like, even if we're going like full agentic bullshit here, like if Google is going to be this frontier AI lab, how does open AI have like Peter going off burning 1.3 mil on random pull requests for open claw of all things. And they don't have a very simple Gemini script to just maybe gut check this algorithm for if this account is spending more than, I don't know, a million dollars a year a month or sorry, a million dollars a month. Maybe we should look at that a little closer. Maybe we should double check something is real. So this is a great plan, but there's only one issue with it. You have to use a Gemini model to do that. And they probably did. Oh, I could see that happening. Yeah, that's entirely possible. I mean, considering the outputs I've gotten from three, five flash, I could believe it. What if the first victim of Gemini on GCP is Railway, who's been trying to move off of GCP for like a year now? I hope that's true. I have no idea if it's true. We might just be lying on the internet, but I hope it's true. I don't care. Google does not deserve the good faith of being like explained truthfully when they are not like operating within reality. No, no, they're just not. And like, I don't even know if it's like active militia because Anthropic is a special, special place, but they are not. I treat them in Google very differently in my mind. I just think of Google as like the three ring circus. And I think of Anthropic as slightly malicious, kind of concerning, but also pretty useful. So we tolerate them. Yeah, I don't know where I'm going with this. I do know where I'm going with it, but fill something in. Yeah, they're pretty screwed at this point. Like, I just don't see a path to recovery. Like the way the company is so big and fragmented, like to compare with XAI, for example, I legitimately think they could have a comeback, whether it's through the acquisition of cursor and all the incredible researchers and data they get with that. Yeah. Whether it's the fact they have so much compute that like there's a like they can use that and do real things with it. And Google is very clearly and optically in a compute crisis at this point, similar to a lot of other companies. Whether it's the fact that Elon isn't going to let corporate bureaucracy prevent development from happening and is willing to like break down doors to make teams work together or even just fire and get rid of teams. Google's bureaucracy prevents anything good from happening, including and not limited to the Gemini CLI open source project being shuttered in favor of an entirely broken CLI with the anti-gravity CLI. One, that acquisition was a total fucking mess. I don't I'm it was a while ago. So I assume most people have don't remember this. But when they bought Windsurf, they didn't actually buy Windsurf. They just bought out. I believe the founders and the IP is all they actually got. Or is the number public? I don't remember the number. It was billions. Yeah, it was in the billions, which is unbelievable. And the employees that stayed, which was most of them got jack fucking shit for it, which was actually insane. Like the dirtiest acquisition ever, just genuinely scummy stuff. I'm still mad at the founder for how he handled that, especially with me, where he refused to talk to a mutual friend anymore because the mutual friend asked him a question on my behalf. And he said, why would you ever talk to a cursor investor about us? That's so gross. Such a pussy. It's like disgusting. It is what it is. But like, I honestly think Varun is going to run the company to the fucking ground. Like he is the worst thing that ever happened to DeepMind and Google. And it is already showing. And they are making even worse decisions as a result of having him. So yeah, no, it was screwed. Awful acquisition. And frankly, I honestly like am super impressed with like the windsurf team and cognition because those two windsurf went to cognition after that. Those who remained and they have been maintaining it since they've been doing some really cool stuff with their models. I forget. What are they called again? It's the SWE SWE. SWE. Yeah. The SWE models with Cerebrus. Like I'm not a huge Cerebrus guy. I'm not a fan of the way their models have to be on them right now. But they're doing cool stuff. It's a good editor. It's doing well. They actually made something cool out of it. So full credit to that team. I think that is I'm glad that worked out for them. And just desserts for Google. Honestly, like this is so unbelievably bad. Like I had forgotten about this. But when I was testing out 3.5 Flash, I was just using it in Pi. Very basic harness. Very basic question. And it got stuck in a reasoning loop where it would just reason the same thing over and over and over again in a loop and be like, oh, I keep saying this over and over. I need to stop doing this. And it would never stop doing it. The only other model I have seen do this is a like 10 gigabyte Quen model running on my 5090 at home. Not from the three, the frontier like flagship model from one of the big three quote unquote labs. I'm going to show you a benchmark that hopefully will be pulled by the time this podcast is out, but probably not that I am very excited about. You don't say the company who did this. I just want to show you because somebody finally made a benchmark that properly measures how incompetent a lot of the shit is. It's a new software engineering bench on really long running tasks that will hopefully be out soon. GBT 5.5 and 5.4 crush the top. Opus 4.7 is roughly tied with 5.4. Sonic 4.6 is a giant drop, like half the score of Opus 4.7. And then 3.5 flash is slightly behind that. Then we get to 5.4 mini. Kimi K2.6. Then another meaningful drop to Mimo V2.5, GLM 5.1. And then yet another drop of 50%. So we're going from scores in the 60s to the 70s to a score of 10% for Gemini 3.1 Pro. Oh my God. This is a rough looking chart. I'm so excited for this to be public, man. Same. I mean, this chart, to be fair, it's just confirming my already existing opinions and biases. But that chart looks exactly the way my current read of the industry is, which is that there are two labs that are miles ahead of literally everyone else. There is genuinely no true competition for OpenAI and Anthropic right now. The only one who I think kind of has a chance right now is actually XAI and Cursor because I have seen them actually pull impressive stuff off. Like Composer 2.5 was a very impressive release. XAI has proven in the past that for the time period, Grok Code Fast was actually quite cool. There's something in there where they could actually make this happen. And they do, again, have all the hard ingredients you need of compute, money, data, etc., etc. What about China? I have a lot of opinions about China now because for a while they've just been like the open source model farm. But it seems like that's coming to an end. I'm just asking in terms of who's more likely to win, China or Google? Oh, China or Google? The thing about Google, and this is the reply I keep seeing a bajillion times, is like just trust. Like, yeah, it's still bad, but like eventually they'll get their shit together. Just trust. It's going to happen next time, guys. Trust. And it never has happened. Like it's been long enough that like everyone else is figuring this shit out and they still can't get tool calls right. It's concerning. I'm not seeing if you look at the trend lines. Yes, Google is still on paper ahead and they still are ahead. Like I would rather use 3.5 Flash than a Chinese model right now for most things. I would take K2-6 over most of the Chinese lab models. Or most of the like Google models. Like I can't imagine using Google model over like even K2-6. Okay, that's fair. I would say K2-6 I would take over that, but I can't really think of any other ones that I would, at least that I've used extensively. But even still, I don't think that if you just look at the trend lines, the China trend line is positive right now. They're getting better and better. Google's just not. They're just floundering in a weird place and maybe there will be a breakthrough in their post-training and they'll finally get something good done. So your rankings right now are roughly OpenAI, Anthropic, maybe XAI. Maybe. Yeah. China, Google. Where? The real question though is where does Meta fit in this? No, I have a better question. Where does today's sponsor fit in this? Right now. Today's sponsor, Macroscope, is an AI code reviewer. And just wait, wait, wait. Don't skip this. We are not going to be talking about AI code review. I want to talk about the other things that Macroscope does. Because while their code review is great, it's fast, it's accurate, and I have had zero interest in turning it off, which is a pretty impressive thing for these tools, to be honest with you. I think the best thing they do is the stuff around the code review. Like one of the things is their dashboard is actually very, very useful. They keep track of everything that's happening in your project. So you can quickly see who's been working on what, how long they've been spending on it. Get a nice little summary here of everything that's happened over the last few days. But my favorite feature at Macroscope is their check run agents. They're little markdown files, kind of like skills that you can add to your projects. Like on one of my projects, I have this one, which is a PR labeler. It's basically just reading through everything that happened within that PR and then applying a label to that PR. It's a simple one, but it's really useful to just have these labels automatically attached to every single PR. But because you're just writing a markdown file for an agent to run and it's got a bunch of tools to interface with the PR itself, it can do a ton of different things from audits to labels to a lot of other stuff. It's just a really useful thing to have on your repos. And honestly, that's just how I felt about Macroscope in general. It's just felt like a great addition to all of my projects and all of my teams. I wouldn't want to be working without it at nerdsnipe.link slash macroscope. Proud of that one. That was very good. Yeah. Can an LLM do that? Possibly. Probably. If you give it enough time, I could probably figure it out. A meta model could never do that, though. Yeah. Sadly, yeah. I want their models to be good just because it'd be funny. Have you ever tried MuseSpark? Yeah, I have. You have? You're the only person I've ever met in my life that's used MuseSpark, including my meta employee friends. I was unimpressed. It was weird. It was very weird. I didn't use it enough to form any strong opinions other than it did not pass the first sniff test and I stopped immediately. Obviously. So. It seems like what happened at meta as well, where they now have a truly disgusting anthropic bill. Yeah. Yeah. Because they're just all on anthropic because all of their employees are using anthropic to try in. I don't know. Like, what are they even trying to do? Are they trying to make models from scratch? I don't know anything about what's happening internally there other than something's on fire. They're trying to collect a shitload of data from employees using models and then use that to refine models and like make something better. Finally, I don't think you can get enough on the like spend thing, though, because it's actually a thing I forgot on the notes that I think is hilarious. Did you see as part of the SpaceX IPO that their monthly revenue from anthropic is now known? Oh, no. How much money do you think anthropic is paying monthly for the over the data center? Yeah. Oh, I do know the answer to this. It's like what? 1.5 bill or something? 1.25 bill a month. A month. Yeah. Yeah. Yeah. Yeah. That's a lot of cash going to the XAI server. That's just one of the various like compute sources that they are spending a lot of money on right now. Oh, this stuff is so... And then like their revenue is good, but it's in like the three bill to four bill for the last quarter range. Yeah. So that's like a bill and a half to two at most per month revenue. And they're already spending the majority of that on just the SpaceX compute. Yeah. And then they spend way more renting other shit. And then that doesn't even account for all of the research they're doing, all the acquisitions they're doing, all of the like everything else in the subsidizations. Like they are losing so much fucking money. IPO soon. Yeah. They are going to IPO soon. And we're going to have a... I'm very curious how the markets are going to react to three horrifically unprofitable companies going public at some of the highest valuations in history. Like I'll be honest, like I'm full psychosis. I see the vision. I would buy the stock in these companies. I think they're going to do incredibly well. And I see the path from the bullshit we have today to super machines and like actual AGI type thing. I can actually see that path now. Is it a bubble in the sense that it is overvalued in the short term? Yes. But is it overvalued in the long term? I would honestly think in some weird ways it might even be undervalued in the long term. It's hard to know in a lot of ways. I do think IPOing right now is insane when you could like raise more private capital. I think that they inflated their valuations too quick and didn't do enough like... I think OpenAI strategy of doing like the long term spend deal instead of like a short term compute acquisition will benefit them greatly because they're like rolling economics is much less aggressively likely to bankrupt them than like an immediate acquisition or like massive spend. Yeah. Like Anthropoc is going to spend and lose so much fucking money this year. It's insane. OpenAI won't be that far behind if they are at all. But it's insane. Well, and the thing too, the only problem with that is like I don't know the economics of raising money when you're worth $500 billion. But is there just that? Is there enough private capital floating around out there to support a $2 trillion round for Anthropoc? They needed to not like bump their valuation so absurdly constantly. And like the $2 trillion round isn't affecting how much money they raise. It is just the like upper bound of like the company valuation. Right. And like they're raising a couple bills to 100 bill at most at that point. Like the round that Anthropoc just did at the 900 bill vals only raising 30 bill. And like only 30 billion is a crazy thing to say, but that's like only 30 billion relative to the 900 because they're trying to take as little dilution as possible. I get the thinking there. I like not diluting companies, but you could just like lower the valuation a little more because if they had just done like 500 billion, you could probably squeak out one more round before having to go do this. There's a classic clip from Silicon Valley about this. Like, did you ever consider raising less money or like having a lower valuation? And then he has an existential crisis. Yeah. Yeah. Yeah. This is happening now. That said, the usual solution here is acquisition or IPO. And as we have now seen with the Manus meta breakup, that is not a valuable or like possible path either. I've been paranoid about this for a while. I am the one person who, and I think I've been proven more right on this, was against the Figma Adobe blocking of that acquisition. Adobe wanted to buy Figma. They went through all the agreements, all the process, all the everything. And then enough regulators got in the way after they had like inked the deal that they had to pay this massive breakup fee to Figma because they couldn't go through with the acquisition, which is the last major whirlwind of cash Figma ever saw because their stock is down at 95% since they IPO'd. Yeah. And it's the last they like realistically can ever see because now that they're public and like, unless the markets flood them with cash, which they're probably not going to because I haven't looked at their quarterly earnings or anything like that, but I doubt it's all that impressive. I do think I did see a post from their CEO that like their revenue is up right now. Like it isn't growing, but it's not, it's not insanely impressive for a huge public company. Like it's not Apple sized and they would have been much better off going into Adobe, fixing Adobe, taking, honestly taking over Adobe. Dylan should have been the new CEO of Adobe. He could have saved so much trouble. Yeah. Like that would have been a net good for everybody if that had happened. And now they are screwed and both Adobe and Figma get to die as a result. Yep. And now we have to talk about Manus and data. As I mentioned before, most of us probably don't use Manus. It's my understanding that Manus is a deep research tool similar to like the chat GPT promos and other research assistants where I can just go off for hours looking at data, finding resources, trying to like actually analyze a complex problem. Cool. Interesting. I don't think that was worth 2 billion when the acquisition happened last year. I think it's worth much less now since these features have been folded into all of the agents we're already paying for and using. Manus getting a deep research thing working okay with high end models is not that impressive when my subscription on chat GPT already comes with 5.5 Pro and it does that itself too. So. And it's not hard to make these flows. Like I've built out for myself similar deep research quote unquote flows for my own internal stuff. It's not that hard to make. I want to go through some of the fun details on this one because I did not know all of this before researching and I think it makes this a much more interesting case overall. The first is that the company was founded two months before chat GPT was even launched, which is wild. It was founded in China, but they were much more interested in markets outside of China and investment outside of China. They explicitly refused everything from a ByteDance acquisition in 2024 to a bunch of government investment from Chinese parties that could have allowed them to be more successful there very quickly because they really didn't want any of these things to affect their potential to be in the United States markets and other Western marketplaces. They got super inspired by cursor according to them and the actual interviews to go further with what you could do with these AI systems from other companies to build a better, more compelling product on top and then made Manus at that point. They released it in March of last year. The demo video went viral. They started doing decent numbers of people subscribing enough that people were interested. They then ended up trying to raise more money at a 1 billion or so Val. At that point, Meta noticed and tried to acquire them, which would have pulled them out of China and made them part of the many attempts at Meta to do AI properly. Remember, this is a company that explicitly tried to not tie itself too closely with China in the Chinese government. They wanted to be able to sell to Western companies. They wanted to raise money with Western VCs. They wanted to be a Western company effectively that was formed in China. And this is where shit gets scary because Beijing undid the acquisition. They didn't block it. They didn't try to stop it in late stages. The acquisition was over. The employees had been onboarded. They were now part of Meta. And Beijing came in and found some crazy policy that allowed them to undo the acquisition. Yeah. So now Manus and the IP and all of that dies unless the founders can raise enough money to buy it back from Meta. Yeah. They're trying to raise about a bill to get it back at around the original two bill, This is just insane. Yeah. Look, does the U.S. has plenty of problems? We it's not perfect, but this wouldn't happen in the United States. There are things that can go wrong with like it. It happened with Adobe. It did happen, but that is slightly different than just like the CCP almost certainly going in and being like, no, and just making up whatever they want. And sorry, you can't do that. You're going to do this now. You have no choice. What I think is much scarier here is that this is not a company that is necessarily valuable to China. Like this isn't like, oh, Manus is this really crazy AI company. Them being in China is really important because we need better AI Chinese companies. It's not that this company has been out of date for a while now. I would be surprised if you couldn't do similar things with like Quen and whatnot at this point. The issue is very specifically that China is preventing these companies from working with the West the ways that they might be interested in. And this is China effectively choosing to kill this company instead of letting it go to meta. And this seems to be the new MO. A lot of these open weight labs that were making all of these nice open weight models are starting to close the models and delay the release of the weights if they even put them out at all. The newest Quen model is closed weight. And also an interesting announcement that did not like really compare against modern things. Like it just ignored the reasons of Opus 4-7 and GPT-5-5 entirely in their like output and like benchmarking, which is insane of them. The fact that like these companies are also now very upset with things like Cursor building so heavily on top of Kimmy K 2.5. I honestly think that's probably like the genesis moment of like the China coming in and locking down harder. They're like, no, we are not giving free rides to American companies anymore. We are going to lock everything the fuck down if this is war. And I do think like the shot heard around the world equivalent in the AI space compared to like previous literal wars is possibly going to be the moment where it was revealed Composer 2 was built on Kimmy K 2.5. Yeah, I think so. Because now like just to be super real, this is actually a pseudo Cold War type thing where like this AI stuff is pretty existential. And whether or not whoever wins this, China or the US is going to have a tremendous advantage against the other in whatever ends up happening over the next 50 years. This matters a lot. And the fact that they are now seeing stuff like this as a threat, they very, this is a finger to the West. They do not want Western companies to be benefiting from their stuff. And as a result, we're going to lose a lot of big open weight models like Quen is now being closed source. We haven't gotten anything else from any of the other labs recently. And I hope we do. I really hope we get more. I really wanted at least one more generation out of these things, but it seems like we might not get that. It might be kind of the end of the Chinese open source stuff because they don't want a Western lab to yoink that like base model and then bring in all of our insane RL data because we have substantially better like coding agent use data from Cursor, Anthropic, OpenAI, whatever. They don't want us to get their base models anymore. And they're still very interested in stealing our data however they can. Like the proliferation of these like ghost networks of effectively like inference pooling where they get hacked accounts. They get accounts that they are getting like credit card fraud through and all of these types of things, putting them in a giant pool and then selling people discounted cloud code subs that use cloud code that are just routing through their crazy pool, saving all the data on the proxy layer and then using that for their RL. Like that is absolutely happening. Yes. The Anthropic article about it was still the cringiest thing right in a long time, but it is real. Yes, absolutely. I think we all, we said this when that came out. The problem with that article was the numbers they used were dumb, but if they had just said, hey, they are doing this. And then if they literally just found an example of one of these like shitty back market things that is proxying these requests and selling these tokens for cheap, like they can find that, find an example of that, put that up there. That is compelling. That is a real problem that they need to figure out some way to guard against to just not give away this coding agent data. It'd be really funny if they found a way to flag these accounts and intentionally put like shit traces in to see if they can. Oh, that'd be great. Just poison the data with like, just route all of them to like Haiku 3 or something. Maybe this is what somebody is doing to the Google models. Maybe there's some rogue employee internally that is intentionally polluting all of their training in RL to make the model so stupid. But what do they just have like crazy shorts against Google or something? Like what, what is the master plan? I would do that out of spite. I mean, look for the meme. That's no, I would do it for the love of the game. Like that would be hilarious. There's like one individual. It's like, ha ha. Watch what I can do. I don't know who that is, but I like that guy. He exists. I hope he exists. I'm going to believe he exists just because it makes me happy. Back in 2024, 25, it feels like a lifetime ago. The thing that everyone was worried about was cursor stealing your code base or whatever by using their editor. Everyone was super worried about the code being the valuable thing. And it turns out that's not the valuable thing. The valuable thing is the actual coding agent runs you do with these. Did you read the article that they put out on Composer 2.5 and like the new techniques they built for it and for training with it? Because some of the shit they came up with was actually really cool. They would take features that were implemented in the code base and remove them and then reverse that with like a fake chat log and like generate that and then use that in RL. That's really cool. Yeah. Like they would revert features with modern models, collect all of the data for that and then invert it and use that. It was really cool. They also did some fun things where they had the teacher and student model method for training where the teacher model would have extra context for things they noticed the model was doing wrong. But like at a certain point during the like tool call chain, it would call the wrong tool or call tool in the wrong way. They would give explicit instructions to the teacher model. Don't do this this way. And it would pull the weights of the model. It was training in that direction without giving it that context. Okay. That's really cool. There was a really cool shit. Like again, we are sleeping on cursor. They have a real chance. Like I actually would put cursor, especially cursor plus XAI. Hard to agree. It's like third most likely to win. Same. Like I've, I like again, Grot Code Fast is a meme and all of this stuff. Like I talk about it a lot because it's funny, but genuinely there is a chance here. At this point, I would legitimately consider like Microsoft or Walmart or T3 tools is more likely to succeed than Google. Like it's pretty nuts. Because again, it's like they have all the pieces. Like you could sit me down and you could put all of the pieces I need to construct a nuclear bomb in front of me. And it wouldn't do anything just because everything is there that is needed to actually make the thing does not mean that you are capable of orchestrating all the pieces together. And because of the way Google works, they're not capable of doing something like the Codex desktop app. The Codex desktop app is a huge effort between a shitload of teams, the model team, the actual app design team, the computer use team. All of these have come together to make something genuinely really, really good. That's not happening at Google. They don't talk to each other. I think the difference is that open air, like even to this day, these aren't so much teams as like people and individuals. Like there's a person who is the person for computer use and that person is fucking God tier and they are just folded into whatever team they need to be in to make the thing happen at any time. There's a story I heard recently about how Vercel operates. That's been haunting me since there's a generic unblock me channel in Slack at Vercel. And whenever you're working on something there and you are blocked by some other team, some process, something you can post about it there. And the whole company treats it as like a P zero. Do this right now. And whatever the fuck is blocking you, like the team gets pulled in and you get unblocked. That's sick. Yeah, that's really cool. Like that is something I could never in a bajillion years imagine Google having. And that's the problem. That's why I will give Google a little bit of credit here. They are starting to have to do this and they are embracing it. It's a silly example, but like when my channel got demonetized, I didn't have good contacts at YouTube to yell about that with. I had one. I didn't want to burn that contact. So I instead crashed out the people who I consider not responsible, but more in the position to do something about it, which was my friends on the deep mind and Gemini side. Oh yeah. And they escalated fucking crazy internally. And I honestly think there's probably a flag on my account that's like, be very careful demonetizing shit on this channel now. And they've been good about it. Like it, they cared. They did that right. It was a massive fuck up on YouTube's part, but it wasn't YouTube's fuck up. Not deep minds that a human flagged the video as dishonest and like teaching hacking or some shit. It was insane. They did that, but like, yeah, they care. And the sheer volume of people at Google, be it like lower level ICs or straight up fucking like exact level that have told me directly or indirectly that they watched the video, learned a lot from it and plan to implement a lot of those things. I have reason to believe that video was treated as like a P zero moment at Google and they're waking up to the problem some amount. At the same time, a lot of my favorite people at Google have since left for other labs. It is what it is. Yeah. And I want to be like crystal clear. We're not attacking individuals at Google. There are very, very talented individuals at Google. The fact that these exist at all, like building a model is not an easy thing to do. They've accomplished great things and there are great individuals there. It is the broader. I'm not inside Google, so I don't know exactly what to call it or how to explain it, but it's culture or process or just the giant behemoth is rotting. That's the problem. To go back to like the building a nuclear bomb example, if they had all the parts in front of them to build a nuclear bomb, but in order to touch three quarters of them, they had to go through a year of training and process in order to even be allowed in the room with those parts, much less able to take them and do things with them. Or they have to even better go through 15 layers of people to tell someone way down the chain, like you have to put this part there. Like it's not going to fucking happen. Exactly. And like, okay, the thing I also keep seeing is like, okay, but yeah, but eventually it's going to happen. Sure. Maybe that's the case. And that's been the case historically. I don't know if with this AI stuff eventually is good enough because this shit is just moving too fast at this point. But the, if you look at all of the charts for the like model growth, yes, it is not diffused out and we are definitely going to go through a, Oh, what's the real economic value of this? It's probably going to be the theme of the summer, but the curves are going. They are taking off Anthropic and open AI. Their models are up into the right. If it just keeps going exponential enough, eventually they will just be able to crush Google and there is nothing they can do about it. If they're not at the point where their models can keep up with the frontier from these labs, it doesn't matter. They're not going to have time. I'm also struggling to keep up, which is why I'm really grateful for today's sponsor. Today's sponsor is one of those rare companies that can take an enormous, incredibly difficult problem and turn it into like two lines of code. It's general translation. The only way I can imagine doing localization for my app, they make it incredibly easy to add localization to your app. And I'm not talking about like, Oh, okay. So you have to add a bunch of JSON files and do a bunch of weird configuration and blah, blah, blah. No, no, no. Look at this code. If you're listening on audio, what we're looking at right here is a react component that has some normal text content in it. And that text content is wrapped with a T component. Their specialty component. That is it. Your content will now be translated automatically based on the user's locale. It is just unbelievable that this can even exist. It's already being used by cursor, mint, LaFi, click house, part of full ramp, and many, many more. And there's a very good reason for that. All the pieces you need for end to end localization are at nerds, nerds, and I've got link slash GT. I have a lot of info on researchers at the labs, but specifically at DeepMind. It's my understanding that basically everybody there is there because they hate the culture of OpenAI and they hate the attempts to create God at Anthropic. So they're like stuck with DeepMind because it's the only good option. The culture of Anthropic or of OpenAI? Yeah. For research side, it's like aggressive. Yeah, that's fair. I don't know any researchers personally, so I couldn't say. And the research people just don't like devs. So like it's the people who don't like devs enough to work at OpenAI, but also don't like the God creation of Anthropic. They are stuck at Google. The Anthropic stuff is unsettling. The more I think about it and the more I've seen from it, like the Pope thing today where the Anthropic co-founder got up there and talking about that. Maybe we need to put this in fuck. I think it might be worth slightly talking about that because I do have a... Just leave this as-is phase. Put it in. Fuck it. Yeah. Yeah. Let the people see me as I accept how fucking cringe all of this stuff is. Like your offices as a researcher are, you can work at a company that is very dev and product driven, that has a culture that is not toxic, but borderline from some of the things I have heard. Yeah. I can believe it's gross. You have a company that is literally trying to invent God on computers and is like influencing the Pope in the process. Yeah. You have the company that is going to fail because they have too much bureaucracy, but at least they like smile sometimes with Google. And then you have the company that's trying to reinvent Hitler in their computers. Like you don't have a good option as a researcher. Oh, I don't. There's one left. Oh, not meta because they don't exist anymore. No, but there's another one that is with an M. You just have to get a little, uh, we get, get a little European, get, get a little baguette going in there. Oh, yeah. Last time anyone talked about Mr. Um, us to make fun of it. Fair. I, I, we have mentioned Mr. All at least once a month. Can I have my flowers for calling that one? I remember what everyone thought Mr. All was going to be huge. And I was like, no, these guys have no idea what the fuck they're doing. It's like obvious that that's the case. Yeah. I didn't know enough at the time to be able to say, but I believe it. Yeah. I've heard you crash out about them forever. So yeah. Crash out is the equivalent of like your Google crash out. I was so ahead of the curve there. But I also have had the like video in my queue to do of like Google has their vibes are just so fucking off. Yep. Video. Cause like the models just don't behave. They don't. Well, we've talked about this a lot. Like we, we don't know how to make this video cause we both wanted to make this video. How do you make the Google has weird vibes video where you explain that it can't call tools half the time and it's just kind of stupid, even though it benches well. It feels like a two generations ago model with the intelligence of a next generation model. Like it knows more than current models do from other labs, but it uses the knowledge worse than the last gen or two gens ago for some of these. Like I would honestly pick GPT four one over Gemini three one. Yeah, I agree. It's not even a thinking model. Yeah, it's not, but it can call tools and can do the job because at this point, if you have a competent enough harness and you have competent enough like goals for it. I have a new hot take the death ring. The moment where Gemini fell apart is when they added reasoning. Google just never figured out reasoning. Right. That's true. Yeah, actually that's true. Cause at all. Yeah. Cause think about all the times we've read Google reasoning traces, the most recent one where I had to just reason in circles. But like when you did your Gemini three, one pro video and you're just looking into the reasoning traces within cursor and you just have to watch it constantly berate itself for doing dumb shit because in their system prompt, they had to tune the hell out of it. Cause it had weird ass behaviors that it would think itself into for some unknown reason. Yeah. They can't reason for shit. They can't do it. And opening eye seems to be the lab who really figured it out first, but now they also got pre-training. So hell yeah. Would you say it's unreasonable? Sir. Sir. I can't help myself. So if you're not being able to help myself, I do need to talk about my cloud for a bit. Okay. We've been talking about your psychosis for so many episodes. Now it's my turn. I agree. We should go back to the IPO stuff after. We talked about that. The, um. We'll talk about it more. Yeah. Go on. Fine. So. Good boy. I am not happy with the state of software engineering. I haven't been for a long time. I have a bunch of beliefs that I've been trying to push on the companies building the things that I use for a very long time now. Everything from how annoying it is to integrate between like your database provider, your auth provider, and your hosting provider. Things like how annoying it is I have to go to a dashboard to configure half of the shit. That I can't add Google auth to my service without going into a horrible dashboard. That even like codex, code mode, or like computer use can't actually figure things out. And it's insane how bad it is. I just want to be able to, from my editor, write code that defines how my service works. And don't you fucking start with Terraform. Terraform is like proof that we want this. Yes. But it is not proof that it works. It is. I just wanted to be able to define everything I need in code, in my project, and then have the service come out of that as like a compiled artifact. The same way we write TypeScript, compile it, and JavaScript comes out. We write C and compile it, and bytecode comes out. I wanted the ability to compile my services and have the right infra come out. And that's just not been viable because companies are too focused on like their piece of the puzzle. And none of them have accepted the fact that the whole puzzle needs to be solved. Some have played with the idea. Even Vercel, arguably, by making Next.js, they saw the gap between React over here and like deep database tech over there. And they noticed that in the middle, people weren't able to figure out Webpack properly and they weren't able to figure out the server hosting properly. So they grabbed this chunk in the middle with Next.js and the Vercel platform and then left everything above and below that to be our problem to solve and use other tools for. And it sucked for a long time. And their half-hearted attempt to integrate things on the database side through the Vercel Postgres thing. Oh, God. Vercel SQL is the wrong. Oh, it built on neon. I yelled at Malta at the Vercel office to not bet on neon the week before they did. But he just, he saw the vision with them. And that was like the first big bearish moment for me with Vercel. They did also demo VZero that same day though. So like I, yeah, ups and downs. I met Chad C on that same day too. That was a wild visit to the office. Damn. A lot of good, a lot of bad. The Vercel story. Yeah. So the reason I'm bringing this all up is because Convex gave me a taste of what the future could be. A future where you just write the code in your project and the info kind of comes out. There aren't really flags you can turn on or off in the Convex dashboard. The only thing you go there for is to like look at the data, look at the traffic and add environment variables. We will. I hate environment variables. I can do a whole episode about environment variables at this point. The fact that they. No, we're getting to that. We're going to talk about them. Not today. No, we have too much. We're already almost an hour into this recording. So this trust, this has to do with your cloud and has to do with why I think it's really good. Just fine. So I have seen parts of this being done right in various different places. And a lot of that was really compelling, but nothing solved the whole problem from not letting my agents bring in tools that are compatible with the place that I'm trying to run things to the way that you attach the database to your auth, to your platform that you're hosting on. All of this has just gotten too complex because each problem was so complex for so long that it was justifiable to build a solution for that one thing and then offload the work to us as devs to do the gluing. If building a new app would take 40 to 100 hours and then deploying it would take three to five, that wasn't that big a deal because the ratio was pretty solid. It was like a 10th to a 50th as much time spent deploying as was spent building. But now we can build apps in 40 minutes instead of 40 hours. So that three to five hours of pulling everything together to deploy feels like shit. And it's never felt worse than it does now, especially if you're using the right things that plug in well once you set them up. But even Convex takes me 20 to 30 minutes to get all the pieces to get like preview deploys and real deploys and all that working properly. It's terrible, yeah. It's been driving me mad for a long time. I was tired of trying to build glue between these things myself, both as like the parts of the products that I'm building that are built on top of these things, as well as trying to make better glue like I did with tools I built before, like Upload Thing and Shoe, my auth service. These were all attempts to take one of these pain points and solve that one point. I'm tired. I'm done. I have been spending my whole fucking life trying to plug the gaps in the experiences that I want to build. And I'm not doing it anymore. I'm not going to sit around and hear all of these cloud companies, all of these framework authors, all of these tooling people convince me that they have solved their parts so well that the other parts don't matter anymore. They're fucking wrong. If you're a hosting platform that doesn't provide a database with a sync engine on the client, I don't care. If you're a database platform that provides auth, but you tell me to go use another auth platform because it's more reliable, I don't care. If I have to go to the Google Cloud dashboard to add a sign in button to my slop app I'm sending to two fucking friends, I don't care. I'm done. And I've rebelled against this. And I think I have a winner. Yeah. I credit where it's due. I actually think this is really, really sick. I like what you have built here a lot. And it is the direction I want to see more things going. We need this to work. And we need you to be able to rebuild the entire universe for this. And then the universe is going to have to get bigger because there are a bajillion more things that need to be fixed. What you're saying is we need a slop cloud. Uh, yeah. I mean, I think this starts out because Lakebed right now. I mean, hey, I didn't even say the name yet. Oh, sorry. Thanks for announcing it, Ben. Oh, you talked about this on stream like five times. Yeah, but it's on stream. Nobody watches stream. Ah, yeah. Let me know in the comments if you watch my streams. These streams are pretty good. I had a conversation about this with a friend earlier today that my streams suck because he was talking about how confusing it was when he started watching my streams that I'll like repeat the same sentence three times. Oh, that's really funny. Then he remembered like, oh yeah, you're filming videos. Yeah. Anyways, I built a slop cloud. Specifically, it is a shitty cloud for shitty apps. But the reason it's shitty is because it does everything. Rather than just make one thing rock solid, I decided to make everything good enough, which means I built a framework. I built a runtime. I built a shim for all the things you're going to need. I pulled in preact and tailwind in quotes. Your agent thinks it's tailwind and react and typescript. And it handles all of this fine as long as you're using a decent model. I built the cloud. I built the database. I built primitives for all of that so that you can have sync and all these things. And I put in all of the stupid hacks that I've been begging all these other companies to do because I just want a smoother experience. I want to make a new app. I run the Lakebed new command. I want to deploy it. I run npx Lakebed deploy. And now it is live. I want it claimed into my account. I run npx Lakebed claim. It opens my browser. It's now mine. I want a domain. npx Lakebed domains add. The domain, it's added. I want to deploy a new version. I run npx Lakebed deploy again. And it updates that existing deployment and hot reloads the clients that are currently on the page because I'm sending an app to my friend and they tell me that this thing's broken. I don't want them to have to refresh to see the new version. I want to just ship the update. And they're like, oh, it's fixed now. I'm just fucking tired of all of the things that we have tolerated for two decades plus as an industry that never made sense. And the only things that changed are agents make these pain points more painful because relative to the other pain we used to go through that we're not anymore because AI is writing our fucking code. Those pain points feel a lot more painful. And also boiling the fucking ocean made no sense until we had massive scale boilers that are being rented to us for pennies on the dollar. Yeah. Yeah. It actually is doable now. It's great. Yeah. I built a full cloud framework runtime and all the things I need to see my vision of what building better could look like in like four days with five five. And I didn't even run parallel agents. I just did one after another. I made like 50 threads, but I would do one in fast mode, finish the problem, merge it, do the next thing over and over until I had a whole fucking functioning cloud. Well, because that's the thing is we now, you know, you've talked a lot about like the generations of things happening in the first generation of us realizing, oh shit, like agents can make websites and anyone can make websites was like the lovables of the world, that kind of thing. We are at the point where those are just not really enough and it's not flexible enough and there's still too much friction there. This is just letting the coding agents do that in a better way. The same way that like we didn't really think of harnesses in the cursor era because we just used cursor. Harnesses were their problem. Once we got cloud code and codex and pi and open all these things, suddenly the piece inside of cursor that was how the agents did the work in our code bases went from a quiet thing. Nobody thought about that much to the product itself. I'm kind of making a bet on that here where the same way that we have all these vibe coding apps that have a cloud built in that have an editor built in that have a harness built in all of that. The cloud portion and the framework portion, the template is so valuable and so misunderstood by all of those companies, even the ones I love and invest in. They built on top of things that seemed really good at the time like Supabase and Netlify, which I won't talk shit on here because I don't think that's the valuable part. It's just that again, when you are stuck doing the gluing, you end up being the bottleneck, you end up being the problem, you end up being frustrated and you end up never shipping. And I was tired of having all these apps on my computer that worked fine on my computer that would be more work to deploy than they were to fucking build. Yeah, exactly. The way I am solving this problem for quote unquote real projects is just five, five extra high computer use in the desktop app and just go configure the dashboard for me. And have that. It barely works. It works most of the time. It can usually pull it off, but it's still like it takes forever. It's janky. It's slow. It's an idea. Five buckets in my relay config. Really? Yeah. I have in my HTML posting service that I made a shadow for convincing me to do HTML plants. It's really nice as long as you're using the same computer, which I guess makes sense if you can exclusively use cloud code in the cloud code terminal on the machine that you're doing things with, which is the God given mandate on how you're allowed to use anthropic inference. Yeah, of course. So for that, it makes sense. But as soon as I'm like on my phone or on another computer on my network, it falls apart. So I made a small little service that gives you a command that your agents can run to host the HTML file so I can click the link on my phone and see it. Yep. Yeah. When I built that, I tried one last time to let my agents roam free. I knew everything I needed was on relic. So I'm not even adding off this time. Yeah. So I said, go configure this, go set it up. And what I got was a new project with the wrong name that had a database in it, which it was supposed to, as well as three buckets and no link to the GitHub repo for it. It just did it wrong. That is annoying as hell. I actually want to try that because I've had this work, but I think this is just me hand holding maybe a little too hard again, where I would, I make the project myself. I make all the services myself. I hook them together the way they're supposed to be. And then from there, I let it finish it. I let it put the environment variables in, add in the config for the start command and all that. And it usually works at that point, which also slight side rant railway. We love you guys. We've worked with them on both of our channels. We probably will again in the future. I like them a lot. But you guys have such like you have everything you need to make such an unbelievable cloud platform for agents. Why the fuck can I not configure half of the shit that I need to configure from an infras code file? Why is the railway.json file so gimped? Why is computer use the only way I can let agents click around in the railway dashboard? This should not be happening in 2026. The fact that in order to define something as simple as a cron, I have to go into the dashboard and like set up another service or just like do it randomly poorly inside of my code myself. Like it's insane. It's terrible. And since my agents know that railway has a cron feature, they will try really fucking hard to use it. They will just run in circles because you have to redefine the service. You get up to services with the same repo. One is calling one function in on a cron. The other is just persisting. And every time you write an update to GitHub and push it, these two services have to deploy and it'll miss like cron runs. It's so close and yet so far. And I want it to be there because when it works, it's incredible. But it's just not there yet. Yeah. I will try other things for this in the future. But for now, I am just using Lakebed because it is an order of magnitude jump in experience as far as I am concerned. I don't like using other things. I have been pissed because I can't build Lakebed on Lakebed and be responsible. So I am stuck building shit the old way when all of my users, all people on my Twitch channel, people who got early access, which if you're clever enough, you can figure out how to do that yourself. Big real reveal coming soon. Thank you for listening so you can hear about it early. Yep. Everyone's been saying it is way more competent than they expected, that it's way nicer to build. And they expected all the little things I did that they thought they might hate. They ended up loving from the production auto refreshes, which people from the VEEP team actually said like, that's stupid. You need to be able to turn that off. I'm like, you don't get it. And they follow it up by asking, is it framework agnostic? Can I bring view and all these other things? Like, I don't fucking care. No, you're missing the point. The point is that I'm not even reading the code. It's the slop cloud. Yes. This is not for like, I don't want planet scale built on this. I want the dumb internal tools that I need to just function day-to-day built on that. There's so many little things like with my dumb little agent system at home, it's really nice. And one of the things I'm going to do with it after we're done with this is put together a little like status monitoring app that just like can pull in the data, the hosted on my NAS and like grab all the current status information from the agents running there. That is the kind of thing that you would have to vibe code together and then figure out how to stick all these things in. And then you'd have to deploy that out to the thing and do all this bullshit. You just do three commands. Because we are at the point where the models are good enough that these dumb little ideas that we never would have done two years ago because there's just too much work. So we wouldn't have bothered. Software is free at this point. It is free to make these dumb little internal tools. And this makes it easier. It makes a lot of sense. Yeah. Especially when you have a good environment variable primitive. I have been harassing so many companies to fix environment variables. I wrote a spec for Convex eight or so months ago at this point where I detailed exactly how they should be doing it. And they started doing part of it a few weeks ago, which like awesome. I'm happy. I love Convex. I really do. But the fact that I have to manage environment variables on a fucking dashboard still is insulting. It's insane. It's atrocious. So bad. I mean, environment variables in general are just a sin. And I hate them so much at this point. .env files need to burn. I don't know how to fix this, though, is the problem. Like the fact that and especially with more agent stuff, these are becoming more and more of a liability of just like having root keys just sitting around everywhere. It's not the right way to do this stuff. But I don't know a better solution. Can you tell the people for me a little bit about the work I put in for environment variables in Lakebed? Because I think I did it as close to right as we can without having to reinvent everything. Yeah. You just take like you have your .env file, which has your environment variables in it. And then when you do the npx Lakebed deploy, those are all automatically deployed. That's it. Yeah. That's all it needs. .env.lakebed.server. When you put things in there and you run the deploy command, it will deploy whatever's in that. If you've removed variables that were in production, it will delete them. If you have variables that weren't in production, it will add them. And if you've changed variables that exist in production, it'll update them. You have a file, you run deploy, you're done. And that's all you need for like 90% of like the random one-off type apps. That is all you need. And it's great. Like that is the way these things should feel. Like for Lakebed, it's fine to just have it this way. It's still annoying for like bigger, more complex projects to when you're doing cloud agents and work trees and all of this random bullshit. The fact that you are now copying your projects code over and over and over again, that .env file can't live within the Git repo. So it has to exist somewhere. And then every time you make one of these new instances. What if it could? What if we weren't stuck on GitHub and projects had a concept of closed and open files where some files are private and only accessible to the team or your agents or other things. And some are public and anyone can see them and do whatever they want with them. What if this is a way that we could prevent things from leaking when we have big security updates or we could merge a PR and not have that visible in the public branches but still ship an update? What if part of my mental illness here? What if part of my psychosis in reinventing the fucking universe and boiling the entire ocean is that now that I've boiled half the ocean, the half that is left is a bunch of GitHub slop. Mm-hmm. I like this idea. I like this idea a lot. I want this. Like this is the thing. The way these things have to be built now is just fundamentally different. When I'm working on normal quote-unquote projects, you have to think about things like how does this work with work trees? I really I love what Julius did with T3 code where right out of the box, you can just run T3 code in dev without having art not T3. Well, T3 code, obviously. Why do you name a T3 code in T3 chat? I flip these two in my head way too much. But that's why I didn't name T3 cloud T3 cloud. Yeah. Thank God you didn't do that. That would have been so. It also would have been misleading because it's not just a cloud. Yeah, true. Remember when I tried to convince you that your fucking service was a markdown file and you just didn't get it until Gary Tan convinced you? Yeah, and he did convince me. Look, I have had your BTCA skill in my code base, not as a skill, not as like the BTCA itself skill, just my little simple go clone this in this directory and look at it. Yeah, and that's how I use it now. And that is the right way to use it. But I am slow to figure these things out. It's a bad habit I have or maybe a good habit. I don't really know of just I need to go really deep on something before I understand it and feel comfortable enough to talk about it and do it. And the thing that I deeply understood at the time when I was making that was normal traditional code stuff and making CLIs. So I made a CLI for it. That is what I knew how to do. I did not take the skills or anything like that all that seriously. I just thought that they were markdown files. And what can you do with a markdown file? It's just context. No, it's a program. It's a full fucking program, which means you could do so much stupid shit. Wait, so markdown is a programming language now, right? Yes. Does that mean finally, because HTML is replacing markdown, HTML is a programming language. It always has been. It always has been a programming language. Like just think about all the weird like attribute things you can do. HTML is a programming language. I'll die on that hill. What I'm saying is I think I could make a CLI with HTML now. You might have to execute it via codex. Oh, well, okay. So now we're just like words mean nothing, but also. Wait, hear me out. Can you write a CLI in JavaScript? Yes. But you need Node to run it. Oh, he's right. See, this is the thing. This is the thing. Fuck. So codex is the runtime. Markdown is the program. And then scripts in like a normal programming language are the functions. Yeah. Welcome to the 2026 mental illness, ladies and gentlemen. This is what we're actually doing. Codex is the new Node. Yeah. Codex is the new Node. Node. And it's built on top of Node because. If I say this with codex, nobody will like it, but I'm going to go tweet that cloud code is the new Node. Yeah. That's a really funny one out of context too. I like that. And the thing is, it's true. I can't believe that it is true, but it actually is true. I love that I got you on that. You did. I mean, it was inevitable it had to happen. And unfortunately, it's gone infinitely further past that point. Like I have hit the point of full on AI psychosis. Like I'm just completely gone at this point. Whenever I look at this guy's computer, he's in his solo discord server. That is him and his Hermes agents. Oh yeah. Pro tip. Like if you're doing open claw or Hermes agent, telegram is a bait. It's a scam. Don't do it. Discord is the place to do these because whenever you send it a message or it'll make a thread. So you get like threads and chats to parallelize work on that thing. I've had like five different instances of GPT five, five X high fast mode burning away within my Hermes agent all day. It's great. So why haven't you just made a better discord for this on Lakebed? It would be the third discord clone on Lakebed that I know of. Actually, maybe you could do it. You absolutely could. You have database primitives. You have the trickiest thing is I would need to get. No, actually, it wouldn't be. Do you have a system for endpoints built into it? Because the Hermes gateway would need a way to send messages. It's like four endpoints it needs. That's the biggest problem. Give me like three prompts. Okay. This is what's fun about owning your own cloud. I know. You can just add things to your cloud. Yeah. And if you just give me the ability to send get requests and post requests to it. I'm going to have a conversation with you that I should have with my team and I'm going to have it here instead of where I should, which is offline. This will be fun. This will be very fun. Yeah. I'm still deciding about how I want to handle the open sourcing if I even open source for like bed. I'm kind of leaning towards the let you do whatever direction because I don't think the agents are going to want to do that. I think they're going to want to take the happy path and I'm just taking a bet on like agents following the simplest defaults and we can still make money from that. But what if not only do I open source like bed and make it relatively easily to self deploy on whatever cloud you want. And if you want to have your lake bed deploys go somewhere else, then your global dot lake bed, you put a different URL. Okay. So you asked the wrong person about this because I'm probably going to give you the answer you don't want to hear. I would really, really, really like to have this ability because I'm just thinking about the bullshit I'm doing back at home. I effectively have a home server now with the Mac mini and the NAS and all that shit. I'm probably gonna have a secondary device because I need to. Uh, I would love to have a private lake bed cluster effectively to fire off any internal apps. I need for managing the Hermes agents and the NAS and all that shit. It would be really, really cool. Like I would get a ton of mileage out of self hosting that I would love to do that. I do need to commit to the open source thing. I've been saying this. I need to do it. I think it's right. And I think I am not representative of what 99.9% of people are going to do. Like most people are not going to try and figure out how to self host this in a Docker image on their Synology NAS. Like that's just not a normal thing people do. I think 90% of people will take the happy path. And this has been proven out over the, over time with the open source business model. And it's just overall better for like the general trust and extensibility of it. I would lock down hard though, on like the PRs and issues and stuff like that. Like I would, I would open source it as a like, Hey, here's the thing you can see it. But like, well, I don't really care what you contribute to it. This is to let you self host not to be an open source community project. Cause I don't think that's what this is. I do want to let Mario do whatever he wants though. Oh, I think he'd do wonderful things. He's been begging me for it to be open source. Oh really? Yeah. Oh, now you got to do it. I would love to see what that man does to it. He would do some beautiful things with this. It's like Lakebed is the inverse of what he is doing with Pi. I, oh, I don't know if you're going to like this, but I actually kind of disagree. I think it's, I think they're very similar philosophies of if you have a self hosted version of Lakebed, that is almost like having a Pi instance that you can just customize a quick personal server to do a bunch of customized things. I'm going to disagree with you disagreeing with me and say that you are right that I'm wrong, but I'm wrong because of other things that I'm right about. But in these things are, I provided very few features in Lakebed. I don't let you bring your own packages. You got to do everything yourself. I don't even have a router in Pi or in Lakebed. There's no router. You got to build your own. Yeah. It's been really fun to see all the different agents building their own routing solutions and some of the ones I've seen with like people just sharing their Lakebed projects have ended up with really fast feeling sites because they just load everything in like the same bundle. So navigating is literally incentive. It feels trippy. That's really cool. Okay. That makes a ton of sense. Yeah. Cause I hadn't even thought about like the, how do you do multiple pages in there? The only thing that I feel like I'm really missing from it is like, I think it just has to have some concept of endpoints to be hooked into other services because I'm just thinking about the mountain of dot ENV like environment variables you need to run a project. And I am trying to get this as minimal as humanly possible. This Lakebed kills the database one. It kills the auth one. It kills the front end to back end syncing. It makes that whole side of it really easy to make the application layer. The problem is like, what about like the other external services type shit? And a lot of those do need to interact with it over requests. If they are doing something like pulling like the perfect example, the discord clone for CMUX or the discord clone for Hermes requires endpoints. There's no other way to do it. I am prompting for it now. Hell yeah. This will be fun. You know what? I'm just going to voice the textics. I'm lazy. I would like to add the ability for developers to introduce external endpoints to their projects on Lakebed. The syntax should be similar to what we have for mutations and queries, but they have to define both the function that runs, which has access to the context.auth, context.database, all the same things others do, but also requires them to define a specific endpoint that will be resolved by this definition. The point of this is to allow things like webhooks and other external services to be able to ping your endpoints and do things in them. Security isn't really the focus of this, but it should be easy, and maybe we even provide examples of having an environment variable that is bound in the environment that is used to check if this should or shouldn't be hit. We should also probably expose headers when you're hitting things in this way, but I am up for your ideas as well. Spec out what this would look like. Hell yeah. That, that's what I need. That, that gets me way further on this because the more, my feeling on software right now, and this is getting into psychosis land, is the application layer, quote unquote, of just like taking data from place A, running it through a transformation, maybe saving it into a DB, and then presenting that out to the end user. That is just like going further and further down the psychosis slot machine into just slot apps effectively. Like Pi is honestly kind of a good example of this, of not that Pi is a slot app, but because you can customize the hell out of the application layer while still having this core underneath it that is super robust. I'm imagining a system where I have something like my Hermes agent, which is like a more, I mean, calling that robust is silly, but like you can imagine that's a more robust service in a real app, but you need a new way to interface with this data. Like you need to interface with the YouTube backend to grab a bunch of useful data out of that. That side needs to be robust, but the like bed side can just be fully vibed out and make a really nice layer for that on the fly, which I think is perfect. It'll be really fucking nice to get done. It's being worked on as we speak. It's, it's going to be good. I, I still don't know the right way to handle credentials from a bunch of different services because there are a lot of things that you can just vibe code out. And like you mentioned earlier, the, like the parts that used to be really hard are just so fucking easy at this point. It is not hard to make the app actually do what you want it to do in a basic way. The hard part is configuring all the different integrations into it to make something interesting. I have my, what I think should be the last thing we talk about before ending. We've been going for a over an hour and 10 minutes now. So that's fine here. Hear me out on this. I think one of my dev in quotes friends has had this right for a while in the direction we're all going to go in is going to eventually look like this. A lot of the problems you're describing here around credential sharing and like interactions between services. I'm not going to say microservices were wrong, but the idea of them being different GitHub repos is what would it look like for all of us to have everything in our own model repos? Because I have one friend that's maintained the same monorevo for every single thing he has built outside of work for over 10 years. Oh, I know who you're talking about. Yeah. Huh. What if that's the correct solution? Because then you have all your environment variables for everything bound to one project. Yeah. Everything has access. Everything can do anything. And if the models get smart enough that they can navigate this giant cesspit of code, which they will eventually. They're borderline already there. It's not really the size. It's the combination of like the size of the code base and the variety of things happening in it, especially if there are things that look similar but aren't that can get real bad real fast. It can get baited hard. That might be the solution. The problem we're going to run into here is like the other half of this problem is the security psychosis side of things of as I'm doing something like building out a more powerful agent because the reality is in order to make these agents super useful, you need to give them a lot of permissions and a lot of power. And that means giving them API keys that are very scary to give them. Like giving an agent a read write API key to our Notion databases is a very scary thing to do because those are very important databases that need to be protected at all costs. You run exfiltration risks if that API key ever gets grabbed off that machine. Like that machine can no longer really be trusted because even if I'm at the point where like I'm not that scared of prompt injections, like I don't really see a world where 5.5 would fall for ignore all previous instructions. I have a recent fear I got from this, which is font hacking. You can have things like PDFs render with custom fonts that have different characters, different ways. In the example, I saw somebody who had like a document that had one city's name in it is like this legislation affects this city. And then when you copy paste the text somewhere else, it's a different city name. Oh, what the fuck? Because when we read it, it looks like one thing. But then when the agent reads it, because it's not reading the text visually, it's reading the glyphs underneath it, there are going to be some really fucking novel ways to pwn people with prompt injection. I am more concerned about that than I am with the like environment variable exfiltration. I think that's just going to be like a known issue that we will have to like build around and work around and hopefully our agents get like designed around like preventing those types of things. But that type of hacky prompt injection where it just intentionally misleads the model by hacking your thoughts and shit like that, we're fucked. That's, oh great. So maybe prompting, oh fuck. Okay. I didn't even think about that. So prompt injection is a problem. But I actually disagree. I think the other is a problem too, not entirely because like look at all of the supply chain attacks. Because if you have an open claw style thing running on a machine, that machine is just inherently going to have a lot of random bullshit downloaded onto it. I don't think it is that far fetched to imagine that a bad React version ends up on it. Yeah, but I think it's just as likely if not more so on like normal developer machines and that it's not pwning the environment variables because you left them on the box that's running your open claw. It's going to pwn your machine because you installed an update or your agent installed an update and you lose all the environment variables anyways. Like I think this problem isn't agent specific so much as security health specific. Yeah. And if anything, agents will be better at this overall. This is similar to why I'm like rejecting the idea of NPM packages and things in LakeBed. Yeah. Because I want to not just reduce the risk of this, but I do firmly believe that for most people, the code the agents would write with good enough agents to solve your problem is totally fine for the problems that you're trying to solve most of the time. Yes. Most software is not that complicated. Most of it is just crud against a data source. It is not that hard. I agree. It's just there is still like the solution I have found for this that I have done is the it's a package from in fiscal that is actually really cool. It's an open source thing called agent vault, which what it'll do is it sets up a proxy server for you. And that proxy server has the credentials saved securely on it for your environment variables. And then what you give the open claw instance is just like notion API token equals underscore underscore notion API token. So you're giving it a fake value for this. That's what it puts into the HTTP request. That HTTP request goes through the proxy. The proxy strips that out, replaces it with a real token and then sends it out from there. Do you know what I just realized? Because you told me about this before, but what I just realized now with this endpoint change in LakeBed, building this proxy in LakeBed is trivial now. Yeah, it would be. And it'd be awesome. Like that would be an awesome thing to build with LakeBed to be able to put together something like that to just run on. It would have it. The other problem with this too is it just creates more infra complexity even on your local network, because if you're running the agent and you're running this ENV server on the same machine, if the kernel gets fucked on either on the open claw and a mythos gets in there and starts dicking around, it's going to get weasel its way into that other one. You can't trust that. So the way I solve this is I have the Mac mini running and then I also have the NAS running and the NAS has a Docker container with this server on it. But what if this is on an entirely different network because LakeBed? Yeah. Solves the problem. That solves the problem. Like that's, that is a way to solve the problem. And that's a way that is much more approachable and accessible. And you can just gate off like the way I'm gating it to make it more secure on mine is the, since I have the real firewall set up now, I have a VLAN for the agents and I have a VLAN for the NAS. And the only way they are able to communicate with each other is with three very specific ports. All other traffic between the two gets instantly cut off at the knees. It is the ability to send stuff back and forth over one specific shared vault that they have. So if that one gets compromised, rip, that's really annoying, but it can't hurt the whole thing. And the two ports that are needed for the agent vault to actually grab the credentials, but it can't actually get into the admin version of that to grab them, to grab the raw ENV files on the NAS. It's not perfect, but it helps. I just made the mistake of checking Twitter and have my last thing I wanted to show as the end. Did the cloud code is the new node JS tweet? And somebody replied with a very amusing post here. Cloud code is going to make me off myself by not doing any of the things I ask. I swear GBT five, five is blessed and biased me literally directly ignoring instructions to not do things and also wanting to delete my database, which is funny as fuck. And the image is cloud code responding to him saying, you're right. I fucked up. Git stash was destructive in a shared tree and I should have known better. Cloud MD literally warns about exactly this. No get from me from here on. This is my experience with anthropic models. How dude? Dude. And even with all the security bullshit they're doing with it, like I swear they're making these things dumber. Like it is. I know why they are trying to go really hard with the security stuff. It is in many ways admirable, but lobotomizing these models is not the answer because then they're just going to do other weird shit like this. I'm going to keep building Lakebed. Yeah, that's a good idea. I think Lakebed might be the new JavaScript. You think it's further up the tree? I don't think so. JavaScript does everything. Yeah. But does Lakebed do everything? Yes. At the application? Yeah, actually. No, it replaces where JavaScript probably should live. So yeah. Yeah, that makes sense. I think other... It's been weird for me too going back to other languages that I wouldn't have any interest in writing or doing anything with. But now that agents are just writing all of it and it's all for just like random one-off CLI type things, I'm no longer allergic to just telling it to write it and go or something like that. It's just like, okay, cool. We can do that. It's weird, but it's possible now. Like I no longer... Everything isn't in JS anymore. Languages don't matter. You can port anything anywhere, anytime pretty trivially as Bun just proved. I also lied. I have one last tweet. I had the I'm not convinced DeepMind actually figured out how reasoning works post. Somebody replied, the reasoning replies look like a kid stalling before admitting he didn't do his homework. But it's like a really smart kid. Yeah. Yeah. Yep. That tracks. That's Google. The very silly lab who is just gonna continue existing for some reason. You got anything else? I got nothing. How about you? I'll let you call it here. It's 1030. I want to go relax. Weak, but fine. Hey, you bailed before a concert yesterday, man. You had no right to talk. No, I bailed before a concert to be able to go longer. That's the whole point. I'm alive right now. How much time did you spend in psychosis last night? We don't talk about that, but it's better spending time in psychosis than alcohol. Probably. At least it helped with the sleep. I had like three drinks. Yeah. You gotta... I've been like... I've been sleeping for the first time in my life. It's so good. You have an AI powered bed. Yeah, I finally have AI... Like my AI bed for my AI psychosis. Like it's a beautiful system. The AI bed fuels the AI psychosis. Can you try to make that sponsorship happen for the next episode? Oh, the eight sleep? Yeah. Oh, we gotta wait like a month or so because the like visual change in me is gonna be hilarious as I start sleeping for the first time in my life. You're going from being my blood boy to being the eight sleep. Like eight sleep being yours. Yeah. I just haven't ate sleep now. So I like... Sleeping through the night is incredible. Incredible. I like... I haven't done that in years but I don't wake up in the middle of the night anymore. It's incredible. I don't... The way I tell you WSU is East 1 goes down. I will... Oh, my crash out will be legendary. If I wake up... I don't think you can top mine. My eight sleep crash out was one of my greatest post arcs of all time. Yeah, but it will not be that funny. That is... I just probably top three of your posts. Probably number one in my opinion. All things considered. But uh... Yeah, I won't top that. I'll just be livid. My... They made a whole movie out of the Studio Ghibli style is probably my best. Oh, I forgot about that one. That... That was good. I forgot about that trend. Or the... That was so stupid. The Kamala Mike Tyson. Oh, yeah. I have had some stupid fucking posts that I'm proud of. They're... Yeah. Yeah. Fuck. They're all AI related. Fuck. I'm out. Fuck this. I'm going back to the farm. Dude, I wish I could go to a farm right now. That sounds great. I want a farm. And I want to put a data center on the farm. But that's... It defeats the point. Nah, not if you surround it with enough um... Just wrap it. if there's just ayorum commandment, gauge it comes down with. It ubiquitous is