← Back to search
RS368: Grok Bot & Hermes
Rogue Startups · 2026-09-22 · 50 min
Show full episode description
AI agents are getting a lot more capable, but are they actually making work easier yet? Craig and Colin compare their latest experiments with GrokBot, Hermes, OpenClaw, Claude Code, Codex, and other AI tools, including what works, what breaks, and where agents are genuinely reducing cognitive load. They dig into open vs. closed systems, shared company agents, local AI, model costs, and the idea of an AI chief of staff that keeps founders focused. Plus, Colin shares how he's using AI to overhaul a 15-year-old WordPress site, and they explore what the next generation of AI-powered websites and marketing tools might look like. Highlights from Craig and Colin’s conversation: Open vs. closed AI agents for running a business Why Craig chose Hermes as Castos' shared company agent The security challenge of giving AI agents real power Why AI should reduce cognitive load, not create more work How Colin uses AI for CRM, task management, and daily planning Claude Code vs. Codex, Cursor, and other AI development tools Why expensive AI models aren't always the best choice The case for running AI locally or privately How cheap AI access encourages founders to experiment Can an AI chief of staff actually keep you focused? Using AI to overhaul a 15-year-old content site How tiny AI-built tools could replace traditional lead magnets Resources and Links from This Episode: Produced by Castos Productions : https://castos.com/productions/ The Podcast Host: https://www.thepodcasthost.com/ Colin on LinkedIn: linkedin.com/in/colinmcgray?originalSubdomain=uk Castos Free Tools : castos.com/tools Email me:
[email protected] Find me on Twitter: https://twitter.com/TheCraigHewitt LinkedIn: https://www.linkedin.com/in/craig-hewitt-78386a66/ Email:
[email protected] Chapters (00:00:00) - Intro and topic setup: bots and agents (00:01:23) - Open vs closed agent platforms (00:03:35) - Why choosing Hermes for a shared team agent (00:07:07) - Permissioning and security challenges with Hermes (00:09:08) - Vision for AI teammates across a company (00:10:44) - Personal use: Grokbot vs Claude code workflows (00:12:53) - CRM, task management, and cognitive load reduction (00:16:09) - Mobile workflows, Notion, and GitHub syncing (00:19:06) - App preferences: Claude vs Codex vs Cursor vs Antigravity (00:22:00) - Right-sizing models and cost discipline (00:22:19) - Golf video workflow and downgrading models to save cost (00:25:33) - Local and private AI compute considerations (00:29:35) - Enterprise API costs, subsidies, and sovereignty concerns (00:32:35) - Value of experimentation under usage limits (00:33:56) - Chief of staff model and multi-agent orchestration (00:37:59) - Chief of staff for managing priorities and stale projects (00:41:25) - Old-school business plans teaser (00:42:01) - Renovating lawns4u.com with AI agents (00:45:00) - WordPress vs modern stacks (Astro, Next.js) discussion (00:47:33) - Building tools and lead magnets beyond WordPress (00:50:15) - Wrap up and future topics
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Choosing between open (
Hermes) and closed (GrokBot) agents for a company-wide AI assistant, and safely permissioning team access.
Benefits
- Own the data, models, and cost with an open agent on your hardware
- Shared team surface in Slack instead of one-person personal bots
- Agent pulls Help Scout, Sentry, GitHub digests instead of standup reports
- Per-channel/user permissioning to limit destructive CLI access
Use cases
- Castos runs a shared Hermes agent on a Mac Mini, accessed by the whole team via Slack
- Agent created its first PR and first marketing drafts this week (both 'pretty average' so far)
- GrokBot checks Help Scout, Sentry, GitHub and inbox a couple times a week as a standup-style digest bot
- Worked on permissioning channels/profiles so devs get CLI access without destructive commands
- Vision: replicate manual Claude Code/Codex workflows inside Hermes with process and guardrails
KPIs / results
- First agent-created PR shipped this week
- GrokBot digests run a couple times a week across Help Scout, Sentry, GitHub
Tools / build
- Hermes agent for Castos (Slack-connected, Mac Mini-hosted)
- GrokBot personal digest bot (Help Scout, Sentry, GitHub, inbox)
- Planned Trello-style control panel for orchestrating the agent
[SPEAKER_01] Claude is my least favorite app for sure. And it's been my least favorite model. It's been our team's least favorite model. Everyone's on Codex now. [SPEAKER_02] Really? You moved over. The superpowers of it is that you have to give it superpowers. To get real value, to make it as useful as it possibly can be, it needs to be able to do anything, which obviously gives it bad consequences potentially. [SPEAKER_01] You and I have learned so much by just fucking around that these companies aren't. Because people are like, well, every time I hit enter, it costs 70 cents. And you multiply that times 5,000 [SPEAKER_02] people every day. That's a lot. You and I, we don't want a boss. We want to be our own boss. Boss is generally quite good for getting you to make progress and do the right things. [SPEAKER_01] Yeah. Hello, welcome back to Rogue Startups. Joined once again by Craig Hewitt. Colin, today we're talking bots, bots and agents. We finally are getting into the good stuff. I feel like we've been dancing around the topic du jour for a couple of weeks, huh? [SPEAKER_02] Yeah, I think there was a few things came up after our chat about it. You were mentioning GrokBot. We've both been talking about what we've been doing with building more automated systems, I think, isn't it? So yeah, it'd be cool to catch up on what you're doing. Get some feedback maybe on the systems I've been building too. And also one little update I think I'd love to share around a renovation, not of property. Well, partly of property, but of a digital property this time [SPEAKER_01] around too. I think the first question when you talk about agents is like open versus closed. Yeah. That's kind of one of the big decisions. I think, you know, GrokBot obviously closed. Mm-hmm. I think at this point you could call Codex an agent. I saw something on Twitter yesterday. They're supposed to be coming out with like, you know, they acquired OpenClaw and Peter Steinberger. They're supposed to be coming out with a Codex agent soon, which will very, very much be like a GrokBot kind of thing. Yeah. [SPEAKER_01] Yeah. Surely Anthropic will do the same. So I think from like an initial strategy perspective, like, hey, do you go on someone's turf and just accept the limitations and the issues there? Or do you go, I know you've played around with OpenClaw quite a bit. We're building a Hermes agent for Castos and chose it very intentionally because it is open. We control everything. It runs on my hardware. You know, I was one of the suckers that bought a Mac Mini like February of this year when the OpenClaw [SPEAKER_01] craze and are finally putting it to good use. But how do you think about, yeah, open or closed? [SPEAKER_02] I mean, I'm usually an open type of guy. Like I would, in the olden days, the argument was always like what platform do you build your audience on, wasn't it? And you always were advising people, like don't build it on Facebook. Don't build it on Twitter. Build it on something you own. Like don't build on rented land and the like. But I don't know. I think the question is, or the answer is a bit different this time around, isn't it? Because it's a bit more about, [SPEAKER_02] it's a bit more about sustainability and about reliability and security as well, I suppose, isn't it? I mean, that was always, that was the big thing with OpenClaw when it first came out. I'd be interested to hear how Hermes is with that as well these days since you're building that. But I think it's different, isn't it? I think this time around, it's more about, I want to be able to do the work and not have to worry about maintenance and it going down all the time. So I'm veering more towards the closed, the prepared systems these days. But yeah, [SPEAKER_02] tell me why, why did you go with Hermes then? Because that's open obviously. So you're, [SPEAKER_01] you're obviously going that route instead. Good question. Very real concern. I think security is the biggest concern because otherwise the benefits are all there. You own the data, you control the models, you control the cost, you can run local models. Spoiler alert, going to be speaking with Heaton Shaw actually this afternoon in real time, but it'll be the next week for this podcast about local models and open source. So like, you know, whatever, subscribe and come back if you haven't subscribed already. To me, [SPEAKER_01] the shareability and configurability of Hermes is what makes it a better candidate for a company agent because it's not just me, right? Our whole team uses it. And so we need a surface where everyone can use it. So it's in Slack right now, but you can imagine we're going to build our own little like Trello kind of system to interface with the agent. Probably quite a bit harder to do that with [SPEAKER_01] GrokBot. When we started doing this, it was before you could share bots. Yeah. But even then, I think everyone's kind of like slightly like splintering the same bot and customizing for their own use. [SPEAKER_02] I've not looked into this yet. So you, so you've been using DropBot more than I have. You can't then install that in a publicly available space like Slack or something similar. Is that not the case [SPEAKER_01] yet? No, that's what I'm saying. You can't. Okay. Yeah. From their perspective, it's the whole appeal of it is like, you just download it, sign up and it's, it's wheels included. It's the model, it's the chat interface, it's the harness, it's all the connectors all already there with Hermes. It's like, and I saw the, you know, these guys talking on Twitter yesterday where like, they're going to even strip down more of it to where like, it'll be less kind of batteries included. And Hermes will in the future where like everything is something you have to intentionally add probably [SPEAKER_01] to the point of like security and performance. But yeah, that was the decision for us is like, it's not a personal assistant. It's a, you know what it, we want it to be more like is a shared instance of cloud code. I think that's the better mental model. It's like a shared instance of cloud code that we orchestrate from some kind of community interface like Slack or a Trello board or something [SPEAKER_02] like that. You know, every have been talking about they build their plus ones, I think they call it. And the whole concept of having other, like building AI teammates and having them in Slack. There's something that probably doesn't work about that when it's closed. Like if you're just talking to somebody else's AI teammates. So you're in engineering, somebody else is in marketing, you try and talk to their created bot, their simulated presence or whatever you want to call it. [SPEAKER_02] And it's just like back and forth with that. I don't know if that necessarily works, but when it's put open, suddenly these things have their own personality, their own presence and the person that actually runs them can see it. They can see the interaction, they can potentially feedback and make it better. Or, you know, say, I don't know if I necessarily agree with that and give actually some of their own feedback. Others can as well. Suddenly it's become so much more powerful at that point, isn't it? [SPEAKER_01] Yeah. I mean, interestingly, I saw they shut down their plus ones program every day. [SPEAKER_02] Did they? I didn't notice that. [SPEAKER_01] Which like, you know, for context, plus one is managed hosting for open claw, basically, from every and every dead shippers, like a publication, like a AI publication. They do a for it. It's one of the few things I pay for in the knowledge world. Yeah, they're very good. That's why we chose Hermes. But a very, like a very specific thing that I was wrestling this week [SPEAKER_01] is you've used open claw a lot. You know that it's, I'll say brittle, right? It can break. You just send it one thing and it updates some kind of harness setting and the whole thing kind of breaks. Yeah. I haven't thought so far. Hermes is like slightly better, but, but like not a lot. So initially it was like, only I can talk to our Hermes agent. That's how it was initially in Slack. And I was like, that's dumb because like our team needs to be able to use it. So the gap between it's only me [SPEAKER_01] and it's literally everyone and everyone can do everything. That's basically the step. It's either you or it's everyone, including settings and updating skills and updating the core and all this kind of stuff. So did a fair bit of work this week on like, hey, how can we permission, essentially channels or profiles or users to be able to, or not to be able to do things? But then it gets tough because like, hey, the dev team needs access to the CLI. [SPEAKER_01] How can we give them access to the CLI to spawn like cursor instances to do dev work, but not do, you know, destructive stuff via the CLI? [SPEAKER_02] Yeah. Did you get very far with it? Was it possible to do that? [SPEAKER_01] Yeah, I did. And actually like chat chatted with the Hermes team on Twitter about it. They gave me some like concrete things and some kind of soft things, but it's just like, this is the nature of like open source and, uh, you know, Hermes is a platform. You know, it's not a product, I would say. Like, it's kind of like, it's just like WordPress, which, which kind of fries my ass that I'm using it. But like, it's like the starting point, right? It's getting you like eight steps down the road and then you have to work on it the rest of [SPEAKER_02] the way. It's also the classic sort of catch 22 in a way of the, of open claw when it first came out, which is the superpowers of it is that you have to give it superpowers. Yeah. Like to, to get real value, to make it as useful as it possibly can be. It needs to be able to do anything, which obviously gives it, you know, bad consequences potentially. So as soon as you open that up to your team, then maybe that's even worse, but I don't know, [SPEAKER_02] that really, that does excite me in a lot of ways to think of like being able to have assistants like that in my old team. Like if we could have had, you know, built up a marketing assistant, new our content, new our campaigns and new our approach. And you could just say to it, like we just published an episode, uh, go and do the standard repurposing plan. And then it could set that up and set up drafts and things for us to then go and review and edit. And the team could do that and see it in progress as well. And then same for engineering and same for [SPEAKER_02] growth and sales and all sorts of areas that I love that idea of being able to work with a team. I think there's so many little things, hard parts that that could really grease and get you moving. [SPEAKER_01] Yeah. I mean, I think the reality so far is like, um, yeah, like our agent, you know, created its first PR this week, it created its first bit, you know, bits of marketing and they were pretty average, you know, both, both of them were right. Because like our manual system has a lot more process and guardrails and context built in. And so we're, we're kind of like, what I tell the team is like, what we basically have to do is take all the stuff we're doing in cloud code and codecs right now and put it into Hermes and have it work like we work manually. [SPEAKER_01] It's just not the same yet, but totally like that. That's the vision is like every, literally everything goes through this, uh, harness. [SPEAKER_02] So that'll be running on your Mac mini in your house. Uh, your team will be accessing it from Slack or potentially you mentioned a kind of Trello like. [SPEAKER_01] Yeah. Slack today, some kind of Trello S kind of control panel in the future. Maybe. [SPEAKER_02] Are you still using Grokbot then for personal stuff? [SPEAKER_01] Mm-hmm. [SPEAKER_02] Yeah. What's it turned out a few weeks later? What's, what's it turned out is the actual [SPEAKER_01] daily weekly drivers? Yeah. I feel like, uh, a caveman when I use Grokbot. I mean, I have it set up to like for Castos at like Chex, Chex helps scout a couple of times a week and lets me know of anything that came in that like I should get involved in. Same with like Sentry, same with like GitHub and like, you know, dev stuff that's moving forward. So it's almost like a standup bot that's pulling information from our systems instead of the team reporting it. It does the same with [SPEAKER_01] my inbox, both for like my personal and for my, my Castos inbox. The challenge is I don't use much of that to do stuff. You know, like I get the ping and it's like, Hey, here's your inbox. Here are the three things that need your attention. And I see that. And then I go to Gmail. So like, I'm just not, I'm just not as effective as I should be. And exactly what I said of like, if everything routes through this harness, then it gets better. If I just, if it's just [SPEAKER_01] a notification system, then it's going to stay dumb. And that's my fault. [SPEAKER_02] Yeah. So do you plan to change that then? Or is actually most of that going to just go through your new Hermes? I mean, if you're investing in this Hermes agent, making it actually super robust, secure, useful, knows all of your context, would you shift it over there instead, even for your personal elements? Or is that part of the segregation? Yeah, both. Right. Like I definitely won't have [SPEAKER_01] the Hermes going through my personal stuff. But yeah, like the help scout and the century and the GitHub. Yeah. You know, stuff sure enough should be going through the Hermes agent. I mean, maybe to this point, it's just so easy to do it in Grokbot because you just literally say, Hey, connect to Century, give me a digest twice a week of what's going on there. And then create issues for the three biggest looking problems that you're seeing errors on. Literally, that's literally it. [SPEAKER_02] How about you? Like how are you using both? Right? Yeah. Yes. In a way I've actually, I've kind of come off most of it. I do have an open clause that still runs. I've basically relegated it to a note taker. It's a librarian. So I quite, I use it as a CRM. Actually, it's one of the most used parts. So I was down in Dundee this week, actually met a few old contacts, great to catch up with them. And as soon as I come out of the meeting, I just talk into Telegram and tell [SPEAKER_02] my open clause to update my CRM and say, this is someone else, a couple of things I learned about this person, what we should follow up with, maybe a next time to follow up that kind of stuff. So the CRM side of things, actually, I've found really useful. I'm not sure why exactly I've found it more useful for that than other things. But that's just something I've stuck with actually. And the other parts, I used to use it quite a lot as a task management. So to help me, [SPEAKER_02] because I'm working on like five or six different projects, I always had trouble like logging tasks and knowing where exactly to put them and logging information. And I would just let it do the filing basically, which is what I mean by the librarian aspect. So instead of having to have seven or eight different repositories, it just puts it in the right place. I've kind of fallen off on that a little bit and just come back to using my to do app. So I have found it less useful. The thing I've ramped up though, is actually just using clause code, honestly, for just about [SPEAKER_02] everything, including updates, because I'm in there all day, every day anyway, with my building. So like, whether it's the knowledge work stuff, like working on marketing workflows, working on new content, or whether it's coding and building apps and things like that, I'm in there anyway. And so I've built these, I've got a daily brief that comes through every day, which looks at my calendar, looks at my inbox, looks at all the other kind of tasks. It helps me triage tasks, figure out what's most important that day, sends me that by email, comes into my [SPEAKER_02] inbox in the morning. And that's kind of taken over a lot of what I used to use the open clause for. So I have a feeling I'll probably even change the CRM over to that in the near future, because [SPEAKER_01] Claude could do that too. Yeah. Something that I do along those lines is I use granola to record a lot of calls and it's automatically pulling into my open clause. Yeah. Every call that it checks a couple of times a day, it pulls all the calls in and it updates its context of initiatives that I have based on what's going on. So like for my coaching, [SPEAKER_01] I record all the calls with everyone's permission and it pulls those in and then gives me a briefing based on, so like an email summary to the clients and then a prep message before my next one. So like, hey, last time you talked to Billy Bob, you talked about this and this and this, make sure you touch on these three things that were on their plate. So yeah, I think that's kind of a good [SPEAKER_01] example of like the thing I tell our team is like the goal is like remove cognitive load. Just like you're saying, like you don't want to have to go into CRM, remember all the shit that just happened in that meeting and blah, blah, blah. You just want to like plop something into Claude code that is just easy so you don't think about it. And then all the downstream stuff like works. I think that's like the mental model for me. Yeah. Yeah. It has to just work. That's the [SPEAKER_02] trouble. It has to not get in the way of the actual work, which is so often what the systems we build [SPEAKER_01] do. Yeah. Yeah. Well, what about like if using Claude code is your driver for that? What about like mobile? Do you have it paired with your, and are you on a desktop because it's always on or like this thing? I, cause I don't sit at this desk maybe half the day. And so I'm on the go all the time and do a lot of work from my phone or different places. So I can't really count on something on my computer to be [SPEAKER_02] the driver. Yeah. So I've done a couple of things with that. One is I actually use Notion as the information logging spot. So I've got Notion linked up with Claude via MCP. So when I'm using Claude code, it creates like a build, it creates a build board for all my projects. So I can always see the state of the projects. I can see the to-dos, the in progress, the complete. So I can always just go into Notion and see that. I know a lot of people do that in just Markdown, for example, [SPEAKER_02] like you've just got Markdown files, logs that way. But I like having it in Notion, partly because it's more readable to me. It's just a bit more visual and I can go in and I can make edits really nicely. But the other part is that that then means that I can just open up a normal Claude chat and that same Claude account is linked up via MCP to the same Notion. So actually I can just, I can just drop in CRM updates straight into a normal Claude chat, like any old Claude chat, [SPEAKER_02] and it can go and check that, it can update it. And so it works really nicely that way. So the kind of knowledge work side of things is mostly taken care of in terms of syncing between normal Claude and Claude code. The other part is the development. And I did a lot of work recently actually to make sure that all of the projects that I care about are fully all GitHub synced. So I've been using the Claude code cloud, I don't know what the actual term for it is. What is it? Do they call it someone particular? [SPEAKER_01] Yeah, I know what you're talking about. [SPEAKER_02] Yeah. Claude code, cloud mode. Claude code, cloud mode. That's a good tongue twister. It's a lot. So I've been using that quite a lot now so that I can actually, as long as I am confident, and I'm usually confident in this, as long as I'm confident that I have committed and pushed everything from the latest sessions on my desktop, I can just open up a cloud session on my phone and actually just do some updates that way. And it syncs with the same GitHub. And then next time I'm on desktop, I just synchronize it again. [SPEAKER_02] So I found that actually really useful. So both for app-based sessions, but also knowledge work stuff. [SPEAKER_01] And so do you have like a routine or a script or something to like, when you fire up Claude, it automatically pushes and pulls and syncs and all that with GitHub? Or do you have to remember to like, when you start a session, oh, pull this down, merge, all that kind of stuff? [SPEAKER_02] Yeah, I have not yet, actually. So yeah, that's something I have on the list is to make sure that the sessions, yeah, when I open up the desktop version, it does go and pull it. Yeah, because I haven't done that yet. [SPEAKER_01] Yeah. You know, on this topic, like I just opened the Claude desktop app. And I did a video yesterday about GrokBot and Cursor projects, which we can get into Cursor projects. But I haven't used Claude code in like three weeks. Because for me, just the app of Codex and Cursor are just so much better of a place to work. Like I don't like working in the terminal. [SPEAKER_01] I like, I mean, it ends up being the same thing. Like Codex or Cursor is just a chat interface. But you're able to like view and preview and edit docs right in the app. I don't have to have like VS Code opened up where it's like a little bit in the bottom, then the images are like, it's on the top. I think just in terms of like the way you interact with, I'll call them like agents that you need to be at the computer for. Right. Like Claude code, Codex, Cursor. [SPEAKER_01] Claude is my least favorite app for sure. And like it's been my least favorite model. It's been our team's least favorite model. Everyone's on Codex now. [SPEAKER_00] Really? [SPEAKER_01] On our team. And they're using it at least half the time, depending on the person and like what they're using. But yeah, super competitive space these days. [SPEAKER_02] I keep hearing this and I keep thinking I need to actually switch over some stuff to Codex or anywhere else, to be honest. But I don't know, it's not been letting me down. I'm on the kind of middle max plan on Claude. And I'm using Fable for planning. I'm using that to create like, you know, documents, like project specs and things like that. And then I'm using Opus 5 to actually build the things. And I'm finding that's a really good balance to actually, you know, create something without using up too many tokens and build it all. But I don't know. [SPEAKER_02] I hadn't even come across. Have you come across or tried Google's one yet, Antigravity? [SPEAKER_01] Not since it first came out. Yeah. [SPEAKER_02] No. It was funny because I was talking to somebody else earlier on that's doing a lot of building. They run an agency actually, and they're converting their agency over to AI development as opposed to the old school version. And he's come, they've come to Antigravity as their main method for this. So I didn't know anything. When he first said the name, I didn't even know what it was. So I looked it up and realized that it's Google's kind of agent model, I suppose, if you want to call it that. [SPEAKER_02] But you can even use all the other models as well. Like they've got Opus in there and they've got a version of Codex in there as well, I believe. So it's a funny model, but he swears by that too. I think the main thing with that is it's actually super fast. Like the main thing is it's just speed because it works with the latest Gemini model, I believe it is. And it's just like faster than all the rest. It's not too far behind on thought and reasoning. And it's pretty close on coding, but it's just like 10 times as fast. [SPEAKER_02] So I think that's a big advantage. [SPEAKER_01] Yeah. I mean, I mentioned like talking to Heaton Shaw this afternoon. I think this is one of the bits of discipline that we all are going to need to get into soon is like probably don't need to run Fable or Astra or even Opus for a lot of things we do. If we're not, and no, I don't mean this offensive to you, if we're not lazy, you know, if we're not trying to say like, hey, here's a bunch of stuff, like go figure it out. But like, [SPEAKER_01] if we have skills and evals and systems and a real, you know, deterministic, like defined process for stuff, I don't think you should need the most advanced model. [SPEAKER_02] Yeah, for sure. I went through a kind of a painful process of this recently, actually, because I've got my, I've got my video workflow that I'm using for our golf channel just now. So any golfers out there, go and check out both sides of power in the podcast. But I've been building a way to use all, I grabbed too much video on the golf course, Craig. I know I think I must annoy my partner sometimes, like I'm just trying to take videos of just about every shot. I don't think I slow it down too much, but I've got all this video by the end of a round. [SPEAKER_02] And I've been building loads of workflows to be able to turn that into decent stuff, whether it's, you know, taking shorts and adding typography or putting a shot tracer on a golf, shot or putting together two or three shots into one log of an entire hole and putting a scoreboard up at the top. So I've built little workflows to do all of those and long form as well, and creating thumbnails for them too, for these videos, for, for YouTube. But all of these, I've had the luxury over the last six months of having, I think I got, [SPEAKER_02] so through a startup program I was part of when I was running Alitu, I got about 2000 pounds worth of Anthropic credits, API credits. Wow. So that was sitting in my Anthropic account. So I just went absolutely wild on this. Like I was just like building all these processes that used everything and anything. And I didn't care about the models because I was never running out of tokens. I was, it was somebody else's money. So I didn't know how to care. But recently I just discovered last week, this whole workflow broke. I was like, what's going on? Something's gone wrong. [SPEAKER_02] Looking at the errors and suddenly realized that my API account was empty. So it turned out this money had been given, had an expiry date on it, which I suppose makes sense. Can't keep that forever. But it meant that suddenly I had to put some money in the account and then continue to use this workflow that I had built. And so I ended up spending a day or two just reviewing how much every one of the little processes I had built cost. [SPEAKER_02] And therefore realizing that this was unsustainable for the most part and downgrading them all. But that was exactly where I came to, Craig. It was like, I realized that nearly everything that I was doing could be downgraded at least one or two levels and still run perfectly fine. So there was even the, I thought the image generation was still going to cost me quite a lot, but I even downgraded that at a level, I think, in terms of the latest model. And particularly the analysis, I was doing analysis. [SPEAKER_02] So doing shorts takes a script of 60 seconds and analyzes it for impact and does this for 20, 30 shorts all across a video to try and determine which five, let's say are worth putting time into. I had that running on Opus, I believe. And I got no less good results by putting that right down to, I think I tried it on Sonnet, still great. I think I even tried it on one of the low, like Flash or something. I can't remember the name of it now, but it was a really low level model, cheapest chips, and it was still pretty good. [SPEAKER_02] So yeah, you're spot on. So many things that we do that you don't need that kind of all purpose model for. [SPEAKER_01] So I have a question. I've been wanting to do this for a long time. So I, the Mac, I got the cheapest Mac mini you can get. It's 16 gigs of RAM, which you can't run anything, any, any good model on that. I've had in the cart for a couple of weeks, Mac studio, which is the Mac daddy. It's 128 gigs of RAM. [SPEAKER_02] How much are they these days? Is it like five grand or something like that? More than that? [SPEAKER_01] Yeah, it depends. The, the M4 version is, which is like a one generation old Mac studio is about 6,000 for 128 gigs of RAM and like a two terabyte hard drive. Okay. Interestingly, as we're doing this, chat GBT just took over my browser and is analyzing my YouTube channel. [SPEAKER_02] So pretty, [SPEAKER_01] pretty crazy. It's, it's like on a, it's on a schedule. I kind of want to do it because I think it would be pretty cool. Like a lot of what we want to do that cast us could be run on a good local model. So you need about a hundred gigs of RAM to run Quinn 3.8, which is like the good local model these days. It's on par with like, uh, somewhere between Haiku and Sonnet probably. But if you think about like, what do we use this for? It's like, [SPEAKER_01] I want to write content. I want to do stuff with support, analyze like one of the big things to do right now. We're just in like evaluation mode on support is like twice a week. They evaluate all the tickets that we've had come in that we've replied to and evaluate those versus our, our knowledge base and see if there are areas in our knowledge base. We can improve based on questions that customers are asking. Not super complicated. You know, if you, if you try to like one shot that more complicated, if you break that down into three or four steps, [SPEAKER_01] a simple model, you know, could take care of it. So anyways, like I, I really just want to buy this because, because I just do. The challenge I have is we're running our Hermes agent through my codex plan and have not run out of credits yet. And that includes all of my own codex use. So it's like, gosh, for a hundred bucks a month, the company basically runs its AI and I run mine. Like that's 50, 60 months of at current, [SPEAKER_02] you know, capacity. That's so much more generous than the anthropic plans. I think, I mean, I run out of my max plan all the time. [SPEAKER_01] Yeah. Oh yeah. Unless, [SPEAKER_02] unless I am just completely wrong modeling it, like we just mentioned. [SPEAKER_01] No, I mean, I think, yeah, I tell, I tell like my, my consulting clients like Claude, you know, call it two, three, five times more expensive per amount of work done than, than codex and like 10 times more than cursor and grok. You could never run out of your cursor plan. Basically, I think. [SPEAKER_02] Yeah. That's a good advert. [SPEAKER_01] Yeah. Yeah. It's pretty, it's pretty awesome. [SPEAKER_02] What, so yeah, you had a question about that, but let me ask you something quickly first. You talking about a 6,000 pounds or dollar max studio. There's quite a lot of AI specialized sort of mini computers coming out now though, isn't there for two, 3,000, maybe even less one or 2,000. You're not tempted by that? Or is it really just the big shiny max studio that you want? [SPEAKER_01] I'm not aware of this. Like what's an, example of a specialized computer like that? [SPEAKER_02] They're PC builds, like often in little cases. So they're, they're kind of akin to a Mac mini in many ways, but they're specially designed to run local models. So they'll just be chock full of RAM and a GPU or two, essentially. Yeah. But you're seeing like quite supposedly quite powerful systems for, yeah, a fair bit less than a Mac studio and a lot smaller as well. Like something you could carry around. This is one of the things I thought was quite interesting, like not a laptop, [SPEAKER_02] but a Mac mini style of thing that you could chuck in your bag to take with you to a co-working space or something similar. So yeah, there's a few of them around. It's an interesting sort of, uh, come of the, uh, the local model or the AI development type approach. Yeah. [SPEAKER_01] I mean, I'm in the very early stages of, uh, working on a business idea around this right now, because I think that, uh, whether it's local or just private AI, um, and I'll, I'll differentiate it to have like local, it's like literally sitting here on my desk. Private is like, it's a local model running in a co-located data center or something, but it's, it's, you know, running an open weight model that's only accessible to, you know, whoever pays for it. Yeah. I think there's a massive market for this. [SPEAKER_01] And there will be more and more, I think one of the, one of the risks, and I think there's like a hundred percent chance of it. And we saw, we've seen it already with Anthropic is like, you know, on Anthropic, this is so fucking crazy. You and I are on personal accounts, right? And so we get like a seat and a bundle of credits. If you're an enterprise customer, you just get a seat and everything is API credits. Every time you chat, every time you run cloud code, it's API credits. [SPEAKER_02] Which are a lot more expensive, aren't they? I believe. [SPEAKER_01] 20, a hundred times more expensive. Yeah. [SPEAKER_02] Is it that much? Yeah. Cause you know, often here we are being heavily subsidized just now as our max plans. Yeah. [SPEAKER_01] Yeah. Yep. It has to be that either, and this is me kind of like convincing myself, this is a good idea, but like fear around these platforms, you know, the frontier model companies being compromised, what's really going on with your data? And are they actually training on it when they say they're not, or just the cost, the cost and the sovereignty of your AI compute over time? There has to be a part of the market that will go this way. [SPEAKER_02] Yeah. Yeah. [SPEAKER_02] I think you're right. I looked at it and I wondered about it, but it's just so cheap and easy just now for the amount of, yeah, it kind of, it was in my head, I was thinking, so if we're being so heavily subsidized just now, whereby we're paying a hundred pounds a month for a thousand pounds worth of credits, let's say, that's like, if I work on that for a year, then I'm getting, you know, it's 10K of funding for me, for my business, from Anthropic. I might as well use some of that, [SPEAKER_02] take something back from all the money they are making out of everyone. But yeah. Would you pay a thousand dollars a month for Cloud Code? No, not right now. No, not for what I'm using it for. [SPEAKER_01] Yeah. Oh, if you were running Alitu still? [SPEAKER_02] Possibly. Possibly. But I mean, at that point, you really have to make the, do the calculations, don't you? You have to work out, like, is it actually paying back at this point? Like, is it cheaper than an extra developer? Or is it more effective than an extra developer? Compare it to employees, compare it to, you know, the outcomes, the return on investment, which then becomes a whole ball ache in itself, which you don't really want to have to deal with. So then, there's not only friction around, is it paying off? If there's friction around, [SPEAKER_02] I know it's so expensive that I would need to treat it very seriously. And I just don't have the headspace to do that. So, I don't know. I think it would put a lot of people off. [SPEAKER_01] Yeah. It's interesting, like, you know, I think you and I are just like, oh, I'm going to do this thing. I'm going to create this website. I'm going to build this app, which I, like, I know, like, I have so many. If I look at GitHub, I have so many repos that have literally never seen the light of day. Yes. And that's fine. If you're a company and you're spending API credits for everything, that doesn't happen. And I think one thing that's interesting about that, because they're like, you know, cost conscious, even if they have generous, like, allowances and initiatives to do more with AI. One thing that I think is interesting about that is, [SPEAKER_01] you and I have learned so much by just fucking around. [SPEAKER_02] Yeah. [SPEAKER_01] That these companies aren't because people are like, well, every time I hit enter, it costs, you know, 70 cents. You multiply that times, you know, 5,000 people every day. Like that's. A lot. [SPEAKER_02] There's almost, yeah, there's almost the other effect as well, which is that I feel, and I'm sure you're the same. I've, I've learned more in the last three months because I've been on this limited plan. What I mean by that is I'll find myself checking where I am, how much usage I have burned so far of my five hour window and my one week usage. And if I haven't got close to it yet, I'll deliberately just fire off a few weird and wonderful experiments just to use it up. [SPEAKER_02] Cause I feel like if I'm not using my limit, if I'm not hitting my limit, I'm not getting my value for money. So, so there's all this, there's incentive to experiment, to do funny things that might not work, to just try stuff because it's there as well, which I don't know whether that's a great thing or a bad thing, but I've certainly tried things that I definitely would not have under a usage plan. [SPEAKER_01] Yeah. [SPEAKER_02] One question that I had coming out of the bot stuff that I had been looking at over the last few weeks is the whole model, especially with Grokbot, where people are moving towards a kind of chief of staff type plan, which then coordinates for you. It's something that quite attracts me because I'm getting a little bit, I'm building one app in particular at the moment, mainly for fun, but I've got a kind of eye on whether I could turn it into something in future. It's a kind of fantasy league type app, but for not sport, for other things. [SPEAKER_02] I've got some really interesting ideas around where it could be used. Mainly it's for use with me and my friends, but in building this, it's been great fun building out a whole bunch of stuff. It's one of the more complicated things that I've built with AI so far. And I've often had two, three, four threads open at once, and then kind of wanted to break them into a few separate tasks as well. And therefore I end up, I'm not a developer. I've watched developers work. I know what a work tree is versus a branch, all this kind of stuff, but still I get completely lost and I get a little bit burnt out coordinating it [SPEAKER_02] all as well and seeing where it all is in parallel. But the whole chief of staff model seems to be quite good for managing that type of thing. And then you bring in the fact that it helps you manage like priorities across tasks as well, potentially some schedule tasks in there too. Is that something you're using just now? [SPEAKER_01] That's the idea with GrokBot is everything runs through the chief of staff. I have two, one for Castos and one for personal stuff, just because they each have separate permissions, which I think is healthy. [SPEAKER_02] Yeah. [SPEAKER_01] I think in concept, what you're saying is true. The fear that I've seen is if you pile, and this is probably not unique to GrokBot, but to Hermes or OpenClaw or whatever, if you're firing seven things at an agent at once, that things can get like dropped. You know, if you're like, hey, go build this thing and check HelpScout and write marketing content and do all this from what I've seen, that's when it can just drop something. Okay. Yeah. [SPEAKER_01] So, I don't know what the answer is there. Again, maybe this is where like, is this like open or closed? Like they'll probably just figure it out. Whereas if it's Hermes, you're like, oh gosh, like we have to figure out how to get like a queuing system or something like that. Yeah. Yeah. I'll say two things on the like actual development side, like this video I did yesterday on like cursor projects. And I don't mean to like be repping cursor, [SPEAKER_01] but it is more like that where it is a multi-agent orchestration interface. It's a single chat that then all by itself spawns sub agents to do things. And so you can imagine for your app, you would have a place to go to do everything. So it's like, hey, go do this. And it's going to go fire off an agent to do this. Hey, go build this thing. Hey, go do QA on this. Hey, go do marketing on this. Hey, build me a, you know, [SPEAKER_01] outreach list and interface with instantly or whatever, like whatever, like it's closed. But like, I think the, the difference there is like the, what I came to in my, in my kind of review of this is like crockbot is a personal assistant. It can be for team. It can be for work, but it's like to help you do more work. Cursor projects is to like help you build stuff. Yeah. Okay. [SPEAKER_01] That makes sense. The one, and then this is for kind of everybody. The one thing that I haven't looked at is AMP. I hear our mutual friend, Brian Castle raving about it. And so like I trust him and I would check out AMP. It seems very similar in that it's cloud-based, it's multi-agent and it's model agnostic. So you can use any model. [SPEAKER_02] Yeah. I've heard a few people raving about it too. I've, I've got it on my list to have a look at. I think the combination between that and codex apparently is pretty powerful. [SPEAKER_01] Oh, interesting. Okay. Yeah. So codex can spawn AMP. [SPEAKER_02] I believe so. Yeah. I've not looked into it, so don't hold me to that. But yeah, I've heard talking about that for sure. One of the things when you're talking about chief of staff, one of the things that I think of, and this was something I hired a real chief of staff for years ago in an actual business, was stopping those projects that you mentioned a little while ago, never reaching light of day or stopping those projects, even progressing. [SPEAKER_02] And also bringing back other things that I've never quite finished. Like there's something around supervisor, a manager, bot or agent of some sort that just constantly monitors the things you're working on day to day. And three days later, when you've got this thread that's still open in, you know, cursor, and you're working away in cloud codes or whatever, it says, by the way, just before you start in this new thing that you just asked me to, do you remember that thing you started like three days ago? That's kind of gone a bit stale by this point. [SPEAKER_02] You sure you don't want to just do that instead, rather than start this whole new thread? Yeah. And it potentially could be a little bit annoying, of course, because it's going to try and stop you doing what you want to do. But equally, that's exactly what you want a chief of staff to do in real life. Like if you actually run a company, if you want to achieve your goals, your aims, that's what it would be for. That's a big thing that I've been trying to build in the last little while is a, a bot that helps me actually manage my week to week priorities, my day to day priorities, tell me what to work on next. [SPEAKER_02] Cause it's kind of as much as all of us, like you and I, we don't want a boss. We want to be our own boss. [SPEAKER_00] Yeah. [SPEAKER_02] Boss is generally quite good for getting you to make progress and do the right things. [SPEAKER_01] Yeah. As you're saying that, like the, the two things that like instantly come to mind is like talking about permissioning, like it has to have visibility to literally everything, personal life, work life across all of your kind of work surfaces, which they could be quite hard, especially if you're talking about like cloud code, local development, you know, stuff. And then you're talking to a cloud agent, like the, not, not that what you're saying is wrong, but like that, that's the, I think limitation to it being successful. [SPEAKER_00] Yeah. [SPEAKER_01] And then the other one, and this is where I see the latest models from Anthropic and open AI being good in this respect for the first time is like, I think we've talked about before, like it gets the big picture, it can go down a rabbit hole, and then it can come back up and see the big picture again. I think both Fable and Astra can do this really well. You would have to use to the point of like models, you would have to use a model like that to be like, Oh, [SPEAKER_01] let's get all of the context from all the stuff Colin's doing on a daily basis or whatever. Consider it in such a way that it can keep track of it over time. Cause some of the fuck, some of these things last forever, right? Like, and then have the smarts to surface like the thing that needs to be surfaced. I think that's a quite complex, like reasoning thing. I think you're right. It absolutely is. But it's just once a day, right? It's going to go, you're going to log all the stuff you're doing across everything. And then once a day, it's going to be like, cool, what happened? What do I need to let Colin know about? [SPEAKER_02] Yeah, totally. And I think in some ways doing the plan ahead of time, like the old board model that we used to use, like you've got your board that runs once a quarter, once every two months, once every cycle you plan out. Here's the plan for the, for the two months. Here's our goals and break it down into tasks. Then as long as that's the context, then actually some of the reasoning is not that tricky because does it fit into this? Does it move us towards this goal? No. Well, [SPEAKER_01] yeah, Chuck it out. Yeah, totally. [SPEAKER_02] There's a few things I've been working on this past week, which are completely talking of boards. It just popped into my head, like proper traditional business. You were, you know, you had a corporate business issue just before our call as well, which goes back to old school, how big old, big old corporate businesses run. But I had a couple of potential business plans around completely old school, traditional type business approaches that we could come back to. Like, I know we've, you have to talk about AI a fair bit and running business right now, [SPEAKER_02] but sometimes you're like, I don't know, can we just talk about like a building, you know, property or something like that? You know, it's some, we could come back to some updates around that in the next couple of weeks. Cause I've got, there's a few interesting things. [SPEAKER_01] Yeah, for sure. For sure. You teased like a renovating a property. It sounds, sounds like a digital property. You got to tell us about it. [SPEAKER_02] Yeah. So actually just before I started the pod, well, as I started the podcast host, I had like four or five different things I was playing around with at the time. I was still working in a day job three times a week. So I had a couple of days a week to mess around with some digital businesses. And one of them I did with my dad, actually. So my dad is a ex-green keeper. He used to look after golf courses. He does a lot of lawn care. So he looks after people's gardens, but particularly grass. He knows grass, [SPEAKER_02] like inside out. He's like world-class grass guy. So we started a website called Lawns4U. So lawns4u.com. And I helped him build that first version of it. We put, it was a WordPress site. We put a WooCommerce store on it. And that was like 15 years ago. And he then took it on and just started putting content on it. So he's like written articles on it, kept it updated over the years. But it has barely been updated technically in 15 years. [SPEAKER_02] There's a few times I've gone on and like try to remember what the hell was going on with this site to fix a little bug that came up with a plugin that updated badly and took the whole site down as WooCommerce's want to do. But that is about it. But in the last few weeks, I decided to see how an agent called Codex would be at helping me to overhaul this site and bring it back up to date. And it's been amazing. It's been absolutely incredible what it can do [SPEAKER_02] when you just give it root access to the server, when you give it like a WordPress login, and then access to like Google Cloud Console, Google Analytics, basically all of the tools I used to use to run and grow a content site. So everything from like just literally getting rid of the bugs and making the thing look good, right up to a design side. So like redesigning the site, which hasn't been done yet. So if you go and look at that URL, don't think that's the new version. It's still in its terrible old state, [SPEAKER_02] but I'll be doing that soon. I'll be doing that soon. To like huge SEO audits on this site, like looking at 200 plus articles and saying, well, here's like 12 different orphaned articles. Here's 14 that have broken links that used to link to this old product that's gone out of date. Here's 12 where the affiliate link here is out of date. Here's 19 that I think if we build this little extra article, we could link them all to and build some real authority here because there's a huge gap that you've missed and all these internal links. And you know all this stuff, Craig, [SPEAKER_02] like the SEO side of things. I mean, it's been absolutely crazy what it's been able to do like with barely any input. Not barely any input, sorry. I've been really directing it from a strategic point of view, but it just goes and does all of the actual legwork. So that's been really, really interesting. And you're keeping it in WordPress? Debatable. For now, yes. I thought that was a bigger change. And that's definitely something I'd love to come back to actually. What would you build a content site? We can do that now if you want, [SPEAKER_02] but that's a big question, isn't it? What you'd build a content site and just now? [SPEAKER_01] Yeah, I totally know where you're coming from. We did the, I'll say, research strategy and implementation part of that. Gosh, at the beginning of the year, probably, for Castus, we built a cloud code project that would pull Google Analytics search console and data for SEO, which is like an Ahrefs kind of service. Help us make decisions, create briefs, look at the SERPs, all this kind of stuff, and then write quite good articles. But then it was like, okay, gosh, this is in Markdown, [SPEAKER_01] we have to copy and paste it into WordPress. You could probably MCP it in or something like that. I think WordPress is still a good tool for a lot of people. You need some kind of CMS. I think for most people, our, like the Castus site is just Markdown. It's just all Markdown. It's Astro. [SPEAKER_02] Yeah. [SPEAKER_01] And it's awesome. But that was a month of pretty intense time for me. And we have a reasonably complex site. There's a lot going on, a lot of history, and a lot of stuff to maintain. So yeah, I think keeping it in WordPress for now is reasonable. [SPEAKER_02] Yeah, I've heard a lot of good things about Astro. So my son runs an Astro blog. So what was the pain that actually made you convert it then? What made it worth that month? [SPEAKER_01] Exactly. Kind of what we're talking about, the ability to go from like monitoring, research, prep, writing to publish all from Cloud Code or Codex or whatever. [SPEAKER_02] That's what I've been doing. So I've got Cloud Code actually with a login and it can go, it's updating articles with me. So I'm just doing it all straight in Cloud Code. I'm saying, right, what should we work on next? This Bowling Green maintenance article hasn't been updated in 19 years. So let's update this one. And it pulls in the article just into literally the Cloud Code interface. We both, like Nick says, here's some suggestions. I skimmed through, I think, here's a few suggestions from me too. [SPEAKER_02] We work together a little bit on the edits and then it publishes the results. [SPEAKER_01] Yeah. [SPEAKER_02] Directly. So did you find you couldn't do that easily? [SPEAKER_01] Yeah, I guess the reality is I kind of just don't want to deal with WordPress anymore. [SPEAKER_02] Well, that's fair. [SPEAKER_01] It's more of like a religious discussion. Yes. Yes. [SPEAKER_02] No, that's fair. That's absolutely fair. And I could see me moving off WordPress for sure. I think the design part of it will be more of the issue for me in the end because I'm creating some really nice new designs for it right now, how I want it to look, how I want it to work. And I think a lot of this might not fit into WordPress. [SPEAKER_01] Yeah, like for ours, we have like, you know, imagine, you know, the equivalent of like page templates and elements and stuff like that. So we have like a lot of tools or building a lot of tools. So we go in and just say like, hey, we want to build a tool and maybe this is what would bring you over the edge is like, like we have a tool that strips the audio out of a video file. It loads FFmpeg in the browser and does all of this in the browser and literally was one prompt in Codex. Hey, we want to do this. It needs to do this. Blah, blah, blah, blah, blah. [SPEAKER_01] Use the page template we have. Link it from other places. It thought for about 10 minutes, gave me a preview URL on Cloudflare and I pushed publish. [SPEAKER_02] Yeah, that is cool. But it'll start building tools for that audience. Yeah, absolutely. Yeah, [SPEAKER_01] and like literally running like, especially if you're on Cloudflare, like it's edge workers and stuff. Like you can do some pretty cool stuff without a database or a backend. Like there's literally no backend at all. [SPEAKER_02] I'll put this on our docket for next time around, but I've been playing with a lot of different ways to the modern version of Lead Magnets. We did a bit of this with the podcast host actually was, you know, I mean, it's just a kind of next stage of product-based marketing, I suppose, but how easy it is to do now and the types, different types that you can do. Getting away from the old white paper or the PDF guide or whatever. But yeah, that's absolutely like the start of it, isn't it? Like how you engage an audience by building tiny little tools [SPEAKER_02] but super useful, super targeted for them. And I've been doing one of them for our golf audience actually, which has gone down really well so far. So we should come back to that. [SPEAKER_01] Yeah, but I think that's probably the answer is like when, you know, because like man, I remember when I first started all this, I was like, how can you build a website without WordPress? Like I just had no, I just had no concept. Yeah, yeah. And I think when you go like, oh, I can do more than I can do in WordPress. That's probably when the decision is like, okay, WordPress is not the tool for me anymore because I have things I want to do that it would just be way harder. [SPEAKER_02] Yeah, it just holds it back. Yeah, definitely. [SPEAKER_01] Yeah, I think that's like a good because it is a pain like that it is. There is real pain to not using a CMS, but you could like one of the architecture decisions we made was like, hey, do we want a really lightweight CMS? Like for my personal site, I use, oh man, I can't remember. My personal site is a Next.js app with a very lightweight CMS built into it just for the blog. Cool. So the pages are pages, the blog is a lightweight CMS. And so you could do that. [SPEAKER_02] Yeah, definitely. There's a few nice lightweight ones around I think, isn't there? I saw a couple of little blog systems, but yeah, that's cool. Yeah, let's come back to that. I'd like to know your thoughts of Framer as well, actually. I think there's a lot of talk goes around about it just now too. [SPEAKER_01] Thanks everyone for listening and if you have recommendations for Colin or I for future episodes, give us a shout and we'll see you next time.