← Back to search

Claude Opus 5 Is Here And I Hate Working With It

Authority Hacker Podcast – AI & Automation for Small biz & Marketers · 2026-07-29 · 47 min
relevance 89 10666 words Episode page ↗ Audio ↗
Show full episode description
Send us Fan Mail Anthropic just shipped Opus 5, and on paper it beats their best models on nearly every benchmark. Then one of the first big reviews opened with "Opus 5 is here, and I hate working with it." Both things are true. We spent the week testing it inside our business to figure out which one matters for yours. In this episode we break down: → Where Opus 5 genuinely beats Fable (and where it makes more mistakes) → The one real unlock for marketers, interactive pages that replace blog posts → A full 3D game one-shotted into a single HTML file → Cheap specialist models on OpenRouter for volume work → The new ChatGPT voice mode that runs your Mac from your phone Use it where it's strong. Skip it where it's insufferable. 🔗 AI Accelerator: https://www.authorityhacker.com/ai-accelerator 💻 Learn Claude Code: https://www.authorityhacker.com/resources/learn-claude-code/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Helps small business owners decide when Claude Opus 5 beats benchmarks in practice — and when to stick with Fable, Codex, or open-source Kimi K3.
Benefits
  • Know which benchmarks matter: knowledge work, computer use, novel problem solving
  • Delegate boring setup tasks to computer-use agents while you work
  • One-shot interactive HTML pages and 3D games for marketing
  • Use OpenRouter to test any model per-token without new subscriptions
  • Avoid overpaying: Opus 5 costs 4x less on plan but judges worse than Fable
Use cases
  • Computer-use agent ran a browser for two hours creating API keys and setup unattended
  • One-shot a Counter-Strike-style 3D first-person shooter as a single HTML file with Opus 5
  • Digital PR: build interactive newsworthy pages instead of templated blog posts
  • Codex gave a short, clear answer comparing trigger.dev vs a custom plugin where Opus 5 produced walls of text
  • Route page-generation calls to Kimi K3 via OpenRouter for a dollar or two instead of another subscription
KPIs / results
  • Opus 5 costs ~4x less than Fable on the plan
  • Computer-use agent worked autonomously for 2 hours
  • 3D games one-shot in a few hours, running entirely in HTML
  • Kimi K3 served by ~6 providers on OpenRouter
Tools / build
0:00 / 0:00
What's the use case for this as a small business owner then? If you look at just the numbers, it seems to be a better model for the most part, but that's not necessarily how it is in practice. You need to rethink how you transmit information. To me, these kill the blog post. Like why would you make a blog post following the same format to teach something where you could make custom crazy pages that have like interactive elements, etc. With computer use, like I was logged into all these things and I was like, you take the browser and you go make the API keys, you prepare everything. I was able to do something else the whole time and this worked for two hours and did it. It's the fact that you can give that boring stuff to an AI agent that will set it up for you. Like that, that's game changer. There is a risk though with this. Anthropix new Opus 5 model beats their best models on nearly every benchmark, but one of the first big reviews opened by saying, Opus 5 is here and I hate working with it. The thing is, both things are true. And if you run a small business, the gap between them is exactly what you need to understand. We've been testing it all week and it's quite impressive in certain areas. People have been one-shotting full 3D video games with it, but at the same time, it's buried Gale in walls of text that he couldn't read. So today we'll tell you when to use it, when to stick with what you've got and one genuine unlock that it hands marketers. I'm Mark Webster and I'm joined by Gael Breton, my co-host and co-founder at Authority Hacker, where we test this stuff inside our business so you don't have to. All right. So the big news this week has been that Opus 5 is officially out and available for everyone who has a Claude account. Your first impression, Gale, you were very excited on Friday in Slack when you were chatting about this, but it since kind of changed a little bit. I'm looking right now at the Anthropic press release where they're comparing Opus 5 to Fable 5 to Opus 4.8 to OpenAIS GPT 5.6 Sol. There's a lot. If you look at just the numbers, it seems to be a better model for the most part, but that's not necessarily how it is in practice. Yeah. I mean, it's a good model, right? It's like, I'm not going to say it's one of the best models you can use and it's whatever. It's very good. The problem is like it comes after Fable and it still feels like, you know, you look at these numbers and you're like, Fable looks worse in most aspects, right? Most importantly, I think for the people who are listening to us, I mean, agency coding is nice, but the truth is most people who are not developers, they don't do coding to the level that they need to the absolute best model to do stuff. Which of these are important? Because they've got knowledge work near the top. They've also got business workflows towards the bottom. Are those the two numbers we should be paying attention to? I think novel problem solving is also an interesting one. Like it's basically like, it's kind of like visual puzzles and things that are like, you know, they would be interesting in normal work. I think the computer use as well is an interesting one because computer use is something that people use more and more and more. And if you don't use it, you should start doing it. Agency coding obviously matters because we kind of all do coding at this point. Like I'm sure you create Python scripts with Cloud or you have it write app scripts on Google Sheets or whatever. Even if you don't know that you do it, it probably is doing this in the background. It does some, right? And it's like, oh, like legal and house. House is an interesting one. It's like on purpose. They don't make it too good because they're afraid people design bioweapons. So like, of course, like basically the two things they're nerfing for is hacking and designing bioweapons, basically. And so that's the things they're kind of looking for. And they're on purpose making these models worse than they could be because it's scary, basically. Interesting. But the point is like, yeah, knowledge work. It's probably my favorite one. This one is like really like making presentations, working on spreadsheets, doing things people do, basically. It looks pretty good. Although there are some chart crimes in this. If we actually... Chart crimes. That's the first time I've heard that phrase. Yeah. Ah, they fixed it, actually. So there were some numbers where they would highlight the one on the left when actually the one on the right was better. So I think, for example, like this one, this one actually highlighted opus initially, even though the opus number was slightly below. It's also like they're really misrepresenting visually in the sense that look at how dark this red is. And then the GPT wins here, but it's much lighter. You know, it's marketing. I mean, you're being marketed to basically. That's all I'm trying to say. But one thing that is not really shown in this is, first of all, the model seems to be making much more errors than Fable in its judgment. Like its ability to judge things. It's opus level, right? Which is not as good as Fable. And so if you've used Fable before and you read this chart, you're like, great, I can replace Fable with a model that costs four times less on my plan. In reality, not really. And the second thing is it's insufferable to work with. It's like, it's a smart model. Okay, I'll show you an example, right? It's like I was working on like a plugin that I think is going to be canceled because there's better things. I found this tool called trigger.dev that does most of what my plugin does for very cheap. And essentially it's just better, basically. And so I was like, okay, well, it looks better than my plugin, whatever. Can you compare it, etc.? And that's the answer I got. This is something to read, right? It's a difficult... I'm looking at a pretty intense wall of text here with a lot of words I don't understand. Exactly. And it has custom instructions to keep things clear, etc. Like it has all of that in its instructions. It still gave me a table, which is nice. And I was like, can you visualize the difference in cost between self-hosting and paying the cloud version? Basically, I did that, which is nice. Like this little graph, etc. That's cool. So it's not too bad. However, let me show you what happened on Codex with the same question. This is the answer Codex gave me. And you just look at it. It's a short answer. This is a better workflow. And it doesn't make your plugin pointless, but it makes the current replays up here and it's untenable. Don't push it. It's much clearer, basically. And so that's what I'm saying when it comes to work. It's like, I feel like I understand more what's going on in Codex because it actually explains to me. It did the thing and then I ask it to visualize. It also did the visualization, basically. So yeah. Do you attribute that purely to the underlying model rather than any kind of files or preferences? I think both Fable and Opus, they are really hard to read by default. It's like they're smart. Like if you just judge the output of like build me this thing and build this thing, they're actually pretty good. That's not much to say. If you ask me to make a presentation, if you ask me to make all of that, it's really good. However, the kind of like walk in between to get to that step and the collaboration, etc. I enjoy GPT a lot more right now. I think it's just you can see like it's more enjoyable. All that's to say that as a business owner, if you're running like a small SaaS, a marketing agency, a service business, something like that. Do you even need to like care, pay attention to it? Or do you just change from Opus 4.8 to Opus 5 and just keep going? Yeah, it's still a better Opus. If you were on Opus, you'll find this is probably an upgrade, particularly in the quality of the output. Maybe not in the quality of the exchange you get with the mobile, but in the quality of the output, it's better, I think. If you came from Fable, you will find it makes more mistakes, for example. It goes more in loops, etc. It's a downgrade from Fable a little bit in terms of feeling at least. Yeah, but I mean, you could argue Opus 4.8. I mean, it certainly was also a downgrade. So is this just less of a downgrade? Oh, from Fable, yeah. From Fable, yeah. And also, it's really good at anything visual, like making websites. And people are now making video games with it that are like super good, actually. So I have some examples, actually, that I can show you. Like this is a video game. It took like a few hours, but it was one shot by Opus. Yeah, it's basically a counter-strike level first-person shooter or thereabouts. And people are making multiples. So for example, this 3D experience was also made by Opus, for example. And it's quite nice. You can see like the ability to generate 3D walls and all of that is now... Like you can see a lot of indie games are going to pop from this, right? I mean, like particularly we're, you know, moving through a grass field with long grass and each blade of grass is blowing in the wind. There's a steam train going over a bridge and the smoke is like looking very natural. I mean, yeah, I mean, this is quite some quite complex things that are going on here visually. Yeah. And so like it's interesting because this all runs in HTML, by the way. Both examples I showed you is an HTML file. It's not like a complex game engine or anything. And so you could have something like this on the web, potentially. And so that opens your ability to build things. Yeah. So for small businesses, that means that, you know, interactive elements on your site become much more realistic, like to do some very impressive stuff. I'm not just talking about games, but like, you know, disseminating information in like an interesting way or making it interactive. Or, you know, before the show, you were mentioning about like some digital PR opportunities and things like that you do here. And I think that's potentially quite exciting. Yeah. I think it's like you need to rethink how you transmit information. To me, this kill the blog post. Like, why would you make a blog post following the same format to teach something where you could make custom crazy pages that have like interactive elements, et cetera? It's like now we have these AI models that are able of like creating such nice things. We don't need templates as much. And we can have the model think of like how to make the whole experience more interactive and fun. And I think, yeah, digital PR is going to be a big one for that. If you're a marketer using these models for digital PR, creating something cool around a newsworthy topic, revisualizing things, making things look a little bit interesting, showing a different perspective on it is going to, in my opinion, gather attention. It's going to be the crossover of your ability to execute with this model and the right topic that is trendy at the moment. And if you're able to do something like that, then potentially you could get lots of exposure. So I think you're a marketer. You're trying to get something out of this new model. That's the unlock. To me, that's the one new thing that you get out of this for a marketer. On top of, it makes much nicer presentations. It makes much nicer web pages as well, et cetera. Like in general, you can do better things. Should we be using this for all websites, like over something like Fable? So for websites over Fable, yes, there is an argument that actually Kimi K3, which is an open source model, is actually even- Which we talked about last week on the show. So that was the one where they had blocked off subscriptions. Yeah, but now it's open source. Now it's been released open source. So you can buy the tokens on OpenRouter, for example, and it's usable. And it's arguably even better. Just to interject there, for anyone that doesn't know, openrouter.com is a site which allows you to interface with basically any AI model in the world. And you pay per token usage. There's sometimes a bit of an upcharge for that. But the idea is you have one OpenRouter account and it's just really easy to use any new models that come out. Sorry, openrouter.ai is the website. Yeah. So you can see there's like six providers right now. So Moonshot is the guys who make the model. So you can use them. But now there's a bunch of other providers. And all these providers can now download the model and serve it. So there will be more and more people offering the model. And so, for example, there's a fast version. You see they charge a bit more, but you get more tokens per second, for example. And these can be also, these are hosted outside of China. Because I know there's a lot of like data security concerns from some businesses for things like that. So it's a way around it. Is it like totally secure in that aspect? I mean, you can pick which providers you want to use. And it's like, let's say you don't want to. Moonshot is obvious in China. They're the company doing it. So it's like you can tick them off. And then you use another one. Anyway, that token per second is not very good. Better than digital. But you can just pick which providers you want to be able to use. And then you just basically get it. So it's a really cool thing to have if you want to like test new models regularly. And you don't want to have to like set up a whole account for each one every time. I think also like for unique capabilities. So, for example, this model is very good at content, right? You could just connect your open router API to your cloud code or your codex and be like, hey, whenever you need to just make a page, just call this API just for this. And then you pay for a few tokens, like you pay a dollar or two. And then that's it. Instead of buying another subscription or whatever, sometimes for like these little pieces that you need, like if you're like a codex person, for example, if you use OpenAI, they're not very good at content, right? Maybe instead of taking a hundred dollar cloud subscription, you can just buy some tokens from Kimi and have it do your front end through codex, for example. And then you kind of cover that end of what you're missing in your subscription. And you just pay as you go. What's interesting for me is you said that Opus 5 is better than Fable at front end design. Like how can that be? Isn't it supposed to be a worse model? It's just new training. Data, I guess. Like, you know, Fable is mythos. Mythos came out in March. We're in July. It's like there has been things happening in between and they've worked on new models. They've acquired new data, whatever, to train the models on. And so most likely there's like a set of data that is purchasable right now because also Kimi can also do the 3D very well. So probably like there's some data vendor that just was like, hey, new bundle of data you can buy, a bunch of data, a bunch of 3D training or a bunch of front end training. And it was bought. And now it makes Call of Duty or Counter-Strike in one shot, basically. And that's not in Fable because the model is older, but it's in Opus because it's a newer model. And just to keep on the topic of Open Router, are there any other models and or use cases where it would make sense to say like, you know, for this specific type of task, always use Open Router and use this model? Yeah. I'm thinking of someone who has a Claude and an OpenAI subscription, but nothing else. Let's say you want to label like data, like very simple task, whatever. You can use something like DeepSeek Flash, for example. DeepSeek Flash, if you look at the price, we're looking at per million input token for the cheapest ones. You're looking at $0.09 and then per million output token, you're not even $0.2. Yeah. Remember, like if you go back to Kimi, we were looking at $3, $15, right? So big difference. And so like, let's say you want to process a hundred thousand web pages and we write their title tag, let's say, for example, this is great. You will not have enough tokens in your Claude subscription, in your OpenAI subscription. And you're just going to either have to rotate the subs or you can just connect something like this and use one of these very cheap models. You know, it's like Gemini 2.5 Flash level, but like to write a title tag or whatever, it's fine, basically. And it's worth it. Like you want to do volume work and then you can have your Claude code or your codex write a script that calls this API. And so it still does it all. Identically for you, it's kind of just a terminal command and just calls this API and you pay very little money per API call. Cool. So anything else you want to talk about with Opus then? No, it's just like there's a lot of like people who don't like it. So for example, there's this creator. It's like when she released her review, it was like big news. Opus 5 is here and I hate working with it. It's the way it works. It's like if you blind test the output, people tend to like the output, but it's not a very pleasant model to talk to. And just for that, I'll find myself just going back to 5.6 and I'll just go to Opus like when I have to rather than just when I need to do a task. So would you compare this to maybe having like an employee who, you know, you didn't match personality wise with, but was very competent? Yeah. Yeah. Something like that. That's how I feel at least. And it's still not Fable. Like it will not be as smart and wise as Fable. So it's like, yeah, I still, I like the combo GPT 5.6 and Fable right now. Like yesterday I would talk about this later. Maybe I was like figuring out editing videos with GPT and it's like, it was stuck. And I was like, ask Fable and Fable just unlocked it. And it's like, it did all the rest of the work. It was just like one, it was not two questions to Fable. Like they went back and forth and then it just unlocked it and it went. And so I kind of like that setup right now where I use Fable as the advisor to GPT 5.6 and GPT 5.6 has kind of like the daily work calls. Do this for me, please. And it's easy to talk to and it's good still. Okay, great. So let's move on. Let's talk about the chat GPT app. But first I want to quickly mention our program AI accelerator. We're getting close to a thousand members in there. It's very, very popular. And we have quite a lot of cool stuff that I just want to show you around quickly. So if you haven't heard about it, it's a community that Gail and I run. There are a number of cool courses in there. If you've never used any of this agentic AI stuff, if you're new to something like Claude Code, we have a full course that gets you up to speed with everything there. Everything is pitched for non-technical business owners. So if you're feeling a bit left behind or, you know, want to kind of catch up with this stuff, this is a great place to start. We have a ton of other courses in there, new stuff being added like every few weeks as well. We have a skills library where you can download and work with all of the different processes that we're building within this, including, for example, the new delegate one, which Gail's built, which lets you use only the best model for the correct task. Yeah. So you're not burning through your Fable credits to, you know, do simple web searches. It creates this kind of like smart area. It makes it a delegator, right? You can use Fable as your main thread. It won't burn too many tokens because it just decides what to do. That costs you less tokens. We have lots of members who said like they get like 20, 30, 40% more usage out of their plans from this actually. And, you know, if you're looking for more simpler input output things, we have all the marketing processes that we build. For example, the LinkedIn carousels one, you can see the output of that here. So you can literally just download our skill, use it in your business and away you go. There's the weekly nuggets section where Gail is kind of like Gail's private blog of cool things he's working on. So if you want to learn how to vibe code your first app, that was... This was a Chrome extension actually. I built a Chrome extension where I opened a new tab and it shows a dashboard of the business. So like every time I open a new tab, I see everything that's going on in the business right now with the numbers, et cetera, without doing it. And this was basically one-shotted and I show my vibe code setup for this kind of stuff. And we also have a higher tier cut for plus members where we do like a weekly call where we just talk about like all of the latest things which were happening at the moment. But we actually show you, you know, before they're kind of ready to come into a course, like how we've built it. So this week we're talking about using Hermes Agent to build AI employees inside Slack. We were talking last week about editing videos with AI. We're almost there getting very, very close to being able to ship something that can do that. But you know, the underlying process wasn't there and folks have been kind of building on top of that already. And if you would like to become a member of the AI Accelerator, you can go to authorityhacker.com forward slash AI dash accelerator, and you can find all the details about what's inside, including current pricing on there. So we hope to see you inside. We're actually going to talk about that stuff a little bit later in today's podcast episode as well. So do keep listening to this episode so you don't miss that. But let's talk now about OpenAI's ChatGPT app, because we talked about that a little bit last week, and we've been both using like a lot more. And I'm super impressed with it. You've tried to use it because you were not using it before. And you were like, I don't know. So I tell you, I'm one of these people. I don't, you know, you switch every week. You're like, I'm going to do this, I'm going to do this, I'm going to do this. Like, you're terrible for this. But almost everything in life is kind of useful for us. But yeah, exactly. But I'm the opposite. I'm like, OK, I'm using Claude. I'm just going to stick with that. And let's just grind out and keep doing that. But like, I'm making a video at the moment. I've been forced to use the ChatGPT app. And honestly, like, I prefer it. I think I might switch to it. It's just nicer. Like, you know, the text editor exists or it doesn't really exist in the Claude app. But just like the smoothness of it, the organization, the folder structure, you know, archiving, like you said, that's really nice. And also, if you're brand new to this stuff, right? Out of the bat, ChatGPT's like ethos is let the user do what they want. Whereas Claude is like, let's control this a bit more. A good example of that would be if you want to use like the native connectors. So, you know, where you use OAuth to log into your Gmail account or Google Workspace, you know, you have a sheet that you want to pull some data on. In Claude, they'll only let you read the data or, you know, make a draft in Gmail. Whereas in ChatGPT, it will send the email if you want. And if you tell it to, obviously, it's controlled. And more importantly, in Google Docs and Sheets, it can edit the sheets, edit the data in the sheets. Whereas in Claude, you couldn't do that. Now, of course, you can install the MCP or CLI to do that. But that's like, you know, it's a more complex process. It's more difficult. People start out with this. Yeah. Intimidating for new users. So I think, yeah, they're kind of making it more accessible to newbies, basically. Have you used computer use yet? Have you had it take over your browser or anything like that? No, I haven't used it yet. God. It's like, this is going to blow your mind when you actually use it. I'll give you an example of something that I was doing today. I'm going to share my app. I'll talk about some stuff I'm doing. Okay. First of all, you were mentioning like the text editor. That's like, let's say I'm working on a new plugin that we're releasing soon, basically. And it's like in the Claude app, I can see the files, but I cannot edit them, which is very annoying because when I write skills, it's like, it's cool to have AI, but I like to also edit things. Here, I can just like open the file and I can just type, right? If I want, and it just edits. Or I can also build context. I can select text and just click add to chat and I'll write a prompt and we'll know what I'm talking about. Very, very handy. And then also you can see edits and you can write this and then you can write multiple edits and then it will add them to that. You can basically do something like you can actually send it, send multiple comments. Sorry. Let me be clear. Let me just reject it. But the point is I was setting up like a new social manager profile on our Hermes, for example, as you can see, nice photo. But the point is to set up the profile, I needed to connect it to Notion. I needed to connect it to Slack. And that requires creating API keys, for example, and like logging into the dashboard, creating these keys, et cetera, which is a pain in the ass. Like you need to go and do all that stuff. And then with computer use, like I was logged into all these things. And I was like, you take the browser and you go make the API keys. You prepare everything. And then you grab them and then you put them where they need to be. And you do all that. You prepare everything for me. And this morning there was like five tabs open, like the Notion tab, the Slack tab, the buffer tab that I gave it access to, et cetera. It just went and set up all the access and everything, like how much scope or access, all the stuff you should do for your agents, but nobody does. And then they end up with dangerous agents. Like it did all of that for me and it was done. Whereas if I did it with Cloud Code, it would have asked me to go and open the tab and put an API key somewhere or do something. I was able to do something else the whole time. And this worked for two hours and did it. So that's kind of why I like it. But I think the interesting part, like what was released this week in the Childivity app, because obviously this is a few weeks old now, is the new voice mode. And so before you were able to dictate. So if you type here, I speak, it's going to start dictating. So I'm going to like press stop, for example. And yeah, here we go. But the point is you have now this new option that is like the voice mode. And the voice mode is the live voice mode with ChatGPT. So actually, let me just start one in my main workspace. I'm going to click here. Just to be clear, though, that's not brand new. That's been around for quite a while. It's just this new one is not brand new. You will see why. So hey, ChatGPT, can you hear me? Yeah. Can you start a new thread and check if there is anything I should answer to in the circle community? Like anyone that needs my help. Yeah. Okay. And while you do that, can you also send a message to Mark and tell him what I've been doing with the new Hermes social media agent? And just tell him how he can use that? Oh, and then also, can you send a message to Milos and tell him that he should give me an update on where the retention engine is at? And I'm going to mute now. But the point is, like, it spun a bunch of threads. So it spun the chat that is going to review that. And when it's done, it's going to reply back to me. Oh, it probably sent you a message already. I don't know if it did. But it will send you a message. And it will do all these things. And the point is, this doesn't just work on your computer. This also works on your phone. So you can lock your computer. I just got your message. Cool. Thanks. Let me know when you have an update on the other stuff I asked you. Okay. And you can see it's basically going to babysit the threads. And it's going to basically handle it. And as I was saying, all right, thank you. But the point is that I can actually operate this from my phone on the ChatGPT app. And it will operate my Mac even if it's locked. It needs to be on, but it can be locked. So nobody else can use it. And that also works with computer use. So it can actually open a browser, do things. It can do things on your computer, et cetera. And so you end up with a mini-jab. Yeah, the unlock here is that previously the voice mode in ChatGPT was kind of shit because in order for it to be fast and responsive, they had to use like shittier models basically because they didn't have a long time to like think before it would react. And it wasn't able to do complex tasks because it would need to think. And, you know, the model wasn't capable of that. And it would just create bad experience. But what's changed now is that you still have that like fast layer for the interactive part where, you know, it feels close to like a natural conversation. There's still a very, very tiny delay, but it's really, really not that big. But now it's going to spawn new chat threads or essentially a sub-agent that is actually going to go and do the complex work with a better model that then reports it back to the voice agent, which that's the interactive layer that you have. So it creates this like fast interactive layer for deeper work. Yeah. And most importantly, it's very good on the go. And also like you need something on your computer, like you forgot a file, whatever. So you can see it's kind of like it went for the browser. I'm not sure it's my favorite way of doing it, but it will get there eventually, basically. It will just kind of start clicking around, figure it out and get back to me. I'm not sure I'm going to let it finish, but you get the idea. I actually connected it. Like there's a whole browser operating in the background to just answer my query, basically. So yeah, it's pretty cool. I'll show you the message that it sent to me. It said, hey, Mark, I've been building a new Hermes social media agent. The idea is to give it source material or a rough idea plus the audience and goal and turn that into social drafts that you can review and refine. You know, it's basically explaining what you're doing. It says sent using chat GPT. So I know that's not you. And I don't get kind of confused. And same with Miloš. He just replied, actually. So Miloš got the message and he just replied. So you get the idea. But like it can operate your inbox as well. Be like, hey, check in my inbox. Is there stuff to do? Do I need to make any decisions, et cetera? And you can put your AirPods in your ears. Go for a walk. Turn this on. Have your phone in your pocket. And just kind of like refund, brainstorm, make decisions, have it do research, et cetera. We're not fully there yet. You can see there's a little bit of awkwardness still in the thing. It's kind of cool. It's very cool. I think it almost like changes the work setup in terms of like hardware a little bit as well, because, you know, we both use MacBook Pros at the moment. And it's this kind of like, it's a nice, you know, it's portable if you need it to be. It's still powerful. But it might start making a bit more sense to have a more powerful, you know, always on desktop, Mac Studio, Mac Mini, whatever. And then use, you know, MacBook Airs and even just, you know, like your phone to do a lot of work if you can interact with it with your voice. And yeah, the way I see it is I'm going to get probably a strong desktop computer with lots of RAM because actually these things eat the RAM quite a lot. And MacBook Air probably, I'll switch my laptop to a 15-inch MacBook Air. So it's like a big screen, but a lighter laptop. And then probably a bigger phone as well, because I actually use my phone a lot on this now, like dictating into Codex. And now the voice mode actually makes it quite pleasant to use on the go. Like our support bot on the community was 90% vibe coded on Codex while I went for a walk for two hours. Yeah, it's been working for like two months now and people seem to like it, you know? So yeah. Brilliant. So anything else new in the ChatGPT app that's worth paying attention to? I think this is kind of the biggest one. They've also fixed a lot of small things. So for example, like before it was called like ChatGPT walk and ChatGPT, et cetera. If we actually look at it now. Yeah, just to explain here. So they merged Codex, which was the like Cloud Code equivalent coding app from OpenAI and the ChatGPT app, which, you know, the mass user one that everyone knows and loves into one app. But when they did that, it kind of inherited more of the Codex architecture. This is for the desktop app, by the way. It inherited more of the Codex architecture and it felt a bit off and like, you know, your conversation history wasn't easily accessible. And it was just, it didn't feel like we're still using ChatGPT anymore, but they fixed that now. So there's a selector. You can change between ChatGPT and Codex. And when you're in ChatGPT, you can change between chat, which is the normal chat and work, which is something like a Cloud Cowork equivalent where you're still working in a local folder. Kind of you get your projects on steroids basically with that. Yeah. They also released ChatGPT sites actually, which I should have mentioned. I actually had it try to make like an audience thing, but the point is like you can actually make full websites and they will host it for you. And they actually give you analytics for it now as well. So if you want to share something you've built, it's quite nice. You see, I go into sites. Do I get the analytics here? Yeah, you get the analytics. Obviously, I haven't shared it externally, so there's not going to be a lot of data, but the point is you can actually see what's happening. So you can create like mini dashboards, all of that, et cetera, host them and share them with people like Google Docs, which is very nice. So you can click on share. And see here, I can just basically choose just me or anyone on the internet. I think when you have a team, you can add your team as well. I'm on like an individual plan. That's kind of how they upsell you on the team. If it's on anyone on the internet, is it like public? They can find it or do you have to give them the link to find it? So we need to check that. I know there's a bit of a drama now with Cloud because Cloud has the same feature, but they indexes it. They indexed it. So if you're going to Google, you can find everyone's artifacts, including business data people may have published, et cetera. They've made it public. So it's very bad. So there's a lot of drama about this right now. This one I haven't tested. I hope they haven't made it live. But yeah, they actually give you like a database, et cetera. There's a lot in there. You can make like a mini dashboard for yourself or for anyone. What's the use case for this as a small business owner then? Like when would you want to use this? Company presentation, for example. Let's say you want to present some information, whatever. We used to do PowerPoints because that was the easiest way to design something visual. But now that you can make mini websites. So this was actually audience research based on like some brainstorm. I kind of like found this skill that was interesting. Like how do you brainstorm your audience for social media, et cetera? And it basically, you know, it's like, oh, what changed, et cetera. And just prepare the thing. But the idea is like it's an interactive way of showing like, oh, you should make three posts that are in the center of this circle, this other circle, et cetera. And just explains that. So it's like instead of a presentation, you can make a website or a page, for example, which a lot of people are doing. They're using HTML for presenting. They're using HTML for planning as well. Like they replace plan mode with like HTML plan. I have a skill coming up about that as well. But yeah, you get the idea. That's how it works. And it's more like an internal mini tool as well. Like you can make yourself a dashboard. You can connect your APIs to it. You can do all of that. And it's easier than like having to manage your hosting, you know? Yeah. I see this as like a straight up replacement for PowerPoint for kind of communication. Like, you know, we're doing a presentation on a call. That's not like a webinar or like a presentation to the board. But it's just like, you know, you want to communicate some ideas to your team or a client in a not super formal way, but like, yeah, just on a call. I think this is a straight up replacement and improvement for that. Yeah. So it's like, it's quite nice. Like overall now it becomes kind of like a nice little productivity suite basically all in one roof. And you don't have to deal with like a Cloudflare account or Vercel account. And if you're not technical, like it's very nice and so on. So yeah. It's nicer. It's snappier. It's crisper. Like the interface is cleaner. Oh, I see. You've been converted this week. It's an improvement on the Cloud app for sure. And yeah, I'm going to be spending a lot more time using it. So I'm interested to see if you stick to it or not. Like now it's like 356 is my workhorse, but they still use this for Cloud. Like I will maintain a subscription on Cloud and I will still use it for some stuff. And as we said many times, you can switch your folders from ChatGPT to Cloud and it basically transfers seamlessly. Like you can transfer your skills. You can transfer your Cloud.md to an Agents. So you're not locked in when you pick one of these. And I switch between both. Most people use both. Yeah. Most of my folders, they are both on Cloud and ChatGPT. Like if I go in workOS here, like if I open the files here, so I'm going to do this. Files. You'll see that I have a Cloud.md and an Agents.md basically. And that just allows me to use both. So yeah, it's that simple. All right, let's move on now and let's talk about video editing. So when AI first got popularized, ChatGPT first came out, the first thing to- Sounds old already. Was text content. Images and things like that were really a couple years away. It was really nano banana when that came out. And then GPT Image 2 more recently where it got, for all intents and purposes, solved. We're getting close to that with video editing. Maybe not the whole end-to-end Hollywood production, but the first part of video editing that most video editors do, where it's like cleaning up footage, cutting out um and ah, merging things, just getting it ready to be for the creative part to be done. That is solved, would you say? Or close to being solved? I'm 90% there, I would say. I'll give you a demo, right? It's like people, like it's easy to talk. It's hard to do. So I'll show you a demo. So basically, I'm just going to show you something that we've built for the podcast. So one thing that we wanted to do for the podcast is we wanted to create this like 30, 40 second teaser video at the beginning that shows kind of the best moment, the best clips. It makes you want to watch it, basically. A lot of big podcasts do it for a reason. It helps retention. But it's really annoying to do because you need to really kind of like identify the pieces. Then you need to grab them from the different parts in the podcast. I need to edit them in a way that creates a bit of a narrative, etc. And so I gave it the last podcast that we recorded and I gave it Diary of a CEO that does it. I was like, you look at how Diary of a CEO does it. And then you try to make something similar with this kind of like question at the end and black screen, etc. Like how they do it. And this is what it did. The name of the game in AI is going to be efficiency over the next year. We're kind of drunk on free credits at the moment. Honestly, I can do six hours of this a day. I came back yesterday from holiday and within 40 minutes, I had eight threads running. I think sometimes you have to let go. Sometimes I'm just like, I have no idea what's going on. The point of Fable is you don't look at it. It's not Fable yet. One more generation. And then we're going to have models that are strong hackers that have no guardrails. A lot of people now, they have their code code or co-work connected to their tools. They have a little folder that they walk in with some data about their business. But like, how do you get five times more done than that? Yeah, I mean, it's pretty good for like a fully AI generated version. You know, maybe there's a few things you might clean up here and there. I agree. It's not perfect. That's great. I think it's a seven and a half out of 10, something like this. Like there's still room for like pushing it to 8.5 or 9. I mean, again, we would add some music in the background to make it a bit more epic. Maybe we need some kind of like slow zoom ins on the faces as you're talking, that kind of stuff. But that's easy to add. Actually, it's very, very easy to do. The hard part is like picking the right moment and then doing the right cuts. Doing the cuts. It's not as easy as it sounds. Like the transcripts don't work perfectly, etc. It took me a lot of time to figure out. But now it's like it's clean, the cuts. And so I managed to also make it do shorts and do shorts pretty well, actually. So I suspect like maybe not this week. It's going to take us maybe another two weeks, maybe to really have it to a state where I'm like, this can run easily in one shot without me babysitting it. But I'm pretty sure our short system is going to be overhauled completely from that because I get better results than Opus Pro, basically. So, yeah, we are pretty close to solving video editing. And again, this is kind of like just the cutting, but there is already tools that allow you to do animations. For example, like there's some stuff called Hyperframe that you can install and use HTML to make animated graphics. And then we have video models like C-Dance that we can use for B-roll as well. So if you want to make like eight, ten seconds B-roll that you can put over people talking, that would be a very easy thing to bake into the skill as well. So what my goal is like once I bake all these things together, we're going to have a decently competent video editor that can, you know, I don't think it's going to necessarily going to make you like a YouTube channel with 10 million subscribers. But I think it's going to be great to like edit video testimonials, edit product demos, edit tutorials, edit all of that, like edit this podcast. A lot of the day-to-day stuff you would want to use video for in your business that's not necessarily a YouTube channel. Exactly. And then you could shoot yourself like rambling and then have it edited to decent social content, for example. Like that would be cool, like for stories even, or like, again, eight out of 10, but better than not done, you know? Yeah. Yeah. I think that's the key here is that for anyone who is, you know, like a full-time creator or like is big on this. I don't think this is replacing that like really anytime soon. We've seen that with text content on social media as well. It's about as good of a writer as it's going to be, but it's just, it misses the human element a little bit. And that's kind of what makes things do well on social, I think. You can, but usually it's a lot. Like if you do well on social, it's like you're making like baby podcasts and stuff like that. If you're trying to make like serious content and then it's like fully AI generated, it's kind of difficult. And there's not a ton of examples of people doing that or they try to hide it or they hide it very well. I don't know. There's a lot of channels like we're not touching. So for example, like one thing I'm working on right now is like our authority hackers social channels, they're completely dead, right? Like we're not using them because nobody's managing them. So I'm like, what if I try to have an agent manage them and just at least communicate as the company on things we do, even if it's not perfect, it's like, it's going to reach a few hundred people here and there. And people who know about us will be a little bit more in their feeds and they'll know more. And then because it's the brand talking, it's kind of like not trying to pretend to be a person. And therefore, I feel a little bit less about putting an agent in charge, basically. And in that case, I think it's less about being a thought leader, which case like people don't want to necessarily follow AI as being a thought leader. They want to follow people. But it's more about just communicating what are Mark and Gail Authority Hacker up to? What are they shipped? What have they done? And that's a much more like factual thing. It's like almost like reporting as opposed to thought leadership. Yeah, it's the brand as well. It's like brands did not really people. And so therefore, it's kind of more OK if it doesn't sound as human, basically. So yeah, that's the kind of stuff that I'm working on. And so my goal is to combine both, right? It's like, oh, the social manager for the brand takes the podcast, cuts some clips, puts them on social profiles, et cetera, and start sharing that. And they still are human element. We're not automated yet on the podcast, but at the same time, it's kind of like we don't have to think about it. We just upload the podcast on a Google Drive and it's just like it picks something good enough, basically, which I think it will do. So give me a few weeks. Let's talk then about Hermes Agent, because that kind of underpins a lot of the plan for this as well. So what is it, first of all? And why are you suddenly convinced it's worth paying attention to? Because for a long time, you're like, this is shit. Just use Cloud Code. It can do everything. What changed? Okay. One simple thing changed. And that's something I've shown you already, actually. And it's the codec thing. I showed you this thread on codecs where I told you like it set up all my API keys, access scope, et cetera here. And what that means is if you want to set up these agents, basically, you need to give them access to things. They need to be able to read data. Can we just explain what Hermes Agent is or OpenClaw, these things for people that don't know? They're basically, in our case, the way I see it is I don't want to replace my codecs or my Cloud Code. These are agents that I use when I'm at my computer or my phone. I use them interactively, right? But my goal with this is to really build AI employees, give them an area of responsibility, give them a list of tasks that come with this area of responsibility, regularly communicate about them and refine them, and then have them take care of it, basically. And so these agents live on a cloud server. So my computer does not need to be on to run. They are connected to... Just to be clear as well, you could run it on an always-on local machine as well. There's a lot of people that do that. I just don't want to be dependent on it. Like if I'm traveling, if there's no electricity at home, whatever, I call them ambient agents, which is like agents living in common areas of the company. So in our case, they live in Notion and they live in Slack. And so you can chat with them on Slack and you can assign tasks to them on Notion and they can interact with it. They can create pages, all of that, basically. And so that creates like a workspace where you have access, I have access, whereas like my cloud code and my codecs, you don't have access, right? There's things that you cannot use it for. And so my goal is to kind of... Like I created like one main agent. I'm going to go more in detail when we talk more about this, but I don't want to go too deep. But I created one main agent that is going to be the manager and I'm creating subagents that will have areas of responsibility, basically. And then each agent has different access. So for example, like the main agent has a lot of access, knows a lot of things, but I don't let it write in the outside world just for security because it's dangerous that there's some data that I don't want to leak out. However, like the social agent does not have much data. Like you can read the member area, like stuff that you might pay us for, but there's no like financial data or anything like that. Like none of that. So that agent can actually draft on social media. I don't even let it post, but you can make drafts. And then the idea is when you create these agents, the way you prevent it from posting is you don't give it the access rights to do that or you just tell it, please don't do it. Both. So like the MCP server has access to right now, like I created a custom API key that does not allow to schedule and post basically. So it's like, even if you tried, it can't basically physically. The process of creating that is like, you know, a lot of people skip over and they just say, Hey, give it full access to everything. And that's where there's. That was the problem. And it's, if you want to do all these customized access granular stuff, et cetera, it's a pain in the ass. Like you need to go through all the API setup. You need to like find the right boxes, understand what this means. You know, you get a one liner, you need to understand what that actually means. It sucks. It really does. And it's hours of time, which I was like, I'll just have my codex and I'll just walkie talkie to it in my phone, basically. But now codex actually is in charge for setting up everything. And so my codex is connected to all these things because like it doesn't do anything unless I'm here anyway. So I trust it for that. And so I'm like, Hey, cool, let's just set up this new agent that is connected to this this way. It can do this, but it cannot do this, et cetera. And because it has computer use codex, it will go in my logged in accounts. I will actually create these API tokens and everything and configure the server for me and do all the work. And all I have to do is talk to it. I think if we go back to the beginning of this chat, this is what I told it. I said, Hey, I want to create a new profile on the Hermes that will act kind of like a new employee, a separate profile that should have access to both like a notion. And then we're going to connect it to buffer and we're going to make it a social manager for the Autohackr brand. So we're going to give it profile for Autohackr, Twitter, Instagram, LinkedIn, and Facebook page, all that stuff basically. And it's just going to manage it through the buffer MCP. Can we just plan this setup? It should really act like a separate entity from the main bot. So the way it works is you should read documentation and set it up. And basically that's how it started, right? And you can see this was dictated. I'll just talk to it. And it just asked me some questions and then I answered and then eventually it did all the work. It worked for like two hours and set up everything up basically. And so that's the unlock. The unlock is you have your interactive agent managing the creation and the management of all accesses to these things. And it's a lot less of a pain. Yes. So we'll see how it goes. We're still at the beginning of this. So I don't want to talk too much. I think we'll make a full episode on this, but yeah, I'm excited. It's quite exciting to see that. So there's this thing I noticed as an AI content creator, like someone in the industry, and that's the people, business owners, like consumers, our customers, they react very strongly and very positively when you talk about AI, like an AI employee or like a, you know, you have an AI chief of staff or an AI social media manager in your company on site. These types of things. I think because people's mental models of how businesses operate, this fits in perfectly to it. Whereas like an agent living in cloud code on a server or whatever, like people sometimes struggle to conceptualize what that actually means or is. So when people create content around this, like, Hey, I replaced my social media manager with AI or whatever, it works very well. So you get people kind of playing into it a lot, but not necessarily like covering the fundamentals very well because it's like, yeah. And you're creating custom API keys, you know, to limit access and security protocols and things like that. It's like, ah, send me to sleep. Like, how do I get it to make me social media posts that will make me go viral? Like that's what people care about really. But it's the boring mundane stuff that actually is a big unlock for this and makes it quite, quite. It's the fact that you can give that boring stuff to an AI agent that will set it up for you like that. That's game changer because now you can just be high level, be like, Hey, we're recruiting a new social media manager. So I'm making kind of like a skill now to create AI employees basically. And the idea is like, there will be like a questionnaire or something. And then it will just basically spawn a new one. It will know how to create all the tokens. There'll be added to Notion. I have already selected, added a selector on Notion. So you can select which agent does the work and so on. And the idea is, yeah, you run this as like a recruitment thing and you can build your virtual team. Like you could be a one man company. Probably if I started again, if I was running a solo company, I'd run it on Discord just to not pay for Slack. But you like Slack. So it's fine. And you can have like an entire team basically. And they all work with each other. They all have different access and so on. And then the thing is like, it's very exciting because when new models come out, your employees all get upgraded basically. And like we use 5.6 right now as well. Like we use 25.6 medium on this. It's a good balance. I think I'm excited. And you can use OpenAI subscription for this, but not Claude. Is that right? Yeah. So, I mean, I think you can use Claude, but they don't like it very much. There's a good chance they will cut it eventually. And anyway, because of GPT 5.6 kind of like tenacity to like finish the job, I think it's a much better model for it. Whereas Claude is more likely to kind of like stop in the middle and ask you for confirmation and so on. Whereas this, you kind of like send it and do it. And so we'll talk more about this. Like there's all the Chrome jobs and MCPs and so on. There's a lot. There is a risk though with this. There is a risk in that OpenAI could just say, okay, Hermes agent, you can no longer use subscription on it. Yeah. You have to pay API costs. And then there's like an efficiency question to be had on that. So that's something to bear in mind. But, you know, they could also just 10x the price of their subscription in a couple of years time anyway. So you need to have to deal with the same thing. Yeah. When you're using Codex or Claude, it's the same thing. They could be like, oh, no, now you pay API price. It's kind of the same. And the fact that OpenAI purchased OpenClaw tells me that they are unlikely to cut that. Like if they were allowing it on OpenClaw, but not on Hermes, then it's like antitrust. Like there's a lot of stuff where it's unlikely to happen in the short term, at least in the long term. Nobody knows where this goes. But I don't want to think about long term because it's like this industry changes so fast. It's like we'll adapt when it comes to us, you know? Yeah. So yeah, that's something we're working on. All right. So we'll wrap it up there then. Thanks for listening to this episode of the Authority Hacker Podcast. Please do us a favor and leave us a five star review on your favorite podcast app. That really helps us get the word out and get in front of more people on the audio platforms. So thanks again. And we'll see you next week for another episode.