← Back to search

Claude Opus 5.5 AI Just Changed Everything

AI News Today | Julian Goldie Podcast · 2026-09-23 · 10 min
relevance 62 2046 words Episode page ↗ Audio ↗
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How to switch to Claude Opus 5.5, understand its gains over Opus 5 and Fable 5.1, and prompt it for long autonomous runs without wasting tokens.
Benefits
  • Fable 5.1-level performance at ~40% lower cost than Opus 5
  • More usage out of existing Claude subscription plans
  • Communicates more directly, thinks before every reply
  • Works longer autonomously with fewer stops
  • Hermes Agent now runs on Claude subscription via new plugin
Use cases
  • Running Goldie Bench tests on Opus 5.5 via Claude subscription (it first wrongly tried OpenRouter)
  • Prompting with three sections: whole task, finish line, when to stop
  • ClaudeMD rule: keep going without input, status notes inline, ask only before destructive actions
  • Auditing every service in a codebase with one subagent each, evidence checked before acceptance
  • Built a Remotion promo video with working controls and a lava lamp/black hole visual unprompted
KPIs / results
  • Agentic coding 66% vs Opus 5 at 52%
  • ~40% cheaper than Opus 5: $4 vs $5 input, $20 vs $25 output per M tokens
  • Cache reads $0.20, cache writes $5
  • Beats Fable 5.1 and Opus 5 on Terminal Bench and Frontier code
Tools / build
0:00 / 0:00
Claude Opus 5.5 just dropped and it makes your Claude way more powerful. This is the smartest everyday model Anthropic has ever shipped and here's the best part. You get way more usage out of the plans you already have. That means more work done, more projects finished and way fewer moments where Claude just stops on you. In this video I'll show you exactly what makes 5.5 so powerful, how to turn it on in one click and the one thing you absolutely have to do after switching or you'll miss out on its biggest superpower. I'll show you that near the end so stick with me by the time this video is over your Claude will feel like a brand new super smart AI employee that works longer, thinks harder and gets straight to the point. Let's get into it. Claude Opus 5.5 just dropped. This is the first model in the Claude 5.5 family so there should be more coming and this actually performs at the level of Claude Fable 5.1 for most tasks but it costs about 40% less to run than Opus 5 which means even if you're on the subscription, you should get more usage out of it but we'll come on to that in a second. This is the first model since they mentioned pacing for the Frontier. So as with previous models, they've tested it with external valuations. Basically they were calling for like a slowdown when it comes to AI releases and that's why I think they've talked about the pacing for the Frontier. If you're wondering how it performs versus Opus 5 whether to switch or what the differences are, we'll see the biggest difference here. So for agentic coding, it's scoring 66% versus Opus 5 at 52%. It's a huge step up right there. And you can see the benchmarks versus Fable 5.1. So it's beating Fable 5.1 and Opus 5 or agentic coding on Terminal Bench. On Frontier code, agentic coding 5.5 is beating both Fable 5.1 and Opus 5. Agentic coding 57 which is beating Fable 5.1 and Opus 5. And it's pretty much crushing on all the benchmarks like you can see. So a big difference in the power of this as well. This also comes at a good time because we now have Hermes agent that is compatible with your subscription on Claude. Previously yesterday, you couldn't use this inside the API, but now you can use Hermes agent with Claude on your subscription, which is pretty cool. You just dropped a new plugin and you can see the difference in prices right here. So for input tokens, $4 versus $5. So Opus 5 was a lot more expensive than Opus 5.5. Output tokens, 20 versus 25. Cash reads 0.2 and cash writes $5, right? So it's just overall a lot cheaper, 40% less. And if you're wondering how it performs versus GP6 Aster on benchmarks, you can see that right here. We're actually running tests on Goldie Bench. So once we've done that, we'll test it out in a second. And I've got to say, like, Opus 5 itself, not the best model release. So hopefully Opus 5.5 is a lot better as well. And I think they heard that because, for example, Claude mentioned here, Opus 5.5 communicates more naturally, addressing some of the most common feedback they've heard on Opus 5, as you can see right here. So it's just a lot less for Bose and it should get straight to the point. A lot better. And they've released like a big report on how it works, all the differences, etc. If you're wondering what the differences are and how they compare side by side. So number one, it's a lot smarter than the previous version. Number two, it's cheaper and faster. It's also much better at research. It should explain things more clearly. It's designed to be safer on behavioral tests, more honest about limits. And you can use it right now pretty much everywhere. So if you go to Claude, if you don't already have it, but you're on a pro plan, you can just go to Claude in your settings over here and then check for updates and then just run the latest update. And that should help a lot. By the way, it's still not perfect, right? So I'll give you an example. I actually said, can you run the Goldie Bench tests for Opus 5.5? And I made sure we selected the model 5.5. And then he started trying to use the open router to use Opus 5.5 instead of the subscription and the model already has, which is a bit worrying as a start. But let's see how it goes later. And so I have to correct it and say you should be using my Opus 5.5 subscription, which you're already using right now. So it's beginning to build those out, as you can see. Not a great start, but let's see how it performs later. Also, GPT-6 released two new models today. So they released GPT-6 Soul and Luna, which are pretty good models. And you get a lot more usage out of them. But Opus 5.5 is the biggest release today because if you were using Fable 5.1, you know that it hit usage limits pretty quickly, but also that it struggled to stay on track. It quite often got distracted, as you can see here. Not impressed with its first build, I've got to be honest. Now, Anthropic have also released some guidelines on how to get the most out of Opus 5.5. Opus 5.5 if you're using it. So they've said hand over a task to find done and when to check in. Drop, think carefully. It always thinks first. And after a long run, check what it needs to go further. So let's go straight into that and dive into this guide. So this just dropped to Opus 5.5. But the difference here is that a few things behave differently. So, for example, it can work longer autonomously on its own. It's more direct and it thinks before every reply. So what they've recommended is that when you're using Opus 5.5, hand over the whole task. Tell it what done looks like and when you want it to stop and ask. And then you can just go off and let it work. If you're using, for example, think carefully inside your prompts, you probably don't need that anymore because Opus 5.5 already thinks before every reply. That's not something that I was using previously, but I assume because I've put it inside this guide, a lot of people are doing that. And then also when it recommends when a long run ends, read what it needs from you first. So if you want to explain what done looks like to Opus 5.5, here's an example. So you break it down into three steps. Number one is the whole task. Number two is the finish line. And number three is when to stop. So that's pretty interesting in terms of the way you would actually prompt it now. You split it into three sections. You're like, okay, here's the whole task. Here's the finish line. Here's when to stop. And you go from there. And also if it's running on a long task, you can type a follow-up whilst it's actually working. So because it's running longer, a restart would cost more because it's going to use Op tokens every time you respond to it. So what you can do instead is if you're in Claude code, you can type the message and press enter whilst Claude is working. Let's test if that works inside Claude desktop as well here. We've got some interesting stuff here as well. So you can see it's got the running task here. And then you can add like a quick bit that you remember of it. Let's see how it's performing of it. Created a nice little lava lamp. Actually looks pretty nice. The black hole. Seems very visual when you use it. This is looking nice. Also, if you're using it for design work that's interesting, name the styles you do not want. So when you ask for a page, an app, an artifact, etc. List of design habits you don't want it to use. So for example, you can see a prompt here for designing. You would say build a personal website or placeholder content, but do not use this background or this type of accent or these type of labels, etc. Now also you can steer it. So if it's running long for like multiple hours, for example, inside Claude2Code, you can put a short rule in your ClaudeMD file about when to stop and ask and when to keep going. An opus 5.5 will then keep you posted as it works. So on a long task, it sometimes forgets to report what's going on, how it's going, etc. But if you have to restart, it's going to use more tokens, right? So how do you do this? So you can put a rule inside your ClaudeMD file. So when a step doesn't need my input, just keep going. And that will make it much more autonomous. Put status notes in the same message as your next action. So you can basically see its status and what it's doing as it goes along. Then you can ask it to stop and ask only when you can't continue or before anything destructive. For example, like deleting something or force pushing or changing something outside the repository that it's currently working on. So some interesting tips there on how to use it. You can also see they give some recommendations on how to split big work across subagents with opus 5.5. So let's say, for example, you're doing a review or an audit across a large code base. You could ask it to split the work across subagents and then check each result. And this just makes Claude opus 5.5 a lot more autonomous again, right? So for example here, you could say audit every service in services. Give each service its own subagent. When a subagent reports back, check its evidence before you accept it. And then it's just working in parallel of multiple subagents. But opus 5.5 is going to check its evidence before it comes back to you, which means that you don't have to manage it so much. And it becomes a lot more autonomous. Now you see here is also created. We gave it a test run for creating a video. This is with remotion, which looks pretty nice as well. So basically, you can create animations. You can create nice videos. I haven't literally prompted it. It's just come up with this on its own. As you can see, it created like a promo video. Nice little progress bar. The controls actually work on it as well. So it's kind of an interactive website with a video in the background. But yeah, beautiful animations right there. And it created this cool little fun game as well, as you can see. So that's basically how to use it, how to get the most out of it. Once we've completed all the tests on Goldie Bench, we'll get back to you on that. But pretty good so far. Now, if you want to get more training on opus, how to automate business, how to generate more leads, get more customers with AI automation, feel free to check out the AI Profit Boardroom. Link in the comments description or go to the AI Profit Boardroom.com. And inside the community, you can ask questions and I personally answer them. You can direct message me as well if you want to get help and support personally. Maybe there's something private you're working on. Inside the classroom, you can access all of our best trainings. We have a brand new beginner to expert roadmap, which you can see right here that tells you exactly what to do week by week. All our best playbooks and all the stuff I personally use to grow my business. Inside the calendar, you can jump a weekly coaching course, ask questions, get help and support in real time. In the map, you can meet people in your local area. And it's all available inside the app. Link in the comments description or go to the AI Profit Boardroom.com. Thanks for watching.