← Back to search
Hermes Mixture of Agents DESTROYS Claude?
AI News Today | Julian Goldie Podcast · 2026-06-28 · 16 min
Show full episode description
Hermes Mixture of Agents: Build a Model Council That Beats Single Frontier Models (Goldy Bench #2)Hermes released Hermes Mixture of Agents, which the speaker integrates as a “Hermes Council Engine” tab inside their Agent OS to run a panel of frontier models (e.g., Opus 4.8 and GPT 5.5) that answer privately while a third “chair” model judges and fuses a better final response. They tested it across a 42-task leaderboard on Goldy Bench where Hermes Mixture of Agents ranks #2 overall, behind Fusion, and they show side-by-side examples versus Fusion and Claude Opus 4.8, claiming Hermes outputs are often less buggy and more usable. The episode argues to “stop chasing the model” and instead build swappable systems that squeeze more from existing models, and it promotes the Agent OS/AI Profit Boardroom bundle with the Mixture tab, Fusion and Sakana panels, shared memory dashboard, tutorials, coaching, and setup roadmap.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Getting beyond-frontier intelligence without new model access by combining existing models in a council system.
Benefits
- Council of frontier models fused by a chair for better outputs
- Beats individual models without waiting for new releases
- Models are swappable parts; the system stays
- Workspace shows everything built, easier than juggling tabs
- Runs on cheap or free models with token optimization
Use cases
- Ranked #2 on Goldie Bench across 42 tasks, above Opus 4.8
- One-shot generated 'The Dragon Realm' open-source game
- Built 3D Skyrim-style open-world game from one prompt
- Neon City and racing games beat Opus 4.8 outputs side by side
- Automated video system with avatar, research, script, voiceover
KPIs / results
- Ranked #2 of 42 tasks on Goldie Bench, only behind Fusion
- 191 pages of community wins and testimonies
- Running Opus 4.8 + GPT 5.5 with a chair aggregator
Tools / build
- Hermes Mixture of Agents (Council engine)
- Agent OS mixture tab
- Goldie Bench leaderboard
- Fusion and Sakana Fugu panels
- Automated video pipeline
📑 Chapters — tap a time to jump there
00:00
Hermes Mixture Intro
- Hermes Mixture of Agents: council of agents fused for better outputs
- Response to gated frontier models
00:41
Council Engine Explained
- Council engine: panel of frontier models merged by a chair
- Opus 4.8 and GPT 5.5 answer privately, chair fuses
03:53
Leaderboard Results
- Ranked #2 on Goldie Bench across 42 tasks
- Only Fusion beats it; outperforms Opus 4.8
04:38
Showcase Game Outputs
- Showcase outputs: Dragon Realm, 3D Skyrim-style game
- Visual, works out of the box one-shot
05:44
Fusion Comparison
- Fusion vs Hermes side by side
- Both produce great Dragon Realm outputs
09:20
Stop Chasing Models
- Stop chasing models, build the system
- Mix of today's models already beats everything
10:08
Agent OS Workflow Demo
- Agent OS mixture tab: switch models, run panel
- Also Fusion and Sakana Fugu panels
12:06
Why Systems Win
- Top two on Goldie Bench are systems, not models
- Swap models when cheaper or better lands
14:34
Recap Key Lessons
- Recap: proof via Goldie Bench, side-by-side comparisons
- It's about systems, not the model
15:41
Join The Community
- Join community: coaching, tutorials, 30-day roadmap
16:31
Final Thanks
- Final thanks and call to action
So Hermes just released Hermes Mixture of Agents, which is a powerful way to basically have a council of agents working together to get better outputs. And this is obviously based on like the restrictions with GPT 5.6, Fable 5 being taken down. It's like, okay, how can we achieve better levels of intelligence using systems where agents work together? And then you fuse the ideas to get the best possible outputs. And from what I've tested so far, this is a really powerful system. And I'm going to show you exactly what we've built with it, how it works, etc., how to use it, and also how it performs on Goldie Bench versus everything else. So let's test this out and see how it works. Now, if you want to indicate how does this whole system work together. So basically, I call this a Hermes Council engine. So you have a panel of frontier models merged by a chair. And I actually put it through my whole 42 task leaderboard. So I've tested it with 42 tasks, you can check them out on Goldie Bench, as you can see right here for Hermes Mixture of Agents. And it actually ranked number two on the entire board, pretty much above every single model out there. The only one that's beating it right now is Fusion. And also, Opus 4.8 was outperformed by a long way. And I'll show you the comparisons in a second. So if you're wondering how does it perform on the benchmarks, it's doing pretty well. And I'll show you the examples in a minute. Now, if you look at the quote from news research on this who created Hermes Agent, they said, Hermes Agent now exposes mixture of agent presets as virtual models, capabilities beyond the publicly available frontier. So if you want to indicate what is this? Well, basically, what we've done is we set up a council engine, which is a new tab built inside Hermes Agent, inside my agent operating system. And this is the setup right here. And so it's a new tab we're built inside the agent operating system. And basically, you can pick a few frontier models, as you can see right here, any providers mixed together. I'm running Opus 4.8 and GPT 5.5. And then when you ask a question, both models answer it privately at the same time. Then a third model, the chair, reads both answers, judges them and writes one final answer that's better than either. And we can see the actual outputs here in terms of the stuff created. So the panel of models is the council. The chair is the aggregator. You have one question that you put in and one better answer out. So news research shipped this idea as a mixture of agents. I wired it into a tab, gave it a workspace and put my own leaderboard on it to see if it actually holds up in terms of performance. And basically, the way this works is like a panel of experts will beat one genius. So if you picture one brilliant person answering a hard question alone, or if you now picture a panel, each expert writes their own take privately, a sharp chair, reads all of them and gives you the best combined answer. So the panel would win pretty much almost every single time. So you have one prompt, Opus 4.8, for example, and GPT 5.5 work together. And then you get a chair that works and builds it and fuses it all together. Now, in terms of the performance, you know, you might be wondering, okay, how does this hold up? So if we have a look at Hermes' mixture of agents here, on the leaderboard, it's coming at number two, just below Fusion. And if we pull up the answers here, you can just test yourself out yourself, right? So if you're thinking, oh, this is not that good, or it doesn't look that great, et cetera, just have a look on this website and see what you think for yourself, right? Okay, you can make your own mind up or you can test it yourself. That's really why I created the Goldie Bench is because we wanted to test all of these models on example prompts rather than just, you know, listen to some theoretical benchmark in a lab somewhere about, you know, an AI model we can't even use yet. So that's why we created this system. And you can see here, for example, the stuff we've created is super nice, like looks pretty cool. It's very visual. We've tested out on loads of different stuff. So for example, like this is called the Dragon Realm, which is an open source project, as you can see, which is pretty cool. It looks super nice. The outputs are super nice. And I'll show you how that compares against like, for example, Opus 4.8 in a second and also Fusion in terms of the outputs. But it looks really cool. Pretty much everything that we built works really nicely straight out of the box as well. Here's another example. So this is basically like a 3D version of Skyrim. Like an open world game. And you can see like it looks absolutely amazing. For something that's just like one shot, you know, you give it one prompt, you get two answers fused out. It looks absolutely great. So we tested it on loads of different benchmarks, as you can see right here. Performed very well, pretty much all the tests. And if we compare it versus, for example, Fusion, let's see what we got back. So we can see head to head how they perform side by side. So, for example, this is the output from Fusion, which looks pretty cool. This is the output from Hermes. So it's more like 2D, but it still looks super nice. This is a game called The Dragon Realm. This is the outputs from Fusion, which again is number one on the leaderboards. And if we compare that versus Hermes Agent, in some ways this actually looks nicer. But they've both created like great outputs. Now, if you're wondering, okay, how does that compare versus something like Claude Opus 4.8? Let's have a look here side by side. So Claude Opus 4.8 versus Hermes Mixture of Agents. So this is the Dragon game we created with Hermes Mixture of Agents. Looks super cool. This is the one from Opus 4.8. This is the output from Opus 4.8 on like a Neon City simulation. Here's the one from Hermes. And this one actually looks nicer, I would say. It's a bit more like you can actually move around with it. It feels less linear, if that makes sense. And here's a racing game side by side. So this one is from Hermes. Pretty cool. Pretty nice to navigate. This is the one from Opus 4.8, which just feels like a little bit buggy. See how it's kind of sideways when you're using it? Doesn't quite make sense. And also the graphics in the background don't look quite as good. This is a 3D racer. So this is with Opus 4.8. It's pretty much broken. Doesn't quite work as well. This is with Hermes Agent using a combination of ChatGPT and Opus 4.8 together. Much better outputs. Way less buggy, more fun to play, more useful, etc. So overall, it's performing really well. I actually created this benchmark after Fable 5 came out. So I can't tell you whether it's better than Fable 5. But as soon as that gets restored, we'll test it out. Same with GPT 5.6. And right now at the top of the leaderboards, Fusion and Mixture of Agents are the two best features I've seen to create the best quality outputs. And you can test it for yourself. Like all the stuff that I've created with this is just hosted on that website directly. Now, the thing that I would say about all of this, you know, whether you're using Hermes Agent, whether you're using, for example, Claude, whether you're waiting for Fable 5 to come back, is like stop chasing the model. Build the system, right? Everyone is waiting on the next model, the next Opus, the next GPT. But if you look at what just happened on the benchmarks, a mix of today's models already beats everything else, right? So you don't need a new release. You don't need to worry about gated access with this. You can just use a smarter way based on what's already out there. And the model is a part you can swap. The system is a thing that you own. And the timing makes this huge, right? You've got Fable 5 in preview. You've got GPT 5.6. It delays. You've got the next Frontier models that are not actually available right now to the public. And so the winning move is not waiting. It's squeezing more out of the models you already have by combining them. So, for example, if you check out the Agent OS system that we have over here, we have the mixture set up right here. And this is so much better than using the terminal because you can see everything that you've created inside the workspace. You can switch the models based on what you want to use. You can then ask the panel anything and run the panel whenever you want, right? So you just click run the panel like so. It's the same, for example, if you're using Fusion, which is a very similar process. You can use it exactly the same way. So you can chat with over here and then see everything that you've built on the side panel here. And the same, for example, Sakana Fugu, another alternative from Japan that does a similar sort of thing. Didn't quite perform as well as Hermes' mixture of agents. But this is a similar sort of process where you can compare them side by side. You can build stuff. You can chat with it. And then you can get the workspace over here. So however you want to use these models, whichever system you use, you can build them all into a process like this. And then if we want to achieve better levels of intelligence from the models we already have access to, this is how you do it. And it's super simple and easy with a system like the H&OS. And all these models are swappable parts. You want to think of those as interchangeable. But really, it's a system that stays. Because, for example, these models over here are going to change in three to six months anyway or even next month or next week. But the actual engine you use to run all this, that's going to stay the same. Now, if you want to get my agent operating system with the council engine built in, the council engine lives inside the agent operating system I built. The same stack the boardroom runs every single day. So you get the full agent operating system, Hermes, Claude, OpenClaw, Wired Into One, Dashboard or Shared Memory. You get the mixture tab with your own model council plus Fusion and Sakana panels. You get weekly coaching calls, daily tutorials and a 30-day roadmap to set it all up. And token optimization tutorials, so it runs cheap as well. So if you need to run this on cheap or free models, you can as well. So if you look at the system as well, like if you're using, for example, a mixture of agents, the great thing about this is, you know, for everyone else, like the other 99% who's messing around chasing models, the problem with that is like they're always behind. They pick one model and hope it's the best one. They hit that model ceiling and stop there. They wait months for the next release. They beg for access to the gated frontier models. And, you know, the APIs, when they do get released for this stuff, is super expensive. With this system that I'm showing you, you know, you get the best models today. You mix several models into a virtual council. You beat every single model on its own. You don't need to wait for anything. You can just squeeze from what's already here. You go beyond the public frontier without any extra access needed. And you can swap everything in when a cheaper or better model lands, right? You can switch and change and chop this up however you want. Now, some people say, you know, the best model wins. I just need to pick the right one. But actually, the two best things on my whole leaderboard are systems. So if you look at Goldie Bench and you look at the top two performing models, Fusion and Hermes, mixture of agents, they're both panels. They're not individual models. They're systems. And that's what seems to give everyone the edge. It's not about the models you use. It's about the systems you create. Obviously, say, well, you know, stacking models, that's lab stuff only. Too complex for me. But actually, if you've just got a simple system like this, where you can type anything you want and it's just one click away, that's way easier than, for example, having Claude in one tab, ChatGPT in another, blah, blah, blah. Also, some people say, well, I'll wait for GPT 5.6. I'll wait for Fable 5. But none of those have a confirmed release date. We don't know if they're coming out, if they're coming out at all. It might be a case that Frontier Technology in the future doesn't come out to the public, that nobody gets access to anything that's the latest stuff anymore. We might get access to only models that are one year old, for example. That might be the future that we're living in. So you just have to bear that mind as well. And you might be saying as well, this might be technical to set up, but you can see so many tutorials and so many wins and testimonies. We've got 191 pages of wins and testimonies from people using our agent operating system from the AI Profit boardroom. So I know if I can do it and I'm non-technical, and all of these people can do it as well, then I know that we can all win and learn together. People are getting amazing stuff with this. So overall, what have you learned? You've learned that Hermes mixture of agents is really, really good. But you've also learned how to stop guessing. You've got proof. I've shown you the proof. And if you want to check it out for yourself, it's on Goldie Bench. I've shown you side-by-side comparisons versus Opus. And I've also shown you side-by-side comparisons versus Fusion as well. And we've also learned today that it's not about the model. I don't think you need to care about the model anymore. I think what you need to do is just get great systems. And bear in mind, most things you want to automate, like for example, we have a system where we can automate videos and it plugs in the avatar. It does the research, creates a script. It creates a voiceover and plugs it all together. This whole system doesn't need a frontier model. We don't need GPT 516 to run this. But it creates better outputs than 99% of people. The same with our SEO deployment systems. Like we can easily automate keyword research from our Google Search Console or from OpenSEO. We can generate the content, deploy it to our website. We don't need frontier models for that. So I think like most things that you're working on don't actually need the best models. They just need the best systems. And that's what you get inside the Agent OS. Now, if you want to get that from me, link in the comments description or just go to the theairprofitboardroom.com. And inside the community, you can ask questions. I personally answer all of the questions inside here every single day with a video tutorial. Inside the classroom, you can get access to all of my best trainings with new daily updates. We've got a beginner to expert course over here if you're new to this. And we also have a new daily update section here. So you can get the Agent OS system here with video tutorial, the last update date, as it filed to install it. And also we add new tutorials. If you want to learn more about how to use mixture of agents exactly with a step-by-step roadmap and all the terminal commands and everything else, you can get that right here. We've also got new tutorials on cloud web design and everything else that comes out. You can also jump on weekly coaching calls, meet people in your local area who are using Hermes and other agents like this. And that's all inside the AR Profit Boardroom. Thanks for watching.