← Back to search

NEW GLM 5.2 BEATS Claude?

AI News Today | Julian Goldie Podcast · 2026-06-16 · 9 min
relevance 48 1741 words Episode page ↗ Audio ↗
Show full episode description
GLM 52 vs Qwen 37 Max vs Claude Opus 48: Real-World Tests vs Benchmarks (No Second Chances) The episode compares GLM 52 (ZAI), Qwen 37 Max (Alibaba), and Claude Opus 48 (Anthropic) head-to-head on five one-shot tasks, arguing that benchmark rankings didn’t match real usability. In coding-focused tests like a voxel runner game, a liquid-in-a-bowl animation, a business landing page, and an arcade game, GLM 52 produced the most fun, polished, and feature-rich results, while Claude’s outputs were often basic and Qwen’s were sometimes buggy or incomplete; Claude clearly won the solar-system orbit map task. The script also notes Qwen’s strong reported benchmarks and faster replies, GLM’s slower responses in agents but strong CLI coding, and highlights limitations integrating Claude into agent workflows compared to Qwen/GLM in Hermes and the creator’s agent operating system.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Which AI model to actually pick for business when benchmarks contradict real-world performance.
Benefits
  • Real one-shot testing over unreliable vendor benchmarks
  • Swap models in and out of Hermes agents on coding plans
  • Run multiple coders in one dashboard by strength
  • Orchestrate teams of agents to build full projects
Use cases
  • GLM 5.2 won 'pretty much all' of the 5 one-shot coding tests, losing only the Galaxy Orbit map to Opus 4.8
  • Quen 3.7 Max reports 80.4% on SWE Bench Verified, strongest of the three on paper
  • Quen 3.7 reviewed all Obsidian notes and rankings, returned SEO keyword research GLM 5.2 couldn't match
  • Team of video agents on a Kanban board fully built a finished AI-generated video from one prompt
  • GLM 5.2 built full open-world and Skyrim-style RPG games in the agent workspace
KPIs / results
  • Quen 3.7 Max: 80.4% on SWE Bench Verified
  • GLM 5.2 won 4 of 5 one-shot coding tests
  • 3 models compared, same 5 tasks, one shot each
Tools / build
  • Hermes Agent OS multi-model dashboard
  • Hermes Kanban video-agent crew
  • GLM 5.2 voxel runner and RPG games
  • Memory Galaxy + Obsidian memory system
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00
Head To Head Setup
01:27
Coding Tests Results
  • GLM 5.2 wins voxel runner, liquid bowl, landing page
  • Opus 4.8 nails the inner-system orbit map
04:09
Arcade Game Showdown
  • Arcade game: GLM 5.2 most fun and playable; Quen ball bug
04:50
Benchmarks Versus Reality
  • Quen 3.7 Max strongest benchmark at 80.4% SWE Bench
  • Benchmarks don't translate to real-world results
06:01
Agents Workflow Tradeoffs
07:59
Final Recommendations
  • Run all coders in one dashboard; Claude orchestrates, GLM 5.2 codes
GLM 5.2 vs. Quentin 3.7 vs. Claude Opus 4.8. I put all three head-to-head, same five tasks, one shot each, no second chances, and the results flipped everything I thought I knew about picking an AI. So here's the part that gets me. The model that wins on paper, the best one with the best scores, came in last when I actually used it, and the one that looked the best, that actually shipped with no official scores at all. So if you've been sitting there trying to figure out which AI to actually use for your business, and you keep seeing these big benchmark charts that say, you know, something different every single time, well, we're going to look at what performs the best out of Quen GLM 5.2 and Claude Opus 4.8. Now, these are three of the top AR models right now. GLM 5.2 is from a Chinese company called Geopu ZAI. On this coding plan, then you've got Quen 3.7 Max from Alibaba, and Claude Opus 4.8 from Anthropic. Now, I ran all three through the same jobs, and we'll start this off. So the first job that we gave it here was a sort of voxel runner game. So we've got GLM 5.2 over here, Quen 3.7, and Opus 4.8. So this is the one from Quen 3.2, sorry, from GLM 5.2, as you can see here. It's a lot of fun, pretty interesting game, pretty cool, and a lot of fun to play. If we look at this one from Quen 3.7, this is Quen 3.7 Max, by the way, and you can see here that it's quite boring to play with, right? I mean, look at that. It's kind of buggy, but it does the job. It does the job. Then we have Claude Opus 4.8, and look how basic that is. That is not so much fun at all. So on the first coding test here, we can see clearly the GLM 5.2 is winning, and Quen 3.7 comes in second. Next up, we have the Inner System Orbit Map. And actually, if you look at all three of these, undeniably, Opus 4.8 wins this, right? It's doing the best job here, and it's done something amazing, as you can see. And we can change the speed, we can change the rotations, etc. This one is okay. This one is pretty bad. Now, you can zoom in, and it looks better, and that sort of thing. But on the outside, it doesn't look that great. I mean, it's kind of cool to play with, but I think out of all these, you know, Claude Opus 4.8, absolutely nailed it. Now we've got the Liquid in a Bowl test. So we have Quen 3.7 over here, GLM 5.2, and Opus 4.8. Now, if you look at these, I mean, it's kind of a boring test, but you can see here that the animation from GLM 5.2 is really nice. Like, we can change this, we can change the theme. It's pretty cool to play with. We have a look at, for example, the one from Quen 3.7 Max. Not quite as fun. It just kind of, you know, fades out very quickly. Then if we have a look at the one from Opus 4.8, look how boring that is compared to what GLM 5.2 created. This is way better, way more fun, way more interesting. And that's what we want, really. So on test 3 and 1, GLM 5.2 1, and then on test number 2, which is the Galaxy Orbit, you can see the Opus 4.8 1. Then we've got the landing page test. So this is useful if you're checking, for example, you know, if it's actually creating something useful for business. So creating a website. Now, let's have a look at this. This is Opus 4.8. Super basic, super boring. Not much to it at all. Not that interesting. If we have a look at Quen 3.7, it's okay. I mean, this is kind of weird because there's nothing here, right? It's kind of just like an empty canvas. But the rest of it was okay. Then if we have a look at GLM 5.2, if we scroll down, it's got some nice animations. There's a lot more to the page. Nice, nicely, but cleanly as well set up. And I like even like the animations on the page. They look super nice. And you see how it's actually filled in the canvas. Whereas, for example, Quen 3.7 didn't do anything. And the one from Opus 4.8, super boring. So that's the difference. This one, the landing page test as well. Then we have the arcade game. So this is pretty cool. Pretty fun from Quen 3.7. But the only issue is you see how the ball doesn't actually bounce off the walls. Like it just disappears completely. If we have a look at Opus 4.8, this has built something better that is more playable and more useful. And then if we have a look at GLM 5.2 here, look how cool this is. This is way more fun and interesting. And so GLM 5.2 won on pretty much all of the tests there, apart from one, which is mind blowing in itself. And so in terms of the actual tests that I've run here, I would say the GLM 5.2, which is a new model from ZAI, is actually beating Opus 4.8 and definitely beating Quen 3.7 max. Now, if we actually have a look at the benchmarks here, Quen 3.7 max is the strongest of all three. You know, Alibaba reports 80.4% on SW Bench Verified. We don't have the benchmarks. We only have GLM 5.1 benchmarks. So it's not really that useful. But if we're comparing side by side, Quen versus Claude, well, it's actually beating Opus 4.7 on a genetic coding. However, one thing to note here is the Opus 4.8 was so new when Quen 3.7 max came out that they didn't include Opus 4.8 on their benchmarks as well. So pretty interesting. For me personally, do I really pay much attention to benchmarks? No, I just test stuff in reality because, you know, just how can you believe in the tests run by the company that owns it? And also do benchmarks from my experience always translate into what you see in reality? No. So there's a big difference here. But the main thing I would say is like, you know, China's GLM 5.2 so far has been really, really good. And if we have a look inside the agent operating system too, we can check the workspace and see what we've built out here. And we created some awesome stuff, as you can see here too. I mean, this is like a full open world game. This is another one kind of like a Skyrim style RPG. And this was a fun one too, as you can see here, right? And this was all created using GLM 5.2. So it's a pretty powerful model. And the other thing I would say here is like, you can actually plug Quen 3.7 max directly into Hermes agent. And you can do that with GLM 5.2 as well on the coding plan. But you can't actually do that with Claude. So Claude doesn't allow your subscription to plug into your agents. Whereas, for example, we can create separate profiles for GLM 5.2, for Quen 3.7 max, and for using Kimi K 2.7, which is another model. So when you're on these coding plans, you can easily swap them in and out of your agents, which I think is super useful in itself as well. The other thing I would say here is like, when I've been using Hermes with GLM 5.2, it seems super slow to reply. Quen 3.7 max does seem a lot faster. And also, if you look at the quality responses here. So we asked GLM 5.2 to take a look at our Obsidian memory. It wasn't as useful as when we actually got the answer back from Quen 3.7. Quen 3.7 actually reviewed all of our notes, looked at what we were ranking for, and then gave us some great SEO keyword research. When you check GLM 5.2, it is super brief and not as useful. So it's pretty interesting to see like, okay, if I'm using AI agents like Hermes, I might not use GLM 5.2. But if I'm coding directly inside the CLI, for sure I'm going to use GLM 5.2 or Kimi K 2.7. These are great models that have created some awesome stuff. So it depends how you're going to use it as well. And another thing to note here is we actually got a team of video agents to work together. And again, you can't really do this with Claude unless you go to Claude directly. But with Hermes, you can easily get them to work together and then create some awesome stuff. So if we go to the Kanban board here, we've got a team of agents to work on some videos and fully create it from scratch. And then if we open this up, this is the finished video, as you can see. And this is fully AI generated. I just gave it the prompt. And then we had the judge tell them to iterate and keep going until it was finally done. But yeah, again, pretty powerful model for having your AI agents work together. So thanks for watching. And we have multiple models. So we have all of these working together, all three coders in one place. And I would recommend that for you because then you get the best out of all of these models depending on their strengths. So if we're actually coding on the backend, like I tend to find the Claude is better for just pulling them all together and orchestrating them. If I'm trying to create something really awesome and cool in terms of like actual coding project, like you saw with the games that we created, then I'll probably go to GLM 5.2. But the main thing is you want all of them in one place. And so if you want one dashboard, we've actually got the H&O operating system with one dashboard, one workspace. You can preview everything live. You can see it built. You can see them all working together with the memory system, the memory galaxy. We've got Obsidian. And that's all inside the AI Profit Boardroom. Link in the comments description or go to theairprofitableboardroom.com. Inside the community, you can ask questions. Inside the classroom, you get access to all my best lessons. Inside the calendar, you can drop on weekly coaching calls. Hope to see you inside. Cheers for watching. Bye-bye.