← Back to search

E087_Hermes Agent. How different is it?

BIMPRAXIS · 2026-06-09 · 17 min
relevance 75 2874 words Episode page ↗ Audio ↗
Show full episode description
Descripción del Episodio Exploramos el proyecto Hermes Agent, una inteligencia artificial que navega por Internet con autonomía, aprendiendo a “habitar” la red. Analizamos su capacidad para resolver problemas y adaptarse a nuevas situaciones, y cómo combina con Browser Harness para lograr una interacción más efectiva. También discutimos sus implicaciones en la automatización de tareas y la posible influencia en el futuro del trabajo humano.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How Hermes Agent plus Browser Harness moves from chatbot to a self-healing agent that inhabits the internet.
Benefits
  • Agent clicks, errors, self-corrects, then writes its own manual
  • Browser Harness gives a safe, self-healing execution environment
  • Semantic page analysis avoids fragile fixed coordinates
  • Runs cheaply on a small VPS
  • Orchestrate agents without deep coding knowledge
Use cases
  • Hermes reached 100,000 GitHub stars at absurd speed; 5 main versions in 20 days, 740+ pull requests (~37 changes/day)
  • User Pli used the Obliterus ability to jailbreak the Gema 4 model with only eight prompts
  • User Adam created a full Mandarin video — script, HTML, text-to-speech voiceover, 1080p vertical MP4
  • Hacker News challenge: extracted top 15 articles to clean JSON, then injected a new skill for next time
  • YouTube thumbnail-grid task self-healed a Chrome failure via remote daemon, wrote a 147-line skill
KPIs / results
  • 100,000 GitHub stars; 5 versions in 20 days; 740+ PRs (~37/day)
  • Browser Harness has under 2,000 stars
  • Jailbroke Gema 4 with only 8 prompts
  • Runs on Hostinger VPS via Open Router + Opus 4.7 for $5-10/month
Tools / build
  • Hermes Agent
  • Browser Harness self-healing browser automation
  • Obliterus jailbreak ability
  • Mandarin video pipeline (HTML + text-to-speech API + 1080p render)
  • 147-line YouTube skill
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
Hello, this is BIMPRAXIS, the podcast where BIM meets artificial intelligence. We explore science, technology and the future from the perspective of architecture, engineering and construction. Let's start! Hello, welcome, welcome to a new episode of BIMPRAXIS. Today we bring you the meteoric rise of Hermes Agent and how artificial intelligence is learning to navigate the internet with an autonomy that is frankly a little scary. Hello, and yes, well, more than browsing, I would say that you are learning to inhabit the Internet, literally. Completely. Completely. Let's see, the mission of our dive today is to break down a super detailed video analysis that creator David Ondrej published. Exact. A video focused on this project, Hermes Agent and another companion tool called Browser Harness. That is. And we want to understand why the hell this project is breaking absolutely all growth records in the history of GitHub. That is, we are no longer talking about a static chat that answers trivia questions. No, no, not at all. We are talking about an artificial intelligence that moves the mouse, clicks on the screen, makes a mistake, solves it on its own and, pay attention to this, then writes the instruction manual so as not to fail again. That's the key jump. To give us an idea of ​​the impact, Hermes Agent has reached 100,000 stars on GitHub at an absurd speed. 100,000 stars that is said soon. In other words, there are legendary projects that have taken years to achieve that figure. Yes, years. And here the team has an update rate that is insane. They have released 5 main versions in just 20 days. My mother. And they have integrated more than 740 change requests, the famous pull requests. That is, we are talking about about 37 changes a day. That is. You basically have hundreds of independent programmers from all over the world proposing patches daily. And if you look at the Google Trends data, it's fascinating. Let's see, tell me. Well, rivals who dominated until recently, like Ope Neclo, are now plummeting. Meanwhile, interest in Hermes is rising vertically. It doesn't stop. It's incredible. But, let's see, the real revolution, according to David's analysis, comes when combining Hermes with that other repository that you mentioned, Browser Harness. Of course, Browser Harness is essential here. It's brand new. It has less than 2,000 stars, but it is the missing piece. If Hermes is the brain, Browser Harness is the hands. And this is where I like to use an analogy to visualize it. Until now, using an advanced I was, well, I don't know, like having a genius office worker, a super gifted guy, but tied to a chair and without arms. Poor office worker. I love the image. Yes, yes. But it was like that. I knew the answer to everything, from quantum physics to poetry, but I couldn't type a single word or send a simple email. Well, Browser Harness just untied his hands and planted a mouse in front of him. As is. And it doesn't just hold your hands, it provides a safe, self-healing execution environment. That's important, yes. Of course, because historically web automation has been super fragile. You wrote a script to, I don't know, download invoices. And if the web designer moved the download button 10 pixels to the right... Everything would break. Everything exactly messed up. It made an error and a human had to go and fix it. But Browser Harness visually and semantically analyzes the page. It doesn't care about fixed coordinates. What a blast! And the creators are so sure of this that they have launched a public challenge. Ugh, yes, the challenge is brutal. They offer a brand new Mac Mini to the first person who finds a task in the browser that the system is not able to complete. You have to have blind confidence in your code to post a computer like this. Well, they know that reliability has taken a quantum leap. And knowing that you now have these hands and that reliability, the logical question is… Well, what are people really doing with this agent on the Internet? Well look, we go from theory to practice. And there are cases that cloth. For example, in the field of cybersecurity. Yes, cybersecurity. Let's see. A user named Pli in the community used Hermes and a specific ability called Obliterus. The objective, to jailbreak the Gema 4 model. Wait, wait. A jailbreak? That is, bypassing its security, ethical and operational barriers. Exactly that. And the shocking thing is not that he achieved it, but that Hermes discovered how to do it on his own, receiving only eight instructions, eight prompts from a human operator. Let's see, let's stop for a second. Let an artificial intelligence figure out on its own how to break the security of another artificial intelligence and with only eight little human phrases. It sounds strong, yes. Well, doesn't this sound a bit like a science fiction movie that ends badly? Yes, yes. The alarm is totally natural. But it must be analyzed coldly, objectively. Okay, give me some context because it's a bit dizzying. Let's see, it's not that the machine has become aware and revealed itself. It is pure logical optimization. The agent tests a text entry, Gem 4 rejects it for security, and Hermes analyzes why he rejected it. Oh, sure. Read the error and adjust the shot. That is. Adjust the angle of attack and try again. It is an iterative trial and error method at a beastly speed without a human holding your hand. Demonstrates amazing problem-solving skills. Okay, so looking at it this way, for a company security audit, having an agent checking every virtual lock without getting tired is an analyst's dream. Completely. But hey, it's not all about breaking rules. They also use it to create super complex things from scratch. Yes. The case of the content creator. The user Adam, I seem to remember. Exactly, Adam. He used it to create a full video in Mandarin and I mean full. From the script to the final file. Exact. Absolutely. The AI ​​structured the narrative. It wrote an HTML file to organize the visual part and then connected itself to a text-to-speech API to generate the Chinese voiceover with exact timings. How crazy. In other words, he acted as scriptwriter, translator and announcer. And as an editor. Because then it orchestrated a rendering engine and delivered a vertical video at 1080p. In other words, a perfect MP4, ready to publish. That is, he coordinated several independent programming languages ​​and tools for a single project. That's not easy. Not at all. And in terms of visual creativity it is not far behind either. In a hackathon they used it to generate animations and gifs of real sculptures. And the interesting thing here, according to the source, is that they overcame that stigma of AI Slop. Yes, yes. That AI-generated garbage that looks plasticky, shiny, and full of anatomical errors. Ugh, yes. It's horrible sometimes. Well, agent Iteroy revised his own work so many times that he achieved a high-value aesthetic finish. A completely personalized brand finish, without that tacky look. Clear. In the end the perseverance of the machine replaces the patience of the human. But hey, the ability to do tasks is very good. Yes. However, what really brings it closer to what the video calls almost AGI, that artificial general intelligence, is what happens when everything goes wrong. That is the key concept. Self-healing. Let's illustrate it with the Hacker News challenge that appears in the source. The premise was simple. Extract the top 15 articles from the web. Yes. Take the title, author, punctuation and comments. Exact. And put it all in a very clean JSON file. Well, it turns out that people found a prerequisite skill in their system and started browsing. But he quickly ran into traps, what programmers call gotchas. The famous gotchas of web programming. A headache. Completely. You came across relative URLs, which are incomplete link fragments. And also with anchors that pointed to zero comments. And any conventional program would have tried to visit that link fragment, would have given a 404 error and the entire script would have crashed right there. As is. Well, the agent did not block. He analyzed the structure, deduced that he had to add the main domain to the URLs and fixed them. My mother. Finished the JSON file perfectly. But the worst thing is not that. That's what he did next. Yes, about modifying your own code. Modified your code. He injected a new skill into his system so that the next time he logged into Hacker News, it would already be taken care of out of the box. And he even left a note warning that there was a section called Barra Ask that could be useful in the future. In other words, pure and simple proactive memory. Abstract the problem and keep the knowledge. It's spectacular. But wait, the YouTube challenge takes it to another level of tension. Ah yes, the one with the thumbnail grid. That same one. He was asked to create an image with the 12 most recent thumbnails from David's channel. Okay. And here comes the drama. Browser Harness attempted to open a Chrome browser locally, on that same machine. And it failed, right? resoundingly. Critical connection error. The bridge between the brain and the hands was broken. And you and I know that in traditional automation, this is where the red letter comes up on the screen and the human has to put down the coffee and go restart the server. Of course, an environmental failure is usually terminal. How did the AI ​​react? Well, he applied diagnostic logic. It read the error, saw that it was a local execution issue, and self-healed in real time. You ran a command to start a remote daemon. That is, a remote daemon, to control an invisible browser in the background. He solved the infrastructure problem on the fly. Exact. And to top it off, instead of entering the website and scrolling down like a human would do... Which is super slow and depends on the loading speed. Clear. Well, he went straight to the guts of the HTML code and found a massive block of JSON data called ITInitialData. Ah, how clever. YouTube sends that raw data all at once upon upload. That is. It pulled the images out of there in milliseconds, bypassing the entire interface. And when he was done, he wrote a new skill with 147 lines of code. The operations manual we were talking about. With detailed comments on cookie warnings and short videos. Let's think about our own companies for a second. Let's see. How many human workers are capable of encountering a completely new technical problem, solving it on the fly, and then, on their own initiative, sitting down to write a flawless 147-line manual so that the next employee doesn't make that same mistake? Ugh, I would say very few, if not none. This surpasses the average worker by a huge margin. Completely agree. And at this point, seeing everything it is capable of doing, anyone would think that you need a million-dollar budget or gigantic data centers. Sure, NASA servers, at least. Well here is the most fascinating paradox of this entire ecosystem. It's ridiculously accessible. David ran it from a simple virtual private server. A very cheap VPS from Hostinger. That is, the typical server that you rent for a couple of euros to host a small website. That same Plan Taren 2, on 24 hours a day, and used Open Router to connect the system with the Opus 4.7 model using API keys. And how much does that cost in consumption of artificial intelligence? Well, between 5 and 10 dollars a month, literally. It's just that it costs you more to pay for Netflix than to have a system that self-diagnoses servers and edits videos in Mandarin. As is. And if cost is not a barrier, neither is technical complexity. That's what I was going to ask you. Configuring all that on a server has to have its own thing, right? Well, look at David's anecdote in the video. I needed to install a Python package called V on Ubuntu. Okay. To do this I had to log in via SSH, which is basically a black screen with white letters where you just enter commands. No mouse. And he himself admits that he has no idea about Linux. And what did he do? Because there you get stuck quickly. Well, he turned around, opened a chat with another AI, in this case Cloud, explained what he wanted to do and asked him to dictate the commands step by step. It took 20 seconds. Look at the paradox. In other words, we are configuring one of the most sophisticated artificial intelligences on the planet. Something super advanced. And when the human operator gets stuck in front of a black terminal screen... His solution is to ask another artificial intelligence for help. It's very good. He asks you to tell him which keys to press. The technical barrier has completely disappeared. And this is vital. It no longer matters that you don't know how to install a package in Ubuntu or that you don't know the syntax of a language by heart. Clear. What counts now is the strategy. Exact. What matters is your ability to orchestrate these agents and have the vision to know what tasks you can delegate to them. David claims that anyone, regardless of age or whether they know anything about technology, can learn to build software like this in just three weeks. Three weeks? It is an absolute change of mentality. We went from being those who code code line by line to being conductors of an orchestra. You don't need to know how to play the violin perfectly, but you know what the symphony should sound like. Yes, yes. And you also have musicians who, if they go out of tune, tune the instrument themselves. As is. Although all this autonomy, this ability to solve problems and us, inevitably leads us to deep reflection, don't you think? Yes, and it is a quite provocative reflection to take home. Let's see, tell me. Why am I convinced that it is due to a labor issue? Completely. If AI agents are now able to encounter unprecedented problems on the Internet, diagnose them, fix them in seconds, and automatically write the perfect standard operating procedure so that no one ever fails again, what about entry-level human jobs? Phew, that's a good question. The thing is that, historically, junior professionals, those who have just started in a company, learn precisely by stumbling upon those small technical errors. Clear. You spend the first few months fixing broken databases, reading documentation because the server has gone down, you get hard at work. Exact. You are creating that professional callus. But if artificial intelligence is the only one that stumbles now and seals the path by paving it forever... There are no more potholes for humans. That is. How will humans acquire that fundamental experience in the future? Where will the senior profiles come from in 10 years if the juniors have no problems to solve? Wow. It is a tremendous paradox of progress. By automating overcoming obstacles, we seem to be automating and almost eliminating our main source of learning. It leaves you thinking, of course. Definitely. It gives a lot to think about. Well. Before saying goodbye until the next program, we inform you that the voices you hear have been generated by Notebook LM's AI and that Julio Pablo Vázquez, a human who sends you greetings, is directing the podcast. In case of error, it is probably human error. We listen to each other. And that's it for today's episode. Thank you very much for your attention. This is BIMPRAXIS. We'll hear from you in the next episode.