← Back to search

E088_How to master Hermes Agent

BIMPRAXIS · 2026-06-11 · 17 min
relevance 54 2848 words Episode page ↗ Audio ↗
Show full episode description
Episodio de BIMPRAXIS: La Guía Definitiva para Transformar la Inteligencia Artificial en un Sistema Autónomo y Personalizado En este episodio de BIMPRAXIS, exploramos la creación de un sistema autónomo y personalizado que gestiona tu vida entera, desde la gestión de tu calendario hasta el análisis de tus datos biométricos. Se presenta Hermes, un sistema que utiliza la arquitectura de memoria a largo plazo para actuar como tu propio archivista, permitiéndote interactuar con él de manera personalizada a través de Telegram. También se discute la importancia de elegir el proveedor adecuado y la utilización de plataformas de enrutamiento dinámico para optimizar costes. Además, se aborda la cuestión de la privacidad existencial y los límites psicológicos de ceder la auditoría de nuestra vida a las matemáticas de un algoritmo.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How to turn AI into an autonomous, persistent-memory personal system that manages your whole life.
Benefits
  • Persistent long-term memory with write permissions, unlike static RAG
  • Dynamic routing sends cheap tasks to cheap models, hard ones to expensive
  • 24/7 operation on a cheap cloud VPS
  • Proactive heartbeats and multi-agent delegation
  • Private, self-hosted; biometric data never leaves your server
Use cases
  • One user ran Cloud Premium API and racked up $64 in a single week
  • $18/month DigitalOcean droplet (2 CPUs, 2GB RAM) keeps the agent always on
  • OpenRouter Elephant: $10 balance unlocks 1000 free daily requests
  • Heartbeat wrote a Python script analyzing Apple Health, found avg 7.5h sleep with a 5h day
  • Threads social audit via getcookies.ext found ERMS/hardware metrics skyrocketed
KPIs / results
  • $64 spent in one week on Cloud Premium API
  • $18/month VPS with 2 CPUs and 2GB RAM
  • $10 balance unlocks 1000 free daily requests
  • Average 7.5 hours sleep, critical 5-hour day detected
Tools / build
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
Hello, this is BIMPRAXIS, the podcast where BIM meets artificial intelligence. We explore science, technology and the future from the perspective of architecture, engineering and construction. Let's start! Hello, welcome, welcome to a new episode of BIMPRAXIS. Today we bring you the definitive guide to transform artificial intelligence into an autonomous and personalized system that manages your entire life. And we really are not exaggerating with the whole life thing, that is, it is a brutal paradigm shift. Completely. Let's see, to put ourselves in the situation, let's imagine this. Hours before your morning alarm goes off, there is a digital system on a server that has already swallowed all your biometric data from the night. So? In the background? Quietly? Exact. And in a completely autonomous way, it has also evaluated how some social media posts from the previous day worked. He has seen that there is little commitment or engagement and has restructured your entire daily calendar. In other words, he has moved heavy tasks to times when he predicts that you will have better physical performance. That is. And all this happens silently, without you touching a single button. All documented in text files. When you close the tab, the system reboots. Zero memory. It is, by design, a stateless environment. And at an operational level that is a tremendous bottleneck. I always compare it to having a colleague working with you who is brilliant, the best in his sector, but who has severe amnesia. Yes, yes, yes. Like in the movie Memento. And this is where today's sources introduce us to Hermes. Hermes, exactly. Which is based on a concept that André Carpaty named WikiLLM. The WikiLLM. I love the name. Let's see, Carpaty hits the nail on the head. He says that the next leap is not to put more parameters into the models, but to give them a long-term memory architecture. Let it be persistent. In other words, it is not reactive, but rather the system acts as your own archivist. That is. Before the model spits out the first word of response, a system goes and reads a giant block of your own history file. It loads it into the context window and then, and only then, does it reason. But let's see, someone had to say, hey, this is like traditional RAG systems. You upload a giant PDF to the AI ​​and that's it. What is the difference? Ugh, the difference is abysmal on a technical level. In a RAG the information is static. If you modify the PDF, the AI ​​does not learn anything new. But Echarme S has write permissions. Oysters, of course. You can modify your own memory. Exactly. If people see that you always correct a certain tone in your emails, they take action, open their own rules file autonomously, rewrite it, and save that new rule for the future. Pure ammunition. My mother. But of course, there is an elephant in the room here. If every time you interact, the AI ​​has to read the entire encyclopedia of your life, that at the token level has to cost a fortune, a fortune. In fact, reports document that someone tried to do this with the Cloud Premium API and pocketed $64 in a single week. 64 bucks in a week. That, for an individual user, is unsustainable. Unviable, totally. That is why the first rule to avoid going bankrupt is to choose the supplier well. And here the sources highly recommend OpenAI Codex in its version 5.4. Why's that? It's smarter than the new Anthropo or Gemini models. What's up? It is not because of reasoning ability. It's for pure survival. Corporate platforms like Anthropo have very strict anti-cheat systems. If they see a constant, scheduled and automatic flow of requests... They think you are a malicious bot and block your account. Exact. They ban you without warning. However, Codex 5.4 swallows this non-stop volume without setting off any alarm bells. It is super reliable for the base. Ok, but I imagine you can't use a single model for everything if you want to optimize costs. There you have hit it. The secret is in dynamic routing platforms, such as OpenRouter. It's pure magic. He's like an orchestra conductor, right? Who decides who to send each task to. As is. Hermes evaluates how difficult the task you have asked him is. If it's just, hey, categorize this bill or bold this for me, the router sends it to a super cheap and fast model. Like Cloud Sonnet, for example. That is. Which has a ridiculous marginal cost. But if you suddenly ask him to audit an asynchronous code architecture or cross-check complex financial data, then yes, he will bear the expense. And he sends it to the Opus 47 model. He pays for the expensive model only when heavy artillery is needed. And be careful that in the documents they talked about a trick with OpenRouter, the Elephant model. Yes, yes, the free model. Look, you put a $10 balance in your account, a minimum deposit, and the platform unlocks a thousand free daily requests for that specific model. Crazy savings! Okay, we have the software and the invoice under control. But this has to be on 24 hours. If I have it on my laptop and everything falls out under the lid. The ecosystem collapses, yes. You can try to use or call locally with lightweight models like Qn 3.5, which by the way works well, but others like Gamma 3B failed because they do not support tool calls. Yes, but you still depend on having the computer on all day. Clear. The real solution is a VPS, a virtual private server in the cloud. And this is where people usually run. Total. As soon as you tell them command line, Linux or SSH connections, half the audience disconnects. It gives vertigo. And it's normal, right? But the documents show a brutal tool called OpenRouter Spawn. This abstracts you from all that technical hassle. I mean, you don't have to touch code in the terminal? Nothing. Spawn connects to a provider like DigitalOcean, installs the image, sets the variables and sets up the services for you. Everything almost with one click. For about 18 dollars a month you have a droplet with two CPUs and two gigabytes of RAM. More than enough to keep your brain awake all the time. Okay, changing gears. We already have the machine working in the cloud. But how do we interact with it? Why don't you go to the server every time? No joke. Tests confirm that Telegram is the best interface, without a doubt. You use Botfather, create a bot with your identifier and that's it. You have it on your mobile. And they were talking about creating different thematic chats within Telegram, right? Yes. Super important. You have a chat for social networks, another for programming, another for chatting. This way you avoid cross contamination in the context of the model. Clear. You don't mix a metrics analysis with the dinner recipe. But what fascinates me is how this is given personality through Markdown files. The famous .md. The key to all logical configuration is in a local web panel. And especially in those three files. User, Souls and Agents. Let's see, this is literally like making a character sheet in a role-playing game. Completely. Let's see how you approach it. Well look, the user.md file is the background, the lore of your character. That you live in Madrid, that you have a dog, your marital status. It is the passive context. Then, souls.md is the charisma, the tone, the empathy. The soul of the people. Exact. And Agents.mdd is the combat skills. The super strict rules of syntax or code that the AI ​​can never break. And notice that this separation, which sounds like a role-playing joke, at the level of neural networks is a beastly shield against hallucinations. Because? If in the end the AI ​​reads the three files the same. Yes. But by being in separate blocks, you force the model's attention mechanism to weight them independently. If you put in a traditional kilometer prompt, where you mix a fun tone with strict code rules... It gets messy. It gets messy, yes. The statistical weights get contaminated and the AI ​​gives you code variables with funny names or skips tabs because it is in creative mode. Clear. By separating the files, you can tell souls.md to act hype, super motivated or even like a pirate. And he will say good morning super euphoric. But when you generate code based on agents.md it will be flawless. Pure syntax without creativity affecting it in the slightest. Okay. I think it's brilliant. But now comes the most science fiction part of all this for me. We've talked about asking him for things. But is this system proactive? Take the leap, yes. Go from waiting for orders to starting the conversation. And it does so with what architects call heartbeats. Heartbeats. This comes from the philosophy of OpenCloud and the famous surprise me command, right? Exactly. They are basically cronjobs, tasks scheduled at the server level. You tell him, hey, wake up every day at 8 in the morning and see if there's anything interesting. And here the case of health is amazing. Please review that case because when I read it I was amazed. Look. They hook up an app to dump the closed Apple Health data to the server's local directory. The AI ​​wakes up, sees that raw sleep data, and says, okay, this is unreadable to a human. And instead of just summarizing, what does it do? Well, she writes a Python script herself. Create a tool to clean data noise, calculate averages, and isolate bad days. And you discover that the user slept an average of 7 and a half hours, but there was a critical day of 5 hours. Oysters! But the strong thing is that it does not delete the script. He realizes it's useful, saves it as a permanent skill, and the next morning... The next morning, the heartbeat skips again. But AI no longer spends time and tokens thinking about how to process. Run the saved script and it will send you the report via Telegram directly. He tells you, hey, you slept badly today. Be careful with performance. Let's see, I have to play devil's advocate here. A system that monitors your sleep and sends you unsolicited medical notifications. It sounds a bit like a Silicon Valley dystopia. It sounds totally intrusive. It's true, especially if we think about today's commercial apps, which only want your attention to sell you things. Of course, the attention economy. But here lies the brutal difference. This ecosystem is hosted by you. You control the cronjobs. Biometric data never leaves your private server. They are not selling you anything. They only transform passive data into actionable intelligence for you. Okay. Seen this way, being local and private things change. And they were talking about another use case, bypassing restrictions on social networks, right? With threads. Wow. That's a spectacular technical example of how they bypass corporate walls. You already know that the official meta APIs are super restrictive. Yes, they shut you down quickly if you try to automate things. Well, ERMS uses an extension called getcookies.ext. Basically the encrypted session credentials from your own browser are exported to the server. In other words, the AI ​​pretends to be you entering from your computer. Exact. It bypasses authentication barriers as if it were a legitimate human browsing. He got into dozens of threads, took out the text, the replies, the interactions and drew conclusions. And useful conclusions, because he discovered that when they talked about the ERMS agent itself or productive hardware, the metrics skyrocketed. He did a complete social media audit for free, without anyone lifting a finger. It's very strong. But of course, for all this to converge, we arrive at the concept of the definitive CIO that the sources mentioned. The orchestra director. Because of course, ERMS cannot do everything alone without becoming saturated. You need to delegate. It's like connecting the Google Cloud console, so it can read your email and calendar. The AI ​​sees that on Thursday you have an in-person medical appointment from 5 to 6, reads an email from an urgent client and reorganizes your day. And that's where multi-agent architecture comes in. The thing is that we are no longer talking about a single, monolithic brain. It's a company. Literal. ERMS is the CIO. Answer the phone on Telegram, make quick decisions. But if you say, hey, do some in-depth research on the real estate market in Valencia. ERMS says, okay, this is going to take me hours. Well, he delegates it to a subordinate, to OpenCloth. You give it the instructions and OpenCloth spends 5 hours browsing in the background, while ERMS is still free to answer you on Telegram. And when the subordinate finishes, he passes the clean report to the boss. But, let's see, technically, so that these two agents don't step on their memory, how do they do it? That's the magic of Markdown files and using a NAS server. All agents read and write to the same physical folder on the server. Is it the same source of truth for everyone? That is. By using Markdown, which is super light, plain text, both machines and us can read it without problems. There are no closed proprietary formats. And to visualize all this swarm of data without going crazy, do you use Obsidian? Obsidian is the icing on the cake. You connect Obsidian to that NAS folder and it renders all those thousands of text files in a visual map, with connected nodes. You literally see your second digital brain beating live. Not two that update themselves while you talk to the bot. It is a technical marvel. It completely breaks the barrier between machine and human mind. But, of course, reaching this level of invisible monitoring poses a dilemma, quite profound. I already tell you. Because in the end you are setting up an infrastructure that swallows your emails, knows how you sleep, reads your finances, understands how you speak over the months. It is no longer an assistant, it is an absolute mirror of your life. A perfect statistical mirror. The machine begins to see correlations that completely escape you. The thing is, imagine, if this persistent network catches hidden patterns between your level of physical fatigue and the bad decisions you make at work or what you buy. The question is inevitable, isn't it? You will get to a point where this ecosystem understands your biases, your bad habits, and your productivity slumps much better than you know yourself. I think so. The technology is already here. The APIs allow it, the code works. The limit is no longer technical. It's purely psychological. Are we prepared to hand over the audit of our lives to the mathematics of an algorithm? Phew. Building a silent observer who ends up knowing your flaws better than your own brain, which is always trying to deceive and justify itself, is fascinating on an engineering level. But it raises hackles at the level of existential privacy. Of course. We will leave this reflection floating in the air until the next in-depth analysis. Before saying goodbye until the next program, we inform you that the voices you hear have been generalized by Notebook LM's AI and that directing the podcast is Julio Pablo Vazquez, a human who sends you greetings. In case of error, it is probably human error. We hear each other! And that's it for today's episode. Thank you very much for your attention. This is BIM Praxis. We'll hear from you in the next episode. Thank you! Thank you! Thank you!