🌐 This transcript was automatically translated to English from the original.
Magical.fm is a chat podcast hosted by Michiruda, a product manager from the Kansai region, and nobamune, a cool software engineer from the Kansai region. Please, please. Today's topic is: Yes, what is a Hermes Agent? What is it? The timeline of how it suddenly appeared. I wonder when it was, about a month ago. Did it come out? I wonder if it was conceived, or rather, the development itself started about a year ago. Hey, I think I did it. Today's recording date is April 20th, and I suddenly appeared on the timeline this Saturday. Hermes: Ah, at first glance, the spelling is similar to harness, so I thought you were talking about harness, but it turns out it's something else. I usually collect information on that Twitter using that other account with 200 followers and 0 followers, and it's one that no Japanese person follows. Hermes Agent: Yeah, I think it's like it came to Japan late, but I feel like my dad is going to read all the documents this Saturday and Sunday. Lately, I haven't been able to input it because I'm preparing for the technical bookstore, and I haven't been able to input it, so it's great. And for some reason, I wanted to try out the performance of the new model, and that's included. I thought I'd try using Hermes Agent. Basically, I'm currently using open-claws in that house, so I'm not sure how it compares to that. Hermes Agent is also spelled Hermes. It's Hermes. It's the same spelling as that high-end horse Hermes. They come from the same god, and I guess I used to call Hermes Hermes at first, but when I listen to YouTube videos from other countries... It sounds like it's pronounced Hermes. It's probably somewhere between E and H. Well, here it's called Hermes, but if you want to read Hermes, you should read it. Hermes is said to be the guardian deity of traveling merchants and the like. Yeah, yeah, why is this being talked about? Why is it being talked about? Well, it's almost like an open-claw thing, but it's a bit sophisticated, well... When did Open Claw come out? 10 Last year? Around November Isn't it around October? I feel like it was recorded somewhere around the time of Bunburi. It's surprising that it's been about half a year since Open Claw. That's right, so I thought it had only been about three months. That's right. What exactly is Open Claw? Well, Open Claw is an AI agent, but... It's something that's set up like this on everyone's local hand, or PC, and you can use all kinds of tools, and you can even connect to things like Google, so even if you're an AI agent, the more you give it the services of tools that you usually use, the more convenient it becomes.Well, there's a risk to that, but that's why it became so popular, and you can use it with the messenger tools that people usually use, like Slack, Telegram, Discord, etc. I think there are probably a lot of people who have already set it up once, but it's quite annoying to keep using it, and there are quite a few that suddenly break down.That's right.I've been using it every day for the past four months.I imagine there are quite a few people who have been using it every day for the past four months.Did you use it every day? It's broken, or it suddenly stops responding, or something like that. Well, the system runs inside a server, so if another component other than the open claw breaks, there's a possibility that it won't be able to communicate. I see. Even so, it was so useful that you decided to follow it and use it properly. Well, why do I basically no longer know what I am? When I talk to LLM, it's via open claw. So in my personal life, I don't use Chat GP or Claude for anything other than the CLI codex or Claude code. Yeah, I only use it for programming. Hmm, I guess that's what I expected. Even if I listen to it, it won't become something for me in the future. Because it won't, uh, because it will be limited to that session. Oh yeah, and also, wow. When you want to do this as is, if you use Open Claw, you can do anything with it.You can just code it like that, but if it was Codex, Chat GP, or Claw, there would be some limitations, right?However, since Open Claw has already been given a computer, there is a feeling of security that it can do anything.So, I started playing with Hermes Agent, wondering why it was becoming such a hot topic. Well, in terms of normal setup, it's no different from OpenClaw.It's the same thing, with the agent running normally and touching it from Telegram, etc.But the biggest difference in terms of ideology is that it's designed to have a nice feel even from the moment it starts up.Well, one thing is, that's that memory, that memory.I remember that OpenClaw operates in a very simple way.It has files like memory MD. It's like recording what you've done in a text file called End, but this Hermes also has a memory MD, and you write in the same way, but there's a limit on how many lines you can't write, and if you exceed that limit when you add to it, you can summarize it as you like and add the new one.The biggest difference is that all the conversations you've had so far are stored in the database at hand. If you put it in there and say, "We had this kind of conversation," it will search the database and bring it up.The summary file and the details will be saved separately, and you can look at it when the time comes.That's right.The search will search the database, and it'll probably come up with something that caught a lot of questions, but it doesn't return it to the main session, it just passes it to the LLM for conversation summaries, and then it comes back here. He's certainly very attentive in taking past history without overpowering this text. I haven't done anything like that, Oppa. I wonder if it's something like saving everything in a database and going to look at it. Is that another one? It's like something I made myself. I created it with Open Claw, and it seemed like it was fun. Did you do that? Huh? Super Memory? Akana: There are a lot of services like M0 and other services like Memoriaza Service, which specializes in memory. Yes, if you use an external service, you can use an external service like Open Claw, Hermes Agent, and another Claude Code, for example. If you go to connect to that super memory service, no matter which agent you use, the same memory will be referenced. So, Hermes Agent, of course, is also a third party service like that. It supports memory groups, but there is a memory session search function using SQ Lite as a standard, so it seems to be nice, but at the moment there is a bit of a bug in the Japanese language.I wonder what Japanese is.Sometimes there is a space between English words and Japanese words, and sometimes there is not.If you can't do that, try using the alphabet, for example, Hermes, and later when you search for Hermes during a conversation. Don't you want to be caught by that? Maybe it's not happening right now. Why is this just an implementation bug? But it seems like things like that do happen. Officially, there's a memory service called Honcho, and I kind of recommended that you use it, so now I'm thinking about using that. Honcho. What's interesting is that it connects to various services such as Super Memory and MEM0, but it's the standard built-in memory MD that I just mentioned. It's designed to be used in addition to the SQLite session search, and Memory MD and SQLite conversation session searches are already based on Hermes, which is already the same worldview.The reason why everyone thinks it's good is probably because the creation and updating of skills is automated.As I said before, skills are just the same AI procedure manuals that you do over and over again. If you want to do this separately, in this way, in this way, the AI will refer to that skill the second time you do it.Yeah, that's the story.And so, normally, I did a certain task, and then I said, ``Make this a skill,'' and I would like it to be made into a skill.Also, I usually used this skill, but somehow, I think it would actually be better to do it this way.I run the skill once. After all, it seems like I want to do this here.Okay, so I have to say after that to reflect this in the skill.The next time I do it, the same thing will happen.As I said before, I want to do this like this.I'll leave the original saved skill as it is and do it like this.Yes, why do I do that in the first place?In response to the LLM system prompt, there is a condition for HelmS Agent to automatically create and update skills.Yes, just to put it simply. For example, if a user requests something like this, you say, ``I understand,'' and then you try to do it, but when you use that super tool, for example, 10 times, you realize that this is a complicated procedure.Why do you create a skill?Yes, yes, yes, yes, yes. It feels like LLM creates it when it makes a decision, and then it gets created a lot, but I don't create it all the time, I look at the existing skills, and if it's similar to it, I'll update the patch in it, so I'm given a tool called skill manager, and it's built into the system prompt, so I can create it, edit it, or update only a part of it, so I guess that's how it works for ordinary people. These two are the biggest ones, and I've read most of the documentation, and it's amazing. There's a lot of differences. But the other thing that's really cool is not the direct functionality, but as I said earlier, the usability of an AI agent depends a lot on the number of tools you can use. For example, it's speak-to-text, so you can transcribe from voice to text, and then you can get the web page. If you want to use something like an AI agent that converts it into easy-to-read markdown and retrieves it, or generates an image, or uses a browser in a nice way, you end up doing something like this. Um, rather than preparing something on hand, like the memory other service I mentioned earlier, some specialized service has prepared it for you, such as image generation, text to pitch, etc. Also, the same goes for browsers.These days, browsers themselves are provided as services, or in the case of text to pitch, open AI, etc.In the end, it's better to use services provided by various services, right?But the problem with that is, for example, if you have 5 services you want to use, You sign up for 5 services, but you pay out 5 API keys and set it up. It's like tsu, right? Shrug. I don't know if you're familiar with this, but tsu is katakana tsu, but I don't understand what it means. It's a katakana tsu that foreigners see. To a Japanese person, it looks like a tsu. To a foreigner, it looks like a face. Katakana tsu. Only half of the corner of the mouth is raised. This is the name of the developer of Hermes Agent. What do you think? Somewhere like North Research has issued a subscription service. If it's a paper function, for example, it's a service called Firecrawl for scraping, or it's called FAL for image generation. I think FAL is a service like Nanobanana and other image generation services, but if you subscribe to this one service, you can use a variety of services. Is it okay to do something like Open Router? Yes, it's like Open Router. Open Router can use a variety of LLM models, but there is a service called Firecrawl for crawling, FAL for images, Open AI for text speech, browser use for using browsers, and modal that provides a nice secure environment when executing commands.In the first place, to use LLM, don't you need LLM even though it doesn't create more than AI? They offer an all-in-one subscription. Isn't it a bit too generous? But it's not particularly low, and you pay for what you use. All I do is the LLM part, and I sign up for an unlimited service from another Chinese service, and the rest, like fire crawling and image generation, is paid to that subscription for about $20 a month. This is really convenient. With one contract, I'm really an AI agent. Tools: The more good tools you give them, the more useful they become. Do you actually use them? When do you use them? Huh? I wonder what it is. I think it's almost all the use cases that I use for things like chat GPT right now. For example, I usually have someone record what they see. Also, of course, task management, things like adding this to Google Calendar, taking a screenshot of the image they give me, sending it to that person, putting this in Google Calendar, and things like recording this in Obsidian. With open claws. It's hard to use my skills. Ah, that's right. I have already decided to put the notes created by the AI in Obsidian in a directory called AGC. They won't put them there, or even if they do, they won't follow the file naming rules. So far, they've been following the rules 100%, which is nice. They're supposed to be using the same model, so it's strange. I feel like the harness is important after all. In the explanatory article about Hermes Agent, I write it as if it were a self-improving agent that gets smarter the more you use it, but it's like the skill you mentioned earlier evolves.Also, memory is probably better than open claw.Memory doesn't even need to be digested as a skill.When I say write code like this, it seems like it's JavaScript, and if I say it and memorize it, then the next time I write it in JavaScript instead of Python. Well, I guess the selling point is that it automatically becomes smarter if you use it without permission.I see.The more you use it, the more you use it.I thought about it this time, but when I was using Open Claw and migrated to this agent, there is a command in the Hermis agent that is like a migration from Open Claw, and it also transfers skills, so I don't think it's a big deal if I have this skill.I don't have any problems with that, or rather, I'm also recording on Obsirian. Obsirian has nothing to do with OpenClaw, so it was easy to switch to it.I see, Obsirian is like a personal document management system.I see.I see.Even if you use it, it feels overwhelmingly different.I don't mean it's overwhelmingly different, but that's probably what it is.The more you use it, the more it becomes optimized for you, which is probably a good thing.It's also because calling tools is less likely to fail than OpenClaw. What's the difference? I wonder if it's from the beginning Text to Speech and Speech to Text You can use it if you sign up for one subscription to Open AI or something like that. Why with Telegram or something like that, if you send it by voice, it's voice, but it turns into a string once and goes to the agent on the other side. Another surprising feature is Discord's voice chat room. You can invite coids to the voice chat room. Tell it to the air agent, and the air agent will return the string as a voice, using text and speech. Oh, that's what it sounds like. You mean voice? With voice. You can talk to your assistant through the glasses. Yes, this is convenient. Definitely one of Open Claw's favorites. When we first introduced Open Claw, there was talk that it was convenient, but security had to be pretty good, so maybe that's what Pumpy would use. Is it the same for Hermes Agent? Open Claw is doing a lot of work on security, and that's right. Hermes Agent is doing the same with Sorchi, but this one is doing a lot of work. It's working hard. First of all, when you install a skill, it does a security scan, of course, and even the skills you created can sometimes fail when you register it, so it's a dangerous skill. When you execute a command, it retrieves something common and uses it directly in your shell. It seems like they're paying a lot of attention to security, as they automatically judge whether it's a dangerous command or not.But since it's an LLM, they can do anything no matter how far they go.They've also incorporated something into the Convenience Desktop App, which allows everyone to use the company's MCP.What's this?I thought it would be useful for everyone to automatically turn it into a skill, so I started by managing that skill.Create Update If we provide an MCP tool that allows you to add a file and update only a part of it, and then do the rest from that tool, there is a predetermined system prompt called Claude Code, and you can use command flags to add just this part to the predetermined system prompt.Yes, so by adding the part about when to manage skills to the Claude Code system prompt, the user doesn't have to know anything. It's like having the added skill tool create and update it without permission. So you don't have to tell Claude MD or Agent MD when to update the skill. Thank God, Codec probably can't do this. Right now, it's like adding it to the system prompt. Yeah, that's why they're doing things like automatically rewriting the agent MD. Thank you. I thought it wouldn't work, but it's working properly. I feel like it's been helpful. I think the AI agent is able to use high-quality tools that users usually use. So, with the subscription integration, it seems like that's the point. It feels like it can be done. Yeah, you have to have a lot of cooperation. Yeah, that's right, so it's okay to do it. But I feel like Fire Claw and Modal really cooperate, but I don't understand open air. There's also a theory that this company signs a single open air subscription, takes an account, and then uses an API to manage costs internally.I hope they do their best.So if it was a text-to-pitch API, that would be fine, but if it was shared by someone else's account using a browser or something, it would be a big problem.I believe that this is separated, but yeah, this is really convenient.So in the case of open claw, If you don't sign up for a lot of things yourself, you have to start with the LLM model first, but this time it seems like if you sign up for this much, you can theoretically use it.That's a good thing, too.It seems like you can use 300 models.That's amazing, there are so many.There are a lot of people I don't know about, like Chinese ones.I thought it was just Lipseek, but you and GLM like the old Quen, etc. Minimax or Minimax? Which one is it? I also forgot about Xiaomi, but when I registered for Kuen in private and tried to use it, I was asked to submit my passport to verify my identity.I was out of town for a while, so I stayed there wondering what to do.Alibaba took my personal information.What is it for?It's for Yabo.I'm saying it's bad, but I'm also like that.I've been saving up all the recent articles on Gomshilian to read later, but I want to make a summary of them.Is Quen cheap? Yeah, so I asked LLM which model is cheaper, so I guess there's no cheap model for Quen or Lipseek.Does Quen have a flat rate?Does that mean it's not a flat rate?Isn't it a flat rate?Maybe that's the case.I don't think I'll be writing thousands of articles.I see.The recommended option in such a case is an open router contract.What is that? Open Router is an LLM model, but it has all the LLM models on the back. So I tried using it because it's cheap, but if it doesn't seem right, there's no need to change from Open Router to something else. I'll change the model name in the summary to you next time. If I change it, it'll just route to you, so it's convenient, so I'll go with that. By the way, there are a lot of free models in the back, so can I try that? You can try it. It's useful. Open Router also has gift cards, so instead of a Starbucks card, everyone will be happy if you give them a $20 ticket to Open Router. It's too weird. I wonder if even engineers use it? I wonder what everyone thinks? It's been around for quite a while. Only engineers who like agents use it. Before, I used to be like Lou Cord and Klein. Lou. When I contributed to that, people said I was ecstatic. I received something like a $100 Open Router voucher, and it was quite convenient. That's why I can set the Open Router model to Open Router code or Klein, so I can use it as an override. It was also good to provide useful information and go beyond LLM. That's great. Who are the people who say they want an Open Router gift card? I hear you like agents too. If anyone wants one, please write to me. Please write to me. What is that? I'm not saying I'll give it to you. I don't have anything to give, so I just need to buy it. Why did you ask me? I just want you to show that you want it. That's right. I was wondering if everyone was using it. I see. I see. If there are people who think the Hell Mace Agent is convenient or who use it in this way, please let me know. If you have any thoughts, questions, or feedback, please send them using the X hashtag Magical FM in all lowercase letters or the letter format in the summary section. If you press the Spotify bell mark, you will receive an update notification, so please do so as well. Thank you very much.