← Back to search

EP308 – Hermes Agent true battle: travel, change, speed, change, expand 24/7, adjust the number of workers to the top of the road!

AI懶人報 · 2026-05-11 · 14 min
relevance 54 111 words Episode page ↗ Audio ↗
Show full episode description
🚀 特別感謝贊助 本集節目由 VoAI 絕好聲創 提供技術支援。 🎤 VoAI 提供最有「台灣味」的 AI 聲音,支援情感語音、台式口音,甚至能一鍵生成虛擬人! 🎁 AI懶人報聽眾專屬優惠: 👉 輸入優惠碼 AILRB26 立享 95 折! 👉 API 方案用戶:透過專員聯繫並告知從「AI懶人報」來的,額外加贈 10% 使用額度! 立刻體驗: https://www.voai.ai/ 告別手動敲指令的時代!本集帶你從單次問答進化到「自主式 AI 協作」,教你如何升級為數位團隊的指揮官,讓 AI 成為 24 小時為你工作的數位員工。🚀 💡 AI 互動的四層級演進 👉 從單次問答跨越到自主協作,讓你的角色從「執行步驟」轉向「定義目標」。 💡 部署 24 小時運行的數位員工 👉 透過 Hermes Agent 讓 AI 脫離螢幕,具備記憶與自我修正能力,實現真正的自動化執行。 💡 容器化部署的安全性架構 👉 將 AI 放入獨立的 Docker 容器,實現任務分工,確保敏感權限與資料的完美隔離。 💡 模型路由的降本增效策略 👉 透過 OrcaRouter 動態分配算力,平均降低 40% 以上的 API 成本,避免大砲打蚊子。 💡 從操作者轉型為指揮者 👉 未來 AI 的核心價值在於系統設計,建立具備自我進化循環的數位團隊,讓工作流程更聰明。 --- 📎 參考資料: Hermes Agent: Zero to Personal AI Assistant (1 Hour Course)(2026/5/10・55,181 views) https://www.youtube.com/watch?v=gb5TlGw6Uks Hermes Agent is blowing me away…(2026/5/9・32,213 views) https://www.youtube.com/watch?v=cEo95olh2j4 OrcaRouter AI Review 2026 — Best AI Routing Platform?(2026/5/9・16,446 views) https://www.youtube.com/watch?v=DzsW88b_KTc Agentic AI Systems, Clearly Explained(2026/5/9・11,781 views) https://www.youtube.com/watch?v=kwRTUw8pb2c Codex Just Became THE BEST Long Running Agentic Harness(2026/5/10・10,274 views) https://www.youtube.com/watch?v=nOFordZCyzs 歡迎請我喝杯咖啡,幫助我繼續把節目做得更好唷~! 👉 https://buymeacoffee.com/ailanrenbao -- Hosting provided by SoundOn
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How to move from static automations to autonomous, collaborative AI agent systems that run 24/7 without supervision.
Benefits
  • Four-level framework from chatbots to collaborative agent teams
  • Codex Ghost runs long tasks with built-in coordination loop
  • Hermes persists memory, skills, personality, scheduled tasks
  • Self-improvement loop rewrites skills after corrections
  • Containerized Hermes isolates finance vs marketing agents
Use cases
  • Codex Ghost ran 45 minutes to an hour autonomously with no intervention
  • Hermes fetches AI news at 6am daily, filters noise, posts summary to a channel
  • Hermes monitors YouTube comments and replies in brand tone; nightly server security scans
  • Ocarouter cuts costs on average 40% below list price, API fees 60% to 80% in some flows
  • Routes across 200+ models and 10 suppliers with under 1 millisecond latency
KPIs / results
  • Codex Ghost runs 45-60 minutes unattended
  • 200+ models, 10 suppliers checked by Ocarouter
  • <1 millisecond routing latency
  • 40% average savings; 60-80% API fee cuts; zero markup
Tools / build
  • Codex Ghost (planning mode + integration loop)
  • Hermes (memory, skills, personality, scheduled tasks)
  • Containerized Finance Hermes and Marketing Hermes on a VPS
  • Ocarouter (OpenRouter-style model routing)
  • Telegram bot control of agents
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
Welcome back to share some good news with you. AI Lazy Newspaper also has a YouTube channel. If you want to know more about the topics discussed today, you can directly search AI Lazy Newspaper on YouTube. There are reference links in the information column under the video. Okay, let’s return to today’s topic. There are a bunch of new AI tools popping up every day. Do you often don’t know where to start? Don’t worry, here is AI Lazy Newspaper to help you select the five most discussed AI tool videos every day. Let AI enter your daily life in ten minutes every day. Let you unknowingly become an efficiency master. I think what I am going to talk about today will make you rethink the question of what exactly am I using AI? Because we are no longer just asking AI questions, but directing an automatically operating system. There is a tool that allows you to enter a line of instructions and then make coffee, lure dogs, and stare at the ceiling. When you come back an hour later, not only have you completed the task, but you have also tested it yourself and made your own mistakes. The visual materials have been generated. My first reaction is, wait, do I still have a job? The second reaction was immediately, this is exactly what I have been waiting for. Before talking about the specific tools, I want to help you build a framework first, because if you don’t understand the background, you will hear the following things. You can think of it as four levels of love and interaction. It's purely a suggestion, with no action at all. The next level up is that you string together tools and use platforms like M8N or CAPEER to set up data to automatically flow between different services. I've done this myself. You obviously just don't want to spend 10 minutes a week to manually compile a report. In the end, it took 6 hours to build a super complex automated process. You still feel that you are very smart. It's really fun when you run it, but the problem is that this static process doesn't think at all. As long as a situation occurs that he didn't expect, for example, the data format has changed a little bit. The whole process will be broken directly. He will only follow the steps you realistically followed. You are still the one who defined each step. But what is really scary is that the third layer AI starts to judge what to do next. You no longer define each step. You only define the goal and let the model figure out how to do it. You tell it to help me convert this script into another language and then write tests. He will push it by himself, execute it by himself, observe the results, adjust it over and over, pull the required files by himself, plan the instructions by himself, see the errors, and fix them by himself. You can think of it as the model is the brain. The outer execution framework is the hands and feet. It is responsible for actually opening files and planning instructions. It was changed to 4, but there is also a ceiling on the third level. Usually it is the AI that does one thing. In an independent working stage, if you close the window, it will forget all the things it has just learned. The fourth level is today’s focus. Autonomy with collaborative capabilities is the AI system. It is not just an AI, but a whole collaborative team to help you operate. AI has professional skills at this level, long-term memory across working stages, and continuous context. You are not giving the AI a task. You are managing an asynchronous digital employee. Okay, the framework is built. Let’s talk about a few tools that make this fourth layer really possible. Let’s start with Codex. We mentioned him before. At that time, he could already test web applications by himself and control the virtual mouse to click on the interface. But recently Codex launched an experimental feature called Ghost. This thing completely redefines how we handle long-running tasks. Have you ever tried to let AI build a complex project from scratch? You know the pain. The upper and lower micro-capacity began to be insufficient, and the model began to get mixed up, or it ran out of time. In the past, to make the AI run for several hours, you had to build a complex coordination layer yourself. Write a script, which probably means reading the command, looking at the current status of the project, performing an action, updating the status file, and repeating until the completion conditions are met. This method can be used, but you have to build all the printing racks yourself. What to do if the API rate limit is encountered? What to do if the AI is tuned? The basic loop will die directly. Codex calls will build this entire coordination layer directly into the tool. All you need is a slash command. Let me tell you what the actual process looks like. Suppose you want to build a complete application, like a first-class fighting game. You don't directly say to help me build a game. The first step is to use Codex's planning mode to put together a very specific blueprint. You define the game mechanical material requirements. The most important thing is to define clear verification standards. What your command says is that the definition of success is that the application can be built normally, the local development server can be started, and an automated test script can successfully load the screen. Simulate keyboard operations to trigger damage and then enter the victory state. After the blueprint is confirmed, you give an instruction and tell him to follow the plan. Then you can do other things. Codex runs a sophisticated integration loop in the background, constantly checking the internal state to confirm the progress. But what makes Ghost really smart is how he handles the budget limit. Let an autonomous AI run for an hour. Tokens will burn quickly. If you hit the preset usage limit, the script will usually abort, leaving you with a pile of half-written files. and broken code, but Codex Goals is smarter. At the end of each round, he will evaluate what to do next. If he finds that he is about to hit the limit, he will not block it directly. He will end the current round and produce a final report detailing which parts have been completed. After you authorize more degrees, what to do next. It is like a public student. Instead of just writing half the sentence and leaving when the work time is up, you will leave a detailed handover instruction on your desk. The results of the actual test really surprised me Codex You can run it for 45 minutes to an hour without your intervention at all. Use AI to generate original proof diagram materials, with transparent background layers and enemies generated by fine logic. You can also write tests to verify your results. All done automatically in the background. You can use Code in one window to do your main development work and at the same time in another window, let Codex run a large target, handle large-scale refactoring, or build a basic architecture. The two things can be done at the same time without interfering with each other. But Codex is still essentially a tool that lives on your computer. It is designed for you to sit in front of the keyboard and do in-depth engineering work. What about when you leave the computer? When you need a continuously operating system to monitor your infrastructure, manage your content, or run scheduled automation. You cannot leave the laptop with the window open 24 hours a day just to grab some information at three in the morning. This is why next we will talk about the latest evolution of Harmicization. If you are paying attention to the field of open source AI, you may also be following OpenCode OpenCo. de is very useful, but recently it has started to have a symptom of wanting to shoot everything, pushing large updates very frequently, and almost every update will break the system. If you are like me, what you want is to use your AI directly instead of spending 20 minutes every morning debugging your AI execution environment. Harmic has taken a different path, lighter and faster, and very focused on reliability and self-improvement. To understand why Harmic is suitable for situations where you are not next to the computer, you need to understand its five core things: memory, skills, personality scheduling tasks, and self-improvement loops. Let's talk about memory first. When an AI starts up, it's basically in a state of frustration. Harmic uses two persistent files to solve this problem. One is about you, recording your preferences, the way you work. You don't like long-winded answers, or you prefer a certain programming language. The other is about the environment and projects. You don't need to update these manually. When you interact with Harmic, it will automatically intercept important information in the background and update these files. Then there are skills. Think of skills as program memory. A reusable operating manual tells the AI how to perform a specific task perfectly every time. If you ask Harmic to generate a flow chart, it doesn't guess what to do. It will pull a specific operating manual and follow the steps. Then there is the personality setting. This file is to set the personality of the AI. If you want an AI to reply to social messages, you may want it to be enthusiastic and helpful. If it is your personal assistant and wants to communicate with you on Telegram, you may want it to be direct, concise and even with a sense of humor. But the function of scheduling tasks is This is the key to turning Harmic from a passive chat tool to an active employee. Because Harmic is designed to run on a virtual private server, it is always online. You can open Telegram while waiting in line to buy coffee and send a voice message saying Harmic. It will help me catch the latest AI news at 6 o'clock every morning, filter out the noise, and then send the summary to my slide channel. Harmic will automatically write the skills, schedule the tasks and then execute them. Every morning it will start an independent work phase. After running the task, the results will be sent back to you. Can you let it automatically synchronize every night and back up your program for it? You can have him monitor your YouTube comments and respond to you based on the tone of your brand. You can have him run a security scan on your server every night. This changes your infrastructure from static to dynamic. And finally, the self-improvement cycle. This is the most interesting part of the whole system to me. When you correct Harmic, let's say he used the wrong tone or messed up a script, he doesn't just say sorry, he actually rewrites his skill profile and updates his memory to ensure that he does not make the same mistake again. The more you use it, the more accurate it becomes. It sounds complicated to set up, but it's actually very intuitive. You open a cheap virtual private server on a platform like Hoster and deploy Harmic in a large container. This containerization method is very important for security and scalability. Imagine that you want to use Harmic to manage your personal finance and social media marketing at the same time. You definitely don't want a super big AI to have access to your bank account and your Twitter account at the same time. That is a time bomb waiting to explode. So you open a container for Finance Harmic and another completely independent container for Marketing Harmic. They each have independent profiles, their own memories, and their own tools. Your marketing AI does not need to know the highest-privilege password of your server. Your financial AI does not need to know how to generate dynamic graphics materials. Through these independent strengths, you can use different Telegram robots to communicate with each AI. You have a small board that can be put in your pocket. No matter where you are, you can directly give instructions in natural language. But here we go. I'm going to talk about a real-life issue directly because many people don't really think about it when planning this kind of system. You have Codecode running in-depth development work locally. You have Codex running 45-minute autonomous tasks in the background. You have three different Harmics running scheduled tasks every hour on the server. You know what this sounds like? Let an AI classify a customer service ticket, format a JSON file, or summarize an ordinary news article. This kind of task does not require the strongest model at all. This is the problem that Ocarouter is solving. And I think it may be the most important underlying tool that most people don't know yet. Ocarouter is a model routing platform. The concept is very straightforward. You don't point your program directly to Anthropic OpenAI, you point it to Ocarouter. You change a URL and a key. You don't need to understand the other code at all. When your AI sends a request, Ocarouter will analyze it instantly. It's a bit like automatically comparing prices when calling a car for you. It will call a cheap car for a short distance, and call a high-end car if you are in a hurry or have complicated road conditions. It will check more than 200 different models, the real-time status and pricing of 10 different suppliers, and then automatically send it to the cheapest model that is capable of handling the task. And the delay of the whole process is less than one millisecond. The cost savings are really exaggerated, on average 40% lower than the list price. In some processes, API fees can be cut by 60% to 80%, and the best part is that their business model is zero markup. You pay the original rate disclosed by the supplier. Ocarouter does not take a commission from your usage charges. You will also get audit records of each request, automatic switching of backup sources when the API hangs, and cost tracking for each request. Combine this with an autonomous AI system and think about it. If Harmic runs a scheduled task at three o'clock in the morning, it just reads a text file and extracts three keywords. Ocarouter It will dynamically send this request to a super-fast and super-cheap model. But if your Codex is running a complex target loop and trying to make an error, there will be a very troublesome asynchronous state problem in your application. Ocarouter will know to send this to the strongest model. You get two benefits at the same time. Use the top model when you really need top wisdom. Automatically save money when you don’t need it. You no longer have to manually change the model in the configuration file. Just set it and forget about it. Let it help you take care of your wallet. Okay. Now put all the things I talked about today together. We are no longer just chatting with AI, we are building and managing systems. You have tools like Codex, which can receive a high-level goal and then execute it autonomously for an hour, elegantly managing your own budget and status. You have digital employees like Hermis, who are constantly operating and improving themselves, living on your server and running automation to communicate with you through Telegram. You have smart low-level tools like Ocarouter, working under the hood to ensure that this large wave of autonomous AI requests will not bankrupt you. Enter software development The threshold for workflow automation and scale has never been so low. But the real change is not in the tool itself, but in the way you think. Your job is no longer to manually trigger backups or format data. Your job is to define the target design architecture and coordinate these AIs. You are changing from operators to commanders. And for those of us who have always believed that we can lie down instead of sitting down, this is the future we have been waiting for. If you feel that today's content has gained you something, please help me follow the like track on Apple Podcasts and leave a five-star review. If I have helped you clear up some pitfalls, please leave a five star so that more lazy people can find us. FBIG Feds Soul AI Lazy Newspaper can find me. The recent AI development is still so happy that it scares me. I don’t want to miss it and continue to listen to AI lazy people tomorrow. I am Kang Lalan. We will see you tomorrow. Bye.