← Back to search

How to Run Hermes Agent FREE Forever!

AI News Today | Julian Goldie Podcast · 2026-06-20 · 13 min
relevance 100 2813 words Episode page ↗ Audio ↗
Show full episode description
Local AI Agent Engine: Build Private & Free with Hermes Learn how to build a local, private team of AI agents using the Hermes engine to automate tasks 24/7 without API costs or internet. This video demonstrates a full Agent Operating System with Kanban orchestration and optimized local models like Llama 3.1.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How to run a private, free, local team of Hermes AI agents that build and automate without cloud or API costs.
Benefits
  • Fully private — data never leaves your machine
  • Free and runs offline, no API costs
  • Agents work as a team 24/7 on a Kanban board
  • Shared memory across every top model
  • Swap models freely (Claude, GLM 5.2, Kimi K2.17)
Use cases
  • Built an SEO blog about OpenClaw from one sentence
  • Built open-world and Minecraft-style games via GLM 5.2 + Claude
  • Automated daily SEO that previously took ~1 hour per day
  • Content Kanban auto-posts video and blog from a keyword
  • Fusion beats banned Fable 5 on benchmarks via five models
KPIs / results
  • Runs entirely on a Mac, no internet, $0 API cost
  • SEO task cut from ~1 hour to one click
  • Judge loops drafts until passing (rejects 4/10 work)
Tools / build
  • Local Hermes agent engine (Ollama)
  • Agent Kanban board for local agents
  • Content Kanban crew (content, SEO, video agents + judge)
  • Memory Galaxy shared memory
  • Fusion panel (OpenRouter API)
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00
Intro: Local Hermes Engine
  • Build a local, private, free Hermes agent engine
  • Give a goal; it plans, runs tools, builds
00:54
Local Agent Kanban Board
  • Local agents on a live Kanban board
  • Built SEO blog about OpenClaw from one sentence
02:21
Setup & Model Options
04:37
How AI Agents Work (The Loop)
  • The agent loop: plan, run tools, build, verify
05:17
Cloud vs. Local AI Benefits
  • Local is private, free, offline
  • Cloud offers more power; mix both
08:00
Best Models for Your Mac
  • Best local models to run on a Mac
10:08
Honest Limits & Use Cases
  • Honest limits of local models
  • Practical use cases for each
11:26
Agent OS & Community Access
  • Agent OS, shared memory, Fusion vs banned Fable 5
  • Community access via AI Profit Boardroom
Today I'm going to show you how to build a local powered Hermes engine which is a team of AI agents powered by local free private AI and you can basically automate whatever you want with it. So this is something I call the local Hermes agent engine. You can give it a goal that could be a voice or text. Hermes breaks it down into goals and steps. Then it runs the tools with commands and creates the project for you and it's built and verified plus you can come back to inside workspace right. And the great thing about this is number one is private so your data is not going to the cloud. Number two is free because you can and you can have these agents working as a team building stuff 24-7. And number three it is completely running locally which means you know you don't need the Wi-Fi or anything like that to run it. So you can give it a goal it plans it runs the tools it builds the thing every step on your own machine. Pretty amazing. How does this work? So you can see an example right here this is the agent Kanban board local agents. This is a team of local offline agents with Hermes working on a live Kanban board separate Kanban board just for them. And so I could say for example build out an SEO blog about Open Claw which we've typed in here and then it assembles the board. So what it actually does is it comes up with the ideas and what it needs to build for this website. Now if we click on for example run the team what that's going to do is start building this step by step to create whatever we want which is pretty amazing in itself. Now if you're wondering okay like what does this actually look like in action well you can see them building now. So these are AI agents running locally with Hermes to build and automate whatever you want. And then if we go over to the workspace here we can see the stuff we built previously. And we can also see the prompt. We can open this up full screen. You can see an example. These are just simple examples just to show you what's possible. I literally built this this morning into our agent operating system and you can see it's now creating the landing page and everything else using our AI agents together. Now we can also open this up. We can give it feedback. We can tell it to refresh it. We can ask it to improve everything that's built. But that's basically running as a team as a loop together which is pretty amazing. So you can see some examples of what we've built with them and each of these just comes from a single sentence. So this was built by a model running entirely on a Mac. No internet. No API costs. Just Hermes agents teams of agents working with local models. Now if you're running for example Ollama with Hermes you can get this set up pretty quickly. So there's a couple of options here. Number one is you can actually go to Ollama. Just make sure you have Ollama running in the background. And then from here you would go inside whichever model you want to run locally. For example like this one. And you can just copy and paste the command to run this inside your terminal with Hermes. However if you want teams of agents that's why we've created this Kanban board here. So we can have a team of agents working together to build and automate whatever we want. And so the great thing about this as well is like we can speak to our agents inside this section as well. Which is great. So we've got like for every API we can change the model and we can create a separate agent profile. We can also talk to Hermes here. We have Hermes Jarvis which is a voice activated version of Hermes. We have a studio where we can generate images, video and voice. Bear in mind you can generate images for local models like Ernie from Baidu. Then you've got your session section and the workspace here with everything you've built with each model broken down. As you can see. So we can see everything that we've built previously with Hermes agent as well. Now if we go to the manage section here and we go to models you can also change the model this way. So for example you can see that we've got LM studio ready to go which we could run local models with. And then we can also run it with NVIDIA. NVIDIA is another way to run local models. And of course OLAMA. OLAMA is another way to run local models. So we've got OLAMA ready to go with two different models. Over here as you can see you can run Gemma 4 with this for example. And you can have teams of AI agents working together. Now if you want to visualize the teams working together. You could have a Kanban board here. So if we go back to this system you can see everything is built now which is pretty nice. And we can just see everything that we've built inside the workspace here. But we can also assign new tasks and get them working together again on new tasks. And just keep building these teams out. So picture this. You can have a team running 24-7. Building new useful stuff. On a Kanban board you can see exactly what's built. You can view it inside the workspace. And this is all free private and offline. Crazy stuff when you think about it. So how does this work? Step by step. Well if you look at this whole system. It's just one goal in and a finished job out. So you don't hand it a checklist. You hand it the outcome. So you could say okay list the files in this order and write me a summary. It doesn't answer in words. It runs the command. Reads the result. Writes the files and then tells you what's done. And that loop. Plan a step. Run a tool. Look at what happened. Plan the next. Is what makes it an agent instead of a chatbot. So you give it the goal. It runs multiple agents. So for example it will run one for running. One for reading. One for writing. And then it's done. So it plans the steps. Runs each tool. Checks the result. And then it says done when the file was really there. So you can see an example of a real run that we did with this system. Now you might say why do this? So if you look at this. Particularly if you're using agents like agents. They use up a lot of tokens. So a lot of people want to run stuff. And the thing is. If you're not running local models. Then every command file is sent to someone else's server. You pay per token on every single run. If you have no internet. Then you don't have an agent. It's just going to stop. You get rate limited sometimes. And then you get cut off mid-job. We've all seen that with Claude. You trust it. Built the thing. But you can't see it. And the result is a powerful agent. That you don't really control. You know. And we've seen that with Fable 5. Like when Fable 5 got taken down. Loads of people's workflows just ended. So with the new way. Every step runs on your Mac. Nothing is sent out. It's free. To run all day. It can work in a plane. In a cafe. With the Wi-Fi off. It breaks the goal down into steps. As you saw. And then does them itself. It checks the disk. Shows you exactly what it built. And the result is a real agent. You own free and private. That can run 24-7. So the work is the same. But the difference is where it runs. How much it costs. And who sees your files. That's a big difference between the old way. And the new way. Also something to be careful of. Is like local models. Sometimes they say they've built the thing. But they've not actually built the thing. Like it's very common. If you've ever run a local model. So with this system. What it actually does. Is the agent says I've built it. And then it checks. Did it really build a file? If it didn't. Inside the workspace. If it does. Then you actually get everything. And it builds. As you can see. So everything here. Was built with local agents. Now if you want this. Local Hermes engine. With the agent operating system. It lives inside our agent OS. The dashboard. Where all my agents sit on one screen. The local agent runs right next to all our cloud ones. Free and offline. Ready for the quick stuff. So you stop using up tokens. The full agent operating system. With local and cloud agents. Plugged into one dashboard. With builds preview in live. You also get the local Hermes engine set up. The profile. The model. The workspace. Etc. A 30 day road map. To wire agents. That do real work for you. And a room of 3,600 builders. Right. So there's always someone online. Ready to help you. So you can actually get that. Inside the AI Profit boardroom. Link in the comments description. Or go to the AI Profit boardroom.com. And if you go to the classroom. You can grab the agent OS. Inside this section. As you can see. If you're already a member. By the way. And you're watching this. You're like. Well how do I add these new updates? So we have. Inside the new zip file. Which we update daily. We have an update section. Where you can add the new updates. And import them into your existing setup. So. If you want to update this. You're already a member. No problem. You can do that. Using the setup. And the zip file. That we give you inside there. So. This is basically it. Now. Let's talk about wiring it in. What models we've used. Etc. So. There's two pieces to this. A local model. That fits your Mac. And a Hermes profile. Pointed at it. Hermes needed model. With a big enough context window. So. You could use. For example. Like. Even. I mean. We. As this example. We've used Lama 3.1. You could also use. GPT OSS. I've tested a lot of local models. And unless you have a really good setup. Many of them are bad. And super slow. So these are the two models. GPT OSS. Lama 3.1. 8B. That have worked for us. On a Mac Studio. And it really depends. What setup you've got. Like. For example. Someone who was posting inside. The app. Profitboard earlier. Daniel. Who's an absolute legend. And you can see an example. Of what he generated. Using the local model as well. That's with an R5090. So. It's pretty amazing. What you can do. If you have a really good setup. If you don't have an amazing setup. You can just use. Lightweight models. That I've shown you today. So. You can wire it in. In about 10 minutes. You just pull in. A light capable model. So. For example. Like. Lama 3.1. 8B. It's only 5 gigabytes. Runs the most Macs. It's fast. And crucially. For an agent. It can use tools. Instead of faking. Using the tool. Right. A lot of local models. Say they've used this tool. But if not. Then you make a Hermes profile. Point at it. So you have a dedicated. Local profile. The offline agent. Uses. And you can point it. At the local model. And keep its context window. You can keep it ready to go. So. Quite often. If you're swapping. Or if you're opening up. Local models. And then shut them down. It slows everything down. Where. With. With this system. You can have like. Lama 3.1. Ready to go at any time. And then you just give it a goal. From the Kanban board. Which you can see right here. And it's a really powerful way. To. Orchestrate your agents. Locally. I mean. We have a system inside here. Where. And they can be orchestrated like this. We have. The pipeline. Where we can automate. And build. With teams of agents. Anything that we want. And we've got a full gallery. Of what we've created here. It's like a. AI agent powered to-do list. But if you want local agents. Building together. This is the best way. That I've seen. To orchestrate local agents. Running together. You can see an example. Of the four steps. Right here. And how it works. Now. Let's talk about the honest limits here. This is great. For. Running commands. Firework. Quick code. Etc. Not great. For super giant complex builds. So. A small local. Local model. Is like a fast helper. It's not a flagship. Or a frontier model. You can hand the heavy jobs. To a frontier model. But then. The smaller jobs. You can give to your sub agents. Or to Hermes agents. Whatever. And it checks itself. So when a build doesn't land. The engine tells you. And then it just runs again. So some people say. Well. A local model can't do real work. You've seen examples. Of actually building stuff today. Of people say. Well if it's free and local. It's not actually going to be useful. Again. We can actually get them running. With a team of agents. Using this agent. Kanban. Which is pretty amazing. Then other people say. Running an agent locally. Isn't that too slow? And that's only the case. If you run a model too big for it. So you want to pick one that fits. And I've given you a couple of options today. Now. If you were thinking. Okay. Like. That's great for you Julian. But what about everyone else? You can see that we have over 183 wins. And pages of testimonials. And wins and reviews. Right here from AI Profit Boarding members. So I know that. If I'm not a coder. And I can get results with this. And you're not a coder. Then. And these members are not coders. Then we can all get results with this stuff. Right. We're at the point now. Where you can just talk to AI like a friend. And it builds whatever you want. So that's basically. The whole system. Now if you want to get. This agent operating system. For me. With Paperclip built in. The AI Agent Mastermind. The Pipeline. The Agent Kanban. For local agents. Clawed. Open Claw. Hermes. Gemini. Codex. Kimi Code. G2. Grok Build. Free Clawed Code. We've got Fusion inside there. Local agents. We've got a self. Iterating feedback loop. For loop engineering with Hermes Agent. The SEO system. For deploying content. A video agent. For just building. And creating amazing edited videos. Like you see with AI. It's got my avatar inside there too. We've even got a music agent. A game studio. Notebook LM built in. Kanban boards. And a full memory system. That's all inside the AI Profit Boardroom. Which is my community. Focused on helping you scale. And save time with AI automation. And this is an amazing community. Of 3,600 members. And you know. If you ask. If you ask a question inside the community. We get back to you daily. So you can see for example. Like Brian asked a question. Seven hours ago. We've already replied to him. Right. And helped him. And he says. Got it. Thanks. Right. So if you ask a question. We get back to you pretty quickly. I make video tutorials. Helping members inside their daily as well. You get access to all of my best trainings. Inside the community. And inside the new daily update section. You get access to my agent OS system. With the video tutorial. The last update date. The zip file completed. So that you can install it. New daily tutorials. Like you can see. A calendar. We have four weekly coaching calls. So you can jump on this coaching course. Ask questions in real time. Get help and support whenever you need it. Inside the map. You can meet people in your local area. Who are building with AI agents. And this is all available. Link in the comments description. Or just go to the AI Profit Boardroom.com. Hope to see you inside there. Cheers. Bye. Bye. Bye.