← Back to search

Hermes Agent + LM Studio is INSANE!

AI News Today | Julian Goldie Podcast · 2026-05-09 · 7 min
relevance 100 1570 words Episode page ↗ Audio ↗
Show full episode description
Run Hermes AI Agent Locally for Free with LM Studio Learn how to run the powerful Hermes agent locally and for free using LM Studio. This step-by-step guide shows you how to set up a private, offline AI system using high-performance models like Quen and Gemma on your own computer.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Running Hermes AI agents free, private and offline using local models.
Benefits
  • Run Hermes for free with local models
  • Fully private and offline operation
  • Quantized lightweight model versions
  • Switch models and providers easily
  • Works on Windows, Mac and Linux
Use cases
  • Run powerful free AI agents like Hermes locally with no internet via LM Studio
  • Download and run Gemma 4 quantized model on a local server
  • Use Hermes offline on a plane with local models
  • Connect Hermes to Ollama local models like Gemma 4
  • Run Nous Research models built specifically for Hermes locally
KPIs / results
  • Qwen 3.5 1B and 35B model options
  • GLM 4.7 Flash recommended local model
Tools / build
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00
Intro: Free Local AI Agents
00:41
Setting Up LM Studio
  • Download free LM Studio from lmstudio.ai; install models
01:09
Choosing the Best Local Models
  • Pick best local model; quantized lightweight versions
02:03
Starting the Local Server
  • Toggle on and start the local LM Studio server
02:41
Connecting Hermes to LM Studio
03:36
Switching Models & Providers
  • Switch models/providers; also use Ollama locally
05:06
Why Use Local AI Agents?
  • Local agents are free, private, fast and offline
06:01
Implementation & Next Steps
  • Step-by-step guide and 30-day implementation roadmap
Today we are going to be testing out Hermes Agent and how to use it for free with local models with LM Studio. So LM Studio can run with Hermes Agent now, which means that you can run Hermes Agent and use it for free with local models. Bear in mind, Hermes Agent itself is a free open source model. And then if you can plug in a pre-LM Studio local running AI model, then you can use it for free as well. So it has a free sort of way to run it inside Hermes. So let's get started with this right now. And you can see that it's quite easy to integrate these two together. So the first thing that we need to do is run the server from LM Studio. So we're going to open up LM Studio. If you don't have LM Studio already, you can get it at lmstudio.ai. It's a free app. You can download it and then you can install local models. So for example, you can set up a local server and load a model with a local server. And this means you can run local models with LM Studio. So for example, if we search inside the model section here, we can install something like, for example, Gemma 4 and then run that with our AR models. It's pretty simple and easy to set up. So for example, like this. And depending on your setup, LM Studio will actually tell you what's the best model to run. So you can see, for example, here, it doesn't recommend using these models because it says likely too large. But you could run a more lightweight model like, for example, this. And the other thing is that it's important to note here is with LM Studio, you get quantized versions of the same model, which means that you get more lightweight versions of the model that are still powerful for running as a local AI. And then you can also use Hermes offline as well. That's the other benefit here. So we can download Gemma 4 like you can see here. And this is a super small lightweight model. But I just want to show you as an example of how you can get that set up. And then what we can do from there is we need to start LM Studio, right, like this. So what we do to do that is we can go into LM Studio, close that, go to local server, and then we just have to run it. So if we click on running now, LM Studio server is now running locally. All we did was we just toggle that bit right here. So we toggle that on. Once we've done that, we just need to load a model onto it as well. So if we load a model here, we'll just get that model from Gemma 4, which is downloading, right? So then we can get this started. Now, once you've started the local server, as you can see right here, we then need to run Hermes Agent with LM Studio as a model provider. How do we do that? So we're going to go to Hermes Setup inside Terminal here, as you can see. And then what you want to do is complete the LM Studio setup from there, right? So you see how we've got LM Studio on the model so we can select. We can go from there. But it might as well. You can also use this with Ollama, which runs locally too. So, for example, we've got Gemma 4 with Ollama set up locally too. You can also plug in the API key if you want to add an API key. And you can just go with the defaults if you want as well. Then we can restart the gateway to pick up the changes. That's basically how we can set this up, right? Now, you can see that Gemma 4 is now downloaded. So we can use that inside a new chat. And if we go to the local server here, we just have to load the model, which is Gemma 4. I wouldn't recommend using this version of Gemma 4, but it's just as an example. And that's basically how you can get this working with LM Studio. And that's basically how you can set this up. And then from here, if we go back to our terminal, click on Done. We can launch Hermes in chat. Select Model, right? So we can switch Model by typing Model. And then we can switch to LM Studio right there, right? So that's how you can switch to LM Studio. And then we would select Google Gemma 4 inside the Studio there. Pretty easy to set up with Gemma 4. It also tells you what the context window is, the provider, the model you switch to. You can have multiple different models with LM Studio too. Now, if we wanted to run this locally with Ollama as well, let me show you how to do that. And again, like inside LM Studio, I don't recommend using that version of Gemma 4. Just wanted to show you a quick way of how you can pick the model, load it in the server, and then from there, connect that to LM Studio, right? That's an example. But some of the best options that you can do here for this sort of stuff is like GLM 4.7 Flash is pretty good. And there are actually some local models designed specifically for Hermes, right? So for example, you can see a bunch of like Noose Research models here. Noose Research is the founder of Hermes Agent, which means these are some of the best models you can run locally with these agents, right? Also, you could run something like a Quen. Quen is pretty good as a local model. I would use Quen 3.5 for that. And they also have some lightweight model versions, as you can see right here. So Quen 3.5 Imb, 35B if you've got a good setup, et cetera, right? So that's basically it. And then if you wanted to configure this with Hermes using Ollama as a local model setup, you can go over here, Olama Cloud. You can plug in the API key if you want to use the cloud method. And then from there, you're good to go. That's how you can run it. So just to recap on this whole process, you can now run powerful free AI agents like Hermes on your own computer with no internet, right? Using LM Studio. LM Studio is a free app you can install. It lets you download and run powerful AI models. You have to have a good setup for this to work, but it will work on Windows, Mac, and Linux. I do this on a Mac Studio. Honestly, I prefer like cloud-based models. But if you do want to know how to use LM Studio, this is where you can do it. And you might be wondering, okay, how does this work together? So LM Studio acts as an engine. It runs the AI model on your computer. And Hermes engine agent acts as the driver. So it tells the AI what to do and gets tasks done for you. It's a local, fully private, fully free AI agent system. Now, why should you care? It's free. It's private. It's fast as well, right? And it works offline, right? So if you're like on a plane or something like that, you can still use Hermes and you can still get the most out of it, which is pretty cool. So some examples of models you could use, that could be like the QEN 3 series, could be DeepSea, Koda, Llama, as well as another one. And if you want to set it up, we've got a step-by-step guide right here and a full roadmap on how to implement it. And what we'll do is I'll put that inside the AI Profit Boardroom and we'll add the video tutorial as well inside there. But that's the full guide on how to use it, how it works, why you care, how to set up, plus a 30-day roadmap for implementing it into your business. Now, inside the AI Profit Boardroom, we add like new daily tutorials, full step-by-step guides on how to use all this stuff, which is pretty amazing. You also get all of my best trainings on AI SEO, how to go from beginner to expert with AI automation, my best trainings here, all the systems I use inside my business, and a full course for agencies on how to get clients as well. So feel free to check that out. Link in the comment description or go to the AI Profit Boardroom. This is my community that's focused on helping you grow and scale with AI automation. Inside the community, you can ask questions, get help and support whenever you need to. So inside the classroom, you get my trainings. In the calendar, you can jump on weekly coaching calls and get help and support whenever you need to. And also inside the map, you can connect with people in your local city who are also using AI agents just like you. So feel free to check that out. Thanks for watching. I'll see you on the next one. Cheers. Bye-bye.