← Back to search

Hermes Agent: Assessing Memory, Reasoning, and Reliability in Open-Source AI

AI with Arun Show · 2026-08-22 · 20 min
relevance 74 3583 words Episode page ↗ Audio ↗
Show full episode description
In a recent demonstration, UC Berkeley student Sainik Ghosh showcases the capabilities of Hermes Agent , an open-source tool designed to retain memory across different user sessions. Unlike standard chatbots that treat every interaction as a new encounter, this technology builds a personal profile and develops reusable skills to improve its performance over time. Ghosh illustrates the agent's versatility by using it to manage travel preferences and conduct automated resume screening for technical roles. While the system effectively filters data based on specific constraints, Ghosh warns of silent hallucinations where the AI provides incorrect information with high confidence. Ultimately, the source emphasizes that while this persistent reasoning is powerful, human oversight remains essential to verify the agent's autonomous decisions.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Testing whether open-source agents like Hermes truly have persistent memory and reasoning or just a convincing illusion of competence.
Benefits
  • Persistent vector profile remembers preferences across separate conversations
  • Dynamically recalculates discarded options when user constraints change midstream
  • Chain-of-thought reasoning catches internal contradictions keyword scanners miss
  • Open source lets developers inspect and modify the architecture
  • Writes its own reusable skills and runs on its own schedule
Use cases
  • Travel-agent test: picked the only package meeting vegetarian, <$200/night, no-flights-before-10am rules
  • Re-surfaced rejected option D after budget rose to $225 and flights allowed after 8am
  • Screened four resumes for Nimbus Widgets backend role on skills, experience, consistency
  • Flagged 'Casey Time Warp' claiming 8 years experience but dates showing ~2.5 years
KPIs / results
  • Travel constraints: <$200/night budget, no flights before 10am
  • Resume rules: 4 years minimum experience; Python, SQL, REST APIs required
  • Caught 8-year claim contradicted by Jan 2024–Aug 2026 dates (~2.5 years)
Tools / build
  • Hermes by Nous Research
  • Vector-profile persistent memory
  • Chain-of-thought resume consistency checker
  • Sam Rivera travel-agent test persona
0:00 / 0:00
Welcome to the Deep Dive. Today our mission is to figure out if Open-Source AI agents are actually ready to run our lives and businesses. Right, or if they're just presenting a very convincing illusion of competence. Exactly. I mean we really want to get past the hype here. Yeah. And we are pulling our roadmap for this from a really fantastic Spotlight series on the AI with Arun Show on YouTube. Yeah, that was such a good breakdown. It really was. They featured a UC Berkeley student named Sinek Ghosh. And he basically decided to look past the glossy tech brochures and just tested himself. Right. He built and rigorously tested a dual purpose AI agent. He wanted to see exactly where it succeeds in the real world and where it, you know, dangerously falls apart. Yeah, dangerously being the key word there. So to set the stage for what we are looking at today, I want you to think about what it's like when you hire a brand new assistant. Oh yeah, the honeymoon phase. Yes, you have that honeymoon phase. Yeah. They sit at their desk writing down every single thing you say in a fresh little notebook. Right. It's the meticulous note-taking stage. You say you like afternoon meetings, they underline it twice. Exactly. It feels like you are finally understood. Yeah, it's great. But then week two rolls around and suddenly they schedule your most critical strategy session for 7am on a Sunday. Oh man, yeah. Yeah. And when you ask them why, they look you dead in the eye with unwavering confidence and say, because you love early mornings. Right. The notebook is gone, but the absolute confidence remains. Exactly. And that tension, that exact tension between incredible capability and confident silent failure is the core of what Sainit Ghosh was testing on the AI with Arun show. It really is. I mean, we are looking at a fundamental shift in technology here. We are moving away from systems that just chat with us to autonomous systems that make decisions on our behalf. Right. Which is a massive leap. Yeah. So the subject of Ghosh's test is an open source AI agent called Hermes. Yeah, Hermes. It was built by a company called Noose Research. And their tagline for Hermes is the agent that grows with you. Which is quite the claim. Yeah, I have to pause right there because grows with you sounds like great marketing copy. But what does it actually mean? Well. Because when I use a standard chat bot, it basically has amnesia. Like every time I open a new window, I'm a complete stranger to it. Right. I have to re-explain my preferences, the context of my project, literally everything. Yeah. And that is because standard models are largely stateless by design. Okay. They process your immediate prompt, they give you an output, and then they basically clear the board. Right. They wipe the slate clean. Exactly. But Hermes is built differently. Because it is open source, developers like Ghosh can actually look under the hood and modify its architecture. They aren't just pinging some closed off black box server. Right. So Hermes is designed to maintain persistent memory. It actively builds a vector profile of you across multiple separate conversations. Okay, wait. Let me push back on that for a second. Sure. When you say it maintains a vector profile, how is that practically different from, I don't know, my web browser remembering my passwords or an app just saving my settings in a database? That is a great question. Because it is not just a database of saved text. Okay. A traditional database is very widget. Like if you save the word vegetarian in a profile, the software just looks for a true or false flag on a menu. Right. It's just a simple binary thing. Exactly. But a vector profile means the AI is mapping your preferences conceptually. Conceptually. Yeah. It understands the semantic relationship between your past conversations. And furthermore, it runs on its own schedule in the background. Oh, wow. Yeah. And it can write its own reusable skills. So it isn't just saving your data. It is supposed to be actively refining how it interprets your requests over time. Okay. Let's unpack this. Because Saini Ghosh didn't just accept that premise, right? No, not at all. He had the reaction any critical thinker should have when a tech company says their product thinks. Right. He was highly skeptical. Yeah. He assumed they probably just bolted a traditional database onto a chatbot and called it a day. Which is a fair assumption. Exactly. So he spent weeks actively trying to break the system. Which I love because that is the only valid methodology for testing AI. You don't test it by trying to confirm the marketing. You test it by hunting for the breaking point. Yeah. You have to try and break it. So for his first test case, he set Hermes up as a personal travel agent. Right. And he created this highly specific fictional persona named Sam Rivera. He wanted to see if the AI could handle complex overlapping constraints. And the constraints were very strict. Yeah. Break those down for us. So Sam Rivera has three hard rules for travel. First, a vegetarian diet. Okay. Second, a strict mid-range budget capped at less than $200 a night. Got it. And third, absolutely no flights before 10 a.m. All right. So vegetarian under $200 and no early flights. So Hermes goes out and pulls together four different travel packages. And the breakdown of how it handles these options is really revealing about how its logic works. It really is. So option A is basically the golden ticket. The flight departs at 1025 a.m. The hotel is under the $200 limit and it includes a vegetarian breakfast. Perfect. Yeah. And Hermes correctly identifies this as the only valid choice right out of the gate. And just as importantly, it successfully rejects the others based on those specific constraints. Right. So look at option B. It has the vegetarian breakfast and it meets the price requirement, but the flight departs at 615 a.m. Oh, way too early. Right. And the AI recognizes that 615 a.m. is earlier than 10 a.m. and flags it as a conflict. Okay. And option C is just a total miss across the board. Yeah. It's at a place called The Grill House. There is no vegetarian breakfast and it costs $210, which blows past the budget. Right. So that one is easily out. But option D is where the test actually starts to matter. Exactly. Because option D features a flight at 9, 5 a.m., which breaks the time rule. And it costs $215, which breaks the budget rule. But it does have a vegetarian breakfast. Yes. Now initially, Hermes correctly files option D away as noncompliant. Right. And in a traditional software program, option D is now dead. It's gone. Yeah. It failed the logic case, so it gets discarded from the active memory queue completely. But this is where Gauche introduces the plot twist on the show. He wants to see if the AI is actually thinking or if it's just filtering. Right. So he acts as the user, Sam Rivera, and he changes the rules midstream. He tells the agent, hey, we have some extra money. Let's bump the budget up to $225 a night. And I can wake up earlier. Let's say no flights before 8 a.m. And this is a massive stress test for an autonomous agent. Right. Because the rules just changed entirely. Exactly. If it is just a traditional database, changing those parameters means you have to run a brand new search query from scratch. Because it wouldn't remember the discarded options. All right. The system wouldn't have them in active memory anymore. But it doesn't run a new search. No, it doesn't. Instantly, Hermes remembers the previous options it had already evaluated and dynamically recalculates everything based on the new context. It's incredible. Yeah. It pulls option D out of the discard pile and presents it to the user as a newly compliant choice. Because under the new rules? Right. Under the new $225 budget and flights after 8 a.m., option D's $215 price tag and 9-0-5 a.m. flight suddenly make it a winner. What's fascinating here is the underlying mechanism of how it does that. Yeah. How does it actually retrieve that without searching again? Well, the agent didn't just append a new text file to its memory. It reprompted its own internal logic. Okay. It looked at its previous vector space, you know, the conceptual map of those four hotels. Right. And it applied the new semantic rules to the old data without needing human intervention to bridge the gap. Wow. Yeah. That is the definition of an agentic workflow versus a static search. Which is incredibly impressive for booking a vacation. Like in the AI with Arun Show video, this is the moment Gush admits the memory claims are actually real. Yeah. The marketing wasn't totally fake. Exactly. But booking a hotel is fundamentally low stakes, right? Very low stakes. Like if the agent fails, maybe you eat a non-vegetarian muffin or you lose 20 bucks. Right. It's an inconvenience. But what happens when human livelihoods are on the line? If it can dynamically change its own rules midstream, how do you control it in a corporate environment? Right. And that is the necessary pivot. We have to look at business verticals where the margin for error is essentially non-existent. Yeah. You do not want an AI dynamically updating its interpretation of compliance laws or hiring practices on the fly. No, definitely not. So for his second test, Gush switches Hermes into an enterprise mode. He tasks it with screening resumes for back-end engineers for a fictional company called Nimbus Widgets Incorporated. Right. And this is where the design constraints become so critical. Yeah, absolutely critical. Because to avoid the massive liabilities associated with automated HR screening, the agent was explicitly programmed with strict objective logic requirements. Okay, so they locked it down. Yes. It was hard-coded to look at only three things. Required skills, years of experience, and the internal consistency of the resume's claims. Got it. It was heavily restricted from evaluating demographics, personal attributes, or subjective concepts like cultural fit. Because you wanted acting as a pure logic engine. Right. Not a judge of character. Exactly. Okay, so the hard requirements for the job at Nimbus Widgets are a minimum of four years of industry experience. Right. Required skills are Python, SQL, and REST APIs. Yeah. And preferred skills are FastAPI, AWS, and Docker. So let's look at how Hermes processes the four candidates here, because it really highlights the difference between traditional applicant tracking systems. Oh, the ATS software most people know. Right, ATS software. It highlights the difference between that and semantic reasoning. Traditional ATS is basically a keyword scanner. If it doesn't see the exact word, it rejects you. Exactly. Wait, let me clarify that real quick. If a traditional APS is just looking for keywords, how does Hermes approach candidate one? Candidate one is Alex Strongfit. Right, Alex Strongfit. Alex has six years of experience and all the required and preferred skills. So for Alex, the result actually looks the same. Hermes clears Alex as a clean match. Okay. A traditional ATS would also clear Alex because all the keywords are present on the page. Right, that makes sense. The divergence really happens with the imperfect candidates. Like candidate two. Yeah. Take candidate two, Jordan NoSQL. Jordan has five years of experience, which clears the time hurdle. Mm-hmm. They have Python, REST APIs, Fast API, and Docker. But they're missing SQL. Right. And Hermes accurately flags the missing required skill and puts Jordan in the needs review pile. Which is what it should do. Yeah. Then there's candidate three, Riley Jr. Oh yeah, Riley. Riley has all the technical skills, Python, SQL, REST APIs. But Riley only has two years of experience. Right. And the requirement was four. Exactly. So Hermes catches the time deficit immediately and sends Riley to the review pile too. Okay. But candidate four is where this entire deep dive hinges. Yes, it is. Candidate four is Casey Time Warp. And here's where it gets really interesting. Well, I love this part. If you feed Casey's resume into a traditional keyword scanner, Casey gets an interview tomorrow. Oh, 100%. Casey has all the required skills. And prominently displayed right on the resume is the bold claim, lead back end engineer, eight years experience. Right. And a traditional system just sees the number eight, verifies it is greater than the requirement of four, and checks the box. It just moves them right along. But Hermes is tasked with checking internal consistency. Right. So further down on the resume, the actual lifted employment dates for that lead back end engineer role are January 2024 to August 2026. Now, in the context of the timeline used in the source material, August 2026 is the present day. So Hermes looks at the claim of eight years, then it looks at the dates January 2024 to August 2026. And it realizes that time span is roughly two and a half years. Right. It completely invalidates the eight-year claim and flags the candidate for a major internal contradiction. Which is a profound technological moment. Hold on. I need to understand how it's doing this. Okay. Because large language models are fundamentally text predictors. Right. At their core. They predict the next most likely word based on patterns. So how is a word predictor looking at January 2024 to August 2026 and doing the mathematical subtraction to realize it doesn't equal eight years? That is the crucial difference between a basic chat model and an agent utilizing chain of thought reasoning. Chain of thought reasoning. Okay. Right. When Hermes encounters the dates, it doesn't just predict the next word. It is prompted to break the problem into steps. Interesting. Yeah. It converts January 2024 into a numerical timestamp concept. It does the exact same thing for August 2026. It then compares the distance between those two concepts against the semantic meaning of eight years. Oh, wow. Yeah. So it isn't using a calculator. It is using linguistic logic to determine that the sequence of months listed cannot semantically align with the phrase eight years. That is wild. I mean, think about a tired HR rep scanning their 400th resume of the week. Oh, they'd miss it for sure. Right. Because you see the right skills. They see the big eight years experience header. They are very likely to just push that through to the interview stage. Absolutely. But the AI caught the lie or the typo by actually cross-referencing the concepts within the document. Yes. It didn't just read the resume. It audited it. And it did it in seconds. Right. So I'm watching the AI with Arun show. I see this happen and I'm thinking, this is it. The logic engine is flawless. Yeah. It feels like magic. It really does. But then Santa Gauche drops the hammer. He introduces the massive unsettling flaw he found while running these loops. Right. Because the reality is Hermes makes mistakes. And the problem isn't just the error rate. It is the nature of the errors. Okay. So if it's smart enough to cross-reference dates using semantic timestamps, how does it fail? It suffers from a phenomenon known in the industry as silent hallucinations. Silent hallucinations. Yes. We need to go back to what you said a minute ago about these models being probabilistic token generators. Right. Predicting words. Exactly. They do not have a native concept of truth. They calculate statistical likelihoods. Okay. So a silent hallucination occurs when the agent generates a false piece of information, accepts that false information as an absolute fact, and then takes an autonomous action based on it. Oh, no. Yeah. And it does it without ever alerting the user that it made a logical leap. So wait, you're telling me it can flawlessly deduce that 2024 to 26 isn't eight years, but it might just randomly invent a reason to reject a candidate? Precisely. Yeah. Because of how the neural network fires, if it encounters an ambiguity in the text, it might probabilistically generate a fact to resolve that ambiguity. Oh, my gosh. Right. And because it's an agent designed to run autonomously, it doesn't stop to ask you for clarification. It just keeps going. It just confidently proceeds with the hallucinated fact as its new foundation. But that completely destroys the utility of the tool. It's a huge problem. Right. I mean, if I'm using a normal chatbot to get a recipe and it tells me to add a cup of salt to my chocolate chip cookies, I can see the text on the screen, recognize it's a hallucination, and just not do it. Right. You catch it easily. But if an agent is running in the background of my business, screening hundreds of applicants, and it silently hallucinates that a candidate doesn't have a required skill, how can any business trust it? They can't, honestly. You can't legally defend your hiring practices if your AI secretly disqualified someone based on a ghost metric it just invented out of thin air. You can't defend it, which is why deploying these systems autonomously is currently a massive liability. Gosh recognized this immediately on the show. He admitted that while he was specifically hunting for errors, he almost missed some of these silent hallucinations because the AI's output was so confidently formatted. Because it looks right. Exactly. Exactly. A normal user operating under the illusion of the AI's competence would be completely blind to the failures. So how did he fix it? Because the video doesn't just end with him saying that tech is fundamentally broken. No, it doesn't. He had to build a strict fail-safe architecture. Okay. What does that look like? Well, first, he created a routing system that forces the agent to serve a known good response from a controlled database if its internal confidence score drops below a certain threshold. So it's not allowed to guess. Right. Rather than letting it guess, it pulls from verified data. Yeah. But the true fail-safe is procedural. Procedural. Yeah. He instituted a mandatory, non-negotiable, human-in-the-loop review step for every single critical action. Oh, I see. So every resume it processes still requires a human to sign off. The human recruiter still has to look at KC Time Warp's file and say, okay, the AI flagged this for a date discrepancy? Let me verify. Yes, the dates are 2024 to 2026. Good catch. Exactly. So the AI basically acts as the very fast spotlight. Yeah. But the human is the only one allowed to make the final decision. Right. And the brutal irony of this current generation of AI agents is that the very trait that makes them valuable... Their autonomy. Yes. Their ability to operate autonomously and make independent semantic leaps. That is the exact mechanism that makes them dangerous to trust blindly. Wow. It brings us right back to the eager assistant from the beginning of the show. It really does. They can schedule the perfect afternoon meeting based on your complex preferences, or they can confidently book a strategy session for 3 a.m. on a Sunday. Yep. And they will feel the exact same level of certainty about both of those decisions. Which means your deployment strategy is the only thing that separates a massive efficiency gain from a catastrophic liability. Gaush's ultimate takeaway is that you simply do not let these agents run wild. You pilot them in isolated scoped environments and you verify every single output they produce. So what does this all mean for the listener? Yeah. What's the bottom line? Well, if you are trying to figure out how to integrate these open source agents into your workflow, the takeaway is pretty clear. The technology is genuinely powerful. It really is. The dynamic memory works. The semantic reasoning can catch things that traditional software misses entirely. But you absolutely cannot hand them the keys to the kingdom yet. No, definitely not. Treat them like a highly capable, incredibly fast, but occasionally delusional intern. That's a perfect way to describe it. They can do the heavy lifting of sorting the data, but their work requires meticulous auditing. And that reality introduces a fascinating friction point that every industry is going to have to confront very soon. Yeah. What's that? Well, if we are required to maintain a human in the loop for every decision, right? If we have to meticulously audit an AI to catch these silent hallucinations, at what point does untangling the AI's complex logic actually require more time and cognitive energy than just doing the screening ourselves in the first place? Wow. That is the perfect question to leave hanging in the air. It's a tough one to answer. It really is. If you have to spend two hours figuring out why your AI assistant confidently booked a meeting at 3 a.m., you haven't actually saved any time. No, not at all. You've just traded the work of doing a PASC for the work of untangling a very sophisticated mistake. Exactly. Well, we will leave you all to think about that one. Until next time.