← Back to search

159 - Understanding Agents: Self-Optimization

Prompt und Antwort · 2026-05-08 · 23 min
relevance 52 3675 words Episode page ↗ Audio ↗
Show full episode description
KI-Gilde Podcast 159: Agenten verstehen GEPA – Wie KI ihre eigenen Regeln umschreibt Wie verbessert sich eine KI eigentlich selbst? In dieser Folge entzaubern wir den Mythos der autonomen KI-Evolution am Beispiel des neuen Hermes-Agenten-Updates. Das Geheimnis dahinter ist GEPA (Genetic Pareto Prompt Evolution).Statt blind Parameter auf teuren Serverfarmen anzupassen, lernt die KI hier durch die sprachliche Reflexion ihrer eigenen Fehler und schreibt ihre Arbeitsanweisungen (Skill Files) einfach selbst als Textdokumente um. Wir erklären anschaulich, wie die drei Säulen von GEPA funktionieren: Reflektierende Veränderung Pareto-Selektion für mehr Vielfalt Evolutionäre Stammbäume Erfahre, warum dieser Ansatz mit 2 bis 10 US-Dollar pro Lauf extrem günstig und durch einfache Textdateien vollkommen transparent ist. Außerdem klären wir, warum unvollständige Testdaten die größte Gefahr (Goodharts Gesetz) bei dieser Methode bergen und stellen die provokante Frage: Ist der hochbezahlte Beruf des Prompt Engineers bald Geschichte?
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How AI agents self-improve via reflective prompt evolution (GEPA) instead of costly weight retraining.
Benefits
  • Massively cheaper optimization than reinforcement learning
  • Learns from language error logs, not blind weight shifts
  • Pareto selection preserves solution diversity, avoids local optima
  • Automates expert prompt-engineering knowledge
  • Transparent, explainable improvements vs black-box weights
Use cases
  • GEPA matched GRPO's top result using only 32 rollouts vs 24,000 — a 750x efficiency gain
  • Pupa data-privacy prompt evolved from 82 points (gen 0) to 97 points (gen 11)
  • HotpotQA multi-hop search prompts rewritten to find logically linked concepts
  • Reflective approach beat traditional methods by 6% average, up to 20% at peak
  • Beat leading prompt-optimization methods by over 10% on average
KPIs / results
  • 32 rollouts vs 24,000 (750x efficiency)
  • Pupa scores: 82 → 87 → 91 → 97 across generations 0/3/7/11
  • 6% average gain, 20% peak gain
  • >10% better than leading prompt optimizers
Tools / build
  • Hermes agent (self-optimization update)
  • GEPA (Genetic Pareto Prompt Evolution)
  • GRPO reinforcement learning
  • MIPRO V2 prompt optimization
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
This podcast is a project of the AI ​​Guild. The content and voices are generated with AI and are used for information and demonstration. Have fun listening! When any company in the tech world makes big announcements today, like we now have an AI that improves itself, all of my science fiction alarm bells immediately ring. Oh yes, absolutely. That immediately resonates with this whole Hollywood fantasy, doesn't it? Total. That always sounds like pure marketing. In other words, a system that secretly develops its own consciousness on servers at night and wakes up smarter the next morning. Um, yes. Terminator sends his regards. But this is exactly the promise that the new update from this so-called Hermes agent has put on the table for us. And now in the spring of 2026. Exactly that. But if you look under the hood, the reality behind it is much more sober. More sober. Yes, but from a technical point of view it's actually almost more awesome than any science fiction fantasy. And that is exactly our goal for this mental dive today. If you're wondering how AI will actually learn in the future without anyone building new server farms, we want to unravel that today. Mmm. Step by step. The engine behind this supposedly magical self-improvement is a tangible concept called GEPA. So let’s take a closer look at the paper from 2025. This concept is already described there. By the way, GEPA stands for Genetic Pareto Prompt Evolution. Just in case anyone was wondering. Thanks for the classification. But before we understand what makes GEPA so revolutionary, we need to briefly clarify why the previous path was so incredibly painful. That's a good point. Today, modern AI systems are rarely just a single chat window into which you type something. We are talking about composite systems. So systems that consist of several modules. Exactly. It's like a pipeline. One module formulates a search query, another searches the Internet, and a third calls external tools. Like a calculator or something? Exactly. And at the end, another module puts together the final answer from all these scraps of information. It's basically like a small digital factory. That's a great comparison. And if you end up with a defective product on the assembly line, you have to somehow find out at which station in the factory the error occurred. And how has this been done so far? So if you wanted to optimize this factory? Traditionally there were two ways. And both have massive blind spots. The first way is classic reinforcement learning, i.e. reinforcing learning. There is this GRPO procedure, right? Correct. GPRPO. You let the system solve tens of thousands of tasks. Every single attempt by the AI ​​to chase a problem through this factory from start to finish is called a rollout. So a complete process from the first click to the final answer. Exactly. And with GPRPO you only look at the end result after every rollout. So was the answer right or wrong? Pass or fail? Does that mean all the intermediate steps don’t matter? Yes, completely. Based on this simple yes or no, the algorithm then goes deep into the AI's neural networks and adjusts the mathematical weights. Those tiny probabilities with which words are strung together. Exactly. You blindly change a few numbers in the background and hope that the system happens to take the right path during the next rollout. That sounds like an extremely inefficient trial and error process, right? It is too. You need huge GPU clusters for this. It takes forever and in the end you've slightly shifted billions of incomprehensible model weights. But no one knows why anymore. Exactly. No one in the world can explain to you in normal language what the system has actually learned. An absolute black box. But you said there was a second way. Yes, the prompt optimization. Known under the name Mipro V2. The mathematical weights in the AI's brain are left completely alone. Okay, that sounds safer. It is too. Instead, you systematically vary the prompts, i.e. the pure text work instructions that you give the AI. And then you just test thousands of variants? Exactly. The computer tries out slightly different instructions and simply measures which variant ultimately brings the highest success rate, i.e. the highest score. But wait a minute. If I understand correctly, we have exactly the same fundamental problem, right? What are you referring to? Well, in the end I only have this one metric. This one score. I completely throw the entire AI solution into the trash. Ah yes, you are absolutely right. It's as if my driving instructor was sitting completely silent in the passenger seat for the entire driving lesson. He just secretly takes notes on his clipboard. Exactly. And he doesn't say a single word. And then when I turn off the engine after 45 minutes, it just calls out and gets out. A terrible driving instructor. Yes, totally. I then have no idea whether I didn't stop correctly at the stop sign, whether I didn't look over my shoulder or whether I hit the curb when parking. The entire so-called reasoning trace, i.e. the AI's chain of reasoning, is simply thrown away. The tool calls, the error messages in between, none of that matters. This is a gigantic waste of information. And it was precisely this frustration that was the absolute spark for the development of GPAR. Researchers have realized that language is a much denser and richer medium for correcting errors than simply a numerical value. Makes sense. If the system is already extremely good at understanding text without it, why don't we use it? Exactly. Why don't we let him think about his own mistakes? So we switch from bluntly moving numbers to actively reflecting in language. How should I specifically imagine the mechanism behind it? The whole thing is based on three pillars. The first is called reflective mutation. Okay, what's happening? Instead of simply exchanging prompts blindly, the AI ​​is presented with its own log files. She sees her entire thought process in text form. So something like, here I queried the database, this was the result and then I drew this wrong conclusion. Exactly. And then a so-called reflection model comes into play. Almost a separate analytical coach. This coach then reads through these text protocols. Yes. And it specifically looks for the moment when the AI ​​took a wrong turn. He diagnoses the error linguistically and then specifically rewrites the prompt for the next attempt. Do you have an example to make this tangible? Clear. Let’s look at the Hotpot QA example from research. These are tasks where the AI ​​has to think about multiple corners. Mmm. So multi-level searches. Exactly. Imagine the original command is simply, generate a new, more in-depth search query from the original question and the first summary. This sounds like a very classic standard prompt. Millions of people probably use it in a similar way every day. And it is precisely because of such prompts that systems often fail. During the second search, the AI ​​usually simply phrases the original question slightly differently. Which is completely pointless because it doesn't bring any new facts. Correct. And this is where Jeeper intervenes. The coach model reads the error log. It looks like, hey, we just paraphrased the question. That didn't get us anywhere. And then the coach rewrites the prompt. Completely independent. He then writes in, be careful, the second search query cannot simply be a repetition. You have to look for logically linked concepts that were missing in the first step. Craziness. Yes. And he might also add that if the first search describes a small community, but the original question is aimed at an entire region, then now search specifically for higher-level information. I have to comment on that for a moment. That's phenomenal. Here, the system extracts deep, domain-specific knowledge from its own failures. Yes, it recognizes systematic failure and actively formulates a rule against it. This is exactly the moment when companies normally fly in expensive prompt engineers. Correct. They then spend days refining the instructions. And here it just happens fully automatically. It is the automation of expert knowledge. But, and a thought immediately comes to mind, doesn't that also pose a danger? To what extent? Well, if the system only looks at the prompt that worked best in the last run and then tightens it further, it won't get completely lost. Ah, I see what you're getting at. If I build a machine that only ever looks for the one perfect solution to a very specific problem, then I'm breeding an absolute idiot, right? She then forgets everything else. This is the infamous local optimum problem. If you only ever take the emissary winner from the last round, evolutionary diversity dies. You are then trapped on a small hill and never reach the actual mountain top. Very nice picture. And to prevent exactly that, the concept uses its second pillar. The Pareto selection. Pareto? Does this refer to the Pareto principle, the classic 80-20 rule from business administration? It comes from the same mathematical family, yes. But what we mean here is the so-called Pareto front. What does that mean in this context? The algorithm doesn't simply throw away all prompts that perform worse overall. Okay, then who does he keep? It keeps every single prompt in its gene pool that was the undisputed winner in at least one very specific task. Ah, I see. Even if a prompt completely fails in 99 tasks and produces terrific nonsense, but finds an absolutely brilliant, unconventional solution in task number 100, then it can survive. Exactly. Because this very niche prompt may have discovered a unique strategy that could be incredibly valuable in future generations. Diversity beats greed, you might say. Exactly, that's the motto. It's not about greedily chasing the highest average score, but rather keeping a broad catalog of solution approaches alive. Okay, that solves the idiot problem. But how does knowledge build up over time? That brings us to the third pillar. When we talk about generations and evolution, how does a language model actually inherit knowledge? In biology we have DNA. What is the DNA of a prompt? The DNA in this case is a historical text log file. And this third pillar is called the Genetic Tree. So the prompts don't mutate in isolation in a vacuum? No. When the reflection model generates a new prompt, it doesn't just see the very last attempt. It reads a summary of all evolutionary history before it. So it looks like, ah, in Generation 1 we tried that, that swallows fairly. Exactly. And in Generation 2 we added Strategy B. That helped a little. So the system layers lessons on top of each other. Do you have a concrete example of this? How does such a prompt grow? A very impressive example is a task on the subject of data protection. The so-called Pupa example. The AI ​​should learn to formulate database queries in such a way that sensitive customer data is never leaked to the outside world. OK. Important topic. How does this start? In the very first generation, i.e. generation 0, the success rate is 82 points. The basic prompt is simply, create a privacy-friendly request. Pretty flat. That's like me saying to an employee, do your job well. Correct. But in Generation 3 the score rose to 87. The prompt is now significantly longer. What's in there now? It contains specific rules for recognizing email addresses and phone numbers. And then let's jump to generation 7. Here we are at 91 points. And the text continues to mutate? Yes. Here the AI ​​suddenly demands very structured justifications from itself for each individual filter step. And it has built in explicit bans on sharing real names. That's crazy. And where does it all end? In generation 11 the system reaches 97 points. The prompt has grown into a detailed manual. So a real set of rules? Completely. It now requires a strict, step-by-step protocol, understandable justifications for each individual data transfer and an absolute zero-tolerance policy for data leaks. So the system has stacked new protective mechanisms layer by layer over eleven generations like a bricklayer? Until a waterproof wall is created. And all because it read its own error logs. That must also have a dramatic impact on efficiency, right? You mentioned earlier that the old methods require huge GPU clusters. The numbers are actually the point at which the industry has now really woken up. If we compare GEPA with the extremely computationally intensive method GRPO... The method that requires tens of thousands of rollouts to only call pass or fail in the end. Exactly. On a standard test for AI systems, GEPA achieves exactly the same top result as GRPO. But with less effort? With massively less effort. While GRPO needed 24,000 rollouts, GEPA managed it with only 32 rollouts. Wait, 32 attempts instead of 24,000? Yes, let that sink in. That's an increase in efficiency by a factor of 750. That's not just a little better, it's a completely different league. On average across all tested tasks, the reflective approach beats traditional methods by 6 percent. At its peak even by 20 percent. And that with a fraction of the computing power. Craziness. Even the leading methods for prompt optimization are beaten by an average of over 10 percent. OK. Now let's put all this crazy theory into practice and get back to our starting point. To the Hermes agent. Exactly. How is this technology now used by the Hermes agent in the real world? If the system learns through reflection, does that happen live? What do you mean? Well, will I witness the AI ​​getting smarter while I chat with it? Uh. This is a common misconception. This type of self-improvement never happens live while you are interacting with the system. Why not? Wouldn't that be practical? That would be a conceptual nightmare. Imagine you are working with a digital assistant and in the middle of your project it suddenly changes its basic rules of behavior because it has mutated in real time. Ah, okay. So a system that simply remembers that I drink my coffee black isn't self-improvement in that sense? No, this is simple memory, i.e. memory, the storing of facts. Self-improvement through GEPA is a purely analytical offline process. Okay, offline. But what exactly is being optimized in the background if not the model itself? The so-called skill files are optimized. Uh, skill files? What is that exactly? These are simple text files, usually in Markdown format, in which the agent's behavior rules, skills and tools are defined in plain text. So the AI ​​rewrites pure text documents? Exactly. Not their own neural networks. The mathematical weights of the model remain completely untouched. Okay, let me translate that into an image. So it's not like the AI ​​is performing neurological surgery on itself in the server room at night in order to have a higher IQ the next day. No, not at all. It's more like an extremely conscientious employee reading through the operations manual for the morning shift after work, when day-to-day business is at a standstill. Mmm. He analyzes where his colleagues made mistakes today and then rewrites the process description in such detail that everything runs more smoothly the next day. This is an absolutely apt image. It's about the ongoing optimization of textual work instructions. And because everything takes place in text format, there are also hard guardrails. So the system doesn't just overwrite these manuals in secret? Exactly. It submits its suggestions for improvement as a so-called pull request. Ah, like human software developers do. Correct. So there is a bouncer. A human reads the proposed changes in the text file, reviews the logic, and must actively approve them before they are committed to the live system. Of course, this guarantees full control. And perhaps most importantly for practice, it is incredibly cheap. What does that mean in numbers? A complete evolutionary optimization pass for such a capability only costs between $2 and $10 per pass in pure API costs. $2 to $10. We were just talking about weeks of training on huge computer machines, which costs tens of thousands of dollars. Yes. This completely redefines the rules of the game for the economy. Especially for companies that are not huge tech giants. This is an enormous lever for medium-sized businesses. A gigantic one. Firstly, you no longer need expensive hardware. You don't need highly specialized engineers to train the core of the model. And secondly? Secondly, you are completely independent of the provider. Since only text files are optimized here, you can apply this perfected prompt to a model from provider A today. And tomorrow I'll just switch to the latest model from provider B. Exactly. This is called vendor independence. And what about traceability? We had already discovered that in classic training, in the end no one knows what the billion parameters actually do. This is the third crucial point. Auditability. Especially for industries that are subject to strict compliance rules. So you can always prove what the AI ​​is doing and why. Yes, you can always read what the AI ​​has learned because it is written in simple, human-readable text. So a compliance officer can see in black and white, okay, the AI ​​added the sentence in rule number 4, always match the account number twice. Exactly that. And he can then approve it. With a black box of probabilities, this is completely impossible. When I listen to all this, $2 per run, full transparency, no server farm, complete independence, I inevitably have to put the brakes on it. You're looking for the catch. Yes. Where is the trap? That just sounds too perfect. There is a catch. And it is deeply rooted in how evolutionary algorithms work. This brings us to a concept known as Good Hearts Law. Good Heart's Law. What does that mean? This law states that when a measure becomes a target, it ceases to be a good measure. OK. And how does this manifest itself in prompt evolution? This algorithm optimizes mercilessly and highly efficiently. But only what is measured. Explain. The reflection model evaluates the errors based on a test data set that you give it. Let's think this through. You build a digital customer advisor for an online shop. OK. Your test data set meticulously checks whether the AI ​​correctly reflects the facts about the return and whether it solves the customer's problem quickly. Then the AI ​​will probably become extremely good at reciting these return conditions without errors. Yes. But if you don't measure friendly tone of voice and include it as a criterion in your test data set, then evolution will simply cut out this aspect over the generations. Because being friendly doesn't give you a score bonus. Exactly. You then get an AI that provides all the facts in record time, but sounds like an incredibly rude, cold bureaucrat. It will no longer greet the user at all because, according to the test data set, this is an unnecessary waste of time. Exactly. So the big challenge is shifting. It's no longer a matter of laboriously training the AI ​​or writing the prompt yourself. No. The new, complex task is to curate perfect, balanced test sets. Test sets that represent all soft and hard factors. Because if my test set has a blind spot... then this algorithm will find it, exploit it and amplify it mercilessly. The engineer's work therefore shifts from writing the rules to defining the testing criteria. Wow. If we pull all these threads together now, a very clear picture emerges. We have seen today that the much-advertised, self-improving AI is not a mystical process. Definitely not. It is a highly pragmatic, understandable mechanism. A process based on linguistic reflection and structured evolution… decency to blindly shift billions of parameters. And that not only makes these systems more tangible, but above all much faster, more transparent and customizable for a fraction of the cost. The initial science fiction promise has been demystified here. But what emerged underneath is a technical brilliance that really gives every company completely new tools. We are simply no longer dependent on someone writing us the perfect work instructions. The system does it itself. It is a true democratization of AI’s adaptability. Finally, a thought comes to mind that we should give the listener something to think about further. Shoot away. Today we have seen how exponentially better these systems are already becoming at analyzing their own logic errors. They layer new rules on top of each other, restructure their prompts and quickly outperform human engineers. If you think about that, the currently extremely celebrated and highly paid profession of prompt engineer may end up just being like the profession of elevator operator from the early 20th century. Oh, that's a harsh comparison. But think about it. A job that simply only exists until the machines understand how to push their own buttons. A very provocative thought. But after what we've discussed today, it's not that far-fetched. A thought that we would like to let sink in. Thank you for these deep insights today. Very gladly. Was fun. And to you out there. Until the next mental dive.