Dans cet épisode, Vincent Heuschling et Nicolas Steinmetz décortiquent Hermès Agent, le nouveau venu open source qui bouscule OpenClaw. Au programme : son architecture en cinq couches, sa mémoire persistante, et sa capacité à tourner en local pour préserver votre souveraineté. On parle aussi dans cet épisode du nouveau mode de facturation d'Anthropic, de cas d'usage concrets (daily briefing bot, revue de PR GitHub) et d'un sujet qu'on n'anticipe pas assez : la traçabilité et l'identité des agents qui se chaînent entre eux. ## Chapitres
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Understanding how Hermes Agent's structured harness rivals OpenClaw and commercial agent runtimes while keeping data sovereignty.
Benefits
Ships a pre-built five-layer harness instead of assembling everything yourself
Auto-writes skills from observed repetitive prompts and tool interactions
Persistent memory with ~10ms retrieval even over huge document volumes
Sandboxed permission and execution system limits runaway agent actions
Model-agnostic multi-model orchestration, supports self-hosted open models
Use cases
Agent reads a list of blogs and delivers a morning news summary autonomously
Start a conversation on Telegram/WhatsApp and continue it in the terminal, fully synchronized
Use GoCli to list and write Gmail as JSON/markdown without opening broad API access
Podcast production: agent connects recent news to past episodes for content ideas
Memory system as a RAG backend with ~10ms retrieval latency
KPIs / results
Hermes Agent near 100,000 GitHub stars since February 2026 release
Anthropic billing changes push interest toward self-hosted open models
00:26:00
Cas d'usage concrets et niveau de contrôle
~100 bundled skills; morning blog-digest agent built very quickly
Feature set comparable to AWS Bedrock Agent Core
Podcast production idea: agent links recent news to past episodes
🌐 This transcript was automatically translated to English from the original.
Hello everyone, hello everyone, welcome to a new episode of Big Data Hebdo. Big Data Hebdo is the French-speaking data and AI podcast. And in Big Data Hebdo, we try to enlighten you on everything that's moving in terms of technology in data and AI. And so, today, we're going to talk about agents, we're going to talk about AI agents, but we're going to talk about things a little more where we can get our hands on it, since we're going to talk about Hermes Agent. Hermes Agent is the open source agent of the moment. And for this episode, I am accompanied by Nicolas. Hi Nicolas, how are you? Hi Vincent! Well listen, it's going well, hotly, but it's going well. But there, it's okay, there are 6 of us in the morning, so it's okay. Nicolas, you are CTO at Cybeltech, which is therefore a company that is interested in everything that is life science and in supporting culture and everything that follows. Exactly. You presented everything well, nickel. Well yes, I took up what you said last time, you see, I didn't invent anything. Ah yes, you listen, you listen, it’s good. I listen to my interlocutors. And so, for my part, I am Vincent Heuschling, I am the co-founder of this podcast and I am an independent consultant. I support my clients in their Data EIA project to help them imagine, design and also implement solutions that could change their business. And so, let's go. We're going to get to the heart of the matter. In this episode, we finally said to ourselves, we had moved on a little bit, we haven't talked much here in the OpenClaw podcast. And then, there is a new kid who has arrived called Hermès. And Hermès positioned it a little differently. The principle always remains the same as a reminder. OpenClaw is what put Apple out of stock on Mac minis, since the principle was that we were running an agent orchestrator. That is to say, we consumed the resource, for example, anthropogenic to make the agent think. But all the orchestration, all the memorization, the interaction was done on a machine that you hosted at home or also at a cloud host. Moreover, I believe that it was Hostinger that even did a type of deployment for its environment where you could request to have an OpenClaw machine directly installed. And OpenClaw had an interaction mode which was as follows. In particular, you spoke on a Telegram channel with your agent and then your agent did things for you. So, we were in the dialogue of the human with his agent to do things. There have been quite a few things shown on the market. But Hermes Agent and Harry were a little bit different. So already, where OpenClaw was the emanation of a solitary actor, Steinberger, here OpenClaw is an open source lab called Nous Research and which are not necessarily forced to produce scientific papers to be recognized, nor to raise funds. So, in fact, we don't talk about them very much. But they built models. The Hermès 3 range of models which existed in 8 billion parameters up to 405 billion parameters, is indeed one of the super good open source models that you can use. So, and so the thesis that they had is say with an open LLM API and open source tools, you could deploy an agent that rivals commercial offerings. So, it's true that indeed, we could relate this to what Anthropik offers in particular with the ability to have agents who run on their own resources. And what you have to look at is that it was still released in February 2026. So, it's still very recent, this open, not open cloud, but this Hermès agent. And they are already close to 100,000 GitHub stars, I believe, while open cloud, much longer, has around 350,000 GitHub stars currently. So, perhaps these figures are a little bit, are more necessarily up to date, but it still gives you an idea of the traction that there may be in there. And so, it's still quite interesting because, casually, looking at how it was designed and how it works, can even, if you don't want to install agents, help you optimize the way in which you work with your cloud traditionally at home. Typically, one of the huge things that Hermès does is this productized harness engineering. So at some point you have a feedback and feedback loop. They say, if you do the same prompt five times or if you have a sequence of prompts, and so on, write a skill. And one of the great strengths of Hermes is that at the heart of the system, it will observe what you do. And after a certain time, if it sees repetitive actions, if it sees interactions with tools, things like that, it will itself write a skill to help you next time with that. And so that’s something that’s super interesting and that can also help you evolve in your practice. That's why I was saying, when I read and became interested in the subject, I realized that the way in which I could structure the cloud that is on my machine, ultimately, it could really progress thanks to that. So Hermès arrives where, therefore, with the others, with OpenClo and so on, it was necessary to put together almost everything. You had it all. He already arrives with a super well structured harness, which already allows you to do very strong things. And so if we look a little bit, if we break it down a little bit into five layers, the instruction, the constraints, that is to say how you are going to sandbox things, the feedback, how you are going to be able to evolve your system, the memory, since at the heart of all that, there is also the ability to have persistence and orchestration. How do you do a multi-agent, multi-model thing and so on? Hermès comes with super clear answers on this. So on the instructions, typically, when you are working manually, you will write your CloudMD or your AgentMD, well, whatever name you give it, it has its skill system which will auto-create, auto-update. Afterwards, there is this notion of constraint, to what extent you make hooks, you do things which will prevent you from doing certain actions. We saw that one of the things that had been highlighted about OpenClo was that there were people where OpenClo had gone, got carried away, it had started subscribing to additional services, because it had been asked to do things regardless of the way in which it could get there. Finally, there, functioning, there really is a system at the heart of it all. At Hermès, they decided to have a sandboxed permission and execution system. So it’s still really good from that point of view. So, the feedback loop, I talked about it just before, that is to say that where you have to do a manual review, take your prompts a little, put them back in a skills, something, so on, there, automatically, there is a learning loop which is quite automated. It just creates markdown files that you can then edit and modify. But yes, so I didn't look at to what extent, indeed, there was this very manual side of proposing a new thing, and then to what extent he could, behind it, he could have a workflow to be able to activate them or not. I didn't look at that. As for memory, until now you had to serialize the memory a little bit yourself and rewrite in contexts, in specific markdowns, rewrite things like that. There, you have a system, there is even a switch which allows all interactions to persist and which allows you to start interactions on a media. Typically, earlier, I was talking about this thing that was very fashionable, to have a conversation with a... with a... how to say... I'm going to do it with a telegram or a WhatsApp and finish it in the terminal of your machine. Well, for once, this is hyper synchronized since there is a tilt which allows you to have this continuous recording of what is happening. And then, in terms of orchestration, that's also where they made a great effort. This is because you have the possibility of synchronizing tasks, of making multi-model calls, of calling on different models since they are not bound to a specific type of model. We recall that OpenClaw had started and it had a falling out with Anthropik for being very heavily used with Anthropik models. And then Steinberger was recruited by OpenAI. And so, I didn't follow up afterwards, but it's very likely that the following releases of OpenClaw will be very typed things for use with the OpenAI ecosystem. And so, it’s still something. There is really a super interesting feature set from this point of view and which also gives very good ideas if you want to develop an agentic loop or something like that behind it. in particular, there is the memory system which allows us to have things and therefore to also be able to have a retrival latency. If you want to do a RAG, for example, the memory system will be able to really help you since you will be at 10 milliseconds even if you have a volume of documents that you have put in your system which is enormous. So, it's really something that will allow us to build scenarios and resolve lots of different things. And as I said, it's still a big open source alternative to market offers because when we look a little bit, we're listing things and when we look at what's there, for example, in AWS Bedrock Agent Core which is the AWS agentic runtime, we're pretty much on the same feature set. So, it’s really a super interesting system from that point of view. When I dug into this, I really found it super interesting. And so, well, it doesn't happen completely naked and on which you have to build everything since there are overall around a hundred skills which cover lots of different use cases. For example, it's really very, very quick to create something that reads you a list of blogs and gives you a summary every morning to read so you can keep up to date on what's happening in terms of information. And that's something, you'll be able to have it running by itself in a corner on a machine at home and which doesn't require any resources because we'll come to that later. It is very, very well constructed to be able to auto-ostrate models and not be dependent on an external provider. Are there things, Nico, that you have had the opportunity to look at a little, these autonomous agents? No, no more than that. The open-closed part, that is to say they quickly cooled me down in mode where they displayed more security alerts than features at a given moment. So I let it go a bit. It’s true that I also have a little trouble delegating things. I like to control everything. I'm having trouble with this side, just...Hump my back and it'll be fine. I admit that I have a little trouble with this. So I'm learning to relax a little, but not completely. And so, behind closed doors, I quickly let go of the matter. Then since everyone is on it, it's true that I have an inverse relationship. It's more popular. Yes, yes, you're like me. You are like me. You have a certain hatred for things that are too hyped at the moment. That's it. I'll wait until later. And it's true that at this moment, we see a lot of things. It's true that Hermès, we see him quite a bit. I'm going to visit Dolama's blog this morning. He was talking about Open Jarvis who also seems to be a bit in the same spirit. We feel that things are moving. Afterwards, it's true that the use cases, I have a little difficulty saying hey, take all my data and do what you want with it. This is the extreme case. And so, there is still a real subject of everything that is permission. Yeah, you have this sandbox topic which is a real, real, real topic. I remain extremely reluctant. Typically, for example, for things that would interact with my emails or with my Google environment, I use a command line tool called GoCli. It was Nicolas Martignol who showed this, which he had done precisely for his automation around DevOps. He had made quite a few videos on it. And he had actually shown that instead of using cloud stuff, integrated into the cloud, notably in Cowork, to read these emails which are painful, which do not progress, which go at a speed that is completely crazy. There is actually a small binary that can be used on the command line which allows you to output emails and list emails. It outputs that in a JSON format or a markdown, whatever. And then even writing emails and that allows us not to open API access directly to a tool on which we do not control what it will do with it. And there, for once, we are able to say well, wait, this thing, I give it these permissions and so on. And at least I have something I can turn off if something ever goes wrong. And so, it’s still a bit like that subject. So you say yes, you want to control everything. You know what they say. Delegation does not exclude control. Yes, no, but it's the first step and I have a little trouble trusting. Trust does not exclude control. Thank you for taking me back. No worries. No, but it's sure. No, but it’s true that it’s happening little by little. Then we see it with the evolution of the models and then the side... Yeah. But it's true that, for the moment, I can't necessarily find it in the use case. I would have to force myself to do it. I know that indeed, in our professions where we have a need to dissect a quantity of information, to process, even post-process a lot of things. For example, I am firmly convinced that in the production of the podcast, having it go around in a corner and process things, which tries to connect recent news to past episodes and so on, that would make it possible to produce a lot of interesting things. That, I am deeply convinced of. Furthermore, I am also deeply convinced. That was my idea for this episode, before coming across that, it was to come back, to talk again about precisely these tools, these semantic tools, these semantic models and so on, which are very closely linked to an agentic loop which would allow us to have a dialogue with the data, with the data environment. It is clear that these autonomous agents have an interest in being able to process this. We also say, well there you go, we say that agents are things capable of working on boring tasks, those that we, humans, no longer want to do at all and so on. Ok, well the thing in data on which we are still a little bit uncomfortable is the subject of documentation, of validating, of doing a bit of linter work on the data environments, of validating that we have put the comments on all the columns, of validating all these things, finally in quotes that we try to do in CI-CDs classically in code, I know that on data it is more tedious and I remain convinced that this, for once, agents who are circling in a corner, we would be able to do this data stewardship work quite well, in fact. So. Oh yes, I find that there are plenty of uses, after that it's more about a personal age, for the moment, I can't do it. The news at the moment when everyone, we feel that there is tension over the use of tokens, that is to say that that... Yeah, yeah. It's a good time to test which was a few months ago where you, as it was in open bar mode on your cloud subscriptions, OpenAI and others, it was easy. There, when we saw that you are going to pay with the token, it risks being a little less... So, that requires having a machine with a little power, that is, but if you are not doing interactive, if it is agents to whom you give the time to think and to be able to provide a response, typically, we will come to it just after, the variety of models, even with quantified models that you can run on your machine, it is still quite important, that is, since there is indeed the possibility of using local runtimes and possibly delegating some one-off tasks which would require a lot, a lot of diversity of performance and particular features, to send that to a remote model. But I remain convinced, the interest of these agents, where I never bought into the OpenClo story, was, yeah, ok, if it's to run something on a machine at home and in the end, eat the token from an external provider, what's the point, what? What's the point? So, at a given moment, I see the real interest in the subject of how I can run things entirely, locally and possibly say that by extension, you can have things where you also guarantee your sovereignty. For me, it still remains a subject on which I am super cautious, super vigilant and therefore, in quotes, I prefer to have a deep-sic model or a quen or a model like that which runs locally than to send everything outside. Yes, no, but the same, you just have to have the machine. You just have to have the machine. You just have to run a machine. And indeed, in the heat we are currently having, our offices are already overheated and therefore, it is indeed perhaps not the best thing to do. we completely agree. So, we often draw a parallel between Hermès and OpenClaw. We've been doing this for a while now. These are projects that are very simultaneous since, as I said, OpenClaw which was called CloudBot and then MoldBot made by Peter Steinberger. Is it the end of last year? Yes, that's it. It launched in November-December 2025. It went viral in January-February 2026. And today, Hermès Agent arrived in roughly the same time frame a little bit later. And so, where there is one that is simpler, OpenClaw, ultimately, it's not much. It's just a Control Plane. You have your agentic loop and then that's all there is and you integrate lots of things around it. Where Hermes is much more integrated and provides everything and in particular this mechanism that they call the learning loop and persistent memory. and so, it's really a gateway on one side and a full agent runtime on the other. However, it's always quite funny when there is competition between different tools in a market. There is a command called Hermès Claw Migrate which allows you to retrieve what you have already done in OpenClaw to switch. So, this kind of thing is always quite funny. What... We talked a little bit about this security. Hermès, by default, the architecture is designed to be safe. These are not a posteriori patches that allow you to correct CVEs that could be there. So, at a given moment, there is really a noticeable difference and that is perhaps what will make us more inclined to call it, to use it than to use its predecessor. So there really is this important point. And there is really one point that is super important which is the fact that Hermès is model agnostic. and that was still the first thing I looked at. It's the extent to which you are free to do what you want with it and use it how you want. So typically, you can obviously use cloud providers. So, it's going to be Noussportal to use the Hermès models that they use, that they offer. You can use your Anthropik account. So, we will come back knowing that this will change in the coming days since I believe that on June 15, Anthropik will roll out this new billing method since until now, precisely, OpenClaw is one of those things which has meant that the profitability of Anthropik in particular has been seriously undermined since remember at the start, you could have your subscription at 20 dollars and make API calls on the basis of this practically unlimited subscription at 20 dollars. Let's say it, I say it practically, but you managed to type quotas from time to time, but it was very very distant and in particular all the non-interactive modes of cloud were possible by the API call, by the cloud-p command which allows you to pass as a parameter in a line cloud command to pass a prompt and retrieve the output. So, all this stuff was OpenBar and so, Anthropique said no, wait, we have customers who pay 20 dollars and who cost us 1000 per month, so that's enough, we're going to stop there and so, from June 16, normally, there will be an API credit envelope which will be provided for each of the plans that you take and therefore, you will have this envelope and once you have exceeded this envelope, on the other hand, we will be paid to the token of API calls and therefore, it will still be another matter. So there, we're going to realize that people are going to say to themselves yeah, if I can use less efficient models and run them at home, it's still going to be pretty good and so, that's where, indeed, Hermès, they were smart, they said yeah, wait, from the outset, we, you can use behind providers with API key or even self-hosted, that is to say you put an Olamas or an OMLX, the framework from Apple and you run models locally and that's what you're going to use as an engine and eventually, you can do routing to an external engine, that is to say that typically, if you have something that requires the power of a cloud opus 4.8, well, let's go, here we go, you make calls when it's necessary and so, it's still super important from that point of view, this ability to look for broader things and be able to have something extremely open. So, for me, that's really what interested me in this and so, if we look a little bit, I have a little bit, as I said, I dug into the thing a little bit and so, I looked, if you go to Hermès Agent, you will be able to see, they have in particular guides and tutorials where you have precisely explained how to do, for example, a daily briefing bot and this daily briefing bot, he, well, you go into the thing, say, there you go, it there is a cron scheduler which will start the bot at one o'clock in the morning and therefore, it will do web research, it will do a summarization of the content and clearly, if you have, it's really a few minutes to do that, what, that's where we say to ourselves, ultimately, we all perhaps have repetitive tasks that at a given moment, we will have to go, we will have to, we will be able to operate like that and have something that runs quite autonomously therefore, it's still a matter of time afterwards, it's a matter of time to try things but I'm quite... And earlier, we were talking about IWS, you see, you can even go and integrate your Hermès Agent with Bedrock since you have the possibility of using your IWS Bedrock account to be able to run this agent so, clearly, I think that these are things that we should still, that we will have to look at quite closely to be able, quite closely, sorry, to be able to create things. And so, we were just talking about the Telegram Assistant pattern, it's obviously something that they have documented where you create your bot, you create your new bot inside Telegram and you get an ID and then you configure that in the gateway and therefore, you will be able to communicate with this Telegram. So. And closer for us and for the developers, obviously, earlier, we were looking for what could be the good use cases and indeed, the peer review on GitHub is indeed something which again can be something interesting to do, to say, well there you go, you put a GitHub CLI on your, where you run your Hermes Agent, you do it and you ask it to do the reviews and summarize the peer reviews for you to be able to easily do that. So there you have it. Afterwards, as the commits are made by agents now, if we have the peering done by agents, I don't know anymore, I don't really know if we won't have at some point a total lack of control over everything that happens. So it's perhaps... You have a real subject and we feel, annoyingly, that it's really the counterpart of the moment which is to say ok, the agents, finally, the AI tools can do lots of things to us afterwards, that's what the... What level of control do you keep on the subject and also in terms of validation and in terms of knowledge of the model, finally, the model of the world which is represented by your application and it's true that it's... It's tempting to do, well come on, it's okay, I'll let it happen, it'll go well, I'll quickly look at it and then gradually you feel that you're losing control a little and it's true that you have to find good compromises. Yeah, we definitely agree, we definitely agree that at a given moment, just now, as you said earlier, if you actually let him recreate skills, at a given moment, you have a real danger in the fact that you know... At the beginning, you're going to check and then afterwards, the thing runs in a loop on its own. So, it's clear that at some point, you're going to end up with something that can end up being dangerous. To summarize. To summarize. But you, I imagine, you, in what you're building in terms of monitoring agricultural infrastructure and so on, you probably have tons of monitoring tasks for things like that where you have to go and look frequently and so on, on which these are things that are super important. I see with a client, they have a terminology which is keep the light on and therefore, it is a big subject which occupies a lot of people every day to go around lots of assets to check that it is good. It is a complex monitoring which cannot be put in a simple dashboard which requires a lot of actions to check in order to do something else and that, I think that these are things where in the end, we will end up being able to get there and for agencies and tasks that are still tedious and with little added value which can really make a difference in my opinion. The limit of that for the moment is that all that is greenhouse equipment is often automatons which are offline and today, finally, it is rather a bit of old technology. We are not so much on the operational part, we are no longer on the analysis part so we extract the data, we recover them, we grind them and then we possibly send back instructions and then we recover after the... But we are not really at the level of the agent who... We can manage them but we are not really on the purely operational monitoring part. So it's true that we're not completely at that level and for the moment, it's more... Well, when we see the interfaces, it's more old stuff. It's like in industry. In the industrial sector, it's all to nothing. That is to say that extremely recent machines will be hyper connected, will be capable of being integrated with management systems in... But... Old machines, you have in the flora of... Industrial, things that will never be connected. So for which it is indeed very complicated. And then afterward, you also have your serfs who are in certain places. I'm not even sure if they have internet connectivity. That's it, yes. No, but what I mean by that... But yes, no, but... With a flow rate that would allow for an interaction we continue with it. Yes, after that, it’s not necessarily a permanent connection. It's true that there, we have certain projects where it's... We have holes in the... Well, which fill themselves up afterwards, but it's not necessarily real time. So there you have it. Yeah, yeah. But yes, in good potential. Yeah, yeah, absolutely. We can... I think we took a little look at this Hermès Agent. I'm sure we'll talk about it again. We have brought up some little news there, recently in the... In the Slack and in particular... In particular the story of... We can emphasize this because in addition, it is a Frenchman who wrote a LinkedIn post on this. It's Steve Morin who you can also listen to in this excellent French-style podcast that I love listening to. And it actually talks about CUDA and where... I don't know if it was you who didn't bring this up, Nico, in the Slack. In fact, to provide some context, Steve Morin is the originator of the ZML box which is ultimately the idea is to have a layer of what I understand to be a somewhat generic attraction layer to allow the models to be run regardless of the underlying hardware. And it's true that today, well, indeed, the elephant in the room is Nvidia's Cuda which allows... which is a bit of the de facto standard in this world. And it's true that the more things go, the more we actually see that the projects which go so far as to finally free themselves from Cuda which... But even from... I don't know, I listened to a podcast for a long time where Steve Morin said in Cuda there are lots of good things and then there are things that are less pretty. It's an old Cuda, we must not lose sight of the fact that it dates from 20 years ago. Yes. Not far. Version 1.0 of Python has just been released but it was announced yesterday, I think I'm lying. Yeah, yeah. But suddenly, there were things that were specific to Nvidia but it's true that Nvidia had the strength to bring both the hardware and the software and as a result it's a bit the de facto standard and it's true that there with... notably in the work of ZML and I don't know, I think he started... Michel Kardec who also... or no, I don't remember who... in short, who had highlighted other projects but there are plenty of projects which ultimately go towards a slightly more generic side which could allow us to free ourselves from Nvidia and I think that there is still a big threat for Nvidia's monopoly which is that one day we can manage to do without art and no longer do all this part which is a bit low level... that's a big big threat for them. Nvidia is threatened on two levels, that is to say that a first level which actually this software stack on which they were hegemonic they will perhaps be less and less first point then second point more and more still even if it is not yet in quotes in production everywhere alternatives to Nvidia technology for not necessarily for the training of models where it will probably remain for a very long time this brute force that we can have will still be very valid but there are still quite a few of alternatives which are in the process of rising on the inference and moreover even Nvidia in quotes is working on this alternative they bought this company which is called Grock nothing to do with the Grock it's an exclusive partnership ah sorry I was wrong sorry I was wrong I apologize it's an exclusive partnership which is to show good will very linked to the exchange of shares that's it ok officially it's not purchased it's a partnership I I don't know what exclusive I don't know if as we said I think in a previous podcast I think that the only people targeted by the maneuver were the anti-trust commission so basically we can see that there is something that is shifting and ultimately it is perhaps a very good thing because when we see deploying a use case based on LLM and others the infrastructure cost that you have to take to be able to have access in particular to inference based on NVIDIA H200 or H100 even H100 it is insane so the fact of having both a software stack which allows it to be done more easily and having hardware mechanisms which are simpler than this big artillery is clearly something which does not go in the direction of the development of NVIDIA it is clear another article which amused me I read but I find it interesting who is on I believe it is people from Uber who made it f have a bit of a sort of feedback on the identity crisis ultimately of the agents and it is true that there is a real question on this since in the measures where we delegate the response that they make I believed a little I did not go into the detail of their response it seems to me a bit of a gas factory but in any case the problem is interesting which is to say finally when we entrust a task to an agent we lose a little for whom we are doing it we lose this whole notion of traceability authorization of all these subjects and it is true that suddenly in their can it actually lose context whether for debug support or anything or also to have hey if I have Vincent who asks something to an agent we must know that it is Vincent who has the origin of the initial request and as the agents in addition can chain each other we must ultimately know all this all this traceability all this chain of operation to possibly deduce I don't know anything about permissions authorizations whatever we want and it's true that there is a real subject on this that we don't necessarily think about we say to ourselves it's good there is an agent we ask him something and then it will go well but it's true that the subject of this whole notion of identity takes up a real subject with agents yeah yeah yeah and well that's always this story of impersonation of finally we've had this subject since the dawn of time finally to know how to trace when we make multiple epia calls and finally know who is at the origin and there we end up with the same thing you are right when you speak to an agent who calls on sub-agents of the sub-agents of the sub-agents who each have their level of permission and so on what do we do do we carry around the token which could have been put in the first finally the token I mean the security token the token which could have been put in the first request and do we take it from start to finish or do we reset that it is indeed and moreover we can clearly see that there is a small movement that is taking place on the on the place finally how we make the link with the human who revolves around these processes and I had found another article done by Adi Osmani who said that ultimately the bottleneck in all agentic processes was going to be the human since at a given moment there is no that a single person who is capable of judging merger to arbitrate what the agents produce is a human brain and therefore you could deploy 250 Mac Minis in your office to do all the work and so on on both sides at a given moment everything will have to go through a single little pipe it's that of your brain so so it's quite amusing this thing what it's not the exchange the tax the orchestration tax and it's just it's true that it puts a lot of good advice we're going to say to try to protect ourselves from it but here the article is nice yeah the article is nice and well Nico I think that we were both concise and and we were able to talk about lots of things so it's it's nice it's not bad it's for me a little plot and then some news around it's not bad I like it yeah I like a subject on which we spend a little time and then we try to review two or three more things it's still not bad and well listen it was cool to record this episode dear listeners dear listeners we hope you had a good time listening to us as good a time as we had recording it and we'll see you again very soon very soon dear listeners Subtitling Société Radio-Canada
Bonjour à toutes, bonjour à tous, bienvenue dans un nouvel épisode du Big Data Hebdo. Le Big Data Hebdo, c'est le podcast francophone de la data et de l'IA. Et dans le Big Data Hebdo, on essaye de vous éclairer sur tout ce qui bouge en termes de technologie dans la data et l'IA. Et pour le coup, aujourd'hui, on va parler d'agents, on va parler d'agents IA, mais on va parler de choses un peu plus où on peut mettre les mains dedans, puisqu'on va parler de Hermes Agent. Hermes Agent, c'est l'agent open source du moment. Et pour cet épisode, je suis accompagné de Nicolas. Salut Nicolas, comment ça va ? Salut Vincent ! Bah écoute, ça va bien, chaudement, mais ça va bien. Mais là, ça va, on est 6 le matin, donc ça va. Nicolas, tu es CTO chez Cybeltech, qui est donc une société qui s'intéresse à tout ce qui est science du vivant et pour assister la culture et tout ce qui s'en suit. Exactement. Tu as tout bien présenté, nickel. Bah oui, j'ai repris ce que tu avais dit des fois dernière, tu vois, j'ai rien inventé. Ah oui, t'écoutes, t'écoutes, c'est bien. J'écoute mes interlocuteurs. Et donc, pour ma part, je suis Vincent Heuschling, je suis le cofondateur de ce podcast et je suis consultant indépendant. J'accompagne mes clients dans leur projet Data EIA pour les aider à imaginer, concevoir et aussi à implémenter les solutions qui pourront changer leur business. Et donc, allons-y. On va rentrer dans le vif du sujet. Dans cet épisode, on s'est dit finalement, on était passé un petit peu, on n'a pas beaucoup parlé ici dans le podcast de OpenClaw. Et puis, il y a un petit nouveau qui est arrivé qui s'appelle Hermès. Et Hermès, il a un petit peu positionné différemment. Le principe reste toujours le même pour souvenir. OpenClaw, c'est ce qui a mis Apple en rupture de stock sur les Mac mini, puisque le principe, c'était on faisait tourner un orchestrateur d'agents. C'est-à-dire, on consommait la ressource, par exemple, anthropique pour faire réfléchir l'agent. Mais toute l'orchestration, toute la mémorisation, l'interaction se faisait sur une machine que vous hostiez chez vous ou alors, aussi chez un hébergeur cloud. D'ailleurs, je crois que c'est Hostinger qui a même fait un type de déploiement pour son environnement où vous pouviez demander à avoir une machine OpenClaw directement installée. Et OpenClaw avait un mode d'interaction qui était le suivant. Vous parliez notamment sur un channel Telegram avec votre agent et puis votre agent faisait des choses pour vous. Donc, on était rendu dans le dialogue de l'humain avec son agent pour faire des choses. Il y a eu pas mal de choses qui ont été montrées sur le marché. Mais Hermès Agent et Harry étaient un petit peu différents. Alors déjà, là où OpenClaw était l'émanation d'un acteur solitaire, Steinberger, là OpenClaw, c'est un lab open source qui s'appelle Nous Research et qui ne sont pas forcés forcément dans le fait de faire des papiers scientifiques pour se faire reconnaître, ni de faire des levées de fonds. Donc, en fait, on parle assez peu d'eux. Mais ils ont construit des modèles. La gamme de modèles Hermès 3 qui existait dans des 8 milliards de paramètres jusqu'à 405 milliards de paramètres, fait partie effectivement des super bons modèles open source que vous pouvez utiliser. Donc, et donc la thèse qu'ils ont eue, c'est de dire avec une API LLM ouverte et des outils open source, vous pourriez déployer un agent qui rivalise avec les offres commerciales. Alors, c'est vrai qu'effectivement, on pourrait mettre ça en rapport avec ce que notamment Anthropik propose avec la capacité à avoir des agents qui tournent sur leurs propres ressources. Et ce qu'il faut regarder, c'est que quand même, ça s'est sorti en février 2026. Donc, c'est quand même tout récent, cette open, pas open cloud, mais cette Hermès agent. Et ils en sont déjà pas loin de 100 000 étoiles GitHub, je crois, alors que open cloud, en beaucoup plus longtemps, a à peu près 350 000 étoiles GitHub actuellement. Alors, peut-être que ces chiffres sont un petit peu, sont plus forcément au goût du jour, mais ça vous donne une idée quand même de la traction qu'il peut y avoir là-dedans. Et donc, c'est quand même assez intéressant parce que, mine de rien, le fait de regarder comment est-ce que ça a été conçu et comment est-ce que ça fonctionne, peut même, si vous n'avez pas envie d'installer des agents, vous aider à optimiser la manière de laquelle vous travaillez avec votre cloud classiquement chez vous. Typiquement, l'une des énormes choses que fait Hermès, c'est ce harness engineering productivisé. Donc, à un moment donné, vous avez une boucle de feedback et de rétroaction. On dit, si vous faites cinq fois le même prompt ou si vous avez un enchaînement de prompt, et ainsi de suite, écrivez un skills. Et l'une des grosses forces d'Hermès, c'est qu'au cœur du système, il va observer ce que vous faites. Et au bout d'un certain temps, s'il voit des actions répétitives, s'il voit des interactions avec des outils, des choses comme ça, il va écrire de lui-même un skills pour vous aider la prochaine fois là-dedans. Et donc ça, c'est un truc qui est super intéressant et qui aussi peut permettre de vous faire évoluer dans votre pratique. C'est pour ça que je disais, moi, à la lecture et en m'intéressant au sujet, je me suis rendu compte que la manière de laquelle je pouvais structurer le cloud qui est sur ma machine, finalement, ça pouvait vachement progresser grâce à ça. Donc Hermès arrive là où, donc, avec les autres, avec OpenClo et ainsi de suite, il fallait à peu près tout constituer. Vous aviez tout ça. Lui, il arrive déjà avec un harness super bien structuré, qui vous permet déjà de faire des choses très fortes. Et donc si on regarde un petit peu, si on détaille ça un petit peu en cinq couches, l'instruction, les contraintes, c'est-à-dire comment vous allez sandboxer les choses, le feedback, comment est-ce que vous allez pouvoir faire évoluer votre système, la mémoire, puisque au cœur de tout ça, il y a aussi la capacité d'avoir une persistance et l'orchestration. Comment est-ce que vous faites un truc multi-agent, multi-modèle et ainsi de suite ? Hermès arrive avec des réponses super claires là-dessus. Donc sur les instructions, typiquement, quand vous êtes en train de travailler manuellement, vous allez écrire votre CloudMD ou votre AgentMD, enfin, peu importe le nom que vous lui donnez, lui, il a son skill system qui va auto-créer, auto-updater. Après, il y a cette notion de contrainte, dans quelle mesure vous faites des hooks, vous faites des choses qui vont empêcher de faire certaines actions. On a vu que l'un des trucs qui avait été mis en avant sur OpenClo, c'est qu'il y avait des gens où OpenClo était parti, s'était emballé, il avait commencé à souscrire à des services complémentaires, parce qu'on lui avait demandé de faire des choses quelle que soit la manière de laquelle il puisse y arriver. Enfin, là, fonctionnement, il y a vraiment un système au cœur de tout ça. Chez Hermès, ils ont décidé d'avoir un système de permission, d'exécution qui soit sandboxé. Donc c'est quand même vachement bien de ce point de vue-là. Alors, la boucle de feedback, j'en ai parlé juste avant, c'est-à-dire que là où vous devez faire une revue manuelle, prendre un peu vos promptes, les remettre dans un skills, machin, ainsi de suite, là, automatiquement, il y a une learning loop qui est assez automatisée. Il ne fait que créer des fichiers markdown que tu peux ensuite aller éditer, aller modifier. Mais oui, alors je n'ai pas regardé dans quelle mesure, effectivement, il y avait ce côté très manuel de proposition d'un nouveau truc, et puis dans quelle mesure il pouvait, derrière, il pouvait avoir un workflow pour pouvoir les activer ou pas. Ça, je n'ai pas regardé. Pour ce qui est de la mémoire, jusqu'à maintenant, vous deviez sérialiser vous-même un peu la mémoire et réécrire dans des contextes, dans des markdowns spécifiques, réécrire des choses comme ça. Là, vous avez un système, il y a même une basculite qui permet de persister toutes les interactions et qui vous permet de commencer des interactions sur un média. Typiquement, tout à l'heure, je parlais de ce truc qui était très à la mode, d'avoir une conversation avec un... avec un... comment dire... je vais y arriver avec un télégramme ou un WhatsApp et de la terminer dans le terminal de votre machine. Eh bien là, pour le coup, ça, c'est hyper synchronisé puisqu'il y a une basculite qui permet d'avoir cet enregistrement continu de ce qui se passe. Et puis après, en termes d'orchestration, c'est là aussi où ils ont fait un super effort. C'est que vous avez une possibilité de chroner des tâches, de faire des appels multimodèles, de faire appel à des modèles différents puisqu'ils ne sont pas bindés sur un type de modèle précis. On rappelle que OpenClaw avait démarré et ça avait fait une bisbille avec Anthropik pour être très fortement utilisé avec les modèles d'Anthropik. Et ensuite, Steinberger avait été recruté par OpenAI. Et donc, je n'ai pas suivi après, mais il est fort probable que les releases suivantes d'OpenClaw soient des choses très typées pour être utilisées avec l'écosystème OpenAI. Et donc, c'est quand même quelque chose. Il y a vraiment un feature set super intéressant de ce point de vue-là et qui donne aussi de très bonnes idées si vous avez envie de développer une boucle agentique ou quelque chose comme ça derrière. notamment, il y a le système de mémoire qui permet d'avoir des choses et donc de pouvoir avoir aussi une latence de rétrival. Si vous voulez faire un RAG, par exemple, le système de mémoire va pouvoir vraiment vous aider puisque vous allez être à 10 millisecondes même si vous avez un volume de documents que vous avez mis dans votre système qui est énorme. Donc, c'est vraiment quelque chose qui va permettre de construire des scénarios et d'avoir, de résoudre plein de choses différentes. Et comme je le disais, c'est quand même une grosse alternative open source à des offres du marché parce que quand on regarde un petit peu, là, on est en train d'énumérer des choses et quand on regarde par rapport à ce qu'il y a, par exemple, dans AWS Bedrock Agent Core qui est le runtime agentique d'AWS, on est à peu près sur le même feature set. Donc, c'est vraiment un système super intéressant de ce point de vue-là. Moi, en creusant ça, j'ai vraiment trouvé ça super intéressant. Et donc, bon, ça n'arrive pas complètement nu et sur lequel vous devez tout construire puisque il y a globalement une grosse centaine de skills qui couvrent plein de cas d'usages différents. Par exemple, c'est vraiment très, très rapide de vous faire un truc qui vous lit une liste de blogs et qui vous sort un résumé tous les matins à lire pour pouvoir vous tenir au courant sur ce qui bouge en termes d'informations. Et ça, c'est un truc, vous allez pouvoir avoir ça qui tourne tout seul dans un coin sur une machine chez vous et qui ne fait appel à aucune ressource parce qu'on va y venir après. C'est très, très bien construit pour pouvoir justement auto-oster des modèles et ne pas être dépendant d'un provider extérieur. Toi, c'est des trucs, Nico, que tu as eu l'occasion un petit peu de regarder, ces agents autonomes ? Non, pas plus que ça. La partie open-clos, c'est-à-dire qu'ils m'avaient vite refroidi en mode où ils affichaient plus d'alertes de sécurité que de features à un moment donné. Donc ça, j'ai un peu lâché l'affaire. C'est vrai que j'ai un peu du mal aussi à déléguer le truc. J'aime bien tout contrôler. J'ai du mal avec ce côté tiens, fais... Bosse dans mon dos et ça va bien se passer. J'avoue que j'ai un peu du mal avec ça. Donc j'apprends à relâcher un peu, mais pas totalement. Et du coup, à open-clos, j'ai vite lâché l'affaire. Puis comme tout le monde est dessus, c'est vrai que j'ai une relation inverse. C'est plus populaire. Oui, oui, t'es comme moi. Tu es comme moi. Tu as une certaine détestation pour les choses qui sont trop hype à l'instant T. C'est ça. J'attendrai plus tard. Et c'est vrai que dans ce moment, on voit pas mal de choses. C'est vrai que Hermès, on le voit pas mal. Je vais passer sur le blog de Dolama ce matin. Il parlait d'Open Jarvis qui a l'air d'être aussi un peu dans le même esprit. On sent que ça bouge. Après, c'est vrai que les cas d'usage, j'ai un peu du mal à dire tiens, prends toutes mes données et fais ce que tu veux avec. Ça, c'est le cas extrême. Et donc, il y a quand même un vrai sujet de tout ce qui est de permission. Ouais, tu as ce sujet de sandbox qui est un vrai, vrai, vrai sujet. Moi, je reste extrêmement réticent. Typiquement, par exemple, pour des choses qui interagiraient avec mes mails ou avec mon environnement Google, j'utilise un outil en ligne de commande qui s'appelle GoCli. C'est Nicolas Martignol qui avait montré ça, que lui, il avait fait justement pour ses automations autour de DevOps. Il avait fait pas mal de vidéos là-dessus. Et il avait montré effectivement qu'au lieu d'utiliser les trucs de cloud, intégrés à cloud, notamment dans Cowork, pour aller lire ces mails qui sont pénibles, qui n'avancent pas, qui marchent à une vitesse mais complètement dingue. Il y a effectivement un petit binaire qu'on peut utiliser en ligne de commande qui permet de sortir les mails, de lister les mails. Il sort ça sous un format JSON ou un markdown, peu importe. Et après même de rédiger des mails et ça permet de ne pas ouvrir l'accès d'API directement à un outil sur lequel on ne maîtrise pas ce qu'il va faire avec. Et là, pour le coup, on est capable de dire bon, attends, ce truc-là, je lui donne ces permissions-là et ainsi de suite. Et au moins, j'ai quelque chose que je peux mettre en coupe-circuit si jamais il y a quelque chose qui ne va pas. Et donc, c'est quand même un peu ce sujet-là. Alors, tu dis oui, tu veux tout contrôler. Tu sais bien ce qu'on dit. La délégation n'exclut pas le contrôle. Oui, non, mais c'est le premier pas et j'ai un peu du mal de la confiance. La confiance n'exclut pas le contrôle. Merci de me reprendre. Pas de souci. Non, mais c'est sûr. Non, mais c'est vrai que ça se passe petit à petit. Puis on le voit avec l'évolution des modèles et puis le côté... Ouais. Mais c'est vrai que, pour le moment, je ne peux pas forcément trouver dans le cas d'usage. Il faudrait que je me force à le faire. Moi, je sais qu'effectivement, dans nos métiers où on a une nécessité de disséquer une quantité d'informations, de processer, voire de post-processer beaucoup de choses. Par exemple, je suis intimement convaincu que dans la production du podcast, avoir ça qui tourne dans un coin et qui processe des choses, qui essaye de rattacher des news récentes à des épisodes passés et ainsi de suite, ça permettrait de produire beaucoup de choses intéressantes. Ça, j'en suis intimement convaincu. Par ailleurs, je suis aussi intimement convaincu. C'était mon idée pour cet épisode, avant de tomber là-dessus, c'était de revenir, de reparler sur justement ces outils, ces outils de sémantique, ces sémantiques modèles et ainsi de suite, qui sont très intimement liés à une boucle agentique qui permettrait d'avoir un dialogue avec la donnée, avec l'environnement de données. C'est clair que ces agents autonomes ont un intérêt à pouvoir traiter là-dedans. On dit aussi, ben voilà, on dit que les agents sont des choses capables de travailler sur les tâches rébarbatives, celles que nous, humains, on n'a plus du tout envie de faire et ainsi de suite. Ok, ben le truc dans la data sur lequel on est encore un petit peu pas à l'aise, c'est le sujet de la documentation, de valider, de faire un peu le travail de linter sur les environnements de données, de valider qu'on a bien mis les commentaires sur toutes les colonnes, de valider tous ces trucs-là, enfin entre guillemets qu'on essaye de faire dans des CI-CD classiquement dans du code, je sais que sur la data c'est plus fastidieux et je reste convaincu que ça, pour le coup, des agents qui tournent dans un coin, on serait capable de faire ce travail de data stewardship assez bien, en fait. Voilà. Ah oui, je trouve qu'il y a plein d'usages, après c'est plus sur un âge perso, pour l'instant, je n'y arrive pas. L'actualité du moment où tout le monde, on sent que ça se crispe sur l'usage des tokens, c'est-à-dire que ça... Ouais, ouais. C'est un bon moment de tester qui était il y a encore quelques mois où tu, comme c'était en mode open bar sur tes abonnements cloud, OpenAI et autres, c'était facile. Là, quand on a vu que tu vas te payer au token, ça risque d'être un peu moins... Alors, ça requiert d'avoir une machine avec un peu de puissance, quoi, mais si tu ne fais pas de l'interactif, si c'est des agents auxquels tu donnes le temps de réfléchir et de pouvoir apporter une réponse, typiquement, on y viendra juste après, la variété de modèles, même avec des modèles quantifiés que tu peux faire tourner sur ta machine, elle est quand même assez importante, quoi, puisqu'il y a effectivement la possibilité d'utiliser des runtime locaux et d'éventuellement déléguer quelques tâches ponctuelles qui requériraient beaucoup, beaucoup une grande diversité de performances et des features particulières, d'envoyer ça sur un modèle distant. Mais moi, je reste convaincu, l'intérêt de ces agents, là où je n'ai jamais adhéré à l'histoire d'OpenClo, c'était, ouais, ok, si c'est pour faire tourner quelque chose sur une machine chez toi et au final, bouffer du token chez un provider extérieur, quel est l'intérêt, quoi ? Quel est l'intérêt ? Donc, à un moment donné, moi, le vrai intérêt, là, je le vois dans le sujet de comment je peux faire tourner des choses intégralement, localement et éventuellement se dire que par extension, tu peux avoir des choses où tu garantis aussi ta souveraineté. Moi, ça reste un sujet quand même sur lequel je suis super prudent, super vigilant et donc, entre guillemets, je préfère avoir un modèle deep-sic ou un quen ou un modèle comme ça qui tourne en local que de tout envoyer à l'extérieur. Oui, non, mais pareil, il faut juste avoir la machine. Il faut juste avoir la machine. Il faut juste faire tourner une machine. Et effectivement, par les chaleurs que l'on a actuellement, nos bureaux sont déjà surchauffés et donc, c'est effectivement peut-être pas la meilleure chose à faire. on est bien d'accord. Donc, on fait souvent le parallèle Hermès à OpenClaw. On fait ça depuis tout à l'heure. Ce sont des projets qui sont très simultanés puisque, comme je le disais, OpenClaw qui s'appelait CloudBot et ensuite MoldBot fait par Peter Steinberger. C'est fin d'année dernière ? Oui, c'est ça. Ça a été lancé en novembre-décembre 2025. Ça a été devenu viral en janvier-février 2026. Et aujourd'hui, Hermès Agent est arrivé à peu près dans le même time frame un tout petit peu après. Et donc, là où il y en a un qui est plus simple, OpenClaw, finalement, ce n'est pas grand-chose. C'est juste un Control Plane. Vous avez votre boucle agentique et puis c'est tout ce qu'il y a et vous intégrez plein de choses autour. Là où Hermès, lui, est beaucoup plus intégré et fournit tout et notamment ce mécanisme qu'ils appellent la learning loop et la mémoire persistante. et donc, c'est vraiment une gateway d'un côté et un runtime d'agent complet de l'autre. Néanmoins, c'est toujours assez marrant quand il y a des concurrences entre différents tooling dans un marché. Il y a une commande qui s'appelle Hermès Claw Migrate qui permet d'aller récupérer ce que vous avez déjà fait dans OpenClaw pour basculer. Donc, c'est toujours assez drôle ce genre de truc-là. Ce dont... On a un petit peu parlé d'effectivement de cette sécurité. Hermès, par défaut, l'architecture est prévue pour être safe. Ce n'est pas des correctifs à postériori qui permettent de corriger des CVE qui pourraient y avoir dessus. Donc, à un moment donné, il y a vraiment une différence sensible et c'est peut-être ça qui fera qu'on aura plus envie de l'appeler, de l'utiliser que d'utiliser son prédécesseur. Donc, il y a vraiment ce point important. Et il y a vraiment un point qui est super important qui est le fait que Hermès est agnostique au modèle. et ça, c'est quand même le premier truc que j'ai regardé. C'est dans quelle mesure on est libre de faire ce qu'on veut avec et de l'utiliser comme on veut. Donc, typiquement, vous pouvez évidemment utiliser des providers cloud. Alors, ça va être Noussportal pour utiliser les modèles Hermès qu'ils utilisent, qu'ils proposent. Vous pouvez utiliser votre compte Anthropik. Alors, on reviendra sachant que ça va changer dans les jours qui viennent puisque je crois qu'au 15 juin, Anthropik va faire le rollout de ce nouveau mode de facturation puisque jusqu'à maintenant, justement, OpenClaw fait partie de ces choses qui ont fait que la rentabilité d'Anthropik notamment a été mise à mal très fortement puisque souvenez-vous au départ, vous pouviez avoir votre abonnement à 20 dollars et faire des appels d'API sur la base de cet abonnement à 20 dollars pratiquement illimité. Disons-le, je le dis pratiquement, mais vous arriviez à vous tapiez de temps en temps des quotas, mais c'était très très lointain et notamment tous les modes non interactifs de cloud étaient possibles par l'appel d'API, par la commande cloud-p qui permet de passer en paramètre dans une commande line cloud de passer un prompt et de récupérer l'output. Donc, tous ces trucs-là, c'était OpenBar et donc, Anthropique a dit non, attendez, nous, on a des clients qui payent 20 dollars et qui nous en coûtent 1000 par mois, donc ça suffit, on va arrêter là et donc, à partir du 16 juin, normalement, il va y avoir une enveloppe de crédit API qui sera fournie en regard de chacun des plans que vous prenez et donc, vous aurez cette enveloppe et une fois que vous aurez dépassé cette enveloppe, par contre, on sera payé au token des appels d'API et donc, ça va quand même être une autre affaire. Donc là, on va se rendre compte que les gens vont se dire ouais, si je peux utiliser des modèles moins performants et les faire tourner chez moi, ça va quand même être pas mal et donc, c'est là où, effectivement, Hermès, ils ont été malins, ils ont dit ouais, attend, d'entrée de jeu, nous, vous pouvez utiliser derrière des providers avec API key ou même du self-hosted, c'est-à-dire que vous mettez un Olamas ou un OMLX, le framework d'Apple et vous faites tourner des modèles localement et c'est ça que vous allez utiliser comme moteur et éventuellement, vous pourrez faire un routage vers un moteur externe, c'est-à-dire que typiquement, si vous avez quelque chose qui requiert la puissance d'un cloud opus 4.8, et bien, let's go, on y va, vous faites des appels quand c'est nécessaire et donc, c'est quand même super important de ce point de vue-là, cette capacité à aller chercher des choses plus larges et de pouvoir avoir quelque chose d'extrêmement ouvert. Donc, moi, c'est vraiment ça qui m'a intéressé là-dedans et donc, si on regarde un tout petit peu, j'ai un petit peu, comme je disais, j'ai un petit peu creusé le truc et donc, j'ai regardé, si vous allez dans Hermès Agent, vous allez pouvoir voir, ils ont notamment des guides et tutorials où vous avez justement bien expliqué comment faire, par exemple, un daily briefing bot et ce daily briefing bot, lui, bon, vous allez dans le truc, dire, voilà, il y a un cron scheduler qui va démarrer le bot à une heure le matin et donc, il va faire des web research, il va faire une summarization du content et clairement, si vous avez, c'est vraiment quelques minutes pour faire ça, quoi, c'est là où on se dit, finalement, on a peut-être tous des tâches répétitives qu'à un moment donné, il va falloir aller, il va falloir, on va pouvoir opérer comme ça et avoir quelque chose qui tourne de manière assez autonome donc, c'est encore une histoire de temps après, c'est une histoire de temps pour essayer les choses mais je suis assez... Et tout à l'heure, on parlait d'IWS, vous voyez, vous pouvez même aller intégrer votre Hermès Agent avec Bedrock puisque vous avez la possibilité d'utiliser votre compte IWS Bedrock pour pouvoir faire tourner cet agent donc, clairement, je pense que c'est des choses quand même qu'il faudrait, qu'il va falloir regarder de manière assez proche pour pouvoir, d'assez près, pardon, pour pouvoir créer des choses. Et donc, on parlait justement du pattern Telegram Assistant, c'est évidemment quelque chose qu'ils ont documenté où vous créez votre bot, vous créez votre new bot à l'intérieur de Telegram et vous récupérez un ID et puis vous paramétrez ça dans la gateway et donc, vous allez pouvoir dialoguer avec ce Telegram. Voilà. Et plus proche pour nous et pour les développeurs, évidemment, tout à l'heure, on cherchait ça peut être quoi les bons cas d'usage et effectivement, la peer review sur GitHub est effectivement quelque chose qui là aussi peut être quelque chose d'intéressant à faire, de dire, ben voilà, vous mettez une GitHub CLI sur votre, là où vous tournez votre Hermes Agent, vous faites et vous lui demandez de faire les reviews et de vous résumer les peer pour pouvoir facilement faire ça. Donc voilà. Après, comme les commits sont faits par des agents maintenant, si on fait faire les peer par des agents, je ne sais plus, je ne sais pas vraiment si on ne va pas avoir à un moment donné une absence totale de maîtrise sur tout ce qui se passe. Donc c'est peut-être... Tu as un vrai sujet et on sent, chiant, que c'est vraiment le pendant du moment qui est de se dire ok, les agents, enfin, les outils d'IA peuvent nous faire plein de trucs après, c'est quel est le... Quel niveau de contrôle tu gardes sur le sujet et aussi en termes de validation qu'en termes de connaissance du modèle, enfin, le modèle du monde qui est représenté par ton application et c'est vrai que c'est... C'est tentant de faire, bon allez, c'est bon, je le laisse faire, ça va bien se passer, je regarde vite ça et puis au fur et à mesure tu sens que tu te décroches un peu et c'est vrai qu'il faut trouver de bons compromis. Ouais, on est bien d'accord, on est bien d'accord qu'à un moment donné, tout à l'heure, comme tu disais tout à l'heure, si effectivement tu le laisses recréer des skills, à un moment donné, tu as un véritable danger sur le fait que tu sais... Au début, tu vas vérifier et puis après, le truc tourne en boucle tout seul. Donc, c'est clair qu'à un moment donné, tu vas te retrouver avec quelque chose qui peut finir par être dangereux. Pour résumer. Pour résumer. Mais toi, j'imagine, toi, dans ce que vous construisez en termes de surveillance d'infrastructures agricoles et ainsi de suite, vous avez sûrement des tonnes de tâches de surveillance de choses comme ça où il faut aller regarder fréquemment et ainsi de suite, sur lesquelles c'est des choses qui sont super importantes. Je vois chez un client, ils ont une terminologie qui est keep the light on et donc, c'est un gros sujet qui occupe beaucoup de monde tous les jours de faire le tour de plein d'assets pour vérifier que c'est bon. C'est un monitoring complexe qui ne peut pas être mis dans un simple dashboard qui requiert beaucoup d'actions pour aller vérifier pour aller faire autre chose et ça, je pense que c'est des choses où au final, on va finir par pouvoir y arriver et pour des agences et des tâches quand même fastidieuses et à peu de valeur ajoutée qui peuvent vraiment à mon avis faire une différence. La limite de ça pour l'instant, c'est que tout ce qui est équipement de serre, c'est souvent des automates qui sont offline et aujourd'hui, enfin, c'est plutôt par un côté d'un peu des vieilles technos. Nous, on n'est pas tellement sur la partie opérationnelle, on n'est plus sur la partie analyse donc on extrait les données, on les récupère, on les mouline et puis on renvoie éventuellement des consignes et puis on récupère après le... Mais on n'est pas au niveau vraiment de l'agent qui... On peut les piloter mais on n'est pas vraiment sur la partie purement opérationnelle de suivi. Donc c'est vrai qu'on n'est pas complètement à ce niveau-là et pour l'instant, c'est plutôt des... Enfin, quand on voit les interfaces, c'est plutôt des vieux trucs. C'est comme dans l'industrielle. Dans l'industrielle, c'est du tout au rien. C'est-à-dire que les machines extrêmement récentes vont être hyper connectées, vont être capables d'être intégrées avec des systèmes de gestion dans les... Mais... Les machines anciennes, t'as dans le flore des... Industrielles, des trucs qui seront jamais connectés, quoi. Donc pour lesquels c'est effectivement très compliqué. Et puis après, tu as aussi tes serfs qui sont à certains endroits. Je suis même pas certain qu'elles aient une connectivité internet, quoi. Voilà, si. Non, mais ce que je veux dire par là... Mais oui, non, mais... Avec un débit qui permettrait d'avoir une interaction on continue avec. Oui, après, c'est pas forcément de la connexion permanente. C'est vrai que là, on a certains projets où c'est des... On a des trous dans le... Enfin, qui se rebouchent tout seul après, mais c'est pas forcément du temps réel. Donc voilà. Ouais, ouais. Mais oui, dans les bons potentiels. Ouais, ouais, tout à fait. On peut... Je pense qu'on a fait un petit tour sur cette Hermès Agent. Je suis certain qu'on en reparlera. On a remonté quelques petites news là, ces derniers temps dans les... Dans le Slack et notamment... Notamment l'histoire de... On peut le souligner parce qu'en plus, c'est un Français qui a écrit un post LinkedIn là-dessus. C'est Steve Morin que vous pouvez aussi écouter dans cet excellent podcast qui est à la French que j'adore écouter. Et ça parle justement de CUDA et où... Je sais pas si c'est toi qui avais pas remonté ça, Nico, dans le Slack. En fait, pour remettre un peu de contexte, Steve Morin est l'origine de la boîte de ZML qui est l'idée finalement c'est d'avoir une couche de ce que j'en comprends d'une couche d'attraction un peu générique pour permettre de faire tourner les modèles quel que soit le hardware sous-jacent. Et c'est vrai qu'aujourd'hui, bon, effectivement, l'éléphant dans la pièce c'est Cuda de Nvidia qui permet de... qui est un peu le standard de fait dans ce monde-là. Et c'est vrai que plus ça va, plus on voit effectivement que les projets qui vont jusqu'à s'affranchir finalement de Cuda qui... Mais même de... Je sais pas, j'ai écouté un podcast de très longtemps où Steve Morin disait dans Cuda il y a plein de bons trucs et puis il y a des trucs qui sont moins jolis. C'est vieux Cuda, faut pas perdre de vue que ça date quand même d'il y a 20 ans. Oui. Pas loin. La version 1.0 de Python vient juste de sortir mais c'était annoncé hier je crois à mentir. Ouais, ouais. Mais du coup, il y avait des choses qui étaient spécifiques à Nvidia mais c'est vrai qu'il y avait Nvidia a eu la force d'apporter aussi bien le hardware que le logiciel et du coup c'est un peu le standard de fait et c'est vrai que là avec des... notamment au travail de ZML et je sais plus, je crois qu'il s'est mis... Michel Kardec qui avait aussi... ou non, je ne sais plus qui... bref, qui avait souligné d'autres projets mais il y a plein de projets qui vont finalement vers un côté un peu plus générique qui pourrait permettre de s'affranchir de Nvidia et je pense qu'il y a quand même une grosse menace d'ailleurs pour le monopole d'Nvidia c'est que le jour on peut arriver à se passer de coup d'art et à ne plus faire toute cette partie un peu bas niveau que ça... c'est un gros grosse menace pour eux quoi. Nvidia est menacé à deux niveaux c'est-à-dire qu'un premier niveau qui effectivement cette pile logicielle sur laquelle ils étaient hégémoniques ils vont peut-être l'être de moins en moins premier point puis deuxième point de plus en plus quand même même si c'est pas encore entre guillemets en production partout des alternatives à la technologie Nvidia pour pas forcément pour l'entraînement des modèles où ça restera probablement encore très longtemps cette brute force qu'on peut avoir restera encore très valable mais il y a quand même pas mal d'alternatives qui sont en train de monter sur l'inférence et d'ailleurs même Nvidia entre guillemets est en train de travailler sur cette alternative ils ont racheté cette boîte qui s'appelle Grock rien à voir avec le Grock c'est un partenariat exclusif ah pardon je me suis trompé pardon je me suis trompé je m'excuse c'est un partenariat exclusif qui est pour montrer de la bonne volonté très lié à de l'échange d'action c'est ça bon officiellement c'est pas acheté c'est un partenariat je sais plus quoi exclusif je sais pas si comme on l'a dit je pense dans un précédent podcast je pense que les seules personnes visées par la manœuvre c'était la commission anti-trust donc foncièrement on voit bien qu'il y a quelque chose qui est en train de shifter et finalement c'est peut-être une très bonne chose parce que quand on voit déployer un use case à base de LLM et autres le coût d'infrastructure que vous devez prendre pour pouvoir avoir accès notamment de l'inférence basée sur des NVIDIA H200 ou H100 même H100 il est démentiel donc le fait d'avoir à la fois une pile logicielle qui permette de le faire plus facilement et d'avoir des mécanismes hardware qui soient plus simples que cette grosse artillerie c'est clairement quelque chose qui ne va pas dans le sens du développement de NVIDIA c'est clair un autre article qui m'est bien amusé j'ai lu mais je trouve intéressant qui est sur je crois que c'est des gens de chez Uber qui en fait font un peu une sorte de retour d'expérience sur la crise d'identité finalement des agents et c'est vrai qu'il y a une vraie question là-dessus puisque dans les mesures où on délègue la réponse qu'ils font je croyais un peu je ne suis pas rentré en le détail de leur réponse ça me paraît un peu une usine à gaz mais en tout cas la problématique est intéressante qui est de dire finalement quand on confie une tâche à un agent on perd un peu pour qui on le fait on perd toute cette notion de traçabilité d'autorisation tous ces sujets-là et c'est vrai que du coup dans leur est-ce qu'il peut effectivement perdre du contexte que ce soit pour du support du debug ou n'importe quoi ou aussi pour avoir tiens si j'ai Vincent qui demande quelque chose à un agent il faut qu'on sache que ça soit Vincent qui a l'origine de la demande initiale et comme les agents en plus peuvent se chaîner les uns les autres il faut qu'on connaisse finalement toute cette toute cette traçabilité toute cette chaîne de fonctionnement pour déduire éventuellement j'en sais rien des permissions des autorisations tout ce qu'on veut et c'est vrai qu'il y a un vrai sujet là-dessus auquel on pense pas forcément on se dit c'est bon il y a un agent on lui demande un truc et puis ça va bien se passer mais c'est vrai que le sujet de toute cette notion d'identité reprend un vrai sujet avec les agents ouais ouais ouais et ben ça c'est toujours cette histoire d'impersonnation de enfin on l'a depuis la nuit des temps ce sujet finalement de savoir tracer quand on fait de multiples appels d'épia et savoir finalement qui est à l'origine et là on se retrouve avec le même truc t'as raison quand tu parles à un agent qui lui fait appel à des sous-agents des sous-agents des sous-agents qui ont chacun leur niveau de permission et ainsi de suite qu'est-ce qu'on fait est-ce qu'on balade le token qui aurait pu être mis dans la première enfin le token je dis bien le token de sécurité le token qui aurait pu être mis dans la première requête et est-ce qu'on le balade de bout en bout ou est-ce que ou est-ce qu'on a on reset ça c'est effectivement et d'ailleurs on voit bien qu'il y a un petit mouvement qui se fait sur le sur le la place enfin comment on fait le lien avec l'humain qui gravit autour de ces process et j'avais trouvé un autre article fait par Adi Osmani qui disait que finalement le bottleneck dans tous les processus agentiques ça allait être l'humain puisque à un moment donné il n'y a qu'une seule personne qui est capable de juger de merger d'arbitrer ce que les agents produisent c'est un cerveau humain et donc tu pourras beau déployer 250 Mac Mini dans ton bureau pour faire tout le boulot et ainsi de suite de part et d'autre à un moment donné tout devra passer par un seul petit tuyau c'est celui de ton cerveau quoi donc donc c'est assez amusant ce truc là quoi c'est pas l'échange la taxe la taxe d'orchestration et c'est juste c'est vrai qu'il met pas mal de bons conseils on va dire pour essayer de s'en prémunir mais là l'article est chouette ouais l'article est chouette et bah Nico je trouve qu'on a été à la fois concis et et on a pu parler de plein de choses donc c'est c'est chouette c'est pas mal c'est pour moi une petite trame et puis quelques news autour c'est pas mal moi j'aime bien ouais moi j'aime bien un sujet sur lequel on passe un peu de temps et puis on essaye de passer en revue deux trois trucs en plus c'est quand même pas mal et bah écoute c'était cool d'enregistrer cet épisode chers auditeurs chers auditrices on espère que vous avez passé un bon moment à nous écouter un aussi bon moment que nous on a eu pour l'enregistrer et on vous retrouve très bientôt très bientôt chers auditeurs Sous-titrage Société Radio-Canada