← Back to search

An AI agent at home: this is what Hermes does

Risorse Artificiali AI Engineering in italiano · 2026-05-09 · 77 min
relevance 81 9713 words Episode page ↗ Audio ↗
Show full episode description
Un AI agent in casa che prende iniziative da solo: cosa fa davvero Hermes Agent in locale. Stefano Maestri, Alessio Soldano e Paolo Antinori parlano di tutto quello che sta accadendo nell'AI engineering oltre il modello: ottimizzazioni di inferenza (speculative decoding di Gemma 4 con drafter da 76M, D-Flash, Rotor-Quant, P-Flash), Google che vende le TPU e firma 5 gigawatt di datacenter con Anthropic, il sospetto che ChatGPT Image 2 sia Sora declassato, il caso Elon Musk vs Sam Altman e il dibattito sulla sovranita' digitale europea (cloud alla Lidl, tassazione degli agenti AI). Al centro c'e' il case study di Hermes Agent installato in locale da Stefano: gestione mail e calendario, smart home, paper digest autonomi, e il momento in cui l'agente decide da solo di renderizzare un HTML in foto per leggerlo in macchina. Code Wave Anthropic, Antirez che forka llama.cpp, l'AGI come sistema integrato e non piu' solo modello: una mappa veloce delle cose da tenere d'occhio per chi costruisce con AI in produzione. Follow per non perdere i prossimi episodi. #51
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
AI inference is slow and memory-heavy; the episode also warns against humanizing AI models like Claude.
Benefits
  • Speculative decoding speeds up token generation
  • Drafter models cut compute for next-token evaluation
  • Shared KV cache lets small model reuse big model work
  • RotorQuant cuts memory and speeds prefill vs TurboQuant
  • Local inference optimizations available via llama.cpp
Use cases
  • Gemma 4 2-billion model paired with a 76-million-parameter drafter
  • Drafter generates 4x tokens evaluated by target in one pass
  • Bernie Sanders and Veltroni interviewed Claude as a person
  • RotorQuant runs an order of magnitude faster than TurboQuant
  • Lucebox scorer compresses prompts to speed up prefill
KPIs / results
  • Gemma 4 2 billion parameters, 76 million parameter drafter
  • Drafter generates 4x tokens per pass
  • RotorQuant an order of magnitude faster than TurboQuant
Tools / build
  • Gemma 4 speculative decoding (MTP)
  • RotorQuant
  • TurboQuant
  • D-Flash diffusion drafter models
  • Lucebox prompt scorer
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
Hello everyone and everyone, welcome back, welcome back. Let's start, then, with many things from Elon Musk vs Sam Altman, unlikely interviews with Claude, then technical things including the super-fast Gemma 4 and then many other things. Come on, let's start, we have a thousand today. Where do we start from? Shall we start from sadness? Never feel sad, come on. Come on, let's start from the sadness, let's start from the sadness that are the unlikely interviews, that the unlikely interviews are not the nice ones they did, I don't remember who, maybe Chiambretti. Never regular. Huh? Never regular. Never to regulate, yes yes, never to regulate, the unlikely interviews. No, they're the ones they do to artificial intelligence models. We talked about Bernie Sanders's in the past, and Veltroni did it too. He did it to Claude too, I think. Because he is Noartry's Bernie Sanders. We talked about it a few weeks ago that in America Bernie Sanders, also a senator, interviewed Claude, and a former candidate for the presidency of the Council, who is a journalist for his main job, however, that is Veltroni, decided to interview Claude too, with such an innovative idea, right? Which almost seems... To me, then, when I read it I really thought that it's like Little Tony when he was... Elvis. Exactly, yes. The Little Tony that Elvis does is like the Thrones that Bernie Sanders does. So, no. Just as Bernie Sanders's was no, even more so. This is not what we can pass on to the new generations. They were, um... Let's try to understand this artificial intelligence, not... Not necessarily humanize it, right? To ask him you will destroy us... Er... The thing that almost did to me... Or what do you think about the end of life or something. No, the thing that made me more tender, almost the tenderness of ignorance, in the literal sense of the term, eh, if you can, Veltroni tells me patience, however, of ignoring the thing when he asks if you make mistakes. And clearly what tells him, yes, I'm full of gaps, I make mistakes. Sad, eh, in the sense... They are tools, let's use them as such, let's not humanize them, let's not ask them about themselves. The issue of consciousness, long discussed by Antropic, etc., is an interesting field of research, but let's leave it in the field of research. Because then I read them afterwards, right? Already the newspapers, which is humanized artificial intelligence and which therefore leads young people to do negative things, even the worst. Eh, but if that is the image we begin to give without having understood what we have in hand, like this. As usual, my opinion is quite strong, but no, really not. For me a no. And where did they give it? On television, in prime time? Eh, no, no, I interviewed in a newspaper. Ah, too bad, because it was one of those things to put on national TV or something like that, in my opinion. Eh, but we'll get there. Now, I don't know, we'll get to Mara Venera to interview the PT in prime time. But, look, half-jokingly, since you were saying that your daughter has to finish her high school diploma this year, maybe tell her to prepare on the AI ​​track for the theme, which in my opinion... Ah, no, no, we talked about it, but yes, it's likely that they'll give it as the theme, but that would also fit. And I also think kids would say smarter things. Let's leave aside the politicians, the former Italian politicians, let's touch the American ones. They would say smarter things than Bernie Sanders and even his president of the Republic. Because he too has said some, eh, in recent days. Have you seen? I'm referring to Trump, who said... Then we get away from politics and back to technology. But I am referring to Trump, who said he would like to have the power to veto the release of artificial intelligence models by the White House due to real or presumed danger. Which, that is, you will imagine that OpenAI, Google and Antropica didn't take it very well... Yes, also because then I would be asked on the basis of what to make the decision, that is... And who thinks it is dangerous? You are right. You are absolutely right. That is, which experts... which experts does the White House equip itself with in the situation to have the ability to discern what the researchers from Anthropic rather than OpenAI have done? Elon wasn't the expert, sorry. Eh, or even... Eh, but now there's a bit of a rush, so... The expert could become OpenAI, which will certainly favor Cloud releases. Exact. I mean, well... It really gets... Almost bordering on ridiculous. If not, interview the IAI too, ask him what he thinks about that other release that's coming. Ah yes, also... It seems to me... It seems fundamental to me. Let's get away from politics, come on, what do we do. We're just kicking ass. So, no, new models, come on, let's talk about new models. Let's start from... At Google home. Let's go to Alessio's house for a bit, the inference, etc. Gemma 4. Did you see that they did what you wrote in your last Aladdin article? Who read perhaps. Definitely, look. But then, in the meantime... Let's see if I can share something with you. So what happened? It happened that I was talking about speculative decoding, since the world reads me, of course, even Google thought of advertising this technique. Unfortunately they have an automatic translator. For the other they do the translation, so... Exactly. No, joking aside. First let's look at this speech by Gemma 4, then if you'll give me a moment I'd like to think a bit more broadly about what's happening these days, in this field. In Gemma 4 they decided to enable speculative decoding, which would be a technique to speed up the token generation phase, and therefore the responses, when querying a model. So not the first part, which is that of understanding the prompt, but the subsequent part of generating the response. How do you do this optimization, this speeding up? There are various techniques and a group of these techniques are based on the use of drafter models. Practically they are smaller and consequently faster models, which are asked to make predictions about the next token or tokens to be generated and the target model, which would be the large model you are working with, instead of generating the next token itself, first makes an evaluation of the prediction made by the small model. If the small model was good enough at predicting the next token well, it saves time because evaluating the prediction is substantially less computationally heavy than actually calculating the next token. Or in any case, as is done in this case here on Gemma 4, it is possible to parallelize and evaluate substantially multiple tokens in a single pass. So when the small model catches us, you have a big profit. What those at Google did was essentially optimize this idea a lot and as they did it with a very small model so to say I seem to have gained some points the 2 billion model of Gemma 4 however it is already relatively small it has a drafter model with 76 million parameters therefore 2 billion 76 million therefore extremely faster and this small drafter model generates 4 times let's say the tokens that the large model would generate and the target model does the evaluation of these of this generation in a single pass. What did they do to improve things further? They have basically invented tricks such as sharing the KVCache of the two models so the small model draws on processes that the large model has already done for the cache and furthermore when the embeddings from which the generation operation for the draft model starts are calculated these embeddings are hung concatenated after the result let's say the activations of the last layer of the larger model so it is a way to allow the small model despite it being equipped with just a few parameters and therefore not very intelligent let's put it this way, starting from a pre-processing of the current state that the large model had arrived at, this obviously if you read the paper it is explained much better, this essentially allows the draft model to catch us quite often and however and here you will tell yourself for a moment let's tell you my thoughts of the last few days this whole thing fits into a much more extensive reasoning that is, we are noticing the research that is addressing the problems of efficiency of the inference below in various in various fields this which we have just described has to do with the speed of speculative generation decoding is not only this approach used by Google which by the way is called mtp multi token prediction but there are other other techniques such as the one I talked about in my article which is ngram which essentially allows the model to see what it has generated in the previous steps and make predictions based on that it is possible to match draft models developed let's say independently with respect to the target model that is being used clearly they must be matched well i.e. it is not that you can take any draft model but without them also being embedded as in this case of gem there is research to create draft models for quen models for example with various techniques and among other things there is a technique called D-flash which is abundantly researched in this period which allows us to use diffusion models I don't know if you remember that months ago we were also talking about it in the podcast we mentioned the existence of models for the generation of text which are not autoregressive but based on the idea of diffusion the same that is used for the generation of images and these models essentially do as in the case of the image a noise reduction starting from something that represents the total noise and generating several tokens in parallel many tokens in parallel this idea this approach is exactly what goes well with the construction of a draft model that makes predictions of the next tokens because the downside of the fusion models was precisely that of being fast but not excessively accurate compared to the best autoregressive models and this is exactly the condition we are in now with the draft models so we are interested in speed we are willing to accept lower quality because then there will be the target model that will evaluate the prediction therefore they exist there are draft models that are being developed at the moment precisely to do this thing called flash in the meantime the research is also obviously trying to tackle the KVCache phase and therefore the cache we had talked about TurboQuant several weeks ago many other ideas for optimizing the cache have come out where the objective is obviously to reduce memory occupation therefore allowing the use of relatively large models even in the case of few resources and memory resources there has come out among the various other ideas for optimizing the way in which the cache is made something called RotorQuant which basically goes to try to improve one of the defects of TurboQuant which was the fact that to build the cache in the way they explained in the TurboQuant paper different computational resources were essentially used so if it is true that the memory used was reduced the prefill phase was still slowed down the idea of these people from RotorQuant is quite complex to explain but essentially they make different transformations of the input vectors they divide them into smaller vectors and then they have an intelligent way to process these smaller vectors moral of the story orders an order of magnitude more faster than TurboQuant in different usage scenarios good therefore optimisation ion in generation in let's say on the cache part but news from the last few days also in the prefill prompt processing phase which is as we were saying before what happens when you go to process the prompt the basic prefill phase if you want natural optimization is precisely the use of the kvcache because the idea is that instead of redoing all the attention one goes to see what was calculated previously but we need to use a technique of using a technique that we can share something maybe Stefano can you do it yes of course basically the idea of lucebox is the group that developed this thing is to have a new small drafter model that functions as a scorer that is to make an evaluation of the prompt of the prompt tokens that are passed trying to understand which are the interesting parts of the prompt token we are talking about the classic problem which is to find lake in the haystack that is to understand in a very large context in which we have a lot of information and not all of it is super salient to understand where the most salient parts really are on which you can actually do the prefill and therefore downstream of this scoring phase a prompt is built more compressed and with that prompt you do the actual prefill clearly this approach unlike flash and in short in generation this approach is not lossless i.e. you lose information in doing this compression but the idea is that in the correct use case for example the one in which there has been a lot of discussion in the context the previous calls and previous prompts and not all the information that has been passed is salient because various things have been tried in the end you have obtained a let's say a certain result it is possible to compress the context in this way and have a speed up of the prefil without substantially losing too much in terms of accuracy, all these improvements on these three main cash generation prefil areas are partly available for example the ama cpp but in some cases there are actually forks of the ama cpp there is an effort to try to unify all these developments together let's see I am quite confident and I was struck by all this research which is converging to improve the local difference too give me your opinion Alessio as the major expert having not chosen many these things are very interesting and I was wondering what what they will bring will lead to better techniques and therefore the state of the art or in reality since some go in one direction others go in another there will always be room for divergent algorithms and approaches as there is now with data structures or algorithms in general for which there will always exist that case study for which this is the best case so maybe there will be I don't know small shops of people who optimize the model for that specific thing for that specific language for that specific hardware for that specific slack request that it must have so all these techniques will not disappear but each one will have the its because in the specific case what do you think the scenario is but in my opinion it is a mixture of the two things i.e. when there is an optimization which for a given phase of the inference is better than all the others in say 95% of cases it is likely that that will become the de facto standard various things come to mind in attention now the so-called psique moment in which they defined fast sparsa attention then they all used it because it was clearly better than dance attention which dance was no longer used however okay, a different approach was used, but there are things that everyone will use, in my opinion, others that are more niche, there are plenty of start-ups that are betting on this thing in their heads, and Mira Morati, his start-up is betting on the fact that companies will need specific optimizations and specific fine tuning, that is, they don't make models, they are trying to democratize the tuning phase, they are more on the models, but in my opinion they will also get a little bit about inference if we go as it seems at the moment at least more and more to have hybrid solutions where local inference and cloud inference coexist and solve different problems yes, however, I mentioned these things here because they are clearly a whole series of optimizations that can make beautiful models usable, let's say that they have interesting capabilities on non-exhaustive hardware therefore created for local inference, including consumer ones, however in reality all these things are also of great interest to those who offer these services in the cloud because it is all a way to save resources for if you still want to have a better return on investment because the moment you can offer the same performance with a fraction of the resources invested without compromising the quality it's all profit yes look I confess my inspiration was thinking about the Linux kernel and the infinite number of parameters and configurations that it hides that most of us don't use in the sense that we use the default and from time to time anyway we touch on the most obvious things but that thing there does a lot of stuff imagine what do I know it has all the possibilities for managing memory such as killing processes or the infinity of different file systems that exist who knows why they exist I mean it is justified that it continues to exist I don't know anyone uses them for sure I wonder if it will be the same thing with all these technologies optimizations for AI well I know about inference engines those let's say commercial or even non-commercial but cloud the level at which you can do the tuning is really high maybe we also invite we have among our contacts also those who work on VLLM maybe sometimes we invite them to talk to us about how the white rabbit theory is based because then on that thing we are at a level similar to what you describe of the Linux kernel and even less friendly than that perhaps because we are younger in terms of technology but yes there won't be much space left because all the very vertical technologies reserve all that space there this is my opinion at least I don't know if Alessio sees it differently no no in my opinion too and also the fact that there is so much different hardware means that it will take a phase of settling that is, in my opinion we will have for some time different parameters as you say Paolo which lead to clear improvements with an architecture and are not good for others I was looking to say there are remaining at the masses pp different plans open to optimize the inference for example on Strixalo on the hardware that I have and one of the recurring comments that I have seen from the maintainers of lama cpp is yes it is fine but this thing that you are proposing is not an improvement rather vaguely worsening on this other architecture or until the let's try on all of them, let's not take it until these problems are also resolved here clearly we will have n.000 parameters n.000 approaches and no because then on the one hand you add the complexity of this thing the cpp is an extremely complex project on the other hand you also add some of their strong choices more or less agreeable for me not agreeable which is the first thing they write in the contributing .md and that they don't accept pull requests completely or mainly AI generated which frankly leaves a bit of time in the current times and in the type of project that is, these people make AI and don't want AI to be used for their project, it seems like a bit of a contradiction in terms to me and it's the thing he wrote on there is a need to make an inference engine that is optimized in particular for an architecture in his case that of the Mac Mini and he says okay I'll do it and I don't know if it started from a fork or if it started from scratch because the software is not yet public he said it will obviously become public but at the moment it isn't however VLLM blade CPP or blade all these that try to keep all types of hardware together certainly have an extremely higher complexity as it is for Linux the example of Linux is fitting because even in PC architectures there are 1500 differences and in fact the audio never works look while we were talking about it and it wasn't where I wanted to get to but I found myself thinking that perhaps there could be room for a company that does what Red Hat did at the beginning of Linux that would worry about keeping track that all these little pieces converge from time to time and then someone tries them all together and guarantees you that that's what works what Red Hat does with VLLM it requires a scientist a little different in q This is a case where each of these algorithms that we talked about today will become a consumable aspect of the library and therefore a building block. However, I can't imagine that there could be a demand for something of this type. Yes, Red Hat is betting on it. They have acquired the main company that contributed VLLM to do these things. They are making it their AI offer. One of the predominant ones, Daniele Zunca also told us when he came here for an interview, is precisely that of giving a VLLM for enterprise. So how was Red Hat Linux for enterprise? yes I wanted to say that again you see how the research fades into software engineering in making all the specific ideas for a given case etc. well integrated with the right levels of abstraction etc. so that they can be used by various types of software yes yes yes absolutely we are at the phase on that part there at least we are at the software engineering part we are not the research still exists and it is certainly fundamental predominant but we are in the software engineering phase as yesterday there was on Wednesday there was there was the annual convention which for them is half-yearly in reality of anthropic code and then I heard the keynote at a couple of talks because for us it was evening or night no disturbing news in the sense that they made a list of all the disturbing innovations that they have made in the last period there was no model announcement there was no real announcement of new modes even if there are many rumors around a new mode that should appear in cloud code which the leaks call it orbit which is essentially proactive cloud code a bit like open cloud or Hermes agent which we will then talk about because I am super enthusiastic about Hermes but nothing sensational but a lot of we have engineered this part here so that cloud code is more easily usable with Sonnet having only Opus as advisor for example they have focused a lot on that part there of another part and a whole series of things that we are starting to see how at least in the part of the harnesses of the coding agents or harnesses in general we are moving a little bit towards the software engineering part and also there are those who say that the so-called EGI is no longer something that only does the supermodel but he is a supermodel with an engineered with an excellent inference with an excellent tool properly integrated with everything else let's say and therefore integrated yes because then in the end you know when you say the EGI it means that he must be able to do everything that a man does therefore he must have the same tools that that person has otherwise where do you have but this also brings us a little to the discussion you were making before about the hardware there is an announcement from Google instead which instead is preparing for the Google I.O which is in a few days and if you remember Google tends to make announcements in the two or three weeks preceding Google I/O and then at Google I.O. there are no announcements but there is the presentation of what has already just been announced and one of the announcements which is not so much about the product is not so much consumer but is that it has a selected group of customers selling TPUs which is actually a big announcement in our world because selling TPUs means that envy is a real big competitor because AMD has tried but the distance in the data center world is still very high the TPUs instead of Google that run Gemini, let's remember, are of a level comparable to the Nvidia GPU as regards inference at least on training there are conflicting opinions there are those who say that they are very good even in training those who say less however between this announcement and the announcement of an agreement instead made with Antropic to provide Antropic with data center power for 5 gigawatts Google made a leap on the stock market the other day these are not financial advice but just to underline it how the market realizes the importance of this two announcements that is that it begins to also provide others but in a significant way because 5 gigawatts is a lot of stuff and the fact that you are also starting to sell hardware is also a change of strategy a bit by Google that has always made hardware but has not always sold the service but has never sold the hardware as such so perhaps also something to point out for 30 different ones that are starting to appear and if TPU arrives we go back to what you were saying Paolo before TPU is another architecture on which to add parameters to be set actually sorry you made me come in mind but I should do an internet search for the details that Google years ago had actually produced a USB TPU that had become widespread and I no longer remember what it's called I have it for the Raspberry it's called it's called it's called I don't know if you want to open the box but I didn't even know you had it I've always found them interesting too for doing inference on the edge then in reality I read that at the time in short it was a bit of an embryonic thing it worked but not like that it was no longer a game yes so in in reality they had already gone through it not with probably a business vision like this but more to say let's try it seemed more like a research thing it seemed more to give it to enthusiasts yes it's not for the server side here it was a bit of a toy the ones we talk about now are the giant ones yes yes which I also came across in my path this week because I was always trying to antivocal to see if we could make a more optimized version of some of the models in particular I would like to fine tune parakits with conversational Italian because a apparently they trained him on the formal and set language of journalists but not on friends who slur words at you and so when you get a real voice from a friend who doesn't know what the fuck to say to you it doesn't work so well whisper is better on this side and so I said okay come on let's see if it can be done and I entered a rabbit hole so yes it can generally be done but because of the technology that Nvidia used in that case it can be done if you have Nvidia hardware and blah blah blah in short in the end I don't know if I succeed but in my discovery I discovered that on Google collab in addition to traditional GPUs you can also choose the TPU and the model prediction which told me that if I had done the same operation with the TPU I would have had a gain of 4x compared to the times made with the GPU which was not bad I didn't get to the end because the step before getting to doing that didn't work so I haven't seen it but sooner or later I will solve my curiosity to try to play with it and see a little how it behaves yes then this this is certainly one of the aspects is also understanding how these things behave here when they put them to the test seriously, however all the optimizations that Alessio was talking about including those on Gemma 4 which are actually in their house are trying to go a bit in this direction here among other things always staying in Google always for the ads in preparation for the Io it seems that they are testing a model called Omni which should unify Veo with Nano Banana therefore generation of video generation of images in a single model which is something a little in somehow waiting also because the rumors said another news that I read that in reality it is true that OpenAI has discontinued Sora you remember the video generation model etc. etc. but they also say that in reality Sora was discontinued for that video generation thing but in reality chat GPT image 2 it's not that Sora has downgraded only making images with optimizations on this thing etc. etc. which all in all it would fit because objectively the results of Sora were remarkable from a point of view visual and those of certain GPT images are notable yes then I read the announcement I honestly haven't yet found much information on Omni it is interesting that we are talking about a multimodal omnimodal model I don't know how we want to define it which also generates videos because typically these multimodal models accept various things but then they essentially generate very frontier images or text let's say interesting let's see let's see if they absolutely amaze us it's all to be seen clearly because then between saying and doing there is always a middle ground let's go back come on let's go back to the world for a moment of politics now we have already lost everyone first we lost all the listeners who wanted technology talking about politics at the beginning then we lost all the others we killed them with with the whole part about inference let's go back to politics and I would like to go to one of Paolo's favorite characters if I remember correctly it is Elon who clashes with Paolo's other favorite character which is Sam it is true that you don't know which sticker to choose yes yes yes I confess then yes then for those who don't know what's happening these days I have no idea if it's being talked about in Italy Moreover, I don't watch television or newspapers much so I have no idea if anyone is paying attention to this news but after much bickering we have finally gone to court for an issue which, if you like, with which I also exceptionally agree with Elon, that is, the criticism's attack, Elon Musk's attack on Sam Altman for which he decided to transform an open AI non-profit into a for-profit company to make a business out of it which doesn't sound very good to me, that is, if you are born with a non-profit and then discover that you can make money by not doing it another company but transforming yours in short it's not like you're really on the list of my favorites now maybe I'm fascinated by this whole affair because I just finished watching a legal series on Netflix Lincoln Lawyer and so I'm well taken and I can't wait to see these things and already on my YouTube feed there are a lot of videos that show how the day in court went these things but I haven't had time to dedicate myself to them yet I trust that I will hear echoes of them through American satire programs through John Oliver Colpere and the others so I hope sooner or later to have juicier things what I've read now is that Ailon isn't doing very well first of all the criticism is that he probably doesn't give a damn about the ethics and morals underlying it despite him having put about thirty million dollars into a company that was supposed to be a profit and then changed but it seems that there is a bit of background to questioning Elon's good intentions because if on the one hand he says ah we are worried about the fact that the IGA will kill us all and these things we need a bit of ethics on the other hand it turned out that Elon had tried to buy OpenAI and spread it with Tesla to have everything in one but they told him no and apparently he was a little resentful so he said then I'll take the ball away and so there's this part here then precisely it seems that Elon is also in court a little erratic as they say in English so I don't know what the most appropriate translation in Italian is, a rambling guy, I don't know and when he talks he skips a lot, so much so that serious lawyers like those of the series I watch are roasting him a little because he says what comes to mind and not what he is prepared for so if we base ourselves on the initial conversations it is difficult for this case to end well for Elon one of the revelations that I found most interesting, not necessarily true but more interesting is that one of the reasons that Elon cites for having invested in OpenAI at the time when it was a non-profit that had the aim of bringing the knowledge of AI to humanity therefore absolutely noble was that he had had a chat with the two founders of Google and they had said to him he had asked him listen but aren't you worried that AI will wipe us all out and one of the answers attributed to one of them is apparently this oh yes but it's not a problem as long as AI survives the extermination of humanity and this comment if it's true obviously makes me get a little nervous and I can understand Elon's concern now it has to be said if it is ever been true though whether it was ever true or not however that someone not necessarily them could imagine it like this is quite disturbing and I don't think that the current trial will lead to anything useful in this situation but it makes me slightly more worried that there could be the mad scientist who for the love of science loves the phantom of the opera more who loves computers more than humans and who therefore will bring this thing anyway I imagine that other new super-fucking revelations will come out of the trial of Elon and Samatman and so I will keep you updated as they happen they come out like this instead of continuing to hear only about Garlasco's murder we also hear about the OpenAI trial, no but then that statement there by one of the two founders of Google could easily have been made by Sergey Brin because he is a very slightly over the top character in his statements so I don't believe it enough that someone could have said it who then didn't and I also think who between the two is but certainly whoever has Google's AI in his hands at the moment, Demis Hassabis, has different visions and does nothing but reiterate it in his definitely very human centric interviews more than all the others even more than the good Darione but then I saw another thing Paolo what a self-fulfilling prophecy let's buy let's buy the servers from Lidl so let's buy the servers from Lidl then let's self-quote an episode in which I have no idea what it was I don't care to go and look for it but joking in the past we talked about Chinese models and the fact that they arrive at a low price and all that democratization and we were joking that in the near future someone will buy the model on Aliexpress which is our equivalent of buying it at Lidl and guess what something vaguely similar is actually happening in Europe driven by the perceived need god thank you from the rulers at European level who say maybe it's not that we can really trust the good intentions of the United States our historical ally who always do the right thing and of their companies perhaps it's the case that we start thinking about having a proprietary infrastructure that isolates us from the risks of being excessively dependent on someone from third parties a parallel if you like to what has happened with fuel with Russia at the beginning of Russia's war with Ukraine realistic if you ask me as a scenario and then they said let's organize ourselves let's try to invent something and the first infrastructure that was missing was that of cloud services because right now AWS dominates Microsoft and Google mainly probably also Alibaba but no one ever remembers them and strictly European there isn't much and also the various local European regions that have been guaranteed it was noticed recently how sorry they are I activated Google on the phone which heard me speaking bad about him and he replied I said we don't have anything European and even though there are regions officially with European data centers they are subject to American law and American law someone has finally gone and looked he says that the cloud act if the government needs it for national security issues which apparently happens for every bullshit now in the United States it needs to go to the providers to ask them for your data it doesn't matter where you are under what legislation they have given it and this is a true thing so true that in addition to having been noticed by the European community it has been noticed from someone close to the Europeans, i.e. Switzerland which ended up on the front page because they severed a contract worth 200 million or more, I don't know, I don't remember these fantasy numbers with Microsoft, I think precisely because it couldn't guarantee them this thing, basically when they asked and verified they told them no, your data is very private except for that time when they ask us for it and since the Swiss care about these matters, traditionally having one foot in both shoes, privacy is, let's say, theirs, their brand, they took this situation directly and that's it. announced this this cut in those parts how that will be resolved I still don't know in the sense that I don't know if they will look at a local solution, their European solution or someone else's solution or they will force their hand over there to be able to change the rules for them I don't know but the fact is that we have reached the front page with this thing and it is a question therefore where Switzerland moves Europe also moves Europe has moved in a slightly different way or by issuing a tender to go and ask for European commercial initiatives of being able to provide a cloud service that Let it be dignified, this thing was done in a very democratic way, so much so that instead of emerging a supergroup, a coordination of someone who can give you vague confidence that he is capable of doing it, I don't know, I would have dreamed that Cern would do it, ok we are in Switzerland but someone with a name you can trust instead no various more or less random actors have advanced in Germany in Holland in Holland in other places nothing from Italy because we won't have the postal cloud at least we got one right and one of these is a linked company at Lidl it is not exactly Lidl that is, the title we say Lidl the Lidl cloud is actually a sort of service company that was born as a sub-company for the needs of the Lidl group and then specialized if you want the same relationship that existed between Amazon and AWS but the name came from there and therefore somewhere in a while there will be the cloud of llidl which will be our variant compared to the AWS cloud we joke about it in reality Lidl does its job well it's more to be seen in my opinion the real bet is not so much who manages it but more to see if the ambitions of reaching a scale that can be self-sustaining which is if not at the same let's say the same level as that of AWS but which goes in a direction that makes you have confidence that Europe can build on this is a good thing and let's see if we are capable it is a good and fundamental thing because it's also there I've mentioned this thing many times I also believe here on the podcast and I've certainly talked about it in some interviews with some of the guests there's also a possible scenario that's almost probable today I'm going too far to say that it's almost probable that at a certain point the so-called agents so from here I'll tell you about mine as an agent no joking aside the so-called agents when they stay inside the cloud etc. will reach a point where they will produce value ok they'll produce value it's something that Draghi also says I don't know I'm not sure in mind as I go and when it is an agent that produces value there is a possible scenario in which what is taxed is the value produced traditionally in the European economy the tax is placed on the value produced not so much on those who produce it but because it is more convenient to tax the person than the value but in reality what you tax is the value produced and if the value is produced on a soil other than the European one it could be subjected to a taxation that is different from the European one so the problem is starting to become complex also of sustainability of the welfare of which Europe it rightly has the flagship for which to look forward and be able to produce digital value on European soil is something that is not only important from the point of view of privacy which is of great interest to Europe and has always made it a workhorse that's fine but there is also a much more practical question of where the money ends up then yes this is a legal quagmire not to be laughed at all the issues related to cryptocurrencies to taxation come to mind understanding where the funds reside you are in Europe in the United States where in reality they are not I am in the light the tangle, however, is actually a tangle and as far as the blockchain is concerned it is more difficult because it is a distributed thing and it is built to avoid that thing there in reality one of the ideas is that who is doing and who will do agent-based services will tend to put them more in a cloud where yes that problem exists there because there are many points etc. but a little more manageable and it is in everyone's best interest to get to define who produces who dares as value so perhaps it is a slightly less complex problem but which can become bloody if not do you have anything on Italian European soil that produces value or in any case be able to offer the possibility from the point of view of those who set up this thing here to say okay but all this goes around here so I want it to be treated with the laws over here okay exactly exactly this to point out the thing but talking about agents and also talking about locations of doing things locally I told you that I set up Hermes agent in my house on an old computer using a Chinese model like this so as not to miss anything and it is doing things it is doing things which I will also talk about later in the newsletter that comes out on Monday but that I wanted to tell you about, maybe also show something for those who have gotten to the end here at least sorry if I interrupt you just to clarify you are running the agent locally with the model however offered in the cloud by a correct provider because I don't have a machine on which to run a model with sufficient performance to make it do these things so the model is cloud but instead the agent the tool to use the technical term which is Hermes agent runs locally on one of my machines and the files it modifies the things it does it does it there there but in addition I gave him access to a whole series of my services. I told him about it last time also to my email and so on, trying to take all the precautions possible. I chose Hermes and not open cloud for this reason because it's a little easier to customize access. There's a very clear skill system about which for example I said it the other time about Gmail. I told him okay go to Gmail. Let me see how it goes and he'll tell me. I showed him his script and I told him here are those two methods there, the ones that do delete and the one that does send of the email, we delete them right from the script you use, it's not that I'm telling you not to use them, delete them, you don't have to have them anymore and then I checked that he had done it and so and so it is and what I make him do I make him do a lot of things from the most banal things if you want of cyclical use so in the morning or every hour he checks my email he tells me which emails are only informative and he gives me a summary of a line which instead need my attention by going to read the email and therefore understanding that I need my attention and furthermore for those that require a response I have him generate a draft of the response he doesn't send it I then find them in gmail already drafted then I edit them but I have a starting point at least and he looks at the calendar it tells me which are the appointments of the day and the important ones of the week for me the important ones are those of a certain color i.e. we discussed for a moment how to define these things I find the gaps like I say to him eh I have to talk with Alessio we take an hour you find me a hole in the calendar he makes me three or four proposals I tell him that one is fine and he sets me the appointment management things of this type he manages things in my house so I gave him all the various smart sockets everything smart I have in the house which is not as much as Paolo but I have something too eh and so like from outside the house I tell him to turn on the air conditioner or the heat pump eh to check the lights the cameras these things here eh and then and then he does things then the access is with telegram for me he does quite proactive things that is in the sense that the other day I told him hey look at my personal site eh a personal site like an electronic business card no hey what do you think the answer wasn't very kind but it doesn't matter eh more or less at the level of the answer on how our podcast thumbnails were yes yes yes more or less eh no he was kinder because even there you tell him how he should behave I recommend being extremely frank and extremely direct and he told me yes I understand the attempt but I don't we're here um and I tell him okay what do we do eh this while my son was at basketball training what do we do what don't we do and he says look I'll give you three proposals tell me if you like more of a dark version as a developer or another ending I choose the dark version and he says let's do it like this I'll make you a new website directory I'll push it to you I'll deploy it and then you take a look at it and tell me how it's going sorry I'm basically he told you but what am I asking you to do after you expressed your choice to him that is he was doing so disgusting your preference that he said let's do both of us then decide later no no no he did to me what I asked him but the incredible thing that I can share a slide of because there is the new website wait while I open it while I talk the thing that he did to me left me there is that in other conversations for the podcast for the newsletter etc. etc. I had given him in the newsletter for example an agenda part in which I say the conferences which were the ones who published the video the ones where I will go and anything else and he had this information because we had already talked about it he has a memory management done very well better than others hierarchical and so on and he generated this thing for me here that even if you like it you don't like it but then these three my contents let's say that they are the two newsletters and our podcast were on the old site so it's not his merit the projects were there on the old site and therefore it's not his merit of these I only had the last one the devox because I had done it by hand instead he saw that there were all these talks that I had done that I will do and he put them on the page for me taking both the slides and the videos that everything else is right and this thing here about taking initiative struck me quite a bit it's not the initiative that also who cares that they attributed open cloud to do things alone but initiative let's say intelligent another thing so let's also talk about that topic there is one thing that we had put in the lineup which is an article wait while I recover the title because I risk saying nonsense that of the 100 million what is 12 million context now I don't have it here at hand but c 'It's an article we talked about Alessio right we talked maybe in chat we don't have it yes I haven't looked into it in depth but there is this article that talks about a start up that has created a 12 million model with a long context of 12 million tokens since there have been announcements in the past of people who made 100 million tokens which then didn't go anywhere I wanted to understand it before talking about it for a short time and I said to Hermes I need to understand this article I'm in the car so give me a well done summary him he left and asked me ask me for permission to do anything but because that is my choice I was in the car not driving he is precise and he tells me well but if you want if you give me permission to install a text to speech I will read it to you it seems like a good idea this foul a text to speech model has been installed he told me in the meantime I'll give you a summary in Italian that maybe I didn't tell him that I wasn't driving if you're in the car and you're driving it's better if I do it in your language I'll give you a summary in Italian and I'll read it to you out loud and he did it I'll share it with you for a moment if I can wait until I see that I've actually opened he sent you a vowel so he sent me a vowel yes yes he sent me a vowel then I already sent him the vowels before because speech to text has it by default but now here there's a lot of stuff because he sent me my calendar wait where do you have the vowel sounds wait I can't find them I wanted to let you hear them so here's the reading maybe a little less sexy than others now I can't do it for you hear because we had already tested that things from Telegram don't work, the audio doesn't work, that is, you hear it but then I don't hear it, the listeners don't like it, I'll tell you that the reading is acceptable, it's not the voice of OpenAI but that of Eleven Labs that you can put on, he immediately asked me, do you have an Eleven Labs account, if you have it, I'll use it, no, I don't have it, I don't pay for many things, so I did this with a local model, it took me a few minutes and he sent me these two summaries with two different points of view and so from what he told me I understood a few things I told him yes but listen give me a summary in tabular form while I'm not choosing the right window to scroll ok give me a summary in tabular form of this data and he wrote it to me like this and I told him no I want a table because I don't understand much when I read like this and he made me a table because I told him that I always want everything in markdown and I told him again no I don't understand shit I'm in the car I can't see it and he made this crazy choice he made an html to give me the table done well and then he said to himself oh no but the user told me that in the car how does he read an html wait I'll render it I'll take a photo and send him a photo and he sent me the photo of the table then if you want we'll go into the detail of the paper but maybe not today eh well you and then oh well after that he made me another one on another topic remembering that I told him I couldn't read it and then he gave it to me again like this and in this case it was to me that stuff here was quite wow then I appreciate it but I had already seen it in short in my long conversations via telegram with simply ZAI GLM I too had come to those conclusions in this case perhaps he was proactive that he told you all of them in my cases instead I gave him promptings in the .md cloud telling him look I'm around so when I'm around I need short stuff he uses emojis these things so I he did similar things but I appreciate and understand the wow effect he did no I am particularly satisfied also and above all by the fact that he has access leave aside the cloud services also because being able to access the email to the calendar is a remarkable thing i.e. being able to send a voice message while I'm out and about and say I remembered that you have to add this task and then you can find it in gtask that's a lot it's cool ditto for appointments on the calendar and things like that but also that it has access to the file system for example I told you about the llm wiki the one from carpati that I use to think and to analyze the papers at the moment I am looking at the papers on long-term memory and in this case more to save tokens and time than anything else there is the phase after I have collected all the papers there is the digest phase which means that you download all the papers in pdf the law generates the wiki creates the links etc. which still takes half an hour 40 minutes this on cloud code not only him and I did it in the morning but it annoyed me I launched it and I did another part the number of non-trivial tokens that he uses to do this thing that you pay for and so I did it anyway with glm and so it took a while now I have a task assigned to him that does the Arabase of the github where I told him the things in the evening we do the whole digest part of reflect and when I'm done I create a pull request and I emerge the pull request and then I work on the things already digested and this is a great time saver last but not least and you may have seen it on lince I had him do the pull request review of things I had done with cloud code and there it was fascinating because if they played it they sang it because cloud I told him look you have a pull request review I went to see it and he said ah yes well done this it's interesting I'll do this one I'll do it he committed and then he put a comment no I won't do this one because it's more of a waste of time than anything else I don't know what you saw in the code but you're wrong and so it was nice that they played and sang it which is a bit like what you do when you do slash simplify locally in fact you do a pull request review on your own so full of enthusiasm for Hermes I can't live without it anymore you made me think of a use case for which I need this thing so I'll publicly inform you that I'll also fake communion, I'm the tempting devil, yes it's true and well then if you don't want to talk about that very complicated paper that I have prepared for but the time could be on average long I think that we have more or less reached the bottom of the list let's just say this thing that always in the world of temptations the idea of switching to codex because there are little animals that talk while you write the code is something that I'm trying to resist but I don't know until when I'll succeed listen let me see a screenshot because as I was telling you I activated that thing a while ago I don't have it ready wait I have it let's open the next episode with the screenshot with the animals let's open the next episode with the screenshot I don't have the screenshot ready okay I have it on codex but there is a small subscription and I don't use it much but I tried it and I have to say that in short having the animal because it's a well made animal it's not like the little ugly one that the one in the code had made this is really nice come in he walks in the middle of the screen he also breaks balls a bit if you want but so the metaverse of what's his name from Zuckerberg had his why and the only why was for the programmers to show them cartoons while they program yes no why am I saying it because I tried a bit of codex because I'm trying it with the inchessac there's a lot of new things go and see them on the site I'm not going to tell them here Claude and I have done a lot of things in recent days including the paranoid mode he has from fans so go go and see it if you want okay we say goodbye we say hello to the public see you next time numerous put bells little stars subscribe to the channel and don't miss out watch the shorts don't miss this week's cover which will see Alessio on the cover yes it's up to Alessio it's up to Alessio then I was also saying that not only it's up to Alessio for the cover then we'll try to do the covers again so maybe a little less Bruce Willis because the last cover with me everyone told me he looked like Bruce Willis when he was fit you take what you're told I've been there put on a good face I activate the game and you take everything you're told well good now it's Alessio's turn I'll go straight away to see how to make the cover bye everyone bye bye