Un AI agent in casa che prende iniziative da solo: cosa fa davvero Hermes Agent in locale. Stefano Maestri, Alessio Soldano e Paolo Antinori parlano di tutto quello che sta accadendo nell'AI engineering oltre il modello: ottimizzazioni di inferenza (speculative decoding di Gemma 4 con drafter da 76M, D-Flash, Rotor-Quant, P-Flash), Google che vende le TPU e firma 5 gigawatt di datacenter con Anthropic, il sospetto che ChatGPT Image 2 sia Sora declassato, il caso Elon Musk vs Sam Altman e il dibattito sulla sovranita' digitale europea (cloud alla Lidl, tassazione degli agenti AI). Al centro c'e' il case study di Hermes Agent installato in locale da Stefano: gestione mail e calendario, smart home, paper digest autonomi, e il momento in cui l'agente decide da solo di renderizzare un HTML in foto per leggerlo in macchina. Code Wave Anthropic, Antirez che forka llama.cpp, l'AGI come sistema integrato e non piu' solo modello: una mappa veloce delle cose da tenere d'occhio per chi costruisce con AI in produzione. Follow per non perdere i prossimi episodi. #51
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
AI inference is slow and memory-heavy; the episode also warns against humanizing AI models like Claude.
Benefits
Speculative decoding speeds up token generation
Drafter models cut compute for next-token evaluation
Shared KV cache lets small model reuse big model work
RotorQuant cuts memory and speeds prefill vs TurboQuant
Local inference optimizations available via llama.cpp
Use cases
Gemma 4 2-billion model paired with a 76-million-parameter drafter
Drafter generates 4x tokens evaluated by target in one pass
Bernie Sanders and Veltroni interviewed Claude as a person
RotorQuant runs an order of magnitude faster than TurboQuant
Lucebox scorer compresses prompts to speed up prefill
KPIs / results
Gemma 4 2 billion parameters, 76 million parameter drafter
Drafter generates 4x tokens per pass
RotorQuant an order of magnitude faster than TurboQuant
Tools / build
Gemma 4 speculative decoding (MTP)
RotorQuant
TurboQuant
D-Flash diffusion drafter models
Lucebox prompt scorer
0/0
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
Hello everyone and everyone, welcome back, welcome back. Let's start, then, with many things from Elon Musk vs Sam Altman, unlikely interviews with Claude, then technical things including the super-fast Gemma 4 and then many other things. Come on, let's start, we have a thousand today. Where do we start from? Shall we start from sadness? Never feel sad, come on. Come on, let's start from the sadness, let's start from the sadness that are the unlikely interviews, that the unlikely interviews are not the nice ones they did, I don't remember who, maybe Chiambretti. Never regular. Huh? Never regular. Never to regulate, yes yes, never to regulate, the unlikely interviews. No, they're the ones they do to artificial intelligence models. We talked about Bernie Sanders's in the past, and Veltroni did it too. He did it to Claude too, I think. Because he is Noartry's Bernie Sanders. We talked about it a few weeks ago that in America Bernie Sanders, also a senator, interviewed Claude, and a former candidate for the presidency of the Council, who is a journalist for his main job, however, that is Veltroni, decided to interview Claude too, with such an innovative idea, right? Which almost seems... To me, then, when I read it I really thought that it's like Little Tony when he was... Elvis. Exactly, yes. The Little Tony that Elvis does is like the Thrones that Bernie Sanders does. So, no. Just as Bernie Sanders's was no, even more so. This is not what we can pass on to the new generations. They were, um... Let's try to understand this artificial intelligence, not... Not necessarily humanize it, right? To ask him you will destroy us... Er... The thing that almost did to me... Or what do you think about the end of life or something. No, the thing that made me more tender, almost the tenderness of ignorance, in the literal sense of the term, eh, if you can, Veltroni tells me patience, however, of ignoring the thing when he asks if you make mistakes. And clearly what tells him, yes, I'm full of gaps, I make mistakes. Sad, eh, in the sense... They are tools, let's use them as such, let's not humanize them, let's not ask them about themselves. The issue of consciousness, long discussed by Antropic, etc., is an interesting field of research, but let's leave it in the field of research. Because then I read them afterwards, right? Already the newspapers, which is humanized artificial intelligence and which therefore leads young people to do negative things, even the worst. Eh, but if that is the image we begin to give without having understood what we have in hand, like this. As usual, my opinion is quite strong, but no, really not. For me a no. And where did they give it? On television, in prime time? Eh, no, no, I interviewed in a newspaper. Ah, too bad, because it was one of those things to put on national TV or something like that, in my opinion. Eh, but we'll get there. Now, I don't know, we'll get to Mara Venera to interview the PT in prime time. But, look, half-jokingly, since you were saying that your daughter has to finish her high school diploma this year, maybe tell her to prepare on the AI track for the theme, which in my opinion... Ah, no, no, we talked about it, but yes, it's likely that they'll give it as the theme, but that would also fit. And I also think kids would say smarter things. Let's leave aside the politicians, the former Italian politicians, let's touch the American ones. They would say smarter things than Bernie Sanders and even his president of the Republic. Because he too has said some, eh, in recent days. Have you seen? I'm referring to Trump, who said... Then we get away from politics and back to technology. But I am referring to Trump, who said he would like to have the power to veto the release of artificial intelligence models by the White House due to real or presumed danger. Which, that is, you will imagine that OpenAI, Google and Antropica didn't take it very well... Yes, also because then I would be asked on the basis of what to make the decision, that is... And who thinks it is dangerous? You are right. You are absolutely right. That is, which experts... which experts does the White House equip itself with in the situation to have the ability to discern what the researchers from Anthropic rather than OpenAI have done? Elon wasn't the expert, sorry. Eh, or even... Eh, but now there's a bit of a rush, so... The expert could become OpenAI, which will certainly favor Cloud releases. Exact. I mean, well... It really gets... Almost bordering on ridiculous. If not, interview the IAI too, ask him what he thinks about that other release that's coming. Ah yes, also... It seems to me... It seems fundamental to me. Let's get away from politics, come on, what do we do. We're just kicking ass. So, no, new models, come on, let's talk about new models. Let's start from... At Google home. Let's go to Alessio's house for a bit, the inference, etc. Gemma 4. Did you see that they did what you wrote in your last Aladdin article? Who read perhaps. Definitely, look. But then, in the meantime... Let's see if I can share something with you. So what happened? It happened that I was talking about speculative decoding, since the world reads me, of course, even Google thought of advertising this technique. Unfortunately they have an automatic translator. For the other they do the translation, so... Exactly. No, joking aside. First let's look at this speech by Gemma 4, then if you'll give me a moment I'd like to think a bit more broadly about what's happening these days, in this field. In Gemma 4 they decided to enable speculative decoding, which would be a technique to speed up the token generation phase, and therefore the responses, when querying a model. So not the first part, which is that of understanding the prompt, but the subsequent part of generating the response. How do you do this optimization, this speeding up? There are various techniques and a group of these techniques are based on the use of drafter models. Practically they are smaller and consequently faster models, which are asked to make predictions about the next token or tokens to be generated and the target model, which would be the large model you are working with, instead of generating the next token itself, first makes an evaluation of the prediction made by the small model. If the small model was good enough at predicting the next token well, it saves time because evaluating the prediction is substantially less computationally heavy than actually calculating the next token. Or in any case, as is done in this case here on Gemma 4, it is possible to parallelize and evaluate substantially multiple tokens in a single pass. So when the small model catches us, you have a big profit. What those at Google did was essentially optimize this idea a lot and as they did it with a very small model so to say I seem to have gained some points the 2 billion model of Gemma 4 however it is already relatively small it has a drafter model with 76 million parameters therefore 2 billion 76 million therefore extremely faster and this small drafter model generates 4 times let's say the tokens that the large model would generate and the target model does the evaluation of these of this generation in a single pass. What did they do to improve things further? They have basically invented tricks such as sharing the KVCache of the two models so the small model draws on processes that the large model has already done for the cache and furthermore when the embeddings from which the generation operation for the draft model starts are calculated these embeddings are hung concatenated after the result let's say the activations of the last layer of the larger model so it is a way to allow the small model despite it being equipped with just a few parameters and therefore not very intelligent let's put it this way, starting from a pre-processing of the current state that the large model had arrived at, this obviously if you read the paper it is explained much better, this essentially allows the draft model to catch us quite often and however and here you will tell yourself for a moment let's tell you my thoughts of the last few days this whole thing fits into a much more extensive reasoning that is, we are noticing the research that is addressing the problems of efficiency of the inference below in various in various fields this which we have just described has to do with the speed of speculative generation decoding is not only this approach used by Google which by the way is called mtp multi token prediction but there are other other techniques such as the one I talked about in my article which is ngram which essentially allows the model to see what it has generated in the previous steps and make predictions based on that it is possible to match draft models developed let's say independently with respect to the target model that is being used clearly they must be matched well i.e. it is not that you can take any draft model but without them also being embedded as in this case of gem there is research to create draft models for quen models for example with various techniques and among other things there is a technique called D-flash which is abundantly researched in this period which allows us to use diffusion models I don't know if you remember that months ago we were also talking about it in the podcast we mentioned the existence of models for the generation of text which are not autoregressive but based on the idea of diffusion the same that is used for the generation of images and these models essentially do as in the case of the image a noise reduction starting from something that represents the total noise and generating several tokens in parallel many tokens in parallel this idea this approach is exactly what goes well with the construction of a draft model that makes predictions of the next tokens because the downside of the fusion models was precisely that of being fast but not excessively accurate compared to the best autoregressive models and this is exactly the condition we are in now with the draft models so we are interested in speed we are willing to accept lower quality because then there will be the target model that will evaluate the prediction therefore they exist there are draft models that are being developed at the moment precisely to do this thing called flash in the meantime the research is also obviously trying to tackle the KVCache phase and therefore the cache we had talked about TurboQuant several weeks ago many other ideas for optimizing the cache have come out where the objective is obviously to reduce memory occupation therefore allowing the use of relatively large models even in the case of few resources and memory resources there has come out among the various other ideas for optimizing the way in which the cache is made something called RotorQuant which basically goes to try to improve one of the defects of TurboQuant which was the fact that to build the cache in the way they explained in the TurboQuant paper different computational resources were essentially used so if it is true that the memory used was reduced the prefill phase was still slowed down the idea of these people from RotorQuant is quite complex to explain but essentially they make different transformations of the input vectors they divide them into smaller vectors and then they have an intelligent way to process these smaller vectors moral of the story orders an order of magnitude more faster than TurboQuant in different usage scenarios good therefore optimisation ion in generation in let's say on the cache part but news from the last few days also in the prefill prompt processing phase which is as we were saying before what happens when you go to process the prompt the basic prefill phase if you want natural optimization is precisely the use of the kvcache because the idea is that instead of redoing all the attention one goes to see what was calculated previously but we need to use a technique of using a technique that we can share something maybe Stefano can you do it yes of course basically the idea of lucebox is the group that developed this thing is to have a new small drafter model that functions as a scorer that is to make an evaluation of the prompt of the prompt tokens that are passed trying to understand which are the interesting parts of the prompt token we are talking about the classic problem which is to find lake in the haystack that is to understand in a very large context in which we have a lot of information and not all of it is super salient to understand where the most salient parts really are on which you can actually do the prefill and therefore downstream of this scoring phase a prompt is built more compressed and with that prompt you do the actual prefill clearly this approach unlike flash and in short in generation this approach is not lossless i.e. you lose information in doing this compression but the idea is that in the correct use case for example the one in which there has been a lot of discussion in the context the previous calls and previous prompts and not all the information that has been passed is salient because various things have been tried in the end you have obtained a let's say a certain result it is possible to compress the context in this way and have a speed up of the prefil without substantially losing too much in terms of accuracy, all these improvements on these three main cash generation prefil areas are partly available for example the ama cpp but in some cases there are actually forks of the ama cpp there is an effort to try to unify all these developments together let's see I am quite confident and I was struck by all this research which is converging to improve the local difference too give me your opinion Alessio as the major expert having not chosen many these things are very interesting and I was wondering what what they will bring will lead to better techniques and therefore the state of the art or in reality since some go in one direction others go in another there will always be room for divergent algorithms and approaches as there is now with data structures or algorithms in general for which there will always exist that case study for which this is the best case so maybe there will be I don't know small shops of people who optimize the model for that specific thing for that specific language for that specific hardware for that specific slack request that it must have so all these techniques will not disappear but each one will have the its because in the specific case what do you think the scenario is but in my opinion it is a mixture of the two things i.e. when there is an optimization which for a given phase of the inference is better than all the others in say 95% of cases it is likely that that will become the de facto standard various things come to mind in attention now the so-called psique moment in which they defined fast sparsa attention then they all used it because it was clearly better than dance attention which dance was no longer used however okay, a different approach was used, but there are things that everyone will use, in my opinion, others that are more niche, there are plenty of start-ups that are betting on this thing in their heads, and Mira Morati, his start-up is betting on the fact that companies will need specific optimizations and specific fine tuning, that is, they don't make models, they are trying to democratize the tuning phase, they are more on the models, but in my opinion they will also get a little bit about inference if we go as it seems at the moment at least more and more to have hybrid solutions where local inference and cloud inference coexist and solve different problems yes, however, I mentioned these things here because they are clearly a whole series of optimizations that can make beautiful models usable, let's say that they have interesting capabilities on non-exhaustive hardware therefore created for local inference, including consumer ones, however in reality all these things are also of great interest to those who offer these services in the cloud because it is all a way to save resources for if you still want to have a better return on investment because the moment you can offer the same performance with a fraction of the resources invested without compromising the quality it's all profit yes look I confess my inspiration was thinking about the Linux kernel and the infinite number of parameters and configurations that it hides that most of us don't use in the sense that we use the default and from time to time anyway we touch on the most obvious things but that thing there does a lot of stuff imagine what do I know it has all the possibilities for managing memory such as killing processes or the infinity of different file systems that exist who knows why they exist I mean it is justified that it continues to exist I don't know anyone uses them for sure I wonder if it will be the same thing with all these technologies optimizations for AI well I know about inference engines those let's say commercial or even non-commercial but cloud the level at which you can do the tuning is really high maybe we also invite we have among our contacts also those who work on VLLM maybe sometimes we invite them to talk to us about how the white rabbit theory is based because then on that thing we are at a level similar to what you describe of the Linux kernel and even less friendly than that perhaps because we are younger in terms of technology but yes there won't be much space left because all the very vertical technologies reserve all that space there this is my opinion at least I don't know if Alessio sees it differently no no in my opinion too and also the fact that there is so much different hardware means that it will take a phase of settling that is, in my opinion we will have for some time different parameters as you say Paolo which lead to clear improvements with an architecture and are not good for others I was looking to say there are remaining at the masses pp different plans open to optimize the inference for example on Strixalo on the hardware that I have and one of the recurring comments that I have seen from the maintainers of lama cpp is yes it is fine but this thing that you are proposing is not an improvement rather vaguely worsening on this other architecture or until the let's try on all of them, let's not take it until these problems are also resolved here clearly we will have n.000 parameters n.000 approaches and no because then on the one hand you add the complexity of this thing the cpp is an extremely complex project on the other hand you also add some of their strong choices more or less agreeable for me not agreeable which is the first thing they write in the contributing .md and that they don't accept pull requests completely or mainly AI generated which frankly leaves a bit of time in the current times and in the type of project that is, these people make AI and don't want AI to be used for their project, it seems like a bit of a contradiction in terms to me and it's the thing he wrote on there is a need to make an inference engine that is optimized in particular for an architecture in his case that of the Mac Mini and he says okay I'll do it and I don't know if it started from a fork or if it started from scratch because the software is not yet public he said it will obviously become public but at the moment it isn't however VLLM blade CPP or blade all these that try to keep all types of hardware together certainly have an extremely higher complexity as it is for Linux the example of Linux is fitting because even in PC architectures there are 1500 differences and in fact the audio never works look while we were talking about it and it wasn't where I wanted to get to but I found myself thinking that perhaps there could be room for a company that does what Red Hat did at the beginning of Linux that would worry about keeping track that all these little pieces converge from time to time and then someone tries them all together and guarantees you that that's what works what Red Hat does with VLLM it requires a scientist a little different in q This is a case where each of these algorithms that we talked about today will become a consumable aspect of the library and therefore a building block. However, I can't imagine that there could be a demand for something of this type. Yes, Red Hat is betting on it. They have acquired the main company that contributed VLLM to do these things. They are making it their AI offer. One of the predominant ones, Daniele Zunca also told us when he came here for an interview, is precisely that of giving a VLLM for enterprise. So how was Red Hat Linux for enterprise? yes I wanted to say that again you see how the research fades into software engineering in making all the specific ideas for a given case etc. well integrated with the right levels of abstraction etc. so that they can be used by various types of software yes yes yes absolutely we are at the phase on that part there at least we are at the software engineering part we are not the research still exists and it is certainly fundamental predominant but we are in the software engineering phase as yesterday there was on Wednesday there was there was the annual convention which for them is half-yearly in reality of anthropic code and then I heard the keynote at a couple of talks because for us it was evening or night no disturbing news in the sense that they made a list of all the disturbing innovations that they have made in the last period there was no model announcement there was no real announcement of new modes even if there are many rumors around a new mode that should appear in cloud code which the leaks call it orbit which is essentially proactive cloud code a bit like open cloud or Hermes agent which we will then talk about because I am super enthusiastic about Hermes but nothing sensational but a lot of we have engineered this part here so that cloud code is more easily usable with Sonnet having only Opus as advisor for example they have focused a lot on that part there of another part and a whole series of things that we are starting to see how at least in the part of the harnesses of the coding agents or harnesses in general we are moving a little bit towards the software engineering part and also there are those who say that the so-called EGI is no longer something that only does the supermodel but he is a supermodel with an engineered with an excellent inference with an excellent tool properly integrated with everything else let's say and therefore integrated yes because then in the end you know when you say the EGI it means that he must be able to do everything that a man does therefore he must have the same tools that that person has otherwise where do you have but this also brings us a little to the discussion you were making before about the hardware there is an announcement from Google instead which instead is preparing for the Google I.O which is in a few days and if you remember Google tends to make announcements in the two or three weeks preceding Google I/O and then at Google I.O. there are no announcements but there is the presentation of what has already just been announced and one of the announcements which is not so much about the product is not so much consumer but is that it has a selected group of customers selling TPUs which is actually a big announcement in our world because selling TPUs means that envy is a real big competitor because AMD has tried but the distance in the data center world is still very high the TPUs instead of Google that run Gemini, let's remember, are of a level comparable to the Nvidia GPU as regards inference at least on training there are conflicting opinions there are those who say that they are very good even in training those who say less however between this announcement and the announcement of an agreement instead made with Antropic to provide Antropic with data center power for 5 gigawatts Google made a leap on the stock market the other day these are not financial advice but just to underline it how the market realizes the importance of this two announcements that is that it begins to also provide others but in a significant way because 5 gigawatts is a lot of stuff and the fact that you are also starting to sell hardware is also a change of strategy a bit by Google that has always made hardware but has not always sold the service but has never sold the hardware as such so perhaps also something to point out for 30 different ones that are starting to appear and if TPU arrives we go back to what you were saying Paolo before TPU is another architecture on which to add parameters to be set actually sorry you made me come in mind but I should do an internet search for the details that Google years ago had actually produced a USB TPU that had become widespread and I no longer remember what it's called I have it for the Raspberry it's called it's called it's called I don't know if you want to open the box but I didn't even know you had it I've always found them interesting too for doing inference on the edge then in reality I read that at the time in short it was a bit of an embryonic thing it worked but not like that it was no longer a game yes so in in reality they had already gone through it not with probably a business vision like this but more to say let's try it seemed more like a research thing it seemed more to give it to enthusiasts yes it's not for the server side here it was a bit of a toy the ones we talk about now are the giant ones yes yes which I also came across in my path this week because I was always trying to antivocal to see if we could make a more optimized version of some of the models in particular I would like to fine tune parakits with conversational Italian because a apparently they trained him on the formal and set language of journalists but not on friends who slur words at you and so when you get a real voice from a friend who doesn't know what the fuck to say to you it doesn't work so well whisper is better on this side and so I said okay come on let's see if it can be done and I entered a rabbit hole so yes it can generally be done but because of the technology that Nvidia used in that case it can be done if you have Nvidia hardware and blah blah blah in short in the end I don't know if I succeed but in my discovery I discovered that on Google collab in addition to traditional GPUs you can also choose the TPU and the model prediction which told me that if I had done the same operation with the TPU I would have had a gain of 4x compared to the times made with the GPU which was not bad I didn't get to the end because the step before getting to doing that didn't work so I haven't seen it but sooner or later I will solve my curiosity to try to play with it and see a little how it behaves yes then this this is certainly one of the aspects is also understanding how these things behave here when they put them to the test seriously, however all the optimizations that Alessio was talking about including those on Gemma 4 which are actually in their house are trying to go a bit in this direction here among other things always staying in Google always for the ads in preparation for the Io it seems that they are testing a model called Omni which should unify Veo with Nano Banana therefore generation of video generation of images in a single model which is something a little in somehow waiting also because the rumors said another news that I read that in reality it is true that OpenAI has discontinued Sora you remember the video generation model etc. etc. but they also say that in reality Sora was discontinued for that video generation thing but in reality chat GPT image 2 it's not that Sora has downgraded only making images with optimizations on this thing etc. etc. which all in all it would fit because objectively the results of Sora were remarkable from a point of view visual and those of certain GPT images are notable yes then I read the announcement I honestly haven't yet found much information on Omni it is interesting that we are talking about a multimodal omnimodal model I don't know how we want to define it which also generates videos because typically these multimodal models accept various things but then they essentially generate very frontier images or text let's say interesting let's see let's see if they absolutely amaze us it's all to be seen clearly because then between saying and doing there is always a middle ground let's go back come on let's go back to the world for a moment of politics now we have already lost everyone first we lost all the listeners who wanted technology talking about politics at the beginning then we lost all the others we killed them with with the whole part about inference let's go back to politics and I would like to go to one of Paolo's favorite characters if I remember correctly it is Elon who clashes with Paolo's other favorite character which is Sam it is true that you don't know which sticker to choose yes yes yes I confess then yes then for those who don't know what's happening these days I have no idea if it's being talked about in Italy Moreover, I don't watch television or newspapers much so I have no idea if anyone is paying attention to this news but after much bickering we have finally gone to court for an issue which, if you like, with which I also exceptionally agree with Elon, that is, the criticism's attack, Elon Musk's attack on Sam Altman for which he decided to transform an open AI non-profit into a for-profit company to make a business out of it which doesn't sound very good to me, that is, if you are born with a non-profit and then discover that you can make money by not doing it another company but transforming yours in short it's not like you're really on the list of my favorites now maybe I'm fascinated by this whole affair because I just finished watching a legal series on Netflix Lincoln Lawyer and so I'm well taken and I can't wait to see these things and already on my YouTube feed there are a lot of videos that show how the day in court went these things but I haven't had time to dedicate myself to them yet I trust that I will hear echoes of them through American satire programs through John Oliver Colpere and the others so I hope sooner or later to have juicier things what I've read now is that Ailon isn't doing very well first of all the criticism is that he probably doesn't give a damn about the ethics and morals underlying it despite him having put about thirty million dollars into a company that was supposed to be a profit and then changed but it seems that there is a bit of background to questioning Elon's good intentions because if on the one hand he says ah we are worried about the fact that the IGA will kill us all and these things we need a bit of ethics on the other hand it turned out that Elon had tried to buy OpenAI and spread it with Tesla to have everything in one but they told him no and apparently he was a little resentful so he said then I'll take the ball away and so there's this part here then precisely it seems that Elon is also in court a little erratic as they say in English so I don't know what the most appropriate translation in Italian is, a rambling guy, I don't know and when he talks he skips a lot, so much so that serious lawyers like those of the series I watch are roasting him a little because he says what comes to mind and not what he is prepared for so if we base ourselves on the initial conversations it is difficult for this case to end well for Elon one of the revelations that I found most interesting, not necessarily true but more interesting is that one of the reasons that Elon cites for having invested in OpenAI at the time when it was a non-profit that had the aim of bringing the knowledge of AI to humanity therefore absolutely noble was that he had had a chat with the two founders of Google and they had said to him he had asked him listen but aren't you worried that AI will wipe us all out and one of the answers attributed to one of them is apparently this oh yes but it's not a problem as long as AI survives the extermination of humanity and this comment if it's true obviously makes me get a little nervous and I can understand Elon's concern now it has to be said if it is ever been true though whether it was ever true or not however that someone not necessarily them could imagine it like this is quite disturbing and I don't think that the current trial will lead to anything useful in this situation but it makes me slightly more worried that there could be the mad scientist who for the love of science loves the phantom of the opera more who loves computers more than humans and who therefore will bring this thing anyway I imagine that other new super-fucking revelations will come out of the trial of Elon and Samatman and so I will keep you updated as they happen they come out like this instead of continuing to hear only about Garlasco's murder we also hear about the OpenAI trial, no but then that statement there by one of the two founders of Google could easily have been made by Sergey Brin because he is a very slightly over the top character in his statements so I don't believe it enough that someone could have said it who then didn't and I also think who between the two is but certainly whoever has Google's AI in his hands at the moment, Demis Hassabis, has different visions and does nothing but reiterate it in his definitely very human centric interviews more than all the others even more than the good Darione but then I saw another thing Paolo what a self-fulfilling prophecy let's buy let's buy the servers from Lidl so let's buy the servers from Lidl then let's self-quote an episode in which I have no idea what it was I don't care to go and look for it but joking in the past we talked about Chinese models and the fact that they arrive at a low price and all that democratization and we were joking that in the near future someone will buy the model on Aliexpress which is our equivalent of buying it at Lidl and guess what something vaguely similar is actually happening in Europe driven by the perceived need god thank you from the rulers at European level who say maybe it's not that we can really trust the good intentions of the United States our historical ally who always do the right thing and of their companies perhaps it's the case that we start thinking about having a proprietary infrastructure that isolates us from the risks of being excessively dependent on someone from third parties a parallel if you like to what has happened with fuel with Russia at the beginning of Russia's war with Ukraine realistic if you ask me as a scenario and then they said let's organize ourselves let's try to invent something and the first infrastructure that was missing was that of cloud services because right now AWS dominates Microsoft and Google mainly probably also Alibaba but no one ever remembers them and strictly European there isn't much and also the various local European regions that have been guaranteed it was noticed recently how sorry they are I activated Google on the phone which heard me speaking bad about him and he replied I said we don't have anything European and even though there are regions officially with European data centers they are subject to American law and American law someone has finally gone and looked he says that the cloud act if the government needs it for national security issues which apparently happens for every bullshit now in the United States it needs to go to the providers to ask them for your data it doesn't matter where you are under what legislation they have given it and this is a true thing so true that in addition to having been noticed by the European community it has been noticed from someone close to the Europeans, i.e. Switzerland which ended up on the front page because they severed a contract worth 200 million or more, I don't know, I don't remember these fantasy numbers with Microsoft, I think precisely because it couldn't guarantee them this thing, basically when they asked and verified they told them no, your data is very private except for that time when they ask us for it and since the Swiss care about these matters, traditionally having one foot in both shoes, privacy is, let's say, theirs, their brand, they took this situation directly and that's it. announced this this cut in those parts how that will be resolved I still don't know in the sense that I don't know if they will look at a local solution, their European solution or someone else's solution or they will force their hand over there to be able to change the rules for them I don't know but the fact is that we have reached the front page with this thing and it is a question therefore where Switzerland moves Europe also moves Europe has moved in a slightly different way or by issuing a tender to go and ask for European commercial initiatives of being able to provide a cloud service that Let it be dignified, this thing was done in a very democratic way, so much so that instead of emerging a supergroup, a coordination of someone who can give you vague confidence that he is capable of doing it, I don't know, I would have dreamed that Cern would do it, ok we are in Switzerland but someone with a name you can trust instead no various more or less random actors have advanced in Germany in Holland in Holland in other places nothing from Italy because we won't have the postal cloud at least we got one right and one of these is a linked company at Lidl it is not exactly Lidl that is, the title we say Lidl the Lidl cloud is actually a sort of service company that was born as a sub-company for the needs of the Lidl group and then specialized if you want the same relationship that existed between Amazon and AWS but the name came from there and therefore somewhere in a while there will be the cloud of llidl which will be our variant compared to the AWS cloud we joke about it in reality Lidl does its job well it's more to be seen in my opinion the real bet is not so much who manages it but more to see if the ambitions of reaching a scale that can be self-sustaining which is if not at the same let's say the same level as that of AWS but which goes in a direction that makes you have confidence that Europe can build on this is a good thing and let's see if we are capable it is a good and fundamental thing because it's also there I've mentioned this thing many times I also believe here on the podcast and I've certainly talked about it in some interviews with some of the guests there's also a possible scenario that's almost probable today I'm going too far to say that it's almost probable that at a certain point the so-called agents so from here I'll tell you about mine as an agent no joking aside the so-called agents when they stay inside the cloud etc. will reach a point where they will produce value ok they'll produce value it's something that Draghi also says I don't know I'm not sure in mind as I go and when it is an agent that produces value there is a possible scenario in which what is taxed is the value produced traditionally in the European economy the tax is placed on the value produced not so much on those who produce it but because it is more convenient to tax the person than the value but in reality what you tax is the value produced and if the value is produced on a soil other than the European one it could be subjected to a taxation that is different from the European one so the problem is starting to become complex also of sustainability of the welfare of which Europe it rightly has the flagship for which to look forward and be able to produce digital value on European soil is something that is not only important from the point of view of privacy which is of great interest to Europe and has always made it a workhorse that's fine but there is also a much more practical question of where the money ends up then yes this is a legal quagmire not to be laughed at all the issues related to cryptocurrencies to taxation come to mind understanding where the funds reside you are in Europe in the United States where in reality they are not I am in the light the tangle, however, is actually a tangle and as far as the blockchain is concerned it is more difficult because it is a distributed thing and it is built to avoid that thing there in reality one of the ideas is that who is doing and who will do agent-based services will tend to put them more in a cloud where yes that problem exists there because there are many points etc. but a little more manageable and it is in everyone's best interest to get to define who produces who dares as value so perhaps it is a slightly less complex problem but which can become bloody if not do you have anything on Italian European soil that produces value or in any case be able to offer the possibility from the point of view of those who set up this thing here to say okay but all this goes around here so I want it to be treated with the laws over here okay exactly exactly this to point out the thing but talking about agents and also talking about locations of doing things locally I told you that I set up Hermes agent in my house on an old computer using a Chinese model like this so as not to miss anything and it is doing things it is doing things which I will also talk about later in the newsletter that comes out on Monday but that I wanted to tell you about, maybe also show something for those who have gotten to the end here at least sorry if I interrupt you just to clarify you are running the agent locally with the model however offered in the cloud by a correct provider because I don't have a machine on which to run a model with sufficient performance to make it do these things so the model is cloud but instead the agent the tool to use the technical term which is Hermes agent runs locally on one of my machines and the files it modifies the things it does it does it there there but in addition I gave him access to a whole series of my services. I told him about it last time also to my email and so on, trying to take all the precautions possible. I chose Hermes and not open cloud for this reason because it's a little easier to customize access. There's a very clear skill system about which for example I said it the other time about Gmail. I told him okay go to Gmail. Let me see how it goes and he'll tell me. I showed him his script and I told him here are those two methods there, the ones that do delete and the one that does send of the email, we delete them right from the script you use, it's not that I'm telling you not to use them, delete them, you don't have to have them anymore and then I checked that he had done it and so and so it is and what I make him do I make him do a lot of things from the most banal things if you want of cyclical use so in the morning or every hour he checks my email he tells me which emails are only informative and he gives me a summary of a line which instead need my attention by going to read the email and therefore understanding that I need my attention and furthermore for those that require a response I have him generate a draft of the response he doesn't send it I then find them in gmail already drafted then I edit them but I have a starting point at least and he looks at the calendar it tells me which are the appointments of the day and the important ones of the week for me the important ones are those of a certain color i.e. we discussed for a moment how to define these things I find the gaps like I say to him eh I have to talk with Alessio we take an hour you find me a hole in the calendar he makes me three or four proposals I tell him that one is fine and he sets me the appointment management things of this type he manages things in my house so I gave him all the various smart sockets everything smart I have in the house which is not as much as Paolo but I have something too eh and so like from outside the house I tell him to turn on the air conditioner or the heat pump eh to check the lights the cameras these things here eh and then and then he does things then the access is with telegram for me he does quite proactive things that is in the sense that the other day I told him hey look at my personal site eh a personal site like an electronic business card no hey what do you think the answer wasn't very kind but it doesn't matter eh more or less at the level of the answer on how our podcast thumbnails were yes yes yes more or less eh no he was kinder because even there you tell him how he should behave I recommend being extremely frank and extremely direct and he told me yes I understand the attempt but I don't we're here um and I tell him okay what do we do eh this while my son was at basketball training what do we do what don't we do and he says look I'll give you three proposals tell me if you like more of a dark version as a developer or another ending I choose the dark version and he says let's do it like this I'll make you a new website directory I'll push it to you I'll deploy it and then you take a look at it and tell me how it's going sorry I'm basically he told you but what am I asking you to do after you expressed your choice to him that is he was doing so disgusting your preference that he said let's do both of us then decide later no no no he did to me what I asked him but the incredible thing that I can share a slide of because there is the new website wait while I open it while I talk the thing that he did to me left me there is that in other conversations for the podcast for the newsletter etc. etc. I had given him in the newsletter for example an agenda part in which I say the conferences which were the ones who published the video the ones where I will go and anything else and he had this information because we had already talked about it he has a memory management done very well better than others hierarchical and so on and he generated this thing for me here that even if you like it you don't like it but then these three my contents let's say that they are the two newsletters and our podcast were on the old site so it's not his merit the projects were there on the old site and therefore it's not his merit of these I only had the last one the devox because I had done it by hand instead he saw that there were all these talks that I had done that I will do and he put them on the page for me taking both the slides and the videos that everything else is right and this thing here about taking initiative struck me quite a bit it's not the initiative that also who cares that they attributed open cloud to do things alone but initiative let's say intelligent another thing so let's also talk about that topic there is one thing that we had put in the lineup which is an article wait while I recover the title because I risk saying nonsense that of the 100 million what is 12 million context now I don't have it here at hand but c 'It's an article we talked about Alessio right we talked maybe in chat we don't have it yes I haven't looked into it in depth but there is this article that talks about a start up that has created a 12 million model with a long context of 12 million tokens since there have been announcements in the past of people who made 100 million tokens which then didn't go anywhere I wanted to understand it before talking about it for a short time and I said to Hermes I need to understand this article I'm in the car so give me a well done summary him he left and asked me ask me for permission to do anything but because that is my choice I was in the car not driving he is precise and he tells me well but if you want if you give me permission to install a text to speech I will read it to you it seems like a good idea this foul a text to speech model has been installed he told me in the meantime I'll give you a summary in Italian that maybe I didn't tell him that I wasn't driving if you're in the car and you're driving it's better if I do it in your language I'll give you a summary in Italian and I'll read it to you out loud and he did it I'll share it with you for a moment if I can wait until I see that I've actually opened he sent you a vowel so he sent me a vowel yes yes he sent me a vowel then I already sent him the vowels before because speech to text has it by default but now here there's a lot of stuff because he sent me my calendar wait where do you have the vowel sounds wait I can't find them I wanted to let you hear them so here's the reading maybe a little less sexy than others now I can't do it for you hear because we had already tested that things from Telegram don't work, the audio doesn't work, that is, you hear it but then I don't hear it, the listeners don't like it, I'll tell you that the reading is acceptable, it's not the voice of OpenAI but that of Eleven Labs that you can put on, he immediately asked me, do you have an Eleven Labs account, if you have it, I'll use it, no, I don't have it, I don't pay for many things, so I did this with a local model, it took me a few minutes and he sent me these two summaries with two different points of view and so from what he told me I understood a few things I told him yes but listen give me a summary in tabular form while I'm not choosing the right window to scroll ok give me a summary in tabular form of this data and he wrote it to me like this and I told him no I want a table because I don't understand much when I read like this and he made me a table because I told him that I always want everything in markdown and I told him again no I don't understand shit I'm in the car I can't see it and he made this crazy choice he made an html to give me the table done well and then he said to himself oh no but the user told me that in the car how does he read an html wait I'll render it I'll take a photo and send him a photo and he sent me the photo of the table then if you want we'll go into the detail of the paper but maybe not today eh well you and then oh well after that he made me another one on another topic remembering that I told him I couldn't read it and then he gave it to me again like this and in this case it was to me that stuff here was quite wow then I appreciate it but I had already seen it in short in my long conversations via telegram with simply ZAI GLM I too had come to those conclusions in this case perhaps he was proactive that he told you all of them in my cases instead I gave him promptings in the .md cloud telling him look I'm around so when I'm around I need short stuff he uses emojis these things so I he did similar things but I appreciate and understand the wow effect he did no I am particularly satisfied also and above all by the fact that he has access leave aside the cloud services also because being able to access the email to the calendar is a remarkable thing i.e. being able to send a voice message while I'm out and about and say I remembered that you have to add this task and then you can find it in gtask that's a lot it's cool ditto for appointments on the calendar and things like that but also that it has access to the file system for example I told you about the llm wiki the one from carpati that I use to think and to analyze the papers at the moment I am looking at the papers on long-term memory and in this case more to save tokens and time than anything else there is the phase after I have collected all the papers there is the digest phase which means that you download all the papers in pdf the law generates the wiki creates the links etc. which still takes half an hour 40 minutes this on cloud code not only him and I did it in the morning but it annoyed me I launched it and I did another part the number of non-trivial tokens that he uses to do this thing that you pay for and so I did it anyway with glm and so it took a while now I have a task assigned to him that does the Arabase of the github where I told him the things in the evening we do the whole digest part of reflect and when I'm done I create a pull request and I emerge the pull request and then I work on the things already digested and this is a great time saver last but not least and you may have seen it on lince I had him do the pull request review of things I had done with cloud code and there it was fascinating because if they played it they sang it because cloud I told him look you have a pull request review I went to see it and he said ah yes well done this it's interesting I'll do this one I'll do it he committed and then he put a comment no I won't do this one because it's more of a waste of time than anything else I don't know what you saw in the code but you're wrong and so it was nice that they played and sang it which is a bit like what you do when you do slash simplify locally in fact you do a pull request review on your own so full of enthusiasm for Hermes I can't live without it anymore you made me think of a use case for which I need this thing so I'll publicly inform you that I'll also fake communion, I'm the tempting devil, yes it's true and well then if you don't want to talk about that very complicated paper that I have prepared for but the time could be on average long I think that we have more or less reached the bottom of the list let's just say this thing that always in the world of temptations the idea of switching to codex because there are little animals that talk while you write the code is something that I'm trying to resist but I don't know until when I'll succeed listen let me see a screenshot because as I was telling you I activated that thing a while ago I don't have it ready wait I have it let's open the next episode with the screenshot with the animals let's open the next episode with the screenshot I don't have the screenshot ready okay I have it on codex but there is a small subscription and I don't use it much but I tried it and I have to say that in short having the animal because it's a well made animal it's not like the little ugly one that the one in the code had made this is really nice come in he walks in the middle of the screen he also breaks balls a bit if you want but so the metaverse of what's his name from Zuckerberg had his why and the only why was for the programmers to show them cartoons while they program yes no why am I saying it because I tried a bit of codex because I'm trying it with the inchessac there's a lot of new things go and see them on the site I'm not going to tell them here Claude and I have done a lot of things in recent days including the paranoid mode he has from fans so go go and see it if you want okay we say goodbye we say hello to the public see you next time numerous put bells little stars subscribe to the channel and don't miss out watch the shorts don't miss this week's cover which will see Alessio on the cover yes it's up to Alessio it's up to Alessio then I was also saying that not only it's up to Alessio for the cover then we'll try to do the covers again so maybe a little less Bruce Willis because the last cover with me everyone told me he looked like Bruce Willis when he was fit you take what you're told I've been there put on a good face I activate the game and you take everything you're told well good now it's Alessio's turn I'll go straight away to see how to make the cover bye everyone bye bye
Ciao a tutti e tutti, bentornati, bentornate. Partiamo, allora, tante cose da Elon Musk vs Sam Altman, interviste improbabili a Claude, poi cose tecniche tra cui Gemma 4 superveloce e poi tante altre cose. Dai partiamo che ne abbiamo mille oggi. Da dove partiamo? Partiamo dalla tristezza? Mai di tristezza dai. Dai partiamo dalla tristezza, partiamo dalla tristezza che sono le interviste improbabili, che non sono le interviste improbabili quelle simpatiche che facevano, non mi ricordo più chi, forse Chiambretti. Mai di regolar. Eh? Mai di regolar. Mai di regolar, sì sì, mai di regolar, le interviste improbabili. No, sono quelle che fanno ai modelli di intelligenza artificiale. Abbiamo parlato in passato di quella di Bernie Sanders, e l'ha fatta anche Veltroni. L'ha fatta anche lui a Claude, credo. Perché è il Bernie Sanders di Noartry. Ne abbiamo parlato qualche settimana fa che in America Bernie Sanders, pure senatore, si è messo in intervista a Claude, e un ex candidato alla presidenza del Consiglio, che fa il giornalista per il mestiere principale, però, cioè Veltroni, ha deciso di intervistare anche lui Claude, con un'idea così innovativa, no? Che sembra quasi... A me, allora, quando l'ho letto ho proprio pensato che è come Little Tony quando faceva... Elvis. Esatto, sì. Il Little Tony che fa Elvis è come il Troni che fa Bernie Sanders. Allora, no. Come era no quella di Bernie Sanders, più no ancora. Non è quello che possiamo passare alle nuove generazioni. Erano, ehm... Cerchiamo di capirla, questa intelligenza artificiale, non... Non di umanizzarla per forza, no? Di chiedergli tu ci distruggerai... Er... La cosa che mi ha fatto quasi... O cosa pensi del fine vita o qualcosa del genere. No, la cosa che mi ha fatto più tenerezza, quasi tenerezza dell'ignoranza, nel senso letterale del termine, eh, se puoi, Veltroni mi denuncia pazienza, però, dell'ignorare la cosa quando gli chiede se fai errori. E chiaramente quello che gli dice, sì, sono pieno di lacune, faccio errori. Tristerrimo, eh, nel senso... Sono strumenti, usiamoli come tali, non umanizziamoli, non chiediamo loro di se stessi. Il discorso della coscienza, lungo discusso da Antropic, eccetera, è un interessante campo di ricerca, ma lasciamolo nel campo della ricerca. Perché poi dopo me li leggo, no? Già i giornali, che è l'intelligenza artificiale umanizzata e che quindi porta i giovani a fare le cose negative, anche le peggiori. Eh, però se quella è l'immagine che cominciamo a darne senza aver capito che cosa abbiamo in mano, così. Come al solito, opinione abbastanza forte la mia, però no, veramente no. Per me un no. E dove l'hanno data? In televisione, in prima serata? Eh, no, no, ho intervistato su un giornale. Ah, peccato, perché era una di quelle robe da mettere sulla TV nazionale o quelle cose così, secondo me. Eh, ma ci arriveremo. Adesso, non so, ci si metterà a mara veniera a intervistare il PT in prima serata. Però, guarda, tra il serio e il faceto, visto che raccontavi che tua figlia deve fare la maturità quest'anno, magari diglielo di prepararsi sulla traccia delle AI per il tema, che secondo me... Ah, no, no, ne abbiamo parlato, ma sì, è probabile che lo diano come tema, ma quello ci starebbe anche. E credo anche che i ragazzi direbbero cose più intelligenti. Lasciamo stare i politici, i ex politici italiani, tocchiamo quelli americani. Direbbero cose più intelligenti di Bernie Sanders e anche del suo presidente della Repubblica. Perché anche lui ne ha dette, eh, in questi giorni. Avete visto? Mi riferisco a Trump, che ha detto... Poi veniamo via dalla politica e torniamo alla tecnologia. Però mi riferisco a Trump, che ha detto che vorrebbe, stava allutando il potere di veto al rilascio dei modelli di intelligenza artificiale da parte della Casa Bianca per la pericolosità reale o presunta. Che, cioè, immaginerete che OpenAI, Google e Antropica non l'hanno presa benissimo proprio... Sì, anche perché poi mi verrebbe da chiedere sulla base di cosa prendere la decisione, cioè... E chi è che la valuta che è pericoloso? Hai ragione. Hai assolutamente ragione. Cioè, quali esperti... di quali esperti si dota la Casa Bianca della situazione per avere capacità di discernere su quello che i ricercatori di Anthropic piuttosto che di OpenAI hanno fatto? Non era Elon l'esperto, scusa. Eh, o anche di... Eh, ma adesso c'è un po' di maretta, quindi... L'esperto potrebbe diventare OpenAI, il quale sicuramente favorirà i rilasci di Cloud. Esatto. Cioè, boh... Diventa veramente... Quasi al limite del ridicolo. Se no intervista anche lui le IAI, gli chiedi cosa pensi di quell'altro rilascio che sta arrivando. Ah sì, anche... Mi sembra... Mi sembra fondamentale. Veniamo via dalla politica, dai, che ci mettiamo. Stiamo soltanto pestando dalle cacche. Allora, no, nuovi modelli, dai, parliamo di nuove modelli. Partiamo da... A casa Google. Andiamo un po' a casa di Alessio, l'inferenza, eccetera. Gemma 4. Hai visto che hanno fatto quello che hai scritto tu nel tuo ultimo articolo di Aladino? Che leggono forse. Sicuramente, guarda. Ma allora, intanto... Vediamo se riesco a condividervi qualcosa. Allora, cosa è successo? È successo che appunto io parlavo di speculative decoding, siccome il mondo mi legge, come no, anche Google ha pensato di pubblicizzare questa tecnica. Infortuna che hanno il traduttore automatico. Per l'altro fanno loro la traduzione, per cui... Esatto. No, a parte gli scherzi. Prima vediamo questo discorso di Gemma 4, poi se mi concedete un attimo vorrei fare un ragionamento un attimo più ad ampio respiro su quello che sta succedendo in questi giorni, in questo campo. In Gemma 4 hanno deciso di abilitare lo speculative decoding, che sarebbe una tecnica per velocizzare la fase di generazione dei token, quindi le risposte, quando si interroga un modello. Quindi non la prima parte, che è quella di comprensione del prompt, ma la parte successiva di generazione della risposta. Come si fa questa ottimizzazione, questa velocizzazione? Ci sono varie tecniche e un gruppo di queste tecniche si basa sull'utilizzo di modelli drafter. Praticamente sono modelli più piccoli e di conseguenza più veloci, ai quali viene chiesto di fare delle previsioni sul prossimo token o i prossimi token da generare e il modello target, che sarebbe il modello grande con cui si sta lavorando, invece di fare lui la generazione del prossimo token, fa prima una valutazione della previsione fatta dal modello piccolo. Se il modello piccolo è stato sufficientemente bravo a prevedere bene il prossimo token, si ha un risparmio di tempo perché la valutazione della previsione è sostanzialmente meno pesante dal punto di vista computazionale rispetto a effettivamente calcolare il prossimo token. O comunque, come si fa in questo caso qui di Gemma 4, è possibile parallelizzare e valutare sostanzialmente più token in una singola passata. Quindi nel momento in cui il modello piccolo ci prende si ha un grande guadagno. Quello che hanno fatto quelli di Google è stato sostanzialmente ottimizzare molto questa idea e come l'hanno fatto con un modello molto piccolo allora per dire mi sembra di aver preso qualche punto il modello 2 billion di Gemma 4 comunque già è relativamente piccolo ha un modello drafter da 76 milioni di parametri quindi 2 miliardi 76 milioni quindi estremamente più veloce e questo modello piccolo drafter genera 4 volte diciamo i token che andrebbe a generare il modello grande e il modello target fa in una passata sola la valutazione di queste di questa generazione. Per migliorare ulteriormente la cosa cosa hanno fatto? Si sono sostanzialmente inventate dei trucchi tipo quello di condividere la KVCache dei due modelli quindi il modello piccolo attinge a lavorazioni che già ha fatto il modello grande per per la per la cache e in più nel momento in cui si calcolano gli embedding da cui parte l'operazione di generazione per il modello draft questi embedding sono appesi concatenati dopo la il risultato diciamo le attivazioni dell'ultimo layer del modello più grande quindi è un modo per consentire al modello piccolo nonostante sia appunto dotato di pochi di pochi parametri quindi poco intelligente mettiamola così di partire da un da una pre lavorazione del dello stato attuale a cui era arrivato il modello grande questo ovviamente se si va a leggere il paper è spiegato molto meglio questo consente sostanzialmente al modello draft di prenderci abbastanza spesso e però e qui vi racconterai un attimo diciamo le mie pensieri di questi ultimi giorni tutta questa cosa si inserisce in un in un ragionamento molto più esteso cioè stiamo notando la ricerca che sta affrontando le problematiche di efficienza dell'inferenza sotto in vari in vari campi questa che abbiamo appena raccontato a che fare con la velocità di generazione speculative decoding non è soltanto questo questo approccio usato da google che tra parentesi si chiama mtp multi token prevision ma ci sono altre altre tecniche tipo quella di cui parlavo nel mio articolo che è ngram che consente sostanzialmente al modello di andare a vedere che cosa ha generato negli step precedenti e fare delle previsioni basate su quello è possibile fare abbinare modelli draft sviluppati diciamo in modo indipendente rispetto al modello target che si sta utilizzando chiaramente devono essere accoppiati bene cioè non è che si può prendere un qualunque modello draft ma senza anche che siano embedded come in questo caso di gemma ci sono ricerche per creare modelli draft per i modelli quen ad esempio con varie tecniche e tra l'altro appunto c'è una tecnica che si chiama D-flash che è abbondantemente diciamo ricercata in questo periodo che consente nell'utilizzare modelli di diffusione non so se vi ricordate che mesi fa ne parlavamo anche in podcast si era accennato all'esistenza di modelli per la generazione di testo non autoregressivi ma basati sull'idea della diffusion la stessa che si usa per la generazione delle immagini e questi modelli sostanzialmente fanno come nel caso dell'immagine una riduzione del rumore a partire da un qualcosa che rappresenta il totale rumore e generano diversi token in parallelo tanti token in parallelo questa quest'idea questo approccio è esattamente quello che si abbina bene alla costruzione di un modello draft che faccia le previsioni dei prossimi token perché il minus dei modelli di fusione era proprio quello di essere veloci ma non esageratamente accurati a confronto con i migliori modelli autoregressivi e questa è esattamente la condizione in cui siamo adesso con i modelli draft quindi ci interessa la velocità siamo disposti ad accettare una minore qualità perché poi ci sarà il modello target che andrà a valutare la previsione quindi esistono esistono dei modelli draft che sono in fase di sviluppo al momento proprio per fare questa cosa che si chiama di flash nel frattempo la ricerca sta cercando di affrontare anche ovviamente la fase di KVCache quindi la cache avevamo parlato di TurboQuant diverse settimane fa sono usciti tante altre idee di ottimizzazione della cache dove l'obiettivo ovviamente è ridurre l'occupazione di memoria quindi consentire l'utilizzo di modelli relativamente grandi anche in caso di poche risorse e risorse di memoria è uscito tra le varie altre idee di ottimizzazione del modo con cui si fa la cache una cosa che si chiama RotorQuant che sostanzialmente va a cercare di migliorare uno dei difetti di TurboQuant che era il fatto che per costruire la cache nel modo con cui spiegavano nel paper di TurboQuant si sostanzialmente usavano diverse risorse computazionali quindi se è vero che si riduceva la memoria utilizzata si andava comunque a rallentare la fase di prefill l'idea di questi di RotorQuant è abbastanza complessa da spiegare ma sostanzialmente loro fanno delle trasformazioni differenti dei vettori in input li dividono in vettori più piccoli e poi hanno un modo intelligente per processare questi vettori più piccoli morale della favola ordini un ordine di grandezza più veloce rispetto a TurboQuant in diversi scenari di utilizzo bene quindi ottimizzazione in generazione in diciamo sulla parte di di cache ma notizie degli ultimi giorni anche in fase di prefill prompt processing che è come dicevamo prima quello che avviene nel momento in cui si va a processare il prompt la fase di prefill di base se volete l'ottimizzazione naturale è proprio l'utilizzo della kvcache perché l'idea è che uno invece di rifare tutta l'attention va a vedere che cosa è stato calcolato precedentemente però ci di utilizzare una tecnica di utilizzare una tecnica che possiamo condividere qualcosa magari Stefano riesci a farlo tu sì certo sostanzialmente l'idea di lucebox è il gruppo che ha sviluppato questa cosa è di avere un di nuovo modello piccolo drafter che funzioni da scorer cioè che vada a fare una valutazione del prompt dei token del prompt che viene passato cercando di capire quali sono le parti interessanti del token del prompt si parla del problema classico che è quello di trovare lago nel pagliaio cioè capire in un contesto molto grande in cui abbiamo tante informazioni e non tutte sono super salienti capire dove sono veramente le parti più salienti sulle quali effettivamente puoi fare il prefill e quindi a valle di questa fase di scoring viene costruito un prompt più compresso e con quel prompt si fa il prefill vero e proprio chiaramente questo approccio a differenza della di flash e insomma in di generazione questo approccio non è lossless cioè si perdono informazioni nel fare questa compressione ma l'idea è che nello use case corretto ad esempio quello in cui c'è stata un sacco di discussione nel contesto le chiamate precedenti e prompt precedenti e non tutte le informazioni che sono state passate sono salienti perché si sono provate varie cose alla fine si è ottenuto un diciamo un certo risultato è possibile comprimere il contesto in questo modo e avere una velocizzazione del prefil senza perdere sostanzialmente troppo in fatto di accuratezza tutte queste migliorie su questi tre principali ambiti prefil generazione cash sono sono in parte disponibili ad esempio l'ama cpp ma in alcuni casi ci sono proprio fork di l'ama cpp c'è un effort di cercare di unificare tutti questi sviluppi assieme vediamo io sono abbastanza fiducioso e sono rimasto colpito da tutte queste ricerche che stanno convergendo per migliorare la differenza anche locale dammi una tua opinione Alessio in quanto maggiore esperto avendo non elette tante sono molto interessanti queste cose e mi chiedevo a che cosa porteranno porteranno a una meglio di tecniche e quindi lo stato dell'arte o in realtà siccome alcune vanno in una direzione altre vanno in un'altra ci sarà sempre spazio per algoritmi e approcci divergenti come c'è adesso con le strutture dati o gli algoritmi in generali per cui esisterà sempre quella casistica per cui questo è il caso migliore quindi magari ci saranno non so piccoli shop di gente che ottimizza il modello per quella specifica cosa per quella specifica lingua per quella specifica hardware per quella specifica request di slack che deve avere per cui non spariranno tutte queste tecniche ma ognuna avrà il suo perché nel caso specifico quale pensi che sia lo scenario ma secondo me è un misto delle due cose cioè nel momento in cui c'è un un'ottimizzazione che per una determinata fase dell'inferenza è meglio di tutte le altre in per dire il 95% dei casi è probabile che quella diventi lo standard di fatto mi viene in mente varie cose nell'attention ormai il cosiddetto momento di psique in cui loro hanno definito fast sparsa attention poi l'hanno usato tutti perché era nettamente meglio della dance attention che non si usava già più la dance però vabbè si usava una sparsa diversa però ci sono cose che useranno tutti anche secondo me altre che sono più di nicchia ci sono fior di start up che scommettono su questa cosa in testa tutta e mira morati la sua start up scommette sul fatto che le aziende avranno bisogno di ottimizzazioni specifiche e di fine tuning specifici cioè loro non fanno modelli stanno cercando di democratizzare la fase di tuning loro più sui modelli ma arriveranno secondo me anche un pochino sull'inferenza se andiamo come pare in questo momento almeno sempre di più ad avere soluzioni ibride dove l'inferenza locale e l'inferenza cloud convivono e risolvono problemi diversi sì peraltro io accennavo queste cose qua perché sono chiaramente delle tutta una serie di ottimizzazioni che possono rendere utilizzabili modelli belli diciamo che hanno capacità interessanti su hardware non esagerato quindi per l'inferenza locale anche consumer però in realtà tutte queste cose interessano molto anche a chi fa offre in cloud questi servizi perché è tutto modo per risparmiare risorse per se vuoi comunque avere un ritorno degli investimenti migliore perché nel momento in cui tu riesci offrire le stesse performance con una frazione delle risorse investite senza compromettere la qualità è tutto guadagno sì guarda vi confesso la mia ispirazione era pensare al kernel di Linux e all'infinità di parametri e di configurazione che lui nasconde che la maggior parte di noi non usa nel senso usiamo il default e di tanto in tanto tocchiamo le cose più ovvie ma quella cosa lì fa un sacco di roba immaginate che ne so ha tutte le possibilità per gestire la memoria come uccidere i processi oppure l'infinità di file system diversi che esistono esistono chissà perché cioè è giustificato che continui a esistere non lo so qualcuno li usa di sicuro mi chiedo se sarà la stessa cosa con tutte queste tecnologie ottimizzazioni per le AI beh io sai sui motor di inferenza quelli diciamo commerciali o anche non commerciali ma cloud il livello a cui puoi fare il tuning è veramente alto magari invitiamo anche abbiamo tra i nostri contatti anche chi lavora su VLLM magari qualche volta lo invitiamo per parlarci di quanto è fonda la teana del bianconiglio perché poi su quella cosa lì siamo al livello simile a quello che descrivi tu del kernel di Linux e anche meno friendly di così forse perché più giovane come tecnologia però sì no spazio ce ne sarà lungo ancora perché tutte le tecnologie molto verticali riservano tutto quello spazio lì questa è la mia opinione almeno non so se Alessio la vede diversamente no no anche secondo me e anche il fatto che esista tanto hardware differente fa sì che ci vorrà una fase di assestamento cioè avremo per diverso tempo secondo me diversi parametri come dici tu Paolo che portano a migliorie nette con un'architettura e non vanno bene per altre guardavo per dire ci sono rimanendo alla massi pp diverse piare aperte per ottimizzare l'inferenza ad esempio su Strixalo sull'hardware che ho io e uno dei commenti ricorrenti che ho visto da parte dei maintainer di lama cpp è sì va bene però questa cosa che ci proponete non non è migliorativa anzi vagamente peggiorativa su quest'altra architettura oppure finché non la proviamo su tutte non la prendiamo finché finché non si risolvano anche questi problemi qua chiaramente avremo n.000 parametri n.000 approcci e no perché poi da un lato ci metti la complessità di questa cosa la cpp è un progetto estremamente complesso dall'altro ci metti anche qualche loro scelta forte più o meno condivisibile per me non condivisibile che è la prima cosa che scrivono nel contributing .md e che non accettano pull request completamente o principalmente AI generated che francamente lascia un po' il tempo che trova nei tempi attuali e nel tipo di progetto che sono cioè questi fanno l'AI e non vogliono che l'AI venga usata per il loro progetto mi sembra un po' una contraddizione in termini ed è la cosa che scriveva su X questa settimana o la scorsa questa mi pare antirezza Salvatore Sanfilippo che dice vabbè io mi sono rotto non ho il tempo di scrivere tutto a mano e non ha più senso oggi come oggi in più c'è questo discorso dei tanti hardware che da un lato si capisce che Adamas e PPP voglia supportare tutti dall'altro proprio perché c'è diversità lui dice c'è bisogno di fare un motore di inferenza che sia ottimizzato in particolare per un'architettura nel suo caso quella del Mac Mini e dice vabbè me lo faccio e non so se è partito da un fork o se è partito da zero perché il software non è ancora pubblico ha detto che lo diventerà ovviamente ma al momento non lo è però VLLM lama CPP o lama tutti questi che cercano di tenere insieme tutti i tipi di hardware sicuramente hanno una complessità estremamente più alta come è per Linux l'esempio di Linux è calzante perché anche negli architetturi dei PC ci sono 1500 differenze e infatti l'audio non va mai guarda mentre ne parlavamo e non era lì che volevo arrivare ma mi sono trovato a pensare che forse ci potrebbe essere spazio per un'azienda che faccia quello che ha fatto Red Hat all'inizio di Linux che si preoccupasse di tenere traccia che tutti questi pezzetti di tanto in tanto convergano e quindi qualcuno li provi tutti insieme e ti garantisca che è quello che funziona quello che Red Hat fa con VLLM si richiede uno schienzato un po' diverso in questo caso laddove ognuno di questi algoritmi che abbiamo detto oggi diventerà un aspetto consumabile di libreria quindi un building block però non fatico immaginare che ci possa essere domanda per una cosa di questo tipo si si Red Hat ci sta scommettendo hanno acquisito l'azienda principale che contribuiva VLLM per fare queste cose lo stanno facendo la loro offerta AI una di quelle predominanti ce lo raccontava anche Daniele Zunca quando è venuto qua in intervista è proprio quella lì di di dare un VLLM for enterprise quindi come è stato Red Hat Linux for enterprise si io volevo dire che di nuovo vedi come la ricerca sfuma nell'ingegneria del software nel rendere tutte le idee specifiche per un determinato caso eccetera ben integrate con giusti livelli di astrazione eccetera cosicché siano utilizzabili da vari tipi di software si si si assolutamente siamo siamo alla fase su quella parte lì almeno siamo alla parte dell'ingegneria del software non siamo la ricerca esiste ancora ed è sicuramente fondamentale predominante ma siamo nella fase dell'ingegneria del software come ieri c'era mercoledì c'è stato c'è stato il convention annuale che poi per loro è semestrale in realtà di antropica codice e allora io ho sentito il keynote a un paio di talk perché per noi era sera o notte nessuna novità incratante nel senso hanno fatto l'elenco di tutte le novità incratanti che hanno fatto nell'ultimo periodo non c'è stato un annuncio di modello non c'è stato un annuncio vero di nuove modalità anche se ci sono tanti rumor intorno ad una nuova modalità che dovrebbe apparire in cloud code che i leak la chiamano orbit che è sostanzialmente cloud code proattivo un po' come come open cloud o Hermes agent che di cui poi parliamo perché sono super entusiasta di Hermes però niente di eclatante ma tanta parte del abbiamo ingegnerizzato questa parte qui perché cloud code sia più facilmente utilizzabile con Sonnet avendo soltanto Opus come advisor ad esempio hanno puntato molto su quella parte lì di un'altra parte e tutta una serie di cose che si comincia a vedere quanto almeno nella parte degli harness dei coding agent o harness in generale ci stiamo spostando un pochino sulla parte di ingegneria del software e anche c'è chi dice che la cosiddetta EGI non è più una cosa che fa solo il supermodello ma fa il supermodello con ingegnerizzato con un'ottima inferenza con un ottimo ornese propriamente integrato con tutto il resto diciamo e quindi integrato sì perché poi alla fine sai quando dici l'EGI vuol dire che deve saper fare tutto quello che fa un uomo quindi deve avere gli stessi strumenti che ha quella persona altrimenti dov'hai ma questo ci porta anche un po' a il discorso che facevi prima sull'hardware c'è un annuncio di Google invece che invece si prepara al Google I.O che è tra qualche giorno e se vi ricordate Google tende a fare gli annunci nelle due tre settimane precedenti al Google I/O e poi al Google I.O non ci sono annunci ma c'è la presentazione di quello che è già appena stato annunciato e uno degli annunci che non è tanto di prodotto non è tanto consumer ma è che ha un gruppo selezionato di clienti vendono le TPU che è un annuncio grosso in realtà nel nostro mondo perché vendere TPU vuol dire che invidia è un concorrente reale grosso perché AMD ci ha provato però la distanza nel mondo data center è ancora altina le TPU invece di Google che fanno girare Gemini ricordiamo sono di livello paragonabile al GPU di Nvidia per quanto riguarda l'inferenza almeno sul training ci sono pareri discordanti c'è chi dice che vanno molto bene anche in training chi dice meno però tra questo annuncio e l'annuncio di un accordo invece fatto con Antropic per fornire ad Antropic potenza di data center per 5 gigawatt Google ha fatto un salto in borsa l'altro giorno non sono consigli finanziari questi ma giusto per sottolinearlo come il mercato si renda conto dell'importanza di questo due annunci cioè che cominci a fornire anche altri ma in maniera significativa perché 5 gigawatt è tanta tanta roba e che cominci a vendere anche invece hardware è anche un cambio di strategia un po' di Google che l'hardware se l'è sempre fatto ma non ha sempre venduto il servizio ma non ha mai venduto l'hardware in quanto tale quindi forse anche una cosa da segnalare per 30 diversi che cominciano a vedersi e se arrivano TPU si torna a quello che dicevi tu Paolo prima TPU è un'altra architettura su cui aggiungere parametri da settare in realtà scusami mi hai fatto venire in mente ma dovrei fare una ricerca su internet per i dettagli che Google anni fa aveva in realtà prodotto una TPU USB che si era diffusa e non mi ricordo più come si chiama ce l'ho per il raspberry si chiama si chiama si chiama non so se vuoi aprire la scatola ma non sapevo neanche che ce l'avevi che le ho sempre trovate interessanti anche io per fare inferenza on the edge poi in realtà leggevo che a suo tempo insomma era una cosa un po' embrionale funzionava ma non così era più un giocato sì quindi in realtà già ci erano passati non con una probabilmente una visione di business come questa ma più per dire proviamo sembrava più una cosa di ricerca sembrava più per darla agli entusiasti sì non è per il server side ecco era un po' un giocattolino i più di cui parliamo adesso sono quelle giganti ecco sì sì che peraltro io ho incrociato nel mio cammino questa settimana perché stavo cercando sempre per antivocale di vedere se riuscivamo a fare una versione più ottimizzata di alcuni dei modelli in particolare mi piacerebbe fare fine tuning di parakit con l'italiano conversazionale perché a quanto pare l'hanno addestrato su linguaggio formale e impostato dei giornalisti ma non sugli amici che ti sbiascicano le parole e quindi quando ti arriva un vocale reale di un amico che non sa che cazzo dirti non funziona così bene whisper è meglio da questo lato e quindi ho detto vabbè dai vediamo se si può fare e sono entrato in un rabbit hole per cui sì tendenzialmente si può fare ma per via della tecnologia che ha utilizzato Nvidia in quel caso si può fare se hai dell'hardware Nvidia e bla bla bla insomma alla fine non so se ci riuscirò ma nella mia scoperta ho scoperto che su Google collab oltre a delle GPU tradizionali puoi anche scegliere la TPU e la previsione del modello che mi diceva che se io avessi fatto la stessa operazione con la TPU avrei avuto un guadagno di 4x rispetto alle tempistiche fatte con la GPU che era niente male non sono arrivato fino in fondo perché non funzionava lo step precedente ad arrivare a fare quello quindi non l'ho visto ma prima o poi mi risolverò la curiosità di provare a giocarci e vedere un pochettino come si comporta sì allora questo questo è sicuramente uno degli aspetti capire anche come si comportano poi queste cose qua quando le mettono alla prova sul serio però tutte le ottimizzazioni anche di cui parlava Alessio comprese quelle su Gemma 4 che poi sono proprio in casa loro tentano di andare un po' in questa direzione qui tra l'altro sempre stando in casa Google sempre per gli annunci in preparazione all'Io pare che stiano testando un modello che si chiama Omni che dovrebbe unificare Veo con Nano Banana quindi generazione di video generazione di immagini in un modello solo che è una cosa un po' in qualche modo attesa anche perché i rumor dicevano un'altra notizia che leggevo che in realtà è vero che OpenAI ha dismesso Sora vi ricordate il modello di generazione video eccetera eccetera ma dicono anche che in realtà Sora è stato dismesso per quella cosa generazione video ma in realtà chat GPT image 2 altro non è che Sora ha declassato fare solo immagini con ottimizzazioni su questa cosa eccetera eccetera che tutto sommato ci starebbe perché oggettivamente i risultati di Sora erano notevoli da un punto di vista visivo e lo sono notevoli quelle di di certi GPT image sì allora io ho letto l'annuncio onestamente non ho trovato ancora grandissime informazioni su Omni è interessante il fatto che si parlerebbe di un modello multimodale omnimodale non so come vogliamo definirlo che genera anche video perché tipicamente questi modelli multimodali accettano varie cose ma poi generano sostanzialmente immagini o testo molto di frontiera diciamo interessante vediamo vediamo se ci stupiscono assolutamente è tutto da vedere chiaramente perché poi tra il dire e il fare c'è sempre di mezzo mare torniamo dai torniamo un attimo al mondo della politica adesso abbiamo già perso tutti prima abbiamo perso tutti gli ascoltatori che volevano la tecnologia parlando di politica all'inizio poi abbiamo perso tutti gli altri li abbiamo uccisi con con tutta la parte sull'inferenza torniamo alla politica e vorrei andare su uno dei personaggi preferiti di Paolo se ricordo bene che è Elon che si scontra con l'altro personaggio preferito di Paolo che è Sam è vero che che non sai quale figurina scegliere sì sì sì confesso allora sì allora per chi non lo sa cosa sta succedendo in questi giorni non ho idea se se ne parli in Italia peraltro io non guardo molto la televisione giornali quindi non ho idea se qualcuno presta attenzione a sta notizia ma dopo tanto battibeccare si è finalmente andati in tribunale per una questione che se volete con la quale io sono anche d'accordo con Elon eccezionalmente ovvero l'attacco di la critica l'attacco di Elon Musk a Sam Altman per cui ha deciso di trasformare una no profit open AI in una società profit per farci del business che non è che mi suona proprio bene cioè se nasci con una no profit e poi scopri che puoi guadagnarci non facendo un'altra società ma trasformando la tua insomma non è che sei proprio nella lista dei miei preferiti ora magari io sono affascinato da tutta questa vicenda perché ho appena finito di vedere una serie legale su Netflix Lincoln Lawyer e quindi sono preso bene e non vedo l'ora di vedere queste cose e già sul mio feed di YouTube ci sono un sacco di video che fanno come è andata la giornata in tribunale queste cose qua ma non ho ancora avuto tempo di dedicarmici confido che ne sentirò eco tramite i programmi di satire americani tramite John Oliver Colpere e gli altri quindi spero prima o poi di avere cose più succose quel che ho letto adesso è che Ailon non se la sta giocando tanto bene innanzitutto la critica è che probabilmente lui non frega niente dell'etica e della morale che ci sia sotto nonostante lui abbia messo una trentina di milioni di dollari in una società che doveva essere un profit e poi è cambiata però pare che ci siano un po' di retroscena a mettere in discussione i buoni propositi di Elon perché se da un lato dice ah siamo preoccupati del fatto che l'IGA ci ucciderà a tutti e queste cose abbiamo bisogno di un po' di etica dall'altro è saltato fuori che Elon aveva cercato di comprarsi OpenAI e diffonderla con Tesla per avere tutto quanto in uno però gli hanno detto di no e a quanto pare si è un pochettino risentito allora ha detto allora il pallone me lo porto via e quindi c'è questa parte qua poi appunto pare che Elon sia anche in tribunale un attimino erratic come si dice in inglese quindi non so qual è la traduzione più appropriata in italiano tipo sconclusionato non lo so e parla salta un po' di palo in frasca tant'è che gli avvocati seri come quelli delle serie che guardo io lo stanno un pochettino arrostendo perché lui dice quello che gli viene in mente non quello per cui preparato quindi se ci si basa sulle conversazioni iniziali è difficile che possa finire bene per Elon da questo caso una delle rivelazioni che ho trovato più interessanti non necessariamente vera ma più interessanti è che una delle motivazioni che Elon cita per avere investito in OpenAI a suo tempo quando era una no profit che aveva l'obiettivo di portare la conoscenza dell'AI all'umanità quindi assolutamente nobile era che aveva fatto una chiacchiera con i due fondatori di Google e gli avevano detto gli aveva chiesto senti ma non sei preoccupato che l'AI ci faccia fuori tutti quanti e una delle risposte attribuite a uno di loro è a quanto pare questa oh sì ma non è un problema fin tanto che poi l'AI sopravvive allo sterminio dell'umanità e questo commento se è vero mi fa ovviamente venire un po' gli occhi e palla e posso capire la preoccupazione di Elon adesso c'è da dire se sia mai stato vero però se sia mai stato vero oppure no comunque che qualcuno non necessariamente loro possa immaginarla così è abbastanza inquietante e non credo che il processo attuale porterà ad alcun che di utile in questa situazione ma mi fa essere lievemente più preoccupato che ci possa essere lo scienziato pazzo che per amore della scienza voglia più bene a il fantasma dell'opera che voglia più bene ai computer che non agli umani e che quindi porterà questa cosa comunque immagino che altre nuove rivelazioni supercazzoli usciranno dal processo di Elon e Samatman e quindi vi terrò aggiornati man mano che ne escono così invece di continuare a sentire solo dell'omicidio di Garlasco sentiamo anche del processo OpenAI no ma che poi allora quella affermazione lì di uno dei due fondatori di Google potrebbe tranquillamente averlo fatto a Sergey Brin perché è un personaggio leggerissimamente sopra le righe nelle sue uscite per cui non cioè ci credo abbastanza che qualcuno possa averla detta che poi non e penso anche chi ecco tra i due però di certo chi ha in questo momento in mano le AI di Google che è Demis Hassabis ha visioni diverse e non fa altro che ribadirlo nelle sue interviste decisamente human centric molto più di tutti gli altri anche più del buon Darione ma poi ho visto un'altra cosa Paolo che profezia autoavverante compriamo compriamo i server alla Lidl quindi compriamo i server alla Lidl allora facciamo dell'autocitazionismo di una puntata in cui non ho idea di quale fosse non mi interessa andarla a cercare ma scherzando in passato parlavamo di modelli cinesi e del fatto che appunto arrivano basso corso e tutto quanto democratizzazione e stavamo scherzando che in un prossimo futuro uno il modello lo comprerà su Aliexpress che è il nostro equivalente di comprarlo alla Lidl e indovinate un po' qualcosa di vagamente simile sta succedendo in realtà in Europa spinti dall'esigenza percepita dio grazie dai governanti a livello europeo che dicono forse non è che ci possiamo proprio fidare delle buone intenzioni degli Stati Uniti nostro alleato storico che facciano sempre la cosa giusta e delle loro società forse è il caso che iniziamo a pensare ad avere un'infrastruttura proprietaria che ci isoli dai rischi di dipendere eccessivamente da qualcuno di terze parti un parallelo se volete di quello che è successo con il carburante con la Russia all'inizio della guerra della Russia dell'Ucraina realistico se chiedete a me come scenario e allora hanno detto organizziamoci proviamo a inventarci qualcosa e la prima infrastruttura che andava a mancare è quella dei cloud services perché ora come ora la fa da padrone AWS Microsoft e Google principalmente probabilmente anche Alibaba ma nessuno se ne ricorda mai di loro e strettamente europeo non c'è granché e anche i vari region locali europee che sono state garantite è stato notato recentemente come sono scusate ho attivato Google sul telefono che mi ha sentito che parlavo male di lui e mi ha risposto dicevo non abbiamo niente di europeo e anche nonostante ci siano delle region ufficialmente con dei data center europei sono soggetti alla legge americana e la legge americana qualcuno è andato finalmente a guardare dice che il cloud act se il governo ha bisogno per questioni di sicurezza nazionale che a quanto pare succedono per ogni cazzata adesso negli Stati Uniti ha bisogno di andare dai provider a chiedergli i tuoi dati non importa tu dove sei sotto quale legislazione loro gli hai danno e questa è una cosa vera tanto vera che oltre a essere stata notata dalla comunità europea è stata notata da qualcuno vicino agli europei ovvero la Svizzera che ha che è finita in prima pagina perché hanno reciso un contratto di 200 milioni o di più non lo so non mi ricordo questi fantanumeri con Microsoft credo proprio perché non gli poteva garantire questa cosa fondamentalmente quando hanno chiesto e verificato gli hanno detto no i vostri dati sono privatissimi tranne che quella volta in cui ci richiedono loro e siccome gli svizzeri ci tengono a queste faccende tradizionalmente avendo un piede in due scarpe la privacy è diciamo il loro il loro brand hanno preso direttamente questa situazione e basta hanno annunciato questo questo taglio da quelle parti come si risolverà quello ancora non lo so nel senso che non so se guarderanno a una soluzione locale loro soluzione europea o soluzione di qualcun altro o forzeranno la mano di là per riuscire a cambiare le regole per loro non lo so però sta di fatto che ci sia arrivati in prima pagina con questa cosa ed è una questione quindi laddove la Svizzera si muove si muove anche l'Europa l'Europa si è mossa in una maniera un pochettino diversa ovvero facendo un bando per andare a chiedere iniziative commerciali europee di poter fornire un servizio di cloud che sia dignitoso è stata fatta in maniera molto democratica questa cosa tant'è che anziché emergere un supergruppo una coordinazione di qualcuno che ti possa dare vagamente fiducia che sia capace di farlo non lo so avrei sognato che il Cern lo facesse che ok siamo in Svizzera ma qualcuno con un nome di cui ti puoi fidare invece no vari attori più o meno random sono avanzati in Germania in Olanda in Olanda in altri posti niente dall'Italia perché non non avremo il cloud di poste almeno una l'abbiamo fatta giusta e uno di questi è una società legata alla Lidl non è propriamente la Lidl cioè il titolo noi diciamo la Lidl il cloud della Lidl in realtà è una sorta di società di servizi che è nata come sottosocietà per le esigenze del gruppo Lidl e poi si è specializzata se volete la stessa relazione che ci passava tra Amazon e AWS però il nome veniva da lì e quindi da qualche parte tra un po' ci sarà il cloud della Lidl che sarà la nostra variante rispetto al cloud di AWS ci scherziamo sopra in realtà la Lidl il suo mestiere lo fa anche bene è più da vedere a mio avviso la vera scommessa non è tanto chi te la gestisce quanto più vedere se le ambizioni di arrivare ad avere una scala tale da potersi autosostenere che sia se non nella stesso diciamo stesso livello di quello di AWS ma che vada in una direzione che ti faccia avere fiducia che l'Europa possa basarsi su questo è una cosa buona e vediamo se siamo capaci è una cosa buona e fondamentale perché c'è anche l'ho nominato tante volte questa cosa credo anche qua in podcast e di sicuro ne ho parlato in qualche intervista con alcuni degli ospiti c'è anche uno scenario possibile quasi probabile ad oggi mi sbilanciano a dire quasi probabile che è che ad un certo punto i cosiddetti agenti così da qui vi parlo del mio di agente no a parte gli scherzi i cosiddetti agenti quando staranno dentro al cloud eccetera arriveranno ad un punto in cui produrranno valore ok produrranno valore è una cosa che dice anche Draghi non me lo so non la sto a mente andando io e nel momento in cui è un agente che produce valore c'è uno scenario possibile in cui a essere tassato è il valore prodotto tradizionalmente nell'economia europea la tassa viene messa sul valore prodotto non tanto su chi lo produce anche ma perché è più comodo tassare la persona che il valore ma in realtà quello che tassi è il valore prodotto e se il valore è prodotto su un suolo diverso da quello europeo potrebbe sottostare ad una tassazione che è diversa da quella europea per cui il problema comincia a diventare complesso anche di sostenibilità del welfare di cui l'Europa ha giustamente il fiore all'occhiello per cui guardare avanti e poter produrre valore digitale su suolo europeo è una cosa che non è soltanto importante dal punto di vista della privacy che interessa molto l'Europa e ne ha sempre fatto un cavallo di battaglia va benissimo ma c'è anche un discorso molto più pratico di dove vanno finire poi i soldi sì peraltro questo è un ginepraio legale mica da ridermi vengono in mente tutte le questioni legate alle criptovalute alla tassazione capire dove risiedono i fondi si sono in Europa negli Stati Uniti dove in realtà non lo sono sono nel legger il ginepraio però è effettivamente un ginepraio e per quanto riguarda la blockchain è più difficile perché è una cosa distribuita ed è costruita per evitare quella cosa lì in realtà una delle idee è quella lì chi sta facendo e chi farà servizi basati sugli agenti li metterà tendenzialmente più in un cloud dove sì esiste quel problema lì perché ci sono molte punti eccetera però un pochettino più gestibile ed è più interesse di tutti arrivare a definire chi produce chi osa come valore quindi forse è un problema leggermente meno complesso ma che può diventare sanguinoso se non hai niente sul suolo italiano europeo che produce valore o comunque poter offrire la possibilità dal punto di vista di chi mette in piedi questa cosa qui di dire va bene ma tutto questo gira qua quindi voglio che sia trattato con le leggi di qua va bene esatto esatto proprio questo questo per segnalare la cosa ma parlando di agenti e parlando anche di località di fare le cose in locale io vi ho raccontato che ho messo in piedi Hermes agent a casa mia su un vecchio computer usando un modello cinese così per non farmi mancare niente e sta facendo delle cose sta facendo delle cose di cui parlerò poi anche nella newsletter che esce lunedì ma che vi volevo raccontare magari far vedere anche qualcosa per chi è arrivato fino in fondo qua almeno scusa se ti interrompo giusto per chiarire tu stai facendo girare l'agente in locale con il modello però offerto in cloud da un provider corretto perché non ho una macchina su cui far girare un modello sufficientemente performante per fargli fare queste cose quindi il modello è cloud ma invece l'agente l'arnese per usare il termine tecnico che è Hermes agent gira in locale su una mia macchina e i file che modifica le cose che fa le fa lì le fa lì ma in più gli ho dato accesso a tutta una serie di servizi miei lo raccontavo l'altra volta anche alla mia mail e quant'altro cercando di prendere tutte le accortezze possibili ho scelto Hermes e non open cloud per questo motivo perché è un po' più facile personalizzare gli accessi c'è un sistema di skill molto chiaro su cui ad esempio lo dicevo l'altra volta su Gmail gli ho detto va bene vai su Gmail fammi vedere come ci va e lui mi ha fatto vedere il suo script e io gli ho detto ecco quei due metodi lì quelli che fanno delete e quello che fa send della mail li cancelliamo proprio dallo script che usi non è che ti dico di non usarli cancellali proprio non devi più averli e poi ho verificato che l'avesse fatto e così e così è e che cosa gli faccio fare gli faccio fare un sacco di cose dalle cose più banali se vuoi di di utilizzo ciclico per cui al mattino o ogni ora mi controlla la mail mi dice quali mail sono soltanto informative e me ne fa un risunto di una riga i quali invece hanno bisogno della mia attenzione andando a leggere la mail e quindi capendo che ho bisogno della mia attenzione e in più per quelle che prevedono una risposta gli faccio generare un draft della risposta non la manda io me le trovo poi in gmail già draftate poi le edito ma ho un punto di partenza almeno e mi guarda il calendario mi dice quali sono gli appuntamenti del giorno e quelli importanti della settimana per me quelli importanti sono quelli di un certo colore cioè abbiamo discusso un attimo come definire queste cose mi trovo i buchi tipo gli dico eh devo parlare con Alessio ci mettiamo un'ora mi trovi un buco nel calendario mi fa tre o quattro proposte gli dico va bene quella lì e lui mi fissa l'appuntamento cose da gestione di questo tipo mi gestisce le cose in casa per cui gli ho dato tutte le varie prese smart tutto quello che di smart ho in casa che non è tanto quanto Paolo ma qualcosa ho anch'io eh e quindi tipo da fuori casa gli dico di accendere il condizionatore o la pompa di calore eh di controllare le luci le telecamere queste cose qua eh e poi e poi fa fa cose allora l'accesso è con telegram per me eh fa cose anche abbastanza proattive cioè nel senso l'altro giorno gli dico eh guarda il mio sito personale eh un sito personale tipo biglietto da visita elettronico no eh cosa ne pensi la risposta non è stata gentilissima ma fa niente eh più o meno a livello della risposta su come erano le nostre thumbnail del podcast sì sì sì più o meno eh no è stato più gentile perché anche lì gli dici come deve comportarsi contiglio di essere estremamente franco e estremamente diretto e lui mi ha detto sì capisco il tentativo però non ci siamo ehm e gli dico vabbè cosa facciamo eh questo mentre mio figlio era all'allenamento di basket cosa facciamo cosa non facciamo e lui dice guarda ti faccio tre proposte dimmi se ti piace più una versione dark da sviluppatore o un'altra fine io scelgo la versione dark e lui dice facciamo così ti faccio una directory new website te la te la push te la deployo e poi tu gli dai un'occhiata e mi dici come va scusami praticamente ti ha detto ma che te lo chiedo a fare dopo che gli hai espresso la tua scelta cioè gli faceva talmente schifo la tua preferenza che ha detto facciamo che faccia tutti e due poi decidi dopo no no no mi ha fatto quello che gli avevo chiesto ma la cosa incredibile di cui posso condividere una diapositiva perché c'è il new website aspettate che l'apro intanto che parlo la cosa che a me mi ha fatto mi ha lasciato lì è che in altre conversazioni per il podcast per la newsletter eccetera eccetera io gli avevo dato nella newsletter ad esempio una parte agenda in cui dico le conferenze che sono stato quelle che hanno pubblicato il video quelle dove andrò e quant'altro e lui aveva questa informazione perché ne avevamo già parlato lui ha una gestione della memoria fatta molto bene meglio di altri gerarchica e quant'altro e lui mi ha generato questa cosa qua che al di là possa piacere non piacere però allora questi tre i miei contenuti diciamo che sono le due newsletter e il nostro podcast c'erano nel vecchio sito quindi non è merito suo i progetti c'erano nel vecchio sito e quindi non è merito suo di queste io avevo soltanto l'ultima il devox perché l'avevo fatta a mano invece lui ha visto che c'erano tutte le questi talk che avevo fatto che farò e me li ha messi nella pagina prendendoci sia le slide che i video che tutto il resto sono sono giusti e questa cosa qua del prendere iniziativa mi ha abbastanza colpito non è l'iniziativa quella anche chi se ne frega che attribuivano open cloud di fare le cose da solo però iniziativa diciamo intelligente un'altra cosa così parliamo anche di quell'argomento lì c'è una cosa che avevamo messo in scaletta che è un articolo aspettate che recupero il titolo perché rischio di dire una sciocchezza quello dei 100 milioni cos'è 12 milioni di contesto adesso non ce l'ho qua sotto mano però c'è un articolo di cui abbiamo parlato Alessio giusto abbiamo parlato forse in chat non ce l'abbiamo sì io non ho approfondito però c'è questo articolo che parla di una start up che ha fatto un modello a 12 milioni con un contesto lungo 12 milioni di token siccome ci sono stati annunci in passato di gente che faceva 100 milioni di token che poi non sono andati in nessuna parte volevo capirlo prima di parlarne tempo poco e ho detto ad Hermes ho bisogno di capire questo articolo sono in macchina quindi fammi una sintesi fatta bene lui è partito e mi ha chiesto mi chiedi i permessi di fare qualunque cosa ma perché quella è la mia scelta ero in macchina non guidante è preciso e mi dice beh ma se vuoi se mi dai il permesso di installare un text to speech te lo leggo mi sembra una bella idea questa fallo si è installato un modello text to speech mi ha mi ha detto intanto ti faccio una sintesi in italiano che magari io non gli ho detto che non stavo guidando se sei in macchina e stai guidando è meglio se lo faccio nella tua lingua ti faccio una sintesi in italiano e te la leggo ad alta voce e l'ha fatto vi condivido un attimo se riesco aspettate che vedo di aver aperto veramente ti ha mandato un vocale quindi mi ha mandato un vocale sì sì mi ha mandato un vocale allora io a lui i vocali già li mandavo prima perché lo speech to text ce l'ha di default però adesso qua c'è un mare di roba perché mi ha mandato il mio calendario aspetta dove hai i soni vocali aspetta che non li trovo volevo farveli sentire ecco allora la lettura magari è un po' meno sexy di altre adesso non riesco a farvelo sentire perché avevamo già provato che le cose da telegram non va l'audio cioè lo sentite voi ma poi non lo sento gli ascoltatori non mi piace vi dico io che la lettura è accettabile non è la voce di OpenAI ma quella di Eleven Labs che si può mettere lui mi ha subito chiesto ce l'hai un account Eleven Labs se ce l'hai lo uso no non ce l'ho pago per tante cose per quella no e quindi questo l'ho fatto con un modello locale ci ha messo pochi minuti e mi ha mandato questi due due riassunti con due punti di vista diversi e allora io da quello che mi ha detto ho capito un po' di cose gli ho detto sì ma senti fammi un riassunto in forma in forma tabellare intanto che non sto prendendo la finestra giusta per scorrere ok fammi un riassunto in forma tabellare di questi dati e lui me l'ha scritto così e gli ho detto no voglio una tabella perché così a leggere capisco poco e lui mi ha fatto una tabella perché gli ho detto io che voglio sempre tutto in markdown e io gli ho ribadito no non capisco un cazzo sono in macchina non riesco a vederla e lui ha fatto questa scelta pazzesca ha fatto un html per farmi la tabella fatta bene e poi si è detto da solo eh no ma l'utente mi ha detto che in macchina come lo legge un html aspetta lo renderizzo faccio una foto e gli mando una foto e mi ha mandato la foto della tabella poi se volete entriamo nel dettaglio del paper ma forse non oggi eh beh ti e poi vabbè dopo me ne ha fatta un'altra su un altro argomento ricordandosi che gli avevo detto che non riuscivo a leggerla e allora me l'ha data di nuovo così e in questo caso è stato a me quella roba qua è stata abbastanza wow allora apprezzo ma io l'avevo già visto insomma nelle mie lunghe conversazioni via telegram con semplicemente ZAI GLM anche a me era arrivato a quelle conclusioni in questo caso forse è stato proattivo che te le ha dette tutte lui nei miei casi invece gli davo io dei prompting nel cloud .md dicendogli guarda che sono in giro quindi quando sono in giro mi servono robe brevi usa gli emoji queste robe quindi mi faceva delle cose simili però apprezzo e capisco il wow effect ha fatto no io sono particolarmente soddisfatto anche e soprattutto dal fatto che abbia un accesso lascia stare i servizi in cloud anche perché poter accedere alla mail al calendario è una roba notevole cioè poter mandare un vocale mentre sono in giro e dire mi sono ricordato che devi aggiungere questo task e poi te lo trovi in gtask quella è tanta è una figata idem per gli appuntamenti sul calendario e cose così però anche che abbia accesso al file system ad esempio vi ho parlato dell'llm wiki quello di carpati che io uso per pensare e per analizzare i paper in questo momento sto guardando i paper sulla memoria di lungo termine e per in questo caso più per risparmiare i token e tempo che altro c'è la fase dopo che io ho raccolto tutti i paper c'è la fase di digest che vuol dire che si scarica tutti i paper in pdf le legge genera il wiki crea i collegamenti eccetera che comunque ci mette una mezz'ora 40 minuti questo su cloud code anche non solo lui e la facevo al mattino però mi scocciava la lanciavo e facevo altra parte il numero di token non banalissimo che usa per fare questa cosa che si pagano e quindi la facevo comunque con glm e quindi ci metteva un po' adesso ho un task assegnato a lui che si fa l'arabase del github dove io gli ho detto le cose la sera si fa tutta la parte dei digest di reflect e quando ho finito mi creo una pull request e io emerge la pull request e poi lavoro sulle cose già digerite e questo è un bel risparmio di tempo ultimo ma non ultimo e l'avete forse visto su lince gli ho fatto fare la pull request review di cose che avevo fatto con cloud code e lì è stato affascinante perché se la suonavano se la cantavano perché cloud gli ho detto guarda che c'hai una pull request review l'ho andato a vederla e ha detto ah sì bravo questo è interessante questa la faccio questa la faccio ha committato e poi ha messo un commento no questa non la faccio perché è più una perdita di tempo che altro non so che cosa tu abbia visto nel codice ma ti sbagli e quindi è stato carino che se la siano suonata e cantata che poi è un po' quello che fai quando fai slash simplify in locale di fatto si fa una pull request review da solo così entusiasmo a palla per Hermes non posso più vivere senza mi hai fatto venire in mente un caso d'uso per cui mi serve questa cosa quindi pubblicamente avviserò che lo farò anche io per finta la comunione sono il sono il diavolo tentatore si true e beh allora se non volete parlare di quel paper lì complicatissimo di cui io mi sono preparato ma il tempo potrebbe essere mediamente lungo credo che più o meno siamo arrivati in fondo alla scaletta diciamo soltanto questa cosa che sempre nel mondo delle tentazioni l'idea di passare a codex perché c'è gli animaletti che parlano mentre fai il codice è una cosa a cui sto cercando di resistere ma non so fino a quando riuscirò senti fammi vedere uno screenshot perché come ti dicevo io quella cosa l'avevo attivata un tempo fa non ce l'ho pronto aspetta ce l'ho apriamo la prossima puntata con lo screenshot con gli animaletti apriamo la prossima puntata con lo screenshot non ce l'ho lo screenshot pronto va bene ce l'ho su codex ma c'è un abbonamento piccolo e non lo uso tanto però l'ho provato e devo dire che insomma avere l'animaletto perché è un animaletto fatto bene non è come quello piccolino bruttino che aveva fatto quello del codice questo è proprio bello entra ti cammina in mezzo allo schermo rompe anche un po' le palle se vuoi però quindi il metaverso di di come si chiama di Zuckerberg aveva il suo perché e l'unico perché era per i programmatori per fargli vedere i cartoni animati mentre programmano sì no perché lo dico perché ho provato un po' codex perché lo sto provando con l'inchessac che c'è un mare di novità andate a vedervele sul sito non sto a dirle qua io e Claude abbiamo fatto un sacco di cose in questi giorni compreso la modalità paranoica che ha dagli estimatori quindi andate andate a vedervela se volete va bene ci salutiamo salutiamo il pubblico alla prossima numeroso mettete campanelline stelline iscrivetevi al canale e non perdetevi guardate i short non perdetevi la copertina di questa settimana che vedrà Alessio in copertina sì tocca ad Alessio tocca ad Alessio poi dicevo anche che non solo tocca ad Alessio per la copertina poi cercheremo di farle ancora le copertine così magari un filo meno Bruce Willis perché l'ultima copertina con me tutti mi hanno detto che sembrava Bruce Willis quando era in forma prendi quello che ti viene detto io ci sono passato fai il buon viso che attivo gioco e prendi tutto quello che ti viene detto bene bene adesso tocca ad Alessio vado subito a vedere come fare la copertina ciao a tutti ciao ciao