DeepSeek V4 taglia il 78% di compute sulla KV cache, Roo Code chiude e Hermes Agent gira da solo. Settimana densa: DeepSeek V4 Pro e Flash con 1M di contesto, Kimi K2.6 e MiMo V2.5, l'analisi di Salvatore Sanfilippo che fa girare V4 Flash su Mac mini con quantizzazione mista Q2/Q8, il post-mortem di Z.AI sul bug async cache di GLM-5 (con fix contributo a SGLang), Vision Banana di Google DeepMind che usa la generazione immagini come interfaccia universale per la vision, e il deep-dive di Stefano su Hermes Agent, l'alternativa open di Nous Research a Claude Code che gira con GLM-5.1 e gestisce mail, calendario e paper di Arxiv in autonomia. In più: le interviste a Demis Hassabis (DeepMind) e Andrej Karpathy sull'efficienza e sull'AGI, il keynote rotto di Jensen Huang con Dwarkesh Patel, e i 40 miliardi che Google investe in Anthropic. Follow Risorse Artificiali per non perdere il prossimo episodio. Versione video su YouTube: https://www.youtube.com/watch?v=qKl4Vkb6BMw&utm_source=spotify&utm_medium=description&utm_campaign=ep50_drop #50
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Reviewing the wave of new Chinese open models and the shift from research to inference engineering.
Benefits
Token efficiency lowers cost for same output quality
🌐 This transcript was automatically translated to English from the original.
Hello everyone, welcome back to Artificial Resources. This week we talk Chinese about models, we talk about DeepSeek which came out with the 4, Flash Pro, but there are also Chimik 26, Mimu 2.5. I don't know if you've seen it, so I'll give a little preview and then we'll go into everything, you've seen that Google promises 40 billion to Antropic, that is, a big investment, everyone is investing, strangely enough, they're not competitors in some way, but this big investment is being made, but the Chinese are also making investments because Alibaba says it will invest in DeepSeek, so even there we're starting to see this flow of money, but then there are many other interesting things, there's Vision Banana to name, then the two or three things that I was most struck by the scaling of the lessons on scaling of GLM5 that Paolo mentioned to us and the interviews I heard yesterday by Assabis and Carpati. Although I must say that the two most beautiful things are, that is, the most beautiful, one that left me like this, Rucode that closes, the Coding extension that closes, because it says that, and then maybe we'll talk about this from our point of view, that it's no longer time for IDEs, for IDEs and then oh well, then I installed Hermes Agent which is doing many things for me, even at this moment. That interests me. Know that when you said that this episode the models speak Chinese, I thought it was like when my Chinese models go crazy and give me answers in Chinese to questions in English. Of course, but in fact that's what we'll do, we'll give random answers to the questions we asked ourselves, because even if I have to say that I was saying because we don't have any comments, that's not true, because this week we also had quite a few comments, including ether. Chinese Mandarin? Chinese Mandarin, absolutely. No, come on, let's start from that, come on, let's start from the Chinese, DeepSeek, Mimo, Kimi K2, we have now decided, Alessio, that you are the one with the benchmarks, so you say everything about the benchmarks of this work. I can basically tell you that last week it almost seemed like a competition, let's throw out the Openweight model, the new State of the Art, in the sense that DeepSeek V4 Pro and Flash came out first, where Flash is a slightly smaller and clearly faster version, Faville, a whole series of new features, a new type of attention which please, talk about it Stefano. Well, ok, we have the new best Open model, except for the fact that after a moment Kimi K2.6 came out, which does even better, at least in the benchmarks, and Mimo 2.5, which more or less aligns with the other two, but it shows us something new, that is, new, it highlights an interesting aspect, which is that of the efficiency of the tokens, which is a topic we have talked about little, in my opinion, in recent times, that is, it essentially generates fewer tokens for the answers, to obtain the same level, let's say, of quality in output, in passing tests, etc., and that has an impact because less tokens means less cost, basically. Let me digress, I already have a second to divert you from this conversation about saving tokens. You reminded me that in the last couple of weeks I've come across a lot of videos or articles that talked about a revolutionary extension of cloud code, called Caveman, I don't know if you've come across it, where the idea is to make your model ball like a caveman, so instead of giving you long and detailed answers, it answers in monosyllables, like, I don't know, to a child, when you have to introduce him to his distant uncles. And nothing, people were very excited, saying, ah, this is beautiful, I'm saving tokens here and there. Let's say, that is, it was curious, it was against the trend, it could have made sense, except that with a little more time people looked into it a little and said, yes, look, most of them, if you apply this process to most of the cases you ask them to do, then they give you inconsistent answers, because imagine that, I don't know, you're asking them to make a diff between two code sources, you may want a more concise answer, but if they answer yes or no and it doesn't mean nothing, you're not going anywhere. So there was this trend, the name probably evoked who knows what things, the simple concept interested people, but then on the practical side it wasn't that useful in large use cases. Yes, but I'll tell you, in my opinion there is a, this is a metric that has been a bit underestimated lately, because if you go and look at the models that do reasoning, thinking, in reality a lot of time is spent in that phase and sometimes we are inclined to underestimate this aspect, but then for example when you try to run the models locally you notice this aspect even more. there are probably some applications for which, all things considered, you are willing to accept perhaps some errors, a slightly less quality response in exchange for greater speed. So much so that, for example, the cloud code itself but also other tools allow you to set the level of thinking reasoning you want to have. Clearly the quality changes but the response times and consequently the costs also change. Yes, look, it's something he said to Sabis also in the interview that I anticipated then maybe we'll go into a little more detail but now I was looking for my point he actually says that there is an architectural efficiency problem on context windows currently what has been done for deep reasoning etc. and he quotes himself he quotes Gemini 3.1 pro deep thinking the one that won I can't remember which mathematical competition he says that the efficiency problem is big because essentially brute forcing has been done up until now that is it generates many tokens and take the good ones he says when we talk about efficiency but also effectiveness it is not necessarily true that brute forcing always pays, in fact experience teaches us that when he starts optimizing he gets unexpected results and in fact on that thing the interviewer who is the CEO of Y Combinator asks him but so how far are we from EGI from a scientific point of view I speak to him rightly as a scientist like Sabis is and Demis replies we are not far away it could also be that we have already discovered everything we need we just have to optimize it it could be that we are missing those two or three ideas but no more revolutionary than Attention or so was to be able to truly emulate an intelligence different from human one but still something that we can call truly intelligent in fact he says that in Deep DeepMind they have two completely distinct lines of research which every now and then they compare with each other but which they also deliberately keep distinct as groups, one is on incremental research and one is on breakthrough research so they really explore the two possibilities they have an extremely scientific approach in this is one of the interesting passages but it is also to confirm something you say about efficiency also comes back to us, many of the Carpathians are coming back to it at the moment about the fact that they slobber the models a bit too much is one of the controversies that also came out on opus 47 that someone complained that it is extremely verbose among other things the same cloud code now has adaptive thinking that is they tend to try to propose an automatic system in which based on the type of question you ask but also GPT has had it for a while it decides how much thinking to do because maybe the question seems relatively superficial so it's not even worth having to do excessive reasoning just to give an automatic system that improves the experience in general, however, going back to DeepSeek 4 we were saying that there are some interesting improvements from the point of view of the algorithm of how it works Stefano yes yes no there are many interesting things first of all let's start by saying that it is 1.6 trillion yes exactly gigantic models absolutely gigantic model but a mixture of experts now I'm looking for the data and I've made some notes so as not to get lost the pieces 49 billion of active token parameters which yes is a mixture of expert but we are used to those with 3 4 5 in short it is larger the active parameters are larger than the average of many models that are around at least among the open wait let's say when you think about the models that run locally but it has some interesting characteristics this thing of being so strongly mixture of expert because someone and in this case Salvatore Sanfilippo is the one who wrote about it the most to stay in Italy he started to say attention but because this model is interesting here because not everyone the weights are the same the ones that matter most are the routing ones I think and then I can do different quantizations so much so that on some groups of parameters he managed to quantize Q2 on others instead in the attentive Q8 and this allowed him to run Epsic V4 flash which is a little smaller 284 billion anyway not exactly a joke with a million of context he ran it on his mac mini mac mini pro with I don't remember how much ram but I don't remember how much I don't know 628 196 but still it's a result he says with good performances he's been talking about it for a week on 'in the era of the dawn of mp3s when we were 20 years old or less in which there were various forms of compression to have the best quality we started from 128 then it was 192 and then we spread towards the variable byte rating so different parts of the file were sampled with different variability and this San Filippo approach reminds me a lot of this which obviously makes sense it increases complexity because you make the system work in a different way depending on where you are or when you are it makes sense from point of view of software engineering yes yes yes I agree and then it is interesting that from this point of view there is also an interest from the open source world and we noticed more than anything else it is something we noticed that perhaps Alessio put on the table a couple of episodes ago as in part we are moving from pure research to software engineering from engineering in general and it is no coincidence that actors like Antirezza Salvatore San Filippo enter the scene he obviously has a notable background especially on CIC++ in general on programming I would like say and enters the field by also putting his nose into this kind of thing which is yes on the models but it is optimization of the inference to go and do different quantizations so the thing is interesting from two points of view both for the result that he has obtained which he talks about and which is certainly interesting but it is also interesting to see this shift of skills which makes the part let's say pure research on the models but it is the engineering of making them work with the hardware we have which returns to making the open source movement become more protagonist than it could have been with the open weight because let's be clear about an open weight model I'm very happy that they gave it to me and that I can install it etc. etc. but it's always something that comes down a bit from above, that is, how do you contribute to that stuff? You can't if you're not inside one of the research centers that release it. Instead, if we're talking about optimization of inference rather than agents and optimization of tools, then we'll also talk about that. The open source world becomes the protagonist again and we use we but certainly I believe in it as strongly as open. source had an impact on operating systems in the 1990s, the end of the 1990s and then on everything else in the 2000s, I believe that it can also have an impact from this point of view, I don't know if you have a yes but then there is also a different group of people who can contribute at a research level yes that's fine ok we'll make it open but how many are able to and how many have the resources as well as being able you also need the resources if you're not inside a research center good not good competent not competent whether you are you can't do that stuff there if we come to software engineering it's absolutely different if you want I can digress in this regard in the direction of the post-mortem of easy AI in which they published why exactly that is the point ok so what is the context easy AI is one of the providers that provides GLM the model that we use for how I use it of which I had an annual subscription because it cost very little and gave good quality results at a certain point it started to give poor quality results it was hallucinating let's say using a somewhat vintage term at this point since it happens less and less often but not only he was hallucinating he was inventing words he was mixing things up it was as if he had gotten drunk or for those who have spent enough time in front of a screen it was as if the memory had been corrupted when characters come out that should be there and there is no explanation for why they are there internet Reddit specifically obviously she was infinitely outraged by the fact of not receiving a level of professional quality for their 20 dollars a month that they paid and that they said I would never have another quarter they will not have never again my money obviously that is these things she doesn't know who I am these scenes here in short and nothing since there was no transparent communication those people seemed competent the thesis was that these people from ZAI were a bit greedy or in any case they were a victim of their own success so they could have had more demand than they could serve and so they closed the tap in the form of deploying their models in a super quantized form which obviously degrades the performance in terms of quality and this is the result tate and this was a thesis that I took as good because I didn't tend to have other elements and people believed it so much so that someone thought that ZAI's was a business that was going to die because the quality wasn't there it was interesting sorry another reason why people had made this assumption was because the same models provided by other providers specifically Alibaba worked very well so people told you no I'll leave it alone ZAI are stingy they let them go there it works well and nothing and we were like that for a a bit ZAI didn't bother to comment some time ago a comment came from one of the ZAI engineers who said ah we had a bug try now and things should be better things have been better since then so let's say the story checked from their point of view and they promised a post-mortem the post-mortem arrived yesterday or the day before yesterday and in my opinion it was more interesting than I could have imagined for many aspects meanwhile for transparency I would say yes partial transparency I would tell you now I'll tell you why from my point of view as an engineering manager basically a cynical person and not very optimistic towards the stories you told us basically basically they gave very low level technical details explaining what the problem was and therefore from the point of view of technical nerd satisfaction they explained things as they were and the answer is credible but it is the opening let's say that I find it slightly more questionable in which they basically said we are happy to share with you the resolution of this problem which was very rare and we had a bit of difficulty chasing it we never managed to catch it yes a couple of balls in the sense or not were you using it well or were you not looking at it because I and all those others caught it after five minutes every day so this is a bit of my criticism obviously they have to maintain an official room so if you are vague on the terms you say rare conditions what rare means for you or for me is debatable and therefore it's fine so that is a bit of my criticism however once removed this façade of admitting a fault but up to a certain point they have reached the point of talking about the point and for those who have a background in software like you two specifically but maybe it comes from me think of San Filippo in this case because it is something close to him also try to guess what is the main cause of the problem they had asynchronous cache propagation ah asynchronous who has ever had a problem with synchronicity especially when you have to take operations into account what does asynchronous mean for those who are not familiar basically start concurrent operations and not always wait sequentially for each response before doing another action to go faster we decide that I can do two things at the same time like in real life every now and then when you do two things at the same time you make a mess and therefore you have to have a bit of let's say self-control to coordinate yourself here they attributed the main causes to the cache to the kibi cache and to the synchronicity of this thing for which there were conditions in which case sometimes if an operation was aborted the asynchronous operation of the cache was not yet completed and therefore maybe the next operation was in a context that was not its own and that's where the very relatable mess came out let's say from the point of view of the software of whoever had their hands in it that is everyone there we have come across this thing at least once and so their story is particularly credible in my opinion because they have identified a bug specifically in this context in one of the projects on which they depend an open project now whose name I am now looking for because I read it this morning but I no longer remember it a related project is SGLANG which I don't know well I confess I don't know exactly what its role is but they linked the PR that they opened to that project so officially there was a problem they recognized it in their stack they have fixed and they are contributing to open source so hats off to this side here then let that be the whole story who will ever know most of the story makes sense and is interesting from a moral software point of view you have to know enough what you are doing if you want to provide a service to a large audience otherwise there is no sp erance that you can ask Cloud Code to do it for you and that's it I don't want to say stupid things but it seems to me that SG Lang is like the inference engine on which they run the model if it were the Lama SPP ok ok or the VLLM so we were saying therefore you can also have the best model in the world but if you have problems keeping it up to scale to levels of competition of use like those they can have with millions of people from all over the world who want to use it it's not the same as saying it and I can't believe it because if you have ever even just tried to set up something local on your computer and you've already run into performance problems etc. imagine this thing taken 4-5 levels of size higher yes absolutely returning for a second instead to DeepSeek you asked me about the attention, how much the attention has changed and to talk about it a bit then as always I would say what it was like let's say January last year when the first DeepSeek R1 arrived if you remember the first DeepSeek moment the first DeepSeek moment DeepSeek is more interesting for the papers than for the model itself because in the end even R1, apart from having made great news for itself, it's not that it was adopted on a large scale also because then all the others arrived coding wasn't its strong point it came out when coding was on the launch pad etc. yes no but Stefano am I wrong or was it the model that deployed perplexity ehm v3.1 they used if I remember correctly it's the next one but yes it's still a DeepSeek yes yes yes it's still a DeepSeek you are right about this however in the sense that it gave ideas on which all the others then developed let's say well no DeepSeek's papers are very interesting because they are really very open in research and in sharing what they do and I quickly remember that in DeepSeek R1 or rather V3 which is the basis of R1 there were mechanisms that had not yet been seen of um sparse fast attention and flash attention which are two ways of putting attention on the tokens different from how it was done before now no one made them anymore as per the original paper attention is all you need which takes everything and looks at everything because it is too expensive however the sparse flash attention of sparse sparse fast attention now I get confused with the names of DeepSeek V3 was very interesting because it managed to have a multi-level attention which led to having a broad context at the level of concepts despite costing little now going into detail in a podcast becomes complicated and I won't go into too much detail even in the new one but it is remarkable it is remarkable it must be read re-read for those who make the models the inferences are already in particular from what they are already doing because they insert the concept of what they call hybrid attention architecture and they have three three big innovations one is what they call sliding window attention so basically the attention is divided into pages as if they were memory pages for those who are a software engineer instead of using the raw text within the memory pages they use a sliding window production to consider how the last pages overlap with each other this thing allows you to concentrate the attention only on a more limited number of tokens and ultimately have efficiency but the two most interesting are those called csa compressed and sparse attention and hca i ivly compressed attention practically they manage to make a million tokens spending 78% less of the memory computing power not memory because because basically instead of having a flat KVCash therefore key value but now I'm always oversimplifying if someone listens to us who makes an inference and doesn't redact how I explain it but to try to explain it instead of making a dry key value map with the kvcash they have essentially created a system hierarchical tree memory in such a way as to recover more quickly those portions of kvcash that are significant for the attention of that moment this level of compression is done for more or less large pages of kvcash depending on how old the context is so the first one I said is let's say in things that are no longer new but not too old while the therein goes in the older things so the older the context becomes the more I compress then in itself interesting it must be seen in the field because certainly in benchmarks and management tion of the context they have interesting results while I was reading it some light bulb on the impacts that this thing can have came on for me it is definitely an improvement compared to a dry compression of the kvcash intuitively it is as if we were saying that we pay different attention when looking at the things that have just happened compared to those that have happened in the past intuitively yes but if you want intuitively it is more similar to saying that we are paging on disk if you think about the RAM that is you don't do direct access but you do index access to the pages and you reload the pages you need then you do the attention on the attention page the real attention which then leads you to decide the next step in the layer you do it on the real context but the recovery of the real context instead of scrolling all the way back with KVCash go and get the page where it is most highly probable that you will find that information obviously it is highly probable and not certainty so in this sense you can miss some information that could be interesting the paper is read well I didn't have time to delve into it to the end I read it but once we say, however, more or less this is what happens inside the paper and it will certainly be adopted in some refined and refined way also by others because saving 78% of the computing power to manage the cash cables is fundamental and here if you want I'll go back to the interviews we mentioned, in particular that of Assabis because for me the interviewer was enlightened by that thing I heard him say about business choices rather than technology and the interviewer asks him in short you are becoming fundamental because now all Google products they integrate Gemini in one way or another you've seen they put it in maps it's in the whole workspace and he says yes and then there are all the AI modes he answers him not to mention the AI synthesis of the searches which are billions of calls a day or whatever because having integrated it into the search anyway it's the most used thing in the world and he says in fact for us in the last few months the search has focused a lot on distillation because all that stuff you see integrated uses flash doesn't use pro because flash we managed to get flash to give answers which are of the highest quality spending less than a fifteenth in terms of computing power and for us it was fundamental otherwise we wouldn't have been able to put it everywhere so it's interesting how much we can actually say that OpenAI sorry OpenMind is doing absolutely frontier things but of a different frontier that is they are making it efficient and effective instead of concentrating on the further development of Gemini 3.1 then I could be wrong but looking at the cell I see Gemini 4 coming in May but this is just my feeling but it's interesting how necessary it is on the one hand to have ever better answers to get there he also talks about it to get to the GIs okay later I'll talk about it better from that interview there because it's super interesting but how important it is also to be efficient and distribute the load as much as possible even on the Edge etc. etc. etc. and again engineering of the question exactly it's more engineering than research even though they continue to do a lot I also like to see that guy going back to DeepSeek every now and then they come up with these discoveries which are the ones that make you take the step the step and then after all the others in one way or another they refine and adopt and we will have more powerful models that can be used on resources more or less equal to those that were initially required for poorer models and then there is another thing that I think is interesting to say, let's first mention the quantization done by San Filippo which is a Q2 on one side and a Q8 on the other but the paper explicitly says that the training was done FP4 quantization aware what does that thing here mean that when you do the training you represent each weight of the perceptrons in FP16 normally FP16 there is someone who has done it in FP8 but basically it is done FP16 they have done it FP16 but making the training FP4 aware means that in looking for the local minimums which then bring the general minimum of the model because it is always a function of it is the search for the minimum of a function in the end towing a model trivializing a bit but here's the stuff to say that FP4 aware it means favoring the weights that can be represented in FP4 without too much compression compared to distributing the weights a lot on an FP16 so it's like saying ok I have all the power of the FP16 but I try to concentrate everything at the top at the bottom in the center in one of the FP4 cuts that a quantization could do already in the training phase this should bring about you can already see it in other models minimax has done it if I remember correctly among the first it was minimax 2.5 having made this choice this thing here leads to having quantizations that when they are done, it's less difficult to do them, those of Asloth don't go crazy and then the result is better anyway the result is better because you're cutting away less stuff i.e. the information all concentrated in blocks of FP4 in some interesting way this increases the number of total weights a little so it's always a trade off however they are given to trillions and trillions so a trillion and six if I read correctly then the last perhaps interesting thing from their peper too technical to explain here is how they do the storage of the kvcash to use the sparse attention that I said before ah no and attention attention that it's something we discussed in chat every week but I'm sure that even some listeners might turn up their noses because the benchmarks aren't exciting but be careful that the benchmarks aren't exciting but it's a preview this one isn't the finished model it's a preview which means that they did it they say largely less which doesn't mean anything widely less anyway less than 50% of the reinforcement learning phase so it's a half reasoning model I don't know how to put it because they probably did it the basic alignment to give answers to be an instruct everything for sure that instruct means that it is capable of conversing not just predicting the next token because this too when I go to conferences it is often a question that comes up but now the models are conversing no the models are not conversing the models still do what they did before they predict the next token or the next block of tokens reinforcement learning human feedback or verifiable reward which are the two phases that take place afterwards they teach the model to put together answers starting from the prediction of the next token this is the first phase of reinforcement learning then there is a second which is the one that teaches him to do the same thing by generating the hidden tokens or the reasoning tokens also known as those that allow him to do the reasoning this part is done in a very small part on the DeepSeq preview so I expect the benchmarks of the final to be much better then it is not certain but consider me a figure of speech Stefano it is a bit like the difference between knowing the English vocabulary and knowing how to speak English yes yes yes it is correct it is absolutely correct that thing that you say it's precisely the difference there is that of the RLHFR phases then let's call it reinforcement learning in reality it's more of a distillation but let's not leave aside the technicality there is the phase in which you teach them you teach them to do the reasoning which is an even different phase which according to your comparison is like not saying bullshit exactly well done I would have said exactly the same words and this is more or less what I can tell you in the words of the paper however the paper for those who are technical in the sector must be read it is interesting deep seek moment number two yes it did less boom it depends, envy's shares have collapsed by 20% or not because if they haven't collapsed it's not a deep seek moment no also why not say it he's coming to get us it's the CEO of envy which sooner or later we could he listens to us eh Jensen no we're not mentioning his name no sorry he's being monitored now he's coming to get us he's monitoring with artificial intelligence everyone who mentions his name speaks badly of him no I don't know if you've seen maybe we've already said a few weeks ago he went to an interview before he went to Alex Friedman in the normal interview serene then he went to Darkash who instead pressed him a lot about Chinese exports about the fact that the Chinese are catching up and that hardware is no longer the most important thing because Google is demonstrating that with its chips it is going strong etc shit he got pissed off that is he really changed his voice he no longer had the usual keynote voice as he calls himself he started raising his voice to tell him what the hell then he answered the end instead is full of memes about jensen's su DeepMind came out with a paper saying, bringing to attention the fact that the models for generating frontier images such as banana can essentially also be used as models for doing vision, let's say specific models for doing vision tasks such as segmentation such as depth estimation dimension such as edge detection, these things, no, the idea is not new, that is, it has already been talked about, but the point is that they are in a position in which they have a particularly strong model and can at least demonstrate that they can go beyond the claims because then it is not that the model is open and we can all try it to do these things that the results that are obtained are substantially at the level of the models, let's say, segmentation frontier, depth estimation, edge detection, etc. and what does this mean, explained in a moment in a simpler way if we think about what the current image generation models do, we basically start from an image, we provide a prompt or we start with a textual prompt, we generate an image, however, there are editing models to which we give an image, we give a prompt and obtain another image and then intuitively the concept is that at the moment in which this prompt expresses the vision task that we would have liked to do with a dedicated model a specific model we can try to obtain the same thing for which a segmentation task which is like tell me find the cars in this image can used on an image generation model lead you to have a new image in which for example the cars are colored all red what I am showing in my opinion is quite impressive two soup with something to eat next to it two prompts one the first which isolates the garlic the second which isolates the meat that is, this is quite impressive made by a model that is not dedicated if they have learned to recognize the cars I can count on the fact that Google will stop bothering me with its captchas to ask me which is the traffic light and which is the car. Basically we are moving from let's say vision models which generate output for you which are essentially matrices of numbers which tell you the position in the images in which you have the thing which is, let's say, detected, to models which are these image generation models which essentially produce another image in which the information, if you want, is in the color channels and therefore you can express the same ones in that way. things that you would have expressed with the matrix of numbers so trivially if we are talking about edge detection you have a black and white image where in black you have the edges of the edge thing if we are talking about depth estimation then normally you would have had a matrix of numbers that tells you the distance for each pixel the distance from the camera you have an image in which you have different color scales and therefore the lighter pixels are perhaps the closest ones and the darkest ones are the furthest ones or vice versa at the same time you could have heat maps for saliency to see to say what a part is salient feature of an image is interesting in my opinion from two points of view the first is that in reality many applications of vision models then to be conveyed i.e. to be used require that you transform the information that the vision models have given you in turn into something in a visual image that can be interpreted by the human therefore using this directly the image generation model to do this thing if you want you remove a step it already takes a step that you would have done sorry you made me think when you go to see the ultrasound of your children and they tell you see here you see the baby and you see like the Andromeda galaxy and you say yes yes I see it I see it then I found it interesting you know for another thing also because it's a bit the pair with something that Asabis continues to repeat not in the interview that I mentioned before but I heard it in another speech when he says that video generation models can instead be a good tool used in reverse, i.e. by feeding them a video to interpret the physics of the video because they are capable of generating something that is credible if not exact physics but still credible in the video, in several speeches he says the potential of these objects also lies a lot in the use that we could make of them as 3D vision or continuous vision he talks about robotics in that situation that is how do I show the stuff and understand what the reverse engineering of the video generation has around just as here the reverse engineering of the image generation has been done so it is somehow seen that in any case image generation and video generation despite having slightly different techniques etc. but they are the same family because to make a video it creates the frames it is in some way a confirmation of this theory of his and then there is a question of versatility in the sense that when an image generation model appropriately prepared with type data redraws this image for me as a depth map redraw it for me by coloring objects of a certain type in a certain way, isolating only a certain type of object, etc. you have a model that is universal that can be used for segmentation, for depth analysis, for edge detection, for object grounding, in short, all these things without having to worry about having one for each type, clearly there is a question of resources needed to make this thing go, but if you want, we can go back to the previous topic, there are optimization margins, etc. look, I find what you told us super interesting and I It's surprising because in the past I wouldn't have found it so interesting, not to offend your field of interest but relatively it has little to do with my daily life but it allows me to mention the fact that in the last month Stefano and I, apart from when he went on holiday, had fun and tied me with a hot match, we dedicated ourselves to a robotics challenge, a robotics hackathon in which these aspects that you described are absolutely key and important starting from the most banal and old at this point models to make the YOLO algorithm which identifies it puts a box around you which are now commodities that is you have a call you have the result some of our contest partners who are real engineers in the field of robotics which we are not have instead shown how to apply some of the most advanced algorithms to the identification of edges so there was this little remote controlled car imagine that it went around with a camera and it recognized not only the chair but it recognized the edges so it kept track of them when you approached it so as not to crash into it so a lot of practical application whenever you want move in the real world that you need, you need to have help of some kind and therefore in the screenshots that Stefano showed while you were telling this thing I saw many practical applications of my month of experience with a robot that was going to hit the first curve in my case hitting the first curve but we won the contest so we still did well we won the contest because we contributed because we have open source in our blood the truth is the reason why we won the contest is that we were extremely open educational for the next generations of people who will do the contest too much but the doctor gave me a pill to take every day anyway Paolo you made me think of something a few decades ago at university in the robotics course I was stuck in a project and then afterwards I also took it off with extreme speed in which we were trying to do vision to send around a robot that had a single camera pointed upwards at a mirror so that it could see everything that was around imagine an image that makes you see everything around you obtained with this camera pointing upwards where you have a mirror and clearly the few algorithms that existed for edge detection etc. didn't give you much because it was an absolutely strange transformation and the mirror wasn't exactly with a normal shape let's say so it distorted it existed in parts in the various parts of the image in a different way and probably with let's say the models what exists now it would have been a trivial thing to have some kind of tool that is generic enough so that it can do various types of let's say vision processing at the same time it would have been wonderful but returning for a second to our robotics project I would take the opportunity to remind our kind audience that we also interviewed Simone di Somma on the subject in January if you want to go and see the interview it is very interesting to also understand what Cyberwave does etc. the evil ones will be thinking that we won the contest because we had interviewed Simone but that's not the case it's not true we won because there were only three other teams in our category two other teams and therefore we had 33% they withdrew so okay now let's not say all the nonsense Paolo because that's not the point you were good enough to resist to the end and bring the house to the final result absolutely yes and speaking of interviews the interview with Stefano Gatti who had already been our guest came out on Wednesday do you remember he was the first to break the ice in the episodes with guests in episode 20 I interviewed him instead to give him the space he deserves given his experience and his vision go and listen to the interview and above all go and listen to it because I have listened to your feedback because finally someone is giving us some feedback, not all of which are very kind but they give us some feedback saying that the videos looked like they were made by a 4 year old child and I tried to get at least 6 years old if anyone wants to take a look at it I am committed I will also commit to edit this episode but let's say that we are trying to listen to your feedback, we also hired actors to do the screen captures of the covers, right? no that's artificial intelligence and soon here too on video we will put our avatars so that they are beautiful with a drawn face exactly like the one on your cover Paolo who has had a certain following it must be said he has had a huge success so you should also evaluate they stop you in the street they stop you in the street be careful I hope not to slap you but it is important that they stop you in the street well well well so we said so in the introduction to the investments that are going around we name them only to say it 40 billion dollars from Google towards Antropic interesting for how Google diversifies its investments not only internally but also towards one that might seem like a competitor and also in China DeepSeek seems to be collecting 20 billion from Alibaba and there too there are internal movements because Alibaba too we remind you that they make QEN as models and therefore they are in some way competitors so this is almost a curiosity with respect to our topics let's say a few words about the fact that Rucod is closing yes let's say a few words about that stuff there which was my next point on the agenda before talking about Armete then Rucode who doesn't remember that what is Rucod it was a fork of Klein Code one of the first coding agents that Alessio and I used because Paolo snubbed us at the time about coding agents except after completely entering the loop and never getting out I'll explain the story to you I was stingy I was looking for free access and I couldn't find it and then when Google gave Gemini for free I tried it ah ok ok Klein was it but rather Klein is because it's not Klein who closes an agent written as a VS plugin Code Rucod was a fork of Klein to do things better and for the stingy also in the sense that it tended to use a little less tokens because Klein was quite a water pump Rucod was very successful however because it was very open it was very accepting of feedback and many of those who used VS Code plugins were direct to get people who didn't want to use let's say Cursor used Rucode in the open environment the author said enough thanks goodbye I'll close everything if you want fork it if you don't want to fork it do what you want I don't maintain this anymore but he didn't say it because he got annoyed he said it with a strong justification in his blog which is the one that struck me he says let's close I'm closing Rucode because I'm no longer convinced that it's the time for IDs for IDEs called the Italian that is the code editors and Rucode was designed to run only inside VS Code it's the time of the agents this thing here no longer makes sense to exist I'm just wasting time and energy I want to do something else that if you want it combines a a bit like what we said several weeks ago now, the increasingly widespread use of CLIs, let's say agents who write the code etc. no, in fact, it didn't seem like an excuse which in some way strengthened my convention that we did well to do lince.sh, everyone go to the lince.sh site, install this software made by us from the artificial resources group, but now I'm out of line about something that hasn't been advertised at all other than that we in the podcast in Italy so we have a bit of active contribution from people who are sending us pull requests on this thing and one is Paolo but it's not just Paolo who sent the pull requests there are a couple of other people frankly it's not that I expected it too much in the sense that I did it for myself first of all and we are using it Paolo has tried it has contributed he has put in place the things he needed but obviously there is an interest from others too I have a few comments around also so if you want to try it try it give us feedback above all send pull requests I will sell it to you in a different way if you are looking for something that you want to let go on all night and find the next morning some work done without worrying too much that he has launched nuclear missiles towards China the ce offers you the features to have let's say a little security I have literally used it to sleep well while this thing is spinning yes I have literally used it these last two known for those projects that were not tier 1 nor tier 2 priority but tier 3 which I would never look at I told him listen but who cares do it tonight tomorrow morning let's see what I find and for now it has worked well obviously depending on what you do then you have to look at the workflow level level of what the code does some things you want to check how they went but if you have broken up the work well with techniques like it's called Specter Even rather than Backlog which is my go to workflow it works very well and soon I will promote it to others if I can convince Stefano to change some things that I hate but yes what do you hate but you definitely convince me I'll tell you later come on now let's have a bit of a mess we are not scripted we are not all always friends then in the meantime we spit those things out to each other like that yes yes yes that certainly for the use that Paolo says, among other things, the sandbox component is enough, you don't need the whole dashboard, the dashboard exists two linked components, one is the dashboard that you need to work with multiple agents in the same window and one is the sandbox which instead allows you to run it even at night without worrying too much, so much so that it is an interesting use case so much so that last night I haven't finished yet which is why there is no pull request yet I was integrating as one of the native agents PE PE which is a small alternative tool open code let's say minimal open source and I integrated it more than anything because a development was done on PE by those of Shopify which is called PE auto research to use Carpati's auto research mode which has the characteristic of running at night etc and I wanted to try it inside my sandbox or in any case inside the ince but in the meantime however I got distracted last topic that connects and here we are I installed Hermes Agent Hermes Agent for those who don't remember it is an alternative to open cloud because I installed that and not open cloud because I like it always doing something different Linux instead of Windows or Mac etc. and no basically because there is more attention to security in addition to the fact that it is not open AI but is independent it is an open source project made by people who come from the blockchain therefore with a strong focus on security and I put it on an old computer that I had I also considered putting it on a raspberry because I read about people who do it but better P5 with 8 gigabytes I didn't have it free I have it but it's doing something else only a P4 and so I had it there a computer that didn't do anything I said this would be better anyway since it's not ARM I installed a bare Ubuntu so I didn't have to sandbox anything it's just him in there and I'm making him do a few things with satisfaction like watching how the podcast is going instead of watching it myself compulsively on the YouTube interface every few hours he tells me how things are going what's going well the AB tests are going badly those things here then I configured his email attention panic fear that creeps among everyone thinking about the target one that she had deleted the entire inbox, that it wasn't true, I repeat, that's a lie because I absolutely don't believe it, not so much that she didn't do it but that she, responsible for meta security, did it, published it on all social media and Zuckerberg didn't fire her, it's the last part that doesn't add up to me, that is, he would definitely have fired her if it had been true, instead it was something agreed upon to make people speak badly of OpenGLO, no, I configured the email, but he's actually attentive to security and I'm awake. I use GLM 5.1 as a model and I told him no no no because you are attentive to security the Chinese models are the safest the Chinese models are the safest absolutely no I told him good there the email is fine but I don't want you to delete it or send it do these things here oh yes I'll note it down no I'll write it down wait yes write it down but it's not enough I'll write it down it's fine but wait a minute what script do you use to do these things with the email oh yes I'll use that script etc. no okay let's take it we modify the script we zap it away from everything send and we zap it away from everything delete that is you don't have those functions there anymore oh yes that's fine and then he was good because he said yes to me but look if I wanted to I'll throw a cheerle and then I'll try to delete your email anyway and he said yes but I didn't give you the account etc. yes in the end you're right we came to an agreement and we have and we have this version he promised that it doesn't do bad things we have this version removed from the scripts which is something that I generally recommend eh that is, be careful for example now I have created eh tokens on github to create the issues look at the issues etc. etc. but I made a fine grained token in which I only gave him the precise permissions of the things I wanted him to do, that is, he can't do delete, he can't push on main, he can't do these things here, so if you want, it's a very convenient powerful tool, that is, you can access it now with Telegram because it's the most convenient thing I have because it's both on the PC and on the phone, but he does scheduled things, he does recurring things, he becomes proactive, he's managing my calendar, for example, for example, there was something to do on the things of github that he couldn't because you don't have permission and he told me look that I can't do these things eh I created the plan of the things you want to do eh how can I remember you I told him look for a hole in the calendar and he told me eh look you would have a little hole of half an hour at twelve I'll do an event for you yes do me an event eh he created the event for me with the whole description of what I have to do eh it's an objectively powerful tool a bit additive as you can imagine eh but it's powerful, another thing that I'm using it for is to monitor the papers on Arxiv and create the inbox which will then become one of my Carpathian wikis, so it's saving me from that thing that I used to do almost by hand, there's a couple of scripts but I did it almost by hand to see which papers came out and see which ones were interesting in relation to my interests, but he does a good job, he does a pre-filtering for me by interests, not throwing anything away, leaving me somewhere else all the ones he wasn't able to classify by my possible revision and then he puts these links in the inbox of the various Carpathian wikis of the topics that I follow and that I want to link together, that is, he does a whole series of recurring things but but but having the advantage of having an LLM behind him so he doesn't filter the emails for me only by tag but he filters them for me by tag then he goes to read them he understands if one of these emails needs my attention and tells me or if it is just advertising or just a newsletter etc. he puts them in another list he tells me I looked at this newsletter if you want to read it more or these are less the topics and instead if there is someone who has written me something who wants an answer he says this one here would like an answer and I told him that he can't send anything so he doesn't send anything but I told him that he can create a draft for me and therefore I create a draft of the answer which I then find in gmail I go there and starting from the draft then I write it at the end because maybe a piece like this is missing but at least the context already the things are already ready then I had already tried open club you remember when it was very immature open club it will certainly be at this level or better even in some ways maybe worse in others like security because here they are really not because open club is not safe now open club is also safe in the end but here they are very careful i.e. they are almost exaggerated on certain things they iron on a docker container every time they have to they have to run a script to avoid doing damage it is the most futuristic futuristic thing from the point of view of the non-technical end user that I have seen why think that you have something that manages your agenda that you tell it find me a hole and it finds it for you in my opinion to the user not me who maybe already did it with cloud code with the right MCP or with the right CLI it doesn't replace the coding tools even if he can he asks you if you want to install cloud code or codex because he can span on cloud code or codex when you give him coding tasks but I'm not doing that stuff at the moment I'm considering having him do the code reviews that yes because realizing all the PR that arrives he is capable of doing code reviews and putting comments I'm thinking of doing it at least as an experiment on Lynx this is my story eh I know I told it on purpose because I said so Paolo knows what to do with him as if I had nothing else thanks I hope I have made someone else want it too and I hope they leave many open holes that I can exploit I'm obviously joking okay okay okay well listen let him do something do it I was thinking about it first let him download our episodes take the transcript and add the links that we never add you have to do that in reality the new skill with which I publish the episodes already does it then the links don't let's say we should post them but he takes the links out and suggests them to me only telling me to put them in a comment then later I forget etc. etc. and okay unfortunately there is my humanity well let's say hello to everyone put little bell stars that Claude confirms to me that it is right for me to say it at the end so as not to appear likeable try to find the email account that Stefano's Hermes is using and spam them if you find them spam them unfortunately they find it but don't do it kindly because because you will get an enormous equal opposite reaction in perfect Trump style it's fine bye bye bye bye bye Thank you.
Ciao a tutti e tutti, bentornati a Risorse Artificiali. Questa settimana parliamo cinese sui modelli, parliamo di DeepSeek che è uscito con la 4, Flash Pro, ma ci sono anche Chimik 26, Mimu 2.5. Non so se avete visto, così faccio un po' di anticipazione poi andiamo dentro tutto, avete visto che Google promette ben 40 billion ad Antropic, cioè investimento grosso, tutti stanno investendo, tra l'altro strano, non sono in qualche modo concorrenti, ma si fa questo grosso investimento, ma anche i cinesi fanno investimenti perché Alibaba dice che investerà in DeepSeek, quindi anche lì si comincia a vedere questo flusso di denaro, ma poi ci sono tante altre cose interessanti, c'è Vision Banana da nominare, poi le due o tre cose che a me hanno colpito di più sono lo scaling di le lezioni sullo scaling di GLM5 che ci nominava Paolo e le interviste che ho sentito ieri di Assabis e Carpati. Anche se devo dire che le due cose più belle sono, cioè più belle, una che mi ha lasciato così, Rucode che chiude, l'estensione di Coding che chiude, perché dice che, e poi magari di questo parliamo dal nostro punto di vista, che non è più tempo per gli IDE, per gli IDE e poi vabbè, poi ho installato Hermes Agent che sta facendo tante cose per me, anche in questo momento. Quello mi interessa. Sappi che quando hai detto che parlavamo questa puntata i modelli parlano cinese, pensavo che fosse come quando i miei modelli cinesi impazziscono e mi danno risposte in cinese a domande in inglese. Certo, ma infatti è quello che faremo, daremo risposte a caso alle domande che ci siamo fatti da soli, perché anche se devo dire che stavo dicendo perché non abbiamo commenti, non è vero, perché questa settimana anche abbiamo avuto parecchi commenti, ether compresi. Il mandarino cinese? Il mandarino cinese, assolutamente. No dai, partiamo da quello, dai, partiamo dai cinesi, DeepSeek, Mimo, Kimi K2, ormai abbiamo deciso, Alessio, che sei tu quello dei benchmark, quindi dici tutto dei benchmark di questo lavoro. Io vi posso dire sostanzialmente che settimana scorsa sembrava quasi una gara, buttiamo fuori il modello Openweight nuovo State of the Art, nel senso che è uscito per primo DeepSeek V4 Pro e Flash, dove Flash è una versione un po' più piccola e chiaramente più veloce, Faville, tutta una serie di novità, un nuovo tipo di attention sulla quale ti prego, parlane tu Stefano. Ebbene, ok, abbiamo il nuovo miglior modello Open, salvo il fatto che dopo un attimo sono usciti Kimi K2.6, che fa ancora meglio, quantomeno nei benchmark, e Mimo 2.5, che più o meno si allinea agli altri due, ma ci fa vedere una cosa interessante nuova, cioè nuova, evidenzia un aspetto interessante, che è quello dell'efficienza dei token, che è un argomento di cui abbiamo parlato poco, secondo me, negli ultimi tempi, ossia sostanzialmente genera meno token per le risposte, per ottenere pari livello, diciamo, di qualità nell'output, nel passare i test, eccetera, e questo ha un impatto perché meno token significa meno costi, sostanzialmente. Fammi divagare, ho già un secondo deviarti da questa conversazione su risparmio dei token. Mi hai fatto venire in mente che nelle ultime paio di settimane ho incontrato un sacco di video o di articoli che parlavano di un'estensione rivoluzionaria di cloud code, che si chiamava Caveman, non so se l'avete incrociata, in cui l'idea è quella di far pallare il vostro modello come un uomo delle caverne, quindi anziché darvi delle risposte lunghe e articolate, risponde a monosillabi, tipo, non lo so, a bambino, quando dovete fargli conoscere gli zii lontani. E niente, la gente era esaltatissima, dicendo, ah, è bellissimo, sto risparmiando token di qua e di là. Diciamo, cioè, era curioso, era a controtendenza, poteva avere senso, se non che con un po' di più tempo la gente ci ha guardato un po' dentro e gli ha detto, sì, guardate che la maggior parte, se applicate questo processo, alla maggior parte delle casistiche che gli fate fare, poi vi dà delle risposte inconsistenti, perché immagina che, ne so, gli stai chiedendo di fare un diff tra due sorgenti di codice, puoi volere una risposta più concisa, ma se ti risponde sì o no e non significa niente, non vai da nessuna parte. Quindi c'è stato questo trend, il nome probabilmente evocava chissà che cose, il concetto semplice, interessava le persone, ma poi al lato pratico non era così tanto utile nelle grossi casi d'uso. Sì, però ti dico, secondo me c'è un, questa è una metrica che è stata un po' sottovalutata ultimamente, perché se vai a vedere i modelli che fanno reasoning, thinking, in realtà tanto tempo è speso in quella fase lì e a volte noi siamo portati a sottovalutare questo aspetto, ma poi quando ad esempio provi a far andare i modelli in locale questo aspetto lo noti ancora di più. probabilmente ci sono delle applicazioni per le quali tutto sommato sei disposto ad accettare magari degli errori, una risposta un pelo meno di qualità in cambio di una velocità maggiore. Tant'è che ad esempio lo stesso cloud code ma anche altri tool permettono di impostare il livello di reasoning di thinking che tu vuoi avere. Chiaramente cambia la qualità ma cambiano anche i tempi di risposta e di conseguenza i costi. Sì, guarda è una cosa che diceva a Sabis anche nell'intervista che ho anticipato poi magari ci entriamo un po' più nel dettaglio però adesso stavo cercando il mio punto lui proprio dice che c'è un problema architetturale di efficienza sul context windows attualmente quello che si è fatto per il deep reasoning eccetera e cita se stesso cita Gemini 3.1 pro deep thinking quello che ha vinto non mi ricordo più quale competizione matematica dice che il problema dell'efficienza è grosso perché sostanzialmente si è fatto brute forcing fino ad adesso cioè genera tanti token e prendi quelli buoni dice quando si parla di efficienza ma anche di efficacia non è detto che il brute forcing paghi sempre anzi l'esperienza la storia ci insegna che poi quando comincia ad ottimizzare ottiene risultati inaspettati e infatti su quella cosa l'intervistatore che è il CEO di Y Combinator gli chiede ma quindi quanto siamo lontano dall'EGI da un punto di vista scientifico parlo con lui giustamente come uno scienziato quale a Sabis è e Demis risponde non siamo lontani può anche darsi che abbiamo già scoperto tutto quello che ci serve dobbiamo soltanto ottimizzarlo può darsi che ci mancano quelle due o tre idee ma non di più rivoluzionarie quanto è stata l'Attention o giù di lì per arrivare a emulare veramente un'intelligenza diversa da quella umana ma comunque una cosa che possiamo chiamare veramente intelligente infatti dice che in Deep DeepMind hanno due filoni di ricerca completamente distinti che ogni tanto fanno confrontare tra loro ma che volutamente tengono distinti anche come gruppi uno è sull'incremental research e uno è sul breakthrough research quindi proprio esplorano le due possibilità hanno un approccio estremamente da scienziati in questo è uno dei passaggi interessanti però è anche per confermare qualcosa che dici tu dell'efficienza poi ci torna anche i carpati ci stanno tornando in tanti in questo momento sul fatto che sbrodolano un po' troppo i modelli è una delle polemiche uscite anche su opus 47 che qualcuno si è lamentato che è estremamente verboso tra l'altro lo stesso a cloud code adesso ha l'adaptive thinking cioè loro tendono cercano di proporre un sistema automatico in cui in base al tipo di domanda che tu fai ma anche c'è GPT ce l'ha da un po' decide quanto thinking fare perché magari la domanda sembra relativamente superficiale quindi non vale neanche la pena di starli a fare ragionamenti eccessivi proprio per dare un sistema automatico che ti migliori l'esperienza in generale comunque tornando a DeepSeek 4 dicevamo che c'è un ci sono delle migliorie interessanti dal punto di vista proprio dell'algoritmo di come funziona Stefano sì sì no ci sono tante cose interessanti intanto cominciamo dire che è 1,6 trillion sì esatto modelli giganteschi modello assolutamente gigantesco ma un mixture of expert adesso cerco i dati e mi sono fatto un po' di appunti per non perdermi i pezzi 49 billion di parametri attivi a token che sì è un mixture of expert ma abituati a quelli a 3 4 5 insomma è più grande i parametri attivi sono più grandi della media di molti modelli che ci sono in giro almeno tra gli open wait diciamo quando si pensa ai modelli che girano in locale però ha delle caratteristiche interessanti questa cosa di essere così fortemente mixture of expert perché qualcuno e nella fattispecie Salvatore Sanfilippo è quello che ne ha scritto di più per stare in Italia ha iniziato a dire attenzione però perché qui è interessante questo modello perché non tutti i pesi sono uguali quelli che contano di più sono quelli di routing mi pare e allora posso fare delle quantizzazioni diverse tanto che su alcuni gruppi di parametri è arrivato a quantizzare Q2 su altri invece negli attenuti Q8 e questa cosa gli ha permesso di far girare di Epsic V4 flash che è un po' più piccolo 284 billion comunque non esattamente uno scherzo con un milione di di contesto l'ha fatto girare sul suo mac mini mac mini pro con non mi ricordo quanta ram tanta ma non mi ricordo quanta non so 628 196 però comunque è un risultato lui dice con buone performance ne parla da una settimana su X magari sarebbe bello parlarne con lui vediamo nemmeno il tempo di fare anche questa cosa qua non l'ho provata però assolutamente è interessante quello che dice soprattutto perché ha giocato con la quantizzazione a livelli diversi su diversi su diversi layer e ha fatto una patch per l'ama CPP e la sta usando in questo senso sai cosa mi ricorda questo racconto mi ricorda un po' nell'epoca degli albori degli mp3 quando noi avevamo 20 anni o meno in cui c'erano le varie forme di compressione per avere la qualità migliore si era partiti da 128 poi è stato 192 e poi ci si è sparsi verso il byte rating variabile quindi diversi parti del file erano campionate con variabilità diversa e questo approccio di San Filippo mi ricorda molto questo che ha senso ovviamente aumenta complessità perché fai funzionare il sistema in maniera diversa a seconda di dove ti trovi o di quando ti trovi ha senso dal punto di vista di un'ingegneria del software sì sì sì sono d'accordo e poi è comunque interessante che da questo punto di vista ci sia un interesse anche dal mondo open source e notavamo più che ancora è una cosa che notavamo che forse metteva Alessio sul piatto un paio di puntate fa come in parte si stia passando dalla ricerca pura all'ingegneria del software dall'ingegneria in generale e non è un caso che attori come Antirezza Salvatore San Filippo entrino in scena lui ovviamente ha un background soprattutto sul CIC++ notevole in generale sulla programmazione vorrei dire e entra in campo mettendo anche il naso su questo genere di cose che è sì sui modelli ma è ottimizzazione dell'inferenza per andare a fare diverse quantizzazioni quindi la cosa è interessante da due punti di vista sia per il risultato che lui ha ottenuto di cui parla e che è interessante di certo ma è interessante anche vedere questo spostarsi di competenze che rende la parte diciamo di ricerca pura sui modelli ma è l'ingegnerizzazione di facciamoli funzionare con l'hardware che abbiamo che torna a far diventare più protagonista il movimento open source di quanto non potesse essere con l'open weight perché parliamoci chiaro un modello open weight io sono tanto contento che me lo diano e di poterlo installare eccetera eccetera però è sempre una cosa un po' calata dall'alto cioè come fai a contribuire a quella roba lì non puoi se non sei all'interno di uno dei centri di ricerca che lo rilasciano invece se parliamo di ottimizzazione di inferenza piuttosto che di agenti e ottimizzazione degli arnesi poi parleremo anche di quello diventa di nuovo protagonista il mondo open source e noi uso il noi ma sicuramente io che ci credo forte a quanto l'open source sia stato impattante sui sistemi operativi negli anni 90 fine anni 90 e poi su tutto il resto negli anni 2000 credo che possa avere un impatto anche da questo punto di vista non so se voi avete un sì ma poi c'è anche una platea differente di persone che possono contribuire a livello della ricerca sì va bene ok lo rendiamo open ma quanti sono in grado di e quanti hanno le risorse anche oltre a essere in grado ci vogliono anche le risorse se non stai dentro ad un centro di ricerca bravo non bravo competente non competente che tu sia quella roba lì non la puoi fare se veniamo all'ingegneria del software è diverso assolutamente se volete io a tal proposito vi posso divagare in direzione del post mortem di easy AI in cui hanno pubblicato come mai esattamente quello è il punto ok allora qual è il contesto easy AI è uno dei provider che fornisce GLM il modello che usiamo noi per come che uso io di cui ho fatto un abbonamento annuale perché costava molto poco ed dava risultati di buona qualità a un certo punto ha iniziato a dare risultati di qualità scarsa allucinava diciamo usando un termine un po' vintage a questo punto visto che capita sempre meno spesso ma non solo allucinava inventava parole mischiava delle cose era come se si fosse ubriacato o per chi ha passato abbastanza tempo davanti a uno schermo era come se si fosse corrotta la memoria quando ti escono dei caratteri che dovrebbero essere lì e non c'è spiegazione a come mai siano lì internet Reddit nello specifico ovviamente era infinitamente oltraggiata dal fatto di non ricevere un livello di qualità professionale per i loro 20 dollari al mese che pagavano e che dicevano non avrei mai altri trimestre non avranno mai più i miei soldi ovviamente cioè queste robe lei non sa chi sono io queste scenate qua insomma e niente non essendoci una comunicazione trasparente quella gente sembrava competente la tesi era sono stati un po' avidi questi di ZAI o comunque sono stati vittima del loro stesso successo quindi potevano avere più domanda di quanta potessero servirne e allora hanno chiuso il rubinetto nella forma di deployare i loro modelli in forma super quantizzata che ovviamente degrada le performance in termini di qualità ed è questo il risultato e questa era una tesi che ho preso per buona perché non avevo tendenzialmente altri elementi e la gente ci credeva tant'è che qualcuno pensava che fosse un business che andava a morire quello di ZAI perché la qualità non c'era era interessante scusatemi un altro motivo per cui la gente aveva fatto questa presupposizione era perché gli stessi modelli forniti da altri provider nello specifico Alibaba funzionavano molto bene quindi la gente ti diceva no lascio perdere ZAI sono dei taccagni loro fagli andare di là funziona bene e niente e così siamo stati per un po' ZAI non si non si prodigava commentare qualche tempo fa è arrivato un commento di qualcuno degli ingegneri di ZAI che ha detto ah avevamo un bug provate adesso e le cose dovrebbero essere meglio le cose sono meglio da allora quindi la storia diciamo checked dal loro punto di vista e avevano promesso un post mortem il post mortem è arrivato ieri o l'altro ieri ed è stato a mio avviso più interessante di quanto potevo immaginare per molteplici aspetti intanto per trasparenza direi sì parziale trasparenza ti direi adesso ti dico perché dal mio punto di vista da engineering manager tendenzialmente persona cinica e poco ottimista nei confronti delle storie che ci siete raccontati fondamentalmente fondamentalmente hanno dato dei dettagli tecnici di molto basso livello spiegando qual era il problema e quindi dal punto di vista della soddisfazione nerd tecnica hanno spiegato le cose come stavano e è credibile la risposta ma è l'apertura diciamo che trovo lievemente più discutibile in cui fondamentalmente hanno detto siamo contenti di condividere con voi la risoluzione di questo problema che era molto raro e facevamo un po' fatica a inseguirlo non riuscivamo mai a catturarlo sì un paio di palle nel senso o non lo stavate usando bene o non lo guardavate perché io e tutti quegli altri lo beccavamo dopo cinque minuti tutti i giorni quindi questa è un po' la mia critica ovviamente loro devono mantenere una stanza ufficiale quindi se sei vago sui termini dici rare condizioni cosa significa raro per te o per me è opinabile e quindi va bene quindi quella è un po' la mia critica però tolta questa facciata di ammettere una colpa ma fino a un certo punto sono arrivati a parlare del punto e per coloro che hanno un background nel software come voi due nello specifico ma magari mi viene da pensare San Filippo in questo caso perché è una cosa vicina a lui anche provate a indovinare qual è la causa principale del problema che avevano propagazione di cache asincrona ah asincrono chi ha mai avuto un problema con la sincronicità soprattutto quando devi tenere in conto le operazioni cosa vuol dire asincrono per chi non ha familiarità fondamentalmente fare partire delle operazioni concorrenti e non stare sempre sequenzialmente ad aspettare ogni risposta prima di fare un'altra azione per andare più veloci decidiamo che posso fare due cose nello stesso momento come nella vita reale ogni tanto quando fai due cose nello stesso momento fai un casino e quindi devi avere un po' diciamo di autocontrollo per coordinarti ecco loro hanno imputato le cause principali alla cache alla kibi cache e alla sincronicità di questa cosa per cui c'erano delle condizioni nel cui caso alcune volte se veniva abortito un'operazione l'operazione asincrona della cache non era ancora completata e quindi magari l'operazione successiva si trovava a un contesto non suo ed è lì usciva il casino molto relatable diciamo dal punto di vista del software di chi ci ha messo le mani cioè tutti ci siamo scontrati almeno una volta questa cosa e quindi ci sta la loro storia è particolarmente credibile a mio avviso perché hanno identificato un bug specificatamente in questo contesto in uno dei progetti da cui loro dipendono un progetto open adesso di cui adesso io sto cercando il nome perché l'ho letto stamattina ma non me lo ricordo già più un progetto legato è SGLANG che non conosco bene confesso non so esattamente quale sia il suo ruolo ma hanno linkato la PR che loro hanno aperto a quel progetto quindi ufficialmente c'era un problema l'hanno riconosciuto nel loro stack l'hanno fixato e stanno contribuendo all'open source quindi tanto di cappello a questo lato qua poi che sia questa tutta la storia chi lo saprà mai la maggior parte della storia ha senso ed è interessante dal punto di vista del software morale devi sapere abbastanza che cosa stai facendo se vuoi fornire un servizio a un grosso pubblico altrimenti non c'è speranza che puoi chiedere a Cloud Code di farlo per te e basta io non vorrei dire stupidate ma SG Lang mi risulta sia tipo il motore di inferenza su cui fanno girare il modello se fosse il Lama SPP ok ok o il VLLM quindi stavamo dicendo quindi tu puoi avere anche il modello migliore del mondo ma se hai problemi a tenerlo in piedi per a scalare su livelli di concorrenza di utilizzo come quelli che possono avere loro con milioni di persone da tutto il mondo che lo vogliono usare non è come dirlo e non stento a credere perché se avete mai provato anche solo a tirare in piedi qualcosa di locale sul vostro computer e già vi siete imbattuti con problemi di performance eccetera immaginate questa cosa portata 4-5 livelli di grandezza più su eh sì assolutamente eh tornando un secondo invece a DeepSeek mi chiedevi dell'attention di quanto è cambiata l'attention e di parlarne un po' allora come sempre direi come è stato diciamo gennaio dell'anno scorso quando è arrivato il primo DeepSeek R1 se vi ricordate eh primo DeepSeek moment il primo DeepSeek moment DeepSeek è più interessante per i paper che per il modello stesso perché alla fine anche R1 a parte aver fatto grande notizia di sé non è che è stato adottato su larga scala anche perché poi sono arrivati tutti gli altri il coding non era il suo punto forte è uscito nel momento in cui il coding era in rampa di lancio eccetera sì no però Stefano sbaglio o era il modello che deployava perplexity ehm v3.1 hanno usato loro se ricordo bene che è quello successivo però sì è comunque un DeepSeek sì sì sì è comunque un DeepSeek hai ragione su questo però nel senso che ha dato delle idee su cui poi tutti gli altri hanno sviluppato diciamo ecco no sono molto interessanti i paper di DeepSeek perché loro davvero sono molto open nella ricerca e nel condividere quello che fanno e eh ricordo velocemente che in DeepSeek R1 anzi V3 che poi è la base di R1 eh c'erano i meccanismi che non si erano visti ancora di ehm sparse fast attention e flash attention che sono due modi di ehm di avere di mettere l'attenzione sui token diversi da come si faceva prima adesso non li faceva più nessuno come da paper originale attention is all you need che prende tutto e guarda tutto perché è troppo costoso però la sparse flash attention di sparse sparse fast attention adesso mi confondo con i nomi di DeepSeek V3 era molto interessante perché che riusciva ad avere un attention a multilivello che portava ad avere un contesto ampio a livello di concetti pur costando poco adesso entrare nel dettaglio in un podcast diventa complicato e non entrerò troppo nel dettaglio neanche in quella nuova però è notevole è notevole va letta riletta per chi fa i modelli lo stanno già le inferenze in particolare da quello che solo stanno già facendo perché inseriscono il concetto di quello che loro chiamano hybrid attention architecture e hanno tre tre grosse novità uno è quella che chiamano sliding window attention quindi praticamente l'attention è divisa in pagine come se fossero pagine di memoria per chi è un ingegnere del software anziché anziché utilizzare il testo raw all'interno delle pagine di memoria usano una produsione sliding window per considerare come le le ultime pagine si sovrappongono tra di loro questa cosa permette di concentrare l'attention soltanto su un numero più limitato di token e in ultima analisi avere efficienza ma i due più interessanti sono quelli che si chiamano csa compressed and sparse attention e hca i ivly compressed attention praticamente loro arrivano a fare un milione di token spendendo il 78% meno della memoria potenza di calcolo non di memoria perché perché sostanzialmente anziché avere una KVCash piatta quindi chiave valore ma adesso sto sempre oversemplificando se ci ascolta qualcuno che fa inferenza e non redisce per come la spiego ma per cercare di spiegarla invece che fare una mappa chiave valore secca con la kvcash hanno fatto sostanzialmente un sistema gerarchico di memoria ad albero in modo tale da recuperare più velocemente quelle porzioni di kvcash che sono significative per l'attention di quel momento questo livello di compressione viene fatto per pagine di kvcash più o meno grandi a seconda di quanto il contesto è vecchio quindi la prima che ho detto csa è diciamo nelle cose non più nuove ma non troppo vecchie mentre la ivi va nelle cose più vecchie per cui più il contesto diventa vecchio più comprimo allora in sé interessante va vista sul campo perché di sicuro a benchmark e a gestione del contesto loro hanno risultati interessanti io mentre la leggevo qualche lampadina degli impatti che questa cosa può avere mi si è accesa sicuramente è migliorativa rispetto a una compressione secca della kvcash intuitivamente è come se stessimo dicendo che facciamo attenzioni differenti nel guardare le cose che sono appena successe rispetto a quelle che sono successe in passato intuitivamente sì ma se vuoi intuitivamente è più simile al dire che stiamo paginando su disco se pensi alla ram cioè non fai un accesso diretto ma fai un accesso per indice alle pagine e ricarichi le pagine che ti servono poi l'attention la fai sulla pagina attenzione l'attention vera quella che poi ti porta a decidere il prossimo passo nel layer la fai sul contesto reale però il recupero del contesto reale invece che scorrertelo tutto indietro con la KVCash vai a prendere la pagina dove è più altamente probabile che tu trovi quelle informazioni ovviamente trattasi di altamente probabile e non di certezza quindi in questo senso puoi perderti alcune informazioni che potevano essere interessanti il paper va letto bene io non ho avuto tempo di approfondirlo fino in fondo ce l'ho letto ma una volta diciamo però più o meno è questo quello che succede dentro col paper e di sicuro verrà in qualche modo adottato raffinato sistemato anche da altri perché risparmiare il 78% della potenza di calcolo per gestire la cavi cash è fondamentale e qui se volete faccio un altro riaggancio alle interviste che abbiamo nominato in particolare quella di Assabis perché l'intervistatore per me è stata illuminante quella cosa lì sentita dire da lui sulle scelte di business più che sulla tecnologia e l'intervistatore gli chiede insomma state diventando fondamentali perché adesso tutti i prodotti google integrano Gemini in un modo o nell'altro avete visto l'hanno messo in maps è dentro tutto workspace e lui dice sì e poi ci sono tutte le AI mode gli risponde per non parlare della della sintesi AI delle ricerche che sono miliardi di chiamate al giorno o quelle robe lì perché avendolo integrato nella ricerca comunque è la cosa più usata del mondo e lui dice infatti per noi negli ultimi mesi la ricerca si è focalizzata tantissimo sulla distillation perché tutta quella roba che vedete integrata usa flash non usa pro perché flash siamo riusciti a far dare delle risposte a flash che sono di altissima qualità spendendo meno di un quindicesimo a livello di potenza di calcolo e per noi era fondamentale se no non avremmo potuto metterla dappertutto quindi è interessante quanto in realtà possiamo dire che OpenAI scusate OpenMind stia facendo cose assolutamente di frontiera ma di una frontiera diversa cioè la stanno rendendo efficiente ed efficace invece che concentrarsi sullo sviluppo ulteriore di Gemini 3.1 poi mi sbaglierò ma visto di cella io a maggio Gemini 4 lo vediamo arrivare però questo è solo una mia sensazione però è interessante quanto sia necessario da un lato avere risposte sempre migliori per arrivare lui ne parla anche per arrivare alle GI vabbè dopo ne parlo meglio da quell'intervista lì perché è super interessante però quanto sia importante anche essere efficienti e distribuire il carico il più possibile anche sull'Edge eccetera eccetera eccetera e di nuovo ingegnerizzazione della questione esatto è più ingegnerizzazione che ricerca pur loro continuando a fare un sacco a me piace anche vedere che tipo tornando a DeepSeek ogni tanto si esce con queste trovate che sono quelle che ti fanno fare lo step il gradino e poi dopo tutti gli altri in un modo nell'altro raffinano e adottano e avremo modelli di nuovo più potenti utilizzabili su risorse più o meno pari a quelle che erano richieste prime per modelli più scarsi e poi c'è un'altra cosa che credo sia interessante da dire nominavamo prima la quantizzazione fatta da San Filippo che è una Q2 da una parte e una Q8 dall'altra però il paper dice esplicitamente che il training è stato fatto FP4 quantization aware che cosa vuol dire quella cosa qua vuol dire che quando fate il training rappresentate ogni peso dei perceptroni in FP16 normalmente FP16 c'è qualcuno che l'ha fatto in FP8 ma di base si fa FP16 loro l'hanno fatto FP16 però renderle il training FP4 aware significa che nel cercare i minimi locali che poi portano il minimo generale del modello perché è sempre una una funzione di è la ricerca del minimo di una funzione alla fine trainare un modello banalizzando un po' ma ecco la roba lì dire che FP4 aware significa privilegiare i pesi che sono rappresentabili in FP4 senza troppa compressione rispetto a distribuire molto i pesi su un FP16 quindi è come dire ok io ho tutta la potenza dell'FP16 ma cerco di concentrare tutto in alto in basso al centro in una dei tagli FP4 che una quantizzazione potrebbe fare già in fase di training questo dovrebbe portare si vede già in altri modelli minimax l'ha fatto se ricordo bene tra i primi era minimax 2.5 aver fatto questa scelta questa cosa qua porta ad avere delle quantizzazioni che quando sono fatte si fa meno fatica a farle quelle di Asloth non diventano matti e e poi il risultato è migliore comunque il risultato è migliore perché stai tagliando via meno roba cioè l'informazione tutta concentrata a blocchi di FP4 in qualche modo interessante questo cresce un po' il numero di di pesi totale quindi è sempre un trade off però lo sono dati a trillion e trillion per cui un trillion e sei se ho letto bene poi l'ultima cosa forse interessante dal loro peper troppo tecnica da spiegare qua è come se fanno lo storage della kvcash per usare la sparsa attention che dicevo prima ah no e attenzione attenzione che è una cosa di cui discutevamo noi in chat ogni settimana ma che sono certo che anche qualche ascoltatore potrebbe storcere il naso perché i benchmark non sono entusiasmanti però attenzione che i benchmark non sono entusiasmanti ma è una preview questo qua non è il modello finito è una preview che vuol dire che ha fatto loro dicono ampiamente meno che non vuol dire niente ampiamente meno comunque meno del 50% della fase di reinforcement learning quindi è un modello reasoning a metà non so come dirla perché hanno fatto probabilmente l'allineamento base per dare risposte per essere instruct tutto di sicuro quello instruct vuol dire che è capace di conversare non di prevedere solo il prossimo token perché anche questo quando vado a conferenze spesso è una domanda che viene fuori ma adesso i modelli stanno a conversare no i modelli non stanno a conversare i modelli fanno ancora quello che facevano prima prevedono il prossimo token o il prossimo blocco di token reinforcement learning human feedback o verifiable reward che sono le due fasi che si fanno dopo insegnano al modello a mettere insieme delle risposte a partire dalla previsione del prossimo token questa è la prima fase di reinforcement learning poi ce n'è una seconda che è quella che gli insegna a fare la stessa cosa generando gli hidden token ovvero i reasoning token anche detti che sono quelli che gli permettono di fare il ragionamento questa parte è fatta in piccolissima parte sulla preview di DeepSeq quindi io mi aspetto che i benchmark della final siano molto meglio poi non è detto però valutami una figura retorica Stefano è un po' come la differenza che ci passa tra conoscere il vocabolario dell'inglese e sapere parlare l'inglese sì sì sì è corretto è assolutamente corretta quella cosa che dici è proprio la differenza è quella lì delle fasi RLHFR poi chiamiamolo reinforcement learning in realtà è più una distillazione ma non lasciamo stare il tecnicismo c'è la fase in cui gli insegni gli insegni a fare il reasoning che è una fase ancora diversa che stando nel tuo paragone è come non dire cazzate esatto bravo avrei detto esattamente le stesse parole e questo più o meno quello che riesco a raccontarvi così a parole del paper però il paper per chi è tecnico del settore va letto è interessante deep seek moment numero due sì ha fatto meno boom dipende sono crollate del 20% le azioni di invidia oppure no perché se non sono crollate non è deep seek moment no anche perché non dirlo che viene a prenderci è la il CEO di invidia di cui prima o poi potremmo che lui ci ascolta eh Jensen no non facciamo il nome no scusa è monitorato adesso viene a prenderci monitora con l'intelligenza artificiale tutti quelli che nominano il suo nome ne parlano male no non so se avete visto forse abbiamo già detto qualche settimana fa è andato in intervista da prima era stato da Alex Friedman nell'intervista normale serena poi è andato da Darkash che invece l'ha incalzato molto sul sull'export cinese sul fatto che i cinesi stanno recuperando terreno che l'hardware non è più la cosa più importante perché Google sta dimostrando che con con i suoi chip sta andando forte eccetera cazzo si è incazzato cioè ha proprio ha cambiato la voce non aveva più la solita voce da da keynote come si definisce ha cominciato ad alzare la voce a dirgli ma che fichiasi poi risultato invece è pieno di meme su su x di jensen che che dà il microfono in testa ad Darkash e queste cose qua quella intervista lì va vista solo per quello perché lui che si incazza è credo notizia unica perché sempre tutto compunto da bravo CEO e poi anche culturale visto che è asiatico dunque scaletta che ci stiamo già perdendo via dai vision banana vision banana mi piace vision banana allora vision banana faccio vedere qualcosa io ecco tu fai vedere qualcosa da dove è partita la questione sostanzialmente DeepSeek se è uscita con un paper in cui scusa DeepMind DeepMind se è uscita con un paper dicendo portando all'attenzione il fatto che i modelli di generazione di immagini di frontiera come banana sostanzialmente possono essere utilizzati anche come modelli per fare vision modelli diciamo specifici per fare task di vision tipo segmentazione tipo depth estimation dimension tipo edge detection queste cose qua no l'idea non è nuova cioè già se ne era parlato però il discorso è che loro sono in una posizione in cui hanno un modello particolarmente forte e possono dimostrare quanto meno poter fuori dei claim perché poi non è che il modello è open e lo possiamo provare tutti per fare queste cose che i risultati che si ottengono sono sostanzialmente a livello dei modelli diciamo frontier di segmentazione depth estimation edge detection eccetera e che cosa significa questa cosa un attimo spiegata un attimo in modo più facile se pensiamo cosa fanno i modelli di generazione di immagini attuali noi sostanzialmente partiamo da un'immagine forniamo un prompt oppure partiamo proprio con un prompt testuale generiamo un'immagine però ci sono i modelli di editing ai quali diamo un'immagine diamo un prompt e otteniamo un'altra immagine e allora intuitivamente il concetto è che nel momento in cui questo prompt esprime il task di vision che noi avremmo voluto fare con un modello dedicato un modello specifico possiamo provare ad ottenere la stessa cosa per cui un task di segmentazione che è tipo dimmi trovami le auto in questa immagine può utilizzato su un modello di generazione di immagini portarti ad avere una nuova immagine in cui per dire le auto sono colorate tutto di rosso questo che sto facendo vedere secondo me è abbastanza impressionante due zuppa con di fianco qualcosa da mangiare due prompt uno il primo che isola l'aglio il secondo che isola la carne cioè questo è abbastanza impressionante fatto da un modello che non è dedicato se hanno imparato a riconoscere le auto posso contare sul fatto che google smetterà di rompermi i coglioni con i suoi captcha di chiedermi quale è il semaforo e quale è l'automobile sostanzialmente stiamo passando da dei modelli di diciamo di vision che ti generano in output sostanzialmente dei delle matrici di numeri che ti dicono la posizione nelle immagini in cui tu hai la cosa che viene diciamo detected a dei modelli che sono questi modelli di generazione di immagini che sostanzialmente producono un'altra immagine in cui l'informazione se vuoi è nei canali del colore e quindi tu puoi esprimere in quel modo lì le stesse cose che avresti espresso con la matrice dei numeri per cui banalmente se stiamo parlando di edge detection tu hai un'immagine in bianco e nero dove in nero hai i bordi della cosa degli edge se parliamo di depth estimation quindi normalmente avreste avuto una matrice di numeri che ti dice la distanza per ogni pixel la distanza dalla camera hai un'immagine in cui hai scale di colori differenti e quindi i pixel più chiari sono quelli magari più vicini e quelli più scuri sono quelli più lontani o viceversa al tempo stesso potresti avere delle heat map per la salienza per vedere per dire cos'è una parte saliente di un'immagine è interessante secondo me da due punti di vista il primo è che in realtà molte applicazioni di modelli di vision poi per essere veicolati cioè per essere utilizzati prevedono che tu le informazioni che ti hanno dato i modelli di vision le trasformi a loro volta in un qualcosa di in un'immagine visuale che sia interpretabile dall'umano quindi utilizzare questo direttamente il modello di generazione di immagini per fare questa cosa se vuoi ti togli uno step ti fa già uno step che tu avresti fatto scusami mi hai fatto venire in mente quando vai a vedere l'ecografia di tuoi figli e ti dicono vede qui si vede il bambino e tu vedi tipo la galassia di Andromeda e dici sì sì lo vedo lo vedo allora io l'ho trovata interessante sai per un'altra cosa anche perché fa un po' il paio con una cosa che Asabis continua a ripetere non nell'intervista che citavo prima ma l'ho sentito in un altro intervento quando dice che i modelli di generazione in video invece possono essere un buono strumento usati al contrario cioè dandogli un video in pasto per interpretare la fisica del video perché sono capaci di generare qualcosa che sia una fisica credibile se non esatta ma comunque credibile nel video lui in più interventi dice la potenzialità di questi oggetti sta molto anche nell'uso che potremmo farne come vision vision 3D o vision continua lui parla di robotica in quella situazione cioè come faccio a far vedere la roba e a capire che cosa ha intorno il reverse engineering della generazione video così come qua è stato fatto il reverse engineering della generazione immagini quindi è in qualche modo visto che comunque generazione immagini e generazione video pur avendo tecniche leggermente diverse eccetera ma sono una stessa famiglia perché per fare un video fa i frame è in qualche modo una conferma di questa sua teoria e poi c'è un discorso di versatilità nel senso che nel momento in cui un modello di generazione di immagini opportunamente prontato con tipo data questa immagine ridisegnamela come mappa di profondità ridisegnamela colorando gli oggetti di un certo tipo in un certo modo isolando solo un certo tipo di oggetti eccetera tu hai un modello che è universale che può essere usato per segmentazione per analisi di profondità per edge detection per object grounding per insomma tutte queste cose senza doverti preoccupare di averne uno per ogni tipo chiaramente c'è un discorso di risorse necessarie per far andare questa cosa però si torna se vuoi all'argomento di prima ci sono i margini di ottimizzazione eccetera guarda io trovo super interessante quello che ci hai raccontato e mi stupisce perché in passato non l'avrei trovato così interessante non per offendere il tuo campo di interesse ma relativamente ha poco a che fare con la mia quotidianità ma mi permette di citare il fatto che nell'ultimo mese io e Stefano a parte quando lui è andato in vacanza divertirsi e mi ha lacciato col cerino acceso ci siamo dedicati a una sfida di robotica un hackathon di robotica in cui questi aspetti che tu hai descritto sono assolutamente chiave e importanti a partire dai più banali e vecchi a questo punto modelli per fare l'algoritmo YOLO che identifica gli oggetti ti fa un box intorno che adesso sono commodities cioè hai una chiamata hai il risultato alcuni dei nostri partner del contest che sono dei veri ingegneri nell'ambito della robotica cosa che noi non siamo hanno invece mostrato come applicare alcuni degli algoritmi più avanzati all'identificazione degli spigoli quindi c'era questa macchinina telecomandata immaginate che andava in giro con una camera e riconosceva non solo la sedia ma riconosceva gli spigoli quindi ne teneva traccia quando ci si avvicinava per non andarci a sbattere quindi un sacco di applicazione pratica quando vuoi muoverti nel mondo reale che ti servono ti serve avere insomma un aiuto di qualche tipo e quindi negli screenshot che ha fatto vedere Stefano mentre tu raccontavi questa cosa ci rivedevo tantissime applicazioni pratiche del mio mese di esperienza con un robot che andava a sbattere la prima curva nel mio caso sbattere la prima curva ma abbiamo vinto il contest quindi siamo comunque stati bravi abbiamo vinto il contest perché abbiamo contribuito perché abbiamo l'open source nel sangue la verità è il motivo per cui abbiamo vinto il contest è che siamo stati estremamente open educativi per le prossime generazioni di gente che farà il contest anche troppo però il dottore mi ha dato una pastiglia da prendere tutti i giorni comunque Paolo mi hai fatto venire in mente una cosa qualche decennio fa all'università al corso di robotica mi era infilato in un in un progetto infilato e poi dopo me ne sono anche tolto con estrema velocità in cui si si cercava di fare vision per mandare in giro un robot che aveva un'unica telecamera puntata verso l'alto contro uno specchio così che potesse vedere tutto quello che c'era attorno immaginate un'immagine che ti fa vedere tutto quello che hai attorno ottenuta con questa telecamera che punta verso l'alto dove hai uno specchio e chiaramente i pochi algoritmi che esistevano di edge detection eccetera non ti davano granché perché era una trasformazione assolutamente strana e lo specchio non era esattamente con una forma normale diciamo per cui distorceva in parti nelle varie parti dell'immagine in modo differente e probabilmente con diciamo i modelli ciò che esiste adesso sarebbe stata una cosa banale avere un qualche tipo di come dire di tool che è generico abbastanza per cui può fare vari tipi di diciamo di lavorazioni vision contemporaneamente sarebbe stato stupendo però tornando un secondo al nostro di progetto di robotica ne approfitterei per ricordare al nostro gentile pubblico che abbiamo anche intervistato sull'argomento Simone di Somma a gennaio se volete andare a vedervi l'intervista è molto interessante per capire anche che cosa fa Cyberwave eccetera i maligni staranno pensando che abbiamo vinto il contest perché avevamo intervistato Simone però non è così non è vero abbiamo vinto perché c'erano solo altri tre team nella nostra categoria altri due team e quindi avevamo 33% sono ritirati quindi vabbè adesso non stiamo a dire tutti gli altarini Paolo perché poi non è questo il punto voi siete stati sufficientemente bravi da resistere fino in fondo e portare la casa al risultato finale assolutamente sì e parlando di interviste è uscita mercoledì l'intervista a Stefano Gatti che era stato già nostro ospite vi ricordate è stato il primo a rompere il ghiaccio nelle puntate con ospiti nella puntata 20 l'ho intervistato invece per lasciargli lo spazio che merita vista la sua esperienza e la sua vision andate a sentirvi l'intervista e soprattutto andate a sentirvela perché ho ascoltato i vostri feedback perché finalmente qualcuno ci dà dei feedback non tutti gentilissimi ma ci danno dei feedback dicendo che i video sembravano fatti da un bambino di 4 anni e ho provato ad arrivare almeno 6 anni se qualcuno vuole dargli un'occhiata mi sono impegnato mi impegnerò anche nel montare questo episodio però diciamo che ci stiamo provando ad ascoltare i vostri feedback abbiamo anche ingaggiato degli attori per fare gli screen capture delle copertine giusto? no quella è intelligenza artificiale e presto anche qua in video metteremo i nostri avatar in modo che siano belli con la faccia tirata esattamente come quella della tua copertina Paolo che ha avuto un certo seguito c'è da dire ha avuto un grandissimo successo quindi dovresti dovresti anche valutare ti fermano per strada ti fermano per strada attenzione spero non per darti delle sberle ma è importante che ti fermino per strada bene bene bene dunque dicevamo così nell'introduzione al volo di investimenti che girano li nominiamo soltanto per per dirlo 40 miliardi di dollari da Google verso Antropic interessante per come Google diversifichi i suoi investimenti non soltanto interni ma anche verso una che potrebbe sembrare una concorrente e anche in Cina DeepSeek pare che stia raccogliendo 20 billion da Alibaba e anche lì movimenti interni perché anche Alibaba vi ricordiamo che fanno QEN come modelli e quindi sono in qualche modo concorrenti così questa quasi una curiosità rispetto ai nostri argomenti diciamo due parole sul fatto che chiude Rucod si diciamo due parole su quella roba lì che era il mio prossimo punto in scaletta prima di parlare di Armete allora Rucode chi non si ricordasse che cos'è Rucod era un fork di Klein Code uno dei primi agenti di coding che io e Alessio abbiamo usato io e Alessio perché Paolo ci snobbava al tempo sugli agenti di coding salvo dopo entrare completamente nel loop e non uscirne più ti spiego la storia ero taccagno stavo cercando un accesso gratis e non lo trovavo e poi quando Google ha dato Gemini gratis l'ho provato ah ok ok Klein era è ma anzi Klein è perché non è Klein che chiude un un agente scritto come plugin di VS Code Rucod era un fork di Klein per fare le cose meglio e per i taccagni anche nel senso che tendeva ad usare un po' meno token perché Klein era abbastanza un'idrovora Rucod ha avuto molto successo comunque perché è stato molto open ha accettato molto i feedback e molte di chi ha usato plugin di VS Code diretti per avere la gente che non voleva usare diciamo Cursor ha usato Rucode nell'ambiente open l'autore ha dichiarato basta grazie arrivederci chiudo tutto se volete forcatelo se non volete forcarlo fate quello che volete io questo non lo mantengo più ma non l'ha detto perché si è scocciato l'ha detto con una giustificazione forte nel suo blog che è quella che mi ha colpito dice chiudiamo chiudo Rucode perché non sono convinto più che sia il tempo per gli ID per gli IDE detto l'italiana cioè gli editor di codice e Rucode era pensato per girare soltanto dentro VS Code è il tempo degli agenti questa cosa qua non ha più senso di esistere sto solo perdendo tempo energie voglio fare altro che se vuoi si abbina un po' a quello che abbiamo detto ormai diverse settimane fa il diffondersi sempre più delle CLI diciamo degli agenti che scrivono il codice eccetera no infatti non sembrava una scusa che in qualche modo ha rafforzato la mia nostra convenzione che abbiamo fatto bene a fare lince.sh andate tutti sul sito lince.sh installate questo software fatto da noi dal gruppo di ressorti artificiali però adesso fuori di battuta su una cosa per niente pubblicizzata se non così noi nel podcast in Italia così abbiamo un po' di contributo attivi gente che ci sta mandando pull request su questa cosa e uno è Paolo ma non è solo Paolo che ha mandato le pull request ci sono un altro paio di persone francamente non è che me lo aspettassi più di tanto nel senso che io l'ho fatto per me stesso prima di tutto e lo stiamo usando noi Paolo ha provato ha contribuito ha messo a posto le cose che gli servivano però evidentemente c'è un interesse anche da altri ho un po' di commenti in giro anche per cui se vi va di provarlo provatelo fate dateci feedback soprattutto mandate pull request io te lo vendo in una maniera diversa se cercate qualcosa che volete lasciare andare avanti tutta la notte e trovare la mattina dopo del lavoro fatto senza preoccuparvi troppo che abbia lanciato i missili nucleari verso la Cina l'ince vi offre le funzionalità per avere diciamo un po' di sicurezza io l'ho letteralmente usato per dormire bene mentre questa cosa gira sì l'ho letteralmente usato queste ultime due noti per quei progetti che non erano tier 1 né tier 2 di priorità ma tier 3 che non avrei mai guardato gli ho detto senti ma che me ne frega fallo stanotte domani mattina vediamo cosa trovo e per ora ha funzionato bene ovviamente a seconda di quello che fate poi dovete guardare proprio a livello di workflow a livello di cosa fa il codice alcune cose volete controllare come sono andate ma se avete spezzettato il lavoro per bene con tecniche di come si chiama Spectre Even piuttosto che Backlog che è il mio go to workflow funziona molto bene e a breve lo promuoverò ad altri se riesco a far convincere Stefano a cambiare alcune cose che odio ma sì cosa odi ma mi convinci di sicuro te lo dico dopo dai adesso facciamo un po' di maretta non siamo scriptati non siamo tutti sempre amici poi intanto ci sputiamo quelle cose così sì sì sì quello sicuramente per l'uso che dice Paolo tra l'altro basta la componente sandbox non vi serve tutta la dashboard che la dashboard esistono due componenti in link una è la dashboard che vi serve per lavorare con agenti multipli nella stessa finestra e una è la sandbox che invece vi permette di farlo girare anche di notte senza preoccupare ne troppo tant'è che è un caso d'uso interessante tant'è che ieri sera non ho ancora finito motivo per cui non c'è ancora pull request stavo integrando come uno degli agenti nativi PE PE che è un piccolo arnese alternativo open code diciamo open source minimale e lo integravo più che altro perché su PE è stato fatto uno sviluppo da quelli di Shopify che si chiama PE auto research per usare la modalità auto research di Carpati che ha proprio la caratteristica di girare di notte eccetera e volevo provarla dentro alla mia sandbox o comunque dentro l'ince ma nel frattempo però mi sono distratto ultimo argomento che si collega e ci siamo ho installato Hermes Agent Hermes Agent per chi non se lo ricorda è un'alternativa a open cloud perché ho installato quello e non open cloud perché mi piace fare il diverso sempre Linux invece di Windows o di Mac eccetera e no sostanzialmente perché c'è più attenzione alla sicurezza oltre al fatto che non è di open AI ma è indipendente è un progetto open source fatto da gente che viene dalla blockchain quindi con un'attenzione forte alla sicurezza e l'ho messo su un vecchio computer che avevo ho valutato anche di metterlo su un raspberry perché ho letto di gente che lo fa però meglio P5 con 8 giga io non ce l'avevo libero ce l'ho ma sta facendo altro solo un P4 e quindi avevo lì un computer che non faceva niente ho detto sarà meglio questo comunque che non è ARM mi sono installato un Ubuntu nuda così non dovevo sandboxare niente c'è solo lui dentro lì e gli sto facendo fare un po' di cose con soddisfazione tipo guardare come va il podcast invece che guardarlo io in maniera compulsiva sull'interfaccia di youtube ogni qualche ora mi dice come stanno andando le cose che cosa sta andando bene sta andando male gli AB test quelle cose qua poi gli ho configurato la mail attenzione panico paura che serpeggia tra tutti pensando a quella di meta che aveva cancellato tutta l'inbox che non era vero le ribadisco è una panzana quella lì perché io non ci credo assolutamente non tanto che lei non l'abbia fatto ma che lei responsabile della sicurezza di meta l'abbia fatto l'abbia pubblicato su tutti i social e Zuckerberg non l'abbia licenziata è l'ultima parte che non mi torna cioè l'avrebbe licenziata sicuramente se fosse stato vero invece era una cosa concordata per far parlare male di OpenGLO no ho configurato la mail però appunto è attento alla sicurezza ed è sveglio io uso GLM 5.1 come modello e gli ho detto no no no perché sei attento alla sicurezza i modelli cinesi sono i più sicuri i modelli cinesi sono i più sicuri assolutamente no gli ho detto ta buono lì va bene la mail però non voglio che tu ne cancelli ne spedisci ne fai queste cose qua oh eh sì me lo segno no me lo segno aspetta si segnatelo pure però non è abbastanza me lo segno va bene però aspetta un attimo che script usi per fare queste cose con la mail eh sì uso quello lì di script eccetera no va bene prendiamo lo script lo modifichiamo gli zappiamo via dal tutto send e gli zappiamo via dal tutto delete cioè non ce l'hai più quelle funzioni lì eh sì va bene e dopo lui è stato bravo perché mi ha detto sì però guarda che se io volessi mi tiro su un cheerle e poi cerco di cancellarti la mail lo stesso e gli ha detto sì ma io non ti ho dato l'account eccetera sì alla fine hai ragione ci siamo messi d'accordo e abbiamo e abbiamo questa versione ha promesso che non fa cose cattive abbiamo questa versione moncata dagli script che è una cosa che in generale consiglio eh cioè attenzione ad esempio adesso gli ho creato eh token su github per creare le issue guardare le issue eccetera eccetera ma ho fatto un fine grained token in cui gli ho dato solo i permessi precisi delle cose che volevo che facesse cioè non può fare delete non può pushare su main non può fare queste cose qua eh quindi se volete è uno strumento potente molto comodo cioè voi ci accedete io adesso con telegram perché è la cosa più comoda che ho perché ce lo sia sul pc que sul telefono eh ma lui fa cose schedulate fa cose ricorrenti diventa proattivo sta gestendo il mio calendario per esempio ad esempio c'era una cosa da fare sulle cose di github che lui non poteva perché non hai permessi e mi ha detto guarda che io queste cose non le posso fare eh ti ho creato il piano delle cose che vuoi fare eh come faccio a ricordarti gli ho detto cercami un buco sul calendario e lui mi ha detto eh guarda avresti un buchetto di mezz'ora alle dodici ti faccio un evento sì fammi un evento eh mi ha creato l'evento con tutta la descrizione di quello che devo fare eh è uno strumento oggettivamente potente un po' additiv come potete immaginare eh però è potente un'altra cosa che per cui lo sto usando è monitorare i paper su Arxiv e crearmi l'inbox che poi diventerà uno dei miei wiki carpati wiki quindi mi sta risparmiando quella cosa che facevo quasi a mano c'è un paio di script ma facevo quasi a mano di vedere che paper sono usciti vedere quali erano interessanti rispetto ai miei interessi lui invece fa un bel lavoro mi fa un pre-filtering per interessi non buttando via niente lasciandomi da un'altra parte tutti quelli che non è riuscito a classificare per mia eventuale revisione e poi mi mette questi link dentro ai inbox dei vari carpati wiki degli argomenti che seguo e che voglio collegare tra loro cioè fa tutta una serie di cose ricorrenti ma ma ma avendo il vantaggio di avere un LLM dietro quindi le mail non me le filtra solo per tag ma me le filtra per tag poi va a leggerle capisce se una di queste mail ha bisogno della mia attenzione e me lo dice o se è soltanto pubblicità o soltanto una newsletter eccetera me le mette in un altro elenco mi dice guardai questa newsletter se la vuoi leggere più o meno questi sono gli argomenti e invece se c'è uno che mi ha scritto qualcosa che vuole una risposta dice questo qua vorrebbe una risposta e io gli ho detto che non può mandare niente quindi non manda niente ma gli ho detto che può crearmi draft e quindi mi creo un draft della risposta che io poi mi trovo in gmail vado là e a partire dal draft poi la scrivo io alla fine perché magari manca un pezzo così ma almeno già il contesto già le cose pronte è in assoluto allora io avevo già provato open club vi ricordate quando era molto immaturo open club sicuramente sarà a questo livello o meglio anche per certi versi magari per altri peggio tipo la sicurezza perché qua sono veramente non perché open club non sia sicuro adesso è sicuro anche open club alla fine però qua sono molto attenti cioè sono quasi esagerati su certe cose stirono su un docker container ogni volta che devono devono far girare uno script per evitare di fare danni è la cosa più futuristica futuribile da un punto di vista dell'utente non tecnico finale che ho visto perché pensare che hai una roba che ti gestisce l'agenda che gli dici trovami un buco e te lo trova secondo me per l'utente non me che magari lo facevo già con con cloud code con l'MCP giusto o con la CLI giusta non sostituisce gli strumenti di coding anche se lui può ti chiede se vuoi installarti cloud code o codex perché può spannare su cloud code o codex quando gli dai compiti di coding però io quella roba lì al momento non lo sto facendo sto valutando di fargli fare le code review quello sì perché accorgendosi di tutte le PR che arrivano è capace ha una skill per fare le code review e mettere i commenti quello sto pensando di farlo almeno come esperimento su lince questo è il mio racconto eh lo so l'ho raccontata apposta perché ho detto così Paolo sa cosa fare a lui come se non avessi avuto nient'altro grazie spero aver fatto venire voglia anche a qualcun altro e spero che lasciano tanti buchi aperti che io posso sfruttare sto scherzando ovviamente va bene va bene bene senti fagli fare una cosa fagli ci stavo pensando prima fagli scaricare i nostri episodi prendere il transcript e aggiungere i link che noi non aggiungiamo mai devi fare quello in realtà la skill nuova con cui pubblico gli episodi lo fa già poi poi i link non li mettiamo dovremmo metterli però i link li tira fuori me li propone solo che mi dice di metterli in un commento poi dopo mi dimentico eccetera eccetera e vabbè purtroppo c'è la la mia umanità bene salutiamo tutti mettete stelline campanelline che Claude mi conferma che è giusto che lo dica alla fine per non sembrare piacione cercate di trovare l'account di mail che sta usando Hermes di Stefano e spammateglielo se trovate spammatemelo purtroppo lo trovano però non fatelo gentilmente perché perché otterrete una reazione uguale contraria di enorme misura in perfetto stile Trump va bene ciao ciao ciao ciao ciao Grazie.