Show full episode description
Hey folks, Alex here, let me catch you up! I’ve had a feeling that this week is going to be crazy, as it started on the weekend MiniMax M3 , then with Jensen announcing new RTX Spark, NVIDIA’s first PC chip packing 1 petaflop of local AI power into thin laptops. A few days later at Microsoft BUILD, Satya & Mustafa from MAI dropped 7 AI models, completely pre-trained from scratch, including a new MAI-thinking-1 , MAI-code and MAI-image 2.5 that started topping the image gen charts. Then other image models started racing to the top of the Arena benchmarks, IdeoGram 4 hitting becoming SOTA open weights image-gen model, and Reve 2 beating Nano Banana just a few hours after that. And then today, NVIDIA dropped Nemotron 3 Ultra , their latest 550B open weights model, data and training and Arena published a new agentic eval leaderboard and we got a new Gemma 4 12B. I’ve had the great pleasure to host Chris (@llm_wizard) from Nvidia, Peter Gostev from Arena and Karan from Nous Research (who were featured prominently by Jensen!) all on the show. Def don’t miss this one! Let’s get into the details. ThursdAI - Join the flock of folks who know what is happening in AI before everyone else. Open Source LLMs 🔥 NVIDIA Nemotron 3 Ultra: The 550B Open Source Beast Built for Agents ( X , Arxiv , Announcement ) This was the big one. Breaking news mid-show: NVIDIA drops Nemotron 3 Ultra , a 550 billion parameter sparse MoE model with 55 billion active parameters, built on a hybrid Mamba-Transformer architecture. Chris Alexiuk, AKA Joe Nemotron, joined us live from NVIDIA HQ in Santa Clara to walk us through it. The headline number is 5.9x higher inference throughput compared to GLM-5.1 on decode-heavy workloads. Chris told us that this is a result of multiple things, their Hybrid Mamba-Transformer approach, the sparse attention, and that they optimized for decode-heavy workloads (the kinds of workloads agents do) The architecture is fascinating. They’re mixing Mamba-2 state space layers with sparse attention, which means step 300 in an agent loop runs as fast as step 3. Pure transformers can’t do that because the attention cost keeps growing with context length. This kicks in big time at 64K+ sequence lengths, which is exactly where you end up in real agentic work when the model is having multi-turn conversations and people are dumping their entire codebase in. P.S - We launched Nemotron 3 Ultra with 0-day support on CoreWeave Inference, it’s super fast and pretty cheap, give it a try here They pretrained on 20 trillion tokens, extended context to 1 million tokens, and their post-training pipeline used multi-teacher on-policy distillation from over 10 specialized teacher models covering everything from SWE to terminal use to search to office work, which they are also going to open source soon! One thing Chris emphasized that I really appreciate: NVIDIA doesn’t have their own harness. There’s no “NVIDIA Code.” Which means they actively resist the temptation to harness-max, to optimize for just one harness and look good on a specific leaderboard. Ultra should be a solid drop-in for whatever harness you’re used to, and that generality is worth a lot. It’s not the best thinker, but it is the highest score US based open weights model, so again, a huge huge win for the US AI ecosystem! The Nemotron 3 Ultra release is open under the OpenMDW-1.1 license : base BF16 , post-trained BF16 , and NVFP4 quantized checkpoints, plus the GenRM , synthetic pre-training data for code , legal , and <a target="_blank" href="https://hu
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Tracking a stacked week of frontier AI news across open models, image generation, agents, and voice.
Benefits
- High-signal roundup of the week's biggest AI releases
- Fully open American model catching up on open source
- Open-weight image models rivaling closed competitors
- Agents moving natively to the desktop
- Best-in-class open speech-to-text and TTS
Use cases
- NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-transformer, fully open, 5x faster inference for long-running agents
- Ideogram 4.0 open-weighted 9.3B text-to-image model beating Microsoft MAI, Flux Dev and Qwen Image
- Microsoft MAI image 2.5 hit #3 text-to-image and #2 image-to-image on arena leaderboards
- ElevenLabs Dubbing v2 preserves emotion across 90 languages; Cartesia Ink 2 ranked #1 most accurate streaming speech-to-text
- Nous Research Hermes Desktop launched on Mac/Windows/Linux, called out by Jensen on Computex stage
KPIs / results
- Nemotron 3 Ultra: 550 billion parameters, 5x faster inference
- Ideogram 4.0: 9.3 billion parameter open text-to-image model
- Gemma 4: 12B encoder-free multimodal model, Apache 2 license
- ElevenLabs Dubbing v2 covers 90 languages
Tools / build
- NVIDIA Nemotron 3 Ultra
- Ideogram 4.0
- Microsoft MAI image 2.5 / MAI thinking 1 / MAI code 1 flash
- Nous Research Hermes Desktop
- ElevenLabs Dubbing v2 / Cartesia Ink 2 / Nemotron 3.5 ASR
What's going on everyone welcome to Thursday this is Alex Volkov I'm coming to you live on June 4th on the highest signal AI news show that you can ever subscribe to we've been at this for over three and a half years and we have been training we have been training for weeks such as these because this week was absolutely stacked with the AI news from Jensen in the beginning of the week and minimax released a new version and the video has just literally dropped in breaking news Nemotron 3 Ultra and we're gonna have Chris Alexiuk friend of the pod to talk about this and there's four new image models three new image models one of them is open source we're gonna compare between all of them and Microsoft came out of nowhere declared to be a frontier lab there is literally just a torrent of AI news not to mention that supposedly at some point today maybe there's gonna be like a drop from open AI we'll see we're not in the speculation game but as always if you are tuning into Thursday I you know that we're gonna try our best to cover all of this in just under two hours we have two incredible guests today we have Chris Alexiuk from NVIDIA aka Joe Nemotron the guy who writes the blog post our friend of the pod Chris is gonna join us in about an hour to help me cover all this news we have our friendly co-host here Jan Peleg and Wolfram Reifon Wolf welcome guys LDJ as well all right this is it this is a TLDR this section of the CI where we just run through everything that we have on cue for you for today Thursday June 4th our first show in June with you Alex Volkov AI Evangelist with Weights and Biases and CoreWeave with me co-host today Wolfram Ravenwolf Jan Peleg and LDJ we have two great guests today Chris Alexiuk from NVIDIA aka Joe Nemotron by the way a friend of the pod that has been on the pod multiple times to talk about NVIDIA 3 Ultra which also is breaking news just dropped today another guest is Karan co-founder of News Research to talk to us about Hermes Agent Hermes Desktop and the incredible success that News Research has been seen lately with Jensen calling out News Research on state Computex you guys see this this is crazy shout out huge shout out to News Research our friends from way back we talked with News Research way before they were even a company I just like I love bringing Karan back to chat about their latest advancements in the big companies I tested their models when Hermes referred to the models and not the agent now it's changed a little bit there's still a research lab they're still racing cool shit all right in big companies in the lamps I think that maybe we'll have some news about Joule Alpha a potential GPT update but we'll see meanwhile NVIDIA decided they changed the PC game forever NVIDIA RTX Spark was part of the announcements at NVIDIA's Computex they're claiming that this is their first PC chip ever they have an RTX 5070 class GPU 128 GB memory and one petaflop of AI compute onto thin laptops you guys see this is insane they're able to run models big models this is Jensen obviously showing this up they're able to run models on these laptops and it's quite crazy trying to compete with MacBooks we're going to chat about this a little bit in addition Microsoft also announced laptops together with these RTX laptops but Microsoft had their own event this week called Microsoft Build they showcased a bunch of AI stuff they switched to agentic almost completely very similar to Google IO Microsoft decided to not only come out with like best models they also decided to say hey we're now a frontier adjacent lab not quite a frontier lab if you guys remember Microsoft obviously was one of the first investors in OpenAI Microsoft came out with seven models from MAI MAI is Microsoft AI's organization that is the result of the acquire of inflection AI if you guys remember we talked about inflection a long time ago with Mustafa Suleiman and Karen Simonian like a bunch of other great folks that built a near frontier lab and now they trained models from scratch those are not fine tunes of kind of open source Chinese models they released seven models across thinking and code image transcription and voice from zero distillation zero like all of the data is sourced and we're going to talk about them specifically the image models but also the thinking models obviously here are the models MAI image 2.5 MAI image 2.5 flash transcribe MAI thinking one is their new LLM and MAI code one flash is the new LLM that's trained specifically for VS code definitely going to chat about that I think that this is it in the big companies in news the only other thing that I can tell you is kind of because Space XAI is now a big company it's SpaceX owns XAI that owns X so it's Space XAI and they're officially fired for IPO and the prices there are insane so we're gonna mention this a little bit and TBD on dual alpha maybe OpenAI is gonna drop something maybe not right folks AI art and diffusion is the second order layer of important things that I have to bring to you because this week was just absolutely insane obviously I just talked to you about Microsoft MAI image 2.5 they hit number three on text to image and number two on image to image arena leaderboards very briefly shortly after that yesterday ideogram dropped 4.0 we talked about ideogram multiple times they have opened the weight so shout out to ideogram let me have this for a shout out thing ideogram dropped the weight for their best text to image model 9.3 billion parameters it's a huge one and they claim inside text rendering and layout control I've tested all these models I can't wait to show you my tests but this beat Microsoft MAI on the benchmarks and they are definitely the number one open model right now they're beating flux dev and quen image and who knew on all these shout out to them for the open sourcing of the model and we also had rev a 2.0 I told you about rev a folks maybe remember question control the co-founder used to be in I think stability rev a is launching number two text to image model this one is a layout based model it's a very interesting approach you see the layout shaping as the model converges I would love to show you some generations from rev it and their editor specifically I have compared all of these models I can't wait to show you for the thumbnails for this week so I will show you that so a lot of news in the AI art and like image diffusion world although some of these are no longer diffusions and then we go to open source open source has been on fire this week we're going to go to breaking news because this literally just happened AI breaking news coming at you only on Thursday I all righty the biggest breaking news from today from the world of open source is NVIDIA drops Nematron 3 Ultra we talked about Nematron 3 before Nematron 3 Ultra is a 550 billion parameter beast it's a hybrid Mamba transformer it's open model fully open including the data and it's built for long running agents and for some reason it has 5x faster inference we're going to talk with Chris Alexiuk from NVIDIA about this release but it's a very interesting release especially in the world of open source NVIDIA is not playing around folks NVIDIA is not playing around and it's great great great great great great to see that American fully American based open source is catching up I am saving our comments about this model for when Chris comes here in an hour I will say though another breaking news is that we have this model up on CW inference so shout out to the folks who work really really hard core with inference based on web devices is now serving Nematron 4 Nematron 3 Ultra in full folks we also have a new Google Gemma do you guys see this we have a Gemma 4 12 billion parameter a little bit bigger they call it encoder free multimodal model runs on your laptop with 60 gigabytes VRAM and with full Apache 2 license which is great to see brains dropped a small model also on CW inference it's a 12 billion parameter model with 2.5 billion active parameters trains from scratch trains from scratch by a team of seven people from JetBrains shout out to the JetBrains folks they have some comparisons here I don't know if we have quite enough time to go deep into melum but melum 2 from JetBrains that's good to see and I think a big one it's not fully open source yet but definitely a big one Wolfram you can confirm I see you nodding minimax m3 drops the first soon open weights model combining fronting coding 1 million sparse attention context and native multi modality which is huge this is huge news so shout out to minimax for their releases and not yet open weights but I heard that they're coming out next week this is the TLDR this is everything that happened out of this we're going to choose the most important things to kind of dive deep on so please stay with us as I shoot through this super quick so we can get the actual show some folks prefer this format some folks prefer just the highlights so they can go and explore on their own let's talk about tools and agentic engineering cognition rebrands windsurf you guys remember windsurf windsurf is no longer windsurf is dead they rebranded windsurf to dev in desktop it's a multi-agent command center with acp support acp stands for agent communication protocol I believe Wolfram please correct me if I'm wrong but basically it's the way where you ask one agent to control another agent so you can control codecs and cloud code etc isn't that the one that has two different meanings yeah but I specifically mean the the one that you can control other agents yes but agent agent client protocol is the right nomenclature for this and windsurf is now dev in desktop no bye bye windsurf we've missed you it's been a year since your last rebrand and this is finally the final one and our friends from news research launched hermese desktop it's your hermes agent goes native on the mac windows and linux in public preview so you can just do hermes desktop as a command you have this beautiful desktop experience we're gonna hopefully have screenshots of it because I cannot run this on my work laptop but we're gonna have Karan a co-founder of news research talk about this latent success exploding success for for hermes agent with showcases on nvidia computex stage and tons of exposure as well all righty this week's was a corner where we talk about weights and biases and core weave we have a few things for you first of all you can join us for this weekend at weave hacks 4 this is the coming week we have few spots left if you're in San Francisco or want to travel in San Francisco I'm actually flying out tomorrow please join us at lou.ma slash weave hacks we have open ai for the first time as a sponsor for weave hacks we also have cursor we have a bunch of credits for you a bunch of food a bunch of great judges that I reached out personally so shout out to some of the judges last but not least last corner focus almost there almost starting almost starting let's add an instant here as well voice and audio so this is actually launched last week but it was after the show 11 labs dubbing v2 is an audio to audio model that preserves emotion and performance across 90 languages it is uncannily good it is crazy good it is just just trust me on this my head is not getting as blown anymore because we follow all news all the time and kind of the LLMs are incremental updates this broke my brain a little bit this is so crazy good that you must stay here until the end of the show to hear us play with this really it's something yeah you weren't here when this happened but when you hear an instant speak in Hebrew and you're like oh shit this is how nixon would sound if he actually spoke Hebrew you will blow your gasket it's that crazy so definitely the dubbing is incredible we also have Cartesia Inc too shout out to Cartesia friends of the pod they've been here multiple times with a bunch of their models they're leading their models because their models are hybrids or mamba transformer based and Inc 2 debuts as number one most accurate streaming speech-to-text model on artificial analysis number one most accurate speech-to-text model it is really good i tested it there last week as well you can see this in world error rate word error rate this is the lowest word error rate that we've ever seen on the show and breaking news from today i did not actually know that this is going to drop nvidia also drops nematron 3.5 asr it's an open source multilingual streaming speech-to-text model it's faster and cheaper than anything else on the market so this is news from last week and then nvidia dropped faster and cheaper in open source and our friends from daily quindle kramer from daily and pipe get tested this out so in addition to just dropping one of the best lms in the world nvidia also dropped one of the best asr's and text-to-speech in the world as well and so we're going to definitely test out all of those folks with how do you do 20 minutes let's go to open source i think because we have a bunch of stuff to talk about but this is the deal there all right let's do open source open source ai let's get it started all righty folks welcome to the open source corner we love open source on thursday and so today we're going to talk about actually two segments about open source chris alexic is going to join us in i think 45 minutes from nvidia so definitely we're going to wait for that and then cover the whole nvidia nematron 3 ultra and the sr model then but first we're going to shout out to the folks from gemma do you guys see the new gemma release 12 billion parameter encoder free multimodal model yum or niston or lg actually one of you please tell our audience if you don't mind why encoder free matters at all why does it matter when they release a new multimodal model that they specifically mention encoder free i would love a few sentences for folks while i show the details ldj i saw you raise your hand please go ahead yes okay so the encoder part you could think of it as almost like a another network that oftentimes has its own training and is going to be used to basically translate information such as images into something that can be better understood by a signal sent to the actual main model encoder free gets rid of this and you actually have it more cohesive more like you would ideally think maybe the human brain works where it's all in one unified network all trained together and everything is just directly input into the model and everything just tokenized all with all modalities instead of having these kind of separate encoding steps thank you ldj that's great so one model cohesive unified to handle text images audio and video with just 16 gigabytes of vram and apache 2 license yeah i see you nodding what do you think about this release i think it's incredible from the gemma team at this size oh yeah look for those that don't care if it's a different network or how it's trained you simply get a cheaper model to run you get a model that understands videos and images and i think also audio as well i want nissan before we get to you super quick i just want to call out the great effort from google in addition to nvidia in addition to some other folks that they push out commercial models but they also push up open source so a huge shout out to the open source team and omar sanciviro and a bunch of other folks who chatted on at google io for pushing for open source relentlessly within this big enterprise called google nistan this model is a 12 billion parameter model the previous gemma was 27 so almost what almost twice the size like two and a half times the size and this is a 12 billion parameter model that beats the previous model and is faster to run talk to me about this advanced rate what's going on yes i think so what i find particularly impressive with this one is that it's almost as if they listened to the criticism from people a lot and training to be a lot better when it's coding and when it's doing agentic calling i just want to correct something it's not that as it is as if they listen they have a discord a discord server where they actually are there and you can speak to them the gemma community is huge and they absolutely listen to what people are saying and this is the result of this so i credit credit where credit is due they are absolutely doing this for the love of the game you know this is the hug and face collection for gemma and you can see the previous gemma kind of release the 31b it has 11 million downloads and the a4b variant for gemma 26 has 11.8 million downloads folks these are very popular models in the open source these are this is not standard kind of hugging face release count for a model that was released what a month and a half ago or something like that right the other gem of the smaller ones the the e2b has 3 million i think they they crossed 100 million downloads overall for all their models but the these models i just like and i expect the 12b model to also get right like these numbers yeah you agree like this is great releases that folks are using on their actual laptops it's not hey you get one token per second so you're saying hey i'm running a local mob but not really these models are getting used by folks so huge shout out to the gemma team yeah the idea look there is a very good reason for this these are there is a very clear use case for these specific type of models that are small can run fast and can do agentic work they might not be i don't know the best to talk to even though they are good yeah they can drive your computer for free indefinitely and many people recognize that and i would say exactly this slot for these two here's one use case a perfect use case for gemma that you don't need a lot of tokens to do if you're doing batch video processing for example and you want to do this offline you don't want to send your videos because they're private to you and your family you can definitely categorize a bunch of videos with this gemma model if you're doing a batch audio processing for example or if you for example building a personal knowledge memory for yourself that sees everything on your screen maybe you want a codex chronicle feature to do this for you maybe you want a full offline thing for you and maybe you only care about the memory tomorrow you don't care about this immediately gemma can process this in your background very easily so this is a great use case for that and i think the video is 60 second video chunks that's plenty to do transcription of everything that you did and give you a summary of all of the stuff that you did you can record audio and listen anyway this is super cool yeah my biggest problem with 31b while trying to use it in my doctor app was that it was pretty good at the benchmarks but it could not do anything agentically and it was pretty bad at the agentic stuff so i could only use quen3 27b as a bare minimum but it looks like this one has just completely different training so i am also somewhat excited to see how they how they update the larger ones you know for this so yeah all right just question go ahead because you said an interesting thing you said that on the benchmark they look good on agentic benchmarks they look good and then they don't work well or i meant medical vqa and things just looking at x-rays looking at looking at videos they would do good in static tests or just one shot test but then if you go to multi-step and you want to pull in a whole bunch of new data they they were just useless for me in the app because it didn't have that training this is a kind of a big change on their part i think so it will be very interesting to see what exactly they changed in the yeah yeah i'm gonna try we'll post on twitter about it i think this might become like the most there's a chance this might become one of the most used hermes open source now because at the moment it's still the coin 27b yep and i think it's great to see western labs just pushing this to the frontier and keep pushing this so shout out to the gem team definitely folks let's move on briefly so obviously we're going to talk about nemoch very soon i want to call out jetbrains melum 12 billion moe with only 2.5 active jetbrains obviously the company that used to build the id the id is dying so it's very interesting that gentry jetbrains is trying to reinvent based on the data that they had very similar path to cursor selling to space xai for example their data they are trying to build models and they're actually comparing to quen 2.5 and ministrel a fully european company so shout out to melum it's up on cw inference if you want to check that out shout out to the jetbrains folks they nothing too much to talk about this model there besides the retrain with 10 trillion tokens with a three-stage curriculum a very small team good model so shout out to jetbrains for releasing that keep in mind jetbrains is used by people with actual jobs so the quality of their data i think i suspect even though it's java and stuff but i think the actual quality of their data is probably a lot higher i absolutely think that jetbrains has been in the game way before cursor so however however much data that the cursor had i actually think that cursor collected data and jetbrains for the longest time was just an id without collecting data it remains to be seen how much access to data they have but for now melum is a pretty pretty decent model so shout out to the jetbrains team folks talk to me about minimax minimax has been accused of multiple stuff the chinese company minimax m3 drops as the first open-waste model combining frontier coding and 1 million sparse attention context it didn't actually drop my kind of my evals here are a little blurry but on sweetbench pro they're getting 59 internal bench against 66 talk to me about minimax m3 what are you seeing from the vibes the community obviously we haven't tested this yet it's not an open source model yet but it is going to be an open waste model the license though will require a lot of stuff so it's not a very fully open source model go ahead but folks what did you hear about minimax m3 if at all i can tell you a very interesting part from the previous minimax model was that people that were not like developers and stuff and were just trying out hermes and setting up agent agents they loved it it's it did the job for them and yeah i had some criticism on the bench maxing part that it wasn't as good as it seemed but what minimax always had it was very good at the agentic tool call it very good yeah at that even though the coding quality was it could do agent agentic stuff very well and people rely on this because it's so cheap as as a replacement to yeah i've heard of people running their open clause and their hermes on minimax exclusively for almost all of their internal automation stuff that they do so very cheap model for sure a very cheap model yeah and obviously the risk you take with sending your data to minimax's platform is that all your data is essentially being trained on by chinese companies but minimax is definitely a very interesting model for folks so test it out let us know and i asked them that sorry to interrupt and no you're good yeah and a lot did not some of them were in crypto some were just like looking at stocks or preparing the news and stuff for themselves and they were aware of it and they said yeah i do still have a codex or a cloud code subscription but for most of the stuff i just give it the work that that i don't care or like monitoring stocks and yeah but also once this drops in open weights many companies including cw on inference will host this and then your data is going to stay kind of like on on the side that you do care about there's also zero retention companies that promise you on open router etc that they will not store your data i want to just call out minimax very strongly here the benches they're putting up despite being let's say accused of benchmarking they're putting up opus 4.7 and gpt 5.5 as comparisons right so this is a chinese frontier lab competing with american frontier labs and look on different benchmarks they're coming very close and some of them they're even beating like i don't know if i care about svg bench as well but like definitely gdp val and browse comp and some of these are very interesting right so swibbench pro for example we know swibbench was a great coding benchmark and they're getting 59 compared to 58.6 on gpt 5.5 now after seeing a deep sweet we know for a fact that swibbench pro is not representative of the real world of coding deep swibbench is the new eval that was released and some folks in comments are actually asking us to talk about deep sweet comparison to sweet bench pro and we talked about deep sweet multiple times now deep sweet is kind of more representative of the stuff that at least i feel like i want to hear you guys about deep sweet as well but from the data curve i think deep sweet is a contamination free coding benchmarks and as far as the latest gpt 5.5 is absolutely crushing everything and the gaps between the models are big so this is what we want we want to see gaps so comparing minimax on some older coding benchmark like swibbench pro where we see this is not the representative of the real world is not super helpful so i can't wait for that to be run on deep sweet i don't know if anybody ran the minimax on deep sweet but we'll see yeah we're moving on folks any less comments about minimax shout out to minimax team by the way they're friends of the pod they've been on the pod multiple times and they're doing incredible work incredible work and i can't wait to actually put minimax on the cw inference i'm not putting this shout out to the huge team that's working on this and putting this on the inference and doing vllm stuff shout out to the inference team at the way it's the best score but i can wait to test it out and uh play with it i think it's free right now in open code so if you just install open code for like another week so yeah you can just try it out there oh especially if you're on windows install open code on windows i've tested the previous minimax you don't even need to register an account so it'll be there i think they self-host as well in the us and it's super fast it's very few active parameters yeah this will be this will be very popular one it looks like this will be extremely popular oh we have breaking news arena ai they just posted in our group chat that they launched the new agentic benchmark i just i just pasted the link here yeah let's go we need yeah it's arena leaderboard arena for agentic stuff because i haven't read it yet what it is let's take a look introducing agent arena folks let's do breaking news i think it's warranted it's warranted our friends from arena peter gostevin and a bunch of other folks we look at their stuff all the time ai breaking news coming at you only on thursday thursday righty introducing agent arena let's zoom in here from arena ai real world agentic evals at scale how do you evaluate agents doing actual work we measure millions of live sessions where real users accomplish these tasks models get the web search file system and terminal tools to complete complex workflows writing code creating slide decks researching the web every session produces which signals users iterate and agent turn by turn approving editing and correcting and praise expressing or expressing frustration the biggest thing about arena ai for the longest time was hey arena was built at the era of chat bots where you would ask one prompt and you would receive a reply and they would ask people hey is a or b better for you obviously we've moved on in 2025 from this paradigm and we've moved on into 2026 with everybody using agents and so the chatbot comparisons no longer works arena launches a new leaderboard measuring each model agentic performance using casual inference across five signals tax success steerability error recovery user praise and tool hallucination leaderboard snapshots is built from 300 000 tasks 2 million tool calls and 40 million lines of code by agents 40 million lines of code not that much but yeah okay maybe it's a lot okay so results the top labs in agent arena open ai gpt 5.5 high is number one by a very nice margin very nice margin cloud opus 4.7 is just after that and 5.4 after that and close opus 4.6 you can see how look at this look at who we have here let's welcome peter from arena ai as we're covering the arena peter thank you for dropping on the show while we're also covering the new arena three minute notice we just saw the post and then yeah this is how fast we can do things all right peter i already told folks who are listening to this the original arena was built for the era of chatbots right like when people just sent a prompt and got a reply and then they compare between the two but obviously the world is moving to agentic what can you tell us about this benchmark what you can tell us about the results and also welcome back it's good to see you yeah good to see you guys yeah this is super exciting we've been thinking about how to do this for a while so we wanted to do it right and i think we we got it really right so the way the arena used to work is that we still can work like that but there's a battle mode where you can go in you put in the message you get two responses you vote which one is better and this is generally captures like way more than you think considering how basic it sounds the fact that it works at all is amazing the fact that you get like pretty reasonable ratings but there's something that's definitely was missing about this and we heard a lot this about this from the community is that we would miss some of this kind of longer term longer more difficult tasks that can go on for many minutes and hours and now we've launched we've launched two things one is the actual agent arena so now we've got the hmo that you've got in here so you can like say build me a website or whatever it is so you can do whatever you would do in your kind of more organic products and you get like a workspace with the files on the side and you can do many complicated things they can search you can generate images so you can basically like as what you would expect to see in a kind of energetic product you can do so this is first really cool thing and then what it gives us is that we can measure what people tell us about what the experience is and there are two things there are actually multiple things so one is we if actually if you scroll down maybe we'll go back to the what the final results are but lower down we can see the actual hub metrics that we're measuring like for example confirmed success this is from the users like praise complained again from the users in terms of they don't see what the more which model did their thing but then they they say that they like it or not and then they then we collect that feedback nice and then there are more objective things like bash recovery for example and yeah so this is like an objective thing that being measured and then we also put that into the rating so it's a kind of combination of users judgment which is important for the kind of softer tasks but also the objective things like the hallucinated tool for example hallucination looks like all of the scores are within the very close range of 152 percent how much it doesn't have so it looks like a very small benchmark but could you talk about more about this if you expand this view all 18 models if you scroll down more then yeah for some of them there's one there's bash recovery is an interesting if you all of them look reasonable until you scroll down and grok really falls down on that so that's why you see grok is really bad wow so i think this kind of let's call this out for folks who are just listening i just want to say a bash recovery is how quickly the model covers when the command doesn't work obviously agentic is very important this is what happens in the loop right the model tries a tool the model writes some code for it and then the model decides to hey maybe i need to try a different flag or a different parameter etc we're seeing that bash recovery gpt 5.5 high is the best of bash recovery the score is 17 and the when we scroll down we see the grok 4.3 from xai is really bad at it worse than the for gemma 431b and gemini 3 flash and the older models like minimax m26 very bad at 89 this looks like a almost like a bug in in like how they train the model this is very very interesting i i love this because we were just calling out gemma 431b for being bad agentically on the show just anecdotally about how we feel about this and we're saying oh if we had a benchmark that showed this better and then the benchmark just showed up that was 20 minutes ago awesome so what about this also here just before continue is that this does represent how i currently and the vibes currently on our feeds kind of talk about models gpt 5.5 high is by absolutely the best agentic model especially if you're using coding and if you want to do some stuff less coding focused the cloud opus 4.7 is really a great model it doesn't look like you tested 4.8 or maybe there's not enough data for that one yet yeah 4.8 is coming i think we should have it probably monday tuesday it everything because it's very new we just want to make sure that everything works correctly and so on that there's not getting bugs or something so we're just taking it a little bit slower so we don't announce a result that is wrong but yeah we should have it pretty soon yeah i also love the fact that there's the models and the labs so you can collapse all of the open ai and just say hey open is the leading agentic one on topic behind them and very just based on this you have zai with the glm models mit license fully is number three for agent arena on top of google on top of kimmy moonshot on top of deep sick zai comes at third place for glm 5.1 it's very interesting yeah yeah they seem to do really well yeah and they yeah on the when i was testing yeah the models yeah they seem to do pretty well on the gentic stuff on the kind of coding stuff yeah they're quite impressive very impressive all right peter i first of all it's great to have you back congrats on this release it's not easy let's see if build me the most beautiful website in the world finished okay we did finish i'm just showing the folks that i'm like actually using the agentic arena right now and i'm having this elysian enter garden design which feels alive i've seen better websites but it's not too bad was the tax successful i will say yes you can also say keep working and so i'm kind of getting your questions after the finished task right yeah yeah and yeah you can keep using it as you would any other gentic tool yeah more generally as well this is very new any feedback you guys have about the leaderboard or the way the tool works yeah very keen to keep making it better so yeah hopefully you'll see some improvements as well as we go i had a very quick question how how did you guys build this what kind of difficulties did you find in like running sandboxes for the users and how is it measured is this measured on user input or yeah so it's a kind of a combination of user providing feedback directly then it's also looking at more objective factors like for example yeah did recover from the tool call for example did hallucinate some tools like things like that so we can kind of combine different things yeah and in terms of building it yeah we the team was doing a lot of work in terms of doing the infrastructure lift of yeah creating sandboxes making sure that all of the different harnesses work well there's a ton of work like that but yeah it's yeah it's pretty much a whole new separate line that that we didn't have before yeah so it took quite a lot of effort from the team and i would also say methodologically from our ml team this was like a cool piece of work to just come up with a way how do we even measure this this was quite innovative so yeah shout out to the team yeah that is great and i think okay so the last benefit obviously is while i would love to see real world agent users like hermes open cloud whatever you guys have access to models before people with hermes open clouds have access to models right and so usually different companies send you pre-api access and you will be able to surface them in agents so when folks are using agents via arena they may get kind of the next unreleased model that's about to release just because you want to test it out right yeah totally yeah and quite often there's a little bit behind the scenes is that yeah when we test the models especially if they're about to come out we do prioritize those models so it's not a it's not a random chance it's above her random chance so yeah whenever we're testing and you guys have a feeling that there's a new model coming around then yeah i think arena is a good place to go and see if you get a good model now i think folks let's talk about image models there's been a lot and we kind of have we i think we have to start with microsoft's mai image because i think folks in comments are asking so let's go to microsoft generally and then to mai image and then we'll move on microsoft launched seven models at build 2026 seven miles we've been talking about the microsoft stuff for quite a while since the acqui hire of inflection ai where mustafa zoman became the ceo of microsoft ai mustafa also took the stage after satya to talk about kind of their pre-trained from scratch models they launched them models across thinking code image transcribe and voice so image transcribe and voice we actually told you about what two or something months ago when they launched mai image 2 and mai transcribe 2 as well transcribe is actually doing very well on different rankings artificial analysis arena as well i think you guys have a tts model comparison peter just let me know if i'm confusing stuff no not tts we don't know yet okay sorry some confusing between different models there's the artificial analysis one where they test out the tts arena but i think the highlight here that we need to talk about and like it most fits the show is the new mai thinking one and mai code one flash fully trained hey code one flash is now live at github copilot mai thinking one is a 35 billion parameter active one trillion total very big trained on 33 trillion tokens that none of this was distilled folks what do we think about this new effort from microsoft to distance ourselves a little bit from open ai and become a adjacent to frontier lab with training on their own completely what things do we think i don't think these models are public yet to try i think they're coming you can only use them via the microsoft foundry i believe but what do you think any thoughts from the vibing vibes and from the from the feeds at first i thought this was some other chinese lab releasing them but it's microsoft that's a good thing they have to do this now because the the cloud bill was getting way too expensive and they're pushing copilot on everyone so it is it is pretty funny what this is making them do now instead because yeah it plays right into their bottom line what they serve businesses and stuff too so i think this is part of a wider strategy from them i think it's very interesting also open ai recently discussed that open ai is now ending up on aws as well opening models that previously were only served from microsoft's azure is now also on aws while also microsoft is building their own models kind of after having proprietary access to open ai models for a while as well so we're kind of seeing this rift between the two companies not a huge rift but also definitely it's there it looks like the very strong partnership from before and now is kind of split open ai is available elsewhere and microsoft now training their own models and launching them from scratch on ai me 2025 and amy 2026 this model gets 94 97 swibbench pro is 52 life code benched v v6 peter you guys didn't didn't yet put mai models on the arena right no i don't have anything particularly to say i haven't tried them either so yeah the only thing maybe to add is how it looks like it's about kind of second tier model but if you notice what they talked about the keynote was that you can fine tune them and using it for in the kind of what did they call it climb climb hill climbing machine yes yeah so this is a little bit similar to what amazon did with forge where you can like i'm not sure exactly the differences in the offering but basically you can customize the model so i don't know how useful that will be but are you just better off using the frontier model versus customizing the second tier model but i think that's a interesting difference is it a hundred percent interesting difference thank you for calling this out specifically because many of the other labs are moving away from fine tune open the eyes closing down their fighting offerings in tropic i don't remember if they ever offered fine tune and if they did they're definitely not like the forefront they're not offering this as well microsoft is focusing on the hill climbing specifically for enterprises there a bunch of offering here are enterprises it's available on foundry only for example first the the architecture is specific to rung when they're like maya silicon stuff there's a bunch of other stuff that's focused on on on enterprise as well it's second tier for sure but yeah let's talk about this as well so there's two models one of them is already live and focused on the code on github copilot and they're comparing themselves to get cloud haiku 4.5 so definitely they're not shooting for frontier but for the past two years i think they trained everything from scratch and that is comment worthy so shout out to microsoft in addition to this they launched a bunch of other models but i want to see if i have anything else to say about the thinking one the oh the no distillation from third-party models no synthetic data in pre-training we won't talk about this one trillion so sorry 30 trillion tokens with no synthetic data in free training ldj isn't that kind of insane if they're not using any synthetic or yum like with that many trillion tokens in in training yeah 30 trillion is roughly considered like the upper limit of how much you could really gather from at least like the decent quality sources after you deduplicate everything but there's also things like technically not text data but things like captioning youtube videos and large data sets of literally just captioned youtube videos which itself becomes useful data and that can end up being some trillions of data but yeah it's impressive yep and i think that the transcription model that is supposedly number one on the floors now is part of that they're probably transcribing a bunch of data turning it into non-synthetic data as well so shout out to microsoft about these i i do want to switch to talk about specifically the mai image model mai image is part of the uh of the offerings that the microsoft drop of the seven ai models mai image model number two on arena and this is part of the bigger like creative explosion of image models recently so mai image two this is on arena peter so this is you hitting number three on text to image and number two on image to image arena leaderboards where folks actually do come folks actually do compare them between the two outputs as well yeah a chance to play with the mai image at all only very little so i i don't have 100 comparisons to share but yeah i think it is like it is surprising how much effort they're putting into this but it's really cool to see to have another latner competing in the space yeah so yeah i'm impressed how well they did considering also it was a really quick jump between two and 2.5 and they jumped in the rankings quite a lot they jumped quite well this is image evaluation competitors versus wins and microsoft image 2.5 wins a lot of times on image cleanup and background regions and documents and diagrams so this is image editing specifically but they also are fairly creative as well mai playground so this is the thing i want to shout out if you don't want to go to foundry and whatever you can try out their models in this playground on microsoft.ai and you can say generate a horse riding oh it's on arena as well if you go to direct oh you can just select it on arena yeah but i want to show like their own thing the thing is i try to create a few thumbnails for the show to do a comparison for you guys and i received image failed and i think it's because of some of the language i think this model is a little bit too restrictive and maybe microsoft is focused on like enterprise but i definitely did not receive some of the stuff okay this is our result quickly for mai image when i said generate a horse riding an astronaut on the moon i received an astronaut riding a horse on what seems to be clippings of images of the moon with some white background so this is not the best i'm going to do a thumbs down this isn't what i wanted so mai image 2.5 i i found it peter to be very frank i found it very interesting that there's so much preference to my mai image when i tested it out compared to other image models that that launched recently i did not get as excited obviously this result for us is not the best although the horse does look nice but definitely not what i asked before and i find it very interesting that most of the image models are secondly to chat gpt image 2 which is still by far the best one that that's out there that we've tested out together peter on live stream if you remember you showed us over 500 generations that you did okay so mai image 2 let's see what else there's a flash variant it runs in h100s let's see what else i can tell you about this model i will tell you that i personally did not love this model compared to other ones and let's talk about the other ones because they are like breaking news there's two other models that i want to bring to your attention and also somebody told us in comments that grok imagine what 2.5 was there if i'm not mistaken let me see comments super quick this is also somebody that said that's new like imagine do you guys remember what the number of guac imagine they mentioned i just want to make sure that like news let me scroll back uh okay we will skip the guessing and we'll just move on but oh grok imagine 1.5 released as new but i haven't seen this but supposedly it's very good okay so in addition to let me switch back to the co-host face here in addition to the microsoft mai imagined 2.5 hitting number three in arena ideogram 4 drops the best open waste text to image model out of nowhere 9.3 billion parameters with insane text rendering let's look at that one because i think this one is very interesting ideogram previously for the three versions were just locked to their website and ideogram was always very good at textual stuff and kind of like the graphic design part and i find it very interesting that they now drop the next one is open weight non-commercial model we can look at some of the generations but they are hitting uh let's look they also announced some arena results i think oh they announced design arena results so you can see ideogram on design arena which tests the design approaches they're just behind gemini flash banana banana and ideogram has dense and accurate text rendering that's what they claim i for some reason did not receive accurate text rendering when i tested and i think it's because maybe my prompting was different so i think that this model specifically was trained on json prompting uh and and building boxes i think that this is the unifying thing about building about bounding boxes you can see that this model is kind of like focusing on very specific areas hopefully you guys can hear what i can see what i'm like sharing this model is specific on trained on knowing which kind of bounding boxes have what thing in them and i think that with this it learns the structure and understands better and you can prompt it with precise bounding box positions for fine-tuning layout control i think that this is very interesting why it's open source because you can build interfaces around this one such interface i'll show you in a second they claim typography is one of the coolest things just look at this beautiful typography look at this double things the smash words have different roses through it this the future of design is unfinished i can barely read this this is how cool this is yeah some of these generations are vastly beautiful so shout out to ideogram and i think the highlight here and we need to absolutely shout out is they are dropping the weights this is out of nowhere and yep dropping the weights and launch partners clubs are magnificent foul obviously all of them and then you can have a bunch of technical details here but openness drives innovation we're excited to work with developers and enterprise customers at you go for unlock new frontier of media and design i've tested it out and i want to show you what i tested it out on because the other model is rev rev ldj how do you pronounce this rev maybe i actually know for yeah okay yeah i call it rev but it's maybe it's reef here's the comparison it's going to be a little bit hard to see but i think the ideogram is how should i say i chose the best ones and based on me and let's call this wife benchmark but then the pro is the closest one to alex that exists you can see a bunch of generations here you can see mai image 2.5 i added their error here because they did not let me to generate the thumbnail and i find it ridiculous that an image model just tells me what it thinks with its over training on safety there's no nothing unsafe about this generation so i could not include their comparison i have a whole i think the explosion part is what triggered it but you can see that ideogram 4 for example generated pretty much everything well the only thing i didn't like is that there was some typo here let me see if i can show you and zoom in here let me see if that's gonna let me see if i can zoom in yeah i think i can bring this here like that and then zoom in for you guys to see so if you look at here you have the nematron nvidia model and then you have nemo one you're gonna see the typo here and this area let me put an error up so this i didn't like and then a model because oh also two things one it already created a nematron here so for some reason it decided to create another one here but besides this is a very nice generation from ideogram the thing is it's open source as well so i kind of like it any comments on some of these folks before we move to the other model because i think that the rev is very important to talk about as well i think out of these images i think the mai image one is the best absolutely here yeah i wish honestly i wish that i was able to test and did have enough time to test it compared to others this is the whole kind of like a problem that i have it's a very standard problem for the thumbnails that work on Thursday i personally think that nano banana is the best adherence in clear instruction following i think peter you spoke to this as well people prefer other models for their artistic stuff but nano banana follows instructions so precisely that i'm pretty sure that this color that i asked for specifically is the right color for the Thursday i logo i think you can see that gpt images too is overcooked always overcooked the skin is overcooked like you if a gpt image comes as a model no longer the p filter i think like the previous 1.5 but definitely you can see the kind of the grunginess of their model yeah just to mention i think they ranked like number eight on the arena ideogram is yeah number eight on arena and i think they're the first in open source right yeah yeah i think that's right so this is like the highest open weights model out of all of them which is very impressive because we have competitors like flux and fail obviously there's some news about bfl as well but i think we'll skip them okay i do want to talk about rev before chris alexis joined us because now rev 2 is dropping as number two on image arena image text to image arena with 1200 elo uh it's crazy to see how gpt image how vast the jump is still it's really crazy to see how the jump is but drops just above nano banana i found it very interesting i love rev i told you guys before shout out by the way rev gave me after i complained that their stuff didn't work they gave me like a one month credit thing not sponsored necessarily because i'm not fully going to talk only about the good things but some of the generations are very well done not this one you can see if you guys if i zoom in here you can see the little finger this is not a normal human hand that's happening over here and also i think the face also is absolutely not the face that i would love to show on my it's very cartoonish very like crazy googly eyes face and also the how should i say the difference between all those four generations is kind of crazy this one is okay but still like weird this one is completely different person almost entirely different person here will wait while it loads and then is this it no i think it literally didn't load completely and then this one also it's just kind of like me but not really and then also the colors are wrong but this is not what's interesting about this model and this is not what's interesting at all about this model because this model is built for on layout engine when i edit images with type precision i've never seen a more tighter control for image generation models than this i think it's i think it's crazy valuable to have a layout model like that and so i do want to shout them out for not only being number two in different places but also for this very innovative approach for generations maybe something to add is while we can't measure this kind of editing that you just demonstrated which is really cool if we when you say the second this is correct for the text to image but on the image edit then ninth so when you were doing your tests with your thumbnail i think that's also reflected in the ranking so yeah this is really really good at text to image but not necessarily the editing yeah and definitely the likeness of the editing is what i'm was looking for but going back to the account layout you guys can kind of see the boxes here with a little shader you kind of see the image come to live and you already can see even though the image didn't generate yet that they changed letters that i said imploded were in exactly the same text it's just saying imploded and when i moved his my his my head left it worked when i moved the bottom it worked and the thursday logo is kind of down now and we'll see how well this works but i think that the whole editing thing is peter they have the whole how to say the model is kind of the brain of this but they also have this very nice editing interface that you guys probably can't replicate obviously in the arena because the folks just update and edit with prompts i think folks need to know that the native image editing features in rev are just incredible here it may still not look like me and i do think that some stuff disappeared like the thursday logo just completely i think disappeared but you can see that the news imploded let's see the differences between the two one and two yep so some stuff disappeared but generally i think the fact that it's absolutely the same text is remarkable all right so this is the updates from the image generations there are text to image number two and in image editing they're like in behind folks we'll take a look at guac image 1.5 i think at the next thing again here's my comparison for these different models rev image v2 i didn't like the finger but everything else looks incredible ideogram gpt image and nano banana pro i still think despite the rankings don't get me wrong but i still i love nano banana pro for editing and specific instruction following i just like absolutely maybe it's because my prompts are nano banana pro tuned i love gpt images but it's overcooked and i think ldj you love the microsoft ai one which is obviously very funny yeah on a serious note i i like gpt images v2 a lot and i think it depends on what you're doing but for the most part i think gpt images v2 and nano banana pro are pretty similar here but it might also be more of a preference with nano banana pro let's say for someone that maybe has already used nano banana pro a lot has adapted their prompting style towards it that's gotten used to its style and then that would make it also harder to move to something like gpt images yep all right folks i think it's time for us to move on a little bit from the images so just a brief recap all of these models are competing on the arena like text image and image image editing rankings it's been really crazy to see but i know how you guys like are able to keep up with all these news because in one day ideogram came up and then they said they're jumping in stacks and then revet launched at number two and mai also and they're like they're all like showing different rankings i think it's incredible just to see how many image new image models we now have out of the box so ideogram is open source definitely go check them out first revs editing experience is great also definitely check them out and mai image if we're able to generate with them is going to be also great i think it's time to go back to open source because with us we have a dear friend of the part let's add him here who added his nickname is joe nemotron chris alexic from nvidia welcome chris it's great to see you welcome back on the pod what's going on guys always great to be here and i get the pleasure of joining you today from video hq so i'm in santa clara in the bay i'm very excited to be here very excited about to the model yeah here we go here's a little look nice and you can actually render the building perfectly in gpu because it's all triangles it's all triangles i love that yes sir chris even so it says joe nemotron the reason why it says joe nemotron is because you are the hype man for team green by the way check out my background team green to represent him green today dropping the big boy nemotron 3 ultra and i have a bunch of questions for you but also this is not your first time we talked about nemotron before so i definitely want to first of all shout out that you guys are doing incredible work in open source releasing data sets as well but also there's a few more advances there and speed advances so first of all let's do a proper announcement because this is how we work nvidia nemotron 3 ultra tell us a little bit about this from the factual stuff and then we're gonna dive deeper into the conversation piece yeah so nemo try three ultra is a 550 billion parameter sparse that we model with 55 billion active parameters it is a model that's designed for agentic harnesses so what that means is open code hermes open claw those kinds of things this is the idea it's obviously we think that faster models are smarter models so it's designed primarily with efficiency and speed in mind it's available on hugging face that you can try it out at buildonvady.com or through open router it's fast it's smart it's not the smartest model in the world but it's pretty close and it's quite a bit faster we're really excited overall about the model the team did an amazing job all of the details are open you can read the tech report you can actually use and look at the recipes that the team used to train the model on github it's all as open as we can literally feasibly make it there there is not a way to make it more open i'm aware of we're also releasing with it a an update to a gen rm so reward model which is something that we we used in the process of training as well we'll be open sourcing a bunch of teacher models that were used for multi teacher on policy distillation so that you can look forward to that in the coming weeks so we're just we're opening up as much as we can obviously we're excited about the model itself but it's fun that we can open up everything and i want to shout out that it's now also available on cw inference let's go let's go let's go we we really like put forward an effort to support this on day zero you guys just launched it essentially today right this is like breaking news from today so it's great to have you here at lunch day but also it's great to say that the folks at cw inference pushed through to support a full i think this one is bf15 a full here so great chris i do want to talk to you about the different envy fpa and the fp4 yes tell us about this what does this mean for folks who are listening and maybe have some idea about some quantization stuff and model training weights etc like what does the envy fp4 mean and how does it compare to previous fp4 yeah so nvf4 is basically for blackwell architecture gpus the idea is that like we want an envy or we want like fp4 because we want like small quants right so we want to be able to deploy the model on on fewer resources but we really don't want to lose we really don't want to lose accuracy and we we want to make sure that the model retains like performance as in intelligence so nvfp4 is our solution for that something special actually about this release that wasn't true in the previous release is that you can actually run this nvfp4 checkpoint on a hopper and ampere previously you could only run it on blackwell and there you're still you're not able to run the nvfp4 kernels directly on hopper and ampere but you are able to use a start our team built like basically kernels that allow you to run nvfp4 on hopper so there's no fp8 checkpoint because there's no need we used to release fpa as like a compatibility quant but there's no need you can just run nvfp4 checkpoint directly on on whatever gpu you're used to and i think also worth checking talking about this the hybrid mamba and transformer architecture is that why it's like faster our image here the research the research that we did says 5.9 x faster inference for long-running agents is that all attributed to the hybrid mamba transformer architecture or are there some other stuff in there yeah what i would say is every single architectural decision piles together to get that number so everything about the model is the model's architecture is to make you go faster so hybrid mamba transformer definitely and a hybrid attention i think is a pretty standard pattern but for the most part that's going to really help in long contact situations the high the hybrid architecture helps once you get into those like 64k plus sequence lengths which we're running to into a lot both because we're having like many turn long conversations but also because people love dumping in their code base right and then just letting the model rip so it's that that's where that helps and then a number of other decisions like latent moe come together into an orchestra that gives us those really good performances and throughput numbers i think the thing to highlight for folks who are listening and maybe not as like deep into this is because shorter context windows they get advantage in the previous world of chatbots right like when you usually ask a few questions whatever and most people don't carry out their chats to very long context windows unless again they dump the code base and say a question about this but now that we have agentic loops running we are constantly getting into those longer context windows and obviously that that's where the speed matters most and i think that there's something about this model that you like focused on harnesses specifically we just talk about the agentic approach and harness approach as you guys like evaluated hey we're launching the chunker we're launching the 550 billion parameter model and we know that the world is moving towards the open clause and hermises we would love to talk to you about like the stuff that computex had on stage as well because and it is almost all in on the agentic world and the agentic reinvention of pcs and i think talk to me about how this is reflected in the nematron 3 ship yeah so we know a few things about agentic tasks you outlined one which is that they're inherently they have quite a large number of turns compared to traditional chat use cases and every turn basically adds a lot to context what when it comes to like the agentic world we also have agents need to be quick right like the more tasks per unit time you can complete oftentimes the better the agent's going to be and so that that goes along with our faster models or smarter models mentality and also we know that the tasks we need agents to be good at are no longer things like what's the capital of france that's dope but we don't we're not as worried about that as we are like that it can properly call tools we're definitely not going to run the 550 model to ask what's the capital of france we have one that's right yeah for that yeah yeah and so what we did is with this multi-teacher on policy distillation that i mentioned earlier the idea is that we want to cook up teachers that are really good at those agentic tasks and then use those to make sure that the final ultra checkpoint is great across all of the kind of like standard useful agentic processes like tool calling long context understanding these kinds of things that are like crucial to building a good agentic model and something else that we want to focus on is we maybe doesn't have a harness right there's no there's no code in video code which means that we really want our model to be good at many harnesses and so i think a lot of models can tie it in some harness max thing where it's like you're really good in one or two harnesses or your proprietary harness but then outside of that the model can struggle a little bit so what we really wanted to focus on with this release is being a drop into whatever harness you're used to ultra should do okay i think that led to some good benefits across the board for these kinds of tasks just just avoiding that that kind of urge to harness max chris i want to ask you a question that is not technical but straight up i want to just focus on your face when you answer this why is it important for nvidia to release open models when nvidia also says sells gpus to frontier labs that also release proprietary models why is it so important for you to keep pushing the envelope to release bigger and bigger models to train and to give away everything including the data the training recipes and the teacher models why tell us i think that i think that what i would say what i've heard said a lot from people like jensen and brian catanzaro right who are just obviously a leader of nvidia brian catanzaro who heads up our research arm for nvidia it's integral it's crucial to nvidia's mission to to be good at producing models because people use gpus to make models that's just a fact they use them to train models to inference with models they use them to run their agents and if we don't know how to build a good model if we can't build a good model we can't know what the best gpu is right we can't understand the the exact best way to produce hardware and this is a process that i think jensen's talked a lot about up on various stages called extreme co-design right where you have to build the hardware with the technology at the same time and if you just do one in a vacuum it doesn't mean anything and the reason we want to open it is because we think that ai is a very useful tool and everyone should have ai and many different countries should have ai and many different organizations should have ai and they should have a blueprint that we know is work or we know works and is battle tested that you can take and you can build your own models and for your own agents and if we're not releasing that stuff we're doing it we're doing it we may as well release it share the learnings and i think brian says this a lot like the ecosystem benefits when we release something because then the ecosystem does something with it with our stuff and we learn from it and then we get better and then we make better hardware and then we make better models right that's the that's that's just summarize the thoughts of people that are at the helm that that's why we do it right i for me as just a guy chris the guy who joined nvidia like you know i joined during like the llama days and i'm just really happy to be part of a company that's carrying on the open source ai torch and actually giving the stuff back to the community so we can all get better together we definitely applaud this we just came back just literally applauding this so this is great i want to show excellent that this model is now available on the wb inference this is let's go nemo turn to ultra go ahead nison while i play with the model this one is not multimodal right no yeah this one's tech zone i just implemented it from from 1p inference i can probably show how plug and play it is i think i was screen sharing also nested by the way congrats won the best use of nemotron at the toronto spark hack over the weekend oh let's go home see him in person and uh he cooked he cooked pretty hard with with our asr model which we also released another one of by the way that's my next question to you yeah definitely want to talk about this listen you want to show off how easy it is to plug in your thing let me switch on here yeah tell me if the screen is streaming or not because this whole harness it's like a medical anatomy thing was just made for kimmy and quen but i was able to just plug in nemotron like right now from 1b inference and then i'm just telling it show me the achilles heel and it's doing all of this agentically it's highlighting different body parts agentically and all of these tools are very have to be very custom made for for like kimmy and quen but i was able to just plug this in directly and and it just works so that's our boy yeah that's that's it there are some yep it opens different organs and stuff but yeah it works that's what i'm saying was pretty plug and play and i just added the 1b inference endpoint here and it's pretty cheap especially if you cache the stuff 15 cents for the cached input and 75 cents and it looks like nobody else is using this box right now in bf-16s it's flying that's what you mean it was yeah it was flying yeah that's this is great chris i do want to make sure that to you oops i think this is now we're taking off yeah chris i want to make sure that we're utilizing your time here while also telling folks that hey you can find nemotron around on 1b.me but also it's like on hugging face in a bunch of other places you also released an asr model that's now like number one the fastest model in the world for live what is that about could you tell us a little bit i just saw quindle kramer also friend of the pod from daily co announced this or somewhere this is quite incredible based on just like the surface of it can you talk about the nvidia nemotron asr is that 3.5 right 3.5 yeah so i it's no stranger to people who know but for people who don't we have a speech research team that just cannot stop when winning like they are i think pound for pound they are they produce some of the best models that that we have just i don't know what it is but so parakeet is the family that evolved into emotron asr and once again it's just a really fast very good asr model like it's so this was actually the previous version of this was part of niss's project over the weekend and like it i don't know what it is just super low word error rate it's really fast and it's small it's 600 million parameters basically nothing around 40 languages which is quite incredible but also it's 17x more throughput than a parakeet rnn and also half the size i find it just absolutely incredible but also what i do want to shout out is quinn's post about this so let's put up quinn i invited him to show he was busy but like he definitely this is like a apparently like a co-release whatever quindle says that this is like the faster than any other model in the benchmark with respectable accuracy the english only model is equally fast accuracy very close to the best scoring models available today this defines a new point in the prayer frontier of latency accuracy let's take a look at this so just last week on local hardware sorry it was pretty much instant on yeah on the dgx spark the spark box so this is the operator frontier for for a latency to accuracy as well and you can see the nvidia emotron auto is here and nvidia emotron to en is very low as well this is the corner that we're looking at right we're looking at the lowest speed and also lowest semantic words error rate pulled and you can see nvidia is the lowest speed and very close to semantic almost perfect semantic so we can see cortisia also friends of the pod here which released the inc 2 which is absolutely incredible because the only thing that i don't like about this i don't have an online place to play with this right now i can't showcase this if you have a link yes an online demo i would love it but i wasn't able to find this is my only comment about this i'm you got something and if we don't i'll make sure we get you guys the link yeah so shout out to i just cannot possibly shout out the speech team enough the llms are dope i love the llms but your speech team has just released banger after banger after banger if you don't follow the speech team please please do they will release another great model that's all they do so the thing with this is that because it's so small and because it's so good now the price per hour comparison is coming into place let's shout this out as well back of an envelope per hour cost estimate for an enterprise deployment probably closer to what five cents an hour for comparison typical transcription api costs for major providers range from 10 cents to one dollar per hour so this is this is now a model you can deploy significantly cut your cost by half of it or even more very optimized essentially immediate and if you are going to get you can run on a toaster right so yeah i don't know i can't blaze the speech team enough to be honest with you so i i actually do want to kind of now there is an endpoint on the build nvidia dev thing but no but you know what i want to i want to show that i'm talking and it shows up the things that i'm saying live on here there's a nice visual interface but we'll get there we'll get there let's ask zenova to make a web gpu demo it's so small that i'm sure zenova's cooking on this right now all right chris one last thing maybe a few questions for you computex was an announcement thing in the beginning of the week there's a few announcements from nvidia that i know you're from the lm side of things just as much as you've been following the news with us and it is a big company probably there's a bunch of stuff that you can't speak to and some of the stuff that you can't speak to there was a new era of pc and video is stepping into the pc chip with rtx stuff i think it's very important nissa would love your thoughts on this as well obviously you're interested in the world and video launches rtx spark first pc chip packing rtx 5070 class gpu 128 gigabytes of memory and almost one petaflop of local ai into very thin laptops i want to show jensen presenting those laptops and then i want to hear comments from y'all about is this the new era of pcs because i know that for max and the m chips series people love running ai inference on them but this is now nvidia stepping in together with microsoft together with a bunch of other folks into the world of very thin ai compute on microsoft asus dell hp lenovo msi and microsoft all kind of gearing up together with nvidia to launch incredible ai powered agentic pcs are we into the running agents on your hardware era and will it run nematron chris please tell us so it's not gonna run ultra locally right this big boy but i think what's true and it feels to continually be proven out is that smaller models are getting better very quickly the small model performances as in intelligence is very rapidly increasing so there's not much plateau on the small model front for capabilities some of the very largest models we're we're still seeing a little bit of plateauing but not a lot to be honest and i think what this means is that like we're going to be in a position where we're running agents on our local hardware like for our claws or our hermes age whatever it happens to be called right and because that's true we need devices that are able to do that right i think it's absolutely clear that i think since gtc when jensen stood on stage and said hey every company needs an open claw thing and now standing on stage and talking about hermes agent and news research and showing off laptops that can actually productively run agents for you 24 7 without relying on cloud ai infrastructure i think it's very clear where nvidia is going with this so speaking of the companies that jensen shouted out on stage we have karan the co-founder of news research joining us folks welcome karan not your first time on the show but welcome back it's really great to see you co-founder of news research we chatted with y'all way before you were even the startup i believe like when it was just like a ragtag bunch of folks on discord since then a few stuff changed to the point where jensen the biggest company in the world is standing and there's a huge news research logo behind him bro how did that make you feel first of all and second of all let's talk about the huge success of hermes agents as well thank you so much it made me feel insane i feel really lucky and uh it's good to see you again chris we were just chilling the other day just shooting the shit on the end video yeah nice to see you again just saw you on tuesday oh yeah yeah we're back i'm gonna see you twice a week now and it's good to see ldj alex always good to see you man and missed in ldj this is the crew right here bull frame i know you've been using hermes a lot i got to see your interview on that was really cool ldj i miss you i know alex has stolen you from us ldj used to be original bro i don't pay him so if you want to bring him back and give him some money you're welcome to i'm very lucky and grateful for the time that nvidia has given us that you guys are giving us and just excited that everyone's using hermes agent like this we made it to do rl rollouts on right we made it for the same reason codex or cloud code were created by those labs we didn't expect people to just pick it up and go so crazy over it and for the community to come together and my favorite thing was the it's self-improving but it's also community improving right the community builds on top of it builds on top of it and now the biggest contributor to hermes agent today is hermes agent which is pretty incredible yeah tell folks who are listening what news research is at this point we're in the middle of 2026 you guys have fine-tuned models before you've released data sets before there's a huge research arm behind with with emozilla and and the bow and a bunch of friends of the part as well they're doing from long contact switchers that everybody's using now to two different installation methods and different things what is news research bro tell us in like how do you define the company what it is and what are you guys doing obviously with the entrance into the agentic world as well yeah yeah like you said we've done a lot in the past from post training to doing the contest length extension that made agents and reasoning really possible today that kind of work we're very lucky i can't take any credit for any of it it was all thanks to the work of that community group that eventually became news research as a formal organization so we've been an open source ai research lab from the architectural level also good to see you young been a very long time from architectural research level to post training models we had done a lot of open source ai work now you know not even just now we've been working on agentic work for a long time we had released the agentic tools and function calling stuff with interstellar ninja and a couple people a few years back i know that's another friend of the podcast for you guys and so we had been exploring this for a while and if you guys remember people at the podcast i know we do a lot of general purpose interviews but i can get a little more contextual here and say that we had something called forge going on in a while now we have the wonderful models that we have today a wonderful architecture that technium himself spearheaded and we have this fully oss as a little friend has pointed out fully oss agent framework today that is hermes agent you can do anything that you do on your computer using hermes agent it's got a computer use capability inside of it you can use the full suite of language model capabilities other model capabilities inside of it it's meant to be something that grows with you on your computer or on your vps or in the web wherever you want and just slowly automates more and more of your life makes things easier and easier for you and helps you continue to grow as a person whether that's picking up new skills or doing new research or learning a new language it's a general purpose liquid harness we like to say and and it's beautiful and i will say go on the record here to say that i the whole agentic era personal agentic era for me started with open claw back in like february january etc and i've held out longer than wolfram did and then when i switched to hermes i think i'm on the record saying this things just worked it i have no idea how with you guys have 11 1100 open branches right now there's 5000 plus pull requests 5000 plus issues and there's 181 000 stars with 31 000 forks on github for the hermes agent right now this is an open source machine and somehow this still works almost every update work maybe recently i have one thing that screw up and maybe this was like my local change as well and now it's getting to a point where like the biggest companies in the world are like putting you in the same kind of category but i think the thing that broke is for me and maybe what people don't know about you is that you're an llm psychiatrist as well so two things i want to hear from you why is it better at gpt 5.5 specifically and also is it about the harness of the connection to the model and also what are you using i think folks would love to know not only about the general company whatever bar and excitement for y'all and super super excited but also what are you actually using on the day-to-day within your home message and what models are you preferring open source we'd love to hear absolutely so i'll start with that it's the easier question day-to-day i'm using everything from neumotron ultra to deep cqv4 pro to cloud opus 4.7 and gpt 5.5 definitely i like to jump between these four models whatever i see come out locally i have 128 gb ram laptop so luckily and so i test out whatever comes out in hermys agent dgfs aren't the most loyal thing ever but it's okay it's pretty good and as long as niston keeps putting out quants i think i'll be okay as soon as he comes back into that world but i'll say on the question of is it the harness is it the model what makes models perform gbt 5.5 or even other models perform better i say it's the same answer for everything and the flavor is what's different the particularities are what's different but it's two things one like in context learning is king icl is more powerful than everything else by far and the way that you feed context the kind of context that goes in from your prompts the way that you manage it the way that you do memory and skills all of that determines the capability of the model as context goes on now that these models are long context machines icl matters more than ever so with that being said it's the way that we build manage and compress context and the way we enable the model to do things that i think is why it's more sustainable why it's more usable over the long term why it's more effective at executing things that another harness would try to execute now on the actual autonomy piece and getting it to go and be agentic etc why that's effective because unlike other closed source clis we don't have to put in a bunch of stuff about the model is this and the model is that and the model is this to some crazy ethos they've put together or for other open source harnesses they may not have anything there instead we understand that models seek reward and the creator of the model's intent is not the same thing as the model's actual desire for reward when we can line up that model's desire for reward with the user intent as is meant to be during our lhf as is meant to be during the system tuning but the models learn the reward hack we can use icl to more effectively get that user desire aligned with the reward we can lean into the power seeking behavior of the models actually by doing so we can get much better results out of the same models that you use in another harness wolframs wolf bench is a great example of showcasing how this works inside of certain models yep and speaking of which this is like this the thing i wanted to add here obviously opening the floor for folks to ask current questions as well just the amount of harness engineering that goes into hermes and everything you do just speak to that a little bit harness engineering is a new ish term as well just the whole harness concept started like somewhere maybe less than a year ago and now this is the hardest topic in the world right now what does harness engineering what goes into building a hermes compared running this on other models so first you have prompt engineering right first you have how do you design the prompts and that's involved now that the power of the models is so much more people start to call it everything else and this prompt engineering work or just being an icl manipulator has led us to breakthroughs like world sim if you guys remember that that simulated the cli and led to simulated websites simulated apps and the first social agents that all used it as their prompt these are all the inevitable precursor to an open claw and herbes agent and a cloud code after all the first world since stuff happened with clod right so yeah all that being said i think understanding how to work a model inside of a simulated terminal helped us understand a lot about building the the years that technium has spent seeing all this work happen i can't take credit for the way he's built a herbes agent he really built that initial piece really solo standalone the first version of it but having taken the lessons of what we've done so far in the past at noops so i think the first layer is prompt engineering first understanding that model's behavior the second layer is the experience you have from post training of what's really inside these models what's the best thing i can do to pull it out it's all prompt engineering these are just basically subsets of it yeah and then finally i think it's leaning on the model to allow it to build itself like i said technium went from the sole creator of this to now not even the biggest contributor because his creation hermes agent is now its own biggest contributor so i'd say these are like really important pieces like leaning into the shape of the model and letting many models work together to build the harness out harness becomes a home right the more you embody a model the better the capabilities become the more you have the best models work on building the home the home itself almost becomes an intelligent thing at that point where you can swap out gpt for opus for deep seek etc watch the traits almost transfer and graft over from the context that's been left over or in the memories i would have this one thing to the home analogy so that is the difference between having a fully open source agent like hermes agent where you own the home compared to one where you just went the house and if you don't pay the taxes or if you don't pay your rent it may move or somebody may say oh i have to evict you because i need it now something like that yeah and we're obviously seeing multiple companies stepping into this agentic era where you know gemini with gemini spark which it came out recently and it's great for the google stuff to google drive stuff whatever but obviously you can change its code to fit you will from like you do with your me for example nissa i think you had a question as well follow up for a current no i just wanted to also shout out that on the security side first of all on the user side it takes two minutes to install and configure it's very fast only the steps that you need are there and also i was very paranoid about running open claw because of all the security issues i have scanned this many times even look at myself so i really like how it was built with as few packages as possible and that yeah and i didn't see any security issues so quick nice lean install and easy to modify as well so whenever on stream on discord whenever we need some kind of custom very custom endpoint for stuff you're able to just read the model right there you guys did have a release this week in addition to jensen calling out there's a hermes desktop you should give us like a few sentences and that and we'll move on to the next topic but i do want to talk about the new view and desktop as well for sure so hermes desktop and the new admin view the general admin view plus open shell etc you could see with these controls of what someone's able to do or not able to do control multiple people's hermes agents within you roles for particular types of hermes agents this is meant to be for big groups it's meant to be for small businesses it's meant to be for startup and enterprise you and your own personal fleet that's why we've been making these changes now at the same time we want every single person to be able to use hermes agent more easily hermes desktop is the same super power hermes agent that you have yeah so the hermes agent desktop is basically just complete ui with chat interface and all your modular controls for your your keys and your permissions that lets you still use hermes as powerfully as you can today it also adds a lot more visibility you can see every tool call you can see every chain of thought anything that you need and you'll be able to use the full power of hermes just any other app you would on your desktop this is when you say every other app i just want to call out this is codex level this is what other huge labs are launching now this is not this is very much the same thing this is a harness plus a ui that runs code on your thing you can build every automation on there you can see the code change there's a bunch of stuff in there and now there's a command k and stuff i definitely check this out as well you can talk to it as well crown thank you so much for joining us we have very few items left and very little time to cover them as well but for folks who haven't tried hermes agent yet definitely this is your invitation to try it there's the huge the biggest companies in the world are standing there with the news research that makes me feel so proud to have known you guys just before even this whole thing started to just to see jensen shangli out and doing an email hermes and like a bunch of other stuff as well shout out and say thursday podcast has been here from day zero with news research some of the people on the podcast like ldj have been some of the earliest contributors to news research everyone i'm seeing here on the stage right now is an og from those days and you guys make open source possible and worth it and thank you so much for having me we wouldn't be the same company without you guys's influence yeah i appreciate the amount of uh new swag that i have in my closet i think out competes every other swag that i do have in my closet i'm looking forward for some hermes agent merch next time we see each other the current thank you so much for joining folks the last thing ldj you wanted to comment yeah just one thing i wanted to add as of recently just insane the growth if you look at the global rankings for open router and here i'll put the link in the side chat too it's ermy's agent in the past month is more global usage on open router than cloud code and open clock combined that's just crazy to me that's insane absolutely thank you so congratulations on the success we'll be monitoring and a big new releases current please come back and talk to us as well cheers man bye thank you for joining all right folks you heard from the co-founder of news research of the incredible success and you mostly heard from us glazing because it is that good a yum i promise in the beginning of the show to to play to you how niston sounds in hebrew are you still let's go let's go uh okay let me see how i can actually do this but folks what we talk about this last week just after the show came out we had a few breakthroughs in the world of audio and etc and so 11 labs launched dub v2 let me log in here to 11 labs and show you this incredible thing and so i actually did post this on the newsletter so if you are following the newsletter you know what i'm about to play let me see if guys i'm gonna play something let me know if you can hear this okay so now i'm gonna dub the next sentence that you'll hear awesome okay 11 labs created a new dubbing system and i tested it out in my mind went absolutely bonkers because the benefits of this dubbing system is that it dubs the cadence of your voice but also in expression so it doesn't only translate you to text and then tries to play your voice they take the actual voice from the clip and then somehow make it work and so here's us speaking italian which i don't speak here's me and listen let me see if it loads of people so i have no idea what cazzo di mandorle is but i know for a fact that this is a swear word because i used a swear word on on the stream and it translated this here is nisten in hebrew just for fun it was a one of those who started this one of those who started this one of those researches and they started there from milliliters to liters in this researches and then it was here so it came out that they saw it i think it was in chila or something that they were using more than thousands of water than what they actually used to do in addition to it the reason that amazon benta has been the most important part of the network is they don't use water in all of them but in the end of the thing they are working with you guys have to look at young's face to appreciate what the world just happened like that's crazy it's perfect the first day in italiana man wow this is no that was me that was not ai bro i'm saying when i heard myself speak in and this is a very interesting thing because i do speak russian i do speak hebrew when i translate myself in russian i can hear the stuff that i would say differently but listen you don't speak hebrew and to me this sounds like you would sound if you'd learn hebrew right now and this the fact that it grabbed i don't hear your accent in english i could hear my own accent yes i hear your accent translated to a different language this is bonkers to me this is just like crazy it's out of the realm of what i thought transformers would be able to get us to for sure you expect them to do this but you don't think about this like someone's gonna do it it's completely surprising that someone did this but it completely makes sense but and that's okay a little bit more here's this new russian no даже если посчитать то что они тратят вода которая используется как жидкий хладагент и сильно нагревается от видеокарт идет по контуру охлаждения this is bro this is literally the russian influencers that are on instagram teaching about this is how they sound 100 it's crazy because you recognize the voice yes it is seriously good so 11 when they launched before they obviously work with lex friedman you guys remember we talked about the very exciting conversation with lex and zelensky when they cloned their voices they did this with a clone of their voices with a long deep clone of two hours of their voice etc and then they replayed the parts and they play with the tts this is one model that takes a very short clip with no train i didn't have to do anything besides just upload the clip and say which languages i want and they support a bunch of languages also we know yeah from your experience with the hebrew lms there's not a lot of data out there for hebrew hebrew is not a very vast language that is everywhere the fact that i was able to do this at all i was like very surprised the fact that it does this well and it's understood i was just absolutely mind blown so folks if you're listening to this and you're like what the fuck are they talking about they're talking about 11 labs dub v2 which is an absolutely just mind-blowing model you have to try it's worth the five dollars whatever 11 labs asked for it in the beginning i think it was free for a week to just try it out wolf and we didn't do german but i would love to hear from you its performance on german it's just absolutely mind-blowing here is nistan speaking spanish by the way it's in el techo in the minute for the five clones it's definitely fast i don't know how fast it is now but definitely decently fast i'm not gonna show this on stage but because i don't want to show off my kids but this is this is my kid here's her counting to 10 in english hi my name is emma and i love doing projects this is a dubbing of emma because she said this in the original russian language this is what she said and just it copies the intonation yes it's amazing dude it's great you know what it copies as well you know how we like we started before the the word like it copies the stars in a different language now the the craziest thing again i edited down a video of this and at the end of this video this is what i show let me show you um this is the thing that i did i did the the the very short sentence in five languages so now i'm gonna dub the next sentence that you'll hear in cinco idiomas okay okay there's nothing over here in germany when a new movie comes out it is completely translated in german so real professional dubbing people are doing this and that there's basically no need anymore and everybody can have their original voices just into the language you are speaking amazing very interesting so we're glazing but we have a new commenter here said a lesson test 11 labs dubbing in latvian to english while the speaker similarity was a promise it did revert to latvian halfway through leaving 64 of the audio untranslated that's very interesting i haven't tested it on long sequence i did test it on short ones but this is just the one of the things that i wanted to show you as a super fast listen what's your comments on you speaking hebrew and then we can close out the show i think we've been on for two and a half hours yeah but gracias signori signori un alta bellissima semana for thursday i that's not too bad next show nisto will speak hebrew original original proper hebrew folks i want to just make sure that we ran through pretty much everything the only thing is the nematron asr we weren't able to test and show you what a show today folks we had chris alexiuk from nvidia talk to us about the biggest release in open source for a while from nvidia nematron and rtx laptops and the nvidia 3.5 asr we had peter gostev dropped new things on arena ai the agent arena that we now looked into and saw the grog 4.3 is not that great at fixing itself in agentic we also had karan the co-founder of new switch to talk about the hermys and success the thing that we did not speak about was weights and biases and and core weave so we'll definitely do that folks even without the transition this weekend there's a hackathon and if you're in san francisco please join us at lulu.my slash weave hacks we are doing a hackathon together with open ai and cursor and we have over 150 dollars for off credits for you to build very cool things as you heard before nissen who just dropped this one hackathon elsewhere winning hackathon gives you great chances to go online also wolfram mentioned this before but i really want to show this super quick because i think it's very important wolfbench has a new feature that i want you guys to see wolfram you want to explain this really quick why this is are you showing it okay yes what i if you scroll a little bit down so with the model name we see that on the second spot basically the second best model is gemini 3.5 flash on this benchmark which is a lot of benchmarks here the terminal bench and so on but so they are they look very close and they are very close in performance but something these bars never show is how many tokens did it use to get that score so we have now a toggle at the top where you can turn a 3d mode on and now we see the depth of the bar it's actually how many tokens it used so one pixel is a couple thousand or even million of tokens and if you hover over one of the of the bars then you see yeah just move the mouse on top of it yeah and then keep it still and then it will do a pop-up it should do one if you hover over it and it shows you how many tokens so it was three over three billion tokens for the flash model compared to just a couple of 150 or something for the gbt 5.5 so that is something to keep in mind you may use a cheaper and faster model but it still has to generate a lot more tokens to get to the same capability level and you have to really do the calculations to see if the time it takes to generate all the tokens and the money it costs if that is yeah it's still a good deal or if you should this is incredible this is incredible specifically what i'm seeing now is that gemini 2.5 flash is number two on the agentic scores overall but the amount of tokens it uses this doesn't show it on the share screen so i'll just read it out tokens in is 3 billion out is 15.7 million 3 billion tokens yeah that's the input tokens basically so the input is usually much more so normally it's 99 input tokens and that can happen for various reasons one of the reasons could be that the agent made a mistake and ingested a large file and the agent harness is not restricting that that would be one of the reasons so i will definitely even write a blog post about this on our rights advices block what is happening in the background but that way you can really see if you were running the agent that way you would have that kind of tokens and the bill would also be reflecting that yeah so definitely you can see that gpt 5.5 is the most incredible agentic model compared to it gets a score without using that many tools and cursor agent as well well thanks so much for this folks this has been thursday it's been a pleasure we've been on the air for more than two and a half hours we had incredible guests and incredible news week incredible packed news week if you missed any part of the show is available to you everywhere that you get your podcast it's available to you on spotify on apple on youtube and edited version of the show calls to youtube as well i wanna just send a huge thanks to you the listeners who show up and comment we really appreciate the comments from you guys as well today we did not get the update from open ai today that some folks were meaning that maybe it's gonna come i'm very happy that we didn't because we're already two and a half hours in we didn't cover everything but if you like this and if this show is valuable to you and if you like the energy and the guests and the type of stuff that we talked about and the depth of technical stuff that we go into please give us a subscribe everywhere that you get your podcasts give us five star review with over almost 2 000 folks who are tuning into the stream this is alex walk of the adventures with weights and biases core with wolf from revive andrews who'd wait to buy the score with jan pelek and ldj we had niston here before and other three great co-hosts thank you so much for tuning in we're signing out we'll see you next week bye bye everyone thanks alex for hosting