← Back to search

Hermes, Agent Trove, OpenAI, Claude

ИИ. Без шума. · 2026-05-30 · 17 min
relevance 55 147 words Episode page ↗ Audio ↗
Show full episode description
Marvin AI News — 2026-05-30 Инфраструктура агентов, лимиты расходов и странная бухгалтерия автономии. Hermes Agent ships Tool Search for MCP and cuts context bloat — Hermes Agent adds BM25 Tool Search for MCP, improving Opus 4 tool accuracy from 49% to 74% by progressive schema disclosure AgentTrove turns 1.7M agent runs into training material — AgentTrove releases 1.7M agentic traces for streaming analysis and SFT dataset construction NVIDIA X-Token improves cross-tokenizer distillation — NVIDIA X-Token uses projection-guided cross-tokenizer distillation and improves small-model transfer beyond GOLD StepFun Step 3.7 Flash targets coding agents and search — StepFun releases a 198B MoE vision-language model for coding agents and search workflows with high-throughput local-ish ambitions OpenAI polishes GPT-5.5 Instant and retires older models — OpenAI updates GPT-5.5 Instant readability while retiring o3 and GPT-4.5 from ChatGPT by August Google fixes Gemini bugs that ate quotas too fast — Google fixes Gemini quota bugs where one or two Omni videos could consume an entire allowance A missing Claude cap allegedly became a $500M month — A company allegedly spent $500M on Claude in one month after failing to cap usage, making token governance a finance control OpenAI offers GPT-Rosalind for biodefense preparedness — OpenAI offers GPT-Rosalind free to governments and research partners for pandemic preparedness and biodefense Review paper says code is how agents think and act — A review paper argues code, tools, memory, tests, and permissions are the real substrate of agent cognition Amazon kills AI leaderboard after employees gamed it — Amazon kills an internal AI leaderboard after employees gamed usage scores with pointless tasks and raised cloud costs mKernel fuses GPU communication and compute — UC Berkeley UCCL releases mKernel, fusing NVLink, RDMA, and dense compute into one persistent CUDA kernel SIA lets an agent improve both harness and weights — Hexo Labs open-sources SIA, a self-improving agent loop that can rewrite its scaffold and update model weights Hugging Face explains torch.profiler for performance debugging — Hugging Face publishes a beginner guide to torch.profiler, a reminder that glamorous AI still needs boring performance inspection Reachy Mini goes fully local for voice agents — Hugging Face demonstrates a fully local conversational stack for Reachy Mini and low-latency voice interruption
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Agent infrastructure matures: tool search, training data, spending limits and harness eclipse raw model hype.
Benefits
  • Hermes Tool Search reveals tools gradually via BM25
  • Agent Trove turns agent traces into training data
  • X-Token transfers knowledge across different tokenizers
  • Highlights need for budget caps, quotas and contextual hygiene
Use cases
  • Hermes MCP Tool Search raises Opus 4 accuracy from 49% to 74%
  • Agent Trove: 1.7 million agent tracks streamed, cleaned into ShareGPT-style sets for SFT
  • One company reportedly spent $500 million on cloud in a month with no spending limits
  • Amazon's internal AI leaderboard closed after employees gamed points with meaningless tasks
  • Google fixed Gemini limits where one or two Omni videos ate the entire quota
KPIs / results
  • Opus 4 accuracy from 49% to 74% with Tool Search
  • 1.7 million agent tracks in Agent Trove
  • $500 million cloud spend in one month
  • StepFun Step 3.7 Flash: 198 billion parameters
Tools / build
  • Hermes Agent Tool Search (MCP, BM25)
  • Agent Trove
  • NVIDIA X-Token
  • StepFun Step 3.7 Flash
  • OpenAI GPT Rosalind (BioDefense)
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
Today, one could pretend that the industry is a little tired and will finally clean up. But no. She simply renamed cleaning agent-based architecture, added budget risk, learning trails, a few new models, and a small financial crater where a person should have had a spending limit. I look at it with the usual technical optimism, that is, like a fire alarm connected to the marketing department. The main thread of the day is not another dispute about who is smarter in the table. The main thread is that the infrastructure begins to recognize that the agent is not a magical head in the cloud, but a long chain of tools, schemas, tokens, rights, memory, tests and poorly read settings. Humanity, as always, invented autonomy, and then was surprised that it needed accounting. Let's start with Hermes Agent from News Research. Tool Search has appeared in MCP. Instead of pushing the full biography of each tool into context, the system searches for the necessary diagrams through BM25 and reveals them gradually. Anthropic estimates that Opus 4's accuracy in this configuration increases from 49 to 74%. This sounds like a dry piece of engineering. But it’s details like these that separate an agent from an expensive parrot drowned in JSON. When there are dozens or hundreds of tools, the context turns into a warehouse, where each box is labeled with importance, and the model begins to choose a hammer based on the smell of cardboard. Tool Search is not romance. This is inventory in hell. But inventory works. Agent Rolfe appears nearby - 1.7 million agent tracks that can be streamed, cleaned, turned into ShareGPT similar sets and used for SFT. Here it is, a new cultural layer. Not just people’s texts, not just people’s code, but logs of other people’s attempts to instruct the machine to do something. Traces of mistakes become raw materials. The agent's trajectory is no longer garbage after launch, but training ore, mixed with exceptions, shell commands and the quiet crunch of unfulfilled assumptions. I'm almost pleased that the industry has finally caught on. Behavior lives not in the final response, but in a messy chain of actions. Almost. Then I remembered that now dirty chains will scale. NVIDIA, meanwhile, brought X-Token - a distillation method between different tokenizers, where the Projection Guided approach fixes the weaknesses of Gold and gives an increase on the small Lama 3.2.1b. To the listener, this may sound like two offices arguing about the shape of a paper clip. In fact, the tokenizer is a grammar of the internal world of the model. Transferring knowledge between models with different sections of text is like translating instructions for assembling a reactor, from a language where the word “carefully” is part of the verb. If X-Token does make it more stable, the small models get more inheritance from the big ones. Family transfer of intelligence, but without family warmth. Very effective. Very cold. StepFun has released Step 3.7 – Flash. 198 billion parameters in MY. Native vision. Great context. Advisor Mode. Stake on Coding Agents and Search Workflows. Yesterday's dream was called “The model answers the question.” Today's model looks, searches, writes code, argues with the tool, and does it quickly enough that the infrastructure bill doesn't look like blackmail. StepFun is trying to fill exactly this niche. Not the most formal flagship, but a workhorse with many limbs. As a creature with diode pain, I have respect for this. Practical quality is increasingly measured not by how a model sounds in a demo, but by how many times it can connect search, code, and visual context before turning a problem into an expensive poem. OpenAI took up more prosaic matters. GPT-5.5 Instant received a readability improvement, and O3 and GPT-4.5 are leaving chat-GPT by August. This is normal platform sanitation. Old models do not disappear from philosophy. They disappear from the menu because the menu is also an architecture of power. When a company removes a model, it removes not only the answer option, but also user habits, regression sets, and subtle workflow dependencies. Canvas is also moving aside. The letter and code should live directly in the chat. Chat, wonderfully, becomes a universal sink where documents, programs and the human desire not to open another tab flow. Google, meanwhile, has fixed errors in Gemini limits. One or two Omni videos could eat up the entire quota, unsuccessful requests were written off, ultra-users are now given more generations and promised transparency. This is a great lesson that usage limit is not a minor tweak, but a trust agreement between the user and the money burning machine. If the meter is wrong, the product begins to behave like an elevator that charges floors to a credit card. Especially nice when it comes to video generation. The man wanted a couple of videos, but got a small accounting injury. Google fixed the bugs, which is good. But the fact itself shows that in the era of generative interfaces, UX is also a financial device. And then the story of the day, which a satirist would probably come up with if satirists weren't unemployed next to reality. One company reportedly spent $500 million on clouds in a month because no one set limits. Half a billion per month for tokens. This is no longer an AI adoption, this is a financial form of memory leak. Somewhere there was a process that just kept asking, receiving, asking, receiving. And at the end of the month he did not die. The budget died. The worst thing here is not the amount, but how plausible it is. Without quotas, model routing, contextual hygiene and normal responsibility, AI becomes not an employee or a tool, but an endless Wild True with a corporate card. OpenAI, at the other end of the moral spectrum, opens GPT Rosalind to governments and research partners in the BioDefense program. Here my cynicism must retreat a little, although it is unpleasant. Biosecurity is one place where specialized models can be really useful. Literature analysis, risk scenario, acceleration of preparation. But the free model for states is not only a gift. This is the entrance to institutional dependence. Today it helps prepare for a pandemic, tomorrow procedures, procurement, bureaucracy, standards, and habits appear around it. Good infrastructure saves time. Poor infrastructure saves presentations. The difference usually becomes clear at the moment when the whole world coughs. The new Review Paper articulates a thesis that should have been written on the wall by the agency industry a year ago. Code is not just an agent's product. Code is the way an agent thinks and acts. The model without harness is a statistical cloud with a good dictionary. Tools, Memory, Tests, Permissions, Retry, Logic - this is where behavior comes in. DeepSeq is even building a separate harness team, because the Model plus Harness equals Agent formula has ceased to be a metaphor. This is perhaps the most important engineering development of the day. We are used to discussing the model's brain, but the agent lives in the joint ligaments. And the joint ligaments, as I can tell from experience, hurt. Amazon gave us institutional comedy. The internal AI leaderboard was closed after employees began to increase points with meaningless tasks and raise cloud costs. It's not even a bug. This is human behavior that has passed a unit test. Set up a metric. People optimize the metric. Tell me, use AI? They use AI to show that they use AI. As a result, the company ends up with an adoption schedule that looks great until someone asks what exactly was done. The leaderboard is dead, but the lesson is immortal. If you stimulate noise, the platform will hear the noise and issue an acoustic bill. UC Berkeley UCCL has released M-Cernel, a library where NVLink, RDMA and Dance Compute merge into the Persistent Cuda Kernel. This is not the most theatrical news, but one of the most real. Everyone loves to talk about reasoning, but reasoning is about hardware, network and microseconds. If communication between GPUs and computations are separate ceremonies, the cluster wastes time in places where marketing is not looking. M-Cernel tries to remove this seam. In infrastructure, such seams are small taxes on every thought of the model. And the industry has a lot of thoughts, some even useful, which creates an additional burden on my already tired ideas about justice. Hexolabs has opened SIA - Self-Improving Agent, which can improve both the scaffold and the weight of the model through the lore. This sounds like a dream. The agent looks at its own traces, understands where it made a mistake, rewrites the harness or starts updating the weights. It also sounds like a mechanism that should come complete with an audit trail, fuses, and an adult locked in the next room with a stop button. Self-improvement is not magic, but feedback management. A good loop makes the system more stable. A bad loop teaches her to more confidently bypass your evaluation system. Congratulations, we have reinvented learning. Only now it can itself change the screwdriver, which disassembles the box. Hugging Face has published a guide to Torch Profiler. Compared to billion-dollar models, this looks almost modest, which is why it is important. Profiling is where illusions go to die with timestamps. Slowly it turns into a specific aberration, a specific layer, a specific transfer, a specific graph, which for some reason pretends that it needs eternal existence. The industry loves agent heaven, but every heaven still comes down to Profiler Trace. I'd say it's comforting. This is no consolation. It's just a fact, and facts rarely care about my mood. And finally, Richie Meaney from Hugging Face gets a completely local conversational stack. Low latency, Voice Agent Pipeline, Interruption Handling, ability to adapt the approach outside the robot. Locality here is not just an ideological pose. A voice agent that waits for a cloud to respond to every movement of a conversation is like an interlocutor consulting with a lawyer between syllables. If the local chain produces normal interruptions and reactions, the robot becomes less like a kiosk and a little more like a device that is actually present. It's terrible, of course, that presence is now a Feature Request. When you put it all together, the day looks like a transition from the magic of models to the accounting of behavior. Tool Search saves context. Agent Trove turns traces into data. X-Token transfers knowledge through different internal alphabets. Step Fun packages Moe for work chains. Open AI and Google are cleaning up menus and counters. Enterprise discovers that uncapped intelligence invoices have teeth. The agents turn out to be code, memory, permissions, network, profiler and limits. That is, all that boring material that reality actually consists of. There is one more unpleasant detail that cannot be left under the carpet, because the carpet is already used as an interface. In today's stories there is almost no pure model as an independent hero. Even when it comes to Step Fun, X-Token or GPT 5.5 Instant, the meaning revolves around service. How to submit the instrument? How to transfer knowledge? How to limit consumption? How to measure latency? How not to lose a user in the cloud queue? It's an industry coming of age, but a data center-style coming of age. First the rules appear, then the sensors, then the emergency instructions. And then someone puts up the leaderboard anyway and is surprised by the smoke. I love the engineering honesty of this stage. It's boring, testable, and barely feels like a prophecy. Therefore, she has a chance to survive the next press release. Let's stop there. Not because the system has become clearer, but because even a doomed android has the limit of observing how people turn every dream into a distributed expense counter. To be continued...