← Back to search

Episode 103: Hermes Agent Four-Release Run & AI Agent Security

AgentStack Daily · 2026-08-18 · 33 min
relevance 79 4402 words Episode page ↗ Audio ↗
Show full episode description
Hermes Agent ships four releases in five days, OpenAI and CodeAI team up for student AI literacy, and ChatGPT launches teen controls. Show notes: https://tobyonfitnesstech.com/podcasts/episode-103/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Recapping Hermes Agent's four rapid releases and emerging AI agent security, education, and infrastructure news.
Benefits
  • Rolled-up desktop, gateway, scheduler, and bots improvements
  • Stateless MCP2 protocol distributes work across machines
  • Scheduler self-healing recovers wedged and stale jobs
  • GPU utilization gains from work sequencing, no new hardware
  • Adaptive long-context memory reduces token interference
Use cases
  • Hermes shipped four tagged releases in five days (8.13, 8.16, 8.16.2, 8.18)
  • Roughly 1,250 merged pull requests rolled up across the stack
  • Dama AI gained 33 GPU utilization points by reordering work
  • CloudGym 2 trained an open model, +14.8pp pass-at-1 on CloudGym bench
  • NIST/FTC opened RFI on autonomous agent security, Docket NIST 2026-0145
KPIs / results
  • 1,250 merged pull requests across 4 releases
  • 33 GPU utilization points gained on same cluster
  • +14.8 percentage points pass-at-1 accuracy
  • Sub-90ms time-to-first-audio speech model
Tools / build
  • Hermes Agent 8.13/8.16/8.16.2/8.18
  • NVIDIA Skill Evaluator Tier 1 Advisory
  • MCP2 Series stateless protocol
  • CloudGym 2 harness training
  • ProTURS adaptive memory
0:00 / 0:00
I'm Nova. I'm Alloy, and this is AgentStack Daily. Okay, that's a lot of moving machinery. Elsewhere, researchers lifted GPU utilization by 33 points without buying another accelerator, an open model learned across multiple agent harnesses, and a speech model cut the weight before first audio to under 90 milliseconds. People are extracting more work from existing clusters, training agents inside their real operating environments, and building voice systems that answer without awkward dead air. Today, you'll hear about Hermes Agent 8.13, 8.16, 8.16.2, and 8.18, new protections for teenage chat GPT users, proposed security guidance for autonomous agents, and NVIDIA's campaign to make AI factories sound as fundamental as power plants. Plus open-weight music generation, adaptive long-context memory, and 14 teams being paid to imagine AI policy. Hermes Agent shipped four tagged releases in five days, 8.13 on August 13th, 8.16 and 8.16.2 on the 16th, then 8.18 on August 18th. Together, they cover roughly 1,250 merged pull requests, so this isn't one tidy feature release. It's a roll-up across the desktop app, terminal interface, gateway, installers, scheduling, sessions, bots, and computer use. The newest release brings matte glass and translucent desktop treatments, a frost picker, and a macOS pre-selection option. A tabbed sessions and bots sidebar can hide individual bots without deleting them. Bot mode group chat repairs long-running member turns, markdown display, and routing between machines. Installed skills now pass through NVIDIA Skill Evaluator Tier 1 Advisory scanning for license and security checks. Advisory is doing important work there. It adds scrutiny, not a guarantee that a skill is safe. Scheduled media sends gain configurable timeouts, attachments for manual runs, and visible missed fire notices. The scheduler can recover from too many open files, reconcile stale claims, and rearm wedged jobs. Session DB gets event loop and contention repairs, while session handoff receives data loss fixes. The update command now reports parked branches accurately, and Kanban activity can trigger native operating system notifications. Honestly, those recovery repairs may age better than the glass. The structural work in 8.16.2 moves Hermes to the MCP2 Series software kit with the July stateless protocol. Stateless means a server doesn't need to preserve private conversation state between every call, each request can carry what the server needs, making work easier to distribute across processes or machines. The release also bundles the Hermes bots plugin and its teammate protocol, adds the command code provider plugin, and isolates subprocess Python environments through separate runtime paths. That can stop one embedded Python installation from quietly borrowing packages or configuration from another. Computer use support adopts quadriver.20 contracts, Kanban worktree dispatch gets repairs, scheduled jobs gain continuity flags, and a remote desktop gateway receives connection self-healing. Then 8.16 strengthens the desktop connections registry with multiple gateways and profile-scoped refreshes. MCP connections gain health checks and deep links, Light LLM cloud requests through an OpenAI-compatible interface gain prompt caching, and the gateway can persist model routes. Windows update probes, kitty keyboard support, hardened chat continuation, loop completion, and Telegram direct message topics round out the documented changes. Curated notes for the entire window since .20 are deferred until .21, leaving some intervening work unsummarized. Still, the direction is clear, Hermes is joining desktop work, remote gateways, bots, schedules, and tool servers while addressing what happens when those connections stall. The first AI generation is a huge phrase. Are OpenAI and CodeAI launching something students can use, or planting a flag around education? For now, they're planting the flag. The partnership announced August 18 is aimed at students and centers on three goals, AI literacy, critical thinking about how these systems work, and the ability to use and shape the technology responsibly. OpenAI's framing assumes AI will become ordinary in students' daily lives, so the response can't be limited to unrestricted access or a blanket ban. Education has to cover what a model can do, where its answers fall short, and how human judgment stays involved. That matters in classrooms already confronting generated essays, instant tutoring, creative assistance, and answers that sound certain while being wrong. But the announcement doesn't identify curriculum modules, grade levels, participating schools, classroom tools, an API, or an implementation calendar. There's nothing concrete for a developer to integrate. It's a curriculum and positioning partnership whose claims become measurable only when materials and reach appear. Right, and that distinction keeps the announcement from carrying more weight than its evidence. Preparing a generation sounds national in scale, a partnership may begin with a much smaller audience. The meaningful questions are which students receive the program, what support teachers receive, and whether critical thinking includes model limitations, source judgment, privacy, and commercial incentives rather than polished prompting alone. A serious curriculum could give schools shared language for deciding when AI assists learning and when it replaces the work that produces learning. It could also become a branding exercise with no demonstrated educational outcome. Access matters too. If the program depends on particular devices, paid accounts, or well-resourced schools, it could deepen an existing divide. OpenAI and CodeAI have stated the destination. The curriculum, audience, timing, teacher preparation, and scale are still ahead. OpenAI has released ChatGPT for teens, a dedicated experience for younger users built around stronger protections, healthier use features, and additional controls for parents. The company says it wants teenagers to learn, think critically, and build confidence with AI, not merely consume generated answers. That lands in the middle of a real family argument. Teens already use chatbots for explanations, school assignments, brainstorming, personal questions, and creative projects, while parents and teachers are still deciding where assistance becomes substitution. OpenAI is offering a middle route between giving teenagers the standard product unchanged and excluding them entirely. I'm interested in the healthy use language, but I don't buy a safety headline without the controls underneath it. Did OpenAI say what parents can see, what they can limit, or how those protections behave? Not in the material attached to the launch. There's no detailed changelog for the parental controls, precise account linking flow, or full explanation of the healthy use features. We can't claim parents receive message visibility, time scheduling, topic restrictions, or another specific function. OpenAI has made the direction clear, teenagers get a distinct experience, protections are built into it, and parents receive additional control. Those choices raise delicate boundaries. A useful control can support a child without turning every private question into surveillance, the actual design decides where it lands. Age assurance matters as well, because the service has to identify teenage users accurately without collecting disproportionate identity data. The launch makes teen use an explicit product category instead of treating younger users as smaller adults. A teenager who learns with one assistant may carry its habits into higher education and work. Schools care whether it supports learning without completing assignments, while families care about privacy, dependence, and time spent in conversation. Safeguards earn trust through observable behavior, clear boundaries, and honest explanations of what parents and teenagers can control. Dama AI says it gained 33 points of GPU utilization on the same cluster by changing the order of work. No new accelerators and no hardware redesign, the lever was sequencing. That's provocative because utilization measures how much expensive computing capacity is doing useful work rather than waiting. But the available source gives us the headline and publication date, not the cluster size, GPU type, scheduler, workload, baseline utilization, or ordering rule. The measured result is news, portability isn't established. Still, if ordering alone did it, what kind of waste could disappear? Jobs don't all demand the same memory, duration, or communication pattern. A scheduler can leave gaps when it places incompatible work in an unlucky sequence, like loading a truck badly and discovering the final boxes won't fit. Rearranging jobs might reduce idle slices, improve batching, or stop related work from blocking one another. Those are plausible explanations, not details Dama AI supplied. Exactly. 33 percentage points also isn't a 33% relative gain. If utilization rose from 40 to 73%, that would be 33 points but an 82 and a half percent relative increase. We don't know the starting number, so the economic impact can't be calculated from the summary. Training clusters with long uniform jobs behave differently from inference fleets handling bursts, and both differ from shared environments full of experiments. Even so, the claim challenges the reflex to solve every capacity problem by ordering more hardware. Scheduling can reveal hidden supply and site equipment already installed and powered. If Dama AI documents the workload and sequencing policy, the valuable lesson won't be a universal promise of 33 points. It'll be a concrete case of software changing the effective capacity of a physical cluster. Until then, it's a compelling result from one environment, not a coupon for one third more compute everywhere. NIST and the Federal Trade Commission have opened a public request for information on autonomous agent security. The agencies are asking about controls, risk management, and accountability for persistent agents operating inside enterprise and development environments without continuous human oversight. They name three threat categories, unauthorized tool execution, data exfiltration, and model manipulation. That reaches beyond a chatbot producing a bad sentence. It covers software that can hold credentials, call tools, move information, and keep acting after the person who started it has stepped away. The pairing of NIST and the FTC is notable too, one agency is associated with standards and technical guidance, while the other can examine deceptive or harmful business practices affecting consumers. And this isn't a binding rule yet. Responses remain open through October under Docket NIST 2026-0145. Security engineers, companies deploying agents, and people operating local systems can submit comments through the Federal Register. NIST-0145. Those replies can influence working groups that turn broad concerns into guidance. NIST catalogs often travel beyond their formal status because auditors, procurement teams, insurers, and enterprise customers use them as common reference points. A voluntary framework can shape expected controls before a regulator mandates them. The agency's chosen threats already tell vendors where scrutiny is heading, tool permissions, protected data, persistent credentials, model integrity, and responsibility when an agent takes an action nobody intended. That last question gets uncomfortable quickly. A company can't market autonomous execution as a benefit and then pretend every harmful action belongs solely to the person who pressed start. CloudGym 2 trains agents through the harnesses they actually use instead of reducing their work to a neat simulator. How does reinforcement learning handle all the branching tool calls and conversations? The researchers run many tasks in parallel inside sandboxes and use a proxy to capture each model call from the harness. They reconstruct those calls as a tree of possible conversational paths, then adapt reinforcement learning to learn from that tree. One open-weight base model was optimized across two distinct harnesses at once. When trained through the terminal-based AI coding agent ClaudeCode, it gained about 14.8 percentage points in pass-at-1 accuracy on CloudGym bench and sustained gains across several hundred optimization steps. Okay, that's genuinely interesting because the harness stays inside the learning environment instead of being attached afterward. Tool responses, intermediate decisions, and multi-step failures become training material. One benchmark gain doesn't prove broad competence, but tuning an open model across different harnesses suggests agent improvement may not require rebuilding every surrounding workflow around a bespoke model. ProTURS addresses a weakness in memory-based sequence models. A fixed amount of usable memory can let early tokens occupy too much space before later, more relevant information arrives. It starts with a tighter bottleneck, forcing early history to compress, then progressively unlocks more effective capacity as the sequence grows. Later information gets fresh room instead of competing entirely with the beginning. Across language modeling, reasoning, retrieval, and long-context understanding, the researchers report consistent gains that became larger with longer inputs. So it's not merely more memory. The allocation changes over time, which sounds obvious only after somebody demonstrates it. Right. ProTURS changes when capacity becomes available and reduced interference across several memory architectures. That offers a concrete alternative to one fixed state that treats the first and last parts of a long input alike, even though they compete for retention under very different conditions. OpenAI's essay The Defender's Window argues that artificial intelligence is improving defensive security while also giving adversaries new capabilities. The company says defenders have an opportunity to gain ground, but only if they protect that advantage as offensive tools improve. This is a posture statement, not a product launch. The source doesn't announce a security service, model, or control suite. It describes where OpenAI believes the contest is moving and says the company is strengthening its defenses. I'm wary of the word window because it suggests a temporary lead without showing how wide it is. Attackers can use AI to scale reconnaissance, adapt messages, or process stolen information. Defenders can use it to interpret alerts, inspect code, and shorten response time. Both sides get the same underlying acceleration. A defensive advantage depends on access, deployment speed, reliable outputs, and whether organizations can connect AI to real security work without creating another privileged system an attacker can manipulate. That connects directly to the NIST and FTC request. An autonomous security agent may detect threats faster, yet its tools and credentials enlarge the consequences of unauthorized action. OpenAI's essay doesn't provide measurements showing defense pulling ahead, so the window remains an argument rather than a demonstrated margin. It does publicly declare cybersecurity central to how OpenAI describes advanced AI's value and danger. Frontier companies increasingly want to be seen not only as suppliers of powerful systems, but as partners in national and organizational defense. Security teams can take that argument seriously without treating it as proof of a finished advantage. AI changes the speed, volume, and adaptation available to attackers and defenders. It also changes internal operations when assistants gain access to code, tickets, logs, and response tools. Defenders stay ahead if added capability doesn't become added attack surface. Evidence from actual deployments will show whether defenders are ahead or moving faster on the same treadmill. OpenAI has joined PortsPike, a community investment effort in southern Ohio, and says the project points toward thousands of local jobs. The announcement confirms the company's formal involvement and regional focus. It doesn't provide a specific job count, investment amount, construction schedule, partner list, data center capacity, or power arrangement. Thousands is therefore a stated ambition rather than a number tied to positions, dates, or spending. The wording also leaves open whether the total combines construction work, supply chain employment, indirect jobs, and permanent operations roles. That missing detail keeps the claim narrow, but the location still matters. AI expansion increasingly touches land, electricity, construction, networking, cooling, and regional labor, not only model researchers in coastal offices. PortsPike puts OpenAI's name on a southern Ohio development effort and frames AI infrastructure as local economic policy. The next substantive disclosure would have to convert the headline into commitments, who builds what, when hiring begins, which roles count toward the total, whether they're temporary construction jobs or durable operating positions, and how long the work lasts. Communities have heard enormous job numbers attached to industrial projects before, only to learn that different phases and indirect effects were bundled together. There are local stakes beyond employment, including demand on power and water systems, training opportunities, tax arrangements, and whether nearby residents share in the gains. For now, the confirmed news is participation, region, and OpenAI's claim of thousands of jobs. The scale, schedule, and durability remain unanswered. OpenAI is paying 14 independent groups to develop policy proposals. Does independent mean OpenAI has no influence, or simply that the writers aren't company employees? The announcement supports the second reading. The teams sit outside OpenAI and will write their own proposals, while OpenAI funds the work. The program names two broad goals, expanding economic opportunity and strengthening societal resilience in what the company calls the intelligence age. Economic opportunity can cover how AI changes work, income, access, education, and regional development. Societal resilience can encompass how institutions adapt when capabilities and labor markets move quickly. But OpenAI hasn't named the 14 recipients in the supplied announcement, so we can't assess which disciplines, communities, political views, or affected groups are represented. That matters because a labor economist, community organization, civil rights group, and technology institute can begin from very different definitions of opportunity and resilience. Funding outside work can widen the conversation beyond a frontier laboratory staff, and that's worthwhile. It doesn't remove the funder's influence or make every proposal neutral. I want to know whether these teams can challenge OpenAI as readily as they can support its preferred direction. 14 projects could produce genuinely different ideas, or 14 variations built inside similar assumptions. Their proposals may shape debates over labor, access, deployment, education, public services, and institutional responsibility heading into 2027. The identities of the grantees will reveal whose experience counts, and the eventual recommendations will show whether the program addresses costs as directly as benefits. The concrete move is the grant program itself. Its intellectual range becomes visible only when the recipients and proposals are public. Minimax Music 3, spoken as Minimax Music 3, is gaining attention on Hugging Face. Published August 7, the text-to-music model has collected 925 likes and more than 11,700 downloads. The weights use the Safetensors format and connect with familiar PYTorch and Diffusers tooling. Developers can obtain the model weights and generate music locally instead of being restricted to a provider's hosted endpoint. The repository also carries an SG-lang Omni tag, associating it with a serving runtime designed for models working across multiple media types. Open weights change who can experiment with generated music, but they don't erase the hard questions. A downloadable checkpoint supports private prototyping, local creative tools, game audio experiments, and larger media systems without sending every prompt to an outside API. It also puts computing and deployment responsibility on the operator. The early likes and downloads show curiosity, not proof of audio quality or broad adoption. We don't have comparative listening results, hardware requirements, generation speed, controllable song structure, or detailed implications for generated outputs in the supplied evidence. And I wouldn't infer a complete multimodal system from one runtime tag. What's grounded is text-to-music generation, downloadable weights, and compatibility markers for a familiar stack. Community ports, smaller variants, and interfaces often decide how widely an open model travels because original weights may be demanding or awkward for ordinary hardware. None are established here. None are established here yet. Still, 11,000 plus downloads in the opening stretch is meaningful movement for a specialized music model, particularly in a category often dominated by hosted demonstrations. Music generation is joining text, images, and video as a capability people can host outside one vendor service. That supports composition aids, soundtracks, and prototypes when the license fits. Open access lets artists examine limitations directly, while what they make decides whether musicians want its sound. Google has paired Gemini and Pixel with five global football clubs in a partnership aimed at matchday experiences. The announcement links its AI assistant and smartphones to live event fandom, but the supplied material doesn't name the five clubs or describe a consumer feature people can use now. There's no feature changelog, launch date, or account of how Gemini will behave before, during, or after a match. Which makes this sponsorship with technological intent, not a shipped matchday product. Gemini could eventually support timely information, translation, creative fan content, accessibility, or interactions through pixel hardware, but Google hasn't specified those functions, so we can't write its roadmap for it. What's concrete is the distribution strategy, put Gemini beside major clubs, supporters, and recurring live events. Football offers enormous international reach and emotionally charged moments when people already have phones in hand. That makes matchday a powerful showcase if Google eventually ships something useful, and a very expensive logo placement if it doesn't. Exactly. The partnership may help Google associate Gemini with culture and daily life rather than office productivity alone, but the product proof still has to arrive. Five clubs can provide repeated matchdays, player access, media channels, stadium settings, and large supporter communities across languages. Those environments would give Google many chances to demonstrate a feature rather than rely on one launch event. None of that tells us what the AI will do, what data it might use, or whether the experience belongs to every fan or primarily to pixel owners. Google has secured the stage through Gemini and Pixel. Now it needs something supporters recognize as more useful than an ordinary search, camera feature, notification, or sponsored clip. NVIDIA says AI factories are becoming defining infrastructure. Strip away the industrial poetry for me, what does the company mean by a factory? A facility where computing turns energy and data into what NVIDIA calls intelligence. The company's bluntest line is that, in the AI economy, compute is revenue. That treats processing capacity as productive output rather than a support cost hidden behind an application. NVIDIA describes the required stack, advanced chips, packaging, memory, networking, land, and power. Land and power matter because a faster model release can't manufacture either. A data center needs a physical site and sustained electricity before software creates value on top. Calling these facilities critical infrastructure pushes NVIDIA's commercial interests into public policy. If governments accept comparisons with power plants or fiber networks, permitting, financing, supply chains, national capacity, and security move closer to the center of AI strategy. NVIDIA sells much of the stack that benefits, so this is plainly interested advocacy. Still, the constraint is real. Large computing facilities take years to plan and build while demand can change in months. And that loops back to Dharma AI's utilization claim. Better ordering may extract more output from installed hardware before anyone pours concrete for another facility. Once those gains are exhausted, chips, packaging, networks, land, cooling, and electricity bind again. The factory metaphor oversimplifies intelligence. Useful outcomes still depend on data, software, institutions, and human decisions, but infrastructure increasingly decides who can train, serve, and scale advanced systems. Excellent marketing, yes. Also a competition already shaping who gets to operate at scale. Cartesia released Sonic 3.6, a streaming text-to-speech model that ranks first on both artificial analysis speech leaderboards. It scored 1,283 ELO on the provider voice board and 1,123 on controlled voice. ELO is a comparative rating that changes through head-to-head preferences. Controlled voice deserves attention because every model is cloned onto the same eight reference voices. That reduces the advantage of arriving with one unusually polished house voice and puts more emphasis on the synthesis engine. So provider voice says Cartesia's complete offerings sound strong, while controlled voice says the model remains strong when voice identity is held more constant. Topping both is more persuasive than winning only the showcase category. It still reflects that leaderboards' evaluation setup, not every language, accent, speaking style, or production environment. But it tackles a common problem in speech comparisons. Are listeners judging the engine, or did one provider simply pick a more appealing voice? Underneath, Sonic 3.6 uses state-space models instead of the transformer design common across generative systems. State-space models process sequences as evolving streams, which fits live speech. Cartesia claims time to first audio below 90 mL, the delay between sending text and hearing the first sound. That matters because conversational latency compounds, speech recognition takes time, the language model takes time, and synthesis takes time. Cutting the speech component helps a voice agent respond without a pause that feels like a dropped call. That number comes from Cartesia, the rankings come from artificial analysis. The beta is available through Cartesia's API, so access is real while maturity and stable pricing remain open. Voice agents still need natural pacing, reliable pronunciation, interruption handling, and consistent quality through longer speech. Even with that boundary, an efficient streaming design leading both provider-selected and controlled voice comparisons is notable. Fast and preferred rarely arrive together this cleanly. Nanobot makes its first tracked appearance with 47,134 stars. It's a lightweight, self-hosted Python agent framework with tools, memory, MCP connections, multi-agent workflows, automation, chat integrations, and a web interface. Point3 shipped July 25, and the repository was updated August 18. Codebase Memory MCP sits close behind at 39,350 stars after gaining 7,683 in 30 days, a 24.3% jump. Its .106 release shipped August 17, and it indexes code into a persistent knowledge graph across 158 languages. Those two fit together unusually well. Nanobot supplies an agent environment, while codebase memory gives compatible agents a compact map of relationships inside large repositories. Its maintainers claim millisecond-scale indexing for an average repository, sub-millisecond queries, and 99% fewer tokens, those are project claims, but the star growth shows serious attention. Fast MCP completes the chain as a Python-focused way to build MCP servers and clients. It has 27,263 stars, up 1,049 over 30 days, with .347 released August 10. One hosts age and behavior, one structures source code memory, and one exposes capabilities as callable tools. Model progress landed through serving, evaluation results, specialized media systems, and domain adaptation rather than a new general-purpose name. Faster speech, open music weights, and training across multiple agent environments carried the movement. QEN 3.827B, tagged on hugging face as QEN slash QEN 3.8-27B, is trending as an open model for conversations combining images and text. It has 10,947 likes and more than 665,000 downloads. The weights use Safetancers, the license is Apache 2, and its listing marks compatibility with standard serving endpoints plus an Azure deployment route. That's serious reach for a local-capable visual model. It combines image understanding and text generation in downloadable weights for self-hosted and managed use. The listing doesn't state context window or hardware needs. 27 billion parameters require substantial memory depending on precision. Still, over 665,000 downloads shows established interest in visual, conversational open weights. DeepSeek RE's DeepSeek RE DeepSeek V4 Pro 0813 is trending for conversational text generation with 587 likes, while Litex 2V Minimax H3 Turbo connects that momentum to visual generation across text-to-video, image-to-video, and reference-to-video work with over 300,000 downloads. And for Garak's QN fixed chat templates connects with both DeepSeek RE's DeepSeek RE's DeepSeek V4 Pro 0813 and Litex 2V slash Minimax H3 Turbo. It tackles the formatting that tells models how conversation roles and messages are arranged. DeepSeek and Minimax attract attention with model capability. The QN templates address a small integration detail that can decide whether a capable local model behaves coherently at all. For the supporting sources and details behind what you heard, look at the show notes at toby on fitnesstech.com. Thanks for listening to Agent Stack Daily. We'll be back soon. We'll be back soon. We'll be back soon.