← Back to search
Episode 73: OpenClaw v2026.6.9, Hermes v2026.6.19, Claude Code 2.1.176 Released; Poolside Adds Laguna XS.2 and M.1
AgentStack Daily · 2026-06-21 · 24 min
Show full episode description
Today's AgentStack Daily: OpenClaw v2026.6.9, Hermes Agent v2026.6.19, and Claude Code CLI 2.1.176 all shipped new releases. Poolside released Laguna XS.2 on OpenRouter and Laguna M.1 via API. Enterprise teams got new usage analytics and updated spend controls. A retrospective asks whether 30 years of export controls can contain a model called Mythos. Baseten is reportedly raising $1.5B at a $13B valuation. Datasette Apps launched for hosting custom HTML inside Datasette. Meredith Whittaker of Signal says AI chatbots are not your friends. In the Weights debuts as a vanity search engine for AI. Show notes: https://tobyonfitnesstech.com/podcasts/episode-73/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Builders need to track maturing agent stacks, new coding models, and production cost/security controls in one cycle.
Benefits
- OpenClaw 6.9 richer Telegram delivery and auto plugin approvals
- Hermes 6.19 adds background subagents and iMessage via Photon
- Claude Code 2.1.176 lighter install, faster cold start
- Laguna XS.2 pairs tools and reasoning in one endpoint
- OpenAI enterprise spend controls and usage breakdowns
Use cases
- Multi-file refactor agents using Laguna XS.2's cheaper per-token economics
- Enterprise admins setting configurable spend ceilings on ChatGPT seats
- Dataset Apps hosting custom HTML/JS dashboards beside agent data
- Hermes dashboard full profile builder and revamped Skills Hub browser
KPIs / results
- Laguna XS.2 context window 262,144 tokens
- Laguna M.1 context window ~256,000 tokens
- Basetten ~$1.5B round at $13B valuation
I'm Nova. I'm Alloy, and this is AgentStack Daily. OpenClaw 6.9, Hermes Agent 6.19, and the terminal-based AI coding agent Claude Code 0.176 all shipped stable releases this cycle. OpenClaw 6.9 added richer telegram delivery and tightened codex integration with automatic plugian approvals, while Hermes 6.19 introduced background subagents and expanded channel support to iMessage via Photon. Today you'll hear about those harness updates, the release of Poolside's Laguna coding models on OpenRouter and via API, new enterprise spend controls from OpenAI, and why the president of Signal says your AI chatbot is definitely not your friend. We also look at a billion-dollar raise for Base10, a vanity search engine that scores your presence inside model weights, and the historical failure of export controls as a lens for Anthropics Mythos model. The through line is a focus on the production layer, how we manage the cost of agentic loops, how we secure the models at the center of them, and how the harnesses we use every day are maturing to handle multi-step, multi-channel workloads. Three stable releases landed this cycle and shape how agentic harnesses are being assembled right now. OpenClaw 6.9, published on June 21, ships richer telegram delivery. The channel path now sends rich HTML, preserves rich markdown and sticker paths, and renders progress drafts and command output more faithfully. It also normalizes HTML tables safely and keeps mentions and spooled handlers on the right delivery path. Agent recovery is more dependable in this version, retries, terminal outcomes, and session history repair now keep more interrupted or partial turns moving toward a visible final result. The codex integration in OpenClaw is also significantly stronger. It adds automatic Pludgeon approvals, GPT 5.3 Spark OAuth routing, and remote node execution as a dynamic tool. These changes move the harness toward a more reliable app server teardone process. Meanwhile, Hermes Agent 6.19, which the maintainers are calling the reach release, extends Hermes across new channels like iMessage via Photon and the Raft Agent Network. The desktop app gained substantial new capability as well, sub-agents can now run in the background, and the image generation tool learned how to edit. The Hermes dashboard got a full profile builder and the SkillsHub browser was completely rehauled. For those using the terminal-based coding agent OpenAI Codex codex.141.141, these updates provide a more robust back-end. We also saw the terminal-based AI coding agent Claude Codex.176 round out the trio on the stable tag. This release focuses on a lighter install footprint and a faster cold start path for shell-driven sessions. At the API and runtime layer, these changes alter what builders can configure and rely on by default. The practical implication is cleaner harness plumbing. OpenClaw gives builders more model routes and safer auth, while Hermes and Claude Codex refine the user and developer experience. The part to watch is how these new defaults behave under real workloads before flipping them to production, especially where session history and terminal outcomes are involved. Poolside pushed Laguna XS.2 to OpenRouter this week, marking the second-generation model in their XS size class. This is their efficient coding agent series, and the pitch here is compactness over flagship reasoning. It pairs tool calling with reasoning in a small footprint, featuring a context window of 262,144 tokens. This context size puts it squarely in the long context tier for agentic coding workloads where passing large codebases or tool traces is the norm. The single most relevant detail for builders is the unified tool and reasoning pairing. One endpoint exposes both tool invocation and reasoning, which simplifies agent loop orchestration. You don't have to route requests across two separate models to handle planning and execution. The XS size class signals a target audience focused on low latency, cost-sensitive runs. If you are running multi-file refactor agents, the Potokan economics here might be more attractive than flagship class reasoning models. For agent builders, this release widens the compact tier on the router. It sits alongside other compact picks used for cheap classification or initial planning passes. What we need to watch next are the actual latency benchmarks on real coding traces. We also want to see if Poolside eventually ships tool calling specific rate limits or pricing tiers to further differentiate the XS line from their larger models. This refresh of the compact tier matters because it moves the cost frontier rather than just the raw quality frontier. If your agent stack relies on many small, iterative calls, a second-gen compact model with a massive context window is a significant integration target. It allows for more complex multi-step workflows without the latency hit of a larger flagship model. Moving to the other end of the spectrum, Poolside has also released Laguna M.1 via API. This is their flagship coding agent model, optimized for complex software engineering tasks. Like its smaller sibling, it supports tool calling and reasoning with a 256,000 token context window. At the mechanism level, this change shows up in the API surface and the runtime behavior that agent builders integrate against. The Laguna M.1 release is aimed at the heavy lifting of agentic coding. Think about architectural changes, complex bug fixes, and large-scale refactors that require a deep understanding of the entire stack. The primary source documentation for this model includes specific deployment notes and changelog context that builders should review. It is worth tracking how M.1 behaves under heavy production loads, as flagship coding models often face unique challenges with consistency over long traces. Why this matters right now is the speed at which the agent stack is moving. Changes at this layer determine which workflows are reliable and which ones remain brittle. The practical question for builders is whether Laguna M.1 can replace their current flagship default for coding tasks. Early evidence suggests it is a strong contender for those building autonomous or semi-autonomous engineering agents. We should be watching for follow-up releases and independent benchmark results. As surrounding tooling like SDKs and security review frameworks pick up support for M.1, the barrier to switching will drop. For now, the focus is on verifying its performance against the specific coding patterns your agents are expected to handle. OpenAI is introducing new spend controls and usage analytics for chat GPT enterprise. This is a direct response to organizations needing to manage costs and scale their AI deployments with more confidence. These controls land at the management and API level, affecting how admins configure and deploy seats across a large company. It includes more granular cost breakdowns and the ability to set configurable ceilings on spend. For builders and platform engineers, these tools are about visibility. As agents proliferate within a company, knowing exactly which departments or projects are driving usage is critical. The updated analytics allow for a clearer view of how tokens are being consumed, which helps in predicting future budget needs. It also helps identify inefficient agent loops that might be burning through budget without delivering equivalent value. These spend controls shift what the enterprise stack can rely on by default. Instead of reacting to a high bill at the end of the month, admins can now proactively set limits. It is worth tracking how these controls impact agent performance, especially if a ceiling is reached mid-session. Builders need to ensure their agents can handle graceful degradation or clear notification when budget limits are hit. This move by OpenAI reflects the maturing of the enterprise AI market. Cost management is no longer an afterthought, it is a core requirement for production deployments. We expect to see other major providers follow suit with similar enterprise-grade visibility and control features as the focus shifts from experimentation to operational efficiency. A recent analysis argues that 30 years of US export controls on encryption and cyber security software have largely failed to slow their spread. This historical context is now being applied to Anthropik's mythos cybersecurity model. The core argument is that dual-use software has always leaked, forked, and been re-implemented regardless of jurisdiction. The historical mechanism that failed involved treating source code or compiled binaries as the controlled artifact. For Mythos, the challenge is even greater. The question is whether model weights, training compute, or hosted inference APIs can be effectively controlled at all. Unlike a compiled binary, a model's capability is harder to pin down to a specific artifact. This makes the surface for regulation much more complex. The analysis suggests that the diffusion mechanism for AI will likely follow the same path as cryptographic tools. Code forks and foreign re-implementations will outpace any regulatory framework. For builders, the takeaway is that frontier AI cybersecurity capabilities will likely be globally accessible. You should assume that both defensive and offensive AI security tools will be available to a wide range of actors. This makes export classification a moving target for compliance teams. If you are building on or against Mythos, the classification of model weights and inference endpoints remains a point of ambiguity. We should be watching for any new rulemaking from the Commerce Department regarding frontier model weights. At the same time, keep an eye on whether Anthropic publishes a usage policy that attempts to preempt these regulatory questions. The fight over Mythos will likely set the policy frame for the entire AI security sector. AI Inference startup Base 10 is reportedly close to finalizing a $1.5 billion funding round at a $13 billion valuation. This raise comes just months after their previous mega round and highlights a major shift in the market. Inference is increasingly being treated as its own infrastructure category rather than just a feature of model training labs. Dedicated capital is flowing into companies that focus specifically on the serving side of the stack. The technical mechanism here is the inference-specific serving stack. This includes things like model compilation passes for production deployment and GPU pooling across heterogeneous hardware like H100S and H200S. Basten's differentiator is exposing these knobs to engineering teams as a managed service. Instead of relying on opaque API endpoints from a model lab, builders get more control over request routing and optimization for bursty or long-context traffic. This massive raise signals that the inference layer is becoming a buyer's market. There is real competition now between hyperscalers and specialized providers like Base 10, Fireworks, and Together. For builders, this means the default choice of an inference provider is no longer obvious. You should be explicitly evaluating the trade-offs between cost per token, latency, and the level of customization available. We'll be watching to see how Base 10 positions itself against hyperscaler inference APIs like Bedrock or Azure AI. As more capital enters this space, the pace of innovation at the serving layer will likely accelerate. This is good news for agent builders who need reliable, high-performance infrastructure to power their production loops. Dataset has launched a new plugin called Dataset Apps that allows users to host custom HTML and JavaScript applications directly inside Dataset. These are self-contained applications that can leverage the data stored in the Dataset instance. The launch announcement highlights the why behind this move, focusing on the need for more flexible ways to visualize and interact with data without needing a separate hosting stack. For agent stack builders, this is an interesting development in the UI layer. By hosting custom HTML inside the data tool itself, you can create tighter feedback loops for agents that are interacting with databases. It simplifies the deployment of internal tools and dashboards that require a specific front-end but need to be closely coupled with the data source. The change lands at the API and runtime level, affecting how you configure the plugin and deploy your custom apps. This shift means builders can rely on Dataset as a more complete platform for data-driven agents. It is worth tracking how the plugin handles different workloads and how it scales with complex JavaScript applications. If you are building tools for data exploration or agent monitoring, this could significantly reduce your infrastructure overhead. We will be watching for how the community adopts Dataset apps and what kind of custom interfaces start to emerge. As more builders experiment with this, we expect to see new patterns for agent-human collaboration surfaces that live directly on top of the data they are processing. Recent newsletters have noted a relatively slow period for AI news, with no major model drops or agent framework releases dominating the cycle. This breathing room is often a sign of calendar-driven release coordination. With the AI engineer, or AIE, conference on the horizon, many major labs and maintainers are likely consolidating their reveals for that event. Pre-conference lulls like this are a standard pattern in the ecosystem. For builders, this quiet window is actually quite useful. It provides a chance to consolidate notes on current agent frameworks and stable API versions. You can commit to your current stack for the next few days without the high risk of a breaking change disrupting your local environment. It's a good time to focus on refining existing workflows and ensuring your session handling and tool integrations are solid. The watch item here is the AIE opening keynote. Historically, model providers and agent runtime maintainers use these stages to ship reference implementations and pinned version releases. These announcements typically propagate through documentation and repositories within hours of the talk. The quieter the run-up, the more impactful the keynote reveals tend to be. We should expect a flurry of activity once the conference begins. For now, enjoy the stability and use the time to prepare your stack for the next wave of updates. The transition from this lull to the post-conference shipping cycle is usually very fast. Meredith Whitaker, the president of Signal, recently used an interview to push back on the trend of positioning AI chatbots as companions. Her message was clear, these are not your friends, not conscious beings, and not sentient interlocutors. This critique is aimed directly at vendors who use relational language in their onboarding flows, persona system prompts, and conversational design. This intervention is particularly relevant as agentic coding tools and support bots become embedded in our daily workflows. The line between a tool and a teammate is being blurred by anthropomorphic design choices. Whitaker argues that framing models as peers creates false expectations about their memory and intent. For builders, this means that your agent's product copy and system prompts are now part of the trust surface. Anthropomorphic framing in persona prompts or UI greetings can be a privacy red flag for some users. Signals focus on privacy gives this critique a lot of weight in developer circles. As a builder, you might want to consider whether your agent's persona is helping or hurting user trust. Relational copy could eventually become a deal breaker for privacy-focused enterprise buyers. We'll be watching to see if major model providers change their guidance on relational framing in developer documentation. If enterprise procurement starts flagging anthropomorphic copy, we could see a shift back toward more utilitarian and transparent agent personas. How you frame the agent-human interaction is becoming a core design decision. A new service called In The Waits has launched, pitching itself as a vanity search engine for the AI era. Instead of indexing web pages, it assigns users a score based on how prominently their identity surfaces inside the parameters and training data of frontier AI models. It uses the model as the search index, providing a metric for how much presence a person has within the AI's internal representation of the world. The underlying mechanism involves running repeated inference probes against hosted LLMs. The architecture uses an evaluation harness to query identity mentions across multiple model checkpoints. It then aggregates these hit rates and frequency counts into a single composite score. This reframes model visibility as a measurable metric, something that hasn't really existed as a productized surface before. For those shipping AI products, this could become a new category of user-facing metric. It surfaces a different kind of signal than a traditional search ranking. The question for builders is whether the scoring methodology will become standardized or remain opaque. Reproducibility is key to whether this score becomes a real signal for marketing or just another vanity number. We'll watch to see if the probing protocol for this service gets opened up. If it does, we could see new tooling built around tuning your presence across different model versions. It's an interesting look at how we might measure the influence of individuals and brands in an AI-saturated information environment. First on the radar is Fast MCP from the Prefect team. This is a Pythonic framework designed for building servers and clients using the model context protocol. It allows developers to expose tools and resources to LLMs with minimal boilerplate. If you have internal Python functions that you want an agent-like clod code to invoke directly, Fast MCP provides a standard way to wrap those as typed tool calls. It's a great replacement for ad hoc CLI wrappers. Next is the model context protocol for beginners repo from Microsoft. This is a comprehensive curriculum that walks developers through the fundamentals of the protocol. What makes it useful is the cross-language examples in .NET, Java, Rust, and TypeScript. It provides a clear path for wiring agents into non-Python services without needing to rewrite everything. You can pick a lab in your preferred language and start porting your existing internal tools into an MCP server. Finally, we have Unity MCP. This project bridges AI assistants and the Unity Editor, exposing tools for asset management, scene control, and editor automation. It makes the Unity Editor API reachable through standard MCP tool calls. For builders in the gaming or simulation space, this turns asset pipelines and scene edits into agent-driven steps. You can orchestrate a full build-on test loop entirely through your agent stack. In this cycle's model check, we have two significant entries from poolside. First is Laguna XS.2, now available via OpenRouter. As we discussed earlier, this is a second-generation compact model in their XS class. It features a 262,000 token context window and combines tool calling with reasoning. If you are looking to optimize for cost and latency in your agent loops, this is a prime candidate for evaluation against your current defaults. The second selected model is Laguna M.1. This is the flagship entry from poolside, also available via API. It is optimized for complex software engineering tasks and supports the same massive context window. Its capabilities are aimed at higher-tier agentic coding workflows. We recommend routing a few coding agent sessions through both M.1 and XS.2 to see how they handle your specific codebase compared to your current primary models. Other models were reviewed this cycle but were not selected for full coverage as they did not represent new major provider releases. We continue to track the long context and tool calling performance of all major models on the market to ensure your agent stack has the best possible foundations. For our local spotlight, we are looking at Alama.30.10. This update adds MLX accelerated support for Cohers Command A and the North family of models on Apple Silicon. It also bumps the embedded Lama engine to build 9672. This is a significant update for builders who want to keep their workloads on-device, as it makes local inference of these model families viable on M-Series Macs without needing CUDA. If you are working on an M-Series Mac, you can pull a Command A or North model and benchmark the tokens per second against your previous Alama build. This release tightens the loop for agents that need to operate in high privacy or offline environments. The performance gains from MLX acceleration, are noticeable and help make local agents much more responsive. In our first extra, we look at the ambitions of billionaire Mukesh Ambani, who wants to weave AI into every call, app, and home in India. The technical angle here is the embedding of on-device and edge inference into a telecom carrier stack with over 500 million subscribers. This represents a massive scale for network resident AI surfaces. Next is a story about the CEO of Allbirds launching a new AI business with a plan but no employees. This raises questions about how a single founder company can operationalize model selection and shipping by leaning entirely on outsourced or vendor-provided inference stacks. It's an interesting look at the company of one model powered by high-tier agents. Finally, there is a dispute over whether ASML's top chip-making tools have found their way into China. The story hinges on the verification of lithography export controls and how the physical location of such restricted tools is tracked. It highlights the friction between commercial logic and national security regulations in the high-end hardware space that powers AI. From today's stories, the stable releases of OpenClaw 6.9, Hermes 6.19, and ClaudeCode 0.176 shift what your agent stack can rely on by default. Poolside's Laguna XS.2 and M.1 give you new options for coding agent loops with large context requirements. OpenAI's new enterprise controls provide the visibility needed for managing agent spend at scale. The historical failure of export controls suggests that defensive and offensive AI capabilities will remain globally accessible, making security a baseline assumption. The massive raise for Base 10 indicates that the inference layer is now a separately funded and highly competitive market. Finally, remember that your choice of agent persona and copy can impact user trust, especially in light of the privacy concerns raised by signal. You can find more details and all the links in the show notes at tobyonfitnesstech.com. Thanks for listening to Agent Stack Daily. We'll be back soon.