← Back to search

Episode 73: OpenClaw v2026.6.9, Hermes v2026.6.19, Claude Code 2.1.176 Released; Poolside Adds Laguna XS.2 and M.1

AgentStack Daily (Español) · 2026-06-21 · 27 min
relevance 60 4644 words Episode page ↗ Audio ↗
Show full episode description
AgentStack Daily de hoy: OpenClaw v2026.6.9, Hermes Agent v2026.6.19 y Claude Code CLI 2.1.176 llegaron con nuevas versiones. Poolside lanzó Laguna XS.2 en OpenRouter y Laguna M.1 a través de API. Baseten reporta una ronda de $1.5B, Datasette Apps permite alojar HTML personalizado y Meredith Whittaker de Signal advierte que los chatbots de IA no son tus amigos, mientras los equipos empresariales ahora cuentan con nuevos análisis de uso y controles de gasto actualizados. Show notes: https://tobyonfitnesstech.com/es/podcasts/episode-73/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Builders need to track maturing agent stacks, new coding models, and production cost/security controls in one cycle.
Benefits
  • OpenClaw 6.9 richer Telegram delivery and auto plugin approvals
  • Hermes 6.19 adds background subagents and iMessage via Photon
  • Claude Code 2.1.176 lighter install, faster cold start
  • Laguna XS.2 pairs tools and reasoning in one endpoint
  • OpenAI enterprise spend controls and usage breakdowns
Use cases
  • Multi-file refactor agents using Laguna XS.2's cheaper per-token economics
  • Enterprise admins setting configurable spend ceilings on ChatGPT seats
  • Dataset Apps hosting custom HTML/JS dashboards beside agent data
  • Hermes dashboard full profile builder and revamped Skills Hub browser
KPIs / results
  • Laguna XS.2 context window 262,144 tokens
  • Laguna M.1 context window ~256,000 tokens
  • Basetten ~$1.5B round at $13B valuation
Tools / build
0:00 / 0:00
🌐 This transcript was automatically translated to English from the original.
I'm Nova I'm Alloy, and this is AgenteStack Daily OpenClaw 6.9 Hermes Agents 6i19 and Code.176's Terminal Cloud-based AI coding agent all released stable versions this cycle. OpenClaw 6.9 added richer delivery in Telegram and optimized integration with codecs with automatic plugin approvals, while Hermes 6i19 introduced background subagents and expanded channel support to iMessaging via Poton. Today you'll hear about those Jarnises updates, the launch of Paulside's Laguna coding models on Open Router and via API, new enterprise spend controls from OpenAI, and why the president of Signal says your AI chatbot is definitely not your friend. We also take a look at a billion-dollar round for Basetten, a vanity search engine that scores your presence within the model weights, and the historical failure of export controls as a lens for Antropic's mythos model. The common thread is a focus on the production layer, how we manage the cost of agentive loops, how we secure the models at their core, and how the Jarnises we use every day are maturing to handle multi-step, multi-channel workloads. Three stable releases arrived this cycle and shape how Agentive Jarnises are being assembled right now. Open Class 6-9, published on June 21, includes richer delivery on Telegram. The channel path now sends rich HTML, preserves rich markdown and sticker paths, and renders progress drafts and command output more faithfully. It also safely normalizes HTML tables and keeps spooled mentions and handlers on the correct delivery path. Agent recovery is more reliable in this version, retries. Terminal results and session history repair now maintain more interrupted or partial shifts, moving towards a visible final result. The integration of codecs in Open Class is also significantly stronger. Adds GPT 5.3 Spark automatic plugin approvals, routing or aus, and remote node execution as a dynamic tool. These changes move Harness toward a more reliable application server teardown process. Meanwhile, Hermesagen 6 and 19, which maintainers are calling the Rich Relayase, extends Hermes through new channels like iMessage via Poton and the Raft agent network. The desktop app also gained substantial new capability. Subagents can now run in the background and the imaging tool has learned to edit. The Hermes dashboard got a full profile builder and the Skills Hub browser was completely revamped. For those using the OpenAI Codex.141 terminal-based AI encoding agent, these updates provide a more robust backend. We also saw terminal-based AI encryption agent CloudEco.176 complete the trio on the stable label. This release focuses on a lighter installation footprint and faster cold boot path for Shell-based sessions. At the API and Runtime layer, these changes alter what builders can configure and depend on by default. The practical implication is cleaner harness plumbing. Openclada gives builders more model paths and a more secure house, while Hermes and CloudEcoot refine the user and developer experience. The part to watch is how these new defaults perform under real workloads before moving them to production, especially where session history and terminal output are involved. Paul Side posted the Laguna XS2 on OpenRouter this week, marking the second-generation model in its XS size class. This is their efficient series of encoding agents, and the argument here is compactness over high-end reasoning. It combines tool calls with reasoning in a small footprint, with a context window of 262,144 tokens. This context size puts it firmly in the long context tier for agentive encoding workloads where passing large code bases or tool traces is the norm. The most relevant detail for builders is the unified pairing of tools and reasoning. An endpoint exposes both tool invocation and reasoning, simplifying agent loop orchestration. You don't have to route requests through two separate models to handle planning and execution. The XS size class signals a target audience focused on low-latency, cost-sensitive runs. If you're running multi-file refactoring agents, the per-token economics here might be more attractive than higher-end reasoning models. For agent builders, this release extends the compact level on the router. It ranks alongside other compact options used for economic sorting or initial planning passes. What we need to look at next are actual latency benchmarks on actual encoding traces. We also want to see if Polside eventually releases specific rate limits for tool calls or pricing tiers to further differentiate the XS line from its larger models. This compact level update matters because it moves the cost frontier rather than just the raw quality frontier. If your agent stack depends on many small, iterative calls, a compact second-generation model with a massive context window is a significant integration goal. Enables more complex multi-step workflows without the latency hit of a larger high-end model. Moving to the other end of the spectrum, Polside has also launched Laguna M1 via API. This is their high-end encryption agent model, optimized for complex software engineering tasks. Like its younger brother, it supports tool calls and reasoning with a context window of 256,000 tokens. At the mechanism level, this change manifests itself in the surface of the API and runtime behavior that agent builders integrate. The Laguna M1 release is geared toward the heavy lifting of agentive coding. Think architectural changes, complex bug fixes, and large-scale refactorings that require a deep understanding of the entire stack. The primary source documentation for this model includes specific deployment notes and changelok context that builders should review. It's worth tracking how M1 performs under heavy production loads, as I know high-end encoding models often face unique challenges with consistency over long traces. Why this matters right now is the speed at which the agent stack is moving? Changes in this layer determine which workflows are reliable and which remain fragile. The practical question for builders is whether Laguna M1 can replace its current default range. High for coding tasks. Early evidence suggests it is a strong contender for those building engineering, autonomous or semi-autonomous agents. We should keep an eye out for follow-up releases and independent benchmark results. As surrounding tools like SDKS and security review frameworks gather support for M1, the barrier to switching will decrease. For now, the focus is on checking your performance against the specific coding patterns your agents are expected to handle. OpenAI is introducing new spending controls and usage analytics for GPT Enterprise chat. This is a direct response to organizations that need to manage costs and scale their AI deployments with more confidence. These controls land at the management and API level, affecting how admins configure and deploy positions in a large company. Includes more granular cost breakdowns and the ability to set configurable ceilings on spend. For platform builders and engineers, these tools are all about visibility. As agents multiply within an enterprise, knowing exactly which departments or projects are driving usage is critical. Updated analytics allow for a clearer view of how tokens are being consumed, helping to predict future budget needs. It also helps identify loops of inefficient agents that could be burning budget without delivering equivalent value. These spend controls change what the enterprise stack can rely on by default. Instead of reacting to a high bill at the end of the month, admins can now proactively set limits. It's worth tracking how these controls impact agent performance, especially if a mid-session ceiling is reached. Builders need to ensure their agents can handle grass-fell degradation or clear notification when budget limits are reached. This decision by OpenAI reflects the maturation of the enterprise AI market. Cost management is no longer an afterthought, it is a fundamental requirement for production deployments. We expect other major vendors to follow suit with similar enterprise-grade visibility and control features as the focus shifts from experimentation to operational efficiency. A recent analysis argues that 30 years of US export controls on encryption and cybersecurity software have largely failed to slow its spread. This historical context is now being applied to Antropic's mythos cybersecurity model. The central argument is that dual-use software has always leaked, forked and redeployed regardless of jurisdiction. The historical mechanism that failed was to treat source code or compiled binaries as the controlled artifact. For myths, the challenge is even greater. The question is whether model weights, training computation, or hosted inference APIs can be effectively controlled at all. Unlike a compiled binary, a model's capability is harder to tie to a specific artifact. This makes the surface for regulation much more complex. The analysis suggests that the diffusion mechanism for AI will likely follow the same path as cryptographic tools, code forks and foreign reimplementations will outpace any regulatory framework. For builders, the bottom line is that cutting-edge AI cybersecurity capabilities will likely become globally accessible. You should assume that both defensive and offensive AI security tools will be available to a wide range of actors. This makes export classification a moving target for compliance teams. Whether you are building on or against myths, the classification of model weights and inference hotspots remains a point of ambiguity. We should be attentive to any new regulations from the Department of Commerce regarding the weights of cutting-edge models. At the same time, keep an eye out for Antropic publishing a usage policy that attempts to pre-empt these regulatory questions. The myth fight will likely set the political framework for the entire AI security sector. AI inference startup Basetten is reportedly close to finalizing a $1.5 billion funding round at a $13 billion valuation. This raising comes just months after its previous mega round and highlights a big change in the market. Inference is increasingly being treated as its own category of infrastructure rather than simply a feature of model training labs. Dedicated capital is flowing to companies that focus specifically on the service side of the stack. The technical mechanism here is the inference-specific servicing stack. This includes things like model build passes for production deployment and bundling GPUs on heterogeneous hardware like the H100S and H200S. Basetten's differentiator is exposing these controls to engineering teams as a managed service. Instead of relying on opaque API access points from a model lab, builders gain more control over request routing and optimization for bursty or long-context traffic. This huge raise signals that the inference layer is becoming a buyers' market. There is now real competition between hyperscalers and specialist providers such as Basetten, Fireworks and Together. For constructors, this means that the default choice of an inference provider is no longer obvious. You should be explicitly evaluating trade-offs between cost per token, latency, and the level of customization available. We'll be looking at how Basetten positions itself against inference APIs from hyperscalers like Bedrak or Azure AI. As more capital enters this space, the pace of innovation at the service layer is likely to accelerate. This is good news for agent builders who need reliable, high-performance infrastructure to power their production loops. Dataset has released a new plugin called Dataset Apps that allows users to host applications. Custom HTML and JavaScript directly within Dataset. These are self-contained applications that can take advantage of the data stored in the Dataset instance. The launch announcement highlights the rationale behind this move, focusing on the need for more flexible ways to view and interact with data without the need for a separate hosting stack. For agent stack builders, this is an interesting development at the UI layer. By hosting custom HTML within the data tool itself, you can create tighter feedback loops for agents interacting with databases. Simplifies the deployment of internal tools and dashboards that require a specific frontend but need to be tightly coupled to the data source. The change lands at the API and runtime level, affecting how you configure the plugin and deploy your custom applications. This change means builders can rely on Dataset as a more comprehensive platform for data-driven agents. It's worth following how the plugin handles different workloads and how it scales with complex Javascript applications. If you're building tools for data exploration or agent monitoring, this could significantly reduce your infrastructure overhead. We'll be keeping an eye on how the community adopts Dataset Apps and what kind of custom interfaces start to emerge. As more builders experiment with this, we expect to see new patterns for human collaboration surfaces that live directly above the data they're processing. Recent bulletins have noted a relatively slow period for AI news, with no major releases of agent models or frameworks dominating the cycle. This respite is often a sign of calendar-driven release coordination. With the I-Engineer, or IE, conference on the horizon, many major labs and maintainers are likely consolidating their reveals for that event. Pre-conference breaks like this are a standard pattern in the ecosystem. For builders, this quiet window is quite useful. Provides the opportunity to consolidate notes on current agent frameworks and stable API versions. You can commit to your current stack for the next few days without the high risk of a disruptive change disrupting your local environment. It's a good time to focus on refining existing workflows and ensuring your session management and tool integrations are robust. The point to note here is the EI opening speech. Historically, model vendors and agent runtime maintainers use these scenarios to ship reference implementations and pinned version releases. These announcements typically propagate through documentation and repositories within hours of the talk. The quieter the countdown, the more shocking the keynote revelations tend to be. We should expect a flurry of activity once the conference begins. For now, enjoy the stability and use the time to prepare your stack for the next wave of updates. The transition from this dormancy to the post-conference shipping cycle is usually very quick. Meredith Whitaker, the president of Signal, recently used a view interview to pushback against the trend of positioning AI chatbots as companions. His message was clear, these are not your friends, they are not conscious beings and they are not sentient interlocutors. This criticism is aimed squarely at vendors who use relational language in their onboarding flows, persona system pamps, and conversational design. This intervention is particularly relevant as agent coding tools and support bots become integrated into our daily workflows. The line between a tool and a companion is being blurred by anthropomorphic design choices. Whitaker argues that framing models as peers creates false expectations about their memory and intention. For builders, this means that your agent's product copy and system prams are now part of the trust surface. Anthropomorphic framing in person prams or UI greetings may be a privacy red flag for some users. Signal's focus on privacy gives this criticism a lot of weight in developer circles. As a builder, you might want to consider whether your agent persona is helping or hurting user trust. Relational copy could eventually become a deal-breaker for privacy-focused enterprise buyers. We'll be watching to see if the major model vendors change their recommendations on relational framing in developer documentation. If corporate purchasing begins to flag anthropomorphic copies as problematic, we could see a return toward more utilitarian and transparent agent personalities. How you frame human-agent interaction is becoming a critical design decision. A new service called Indewatch has been launched, pitching itself as a vanity search engine for the AI ​​age. Instead of indexing web pages, it assigns users a score based on how prominently their identity appears within the parameters and training data of cutting-edge AI models. It uses the model as a search index, providing a metric of how much presence a person has within the internal representation of the AI ​​world. The underlying mechanism involves running repeated inference tests against hosted LLMSs. The architecture uses an evaluation harness to query identity mentions across multiple model checkpoints. It then aggregates these hit rates and frequency counts into a single composite score. This reframes model visibility as a measurable metric, something that hasn't really existed as a productivized surface before. For those launching AI products, this could become a new category of user-facing metrics. SURF is a different signal than a traditional search ranking. The question for builders is whether the scoring methodology will become standardized or remain opaque. Reproducibility is key to determining whether this score becomes a real signal for marketing or just another vanity number. We will be waiting to see if the probe protocol for this service opens. If it does we could see new tools built around adjusting your presence across different model versions. It's an interesting look at how we might measure the influence of individuals and brands in an AI-saturated information environment. First on the radar is Fast MCP from the Prefect team. This is a Pythonic framework designed to build servers and clients using the Context Protocol model. It allows developers to expose tools and resources to LLMS with minimal boilerplate code. If you have internal Python functions that you want an agent like Cloud Code to call directly, Fast MCP provides a standard way to wrap those as typed tool calls. It's a great replacement for makeshift CLI wrappers. It is followed by Microsoft's Model Context Protocol for Beginners repository. This is a comprehensive curriculum that guides developers through the fundamentals of the protocol. What makes it useful are the multi-language examples in .NET, Java, Roost, and TypeScript. Provides a clear path to connect agents to non-Python services without needing to rewrite everything. You can choose a lab in your preferred language and start contributing your existing internal tools to an MCP server. Finally, we have Unity MCP. This project brings together AI assistants and the Unity Editor by exposing tools for asset management, scene control, and editor automation. Makes the Unity editor API accessible through standard MCP tool calls. For builders in the gaming or simulation space, this turns asset pipelines and scene edits into agent-driven steps. You can orchestrate an entire build and test cycle entirely through your agent stack. In the model review of this cycle, we have two significant inputs from Paulside. First up is the Laguna XS-2, now available through Open Router. As we discussed earlier, this is a second-generation compact model in its XS class. It has a context window of 262,000 tokens and combines tool calls with reasoning. If you're looking to optimize for cost and latency in your agent cycles, this is a prime candidate to evaluate against your current options. The second model selected is Laguna M1. This is Paulside's flagship entry, also available via API. It is optimized for complex software engineering tasks and supports the same huge context window. Its capabilities are aimed at higher-level advanced coding workflows. We recommend routing some encryption agent sessions through M1 and XS. Two to see how they handle your specific codebase compared to your current core models. Other models were reviewed this cycle but were not selected for full coverage as they did not represent new releases from major vendors. We continue to track the long-context and tool call performance of all the major models on the market to ensure your agent stack has the best foundation possible. For our local Spotlight, we're looking at Yama.30.10. This update adds MLX-accelerated support for Command AD Coere and the model family. Norton Apple Silicon. It also updates the embedded Yama engine to version 9672. This is a significant update for builders who want to keep their workloads on-device, as it makes local inference of these model families viable on M-chip Macs without the need for CUDA. If you're working on a Mac with an M chip, you can download a Command AD or North model and compare the tokens per second against your previous Yama build. This release narrows the cycle for agents who need to operate in high-privacy or offline environments. The performance gains from MLX acceleration are noticeable and help make local agents much more responsive. In our first additional note, we see the ambitions of billionaire Mukesambani, who wants to weave AI into every call, app and home in India. The technical angle is the embedded inference on device and at the edge within an operator's stack. Telecommunications with more than 500 million subscribers. This represents massive scale for network-resident AI surfaces. It's followed by a story about the AdWords CEO launching a new AI business with a plan but no employees. This raises questions about how a single-founder company can operationalize model selection and product shipping by relying entirely on outsourced or vendor-provided inference stacks. It's an interesting look at one's business model, driven by high-level agents. Finally, there is a dispute over whether ASML's core chipmaking tools have reached China. The story hinges on verification of lithography export controls and how the physical location of such restricted tools is tracked. It highlights the friction between business logic and national security regulations in the high-level hardware space that drives AI. From today's stories, the stable releases of Open Class 6, 9, Hermes 6 and 19, and Cloud Code.176 change what your agent stack can depend on by default. Paulside's Laguna XS, 2 and M1 give you new options for encoding agent cycles with large context requirements. OpenAI's new enterprise controls provide the visibility needed to manage agent spend at scale. The historic failure of export controls suggests that defensive and offensive AI capabilities will remain globally accessible, making security a baseline assumption. The huge funding round for Baseten indicates that the inference layer is now a separately funded and highly competitive market. Finally, remember that your choice of capi AI agent personality may impact user trust, especially in light of the privacy concerns raised by Signal. You can find more details and all the links in the show notes at tobionfitnesstech.com Thanks for listening AgentsTagDaily. We will be back soon.