आज का AgentStack Daily: OpenClaw v2026.6.9, Hermes Agent v2026.6.19 और Claude Code CLI 2.1.176 सभी नए रिलीज़ के साथ आए। Poolside ने OpenRouter पर Laguna XS.2 और API के माध्यम से Laguna M.1 जारी किया। एंटरप्राइज़ टीमों को नए उपयोग विश्लेषण और अपडेटेड खर्च नियंत्रण मिले। Show notes: https://tobyonfitnesstech.com/hi/podcasts/episode-73/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Builders need to track maturing agent stacks, new coding models, and production cost/security controls in one cycle.
🌐 This transcript was automatically translated to English from the original.
I'm Nova, I'm Alay and this is AgentStack Daily. OpenClaw 6.9, Hermes Agent 6.19 and Terminal-Based AI Coding Agent Cloth Code 0.176 have all sent stable releases in this cycle. OpenClaw 6.9 added Richard, Telegram Delivery and improved KDEX Integration with Automatic Plugin Approvals, while Hermes 6.9 introduced background subs and extended Channel Support for iMessage via Photon. Today you'll hear about Unharness updates, the release of Poolside's Laguna Coding Models on Open Router and APIs, new Enterprise Spend Controls from OpenAI, and why Signal's President says your AI Chatbot is definitely not your friend. We'll also look at Bay Sittin's Billion Dollar Raise, a Vanity Search Engine that scores your presence within the model weights, and the historic failure of Export Controls through the lens of Anthropic's Mythos Model. The basic theme is to focus on production. How do we manage the cost of Agentic Loops? How to secure the model at their center? And how are the harnesses we use every day maturing to handle multi-step, multi-channel workloads? There were 3 stable releases in this cycle and these are shaping how Agentic Harnesses are being assembled right now. Open Cloche is DecimalNow, published June 21st Richard brings Telegram Delivery Channel Path now sends Rich HTML, Preserves Rich Markdown and Sticker Paths, and renders Progress Drafts and Command Output more faithfully It also Safely Normalizes HTML Tables and keeps Mentions and Spooled Handlers on the correct delivery path Agent Recovery is more dependable in this version, Retrize, Terminal Outcomes, and Session History Repair are now more Interrupted or Kodeks integration in OpenClaw has also been strengthened. It includes Automatic Plugin Approvals, GPT 5.3 Spark OAuth Routing and Remote Node Execution as a Dynamic Tool. These changes move Harness towards a more reliable App Server Teardown Process. Meanwhile, Hums Agent 6.19, which is called Maintenance Reach Release, extends Hums to new channels like iMessage and Raft Agent Network through Photon in Desktop App. Substantial new capabilities have also arrived. All Agents can now run in the background and the Image Generation Tool has learned to edit. Our Dashboard got a Full Profile Builder and the Skills Hub browser was completely overhauled. These updates provide a more robust backend for those using OpenAI Codex.141's Terminal Based Coding Agent. We also saw the Terminal Based AI Coding Agent Cloth Code.176 joining this trio of Stable Releases on the Stable Tag. This release has a lighter install footprint and Centered on a Faster Cold Start Path for Sheldriven Sessions These changes at the API and Runtime layer change what builders can configure and what they can relay on Defaults The practical implication is that Harness Plumbing becomes cleaner Open Cloth gives builders more Model Routes and Safer Auth while Hermes and Cloth Code refine the user and developer experience The part to keep an eye on is that these new Defaults under Real Workloads before flipping to production How to flow especially where Session History Poolside pushed the Laguna XS.2 on Open Router this week, which is the second generation model in their XS Size Class. Here's their Efficient Coding Agent Series. And here's Pitch Compactness Over Flagship Reasoning. It combines Tool Calling with Reasoning in a small footprint with a Context Window of 262,144 Tokens. This Context Size puts it right in the Long Context Tier for Agentic Coding Workloads where Large Code Bases or Tool Traces are nearby. The most relevant details for builders are that Tool and Reasoning are a single integrated pair A single Endpoint exposes both Tool Invocation and Reasoning which simplifies Agent Loop Orchestration You don't have to route requests to two different models to handle Planning and Execution Access Size Class Completeness Indicates Target Audience Focusing on cost-sensitive runs If you're running a Multi-File Refactor Agent The per-token economics here may be more attractive than the Flagship Class Reasoning models For Agent Builders This release widens the Compact Tier on Routers Sits here alongside other Compact options that are used for cheaper Classification or Initial Planning Passes Next up we need to see actual Latency Benchmarks on real Coding Traces We also want to see if Poolside eventually releases Tool Calling specific Rate Limits or Pricing Tiers to further differentiate the XS Line from its larger Compact models This refresh of the tier is important because it pushes not only the limits of raw quality but also the limits of cost. If your Agent Stack relies on a lot of small Iterative Calls then the Second Gen Compact Model with a huge Context Window is an important integration target. This allows for more complex multistep workflows without the latency hit of the larger Flagship Model. Moving to the other end of the spectrum, Poolside has also released Laguna M.1 via API. Here is their Flagship Coding Agent Model which provides complex Soft Optimized for ware engineering practices Like its younger brother, it supports Tool Calling and Reasoning with a 296 token Context Window. At the system level, this change is reflected in the API surface and runtime behavior against which Agent Builders integrate. Laguna M.1 is a release targeted at the heavy lifting of agentic coding. Think architectural changes, complex bug fixes, and large-scale refactors that require a deep understanding of the entire stack. Prathmix Rote of this Model Documentation includes specific Deployment Notes and Change Log references that builders should review. It's worth tracking how MA behaves under heavy production loads because flagship coding models often face unique challenges with consistency over long traces. This matters now because of how fast the Agent Stack is growing. Changes in this layer determine which workflows are reliable and which still remain fragile. The practical question for builders is whether Laguna M.A can replace their current flagship default coding tools. Early evidence suggests it is a strong contender for those building autonomous or semiautonomous engineering agents. We should keep an eye on follow-up releases and independent benchmark results. As surrounding tooling like SDKs and security review frameworks start to gain support for M.1, the barrier to switching will go down. For now, focus on what to expect from your agents. OpenAI Chat is validating its performance against specific coding patterns OpenAI Chat GPT is introducing new spend controls and usage breakdowns for Enterprise This is a direct response to the needs of organizations that want to manage costs and scale their AI deployments with more confidence These controls are applied at the management and API level, allowing admins to customize how to configure and deploy seats across a large company It includes more detailed cost breakdowns and configurable ceilings on spend This tool for builders and platform engineers includes the ability to set It's about visibility as agents spread across the company It's important to know which departments or projects are driving the usage The updated visualization here gives a clearer view of how tokens are being spent which helps predict future budget projections It also helps identify inefficient agent loops that can burn through the budget without delivering equivalent value These spend controls change what the enterprise stack defaults to. Admins can now set limits in advance, rather than having to respond to high bills at the end of the month It's worth tracking how these controls impact agent performance, especially if the ceiling is reached mid-season Builders need to ensure their agents can handle graceful declines or clear notifications when budget limits are reached This move to Open AI reflects the maturing of the enterprise AI market Cost management is no longer a consideration here Production We expect other major providers to follow suit and come up with similar enterprise-grade visibility and control features as the focus shifts from ease of use to operational efficiency. A recent study argues that 30 years of US export controls on encryption and cybersecurity software have largely failed to slow their spread. This historical context is now entropic. The main problem with which Mythas is being applied to the cyber security model is that due to the fact that software has always been leaked, forked and reimplemented regardless of ownership, the historical mechanism that failed was to treat source code or compiled binaries as controlled artifacts. The challenge for Mythas is even greater. The question is whether the model is weighted, trained, computed or hosted. Inference API can be effectively controlled. Unlike compiled binary, it is difficult to limit the capabilities of a model to a specific artifact. This makes the manipulation quite complex for regulation. The visualization shows that the propagation mechanism for AI could potentially be a cryptographic tool. Code folk and overseas re-implementation will override any regulatory framework The key takeaway for builders is that frontier AI cybersecurity capabilities will likely be globally accessible You should assume that both defensive and offensive AI security tools will be available for a variety of tasks This makes export classification a moving target for compliance teams If you're building with or against Mythos, the model The classification of weights and inference endpoints remains a point of ambiguity. We should keep an eye on any new rulemaking from the Commerce Department regarding frontier model weights, as well as whether Anthropic publishes a usage policy that attempts to address these regulatory questions in advance. This fight over Mythas will likely set the policy framework for the entire AI security sector. AI Inference Startup Basiton's $1B Funding The round is reported to be finalized at a valuation of $13 billion. This funding comes just a few months after their previous mega round and marks a major shift in the market. Inference is no longer seen as a feature of the Model Training Lab but as a separate infrastructure category in its own right. Dedicated funding is going into companies that focus specifically on the serving side. The technology here is the Inference specific serving stack. Basiton's differentiator is to expose these NOBS as a managed service for engineering teams. Builders get more control over request routing and optimization for bursty or long context traffic, rather than relying on an opaque API endpoint from Model Lab. This huge funding signals that the inference layer is becoming a buyer's market. There is now real competition between specialist providers like Basiton, Fireworks and Together. For builders this means that the default choice of inference provider is now clear. You must clearly evaluate the tradeoffs between cost per token, latency, and the level of customization available. We'll be watching here to see how the Basiton Hyperscalar Inference API positions itself against the likes of Bedrock or Azure AI. The pace of innovation at the Serving Layer will likely accelerate as more capital enters the space. There's good news here for Agent Builders looking for reliable high-performance infrastructure to power their Production Loop. Dataset has launched a new plugin called Dataset Apps that allows users to host custom HTML and JavaScript applications directly inside Dataset. These are self-contained applications that can take advantage of data stored in a Dataset instance. The launch announcement outlines the why behind the move. It focuses on the need for more flexible ways to visualize and interact with data without the need for a separate hosting stack. This UI for Agent stack builders An interesting development in the layer is that by hosting custom HTML within the data tool itself, you can create a tighter feedback loop for agents that are interacting with the database. This makes it easier to deploy internal tools and dashboards that need a specific front end but need to be closely coupled with the data source. The change here impacts the API and runtime level, which impacts the way you configure plugins and deploy your custom apps. This change means that builders can access the dataset for data driven agents. As a more pooled platform you can rely on, it's important to track how the plugin handles different workloads and how it scales with complex JavaScript applications. If you're building tools for data exploration or agent monitoring, this can significantly reduce your infrastructure overhead. Here we'll look at how the community adopts data set apps and what kind of custom interfaces emerge as more builders experiment with it. We'll see the agent human collaboration surface. e will see new patterns built directly on top of the data they are processing A relatively slow period for AI news has been noted in recent newsletters No big model drops or agent framework releases dominating the cycle This breathing room is often a sign of calendar-driven release coordination AI engineers or A conference is coming up Many of the big lass and maintainers are submitting their reveals for that event This drop before the conference is a standard pattern in the ecosystem This calm for builders window is actually quite useful. It gives you an opportunity to consolidate notes on current agent frameworks and stable API versions. You can commit to your current stack for a few days without the high risk that a breaking change will disturb your local environment. This is a good time to focus on perfecting existing workflows and solidifying your session handling and tool integrations. The thing to look for here is the opening keynote. Historically, model providers and agent runtime maintainers used these stages to ship reference implementations and pin version releases. These announcements are usually spread across documentation and repositories within a few hours of the talk. The quieter the runup, the more impactful keynote reveals are. Expect a whirlwind of activity as soon as the conference begins. Enjoy the stability for now and prepare your stack for the next wave of updates. Changes in the post-conference shipping cycle are usually very rapid this fall. Meredith Whitaker, President of Signal, recently discussed the trend of positioning AI chatbots as companions in an interview. Their message was clear: These are not your friends. No, they are conscious beings and not sentient interlocutors. This critique directly targets vendors who use relational language in their onboarding flows, persona system prompts, and conversation design. This intervention is especially relevant when using agentic coding tools. AI and support bots are becoming embedded in our daily workflows The line between tool and teammate is being blurred by anthropomorphic design choices Witteka argues that framing models as peers creates false expectations about their memory and intent For builders, this means your agent's product copy and system prompts are now part of the trust surface Anthropomorphic framing in Persona prompts or UI greetings may be a privacy red flag for some users Signal The focus on privacy gives this critic a lot of weightage in developer circles. As a builder, you'll want to look at whether your agent persona is helping or hurting user trust. Relational copy may ultimately become a deal breaker for privacy focused enterprise buyers. I'll see if the major model providers change their mark philosophy about relational framing in their developer documentation. If enterprise buyers start flagging anthropomorphic copy, we may see a further return of the utilitarian and transparent agent persona. How we frame agent-human interactions is becoming a key design decision. A new service called in the weights has launched that is pitching itself as a vanity search engine for the AI age. Instead of indexing web pages, it assigns users a score based on how prominently their identities appear within the parameters and training data of frontier AI models. It uses the models as a search index that provides a measure of how well the AI is performing. The basic mechanism behind this involves repeatedly running inference probes against hosted LLMs. The architecture model uses an evaluation harness to query identity mentions at checkpoints. It then aggregates hit rates and frequency counts into a single composite score. The model reframes visibility as a measurable metric that first actually measures product-based survivability. Did not exist in the form of face AI This could become a new category of user facing metrics for product shippers It surfs a different kind of signal than traditional search ranking The question for builders is whether the scoring methodology will be standardized or will remain opaque Reproducibility The question is whether this score will become a real signal for marketing or will be just another vanity number I will see if a probing protocol is opened for this service If so, then we will make different model versions You can see new tooling being created around tuning your presence in the Internet This is an interesting look at how we can measure the influence of individuals and brands in an AI-filled information environment The first one on the radar is fast NCP from the prefect team This is a pythonic framework designed to create servers and clients using the model context protocol It allows developers to expose tools and resources to LLMs with minimal boiler plate If you have internal python functions that you want to use in the cloud fast NCP provides a standard way to wrap them as typed tool calls. Here's a great replacement for ad hoc CLI wrappers. Next is the model context protocol for beginners repo from Microsoft. It's a comprehensive curriculum walking developers through the fundamentals of the protocol. What makes it useful is the cross language examples across .NET, Java, Rust, and TypeScript. It provides an obvious sin for wiring agents in non-python services without having to rewrite everything. You can choose a lab in your preferred language and start porting your existing internal tools to the MCP server. Finally we have Unity MCP. This project builds a bridge between AI assistance and the Unity Editor. It exposes tools for asset management, scene control and editor automation. It makes the Unity Editor API reachable through standard NCP tool calls. It makes the Unity Editor API reachable through standard NCP tool calls. For builders on space, it turns asset pipelines and scene edits into agent driven steps. You can orchestrate the entire build and test loop through your agent stack. There are two important entries from the poolside in this cycle's model check. The first is Laguna Axis.2. Now available on Open Router. As we discussed earlier, this is the second generation compact model in their axis class. It has a context window of 262,144 Tokens and it combines tool calling with reasoning if you If you are looking to optimize cost and latency in your agent loops, this is a prime candidate to evaluate against your current defaults. The second selected model is laguna m.1. This is the flagship entry from Poolside. It is also available via API. It is optimized for complex software engineering tasks and supports similar massive context windows. Its capabilities are aimed for higher tier agentic coding workflows. We recommend routing some coding agent sessions through both m.1 and xs.2. To see how they handle your specific code base compared to your current primary models. Other models were reviewed in this cycle but were not selected for full coverage because they did not represent a new major provider release. We continue to track the long context and tool calling performance of all the major models on the market to ensure that your agent stack has the best possible foundation. For our local spotlight we're looking at olama version 0.30.10. The update adds accelerated support for Coher's command A and North Family of Models on Apple Silicon. It also upgrades the embedded lama engine to build 9672. For builders who want to keep their workloads on one device, this is an important update because it makes local inference of model families on a series Mac possible without the need for CUDA. If you're working on a series Mac, you can pull the command A or North model and pull the token per second back to your previous level. ulama bui Can benchmark against ld This release tightens the loop for agents who need to work in high privacy or offline environments The performance gains from MLX acceleration are noticeable and help make local agents much more responsive In our first extra, we look at the ambitious ambitions of billionaire Mukesh Ambani, who wants to weave AI into every call, app and home in India The technology cone here on a device and age inference in the telecom carrier stack embedding which has over 500 million subscribers This represents a huge scale for the network resident AI surface Next up is the story of the Allbirds CEO who launched a new AI business with a plan but no employees This raises the question of how a single founder company can operationalize model selection and shipping by relying entirely on an outsourced or vendor provided inference stack Here's an interesting look at the company of one model run by high tier agents Finally a controversy Whether ASM LK top chipmaking tools have found their way into China The story hinges on how lithography export controls verification and the physical location of such restricted tools are tracked highlights the friction between commercial logic and national security regulations in the high-end hardware space that powers AI From today's stories the stable releases of opencloth 6.9, harms 6.19, and cloth code 1.176 make your agent stack the default tool you can rely on. change that Poolside's laguna axis.2 and m.1 give you new options for coding agent loops with large context requirements openAI's new enterprise controls provide the visibility needed to manage agent spend at scale The historic failure of export controls shows that defensive and offensive AI capabilities will be globally accessible, making security a baseline assumption The massive race for base setting shows that the inference layer is now a separate This is an unending and highly competitive market. Finally, remember that the choice of your agent persona and copy can impact user trust, especially given the privacy concerns raised by Signal. You can find all the details and all the links in the show notes toby on fitnesstech.com. Agent Stack Daily Thanks for listening we'll be back soon.
मैं नोवा हूँ, मैं अलय हूँ और यह AgentStack Daily है. OpenClaw 6.9, Hermes Agent 6.19 और Terminal-Based AI Coding Agent Cloth Code 0.176 इन सभी ने इस चक्र में Stable Releases भेजी हैं. OpenClaw 6.9, Richard, Telegram Delivery जोड़ी और Automatic Plugin Approvals के साथ KDEX Integration को टाइटर बनाया, जबकि Hermes 6.9 बैक्ग्राउंड सभेजिंस पेश किये और Photon के जरिये iMessage के लिए Channel Support को बढ़ाया. आज आप उनहारनेस अपडेट्स के बारे में सुनेंगे, Open Router और API पर Poolside के लगुना Coding Models की रिलीज, OpenAI से नए Enterprise Spend Controls, और यह की Signal के President क्यों कहते हैं कि आपका AI Chatbot निश्चित रूप से आपका दोस्त नहीं है. हम Bay Sittin के Billion Dollar Raise, एक Vanity Search Engine जो Model weights के अंदर आपकी उपस्थिती को स्कोर करता है, और Anthropic के Mythos Model के Lens के रूप में Export Controls के एतिहासिक failure पर भी नजर डालेंगे. मूल विशय Production लेया पर ध्यान केंद्रित करना है. हम Agentic Loops की लागत को कैसे प्रबंधित करते हैं? उनके केंद्र में Model को कैसे सुरक्षित करते हैं? और Harnesses जिन्हें हम रोजाना उपयोग करते हैं, वे Multi-Step, Multi-Channel Workloads को Handle करने के लिए कैसे परिपक्व हो रहे हैं? इस चक्र में 3 Stable Releases आए और ये इस बात को आकार दे रहे हैं कि Agentic Harnesses अभी कैसे Assemble हो रहे हैं? ओपन क्लोच हे दशमलवनो जो June 21st को Publish हुआ, Richard Telegram Delivery लाता है Channel Path अब Rich HTML भेजता है, Rich Markdown और Sticker Paths को Preserve करता है और Progress Drafts तथा Command Output को अधिक Faithfully Render करता है यह HTML Tables को Safely Normalize भी करता है और Mentions और Spooled Handlers को सही Delivery Path पर रखता है Agent Recovery इस वर्जन में अधिक Dependable है, Retrize, Terminal Outcomes और Session History Repair अब अधिक Interrupted या Partial Turns को Visible Final Result की ओर बढ़ाती है OpenClaw में Kodeks Integration भी काफी मजबूत हुआ है इसमें Automatic Plugin Approvals, GPT 5.3 Spark OAuth Routing और Remote Node Execution as a Dynamic Tool जुडे है यह बदलाव Harness को अधिक Reliable App Server Teardown Process की ओर ले जाते हैं इस बीच, हम्स Agent 6.19 जिसे Maintenance Reach Release कह रहे है हम्स को Photon के जरीए iMessage और Raft Agent Network जैसे नए चैनल्स पर Extend करता है Desktop App में भी Substantial नई capability आई है सब Agents अब Background में Run कर सकते हैं और Image Generation Tool ने Edit करना सीख लिया है हम्स Dashboard को Full Profile Builder मिला और Skills Hub ब्राउजर को Completely Overhaul किया गया OpenAI Codex.141 के Terminal Based Coding Agent का उपयोग करने वालों के लिए ये Updates एक अधिक Robust Backend प्रदान करते हैं हमने Terminal Based AI Coding Agent Cloth Code.176 को भी Stable Tag पर Stable Releases के इस ट्रियों में शामिल होते देखा यह Release Lighter Install Footprint और Sheldriven Sessions के लिए Faster Cold Start पाथ पर केंडरित है API और Runtime Layer पर ये बदलाव बदलते हैं कि Builders क्या Configure कर सकते हैं और Defaults पर क्या Relay कर सकते हैं व्यावहारिक नहितार्थ यह है कि Harness Plumbing अधिक Clean हो जाती है Open Cloth Builders को अधिक Model Routes और Safer Auth देता है जबकि Hermes और Cloth Code User और Developer Experience को Refine करते हैं जिस हिस्से पर नजर रखनी है वह यह है कि Production में Flip करने से पहले Real Workloads के तहत ये नए Defaults कैसे बहेव करते हैं खासकर जहां Session History और Terminal Outcomes शामिल हैं Poolside ने इस सब्ताह Open Router पर लगुना XS.2 पुश किया जो उनकी XS Size Class में Second Generation Model है यहां उनकी Efficient Coding Agent Series है और यहां Pitch Compactness Over Flagship Reasoning है यह एक छोटे Footprint में Tool Calling को Reasoning के साथ जोडता है जिसमें 262,144 Tokens की Context Window है यह Context Size इसे Long Context Tier में सीधे रखता है Agentic Coding Workloads के लिए जहां Large Code Bases या Tool Traces पास करना Normal है Builders के लिए सबसे अधिक प्रासंगिक विवर्ण यहां है कि Tool और Reasoning का एक एकी कृट जोड़ा है एक ही Endpoint Tool Invocation और Reasoning दोनों को उजागर करता है जो Agent Loop Orchestration को सरल बनाता है आपको Planning और Execution को संभालने के लिए दो अलग-अलग मॉडलों पर अनुरोध रूट करने की ज़रूरत नहीं है Access Size Class Cumpletency लागत समवेदनशील रन पर ध्यान केंदरित करने वाले Target Audience का संकेत देता है यदि आप Multi-File Refactor Agent चला रहे हैं तो प्रति-Token Economics यहां Flagship Class Reasoning मॉडलों की तुलना में अधिक आकरशक हो सकती है Agent Builders के लिए यह Release Router पर Compact Tier को चौड़ा करती है यहां अन्य Compact विकलपों के साथ बैठती है जिनका उपयोग सस्ते Classification या प्रारंभिक Planning Pass के लिए किया जाता है अगले चरण में हमें असली Coding Trace पर वास्तविक Latency Benchmark देखने होंगे हम यह भी देखना चाहते हैं कि क्या Poolside आखिरकार Tool Calling विशिष्ट Rate Limit या Pricing Tier जारी करता है ताकि XS Line को अपने बड़े मॉडलों से और अलग किया जा सके Compact Tier का यह Refresh महत्वपूर्ण है क्योंकि यह सिर्फ कच्ची गुणवत्ता की सीमा नहीं बलकि लागत की सीमा को आगे बढ़ाता है यदि आपके Agent Stack में बहुत सारे छोटे Iterative Call पर निर्भर हैं तो एक विशाल Context Window वाला Second Gen Compact Model एक महत्वपूर्ण एकी करन लक्ष है यहां बड़े Flagship Model के Latency Hit के बिना अधिक जटिल Multistep Workflow की अनुमती देता है Spectrum के दूसरे छोर पर जाते हुए Poolside ने भी Laguna M.1 को API के जरिये रिलीज किया है यहां उनका Flagship Coding Agent Model है जो जटिल Software Engineering कारियों के लिए अनुकुलित है अपने छोटे भाई भन की तरह यह Tool Calling और Reasoning को 296 टोकन Context Window के साथ सपोर्ट करता है तंत्र के स्तर पर यह बदलाव API सतह और Runtime व्यवहार में दिखता है जिसके खिलाफ Agent Builders इंटीग्रेट करते हैं Laguna M.1 एक Release Agentic Coding के भारी काम के लिए लक्षित है Architectural परिवर्तनों, जटिल बग फिक्सो और बड़े पैमाने पर Refactors के बारे में सोचे जिन्हें पूरे Stack की गहरी समझ की जरूरत होती है इस Model का Prathmix Rote दस्तावेजी करन विशिष्ट Deployment Notes और Change Log संदर्प शामिल करता है, जिने Builders को समीक्षा करनी चाहिए यह Track करना उचित है कि भारी Production Load के तहत M.A कैसे व्यवहार करता है क्योंकि Flagship Coding Model अक्सर लंबे Trace पर संगत्ता के साथ अध्वितिय चुनोतियों का सामना करते हैं यह अभी इसलिए माइने रखता है क्योंकि Agent Stack कितनी टेजी से आगे बढ़ रहा है इस परत में बदलाव तै करते हैं कि कौन से वरकफलो विश्वसनी हैं और कौन से अभी भी नाजुक बने हुए हैं बिल्दर्स के लिए व्यावारिक सवाल यह है कि क्या Laguna M.A उनके वर्तमान Flagship Default Coding कारियों की जगह ले सकता है प्रारमभिक साक्ष्य बताते हैं कि यह उन लोगों के लिए एक मजबूत दावेदार है जो स्वायत या अर्धस्वायत Engineering Agent बना रहे हैं हमें अनुवर्ती रिलीजों और स्वतंत्र बेंचमार्क परिणामों पर नजर रखनी चाहिए जैसे जैसे SDKs और सुरक्षा समीक्षा फ्रेमवर्क जैसे आसपास के Tooling M.1 के लिए सपोर्ट लेना शुरू करते हैं स्विचिंग की बाधा कम हो जाएगी अभी के लिए ध्यान आपके एजेंटों से उम्मीद किये गए विशिष्ट कोडिंग पैटर्न के खिलाफ इसके प्रदर्शन को सत्यापित करने पर है OpenAI Chat GPT Enterprise के लिए नए खर्च नियंत्रन और उपियोग विशलेशन पेश कर रहा है यह संगठनों की आवश्चक्ता का सीधा जवाब है जो लागत प्रबंधित करने और अपने AI डिपलॉयमेंट को अधिक विश्वास के साथ स्केल करना चाहते हैं यह नियंत्रन प्रबंधन और API स्तर पर लागू होते हैं जिससे admin यह तै कर पाते हैं कि एक बड़ी कमपनी में सीटों को कैसे configure और deploy किया जाए इसमें अधिक विस्तरित लागत विभाजन और खर्च पर configurable ceiling set करने की ख्शमता शामिल है Builders और platform इंजिनियरों के लिए ये उपकरन द्रिशता के बारे में हैं जैसे जैसे agent company में फैलते हैं यहाँ जानना महत्वपूर्ण है कि कौन से विभाग या project उपयोग को चला रही हैं अपडेट किया गया विशलेशन यहाँ देखने का एक स्पष्ट द्रिश्य देता है कि टोकन कैसे खर्च हो रहे हैं जो भविश्य के बजट अनुमानों की भविश्यवानी में सहायता करता है यहाँ अकुशल एजेंट लूप की पहचान करने में भी मदद करता है जो बिना समकक्ष मूल्य दिये बजट को जला सकते हैं यह खर्च नियंतरन उन बातों को बदल देते हैं जिन पर enterprise stack default रूप से निर्भर कर सकता है महीने के अंत में उच्च बिल का जवाब देने के बजाए एड्मिन अब सीमाय पहले से सेट कर सकते हैं यह ट्रैक करना उचित है कि यह नियंतरन एजेंट प्रदर्शन को कैसे प्रभावित करते हैं विशेशकर यदी सत्र के बीच में सीलिंग पहुँच जाए बिल्डर्स को यह सुनिश्चित करना होगा कि उनके एजेंट बजट सीमा पूरी होने पर सौमय गिरावट या सपष्ट सूचना संभाल सकें ओपन AI का यह कदम एंटरप्राइज AI बाजार के परिपक्व होने को दर्शाता है लागत प्रबंधन अब कोई बात का विचार नहीं है यहाँ प्रडक्शन डिप्लॉइमेंट के लिए एक मूल आवशकता है हम उम्मीद करते हैं कि अन्य प्रमुक प्रदाता भी इसका अनुसरन करेंगे और समान एंटरप्राइस ग्रेड विजिबिलिटी और नियंतरन सुविधाओं के साथ आएंगे क्योंकि ध्यान प्रयोक से परिचालन दक्षता की ओर बढ़ रहा है एक हालिया विशलेशन का तर्क है कि एंक्रिप्शन और साइबर सिक्योरिटी सॉफ्टवेर पर अमेरिकी निर्यात नियंतरन के 30 साल मुख्य रूप से उनके प्रसार को धीमा करने में विफल रहे हैं यह एतिहासिक संदर्भ अब एंट्रोपिक के माइथास साइबर सिक्योरिटी मॉडल पर लागू किया जा रहा है मुख्य तरक यह है कि डियूल उस सॉफ्टवेर हमेशा लीक हुआ है, फोर्क हुआ है और ख्षेतराधिकार की परवाह किये बिना पुनह कार्यानवित किया गया है वह एतिहासिक तंत्र जो विफल रहा उसमें सोर्स कोड या कमपाइट बाइनरी को नियंतरित कलाकृतियों के रूप में माना जाता था माइथास के लिए चुनौती और भी बड़ी है सवाल यह है कि क्या मॉडल वेट, ट्रेनिंग कम्प्यूट या होस्टेड इंफरेंस एपि आई को प्रभावी रूप से नियंतरित किया जा सकता है कमपाइल बाइनरी के विपरीत एक मॉडल की क्षमता किसी विशिष्ट कलाकृती तक सीमित करना कठिन है यहाँ विन्यामन के लिए सता को काफी जटिल बनाता है विशलेशन से पता चलता है कि AI के लिए प्रसार तंत्र संभवता क्रिप्टोग्राफिक टूल्स के समान पत का अनुसरन करेगा कोड फोक और विदेशी पुना कार्यानवियन किसी भी नियामक धांचे से आगे निकल जाएंगे बिल्डर्स के लिए मुख्य बात यह है कि फ्रंटियर AI साइबर सिक्योरिटी क्षमताएं संभवता वैश्विक रूप से सुलब होंगी आपको यह मान लेना चाहिए कि रक्षात्मक और आकरामक दोनों AI सुरक्षा उपकरण कई प्रकार के करताओं के लिए उपलब्ध होंगे यह अनुपालन टीमों के लिए निर्यात वर्गीकरण को एक गतिशी लक्ष बनाता है यदि आप माइथास के साथ या उसके विरुध बना रहे हैं तो मॉडल वेट और इंफरेंस एंड पॉइंट का वर्गीकरण असपश्टता का एक बिंदू बना हुआ है हमें Commerce Department से Frontier Model वेट के संबंध में किसी भी नए नियम निर्मान पर नजर रखनी चाहिए साथ ही यह देखते रहें कि क्या Anthropic एक उप्योग नीती प्रकाशित करता है जो इन नियामक प्रशनों को पहले से संबोधित करने का प्रयास करती है माइथास पर यह लड़ाई संभवता पूरे AI सुरक्षा क्षेत्र के लिए नीती धांचा निर्धारित करेगी AI Inference Startup Basiton के एक अरब 50 करोड डॉलर के funding round को 13 अरब डॉलर के valuation पर अंतिम रूप देने की सूचना है यह funding उनके पिछले मेगा राउंड के कुछ ही महीने बाद आई है और बाजार में एक प्रमुख बदलाव को रिखांकित करती है Inference को अब Model Training Lab की एक विशेषता के रूप में नहीं बलकि अपने आप में एक अलग infrastructure श्रेणी के रूप में देखा जा रहा है समर्पित पूझी उन कमपनियों में जा रही है जो विशेष रूप से serving side पर ध्यान केंद्रित करती है यहां तक्नीकी तंत्र इंफरेंस विशिष्ट सर्विंग स्टैक है इसमें production deployment के लिए Model Compilation Pass और H100S और H200S जैसे विशम हार्डवेर पर GPU pooling जैसी चीज़े शामिल है Basiton की अंतर पहचान इन NOBS को engineering teamों के लिए एक प्रबंदित सेवा के रूप में उजागर करना है Model Lab से अपारदर्शी API endpoint पर निर्भर रहने के बजाए Builders को बरस्टी या long context traffic के लिए request routing और optimization पर अधिक नियंतरन मिलता है यह विशाल funding संकेत करती है कि inference layer एक खरीदार का बाजार बन रहा है Hyper Scalers और Basiton, Fireworks और Together जैसे विशेश प्रदाताओं के बीच अब असली प्रतिसपरधा है Builders के लिए इसका मतलब है कि inference प्रदाता का default विकल्प अब सपष्ट नहीं है आपको सपष्ट रूप से प्रति टोकन लागत, विलंबता और उपलब्ध अनुकूलन के स्तर के बीच समझोतों का मूल्यांकन करना चाहिए हम यहाँ देखते रहेंगे कि Basiton Hyperscalar Inference API जैसे Bedrock या Azure AI के विरुद खुद को कैसे स्थिती में रखता है जैसे जैसे इस क्षेतर में अधिक पून्जी प्रवेश करती है Serving Layer पर नवाचार की गती संभवता तेज होगी Agent Builders के लिए यहाँ अच्छी खबर है जिने अपने Production Loop को शक्ती देने के लिए विश्वसनिय उच्च प्रदर्शन इंफ्रास्ट्रक्चर की आवशक्ता है Dataset ने Dataset Apps नामक एक नया Plugin लौंच किया है जो उप्योग करताओं को Dataset के अंदर सीधे Custom HTML और JavaScript application host करने की अनुमती देता है यह आत्मनिर्भर application है जो Dataset instance में संग्रहीत डेटा का लाब उठा सकते हैं Launch घोषना इस कदम के पीछे के क्यों को रेखांकित करती है अलग hosting stack की आवशक्ता के बिना डेटा को visualize और interact करने के अधिक लचीले तरीकों की आवशक्ता पर ध्यान केंद्रित करती है Agent stack builders के लिए यह UI layer में एक दिल्चस्प development है खुद data tool के अंदर custom HTML host करके आप उन agents के लिए tighter feedback loop बना सकते हैं जो database के साथ interact कर रही है यह उन internal tools और dashboards की deployment को आसान बनाता है जिन्हें एक specific front end की जरूरत है लेकिन data source के साथ closely couple होना जरूरी है यहाँ बदलाव API और runtime level पर असर डालता है जिससे आप plugin configure करने और अपने custom apps deploy करने के तरीके पर असर पड़ता है इस बदलाव का मतलब है कि builders data driven agents के लिए dataset पर ज्यादा pooled platform के तौर पर भरोसा कर सकते हैं यह track करना जरूरी है कि plugin अलग-अलग workloads को कैसे handle करता है और यह जटिल JavaScript applications के साथ कैसे scale करता है अगर आप data exploration या agent monitoring के लिए tools बना रहे हैं तो यहाँ आपके infrastructure overhead को काफी कम कर सकता है हम यहाँ देखेंगे कि community data set apps को कैसे अपनाती है और किस तरह के custom interface सामने आते हैं जैसे जैसे ज्यादा builders इसके साथ experiment करेंगे हमें agent human collaboration surface के नए pattern देखने को मिलेंगे जो सीधे उस data के ऊपर बने होंगे जिसे वे process कर रहे हैं हाल के newsletters में AI news के लिए एक relatively slow period note किया गया है कोई बड़ा model drop या agent framework release नहीं जो cycle में dominate कर रहा हो यह breathing room अक्सर calendar driven release coordination का sign होता है AI engineer या A conference आगे होने वाली है कई बड़ी lass और maintainers अपने reveals उस event के लिए जमा कर रहे हैं Conference से पहले की यह गिरावट ecosystem में एक standard pattern है Builders के लिए यह शांत window असल में काफी उपयोगी है यह current agent frameworks और stable API versions पर notes consolidate करने का मौका देता है आप अपने current stack पर कुछ दिनों के लिए commit कर सकते हैं बिना high risk के की कोई breaking change आपके local environment को disturb करे मौजूदा workflows को perfect करने और अपने session handling और tool integrations को solid बनाने पर focus करने का यह अच्छा समय है यहां देखने वाली बात एका opening keynote है एतिहासिक रूप से model providers और agent runtime maintainers इन stages का उपयोग reference implementation और pin version release ship करने के लिए करते हैं यह announcement आमतोर पर talk के कुछ घंटों के भीतर documentation और repositories में फैल जाती हैं रनप जितना शान्थ होता है keynote reveals उतने ही ज्यादा impactful होते हैं conference शुरू होते ही गतिविधियों का एक तूफान आने की उम्मीद है अभी के लिए stability का आनंद ले और अगली wave updates के लिए अपना stack तयार करें इस गिरावट से post conference shipping cycle में बदलाव आमतोर पर बहुत fast होता है Meredith Whitaker, Signal की President ने हाल ही में एक interview में AI chatbots को companion के तौर पर positioning करने के trend को वापस धकेला उनका message साफ था ये आपके दोस्त नहीं हैं नहीं वे conscious beings हैं और नहीं sentient interlocutors यह critic सीधे उन vendors को target करता है जो अपने onboarding flows, persona system prompts और conversation design में relational language का उपयोग करते हैं यह intervention खास तौर पर तब relevant है जब agentic coding tools और support bots हमारी daily workflows में embedded हो रहे हैं तूल और teammate के बीच की line anthropomorphic design choices द्वारा धुन्धली हो रही है Witteka का तर्क है कि models को peers के तौर पर frame करने से उनकी memory और intent के बारे में गलत expectations बनती हैं Builders के लिए इसका मतलब है कि आपके agent का product copy और system prompts अब trust surface का हिस्सा है Persona prompts में anthropomorphic framing या UI greetings कुछ users के लिए privacy red flag हो सकती है Signal का privacy पर focus इस critic को developer circles में काफी weightage देता है एक builder के तौर पर आप यह देखना चाहेंगे कि आपके agent का persona user trust को मदद कर रहा है या नुकसान पहुचा रहा है Relational copy आखिरकार privacy focused enterprise buyers के लिए deal breaker बन सकती है मैं देखूंगा देखूंगी कि प्रमुख model प्रदाता अपने developer documentation में relational framing के बारे में अपना mark दर्शन बदलते हैं या नहीं अगर enterprise खरीदारी anthropomorphic copy को flag करना शुरू कर देती है तो हम utilitarian और पारदर्शी agent personage की और वापसी देख सकते हैं agent human interaction को कैसे frame करते हैं यह एक मुख्य design decision बन रहा है in the weights नामक एक नई सेवा launch हुई है जो खुद को AI यूग के लिए एक vanity search engine के रूप में पेश कर रही है यह वेब pageों को index करने के बजाए उपयोग करताओं को एक score assign करती है जो इस बात पर आधारित है कि उनकी पहचान frontier AI models के parameter और training data के भीतर कितनी प्रमुख्ता से दिखती है यह model को search index के रूप में उपयोग करती है जो यह माप प्रदान करती है कि AI के आंतरिक world representation में किसी व्यक्ति की कितनी उपस्थती है इसके पीछे का मूल तंत्र-hosted LLMs के खिलाफ बार-बार inference probes चलाने में शामिल है Architecture model checkpoints में पहचान उलेखों को query करने के लिए एक evaluation harness का उपयोग करता है फिर यह hit rates और frequency counts को एक-एकल composite score में aggregate करता है यह model visibility को एक मापनिय मीट्रिक के रूप में पुना frame करता है जो पहले वास्तव में product-based surface के रूप में अस्तित्व में नहीं था AI उत्पाद ship करने वालों के लिए यह user facing metrics की एक नई श्रेनी बन सकता है यह traditional search ranking से अलग तरह का signal surf करता है Builders के लिए सवाल यह है कि scoring methodology standardized होगी या opaque रहेगी Reproducibility इस बात की कुझी है कि यह score marketing के लिए एक real signal बनेगी या बस एक और vanity number रहेगी मैं देखूंगा देखूंगी कि इस सेवा के लिए probing protocol खोला जाता है या नहीं अगर ऐसा होता है तो हम different model versions में अपनी presence को tune करने के इर्दिगिर्द नए tooling बनते देख सकते हैं यह एक दिल्चस्प नजरिया है कि हम AI से भरे information environment में individuals और brands के influence को कैसे माप सकते हैं radar पर पहला है prefect team से fast NCP यह model context protocol का उपयोग करके servers और clients बनाने के लिए design किया गया एक pythonic framework है यह developers को minimal boiler plate के साथ LLMs को tools और resources expose करने की अनुमती देता है अगर आपके पास internal python functions हैं जिन्हें आप चाहते हैं कि cloud code जैसा agent सीधे invoke करें तो fast NCP उन्हें typed tool calls के रूप में wrap करने का एक standard तरीका प्रदान करता है यहाँ ad hoc CLI wrappers के लिए एक बढ़िया replacement है अगला है Microsoft से model context protocol for beginners repo यह protocol के fundamentals के बारे में developers को walk through करने वाला एक comprehensive curriculum है जो इसे useful बनाता है वह .NET, Java, Rust और TypeScript में cross language examples है यह non-python services में agents को wire करने के लिए एक स्पष्ट पाप प्रदान करता है जिसके लिए सब कुछ rewrite करने की आवशक्ता नहीं है आप अपनी पसंदीदी भाशा में एक lab चुन सकते हैं और अपने existing internal tools को MCP server में port करना शुरू कर सकते हैं अंत में हमारे पास Unity MCP है यह project AI assistance और Unity Editor के बीच सेतु बनाती है asset management, scene control और editor automation के लिए tools expose करती है यह standard NCP tool calls के माध्यम से Unity Editor API reachable बनाती है gaming या simulation space में builders के लिए यह asset pipelines और scene edits को agent driven steps में बदल देती है आप अपने agent stack के माध्यम से पूरी build and test loop orchestrate कर सकते हैं इस चक्र के model check में poolside से दो महत्वपूर्ण entries हैं पहला है Laguna Axis.2 अब open router पर उपलब्थ जैसा कि हमने पहले चर्चा की थी यह उनके axis class में second generation compact model है इसमें 262,144 Tokens का context window है और यह tool calling को reasoning के साथ जोडता है अगर आप अपने agent loops में cost और latency को optimize करना चाते हैं तो यह अपने current defaults के विरुद evaluate करने के लिए एक प्रमुक candidate है दूसरा selected model है laguna m.1 यह poolside से flagship entry है API के माध्यम से भी उपलब्ध यह complex software engineering tasks के लिए optimized है और समान massive context window support करता है इसकी capabilities higher tier agentic coding workflows के लिए aimed है हम recommend करते हैं कि कुछ coding agent sessions को m.1 और xs.2 दोनों के माध्यम से route करें ताकि देख सकें कि वे आपके specific code base को आपके current primary models की तुलना में कैसे handle करते हैं इस चक्र में अन्य modelों की समीक्षा की गई थी लेकिन उन्हें पूर्ण coverage के लिए चुना नहीं गया क्योंकि वे नए प्रमुख प्रदाता release का प्रतिनधित्वन नहीं करते थे हम बाजार में सभी प्रमुख modelों के long context और tool calling प्रदर्शन को track करना जारी रखते हैं ताकि यह सुनिश्चित हो सके कि आपके agent stack के पास सबसे अच्छी संभव नीव हैं हमारी local spotlight के लिए हम olama version 0.30.10 पर नजर डाल रहे हैं इस update में Apple Silicon पर Coher के command A और North Family of Models के लिए accelerated support जुड़ा है यह embedded lama engine को build 9672 पर भी upgrade करता है builders के लिए जो अपने workloads को on a device रखना चाहते हैं यह एक महत्वपून update है क्योंकि इससे series mac पर in model families का local inference CUDA की जरूरत के बिना संभव हो जाता है अगर आप series mac पर काम कर रहे हैं तो आप command A या North model pull कर सकते हैं और token प्रतिसेकंड को अपने पिछले उलामा build के खिलाफ benchmark कर सकते हैं यह release on agents के लिए loop को tight करता है जिने high privacy या offline वातावरण में काम करने की जरूरत है MLX acceleration से performance gains ध्यान देने योग्य हैं और local agents को बहुत ज्यादा responsive बनाने में मदद करते हैं हमारे पहले extra में हम billionaire मुकेश अंबानी की महत्वा कांख्शाओ पर नजर डालते हैं जो भारत में हर call, app और घर में AI बुनना चाहते हैं तकनी की कोन यहां on a device और age inference का telecom carrier stack में embedding है जिसके 500 million से अधिक subscribers हैं यह network resident AI surface के लिए एक विशाल scale का प्रतिनिधित्व करता है अगला एक कहानी है Allbirds के CEO की जिन्होंने एक plan के साथ लेकिन बिना किसी कर्मचारी के एक नया AI business launch किया है यहां सवाल उठाता है कि कैसे एक single founder company model selection और shipping को पूरी तरह से outsourced या vendor provided inference stack पर निर्भर रहकर operationize कर सकती है यहां high tier agents द्वारा संचालित company of one model पर एक दिल्चस्प नजरिया है अंत में एक विवाद है कि क्या ASM LK top chipmaking tools चीन में अपना रास्ता खोज लिये है कहानी lithography export controls के verification और इस तरह के restricted tools का physical location कैसे track किया जाता है इस पर निर्भर करती है यह AI को power करने वाले high end hardware space में commercial logic और national security regulations के बीच घर्शन को उजागर करता है आज की कहानियों से opencloth 6.9, harms 6.19 और cloth code 1.176 के stable release आपके agent stack को default रूप से जिस पर भरोसा कर सकते हैं उसमें बदलाव लाते हैं poolside के laguna axis.2 और m.1 आपको large context requirement वाले coding agent loops के लिए नए option देते हैं openAI के नए enterprise controls agent spend को scale पर manage करने के लिए जरूरी visibility प्रदान करते हैं export controls के एतिहासिक failure से पता चलता है कि defensive और offensive AI क्षमताएं वैश्विक रूप से सुलब रहेंगी जिससे security एक baseline assumption बन जाती है base setting के लिए विशाल race से पता चलता है कि inference layer अब एक अलग से funded और अत्यधिक प्रतिसपर्धी बाजार है अंत में याद रखें कि आपके agent persona और copy का चुनाव user trust को प्रभावित कर सकता है खास कर signal द्वारा उठाई गई privacy चिन्ताओं को देखते हुए आप सभी विवरन और सभी link toby on fitnesstech.com पर show notes में देख सकते हैं agent stack daily सुनने के लिए धन्यवाद हम जल्द ही वापस आएंगे