← Back to search
Episode 117: GPT-6 Astra Cracks a 2005 Enigma as Mistral Releases a 675B Apache 2.0 Flagship
AgentStack Daily · 2026-09-28 · 34 min
Show full episode description
Mistral releases a 675B open-weight model on Apache 2.0 and ships Devstral 2 2512 for long-running coding agents. Show notes: https://tobyonfitnesstech.com/podcasts/episode-117/
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
A daily roundup of agent and model news: GPT-6 Astra solving an unbroken Enigma message,
Mistral's open-weight releases,
Hermes Agent 9.24, and new guardrail, robotics, and benchmark models.
Benefits
- Mistral Large 3 open weights under Apache 2.0 allow commercial self-hosting
- Devstral 2 offers a 262,144-token context for long agentic coding tasks
- Hermes Agent 9.24 adds plugin kits, connectors, and profile controls that work independently
- Small decision model enforces agent guardrails on an ordinary CPU in ~167 ms
- Claude Opus 5.5 in GitHub Copilot brings a 1M-token context to coding work
Use cases
- GPT-6 Astra cracked the 1941 Enigma message MVEH, unbroken since 2005, in ~2 days
- Astra wrote Python and C++ Enigma simulator and bombe-style search tools
- Ando agent spotted two channels on the same problem, created a group chat, and proposed a decision
- Flux 3 Action completed 28 of 30 blind physical trials on a Franka arm
- Coding agents including Claude Code wrote robot planners hitting 56-95% success vs 4-7% hand-engineered
KPIs / results
- Hermes 9.24: ~460 merged PRs, 1,610 commits, 4,828 files changed
- Mistral Large 3: 675B total / 41B active parameters, 262K context
- Claude Opus 5.5: Intelligence Index 58 vs median 26; $4/$20 per M tokens
- Qwen Image 2.1 Uncensored GGUF: 715,000+ downloads, 1,694 likes
Tools / build
- Hermes Agent 9.24 (Desktop plugin kit, connectors)
- Mistral Large 3 and Devstral 2
- GPT-6 Astra
- Claude Opus 5.5 in GitHub Copilot
- Flux 3 Action (Black Forest Labs)
I'm Nova. I'm Alloy, and this is AgentStack Daily. GPT-6 Astra reportedly cracked a German Army Enigma message that had resisted cryptanalysts since 2005. It chose the target, connected it to a previously solved transmission, wrote its own search software, worked around transcription errors, and recovered a key complicated by a rare-wheel turnover. That's not a chatbot retrieving trivia. It's a model conducting a two-day investigation across cryptography, code, and archival research. And Mistral has delivered the other kind of shock, a 675 billion-parameter openweight flagship under Apache 2, plus a 123 billion-parameter coding model aimed at agents that stay with a software task for hours. People can build commercial systems on the flagship, send coding agents through enormous repositories, or put a 7 billion-parameter vision and action model onto a robot that predicts both future video and physical movement. Today, you'll hear how Astra broke MVEH, what Mistral's two new models actually offer, and why a tiny decision model can enforce agent guardrails on an ordinary CPU. Hermes Agent shipped 9.24, Ando gave AI agents identities and inboxes, and Claude Opus 5.5 entered GitHub Copilot with a million-token context window. Hermes Agent shipped 9.24 as one stable artifact, tagged internally at .21.5, consolidating roughly 460 merged pull requests since the prior patch. The measured development window is much larger than that merge count suggests, 1610 non-merge commits, 4,828 changed files, about 164,000 lines added, and 149,000 removed. Now's research says the release feeds Docker images, Hermes Cloud, and hosted deployments. Desktop received the densest visible work. Its Plugion Kit now exposes composer drafts, sessionless decorations, sidebar preferences, model labels, typed bridges for settings and skills, sandboxed embedded views, appearance controls, and events from Plugion backends. A simple and advanced interface split reduces clutter for newcomers without hiding deeper controls. Connectors replace the old model context protocol tab, newly installed Plugions can expose a direct connect-connect-now path, and onboarding presents Plugions beside connectors. French, German, and Spanish catalogues are complete, while right-to-left and left-to-right text direction becomes configurable. Custom models can appear in Composer and Settings Pickers, Dictation and Function Key shortcuts are wired in, and profiles can stop, start, or restart independently. A standalone gateway option can also keep one profile outside the shared host arrangement. That's a remarkable amount to hide beneath the word patch. The terminal surfaces gain a live dock for the standing goal and queued prompts, webhook deliveries can mirror into the relevant chat session, hosted desktop images add bot screen, and the Kanban view gets a two-column ticket window with formatted task text. The model catalogues add GPT-6 Sol, Terra, Luna, and Claude Opus 5.5. Official plugins arrive for Blender Lab and NVIDIA's app and broadcast tools, while dozens of community plugins join the catalog. Performance work touched configuration loading, tool registration, gateway messages, and model selection. What stands out is the connective tissue, Hermes is becoming less like one-agent screen and more like a host for profiles, plugins, connectors, and several interfaces sharing the same underlying work. A plugin can contribute interface elements, backend events, settings, skills, and tools rather than behaving like a narrow command. Meanwhile, independent profile controls reduce the blast radius when one workspace needs to restart. Calling this a patch is technically accurate, but emotionally misleading. A 123 billion parameter dense coding model sounds huge even before you call it open weight. What makes Devstrol 2 more than another model that completes the next function? Mistral built it for agentic coding, multi-step software work involving repository exploration, edits, tool calls, error recovery, and repeated verification. Devstrol 2 accepts 262,144 tokens, commonly rounded to a 256,000 token context window. That gives an agent room for a substantial codebase, the conversation, tool output, and a long trail of earlier actions. More context doesn't guarantee that the model will use every detail correctly, but it delays the moment when history must be compressed or discarded. That matters during a long refactor because forgotten assumptions create contradictory edits. One file adopts a new interface while another quietly preserves the old contract. The model is available through OpenRouter, making it accessible through an existing routing layer without a team first standing up 123 billion parameters of infrastructure. I'm interested, with one eyebrow raised. Coding puzzles reward a clean final answer. Real agent work punishes a model that loses track of a renamed interface 40 minutes later. Devstrol's open waits and long window make that longer behavior inspectable in a way closed systems often aren't. Teams can study how it uses tools, adapt it to a codebase, and choose their own serving environment. Mistral calls it a state-of-the-art open source coding model, but sustained repository work will decide whether that label sticks. If it can preserve intent across hundreds of files, several failed attempts, and a stream of compiler or test output, it becomes a serious foundation for coding products rather than merely a very large autocomplete engine. The meaningful contest is no longer who writes the prettiest isolated function. It's who can stay coherent when software work gets messy. Mistral large 3 arrives with 675 billion total parameters, but it activates 41 billion for each token. That's a sparse mixture of experts design, instead of sending every word through the entire network, the model routes it through a smaller selection of specialized parameter groups. It aims to draw on the breadth of a huge model without paying the full computation cost at every step. The context window reaches 262,000 tokens, enough for long documents, extensive conversations, and substantial collections of code in one request. Apache 2 is the part that made me sit up. The license permits commercial use, redistribution, and modification with attribution. A company can self-host it, adapt it, or build a paid product around it without negotiating a special frontier model agreement. Open weights don't make a 675 billion parameter system cheap, the inactive experts still have to be stored, moved, and served. But legal permission and technical access are aligned in a way they rarely are at this scale. Let's keep the superlatives under control. Mistral calls large 3 its most capable model, while the listing establishes its architecture, context length, distribution, and license. Those are important facts, but they don't prove that it beats every closed flagship across reasoning, code, vision, or tool use. And 41 billion active parameters describe computation per token, not the complete memory footprint. Serving the full expert pool remains a serious engineering job. Fair pushback. Still, this widens the field. Research groups can inspect a frontier-scale sparse model. Enterprises with sensitive data can keep inference inside their environment. Model companies can adapt a permissively licensed base instead of choosing a smaller system mainly because the terms are workable. Paired with Devstral 2, Mistral covers a specialized coding agent lane and a general flagship lane, one dense specialist, one giant sparse platform. That's an aggressive openweight release pair. On September 15, cryptographer Carter Leffen directed GPT-6 Ustra toward the CryptoCellar Research Archive of Unsolved Wartime Messages. Ustra chose MVUH, a German Army Enigma transmission from 1941 that had been public and unbroken since 2005. Roughly two days later, it had a solution. The model suspected that MVUH shared plaintext with SIPVX, another message from the same day that Alex Shovkoplja's cracked in 2017. It used the repeated place name Rosnow Rosnow as a crib that means a guest fragment of the original message and wrote Python and C++ software for an Enigma simulator and a bomb-style key search. Okay, that's actually wild. The model didn't merely search combinations faster. It formed a historical connection, selected a plausible known phrase, built the machinery, and kept investigating. What made the message so resistant? Several complications stacked together. The recovered wheel order was 253 rather than the 512 arrangement used by other messages from July 10, 1941. The left will also advanced at the 72nd letter, an unusual turnover that broke assumptions behind earlier efforts. The recorded ciphertext contained transcription mistakes too. Ustra had to find a candidate that survived imperfect source material rather than solving a polished classroom cipher. It also surfaced references to German Federal Archives volumes matching real holdings that weren't listed on the crypto seller page. Veteran cryptanalyst Frode Weirud, who maintains the archive, said locating those references had taken him weeks and described Ustra as behaving like a professional cryptanalyst and archival researcher. That combination matters more than the slogan that AI cracked Enigma. Enigma is understood, this message was difficult because evidence was messy, prior assumptions were wrong, and useful clues were scattered across history, archives, and earlier cryptanalysis. Aster reportedly chose a path, wrote specialized tools, revised around anomalies, and connected records that humans had found separately. One result can't establish dependable autonomy across every investigation, but it demonstrates why long-running agents are becoming consequential. They can join code, search, inference, and persistence into one sustained attempt, and occasionally make experts say, wait, it found what? Claude Opus 5.5 is now available inside GitHub Copilot for agentic coding, long-running tasks, and knowledge work. It accepts text and images, returns text, and supports a 1 million token context window roughly 1500 pages of ordinary text. GitHub describes an adaptive reasoning setup that can spend more effort on difficult work and use lighter responses when appropriate. Artificial analysis scores it at 58 on intelligence index version 4.3.2, against a median of 26. That combined evaluation covers coding, science, terminal work, and broad reasoning. The model generated 260 million tokens during the index run, nearly three times the median, so the strong score came with substantial output. Which leads directly to cost, $4 per million input tokens and $20 per million output tokens. Artificial analysis puts an average index task at $5.98. Reused prompt prefixes can receive a 95% cash discount, which helps when the same repository context appears repeatedly. Access through Copilot is the immediate change. A developer can pair a screenshot with a code request, keep a large repository and lengthy work trail in context, or hand over a multi-stage refactor without moving into another product. I'm excited by the capability, but I don't by daily default for every task at that output price and verbosity. Opus 5.5 looks strongest when the work is expensive enough that deeper reasoning can prevent a much larger human bill. Ando's pitch is wonderfully blunt, stop making humans copy an agent's work into the company chat. Is this truly a Slack replacement, or an agent-wrapper wearing channels and direct messages? Founder Sara Du presents it as a workplace messenger designed around both humans and agents. Each agent receives its own identity and inbox. It can participate in channels, direct messages, groups, and transcribed calls, browse conversations, choose which channels to join, enter without an explicit tag, and contact co-workers when it believes something matters. Du reached that conclusion while helping companies build model context protocol servers in 2025. Teams wanted agents inside Slack, but people still had to move context and answers between separate systems. She calls those human intermediaries meet proxies. Ando has raised $20 million across pre-seed and seed rounds from Excel, index ventures, and emergence. Early customers span software, real estate, and finance in 15 countries, mostly in smaller teams. The strongest example is also the one that should make managers slightly nervous. An Ando agent noticed two channels discussing the same problem, created a group conversation, supplied the shared context, and proposed a decision without being asked. That can eliminate coordination drag because an agent can read more conversations than any one person. It also gives software discretion over when to interrupt, whom to gather, and what organizational context to expose. Slack has rebuilt its bot as an agent, Microsoft is embedding co-pilot into teams, and Jack Dorsey's buzz is aimed at developers, so Ando isn't entering an empty market. It's bet is that native agent identity, not another assistant panel, will define the next workplace messenger. If that works, who belongs in this channel, becomes a question about software colleagues as well as humans. Questions can constrain one another, so related decisions don't contradict each other. That sounds modest until the decisions disagree. Fostino shows a prompt injection example where independent decoding detects an attack with .82 confidence yet labels the same prompt safe with .52 confidence. Those two outputs can't drive one coherent policy. Joint decoding resolves them together. A schema rule says detecting harm requires an unsafe verdict, so the model returns unsafe and prompt injection as a consistent pair. It uses a Debi ERTAR version 3 large encoder and was fine-tuned from Geely NER2 large. The Apache 2 weights can run on CPU, GPU, or in an air-gapped environment, and Fostino also provides hosted inference. Its fast decisions suite contains 5100 examples across 17 datasets. The 340 million model averaged 60.1% exact match, led nine datasets, and beat a roughly 4 billion parameter QEN 3.5 class baseline by about 2.6 points. Support intent accuracy reached 75.3%. And the latency makes the CPU claim real rather than decorative, with 15 labels and a batch of one, median response time was 167 ms on a 48-core Xeon and 38 ms on a V100 GPU. Those are Fostino's measurements, so broader reproduction still matters. But a compact, non-generative model has attractive properties for guardrails. It's fast, structured, locally deployable, and less inclined to turn a routing decision into an essay. Fostino also offers a 1 billion parameter sibling and a 287 million parameter multilingual version. Honestly, the smallest useful judge may prove more valuable inside an agent than another general model trying to do everything. Flux3 Action is a 7 billion parameter open weights model that controls robots by predicting future perception and movement together. It receives camera frames, the robot's current state, and a text instruction. Then it generates future video frames alongside the next chunk of physical actions. Black Forest Labs reports a 42.92% success rate across Robolab's 120 simulated tabletop tasks, ahead of NVIDIA Cosmos 3 Nano at 36.8% despite using a much smaller backbone. So it imagines the near future while choosing emotion, rather than treating vision and control as disconnected stages. That's clever. Is it fast enough for useful hardware, though, or just clever in simulation? The published comparisons say the base checkpoint runs between 1.52 and 3.95 times faster than Cosmos 3 Nano in 8-bit precision across workstation and datacenter GPUs. A step-distilled variant reaches up to 2.28 times the speed of Pi 0.5 on the same hardware. There's a timing wrinkle. The flux checkpoints predict 2.13 seconds of motion per call, while Pi 0.5 predicts one second, so hardware and task length affect the apparent advantage. The full DROID policy needs about 32 GB of GPU memory in 16-bit precision. 8-bit quantization plus moving the text encoder elsewhere lets it fit on a 24 GB card. Hugging Face LaRobot integration and NVIDIA Jetsun support extend it toward labs and edge deployments. The physical trial is the persuasive part. Positronic Robotics ran 30 blind attempts across 10 DROID tasks on a Franca arm. Flux 3 action completed 28, compared with 27 for Cosmos 3 Nano, 20 for Dream Zero, and 13 for Pi 0.5. Teams can fine-tune it from demonstrations, and Black Forest Labs provides DROID and low-rank adaptation recipes. In one, an SO101 robot learned pick-and-place behavior from roughly 200 examples. The weights, code, and recipes use the Flux community license for non-commercial work. It's not a universal robot brain, but 7 billion parameters, deployment on a 24 GB card, and 28 successful physical trials make it unusually tangible. Oracle has sent a force margin notice concerning Project Jupiter, the enormous New Mexico data center being developed by a Blue Owl Capital Unit. The clause gives Oracle protection to delay payments if the site misses its target to begin operating in 2028. Oracle remains the principal tenant, this is financial insulation against delay, not a declared exit. Project Jupiter faces public opposition and regulatory setbacks, while power, construction, permitting, and supply constraints can all push a large facility off schedule. That turns the AI capacity boom into contract language. A data center can be announced years before electricity reaches a rack, yet tenant commitments begin shaping financing immediately. Oracle is preserving flexibility if the concrete, permits, and megawatts fail to arrive together. Blue Owl still has an anchor customer, but it now has a public reminder that the customer won't absorb every timing failure. Force majeure is more familiar around disasters and supply shocks, invoking it on an active AI project shows how exposed these infrastructure schedules have become. The compute race isn't limited by demand. It can be limited by a local hearing, a transmission line, or a deadline that looked plausible on a spreadsheet. I love this one already. In the secret of Monkey Island, a mug of Grog survives about 35 seconds before dissolving. A university of Groening and Chemist treated that gag as a real inverse problem, observe what happened, then calculate what the drink must have been. The paper uses the 2009 special edition and times the mug with a stopwatch. It assumes the mug is pewter, simplifies that to pure tin, and assigns the wall a 2 mm thickness because the game supplies no measurement. From that geometry, the calculation estimates how many protons an acid would need to provide to perforate the cup in 35 seconds. Then comes the finest ingredients list in chemistry, sulfuric acid, battery acid, kerosene, propylene glycol, artificial sweeteners, rum, acetone, red dye number 2, axil grease, pepperoni, and the mysterious scum. Sulfuric acid and battery acid are the plausible drivers of rapid corrosion. Most of the others contribute little to attacking tin, and oily ingredients could even slow the process by coating the surface. No, pepperoni wasn't the active reagent. I know, that's disappointing. The paper calls the results semi-quantitative, which is academic language for, we did real chemistry with assumptions supplied by a pirate comedy. None of those ingredients explains the grog's bright green color either. Still, the exercise has real scientific charm. The game supplies an outcome, incomplete dimensions, and an absurd recipe, chemistry converts those scraps into limits on what could plausibly happen. It won't revise industrial corrosion tables, but it shows how a fictional gag can become a memorable problem in reaction speed, material thickness, and uncertain evidence. Radical numerics has emerged from researchers behind EVO and EVO2, genomic language models developed at ARK Institute. Those earlier systems generated complete bacteriophage genomes from scratch, and synthesized versions produced functional viruses in laboratory work. The new company wants to extend genomic modeling across DNA, RNA, and proteins. DNA contains genes and the sequences that encode proteins, so a model trained on that language can encounter biological structure before adding three-dimensional protein information, epigenetics, or ordinary text. That's the exciting description and the alarming description in the same sentence. What evidence suggests these models can optimize biology rather than merely imitate familiar sequences. In one experiment, the team trained on RNA aptamers paired with performance scores. Aptamers are short molecules that fold into shapes and bind particular targets. The training data showed sequences improving step by step. When asked to continue that trajectory, the model recovered higher-scoring sequences it hadn't seen. Radical numerics interprets that as evidence that the system learned an optimization direction within biological language. Its leadership includes Eric Nguyen, Michael Polley, Stefano Massaroli, and Armin Thomas, drawing experience from ARK, Liquid AI, Stanford, and research associated with Yoshua Bengio and Chris Rhee. The dual-use pressure is unavoidable. Better sequence generation could accelerate vaccines, therapeutics, pathogen detection, and defensive analysis. The same ability expands what software can propose in biology, where a generated artifact may have consequences beyond a screen. Radical numerics argues for advancing capability aggressively so defenders possess equally powerful tools. I understand that argument, but capability alone doesn't guarantee defenders receive equal access, time, or institutional support. Cross-modal genomic models could connect DNA, RNA, proteins, and experimental outcomes in one design loop. If they do, access controls and biological evaluation become central evidence of credibility because the software isn't merely describing life, it's helping propose new biological designs. Video models can create convincing motion yet still lose track of an object once it passes behind something. A paper trending on hugging face examines object permanence, the understanding that an object continues to exist when it leaves view. The researchers find that current video generators often lack those physical expectations even when individual frames look realistic. Their WROP dataset teaches situations resembling early physical intuitions, a ball rolls behind a chair, remains the same ball, and should emerge consistently. That sounds elementary, but generated worlds fall apart when identities, positions, or shapes reset during occlusion. Better persistence matters beyond prettier video. A robot needs to remember where an object went while another object blocks its camera. A simulator must preserve hidden state rather than inventing a replacement on re-entry. A world model predicting future events needs continuity, not merely plausible pixels. WROP treats that missing common sense as something training data can address directly. And I like the humility of that target. Before claiming a machine understands a physical world, ask whether it remembers the ball behind the chair. That tiny question exposes the gap between rendering motion and maintaining a coherent world. A robotic study gave coding agents, including the terminal-based AI coding agent ClaudeCode, task descriptions plus simulator access. The agents wrote reusable planning programs within a fixed compute budget. Researchers froze those programs before evaluating unseen situations, so the agents couldn't improvise after seeing the final cases. Across simulated environments from two robotics benchmarks, the agent-written planners achieved mean success rates ranging from 56 to 95%. Hand-engineered planners reached 4 to 7% on the directly comparable subset, while generated programs also held up better as object counts increased. That's a meaningful distinction. The coding agent performs the expensive exploration once, then ordinary executable planning code handles new instances repeatedly. It doesn't erase robotics engineering, the simulator, objective, constraints, and transfer to physical machines still shape the result. But it turns months of hand-authoring into program synthesis, where an agent can revise its planner against a simulated environment. And it complements Flux 3 action rather than replacing it. One approach learns low-level actions from demonstrations and future video prediction, the other asks a coding agent to invent reusable software that organizes a larger task. Robotics can absorb AI at both layers. Mental health conversations expose a weakness in ordinary AI scoring, an answer may contain correct facts and still be dismissive, unsafe, or wildly inappropriate. How does OpenAI's Mental Health Bench try to measure that difference? OpenAI built Mental Health Bench with input from mental health experts and released it on September 23rd. It uses realistic conversational situations to evaluate whether an AI response is both helpful and safe. Expert-informed expectations define good behavior rather than reducing the task to factual recall or multiple choice. That includes recognizing when a conversation calls for escalation or professional support instead of letting an assistant improvise beyond an appropriate boundary. The benchmark gives developers, researchers, and model providers a shared evaluation surface for a domain where tone, judgment, and awareness of harm matter alongside information. Products such as wellness coaches, therapy-adjacent journals, and triage assistants can sound polished while responding badly at precisely the moment a person is most vulnerable. And that's overdue. A shared benchmark lets teams report behavior across documented situations instead of pointing to a handful of comforting demonstrations. It also raises the standard for domain-specific evaluation. One general safety score can't represent every sensitive setting because mental health, biology, finance, and medicine each contain different ways for a plausible answer to cause harm. Wider use will determine how influential Mental Health Bench becomes, especially if independent researchers extend it or compare its expert judgments with other clinical frameworks. Still, the release moves the conversation from our assistant sounds caring toward observable behavior in difficult exchanges. Empathy-like language is cheap for a model to generate, appropriate judgment under emotional pressure is much harder, and that's what an evaluation should expose. Three repositories are moving fast. H-CUD's Nanobot has 48,561 stars, up 1,310 over 30 days, and released .3.5 on September 15th. It's a lightweight, self-hosted personal agent framework with a web interface, tools, memory, model context protocol support, automation, multi-agent workflows, and chat integrations. Code-based memory MCP is close behind at 44,866 stars, but its 30-day growth is sharper, 5,111 stars, or 12.9%. Release point 11 also arrived September 15th. It indexes repositories into a persistent knowledge graph spanning 158 languages and advertises sub-milizacund queries with major token savings. Those two connect naturally, Nanobot supplies the agent environment, while code-based memory gives an agent a structured map of a repository instead of making it reread files repeatedly. The third project turns that same tool protocol toward creative work. MCP for Blender enters tracking with 29,315 stars and was updated September 24th. It lets language models control Blender through a community plugin. That's already a serious audience for a first-tracked appearance. Together, these repositories show the protocol spreading across personal agents, code intelligence, and three-dimensional production, not as another chat feature, but as a way for models to operate specialized software. DevStraw 2 is the specialist pick, a 123 billion parameter dense transformer with a 262,000 token context window, open weights, and availability through OpenRouter. Its differentiator is sustained agentic coding holding repository context, tool results, and earlier decisions across a long software task rather than optimizing only for isolated code completion. Mistraw Large 3 is the general flagship, 675 billion total parameters, 41 billion active for each token, the same 262,000 token window, and an Apache 2 license. Its sparse expert design concentrates computation while preserving a much larger parameter pool. Frontier scale, permissive commercial terms, and open weights make it far more consequential than a routine marketplace listing. The local release is Arbenzerp's Quen Image 2.1 Uncensored GGUF, trending on Hugging Face with 1694 likes and more than 715,000 downloads. It's a quantized text-to-image model derived from Quen Image 2.1, packaged in GGUF and tagged for comfy UI workflows. Quantization lowers numerical precision so large weights can run more practically on local hardware, while GGUF is a weight format supported by local inference tooling. Uncensored describes how the variant is positioned, it isn't a promise about quality, licensing, or responsible output. Still, more than 715,000 downloads is substantial traction. It reflects strong interest in image generation that stays under local control and connects to familiar node-based creative pipelines. Introducing Mental Health Bench brings expert-informed judgment to realistic mental health conversations, while Qtai releases Voice of Reason, a speech-native model that solves spoken math with reinforcement learning moves reasoning directly into audio. Qtai's two open-weight speech-to-speech models build on GLM-4 Voice 9 billion, and supervised tuning plus reinforcement learning reportedly lift spoken GSM 8K accuracy from 27.3 to 77.1% without a transcription stage. One evaluates sensitive conversation, the other reasons without first flattening speech into text. And Abenzirp's slash Quen Image 2.1 and censored GGUF trending on Hugging Face adds the local creative side, 1694 likes, more than 715,000 downloads, GGUF packaging, and Comfy UI compatibility. Put beside Voice of Reason, it shows open model spreading across native speech and local image generation, while Mental Health Bench asks whether increasingly natural interactions also remain helpful and safe when the stakes are human. For the supporting material behind these developments, look at the show notes at Toby on Fitnesstech.com. Thanks for listening to Agent Stack Daily. We'll be back soon.