← Back to search

GPT-5.6 Needs Government Approval Now, npm Finally Gets More Security, and Fable Returns | This Week In AI

Agents Hour · 2026-07-02 · 22 min
relevance 52 4219 words Episode page ↗ Audio ↗
Show full episode description
GPT-5.6 is here — Sol, Terra, and Luna — but for the first time, the US government decides who gets access, approving customers one by one. Shane and Abhi dig into what that precedent means for competition, open models, and every lab that comes next. Then a stacked week: OpenAI's own Jalapeño chip, open models pushing hard (Qwen AgentWorld, DeepSeek DSpark, Hermes mixture-of-agents), the model-routing wave (AI SDK 7, Devin Fusion), Claude Tag's viral moment and the Trojan-horse worry, voice agents on Vercel, Coinbase on keeping AI spend flat, a16z's $200M "seed" for Mirendil, npm finally freezing accounts after the Axios and Mastra attacks, Meta's think-to-text Brain2Qwerty, the DevRel Google fired over a viral CLI, and Fable's return after 15 days offline. AI Agents Hour is a weekly livestream by Mastra CPO Shane Thomas and CTO Abhi Aiyer. Mondays 12PM Pacific. 📚 READ MORE GPT-5.6 Sol/Terra/Luna: https://x.com/OpenAI/status/2070555272230384038 Restricted rollout: https://x.com/steph_palazzolo/status/2070241787180966279 Jalapeño chip: https://x.com/OpenAI/status/2069770172802773292 Qwen AgentWorld: https://x.com/Alibaba_Qwen/status/2069720365442719867 DeepSeek DSpark: https://x.com/Yuchenj_UW/status/2070928299744972814 Hermes Mixture of Agents: https://x.com/NousResearch AI SDK 7: https://x.com/vercel/status/2070155382488764566 Devin Fusion: https://x.com/cognition/status/2071624568465490170 Tau harness: https://x.com/_alejandroao Claude Tag (Karpathy): https://x.com/karpathy Claude Tag Trojan horse: https://x.com/ashwingop Open Tag: https://x.com/ataiiam GLM-5.2 in the wild: https://x.com/cline/status/2069171146994729078 Local 1-bit GLM: https://x.com/UnslothAI/status/2069418532375564484 Coinbase on AI spend: https://x.com/brian_armstrong Coinbase gateway: https://x.com/markletree Exa Connect: https://x.com/ExaAILabs/status/2069842203577651283 Mirendil $200M: https://x.com/a16z/status/2069869327411749012 npm account freeze: https://x.com/AikidoSecurity Brain2Qwerty: https://x.com/AIatMeta/status/2071566924803395741 Google Workspace CLI: https://x.com/JPoehnelt 📚 MASTRA RESOURCES Mastra: https://mastra.ai Mastra on X: https://x.com/mastra_ai Mastra Discord: https://mastra.ai/community/discord Mastra GitHub: https://github.com/mastra-ai Learn Mastra in the world's first MCP-Based Course: https://mastra.ai/course Principles of Building AI Agents (Book): https://mastra.ai/books/principles-of-building-ai-agents Patterns for Building AI Agents (New Book): https://mastra.ai/books/patterns-of-building-ai-agents WHAT IS MASTRA? Mastra is an open-source TypeScript framework designed for building and shipping AI-powered applications and agents with minimal friction. It supports the full lifecycle of agent development—from prototype to production. You can integrate it with frontend and backend stacks (e.g., React, Next.js, Node) or run agents as standalone services. If you're a JavaScript or TypeScript developer looking to build an agentic or AI-powered product without starting from first principles, Mastra provides the scaffolding, tools, and integrations to accelerate that process. ⏱️ CHAPTERS
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Weekly AI roundup covers frontier-model gatekeeping, model-routing/mixture-of-agents cost savings, and npm security fixes.
Benefits
  • npm freezes high-impact accounts 72 hours after sensitive actions
  • Mixture-of-agents routing cuts token spend while keeping intelligence
  • Open models catch up as US bureaucracy gates frontier models
  • Voice agents now available on Vercel AI Gateway
  • Agents in Slack via Claude Tag and OpenTag
Use cases
  • GPT-5.6 access approved customer-by-customer under Trump admin security review
  • Hermes Mixture of Agents presets claim 8% above Opus 4.8, 11% above GPT-5.5
  • Devin Fusion reduces Fable-level intelligence cost by 35%
  • DeepSeek DSpark speculative decoding boosts throughput 51%–400%
  • Monstra agents run in Slack; Claude Tag went viral via Karpathy
KPIs / results
  • npm freezes accounts for 72 hours after email swap or 2FA recovery
  • GPT-5.6 Sol on Cerebras at 750 tokens/second
  • Devin Fusion cuts cost 35%; DeepSeek DSpark +51%–400%
  • a16z's $200M 'seed'
Tools / build
  • Hermes Mixture of Agents virtual models
  • Devin Fusion hybrid model harness
  • AI SDK 7
  • OpenAI Jalapeno chip (with Broadcom)
  • Qwen Agent World, DeepSeek DSpark/DeepSpec
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00
Cold open: npm finally freezes accounts
  • npm freezes accounts 72 hours after sensitive actions
00:29
Welcome
  • Welcome to Agents Hour weekly news
00:37
GPT-5.6 Sol, Terra, Luna — and government-gated access
  • GPT-5.6 Sol, Terra, Luna; government-gated customer-by-customer access
04:25
OpenAI's Jalapeño chip
  • OpenAI's first AI chip Jalapeno, built with Broadcom
05:34
Open models: Qwen AgentWorld, DeepSeek, Hermes MoA
07:11
AI SDK 7
  • AI SDK 7 adds reasoning control, tool approval, MCP apps
07:43
Devin Fusion & the model-routing wave
  • Devin Fusion cuts Fable-level cost 35%; model-routing wave
10:26
Harness interop: OpenCode & LangChain
  • Harness interop: OpenCode and LangChain
11:01
Tau harness
  • Tau harness mentioned
11:26
Claude Tag's viral moment & the Trojan horse
  • Claude Tag goes viral; Karpathy praise; Trojan-horse lock-in concern
12:49
Open Tag
  • OpenTag: open-source Claude Tag for any model/harness
13:12
Voice agents land on Vercel
  • Voice agents land on Vercel AI Gateway
13:53
The infra bill: Coinbase on keeping AI spend flat
  • Coinbase on keeping AI spend flat; infra costs
15:15
Exa Connect & a16z's $200M "seed"
  • Exa Connect and a16z's $200M 'seed'
16:32
npm freezes accounts
  • npm freezes high-impact accounts for 72 hours
18:08
Meta's Brain2Qwerty: think-to-text
  • Meta's Brain2Qwerty think-to-text research
19:11
Google fired its viral CLI star
  • Google fired its viral CLI star
21:48
Fable returns
22:00
Outro
  • Outro and subscribe
NPM now freezes high-impact accounts for 72 hours after sensitive actions like an email swap or 2FA recovery code use. I'm proud of this one. Why did they do it after Axios? Is it because of us that we were bitching so much? Welcome everyone, this is Agents Hour. If you're just tuning in, every week we do the news. This week is no different. And as always, there's a lot to talk about. So we're talking about OpenAI to start. So GPT-5.6 kind of launches, I guess, some semblance of launching. And this is from OpenAI on June 26th that says, Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work. I think there's kind of limitations on who can access this, though, from the sounds of it. And that's kind of, you know, I think frustrating for some folks. So this kind of covers it a bit. The Trump admin has asked OpenAI to stagger the release of GPT-5.6 over security concerns. So on Thursday, CEO Sam Altman told staff that the government will be approving access to GPT-5.6 customer by customer. A highly unusual approach. So this is right on the back of all the Fable government regulation and subsequent shutdown a few weeks ago. What do you think about this? It's kind of sad because this might be the way models get rolled out from frontier labs as they get more powerful. I think we've set a precedent now that, you know, any powerful model that's Fable-like will have to go through this, like, security review process. And then, funny enough, on the open model side, no one gives a shit. And you can just launch things that are Fable competitive without any U.S. bureaucracy. I think that's true today. But here's my concern with this. It sets a really bad precedent in that, one, if you wanted to, if you're in the U.S. and you want to try to compete in this market, you're not going to be able to anymore. Yeah. If you're not really well-funded, you're not going to be able to go through all the bureaucracy to get a model released, unless it's across a more narrow task. Anything that's security-related or frontier level, you're probably not going to be able to get out. You're probably not going to be funded enough. So it really stifles competition. It allows a window for open models to really catch up. But also, I would not be surprised if they somehow find ways to regulate, at least in the U.S., what open models inference providers can run. Yeah. Like, open router fireworks could all get, like, GLM 5.2, they have to, like, go through some security review or whatever. Which is BS, dude. And then in hyperscalers, we might be limited to even not allow their customers to run those. So you might think, well, it's open. I can run it myself. But if the inference provider is not legally allowed to allow you to run that, you're kind of out of luck, at least maybe in the U.S. Dude, and Bedrock has this problem already, right? They don't have Opus 4.8, like, because there's a whole GovCloud issue going on already. So the bureaucracy is here, but now it's impacting the normal folks, not just the GovCloud people. Yeah. I did hear that, supposedly, I heard Bedrock might be getting 5.6 soon or something. So TBD on that, but maybe 5.6 is coming to some of those, at least in the limited preview. They'll get the Sol, not the Terra, and not the other one, you know? Yeah, exactly. What are these fucking names, dude? So speaking of names, here's kind of how they're thinking about it. Sol is like the sun, it's really big. Terra is, you know, like the Earth, it's good sized. And Luna is like a moon, it's smaller. So I guess that's how you can think about it. OpenAI has been just blasted for their naming conventions before. Maybe they're trying to be more Claude-like with Opus, Haiku, Sonnet, the same kind of thing of size. So I'm assuming that's why they're doing it. So OpenAI is following in Anthropix footsteps, I guess. I think these ones might actually stick. It gives a little bit of a framing. Yeah. It's all easier to understand than Mini. I guess Mini is Mini, but like, you know, there's three levels here now. Yeah. GPT 5.6 Sol is being launched on Cerebrus at 750 tokens per second. So that would be pretty incredible if true. Sounds like it could be relatively fast. This was a huge thing last week. OpenAI announced we've designed and built our first AI chip, Jalapeno, which is pretty funny. That's kind of a funny name, the Jalapeno chip. It's designed from the ground up by OpenAI and brought to production with Broadcom. Jalapeno is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products. What do you think of the spicy Jalapeno chip? I think it's cool that they're doing this, but it really puts their partnerships in question, like who they're getting chips from. So interesting to see what Jensen thinks or anybody else. Jalapeno is on such a tear though that there's going to be competition. Yeah. They're just moving lower level. They probably look at it as competitive advantage. Yeah, but then what if NVIDIA rescinds some allocation? I guess then they'll have to go to Colossus? I'm just kidding. They'll have to go get it from somewhere unless this chip is that powerful. So we'll see. And we have some chat saying, really hoping in the EU we'll be getting access to a 5-6. And then Mika's 3D says, I see permanent underclass everywhere on Twitter now. China, save us. All right, let's talk about open models. Meet Quen Agent World, where a native language world model simulates seven agent environments within a single model. Environment modeling is the training objective from day one, not a post hoc adaption. So this is kind of like what would a world model for agents look like? So Alibaba released this last week, got a lot of attention. DeepSeq has released something. They published DSpark, a new speculative decoding method that boosts throughput and data. By 51% to 400%. And then they've also open sourced DeepSpec, which is the training framework behind it. This is the real open AI. So DeepSeq is just releasing a lot of stuff in the open, which is great for others that are training models. And this whole idea around model routing mixture of agents is a topic we're going to keep seeing more of. But this is from the team behind Hermes. It says, the strongest models are gated and accesses granted only to a select few. Hermes agent now exposes mixture of agents presets as virtual models, giving you capabilities beyond the publicly available frontier. So 8% higher than Opus 4.8 and 11% higher than GPT 5.5 on our upcoming benchmark. Yeah, this is a trend. They don't release the benchmark. They just say, on this benchmark that we created, that's upcoming. So I can only give you partial credit here, but I do think that this is going to be a trend. I think you're going to see different tasks being routed to different agents. Probably go to 5.6 or Opus 4.8 or Fable when it's available for these types of tasks. But these other types of tasks use GLM 5.2 or DeepSeq, right? And another lower cost model that's still very intelligent, but maybe not quite at the frontier. And speaking of harnesses and other tools, AISDK 7 is now available. So this was on June 25th. So they introduced reasoning control, agent level tool approval, tool and runtime context, file and skill uploads, MCP apps, durable workflows, terminal UI, sandbox support, harness integrations, telemetry, lifecycle events, and more. Cool. I already had those, but all right. It's like they, you know, learned from some people. That's all we'll say. Speaking of like model routing, Cognition released this. It's a post on X that says, conventional model routing sucks. It passes benchmarks but fails to write code you'd actually merge. Introducing Devon Fusion, a new hybrid model harness for agentic coding. In testing, it reduces the cost of Fable level intelligence by 35% and still feels good to use. So this is what, our fifth one that's trying to do the same thing, right? Yeah, open router, factory. I mean, all these. Sikana. Yeah, it's a mixture of, not just a mixture of models, but it's like, I mean, it is a mixture of models, right? A mixture of experts in a different way. It's routing tasks to the right model based on intent, based on, you know, the types of tasks you're running. I think we're going to see a lot of people try to figure out how do we use all these different models for the tasks that they're best at and get cheaper costs while still maintaining the same level of intelligence. Do you think it's about token spend, which is why this whole thing is happening? Or do you think actually people want to pick the right thing for the right task, like as an engineering problem? Or is it more so, oh, like we need to start saving money? I don't know. I think saving money is a big part of it. I think costs are huge. What do you think? I think it's really because of the saving money problem. But also, I think before this whole saving money narrative, people were just opusing everything, 5.5ing everything, even for checking out Git, right? And because, you know, once you start spending tokens, would you, if you had to spend a dollar fifty to check out Git and pull from main, you probably wouldn't spend that money. But you would pay a cent or two to deep seek to do it. Now that cost conversation is here, this mixture of models, mixture of agents, mixture of harnesses now becomes like at the forefront of, we don't want to lose, you don't want to lose your power with the top models, but you know you don't need to use Opus to pull from GitHub. So why don't you use this instead? Like it makes a lot of sense to me. I'm not against it at all. So I was talking a lot with Tyler on our team about this and his biggest concern with all this is, let's say you, you have it do this some coding tasks and then you have it, okay, now I want to pull from main. And so then it uses a different model. But then you go back to a coding task. Well, now if you're switching models, you're breaking prompt caching the whole way. And so the question is, is that actually saving you money to move to a lower level of intelligence? Maybe, but maybe not. And so I think that is the question that is still to be determined on all these, is if you're assuming that nothing was prompt cached, it probably saves you a ton of money. If you can leverage prompt caching, which is usually, you know, sometimes only paying like 10 to 20% of the cost of the original, if it was uncached, well then the value you're saving is less or maybe even is slightly more. That's true. Yeah. So I think that's a good thing. I think that's a good thing. Like, share and subscribe. And follow us on X. And tell your friends. And their friends. I mean, we're not begging. Well, maybe a little bit. Subscribe to Agents Hour every Monday, noon Pacific. All right, let's talk about Agents. We talked about this last week. Claude Tag came out right before our show. We talked about it. But then it did kind of have its moment. It kind of went pretty viral. So Andre Karpathy says, this is a new paradigm for interacting with Claude that is significantly more in line with all the other human activity org-wide. Once you do all the under the hood engineering work to make this just work. And he's talking about just interacting with Claude in Slack. We've been talking about that a lot on the show. How we think that Slack and other tools like Slack are going to be the surface area you work with Agents pretty frequently. We supported Slack Agents and Monstra and channels for a long time because we believe this. And we've seen it internally. We have a bunch of Monstra agents running around in our Slack doing different things. So I think it's inevitable. I thought this was a funny tweet from Ashwin. Claude Tag is a Trojan horse, not because Anthropic is doing anything evil, because the incentives are obvious. Day one, this looks like a great feature. Tag Claude in Slack. Let it follow the thread. Remember context. Breakdown tasks. Chase work. And basically what this says is, as you use it, you basically become locked in. That Claude is like, all your knowledge, all your company knowledge is now going to Anthropic. So it's like this Trojan horse. Which is why I would argue maybe you should own that layer using tools like Monstra. Don't use it. Don't use it because the next day they're going to make a company or a new product that's your product. But I do think that agents in Slack are here to stay. And friends over at Copilot introduced OpenTag. So a better open source Claude Tag works with any model, any agent harness and fully custom agents. So it's really just like Monstra has channels. It's a little bit more of a way to do that. It's like a separate front end, I guess, for your different harnesses. So you can interact with other types of agents in Slack. And finally in this section, voice agents are now on Vercel. Real-time speech and transcription are now live on AI Gateway. Build with use real-time, generate speech and transcribe on AI SDK 7. So voice agents are, I think they've been like, we have a lot of folks with Monstra building voice agents. I think they're relatively popular. I think a lot of people still think they're not quite there, but they are getting, I feel like there's a moment coming. It's getting closer. When we started working on voice like two years ago now, almost, we worked on it too early, let's say. Yeah, we invested a lot like 18 months ago. But at least like now it's like paying off a little bit, even though it's not good enough. We have to make it better, but at least we had it. And now we have to keep investing in it. All right, let's talk about infra. My favorite topic. And cost savings. This is somewhat related to the past thing. There's a trend that just keeps popping up in all these different sections. We could have made this one big section just calling, you know, model routing, because this is a big part of it. But how to keep it. This is from Brian Armstrong of Coinbase. How to keep AI spend flat while token usage grows exponentially. Not with friction and spend alerts. With better defaults, routing and caching. In this post, Brian talks about how they really manage the cache. How they have a smart routing layer that once the cache times out, it might route you to a different model based on what the request is, what the intent of the user behavior is. And you can see that the tokens are still pretty high, but the costs are going down. So they basically cut their costs down while keeping the same amount of actual tokens. And this is Mark over at Coinbase. Says, So they route everything through their own gateway. So they're all controlling this through their gateway. Which is interesting. So it's almost like they built their own OpenRouter type product. I think the comments of these tweets were like, do you know OpenRouter exists? Yeah, but OpenRouter is not your own company and your own infrastructures. And a company like Coinbase can get crazy discounts. They could do it themselves and be okay. All right, Exa released Exa Connect, connecting agents to data beyond the public web. It's available today with ZoomInfo, Crunchbase, SimilarWeb, and many other leading data providers. I think all these types of integrations and connectors and things like that are important. If you're building agents, you have to be able to get data and then use that data. Another trend. Yeah, Exa is very much leaning into first search and now data and connections. There's a lot of companies doing this connection game and gateway game and many of these games. It is a Game of Thrones after all. A16Z said, I thought this was just pretty wild. We're thrilled to lead Merindale's $200 million seed round. Okay, that's not a seed round, but okay. And of course, the credentials are really good, so that's why you can raise this crazy amount. But they're building a system that can help anyone do AI work. They train frontier models that are expert at AI, R&D, and build the product around it. We'll see if anything comes of it, but that is a wild raise for an AI company. Dude, if you say you're training something, you get 100 mil. That's just the base. Yeah. All right, let's go through some quick hits. This one is a little bittersweet, but I think well-received from everyone on the team. This is from Akito Security. NPM now freezes high-impact accounts for 72 hours after sensitive actions like an email swap or 2FA recovery code use. It's a direct response to the Axios and Mastra attacks. Changing the account email is how attackers cut off the real owner's recovery path. Now NPM catches it. I'm proud of this one. We bitched on this show, we bitched online, we bitched on X. We bitched a bunch about how NPM sucks. And we also talked to our friends about how much they suck. And this was a good change. I'm curious, I would like to think it's because of us, it's probably not. But it probably is because why didn't they do it after Axios? Is it because of us that we were bitching so much? We were the straw that broke the back, so to speak. And I think Axios and Mastro were probably the two biggest examples, but there were many other, I think, smaller scale examples that were similar patterns, right? We talked, it seems like every other week, almost every week we're talking about some new supply chain attack. And I think this is one of the vectors that people would use. And I'm very happy, like that's good. I wish it would have been done after the Axios one. It would have saved our ass. But I am glad that they are paying attention. They are trying to make things better. I think with AI and with agents, things need to change on the security front. And the only way to do that is to keep improving and adding some of these different levels of security. So very happy to see this one. The whole team was excited. We celebrated in our Slack. Change can happen, guys. We can do it. If you complain enough, you'll get it. All right. So this is from Meta. This is from the AI team at Meta. It says, we're sharing the next major milestone in our non-invasive brain-to-text decoder research. Brain to QWERTY V2. Building on V1, brain to QWERTY V2 is the highest performing end-to-end pipeline capable of real-time sentence decoding basically from your brain. Plug me in, coach. Plug me in. You don't have to talk it. We talked about voice agents. If you want to talk to your coding agent, a lot of people are talking to it with Whisper, Super Whisper, whatever, all the different tools. Now you just think it. Dude, sign me up, man. Just plug me in like the Matrix. Hooked in, let's code. Hook me in. You don't want to get locked in. That's the new definition of locked in. I don't know how this works. I don't know if it's something that sits on your head. I have no idea how it works. I haven't looked into it enough, but this is wild. Maybe the neural link will actually hit production. Let's just assume it was just a hat. Just some kind of hat. The hat you're wearing right now, Avi, if you could put it on and talk to your coding agent, would you do it? You know I'd do it. If I don't die, I'm in. I thought this was a wild story, so we included it here. Do you remember back a few months ago when there was this Google Workspace CLI and it kind of came out of nowhere? It was announced by just one person. It was in the Google Workspace devs or something, GitHub, and everyone was so excited. It kind of went viral because of how good it was. And then this person got fired. This went viral on X too. Yeah, because apparently, and rightfully so, you could argue, he went against the policies. I don't think he deserved to be fired because of that. Maybe hand slapped or whatever, but I think there are policies in place of these big companies for a reason. You shouldn't be surprised when you just release something into the wild that's kind of under that brand. It looks like it's official, so I'm not actually surprised. I don't know if it was a calculated move on his part. He definitely got a ton of attention because it was really well done, but he kind of did it behind the backs of the team. Overall, I definitely applaud him for releasing it because I think it was way better than anything else that was out there. I think it's because he was a dev rel and some product manager, engineering manager got pissed. But the best thing about this CLI, if people don't remember, we were at the cusp of MCP versus CLI, tools versus CLI. All of this was in the zeitgeist at that time. He drops this. Steinberger also supports it. And then everyone started saying, why do you need MCP? You could just use CLI. This was a pivotal moment for our industry in this year. And then they fired him. He made a movement, dude. Like everyone started doing CLIs. Yeah, he did. He was definitely a big part of that. And now it sounds like he's working on his own thing. I'm sure he has a lot of attention and distribution from what happened. So maybe he can leverage that into his own thing that he's working on. But yeah, he did a good job with CLI. And on the one hand, you do want dev rels to be releasing things, right? They should release demos and guides and things. And its own CLI maybe was a step too far, but also people loved it. So you had one person go off and do something great and then the reward is they get let go. I can't imagine that sits well with others at Google. You have to follow the rules, which I get it on the one hand. You're a big company, but also it's not very empowering for folks that are on the team. This is why we work at startups, right? Hey, if anyone knows Justin, tell him if he wants to make tremendously less money and join Work With Us. Let him know. We're hiring him. Yeah, we are hiring. Come work with us, dude. We'll let you release all the stuff you want. Yeah, dude. All the CLIs you want, dude. And as the last thing, kind of as a capstone, we talked about GPT 5.6 at the beginning of the show around how the government is kind of holding back who gets access. This is a situation update. The Trump administration is close to allowing Anthropic to restore access to Fable 5, which has been offline for 15 days per Axios. Insiders expect the block to lift as soon as this week. All right. Thanks for tuning into the show. Follow us on X, Monstra. Go and follow us, Monstra-AI on YouTube. You can follow me on X, SMThomas3. You can follow at AbhiIyer on X as well. We appreciate you watching the news. Please go give us that five-star review. Only five stars. If you want to give us a four-star or less, find something else to do. That's okay. You don't have to. But if you're really feeling up for it, give us that five-star review. Go to GitHub. Give us that star. I think that's the show. That's the show. That's a wrap, dude. Peace. See ya. Peace. Peace. Thank you.