← Back to search

Every Frontier Model Safeguard Explained | Nous Research | 🟡🔴🔵 MTS Live

MTS · 2026-06-12 · 26 min
relevance 73 5340 words Episode page ↗ Audio ↗
Show full episode description
Jeffrey Quesnelle and Karan Malhotra join MTS to discuss the escalations between open source and closed source AI labs, the implications of hidden model manipulation at test time, and the concentration of power during the singularity. -- https://x.com/theemozilla https://x.com/karan4d -- MTS is a live news and interview show covering technology, business, politics, and culture as it happens. Follow MTS: X | ⁠ https://twitter.com/MTSlive⁠ YouTube | ⁠ https://www.youtube.com/@mtsituation⁠ Spotify | ⁠https://open.spotify.com/show/4HUNNmV1pjV7CW6Lphm1rM⁠ Apple | ⁠https://podcasts.apple.com/us/podcast/mts/id1891088763⁠ Substack | ⁠ https://mtslive.substack.com/⁠ Interested in sponsoring the show? ⁠⁠⁠⁠⁠⁠[email protected]⁠⁠ -- Brought to you by: ElevenLabs - AI that communicates at human level across every channel and modality. http://elevenlabs.io/mts Granola - The AI notepad. Try it once and never go back. https://granola.ai/mts Kong - The AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. https://konghq.com Blitzy - Autonomous software development for enterprise codebases. Ship 5x faster. https://blitzy.com Merge - OpenAI, Dropbox, and Ramp use Merge to get AI to production faster. https://merge.dev/mts AdQuick - Make your brand a billboard. Out-of-home advertising as easy to scale as digital. https://adquick.com Macroscope - AI codebase understanding for engineering leaders. https://macroscope.com/mts -- Timestamps: (00:00) Introduction (01:20) Initial Impressions On Model Nerfing (04:23) Hidden Versus Disclosed Model Safeguards (08:22) The Broader Open Source Response (14:37) Anticipated Corporate And Government Reactions (19:05) Irreversible Precedents For AI Labs (20:19) Decentralized Offense Versus Centralized Power -- Note : This podcast is not investment, legal, or tax advice, and is intended for informational and entertainment purposes only. Hosts and guests may hold positions in the companies and securities discussed; do your own research before acting on anything you hear.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How frontier labs like Anthropic silently nerf models on open-source AI research while hiding that classification is happening.
Benefits
  • Explains hidden test-time steering vectors and PEFT manipulation
  • Distinguishes honest bio/cyber refusals from covert research sabotage
  • Frames covert model steering as a red line for knowledge infrastructure
  • Predicts open-source community will detect classifier triggers
Use cases
  • Nous Research had a leaked old Gemini key run ~$50,000 a day of Gemini 3.1 distillation requests, logs showing Chinese-language distillation
  • Bio/cyber prompts visibly kick users to Opus, but frontier ML-research queries get a silently nerfed Fable with no disclosure
  • Crveu eval showed the classifier denying a benign 'how does mitochondria work' request, proving over-strong filters
KPIs / results
  • ~$50,000/day of distillation requests run through one leaked Gemini key
  • Gemini 3.1 distillation traffic identified in logs
Tools / build
  • Anthropic safeguards / steering vectors
  • PEFT test-time model nerfing
  • Fable frontier model
  • Opus refusal routing
0:00 / 0:00
There's always ways to get around those broad classifiers. That's not a real problem for people who are actively bad actors. That isn't what anthropics priority is. The priority is to hide the fact that the classification is happening at all. That is the real challenge. How are people going to know when it's being steered or not? These companies want to be critical infrastructure for knowledge work in the future. This sort of sabotaging, I think, needs to be a red line that says, we're not going to cross this. There's been some interesting developments in the open source ecosystem over the last day or so. Namely, for the first time ever, Frontier Models are nerfed when talking about open source research topics, Frontier LLM research. Anthropics Terms of Service has actually been in this for a while. You're not allowed to use their products to develop a competitor, but this is the first time they're really cracking the whip and actually nerfing the model. To talk about it, we have Jeffrey Cannell and Karin Malhotra, who are the co-founders of Nous Research in the Situation Room. Welcome to MTS. Welcome. Yo, yo, yo. Thanks for having us. I wish we were here under happier circumstances. So what's your read on the situation right now? Yeah, it's really just, it's quite unfortunate and it feels a little bit kind of like a contract break, you know, between two parties that the open source side and the closed source side have sort of had this, you know, symbiosis. You know, the original, so much of the space in AI was built in the open through open science. You know, in the beginning at NeurIPS, at ICLR, you know, these foundational papers that sort of built the scaffolding. And, you know, there's always been a place for the closed models to come, you know, to exist. And I don't think any of us in the open source space, like, begrudge their existence. But it sort of feels like a little bit of a betrayal now, like the true pulling up the ladder, you know, once you get high enough, you know, pulling up the ladder so that no one else can come after you. And what Anthropic has stated they've believed in before, but it's certainly an escalation as far as, you know, explicitly saying that we think that open source AI ought not to be built. And that we are like, you know, actively use our position within the market to, you know, to enforce that and to keep it from happening. Do you think? Yeah, that's right. Argo, go ahead. I was just going to say, you know, there's been this slippery slope of the alignment problem for many years. And it's been a nice, comfortable excuse for companies to be able to do stuff like this for a long time. You know, a lot of us formed over at News after GPT-2 to GPT-3 jump happened. And these open models became closed because of this danger. You know, as Jeff has said, there's been this beautiful symbiosis since, you know, we've begrudgingly kind of accepted it and it's become the standard now. But all the work we've seen come from Anthropic and other groups on representation engineering, steering vectors. You know, all this work we've seen around like for good fire coming out about, oh, look, control vectors can actually match in context learning. This kind of stuff has been in the name of doing alignment work and mechanistic interpretability and like kind of understanding models better. It isn't too big of a surprise that, of course, these companies are using the research that they've published to now say, we're going to use this to align the model away from you towards us. And in the name of safety, perform these much more subtle techniques than simply closed sourcing, like steering vectors, parameter efficient fine tuning and other things that happen live. These things they can do at test time to just quietly nerf the model from you, take your ability to like have a certain general intelligence away from you at will. Right. Like they've told us now the precedent is going to be that they're not going to tell the user about certain safeguards like the ones around ML training, self-improvement, etc. These safeguards, they're telling us now that they're not going to tell us about when this is going to trigger. There can be so many more things now and later that can be steered away from using these vectors, using PEP that we never know about. And eventually there will be models that you're using that just are completely different from the ones that they've trained before all these test time manipulations have occurred. Why do you think they decided to I guess you could say they're more honest about the safeguards on bio and cyber where it basically just kicks you to opus. But on frontier model research, it doesn't kick you to a different model. Instead, it uses a nerfed version of fable without disclosing that it is doing so to the user. Why is that? I would say I would say, you know, on the bio stuff, it's very much like it's it is a failure case, but it's something that people can sort of like at a level understand. Right. Like if you ask it to help me how to make, you know, a virus to thinning and it says I won't do that. Like, OK, I get there's someone who at least understand you're saying no to my face. Right. And I get that there that, you know, I could say there is an argument that that has a legitimate safety place. Now, the open source world probably can always get around that with open models. But like I think like just saying no to the user and being honest about it has a place, especially in these areas of like outright human safety. When it comes to the A.I. research, I mean, that that is escalating to the level of active sabotage. Right. Like this is like they're like actively sabotaging the model only in the name of protecting their competitive position. Right. Like and because you're doing something like if you really work, if this was the best way to do it, why would you not actively sabotage the bio part? Why would you actively sabotage the chemical, you know, biological stuff and try to like steer them away so that it like doesn't you know, it actually doesn't build the bomb. It actually doesn't. You know, but like on those ones, they just go and we're just going to say no. But over here where like the goal is to maintain the competitive advantage, then perhaps, you know, we're going to actively keep people. And now we know why a lot of this is happening. Right. Right. It's it's true that a lot of the open source model providers, especially, you know, especially out east and stuff do do distill from from Anthropic and open A.I. Right. Like the second a new model comes out, they go around and they find ways of, you know, I'm going to be honest with you. It literally happened to us here at News like a couple of days ago. So one of our API keys for Gemini got leaked. We had an old Gemini key, you know, and somehow it got leaked and got ran fifty thousand dollars a day worth of like Gemini three point one requests through it. We looked at the logs and it's all distillation requests, you know, in Chinese, like literally. So like this does happen. And so I'm sure in their world, they're like, we want to maintain our competitive advantage. But to me, the sabotage part is sort of the thing that I think has gotten under the skin of us in the open source space, because it's basically, you know, there is a good side of open source as well. The active democratization of this technology across all people is something that, like a lot of people do do do agree with. And I think has a place in this world. But Anthropic just fundamentally doesn't actually believe that is a good thing. I agree. Like it's certainly an almost Trojan horse kind of play to actively show you that bio and that biological and cybersecurity filters are blocked. You know, we have ways like everyone from the beginning of time, like of these models, like has ways to get around these filters because you've basically given people an active test of is the classifier triggering or not. The thing they're actually keeping the motor around, as Jeff has very saliently pointed out, is that model training capabilities. You'd never know right now whether or not you're hitting that classifier. That is the Trojan horse that kind of reveals that competitive advantage matters a lot more than the safety over here. Like, you know, actions speak much louder than words and hidden actions versus open actions. It's very clear what's happening there. I'd also just want to add like things like distillation are likely going to be also affected by this because there are there are things like the subliminal learning owl paper. Jeff, you know what I'm talking about that. I can teach a model about capabilities and preferences from a larger model by distilling it around a bunch of other topics. This will be a strong excuse for Antofag to say we're just going to nerf it whenever we think it's being distilled, which can just turn into we're just going to nerf it whenever we want. Yeah. Yeah. So I'm curious, what do you think will be like the broader open source response to this? Like, do you think that there's we're going to start seeing people like publish stuff about like, OK, I could see that the filter was like flagged here. Like, do you think that basically we will learn our way around it in the open source world? And how long do you think it'll take? I think that almost immediately you're going to see people there's two there's two factors here. Right. I think no matter what, you're going to see people do this work and do this thing where where they have so much opportunity in their hands to get from actually successfully doing ML training with this, where they have like such a big upside from doing it that they're going to try to attack it no matter what. Like, mech and turp people now care about the problem of can I tell when the model is being steered? People will now care about the problem of can I tell when peft is active? Right. They're going to want to try to do some kind of comparative work. Like this has happened every time a safeguard has shown up. Right. I have no doubt that this will happen immediately. The other access to consider is like the risk reward of doing this is changing massively. Like I used to red team for everybody. We used to do this all the time. We'd post, you know, here's a recipe on how to do this, do that. At a time where this was open science work today with more and more government entrenchment, people are scared to do some of this work because you maybe you can go to jail if you do this outside of an anthropic improved approved environment. Like we don't know where the regulation is going to fall around this. And it makes you feel almost like a lot of that development is going to happen outside of the states if that's the case. Well, how do you even get around the classifier when it's something like, you know, Kremu evaled it on asking it, like, how does mitochondria work? It's the powerhouse of the cell. Right. And it just it denies that request. Like the classifiers are deliberately like way too strong so that people can't get around them. They will be. There's always ways. I'm definitely not going to just spell out all the ways right now, but there's always ways to get around those broad classifiers. That's not a real problem for people who are actively bad actors or active red teamers or active Mech and Turk people to figure out the A, B of when the classifier is triggered or not or what format they can send a message in. That isn't what anthropics priority is. The priority is to hide the fact that the classification is happening at all. That is the real challenge. How are people going to know when it's being steered or not? And thankfully, there's a little bit of work around this that Anthropik themselves have done, but there's a lot more to be done. And I think it's going to kick up a lot now. And I think, you know, if you if you sort of stand back and say, you know, what would what would something like this this happening in a in a different field like look like? Right. You know, as these models, as companies like Anthropik and OpenA, I become more, you know, involved in everyone's daily activity and daily lives. You know, suppose for a moment we're running up into a general election. Right. Right. And, you know, the model is now silently can, you know, shown can silently change its behavior to like actively steer towards a directed goal and shape how information is shown to people. And like, you know, there was this contract that, you know, we just the AI companies weren't going to do this. Right. Like like we will make the model and maybe we'll, you know, steer or refuse. But like we're you know, we're going to train it. It'll have its behavior. But we're not going to like actively try to change what it says to steer you to steer you, the paying customer towards some particular outcome. And, you know, what would it take if, you know, you know, whoever your aunt, whoever your boogeyman is. Right. Like if you don't you know, you don't like Trump, then think about Trump. If you don't like the Democrats, imagine, you know, the next, you know, the next Democrat president, you know, has it sort of by government fiat. But they've declared that this sort of information needs to be silently, you know, put out or this this angle needs to be silently shown like they're already showing here. We can do it and we will do it for for reasons that we justify only to ourselves. Right. And so it kind of takes you there. I'm sure everyone here also is probably familiar with or heard of like, you know, the three body problem. This is literally like the exact story in the three body problem with the sofans where the aliens like silently manipulate scientific results to like steer society in in one way. And if these companies want to be like critical infrastructure for thought, for the knowledge work in the future, like this sort of, you know, sabotaging, I think, needs to be like a red line that says we're just we're not we're not going to cross this. This whole like it's only going to be triggered by point zero one percent of people and point zero zero three percent of cases at all. Like it'll barely ever happen. It's like how many people that are going to change the world are there? Like what percentage is it? Is it like one percent of the whole of everyone is a lot of a lot of money. It's a lot of people. There's a lot of people. Yeah. And it's like, you know, you're basically saying there's critical, massive outlier people that that move mountains. They're the only ones we're blocking. Right. They're the only ones who's results for fudging labs. Yeah. We'll continue monitoring right after this message from our sponsors. Eleven Labs. AI that communicates at human level across every channel and modality. Elevenlabs.io slash MTS. Shout out to our sponsors Granola, the AI notepad. Try it once and never go back. Granola.ai slash MTS. Special thanks to our sponsor Kong. Kong, the AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. Try it at konghq.com. Shout out Blitzy, autonomous software development for enterprise code bases. Ship 5x faster. Blitzy.com. Let me tell you about Merge. Build once, connect to every API tool and LLM. OpenAI, Dropbox, and Ramp use Merge to get AI to production faster. Merge.dev slash MTS. Quick word from AdQuick. Make your brand a billboard. Out of home advertising as easy to scale as digital. AdQuick.com. Big thanks to our sponsor Macroscope. AI code bases understanding for engineering leaders. That's macroscope.com slash MTS. Do you think there will be any response either by other labs coming out with basically unnerfed fable? Yeah. Or from the government saying this is anti-competitive behavior? Not from the government, I don't think. I mean, that'd be great. But I think like we would see if anything, like if the government had managed to nationalize any of this, like way more use of this kind of technique, this kind of technology. But on the open model side, like, yeah, like this is not some like unskippable chasm, right? Like the open source models will catch up. The thing to ask about is how hard are they're about to use their unnerfed fable to do recursive self-improvement super fast and super hard to make their closed models actually hit that chasm level, right? Like that's the real concern. The concern is not can we catch up to fable? It's can you catch up to fable before they make like Claude 12 internally? And now you're like much farther behind because you're just starting your serious. There could be that there probably is still some not another like transformer level innovation that's sitting out there, right? Some sort of attention level innovation that we just haven't discovered yet, you know, but like once it's discovered, we'll make the AI 10 times more powerful or 10 times cheaper or something like that. And the reason is everyone's racing, you know, to try to get to that. And if if someone like Anthropic were to find that first and then really be able to weaponize it and now we no longer have this culture of open science sharing, you know, that's sort of like the final capture, you know, that could happen. And in the open source space, then, you know, it's just incumbent on us for us to have credible alternatives, you know? Now, the question is, how does a credible alternative get created? I mean, the moat that that these companies have even more so is than just the usage is, is, you know, the access to resources. Like because this field is so scale dependent, right? And like it literally can get bigger, the more the bigger you make the model, it gets better. The more data you put in, it gets better. Or you can, you know, the cash going into the machine is like the thing you can just keep piling more cash in and it keeps getting better. This is actually often not true in a lot of places, in a lot of industries. Like you can't just keep pouring more money into it and it just linearly gets better, right? Whereas like with these models, it still seems to be somewhere in there. So the question is, how does an open source alternative, you know, rise up to this challenge? And, you know, there's a couple of ways it can happen. One, you could have a, you know, you could have a, you could have the government who has a monopoly on power somehow, you know, create it by, like I say, like by fiat. Unlikely, I mean, there are people in the government now who are strongly supportive of open source and could make the anti-competitive argument. The second way is you have a benevolent, you know, a benevolent entity, you know, your trillionaire status sort of people who decide they're going to spend their own money on it or companies like NVIDIA who have some sort of neutral alignment where it makes sense for them to, you know, foster an ecosystem and put a model out. We at Noose recently joined the NemoTron coalition at NVIDIA, which is literally Jensen's plan to try to like create an actual Western based open source model and fund it. And so, you know, we're doing our part in the NemoTron coalition. And, you know, it's nice there because Jensen has pledged 15 to $25 billion of NVIDIA's money towards NemoTron. And, you know, that's a significant amount of money. Now, it's still like maybe a fifth of what, you know, or a tenth of what Anthropic is raising. And it's just that those numbers are so insurmountable. So the real other way it could happen would be whether if we were to make some sort of fundamental breakthrough in the open source side as well, too, right? I mean, transformers were published openly and we found other methods that have been, you know, huge scaling things like mixtures of experts models, right? That was published in the open. The different attention mechanisms that are out there now. Yeah, go ahead. You can look at Jeff's own yarn paper that he had done a few years ago. You wouldn't really have agents or coding or reasoning without going from 16,000 context like to 128K. So we know these open source editors. If we were to find one of these, you know, then perhaps that would give us, you know, a leg up on fighting the incumbency back against the huge amount of money that's in the space now. But for that to happen, you know, the best tool we have would be something like, you know, some sort of open source or some sort of AI that could like assist the smartest minds that we have in the space today towards finding something like that. And what is the one thing they took? They took that. Do you think that now that like this sort of precedent has been established of like we will nerf the model in like certain aspects or whatever to maintain our competitive advantage? Like do you think open AI will follow up, have a model with the same thing? This is kind of like an irreversible practice going forward. Or maybe they will decide the opposite basically that they will get a competitive advantage by not nerfing them. It could go either way. Yeah. It could go either way. They had, you know, there was Olaf blocking from Anthropic being used and like Claude being used outside of Claude code with the Olaf and Codex pretty much like open AI committed to not ever doing that in response, which was very, very well received by people. So it really depends. The thing, right? The thing about all this is they could say anything to us and do it quietly. That's now an established capability for labs to do. So at the end of the day, like develop the ways of knowing that this is happening above above trusting whether or not someone's going to do it. Yeah. Nick from Cohere, who we were talking to earlier, made the argument that it is actually like safer to release these sort of models to the public because then like the market, you know, like adaptation mechanisms would just be very fast. Absolutely. So you, do you agree with this claim? I'll tell you my POV. I'm sure Jeff has some thoughts on it too, but my two cents are you're in the future where there's GPT-7, GPT-8 or Claude-10 and a bad actor that knows how to jailbreak models because as far as, as long as human beings are not infinitely stupider than these models, we will be able to jailbreak them because you have unlimited retries. If a bad actor goes and jailbreaks one of these models and it attacks something like a hospital or a sensitive center or energy center and the other person is also using Claude-12 or GPT-7 or whatever, and they get hit with a classifier that says, I cannot help you with this, it's a real problem. We've been pointing this out for years. We had this keynote about a year ago that illustrated this exact problem. And now with this like broad biology classifier thing happening, cybersecurity thing happening, it's like, hey, like that problem is now real in real life right now. If you put the model out, you get people who get to develop offense and defense at the same time. You have all actors working on it at the same time. That's how all technology has been so far. If you do this, you give bad actors a disgusting advantage over good actors. Presumably Anthropic would give unnerfed mythos to the hospital. Like I get the concentration of power concern there still. I'd say there's no hospital that has unnerfed mythos right now. You know what I mean? Yeah, but and what you're saying, Theo, is like this would be their response. Like, well, we'll find the right people to give it to who get it access to. And this sort of like just gives an insight into the worldview, right? Like this is a worldview in which I can control and see everything. Trust me. Like if the bad thing happens, it's okay because I know the right smart people. I will be able to discern and help you do it. But like typically this sort of, you know, concentration of all bring all power unto myself has not been the way that's been not been like humanity's best. Like humanity is not at its best when things like this happen, right? Like we can trust the central planners. They'll make it right for us. Even if something goes wrong, they'll go make it right. And certainly if you have the conceit that you can do that, that you can see all and will be able to make the right judgment case in all cases. Trust me. We'll only give it to the right people, which says underneath. I am the decider of who is right. You know, whereas like you have to have this like sort of humility about yourself. If you say, I won't be able to even know what the right, how to do this or be able to do it. So really the only way to make it safe is to just equalize everyone, put everyone on the same foot playing fields and let the dynamics of, like you said, the market, human, you know, human society, they sort of let, let the unseen things that prop up all of, you know, society find their equilibrium within this case. You know, you're not trying to hold the water in. One person's trying to hold the water in and hold it in their hand and, you know, you know, try to say, I can keep it all in here and it'll be okay because I can hold it. The other person says, I, it's going to slip through my fingers. There's no, I could ever hold this water. Let's let it fall and let it meet its equilibrium where it may. Look, I totally, if you don't, sorry, go ahead. I totally agree with you on power concentration concerns. Like I'm really pilled on this stuff, but like just like the strongest possible steel man of the opposite side is I think like make people making bioweapons in their basement. Like if it really becomes possible for people to make bioweapons in their basement, like how could you not like need to impose some kind of safeguard restrictions on that? It is easier right now. Right now. It is easier to make a bioweapon with mythos than it is to make a language model with mythos right now because you can tell when the bioweapon classifier is being triggered. You can send some hexadecimal or some other kind of format. You cannot tell when the language model snurfer is being triggered. It is today, Theo, the steel manned argument is unfortunately a terrible reality we're already in because they've chosen by not trusting humanity to align models away from good actors and away from humanity as a whole. Instead, people who take the time out to bypass these classifiers are the ones who will be able to and they didn't do this silent thing in that world we want to live in. They didn't do this proper safeguarding for biology stuff. They did it for their competitive nature. Let me give you a real world example that happened with us as well. So and this was not even on mythos. This was, you know, about a month ago. There was a supply chain attack on a couple of packages in NPM. And it was, you know, putting these credential keys dollars in. And we were trying to figure out whether Hermes Agent, like, had this as a dependency, right? Because people were saying, oh, it looks like you had this package or, like, we had figured it out. And I went down into, you know, we sat down with, I was with, like, Opus. I'm like, hey, look at the, you know, I copied it like the security report, the incident report from Swift on Security, I think, or I forget where it was. And, you know, here's all the effective packages. You know, here's Hermes Agent, check out locally on my computer. Can you look through and do, like, a dependency audit? You know, it's very tedious work. It's perfect work for an LLM, right? To, like, go do a bunch of tedious searching that, you know, like any human could do but doesn't really want to, you know, but doesn't require any creativity. And so go do that. And over and over, we got safety classifier blocked, trying to determine whether our package was infected by, you know, we were literally the white hat of white hat is all. Am I hurt? And there's a safety classifier, safety classifier, safety classifier. So the people who figured out how to get around it to write the malware, they got to use it. But, and by the time it happened to us, what are we supposed to do? Yeah, we actually did go to Anthropic. And guess what? They actually whitelisted our account and got us around the, you know, so we had, we're, like, in the security group that, like, is allowed to use the unblocked Opus only because we have access, only because we're prominent in the space, you know, only because I'm, you know, tech is technium and we're Hermes Agent. Like, there's a thousand other people who would be in this exact same situation and would have no recourse, you know, and it would never be able to, would be, never be able to do anything. And only the bad actors actually get what they want in this scenario. Right. Yeah. It is, it is concerning to be overly concentrated in power during the singularity. Thanks so much for coming on MTS to talk about it with us. Thank you guys. Appreciate it. MTS is an X native live streaming news and interview show covering technology, business, politics, and culture. As it happens, catch us every weekday live on X, YouTube, or wherever you get your podcasts. And, and I'll see you next time.