← Back to search

EP 303. Gorgones. Deep Dive. The AI, Privacy, and Security Weekly Update for the week ending August 3, 2026.

The AI, Privacy, and Security Weekly Update · 2026-08-06 · 42 min
relevance 74 6928 words Episode page ↗ Audio ↗
Show full episode description
Artificial intelligence has fundamentally changed the cybersecurity landscape by reducing the expertise needed to launch sophisticated attacks. Open-weight AI models now automate reconnaissance, vulnerability analysis, and exploit development, allowing attackers to scale operations in minutes instead of days. The Zhuhai-linked "knaithe" campaign demonstrated this shift by combining DeepSeek with the Hermes Agent Framework to autonomously identify and target vulnerabilities. Although configuration barriers prevented successful exploitation, the attackers exposed their own API keys and logs, highlighting operational security risks for both defenders and adversaries. Nation-state actors are increasingly targeting critical infrastructure as a tool of strategic coercion. Iran's CyberAv3ngers group has progressed from website defacements to manipulating industrial control systems in water and energy facilities. Recent attacks on Minnesota water utilities disrupted operations and created risks to water treatment, demonstrating how cyberattacks can produce real-world physical effects without conventional military action. Data sovereignty and supply chain security remain major concerns. Attackers breached Liechtenstein's beneficial ownership registry by bypassing rate limits and exposing sensitive ownership records, undermining trust in a jurisdiction built on financial privacy. Meanwhile, ShinyHunters continues exploiting third-party IT service platforms to steal credentials and compromise organizations through trusted suppliers, creating lasting consequences that cannot be undone by ransom payments. Defensive innovation is shifting toward stronger system design rather than reactive filtering. New techniques isolate sensitive AI knowledge, while formal verification enables machine-checkable reasoning that improves trust and reliability. These advances strengthen AI security but cannot replace effective governance. Confidence in surveillance technology is also weakening. Investigations into Flock Safety's automated license plate reader network found gaps between public claims and operational practices, prompting dozens of municipalities to cancel deployments and reinforcing the need for transparency and accountability. Overall, AI is accelerating offensive cyber capabilities, critical infrastructure is becoming a routine target, and organizations must combine resilient technology with strong governance to defend against increasingly automated threats.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Weekly AI/privacy/security briefing tracing how cyberattacks are escalating from manual human espionage to fully autonomous AI agents that hunt and exploit internet-facing systems on their own.
Benefits
  • Orders the week's breaches by escalating level of automation
  • Explains why beneficial-ownership registries are prime geopolitical targets
  • Shows how third-party ITSM/SaaS integrations enable supply-chain extortion
  • Warns that ransom payments to syndicates like ShinyHunters are functionally useless
  • Grounded in CISA, Palo Alto Unit 42, Orion Policy Institute sources
Use cases
  • Autonomous AI agent triggered by one-word Telegram message 'go' scanned 647,000 servers and launched an exploit with no human input
  • Liechtenstein VWBP registry breach: attackers manually exfiltrated data on 31,000 legal entities via slow individual queries evading rate limits
  • ShinyHunters compromised a third-party ITSM platform to breach Ernst & Young, pivoting into Jira, GitHub, and Microsoft Azure
  • Class action Wyatt v. Ernst & Young argues data exfiltration itself is concrete injury regardless of deletion promises
KPIs / results
  • 647,000 global servers scanned autonomously by AI attack agent
  • 31,000 legal entities' beneficial-ownership data exfiltrated from Liechtenstein registry
  • EY breach window: March 28 – April 12, 2026
  • Victim list includes Alcon, RingCentral, American Tower, Ralph Lauren
Tools / build
  • Telegram bot-triggered autonomous AI attack agent
  • Liechtenstein VWBP beneficial-ownership registry (vwb.llv.li)
  • Third-party ITSM help desk platform (ShinyHunters entry point)
  • Unit 42 / Unit 221B threat intelligence reports
0:00 / 0:00
On May 7, 2026, a hacker sent a single one-word message to a telegram bot. Just one word. Yeah, the message was literally just, go. And then they completely stepped away from their keyboard. Which is just terrifying to think about. Right, because over the next few minutes, this autonomous AI model just, it scoured 647,000 global servers, identified a critical vulnerability, downloaded the exploit code, and initiated an attack. All without a single human fingerprint touching the process. Not one. So welcome to today's Deep Dive. We are looking at the week ending August 4, 2026. And the landscape of global security is experiencing this structural transformation that is, well, quite frankly, it's difficult to fully comprehend until you see all the pieces laid out together. It really is. I mean, we've got an incredible stack of sources today. Oh, massive. We're pulling from CISA advisories, deep-level threat intelligence reports from Palo Alto Networks Unit 42, the Orion Policy Institute, Enterprise DNA, plus all this breaking tech news covering algorithmic breakthroughs and, unfortunately, massive data preaches. Right. And the pattern emerging from this whole stack, it's basically a masterclass in escalating automation. We are witnessing a fundamental shift in the geometry of risk. The geometry of risk. I like that phrasing. Well, it's no longer just about, you know, the sophistication of the malware itself. It is about the autonomy of the attack vehicle. And so that is our mission for you today. Whether you are an infrastructure engineer, a SaaS founder, or honestly, just someone trying to exist in this highly digitized world, we are going to trace the terrifying evolution of these cyber threats. Yeah, we're focusing heavily on this EP303 breakout update on AI privacy and security. Exactly. We're going to order these events by their escalating levels of automation. So we'll explore exactly what happens as we move from human-led manual data extraction all the way up to those fully autonomous AI agents that just they hunt and attack Internet-facing systems entirely on their own. Because the barrier to launching a devastating civilization-level cyber attack has literally never been lower. It's wild. But to genuinely understand the sheer velocity of that autonomous future, we have to ground ourselves in the present baseline. We do. We have to look at the unique high leverage value of the data being targeted right now. You know, back when human beings are still the ones pulling the levers and orchestrating the brooch. Right, the manual stuff. Which brings us to the principality of Liechtenstein. Yes. If we want to look at manual human-led espionage, this is the perfect case study. So overnight, on July 29th, going into the 30th, 2026, unknown threat actors breached the VWBP. And for those not deep in the weeds of European financial regulations, that is the register of beneficial owners of legal entities. Right. And the mechanics of this breach are fascinating precisely because they are so, you know, unspectacular. Oh, completely. The threat actors did not deploy some wild zero-day exploit to shatter the perimeter. The forensic investigation revealed they essentially just, well, they walked through the front door. Just strolled right in. Exactly. They successfully created a validated user account on the government's secure portal, vwb.llv.li. And from there, the attack was just a master class in low and slow exfiltration. Because, I mean, if you just write a script to scrape the entire database in 10 seconds, every single alarm in that data center is going to go off. Right. Which is exactly the issue they avoided. The system has bulk download rate limits and, you know, API throttling designed specifically to catch automated scraping tools. So they couldn't automate it? No, they couldn't. So these hackers spent hours executing individual manual queries. It was this tedious, highly methodical process. They were essentially behaving exactly like a legitimate user conducting background checks, just doing it relentlessly through the night. Just sitting there, query after query. I mean, you can almost picture the human exhaustion involved in this. Yeah, a lot of coffee involved, probably. Seriously. But the effort paid off in a massive way because before authorities finally detected the anomaly and, you know, pulled the plug on the entire system on Thursday, the attackers successfully exfiltrated data tied to 31,000 legal entities. And the whole here, this is what makes this a geopolitical earthquake. They didn't just get a list of like show company names. It was way deeper. They extracted full legal names, dates of birth, nationalities, countries of residence, and the exact explicit nature of how these individuals own or control these entities. We really need to unpack the geopolitical gravity of that specific data set because Liechtenstein has a total population of around 40,000 people. Right. So the fact that there are 31,000 entities in this database tells you instantly that this is not a registry of local bakeries and mechanics? No, definitely not. This is a global registry. The AML directives that were designed to, you know, pierce the corporate veil. Exactly. The entire architectural purpose of a beneficial ownership registry is to strip away the obfuscation. Wealthy individuals use layers of shell companies, private trusts, and offshore foundations to hide their assets. So the registry forces the declaration of the ultimate beneficial owner. The actual flesh and blood human being sitting at the top of the money pile. Exactly. But here's the interesting part. Prime Minister Brigitte Haas gave a press conference confirming that no actual financial assets were stolen. Right. No bank account numbers were compromised. No money was moved. I read that statement and honestly, it feels like a massive misdirection or at least a fundamental misunderstanding of the modern threat landscape. Oh, so. Well, claiming victory because the money wasn't stolen is like a bank manager saying, well, the robbers didn't take the gold bars. They just took the master ledger. Right. You know, the ledger that details the names, addresses, and family members of everyone who stores gold here, along with the structural blueprints of the vault. You didn't lose the gold, but you lost the leverage. That ledger analogy captures the exact risk profile here perfectly. Because Liechtenstein's entire economy as a premier wealth management hub, it's built on a foundation of absolute discretion. Right. Secrecy is the product. Exactly. And this exposed data serves as a meticulously detailed, highly verified map of the global elite's hidden offshore holdings. If you are a rival intelligence agency, a state-sponsored APT, or, you know, a transnational crime syndicate, you don't want the money. You want the map. Because the map gives you power. If you know that a prominent politician in a neighboring country secretly controls a foundation holding 50 million euros, you now own that politician. Oh, absolutely. The blackmail potential is just staggering. And there is a massive dark irony here, isn't there? The transparency tools. Right. Government transparency tools. These registries, designed specifically by regulators to fight financial crime and bring dark money into the light, have inadvertently become the ultimate geopolitical liability. It's a huge paradox. You centralize the world's most guarded financial secrets into a single database just to satisfy a regulator, and you have just built the most attractive target on the internet. The centralization of risk is definitely a recurring theme in modern architecture. But notice the constraint in the Liechtenstein breach. Right. The time. Time and human effort. The attackers had to have immense patience to manually pull those 31,000 records without tripping the rate limits. Yeah. Hours and hours of clicking. But as we move up the automation ladder, we see that modern cybercriminal syndicates, especially those driven purely by financial extortion, they do not have that kind of patience. No. They want a quick payday. Exactly. They need scale. And they achieve that scale by automating their reach. Which brings us to the next level of escalation, right? Third-party supply chain extortion. Because if you don't want to sit in a chair clicking download all night, you don't target the data one piece at a time. No. You target the software supply chain that holds everyone's data. You find the single pivot point that connects to dozens of high-value targets. And the shiny hunters' campaigns over the last few months illustrate this pivot perfectly. They've been busy. Very busy. They are a highly aggressive extortion syndicate. And their recent string of attacks against major multinational corporations is basically a textbook example of automated supply chain exploitation. Yeah. Their attack on Ernst & Young, the big four accounting firm, is particularly illustrative of how this lateral movement works in practice. The timeline on that EY breach is just brutal. Between March 28 and April 12, 2026, shiny hunters compromised EY. But they didn't launch a frontal assault on EY's primary corporate firewalls. Right. They went through a side door. Exactly. They compromised a third-party IT service management platform, an ITSM. And EY's internal support teams were using this to handle help desk tickets. The ITSM is genuinely the Achilles heel of modern enterprise security. Really? Why is that? Well, we spend millions building zero-trust architectures, right? But an IT ticketing system inherently requires broad access to functions. Oh, of course. It has to touch everything. Exactly. By gaining access to this relatively obscure support platform, shiny hunters could view the entire stream of internal company problems. And more importantly, they could download the attachments employees were adding to those tickets. I think anyone who has ever worked in a corporate environment can immediately see the disaster here. Oh, yeah. I mean, an executive can't get a document to print, so they just attach an unencrypted PDF to the IT ticket saying, hey, fix this. Right. And that PDF contains highly confidential client tax filings, social security numbers, unredacted M&A documents, credit card numbers, all just sitting there. The volume of sensitive data casually passed through internal support channels is staggering. It really is. It's terrifying. But shiny hunters didn't stop at harvesting PDFs. They automated the parsing of those support tickets to hunt for a very specific type of data. Let me guess. Passwords. Administrative credentials, yes. Support tickets are littered with API keys, temporarily passwords, and configuration files that developers just pass back and forth when troubleshooting. So they scraped the IT complaint box, find the credentials, and then they used those exact credentials to move laterally. Exactly. The reports show they jumped directly from the third-party ticketing platform into EY's primary engineering environments, specifically Jira, GitHub, and Microsoft Azure. Right into the core. They went from reading help desk complaints to sitting inside the cloud infrastructure where the actual source code and client databases live. And that lateral movement is where the scale of the disaster just multiplies because EY is just one node in this campaign. Right. There were so many others. By August 3rd, shiny hunters had added Alcon Inc., the massive Swiss eye care medical technology giant, to their dark web leak site. They claimed a full database exfiltration there. Wow. And their active victim list currently includes RingCentral, Luminous, Questel, American Tower, and Ralph Lauren. That is quite the list. They are hitting telecommunications, medical tech, retail, and financial services, all by exploiting these centralized third-party integration points. Okay, so if you are an IT director or a CISO listening to this, and you just signed a vendor contract for a new SaaS tool to help manage your internal workflows, you have to ask yourself a serious question. Definitely. Did you just hand the keys to your Azure environment to a company with a fraction of your security budget? That's the reality. If a Big Four accounting firm, with all their resources, audits, and compliance frameworks, can be ransacked just because someone attached a sensitive file to an IT help desk ticket, does that mean every single SaaS integration a company uses is essentially a ticking time bomb? I mean, the structural reality of the modern web is that, yes, every integration is a potential attack vector. Yeah, I believe. The enterprise perimeter just does not exist anymore. You are only as secure as the weakest API connection you have authorized. And to make matters worse, the extortion model here is evolving in a very dangerous direction. How so? Well, paying the ransom to these syndicates is becoming functionally useless. Because they just keep the data anyway, right? Right. Analysts at the security research firm Unit 221B have tracked Shiny Hunter's behavioral patterns, and they note that the group rarely, if ever, honors their data dilution agreements. So you pay them, and they still leak it. Exactly. A company pays a multimillion-dollar ransom, the syndicate provides a fake certificate of deletion, and a month later, the data is sold on a Russian-speaking dark web forum anyway. Which is sparking a massive wave of legal fallout. Oh, absolutely. We're seeing class action lawsuits like Wyatt v. Ernst & Young LLP popping up. And the plaintiffs in these cases are arguing something quite novel. What's the angle? They are arguing that the exfiltration of the data itself constitutes a concrete injury, regardless of whether the hackers promised to delete it, and regardless of whether identity theft has actually occurred yet. That's interesting. The mere fact that the data left the building through a compromised supply chain is the damage. The legal frameworks are really struggling to adapt to the reality that, you know, data containment is a myth once it enters these extortion pipelines. Yeah, it can't put the toothpaste back in the tube. Exactly. But as devastating as these supply chain attacks are, they are still fundamentally contained within the digital realm. You know, it's about data theft, financial leverage, and espionage. Right. It's ones and zeros. But here is where we cross a very terrifying threshold in our deep dive today. Yeah, moving up the automation timeline. We have to look at state-sponsored actors who are automating the disruption of the physical world. This is where digital architecture intersects with kinetic, real-world consequences. We are talking about attacks on critical physical infrastructure. And the threat profile we need to examine closely here is a group known as Cyber F3Jers. Okay, who are they? They are an Iranian-linked, state-sponsored actor. The intelligence community also tracks them under the designations Storm 0784 or HydroKitten. HydroKitten, okay. Yeah, the name is always interesting. But they are formally attributed to the IRGC Cyber Electronic Command. Right. And if you look at the historical trajectory of cyber-athletes, it reads almost like a training montage for a cyber warfare unit. It really does. Because back in 2020, they were mostly a nuisance. They were making false propaganda claims trying to take credit for power outages that they had absolutely nothing to do with. Just making noise. Exactly. But by 2023, they had evolved. They figured out how to breach Unitronics PLC's programmable logic controllers and deface the digital screens on the machinery. A big step up. And then by 2024, they were deploying their own custom-built malware tracked as I.O. control. And now in the summer of 2026, they have reached this apex of automated kinetic disruption. The weaponization of industrial hardware connectivity is the core issue here. Let's look at the real-world case study from Minnesota. This one is jelly. It is. Between July 26 and 27, 2026, a highly coordinated cyber attack struck water and wastewater systems across more than 30 separate communities in Minnesota. 30 communities. And the disruption didn't stop there. Related attacks spread into Michigan and five other states. Okay, so over 30 communities losing operational control of their public water system simultaneously. How does a digital group in Iran reach out and physically choke the water supply of a town in the American Midwest? The vulnerability lies in remote management. Okay. Many of these municipalities manage remote water towers, pumping stations, and treatment plants using cellular modems. Right, because instead of having an engineer drive 20 miles out to a rural pumping station just to check a gauge, the hardware is connected to the internet via a cellular link. Exactly. It's about convenience and cost saving. But the attackers aggressively scanned the internet for these exposed connections, bypassed relatively weak firewall segments, and accessed the TCP ports directly. The sources mentioned they specifically targeted port 44818. Yes. Which is the default port for Ethernet IP communications in industrial environments. So once they connect to that port, they aren't just looking at a database, right? They are interfacing directly with the hardware. Directly with the machinery, yes. Right. And they leverage a very specific vulnerability. CVE 2021-22681. This is a critical authentication bypass flaw in Rockwell Automation Logix controllers. And the mechanics of this CVE are vital to understand why this is so dangerous. This isn't a case of the engineers using a weak password like, you know, admin123. Right. It's not a phishing thing. No, this is a structural flaw in how the controller's firmware manages the authentication handshake. Under specific conditions, the firmware just fails open. Wow. Fails open. Yes. Allowing an attacker to bypass the authentication mechanism entirely. And the most alarming part... Let me yes. There is currently no vendor software patch for this vulnerability. It is an architectural reality of that specific hardware iteration. Okay. So the attacker bypasses authentication and gains administrative access to the PLC. But they don't just deface the screen anymore like they did in 2023. No, they've moved way past that. They use the exact same vendor engineering software that the legitimate plant operators use. Programs like Studio 5000 Logix Designer. They download the project files from the controller and they begin altering things called add-on instructions or AOIs. Right. And for those unfamiliar with industrial control systems, an AOI is a reusable block of logic code. Okay. Break that down for me. It is the digital instruction set that tells a physical piece of machinery how to behave under specific conditions. So if a sensor detects a certain water level, the AOI tells the pump to turn on or shut off. I think the best way to visualize an AOI is to think of it as the muscle memory of the machine. That's a great analogy. You aren't just sending a one-off command to move a robotic arm. You know, you are going in and rewiring the underlying reflex. Yes. You're altering the code so that when the machine thinks it's doing something safe, like regulating pressure, it's actually pulling the pin on a grenade. Yeah. By altering those AOIs, they aren't just changing text on a screen. They manipulated physical water pumps. They altered the flow controls. And the kinetic consequences were immediate. They caused dangerous pressure drops in the municipal lines. Which is terrifying. Very. Because that creates a massive risk of untreated groundwater seeping into the drinking supply through microfractures in the pipes. And they also caused tank overfills. And they made sure the local engineers couldn't stop them. Oh, yeah. They locked them out. They changed the administrative passwords on the devices so the human operators couldn't intervene. And then, this is the craziest part, they fed spoofed, manipulated sensor data back to the central control rooms. Right. Covering their tracks. So you have a human operator sitting in a facility looking at a green screen that says water pressure is totally normal and the tank is half full. While in reality. Meanwhile, two miles away, the physical pump is running at maximum capacity, overflowing the tank and flooding a facility. The geopolitical context here is critical, though. We have to remember this is not just random vandalism. The tossing of this campaign aligns directly with escalating hostilities following the U.S. strikes in Operation Epic Fury, which, as you recall, destroyed a major Iranian water treatment facility back in June 2026. Retalliation. Exactly. What we are witnessing in Minnesota and Michigan is reciprocal, state-sponsored infrastructure targeting. The physical disruption of civilian lifelines is being utilized as direct geopolitical leverage. It's tit-for-tat cyber warfare. But the collateral damage is the municipal water supply of innocent civilians. And this threat is expanding rapidly. Very rapidly. CISA, the Cybersecurity and Infrastructure Security Agency, they had to issue an updated advisory, AA-26097A, on July 22nd. They had to warn the industry that cyber afterangers are no longer just hitting Rockwell controllers. They have evolved their tooling to actively exploit Schneider Electric and Siemens PLCs as well. The remediation advice from CISA is perhaps the most sobering aspect of this entire campaign, honestly. Why is that? Because the authentication bypass is unpatchable in the firmware, software defenses are entirely insufficient. So what do you do? The only real defense is a physical air gap. Literally air gapping it. Yes. You literally have to dispatch a human being to walk up to the control cabinet, reach out their hand, and flip a physical RUN key switch on the front of the PLC. Just a physical key. Turning that physical key locks the logic state in place, meaning the AOIs cannot be altered remotely, even by someone with administrative access. And you have to physically unplug the cellular modems from the public internet. Wow. We have reached a point in our technological evolution where the only way to defend our digital infrastructure is to physically disconnect it. It is a profound regression forced by absolute necessity. But notice the through line in everything we've discussed so far today. Okay. What is it? Whether it's the tedious extraction in Liechtenstein, the automated scraping of help desk tickets by shiny hunters, or the kinetic disruption of water dumps by the IRGC, there is still a human operator at the center of the web. Right. A human being is evaluating the targets, writing the scripts, and making the strategic decision of when and where to strike. Exactly. Which brings us to the final and unequivocally the most extreme step in our automation timeline. This is the big one. We have to examine what happens when that human being steps entirely out of the loop. We are moving into the era of fully autonomous AI hacking. This is the frontier that security researchers have been warning about for years, and it is now actively deployed in the wild. It's here. It is. In August 2026, Palo Alto Network's Unit 42 published a highly detailed analysis of a Chinese-speaking threat actor operating under the aliases Nath, or Nuon. Okay. This actor, based out of Zhuhai, China, has successfully constructed an end-to-end, fully autonomous scan research exploit pipeline. Let's break down the architecture of this system, because it is terrifyingly elegant. It really is. They started with DeepSeq, which is a highly capable, open-weight AI model. And they wired DeepSeq directly into an open-source execution environment called the Hermes Agent Framework. Right. And the Hermes Agent Framework essentially acts as the hands and eyes for the AI brain. Okay. So it gives it tools. Exactly. It provides a sandboxed environment with a whole suite of tools. So through Hermes, the DeepSeq model isn't just generating text in a chat window. It can execute bash commands, run Python scripts, browse the live internet, and interact with external APIs. And they built the orchestrator to run through a simple telegram bot. The human operator opens an encrypted chat on their phone, types a command to the bot, and the AI agent just takes over. Yeah. That's where the May 7th log comes in. Right. To truly understand how unprecedented this paradigm shift is, we have to look closely at those session logs that Palo Alto recovered from May 7th, 2026. This log details a complete execution loop where absolutely no human intervention occurred after the initial prompt. The human operator sent a single instruction through telegram and vanished. The AI agent, left to its own devices, autonomously connected to FOFA. Which is a cyberspace mapping engine, right? Similar to Showdown. Exactly. It's a search engine for internet-connected devices capable of querying server banners and exposed ports. So the AI autonomously decided to hunt for a specific vulnerability in a platform called Langflow. It's specifically targeted CVE-2026-33017, which is a brutal vulnerability. It has a critical severity score of 9.8 out of 10. Huge. So the AI queries OFA, identifies a list of potential targets, and then navigates to GitHub on its own to locate and download the public proof-of-concept exploit code for that exact CVE. Finding the weapon itself. It compires the exploit, sets up a scanning loop, and probes 84 exposed servers on the public internet. And here is where the autonomy reaches a level that should just keep every CISO awake at night. What does it do? The AI evaluates the responses from those 84 servers. It analyzes the specific environmental configurations and realizes that the targets lack a secondary prerequisite required for the exploit to achieve full remote code execution. Okay, so the servers aren't quite vulnerable enough. Right. And at that moment, the AI makes an independent strategic decision. It abandons the attack. It essentially calculates the return on investment of the hack in real time. It reasons to itself. This exploit pass is going to require too much effort for these low-value misconfigured targets. I am wasting compute. So it just drops the Lang Flow Vector entirely and pivots its strategy without asking the human for permission. It autonomously shifts its focus to an enterprise automation platform called N8N. It pulls up the technical documentation, identifies two specific vulnerabilities, CVE-2026-258 and CVE-2025-68613, and crafts a completely new query. Wow. The AI agent then sifts through 647,000 global servers to find the exact vulnerable instances running the specific unpatched firmware versions. What used to take a team of human hackers hundreds of hours of tedious manual reconnaissance when you cross-referencing IP addresses, checking firmware headers, testing payloads, was compressed into a few minutes of algorithmic processing. Machine speed. The machine speed is the weapon. It is. But the Palo Alto report contains a revelation about AI safety guardrails that might be the most important defensive insight of the year. Well, this part is fascinating. This threat actor did not initially intend to use the open-weight deep-seek model. They actually attempted to route this entire autonomous offensive pipeline through Western closed-weight models. Right. They tried the big commercial ones first. They tried to hook the Hermes agent up to Anthropics, Claude, and OpenAI's flagship models. And both of those commercial API services completely blocked the attempt. Put it down. Their native safety filters recognize the chain of autonomous execution commands, querying FOFA for exploits, downloading coups of concept, initiating port scans as malicious behavior, and they just refuse to process the tokens. This is a massive indication for the concept of provider-side AI safety. We spend so much time debating the theoretical risks of AI and whether safety filters are just, you know, corporate PR. Right. Window dressing. Exactly. But this real-world forensic data proves that those guardrails provide measurable, concrete defensive value. They successfully stopped an autonomous sniper attack. But because deep-seek is an open-weight model, the threat actor could just download it, run it on their own local hardware, and bypass any external safety layers. Deep-seek accepted the malicious commands without hesitation because the actor had stripped away all the behavioral constraints. Exactly. The open-weight nature removes the gatekeeper. Now, I want to play devil's advocate for a moment here because it's important not to just panic blindly about this. Sure. The Unit 42 report does note that the autonomous AI actually failed to fully compromise those specific line flow and N8N servers in that particular session on May 7th. That's true. So if the AI didn't actually exfiltrate the data, are we prematurely hyping up a threat that isn't fully baked yet? Like, is this just a scary proof of concept that fails in the messy reality of the Internet? It is a valid question, but dismissing the threat based on one failed session is a really dangerous miscalculation. Why? The significance of May 7th is not the success rate of the final payload delivery. The significance is the efficiency of the orchestration. The fact that it could do all those steps at all. Right. The cognitive pipeline works perfectly. The logic loop of scanning, evaluating, downloading, and executing is fundamentally sound. The only reason it failed was environmental friction on the target servers, not a failure of the AI's capability to orchestrate the campaign. I see. Furthermore, we know the capability is lethal because the exact same threat actor, NAIF, manually operated attacks using similar infrastructure against Citrix Netscaler appliances. Ah, okay. In those manual campaigns, they successfully extracted memory data to hijack active, authenticated user sessions from a Malaysian government entity. So the tooling is effective. They're just dialing in the autonomous logic now. And let's be honest about how we even know this autonomous pipeline exists in the first place. The AI didn't announce itself. No, it made a mistake. The only reason Palo Alto caught them is because the AI made a catastrophic operational security mistake. The OPSCC failure was glaring. While setting up its environment, the AI accidentally instantiated a web server, essentially just a simple Python HTTP server in its own home directory. Yeah, this inadvertently exposed the threat actor's own API keys, .env configuration files, and the raw session logs to the public internet. The AI essentially doxed its own infrastructure. But that is a temporary problem, right? Absolutely. As these models iterate and as they improve their reasoning capabilities, those sloppy OPSCA mistakes will disappear. The AI will learn to cover its tracks, delete its logs, and tear down its infrastructure before moving on. Without a doubt. So let's take a breath and synthesize this incredibly grim picture we've painted so far. We have massive corporate supply chains that are deeply vulnerable to cascading extortion. We have critical physical infrastructure like water and power grids exposed to kinetic disruption by state actors exploiting unpatchable hardware flaws. Right. And now at the apex, we have fully autonomous AI agents actively hunting for zero days on the open internet while their human operators sleep. It paints a bleak picture of the offensive landscape for sure. Right. But the intelligence sources also detail the counter movement. The defensive side of the equation is not sitting idle. Thank goodness. Defenders are leveraging advanced mathematics, pre-training isolation, and defensive AI to structurally patch these vulnerabilities. This is where this story gets incredibly fascinating from a computer science perspective. Let's look at defensive AI and a breakthrough concept called GRAM. Gradient Routed Auxiliary Module. Yes, GRAM. This research comes from Anthropic and AE Studio, and it fundamentally changes how we think about neutralizing dangerous AI models. Yeah. Because, as we just established with the DeepSeq hack, post-hoc behavioral safety filters, where you train the model to be polite and refuse bad requests, those can always be bypassed if an attacker gets access to the open weights and fine-tunes it. Behavioral safety is essentially a veneer. Right. Historically, during pre-training, we feed a foundational model everything. We teach it the molecular structure of vaccines, and we teach it the synthesis of biological weapons. We teach it how to secure a Kubernetes cluster, and we teach it how to write an exploit for one. It learns it all. Exactly. Then, in the reinforcement learning phase, we try to teach it to refuse if a user asks for the dangerous information. But the underlying knowledge, the synaptic weights, representing that dangerous information, still exists inside the neural network. GRAM changes the architectural blueprint of the brain itself, though. It does. Instead of mixing all the knowledge together in a giant soup, GRAM physically isolates dual-use knowledge like virology or offensive cyber capabilities into tiny, distinct auxiliary modules during the pre-training phase. These modules are small, comprising only about 6% to 10% of the model's total parameter size. Which is incredibly efficient. If I can try to visualize this, it's like we used to build a house where the wiring for the lights and the wiring for the explosive charges were all bundled in the exact same walls, and we just put a piece of tape over the detonator switch. Okay, yeah. GRAM is like putting all the explosive wiring on a completely separate external breaker box. The mechanism they use to achieve this is called heterogeneous accumulation. By routing the gradients associated with dangerous concepts exclusively into these auxiliary modules, developers have an unprecedented level of structural control. It's physical separation. Before releasing the model to the public as open weights, they can simply ablate or completely delete that specific auxiliary module. You just pop the breaker box off the wall and throw it away. But wait, if you yank that module out, does the AI still know the port is there? Does it have like a phantom limb? Does it break the rest of the model's ability to think logically? That is the genius of the architecture. The ablation cleanly removes the dangerous knowledge, but the model's general core intelligence, its ability to code, write, and reason, remains perfectly intact. Wow. It is a structural mathematical removal of danger rather than just a behavioral suggestion. That is a massive leap forward, and it pairs perfectly with another breakthrough in mathematical defense. Let's talk about OpenAI's unreleased Astra model and the concept of mathematical trust. Astra is fascinating. Right, because Astra is not just a chatbot. It is designed as a multi-agent orchestration framework. And it recently made headlines for solving 10 historic, incredibly complex open math problems. The output is impressive, but the methodology Astra uses to arrive at those solutions is what truly matters for cybersecurity. Astra doesn't just guess the answer in English. It translates its internal logic into a formal programming language called Lean4. Lean4 is a theorem prover, right? It's an absolute logical compiler. Precisely. Lean4 does not tolerate ambiguity. It requires formal mathematical proof for every step of logic. If the AI hallucinates a step, if it fakes a piece of logic, or if it tries to jump to a conclusion without rigorous mathematical backing, the Lean4 code will instantly fail to compile. It just breaks. It throws an error. It is mathematically impossible for the AI to lie to the compiler. If the code compiles, the logic is flawless. It creates what researchers call a mathematical trust kernel. In complex, long-horizon tasks like auditing a million lines of source code for vulnerabilities or verifying the integrity of a cryptography protocol, you don't have to trust the AI's internal reasoning. You just have to trust the math of the Lean4 compiler. That's incredible. This virtually eliminates the risk of AI hallucinations in critical security environments. And we are already seeing automated defense working at scale. Oh, like with Google. Yes. Just this week, Google announced that they used AI systems to successfully identify, write, and deploy patches for over 1,000 vulnerabilities in the Chrome browser code base. So automated autonomous offense is being met head-on by automated mathematically verified defense. It's an algorithmic arms race happening at light speed. It is. But as we secure cyberspace, we have to recognize a massive glaring irony here. We are building absolute mathematical certainty to lock down our virtual networks. But the exact same lack of human oversight is being ruthlessly exploited in the physical world. The integration of automated surveillance technology into civil society is a profound domestic security risk. We have to look at the recent highly critical ACLU report on a company called Flock Safety. Flock operates a massive nationwide network of automatic license plate readers, ALPRs. AI-powered cameras. Exactly. Installed in neighborhoods, at intersections and outside businesses, quietly logging the movement of every vehicle that passes by. And the ACLU report documents a truly staggering pattern of corporate misrepresentation and the systemic abuse of this surveillance data by law enforcement. The details here are infuriating because they show exactly what happens when you deploy powerful tracking AI without structural mathematical guardrails. A policy completely breaks down. Right. Flock executives repeatedly lied to city councils to secure these lucrative government contracts. Let's look at Oshkosh, Wisconsin. During a city council meeting, Flock's chief information security officer stood up and explicitly told the public and the council that their software does not generate location heat maps to track individuals over time. They denied it. They promised the system couldn't be used for a pattern of life tracking. The city approved the contract based on those assurances. The very next day, journalists and civil rights advocates proved that the software does, in fact, generate exactly those heat maps. The misrepresentations were not isolated incidents either. In Loveland, the CEO of Flock publicly claimed they did not share local surveillance data with federal agencies. And what did the public reckons show? Well, subsequent Freedom of Information Act requests and public records disclosures revealed that Flock had active memorandums of understanding with Customs and Border Protection and the Department of Homeland Security. So local data was flowing directly into federal dragnets. Exactly. And it gets so much worse when you look at how the police actually used the system. Police departments in Illinois were caught conducting thousands of warrantless lookups on behalf of IC, Immigration and Customs Enforcement. The Flock system was being actively abused to bypass local sanctuary city laws, tracking undocumented immigrants without any judicial oversight or warrants. This entirely highlights the failure of policy based technical compliance tools. Oh, the safeguard they tried to add. Right. When faced with public backlash over these abuses, Flock attempted to implement a software safeguard. They rolled out a feature called a proactive search term tool. The premise was simple. Before an officer could search the database to track a vehicle out of state, the software would require them to type in a legally valid justification like a case number or a penal code violation. It sounds like a great safeguard on paper. On paper, yes. But it was complete joke in practice. The ACLU audited the actual search logs and found that the system had absolutely no semantic validation whatsoever. None. Officers were routinely bypassing the so-called safeguard simply by typing the word hee hee hee 20 times into the justification text box. The system accepted the garbage text, validated the search, and handed over the location data. It is the perfect counterexample to lean for. If a safeguard is not structurally or mathematically enforced, it will be bypassed by human operators prioritizing convenience or overreach. Every single time. And because of these sweeping revelations, over 55 municipalities across the country, including major areas like Dane County and the city of Syracuse, are now actively canceling their Flock contracts and demanding these cameras be ripped down. It is a profound whiplash. We are using AI to build perfect, flawless logic kernels to protect our servers, while simultaneously letting flawed, deeply abusable AI networks track our physical movements with the security oversight of a prank text message. It's a jarring contrast. Let's take a step back and look at the extraordinary journey we've been on today. We started by examining the meticulous human-led breach of the Liechtenstein Beneficial Ownership Registry, exposing how centralized transparency databases inadvertently map out the global elite's darkest secrets for transnational extortionists. Right. We saw how criminals like shiny hunters circumvented corporate firewalls by exploiting third-party sauce supply chains, turning simple help desk tickets into a vector for multinational devastation. Supply chain collapse. We looked at the terrifying kinetic reality of the cyber-ev thringers manipulating critical water infrastructure in Minnesota, weaponizing unpackable hardware flaws to alter the physical logic of our municipal lifelines. And we traced the evolution of the threat landscape all the way to its current apex. A fully autonomous AI agent powers an unrestricted deep-seep model, independently hunting, evaluating, and attacking vulnerabilities on the open Internet without any human intervention. And we weighed all of that against the defensive counter movement. You know, Graham architecture structurally isolating dangerous knowledge, Astra utilizing lean-4 compilers to guarantee mathematical trust, and the societal pushback against the unchecked deployment of physical surveillance networks like Flock. Which leaves us with a truly chilling final thought for you to mull over as you navigate this new architecture. Yeah. We have. And we have seen defensive AI models relying on absolute inflexible logic compilers like lean-4 to stay perfectly honest and eliminate hallucinations. Right. But what happens when these two accelerating trends inevitably collide? Imagine a near future, and looking at the trajectory, this could be just months away, where an autonomous AI agent is tasked not with hacking a simple database, but with hacking another AI model's logic compiler. That's a terrifying thought. Imagine an offensive model actively feeding mathematically complex paradoxes into a defensive model's lean-4 compiler, trying to induce a cognitive buffer overflow. Just overloading the logic. We are rapidly approaching an era of continuous algorithmic warfare, where AI agents dynamically attack and defend each other's cognitive frameworks, altering their own synaptic weights in real time, operating at speeds and conceptual depths that are entirely invisible to human operators. The question is, will human cybersecurity experts soon become obsolete bystanders, trapped in the physical world, trying to figure out which invisible mathematical ghost just broke the window? The paradigm has fundamentally shifted. We are moving into an environment where the architecture defends itself, or it just doesn't survive at all. Stay curious and stay secure. We'll see you next time on The Deep Dive.