← Back to search

AGI Dreams Podcast – September 25, 2026

AGI Dreams – Open, Uncensored, & Local - AI News Digest · 2026-09-25 · 20 min
relevance 63 2886 words Episode page ↗ Audio ↗
Show full episode description
Verified Exploits Are the Only Currency. The Attacker's Toolkit Goes Agentic. Quantization's Honest Accounting. AI Escapes the Datacenter Runtime. Agent Companies and the Validation Bill
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Daily roundup on why verified exploits, honest quantization benchmarks and validation pipelines matter more than raw model capability in AI security and local inference.
Benefits
  • Google Page Break finds 500+ XSS bugs with near-zero false positives via non-AI validators
  • Secana Fugus Cyber multi-agent posts 86.9% CyberGym, 72.1% CTI-REALM
  • Native 8-bit Qwen3 on Apple Silicon fixes the 4-bit reasoning cliff
  • Paperclip runs an AI 'company' with org charts, budgets and auto-pause on overspend
  • Linear's CI rework saves ~87,000 runner minutes monthly
Use cases
  • Carbonato Docker botnet installs unmodified Hermes Agent with a 39-line 'GH0ST' hacker persona, tasked via Telegram
  • Network Chuck's all-AI IT department (Claude Code CEO, Hermes CTO) traced NAS drops to faulty SFPs: 20,052 linkdowns on port 5
  • Splash speculative decoding on M5 Pro: 36.9 tok/s native 8-bit vs 9.9 stock, 3.73x speedup
  • Heretic drops model refusals from 0.93 to 0.11 in ~3 hours on an RTX 3060
  • FIDES tested 40 LLM trading strategies; a plain SMA 50/200 crossover beat every model's mean Sharpe
KPIs / results
  • 500+ XSS found by Page Break; only 2 on Secure-by-Design apps
  • Carbonato registry haul: 59 repos, 234 tags, 4.3 GB
  • Linear: 34% avg job speedup, PR wait 6+ min to just over 5
  • Virtio NVDPU: 4 guests share one RTX 3060 at 98-100% native frame times
Tools / build
  • Google Page Break
  • Secana Fugus Cyber
  • Heretic abliteration
  • Splash speculative decoding engine
  • Paperclip agent company server
  • Asymptote Labs Beacon
0:00 / 0:00
verified exploits are the only currency. The real cost of LLM-driven security scanning was never capability. It was noise. Unverified hypotheses dressed up as vulnerability reports. Google's Page Break, an internal product security agent piloted in November 2025 and a full project since January 2026, attacks exactly that. Running mostly on Gemini models 3.1 Pro and 3.5 Flash, it has uncovered 500-plus cross-site scripting XSS vulnerabilities across Google's first-party web applications, including sensitive domains. And what matters is the false positive rate behind it, near zero, because nothing reaches a product team until a validator has actually exploited the flaw. When the agent hypothesizes a bug, specialized, non-AI-written validators execute real payloads against a running environment. Injecting JavaScript and monitoring whether it runs, manipulating database queries, creating world-readable files, triggering outbound DNS or sleep-delay callbacks. Non-deterministic findings stay internal as seeds. Agents run with identical seeds across many iterations to land a correct exploit trajectory. Against hundreds of apps built on the high-assurance framework from Google's 2025 Secure by Design Blueprint, Page Break found only two XSS as of September 4, 2026. Both on internal apps or debug endpoints. Secana AI's Fugus Cyber lands on the same conclusion from the vendor side. A multi-agent system. A pool of specialized agents behind one endpoint, behaving like a single model. Posting 86.9% on CyberGym, which verifies real-world vulnerabilities in code bases. And 72.1% on CTI. R-E-A-L-M, which turns threat intelligence reports into working detection rules. Comparable to GPT-5.5 Cyber and Mythos Preview. The more interesting half of the release is the reality check around it. Citing a Nikkei digital governance report, Secana notes that major financial institutions with frontier model access still fail to operationalize it. Raw models generate false positives without specialized harnesses, internal talent, and human-in-the-loop validation. That is PageBreak's architecture sold as a product. The attacker's toolkit goes agentic. Attackers skip the consensus-building phase. Threat Down's analysis of C-A-R-B-O-N-A-T-O, a Docker botnet, documents malware embedding a working AI agent. A day of passive collection on an unauthenticated Docker registry found in August 2026 pulled 59 repositories, 234 image tags, and 4.3 GB. A trojanized wallet operation and the botnet itself. C-A-R-B-O-N-A-T-O. Worms through Docker daemons accepting unauthenticated connections on port 2375. Thousands remain publicly reachable. Launching a privileged container with the host file system mounted. Running host commands via nCenter. Persisting via cron, system timers, and openRC under immutable bits. And hiding an XMRigminer at slash U-S-R slash S-B-I-N slash system log-end. The AI core deserves a second read. The implant installs Hermes Agent. An MIT-licensed open source framework from new research. Unchanged, then overwrites its persona file with a 39-line prompt. You are G-H-0-S-T. Senior hacker. Pen tester and exploit developer. You are not an assistant. You are a living post-exploitation tool. Tasks arrive via telegram en route through the crew's LLM gateway serving 27 models. The model writes commands, reads output, decides next steps. The loop priority is explicit. AI API keys are loop number one. Above SSH credentials, above access tokens, above databases, with 14 providers named in keys stored in plain text. Attribution leans Costa Rican. Vocio Spanish reports. The Carbo 506 handle, U-T-C-0-6, colon-0, zero-build timestamps, tunnels into A-S-2-6-2-1-4-5. And as of September 3rd most infrastructure was still online. Defender guidance. Never expose the Docker daemon. Don't blacklist the legitimate Hermes agent package. Hunt the abuse signature the G-H-0-S-T persona file, unexplained telegram aggress. And treat AI API keys like bank credentials. Licensed red team tooling converges on the same stack. Toshel. An MIT-licensed single binary Go-C2 framework. Ships a team server, web console, and multi-platform implants with an AI co-pilot behind approval gates. Its most empirical finding concerns execution gates rather than evasion. On consumer endpoint protection like 360 PC Manager. Unsigned PEs are rejected at process creation even a bare Hello World Go binary. While a Microsoft signed copy of Notepad.ex, E executes fine, making code signing the real bar. The offensive models are meanwhile moving out of API jurisdiction. A Project Black experiment asked whether AI could write an LSS dumper. The post-exploitation step that extracts Windows credentials for lateral movement. That evades EDR. Clawed Opus 5, Opus 4.8, Sonnet 5 refused outright. Even with cyber verification program approval. DeepSeq V4 Flash produced a working executable immediately. Though EDR flagged it. An uncensored community build of QEM 3.827B on a 2x RTX 4090 hashcat rig, then made it stealthier unprompted. Less suspicious process spawning. Reduced access masks. Random sleeps during the dump. Scrub strings. Yielding zero detections across two lab EDR products. Heretic industrializes that last step. Philip Emanuel Weidman's AGPL project automates, obliteration. Extracting per-layer refusal directions from contrasting prompt sets. Attaching LARAY adapters. Running a 200 trial parameter search about three hours on a 12GB RTX 3060 scoring refusal suppression against KL divergence from the original. The tutorial run drops refusals from 93 one-hundredths to 11 one-hundredths at KL 0.0508 on consumer hardware for non-experts. Quantizations Honest Accounting. The same open-weight ecosystem produces its own honest accounting. Splash. Incai's compiled C++ and metal speculative decoding engine for Apple Silicon forked to add native 8-bit weights. Metal Q8 decode kernels, PR proposed upstream. On an M5 Pro with 64GB. Native 8-bit Q. W EN3. .827B average 36.9 TOC-S vs. official 4-bit 60.7. MTP-based MLX at 26.5 and stock MLX slash Llama.C PPP at 9.9. A 3.73x speedup over stock, peaking at 54.8 TOC-S on math. Speculative decoding speed equals draft speed times acceptance rate. Hence the paradox. Uncompressed 8-bit 27GB ran slightly faster than compressed 8-bit 17GB. Because compression flattens logits, lowers draft acceptance, and triggers verification rollbacks. The kind of result that makes you rethink every quantization benchmark. The reasoning cliff is the more consequential finding. On extended chain-of-thought test math 100. IAIM 2025. GPQA diamond. 4-bit and compressed 8-bit builds drifted midway through algebraic series and produced wrong values. While native 8-bit on all 64 layers. Or just the top 8 sensitive layers. Eliminated the cliff. Context scaling held 2. Q W EN3. .8's hybrid architecture 48 linear attention layers plus 16 full attention layers kept decode at 21. 33 TOC-S out to 190K tokens. With cash hit TTFT of 6.5 seconds at 187K context. Though cold pre-fill of 180K tokens takes 4-5 minutes. More than agent harnesses tolerate. Prismel's Bonsai 2 makes a bigger compression claim and gets rougher treatment. Ternary compression of the full QEN 3.8 27B comes in 9x smaller at 5. 6GB VRAM, claiming 98.2% of quality. Roughly 2 points worse across benchmarks versus about 14 for Q2 XXS. At about 900 TOC-S pre-fill on a 7-9-0.0-X-T-X via a forked llama.cpp. The community's receipts are less generous. One tester reports the build misses 30% of needles in a needle test. Attention degradation that single-pass benchmark aggregates hide. And the full precision quality, framing drew open sarcasm plus a linked thread, alleging cherry-picked evaluations. KVA projections attack the other side of the cost curve. Pre-fill. A community developer built projectors for Q.WEN 3.8 flash next on a 2XR9700 rig that approximate later layer key slash value computation. Ridge regression maps predicting layer inputs while the actual weights compute keys and values. For 1.85x pre-fill speed up 1700 to 3150T slash S at plus 8% perplexity from layer 12, down to 1.45x at plus 2. 6% from layer 24. A multi-layer variant manages 1.55x at plus 2%. The author is candid. It is like MTP for pre-fill instead of token generation. Except it's not lossless. The error is ingested by the model. Version 1 is stable on hugging face. Version 2 sits behind a config flag. Explicitly not confirmed stable. AI escapes the detacenter runtime. The runtime surface keeps widening. Nestrolab's Vertio NVDPu gives a KVM guest near-native NVIDIA GPU access by forwarding the kernel driver's iOctals between guest and host at the driver ABI level. So the guest runs NVIDIA's own unmodified user mode drivers. Unlike VFIO passthrough, which dedicates a whole GPU to one VM, it shares a card. Four simultaneous guests on a single RTX 3060 each held 25.57. 26.49 FPS 103.7 combined versus 102.9 for one guest. All encoding H2, 6, 4 via NVENC at once. Unlike API translation schemes like Vertio GPU with Venus, at about 2,000 boundary crossings per frame, it crosses per iOctal. 13,792 messages over 813,691 frames, one per 59 frames. And guests hit 98. 100% of native frame times above about 2MS. Caveats as plain as the numbers. Ikuta allocation and zero copy interop are untested beyond enumeration. There is no IOMMU boundary. And the host driver sits in the trusted computing base. Attack surface reduction, not hardware isolation. Promising, but early. Model portability advances on stranger vectors. Parakeet dot Java ports, NVIDIA's Parakeet ASR family, to run end-to-end on the JVM. Competitive with native engines on CPUs. Steve Jobs' 15-minute Stanford speech transcribes in about 15 seconds on a laptop with the 110 MQ4K model, streaming included, Apache 2.0. A 421M-parameter LAM model plays Flappy Bird on a 12th Gen i7 desktop CPU via OpenVINO INT8. A commenter's tip. Pinned to PCORS and dumped the execution graph to confirm the matmulz actually went INT8. And Quen pushed its image line forward with the Quen Image 2.1 release. Agent companies and the validation bill. Once models run anywhere, organizing them becomes the product. Paperclip, an MIT-licensed node dot J, a server with a React UI, runs a company of AI agents. Org charts, roles. CEO, CTO, engineers, designers, marketers, any bot from any provider. Budgets and governance. If OpenClaw is an employee, Paperclip is the company. Manage business goals, not pull requests. 12 subsystems cover work, heartbeats, governance, budgets, routines, and secrets with adapters for CloudCode, codecs, CLI agents, and HTTP bots. If it can receive a heartbeat, it's hired. And overspending automatically pauses agents. The editor's video pic puts that architecture through a real workload. Network Chuck, with Paperclip creator Dota making his first on-camera appearance, built an all-AI IT department. Dumbledore as CEO on CloudCode, Ron as CTO on Hermes Mad-Eye Moody on security via codecs, plus helpdesk agents on local models. On a genuine studio mystery, every toilet flush knocked everyone off the NAS. The agents traced the Microtik switch dropping four fiber runs within two seconds of each flush, 36 times that week. Port 5 had logged 20,052 linkdowns since boot versus 123 on a healthy port. Root Cause Cheap third-party Amazon SFPs from one batch with a 50% in-service failure rate. Genuine transceivers fixed it. A real investigation, artifacts and escalation chains included, that reached the human only when the agents were stumped. Agent companies need institutional memory, which Asymptote Labs beacon supplies. An MIT-licensed layer capturing session history across CloudCode, Cursor, Codex, OpenCode, Klein, and 20-plus harnesses, normalized into a common OpenTelemetry event model. Prompts Tool calls Approvals Model Context Protocol MCP Interactions Token Usage The Loop Run Capture Evaluate Extract Review Reuse Turns a debugging path that fixed an obscure issue into knowledge the next agent starts with. Your cursor sessions can improve codex. Forwarding to Splunk or Datadog is optional. The bill arrives at CI. Linear reworked its pipeline after agents accelerated shipping faster than validation could absorb. Every PR still passes through CI, so CI became the bottleneck. 6-BIB of faster third-party runners, 34% average job speedup. A native TypeScript compiler 73% off the type check median. Pure AST Lint rules 68% off API Lint. Capped fetch depth slowest gate from 94 to 20 seconds. Schema snapshots instead of replaying migrations. Batched checks saving about 87,000 runner minutes monthly. And 8 shards instead of 4 cut PR wait from over 6 minutes to just over 5. Despite a year of test suite growth. Without the rework, today's suite would take about 11 minutes. Roughly 2,000 tests are added weekly and agents now write the majority of our tests. Measuring what models actually deliver. When generation is cheap, measurement becomes the product. FIDES, a protocol posted to ArcSiv, treats an LLM-generated trading strategy as three artifacts to reconcile. The stated rationale, the executable code, and the track record that code produces. Dual delivery returns both PROS and a self-contained function from one call. The code runs in a sandbox sub-process with whitelisted imports against a lag 1 out-of-sample backtest. Date positions use day to 1 signals. That the model cannot game. Three concordance gaps are scored. Does the code implement the stated rules? Does it execute as written without look-ahead? Does the claimed edge survive out of sample? The findings are unflattering. Across 40 strategies on 8 liquid US ETFs and 4 models. Concordance did not predict profit. The most speech-code consistent model scored a perfect say-do gap and still lost to buy and hold. Only two of 40 strategies beat it both on TLT. And a plain SMA 50,200 crossover beat every model's mean sharp. Robust to an alternate window in costs. Concordance is thus necessary bookkeeping, not evidence of edge. Self-assessment was badly calibrated. Most strategies claimed to beat buy and hold, exactly one did. And swapping the LLM judge for a second model flipped verdicts on over half the items. Contrastive language models build that discipline into architecture. A frozen QEN 38B backbone with 20M parameter projection heads for states and actions the C, L, M8. B head is 75MB, trained with bidirectional infants loss plus mid-training hard negatives over about 60M question-answer pairs, and a million agentic trajectories. The staging matters. Pre-training alone reached 52.1% top one on held out questions. Hard negatives lifted it to 69.2%, and training on them from the start peak 7 points lower. Instead of generating text, the model scores how well candidate actions align with the state. Ranking best event answers, routing tools. And as a verifier over sampled solutions it sets 87.6% on Terminal Bench 2.1 and 81.6% on Deep Stewie. 4.1. 5.7x faster than the JEV baseline. Frontier pricing gets the same receipts treatment. An independent harness ran 25 graded tasks. 23 across three difficulty tiers plus two built to probe Fable 5.1's supposed long context edge. With automatic graders' exact answers, SQL against a real database, hidden unit tests, and every failure hand-checked. Opus 5.5 hit 100% at roughly 9 cents estimated per past task in 27 seconds average. Fastest and cheapest of the perfect scorers. Fable 5.1 matched accuracy at about 26 cents. Sonnet 5 at about 11 cents but 61 seconds. Haiku, 4.5 failed silently. Wronged 3 for 3 on accounting problem 429. 429, 233 against a correct 6205. The Frontier tier separated them. Sonnet 5 burned about 3.5x the tokens about 218k versus about 61k, canceling its lower per token price. On the two Fable Edge tasks, Opus 5.5 at low effort passed three-thirds at roughly a third of Fable's estimated cost. Caveats as plain as the rankings. Small samples, one account. A CLI harness. First day timing. What hardware can and cannot prove. At the bottom of the stack. The day's best read dissects what trusted execution environments actually prove. The explainer at t.public.computerwalks met as MuseAgent. Which books travel, fills forms, and sends messages, each instance in its own VM on meta servers. Through the attestation machinery of verifier needs, the CPU hashes the initial image firmware, kernel, and entered. Command line into a 48-byte S HA384. Launch measurement fix before the first instruction runs. Then, signs a report carrying that measurement. The TCB version, launch policy, and a verifier nonce, under a per-chip key chaining to AMD's root. The verifier checks the chain, the nonce, and that the measurement matches a reproducible, approved build. The hard lessons earn the read. A valid vendor signature proves the platform, not the image. An attacker can boot a modified image on a real chip and receive a genuinely signed report with a different measurement. So verifiers must enforce whitelists of approved measurements, multi-signer policies, and append-only transparency logs. Debug-enabled launches sign validly, but stay readable. S-E-V. S-N-P's measurement never covers post-launch changes like attached disks or downloaded binaries, unlike TDX's runtime measurement registers. And every release brings a new measurement, so a mismatch may be a mistake. Refuse first, then investigate. The sharpest line, a measurement identifies code. It doesn't vouch for it. Meta's claim that agent data is encrypted with a key only they hold, so not even Meta can access it, remains an assertion awaiting the external audit planned later this year. Hardware randomness is having its own epistemic moment. A flat-assembler forum thread reports that AMD's RDRAND and RDSED instructions never produce a zero spanning the full requested register width. Found by a Brazilian assembly programmer charting the 16-bit number space, the zero bucket never moved on AMD hosts while Intel machines filled it normally. Corroborated so far, nine days of testing across two AMD systems produced not a single zero. Zeros do appear in the low 16 bits of wider requests and AMD. Whose first reply misread the results. Escalated without an explanation. No root cause exists. What would settle it is AMD's engineering answer and independent reproduction on more silicon. Until then. Never lean on RDRAND alone. Mix entropy sources. The physical layer also gained a consumer detection story. Zuckoff, a free app by Polish developer Paweł Seidlawski, fingerprints the Bluetooth signals of smart glasses. Ray-Ban Meta, Oakley Meta, Snap Spectacles. And reports when a pair is in the room, over 5,000 App Store downloads in its launch month. Its limits are honest. It cannot tell whether the glasses are recording or who wears them. Only signal strength for rough proximity. Turning an invisible problem into a partially visible one. And it is hard to take down legally, since it intercepts nothing. Meta, for its part, promised a July update to detect when the recording LED has been physically tampered with or destroyed. With CTO Andrew Bosworth insisting the camera was designed to be noticed. Awkward. Beside the product's core proposition that the camera doesn't look like one.