← Back to search

Open Weight Models and Open Source Harnesses | Episode 56

AI Security Ops · 2026-06-13 · 37 min
relevance 67 6079 words Episode page ↗ Audio ↗
Show full episode description
In this episode of BHIS Presents: AI Security Ops, the team looks at what it actually means to own your AI stack. Open-weight models and open-source harnesses are no longer just lab toys. They are becoming practical options for security teams that care about where their prompts, code, client data, findings, and tooling actually live. The core question: when your work depends on AI, how much control are you willing to give away? We dig into: - What data sovereignty means for security teams - Why token sovereignty matters in agentic workflows - How provider terms can become a business risk - Open-weight models vs. truly open-source AI - Why harnesses like Hermes and OpenCode matter - Where cloud providers may apply fewer restrictions - The tradeoff between local control and hosted capability - Supply chain risk in models, harnesses, and plugins - Running local models with Ollama, VLLM, and similar tools - Why “local” does not automatically mean “safe” - How to start experimenting without buying expensive hardware - The next risk frontier: local prompt injection Owning your AI stack does not magically eliminate risk. It moves the risk. Hosted models create exposure around data, terms, pricing, and availability. Local models create exposure around maintenance, supply chain, permissions, and prompt injection. The security win is not blindly choosing local or cloud — it is knowing which layer you need to control, and why. ⸻ 📚 Key Concepts & Topics Data & Terms Risk - Prompts can contain code, client data, findings, and operational context - Hosted providers may inspect, retain, or restrict usage - Terms changes can affect entire security workflows - “Allowed yesterday” does not guarantee “allowed tomorrow” Token Sovereignty - Agentic workflows burn far more tokens than simple chat - Rate limits, usage windows, and pricing changes become operational dependencies - Local hardware shifts the constraint from API quota to compute capacity - Cost control is part of architecture, not just procurement Models vs. Harnesses - Open-weight models provide downloadable weights, not always full training transparency - Harnesses provide the tool loop, permissions, memory, and provider adapters - Hermes, OpenCode, Claude Code, Codex, and similar tools shape what the model can actually do - Risk often lives in the harness around the model Local Stack Tradeoffs - Local models improve control over sensitive data - Self-hosting adds maintenance, patching, networking, and monitoring responsibilities - Tools like Ollama, VLLM, and Llama.cpp lower the barrier to experimentation - Expensive hardware helps, but it is not required to start learning Supply Chain & Prompt Injection - Model weights, plugins, skills, and MCP servers are all supply chain decisions - Local agents with shell access can turn prompt injection into local impact - “No provider guardrails” means you own the safety controls - Permissions, sandboxing, and audit logs matter more as the stack gets more autonomous Practical Starting Point - Pick one harness and go deep before chasing every new tool - Test real tasks, not toy demos - Compare hosted and local workflows honestly - Decide which layers you need to own before you need an emergency exit #AISecurity #LLMSecurity #CyberSecurity #ArtificialIntelligence #OpenSourceAI #LocalLLM #AIAgents #SecOps #InfoSec #BHIS #AppSec #PromptInjection #SecurityArchitecture ---------------------------------------------------------------------------------------------- About Brian Fehrman - https://www.blackhillsinfosec.com/team/brian-fehrman/ About Bronwen Aker - https://www.blackhillsinfosec.com/team/bronwen-aker/ About Derek Banks - https://www.blackhillsinfosec.com/team/derek-banks/ About Ethan Robish - https://www.blackhillsinfosec.com/team/ethan-robish/ About Ben Bowman - https://www.blackhillsinfosec.com/team/ben-bowman/ (00:00) - Intro: Owning Your AI Stack (01:43) - Data Sovereignty, Token Sovereignty & Terms Risk (03:38) - Provider Inspection, Prompt Data & Business Exposure (08:09) - Where the Guardrails Live: Model, Harness, or API (12:12) - Open Weights, Frontier Providers & the Innovation Race (14:53) - Local Models, Open Harnesses & Real Hardware Tradeoffs (24:24) - Self-Hosting Reality: VLLM, Ollama, VPNs & Maintenance (31:25) - Getting Started: Pick a Harness and Run Real Tasks Click here to watch this episode on YouTube. Creators & Guests Bronwen Aker - Host <li
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
What it means to own your AI stack via local open-weight models and open-source harnesses, covering data, token, and terms-of-service risk.
Benefits
  • True data sovereignty by running local open-weight models
  • Avoid token cost inflation and rate-limit risk
  • Escape provider terms-of-service restrictions on use cases
  • Open-weight models match frontier models from months ago
  • Bedrock raw API calls are less restrictive than vendor harnesses
Use cases
  • Derek runs a Qwen 3 27B-parameter model locally with a 200K context window (~80-100K usable for the agent)
  • Black Hills now offers agentic AI pen tests via a custom harness on AWS Bedrock using GLM 5.1 for heavy lifting
  • Opus 4.8 blocked a GitLab merge request on offensive tooling and a nuclei results cleanup, violating ToS
  • Anthropic posted its first profitable quarter; new executive order requires 30-day model review before release
KPIs / results
  • Qwen 3 27B parameters, 200K context window, ~80-100K usable
  • Anthropic guardrails: ~60% harness/API, ~40% model (Derek's estimate)
  • Executive order: 30-day model review (originally 90 days)
Tools / build
  • Hermes (agent harness)
  • OpenCode
  • AWS Bedrock agentic pen-test harness
  • GLM 5.1
  • Qwen 3 27B local model
0:00 / 0:00
Welcome to AI Security Ops, the podcast where we cut through the hype and explore the real-world intersection of artificial intelligence and cybersecurity. Each week we examine how AI is reshaping both sides of the security landscape, the threats we're facing, and the defenses that we're building. I'm Bronwyn Aker, and today I'm joined by Ethan Robish and Derek Banks. In this episode, we are going to look at local open weight models and open source agent harnesses like Hermes and OpenCode. And we ask what it really means to own your own AI stack. We'll dig into data sovereignty, token sovereignty, and a quieter risk. What happens when the provider reading your prompts decides your line of work is no longer covered by its terms of service? Yeah, that's a biggie. This show is brought to you by Black Hills Information Security and Anti-Siphon Training. BHIS helps organizations like you identify and close real-world security gaps through penetration testing, adversary emulation, purple-tie engagements, and managed detection and response. Anti-Siphon delivers hands-to-yout. Anti-Siphon delivers hands-on, practitioner-led training built around real attacks and real tools so you can apply what you learn immediately. Learn more at blackhillsinfosec.com and antisyphontraining.com. All right. Derek, you've been doing lots of experimentation with this stuff. I'm going to hand this off to you with a question. What is all this about? Data sovereignty, token sovereignty? Seems like there's a lot of sovereigns in this room. Sovereignty. Sovereignty. It reminds me of my pastime of reading. Sounds like some magic spell or something. Token sovereignty sounds like how to conserve your mana or something, right? Anyway, sorry I'm a nerd. So data sovereignty, the definition would be your data, right? Your prompts, context code, client data, findings and a report, like anything that you might put into input into an AI model is data. And the sovereignty would refer to where that data is, who has access to it, who owns it. And so, you know, I think that I've been doing AI and security work now for a couple of years at the intersection of this. And, you know, it's kind of less than it used to be. But, you know, the general theme from security practitioners and hackers is, you know, you kids get off my lawn when it comes to sending data to third party providers such as OpenAI and Anthropic. Maybe it's a little less now, but I do know that there are agreements in place, you know, with, you know, or sorry, the privacy policy that's stated by specifically those two companies, you know, says that they won't use your data to train models. But that doesn't mean that they're not inspecting your data and looking at it and storing it at all, right? So it's kind of something that happened to me here over the last couple of days, kind of made me want to do this episode. And then token sovereignty would be the same kind of idea as we move into the agentic era. I mean, I guess we're already here. Everything is basically in the AI world revolving around tokens and how much tokens cost. You might have heard the term token maxing, like with the subscription, like with quad or the codex. Basically means using all your allocated tokens in the window of time to get the most out of it, right? And so there's a we're in an era of GPU, therefore token shortage. And so I think that people are starting to realize that this stuff gets expensive, the more tokens that we use, especially if you're paying API rates. And I think that it won't be long before, before, you know, folks start to realize that, you know, inference providers start to jack up the rates a little bit because, you know, they've been subsidized for some time now. However, Anthropic just posted its first profitable quarter. And so at least that's what I heard. So I think that, you know, can you what compute do you have access to? Like if you have these business processes that are relying on, you know, a certain cost of token, what happens if that gets jacked up? Right. So token shortages, rate limits, usage windows, token cost inflation. I mean, your capacities, you're either your API and, you know, your provider or your hardware, which is a whole nother ball of wax. Right. And so and then the third one is something that just kind of happened to me. Right. Where I guess, you know, Ethan mentioned in chat earlier that Anthropic mentioned that they're going to start tightening like cybersecurity use cases in terms of services. And just, you know, the last couple of days, I've had difficulty doing the exact same thing that I've been doing, you know, coding wise on a project that is cybersecurity related. It's offensive tooling. And I was hitting a lot of errors that I was now violating terms of use on Opus 4.8. And yes, I know there's the cybersecurity verification program. We're in it. I mean, we're in it and we're still getting the alerts. Yes. And I will say, I mean, look, I'm not being super critical because surprisingly, Anthropic has responded to us this morning, which like I was actually kind of surprised. They're like, well, if it happens again, just like, you know, here's what you do. And I was like, I didn't know that. And so, but I mean, you know, it's, it's, you know, not just cybersecurity. I mean, I found, uh, was it last summer or last, last spring, actually, I think, uh, I think it was Anthropic, uh, also did that like in the medical world, like can give out medical advice to some degree, like you have to be licensed. And so anyway, I'm just saying that they're reading your prompts. And so, you know, they can decide, uh, you know what, we're going to open up a, our own, you know, hedge fund and we're not going to let you use this for financial transactions or finance trading anymore. I'm being arbitrary. Right. But I mean, that's the risk is there. That's what we do in information security. We point out risk and it's your ability, your, your, you know, your responsibility to interpret that risk as it applies to you. And so, and then, um, go ahead, Ethan. Oh, that it just seems like the common theme between these that ties it all together is risk. Like you, you just said it like data sovereignty, token sovereignty, terms exposure, all different types of business risk. And, you know, with, with the service you're using and relying on with the data that your customers like expect have a, uh, maybe a legal obligation to protect. Um, and then just, yeah, uh, I, I do have a question about the kind of guardrails that you're talking about, like that you hit it with Anthropic. Does, how much of that do you think is baked into the training of the model and how much is it like, um, protections they put around their API? So for instance, if you tried to do the same thing using Anthropics API versus using the Anthropic models in, um, Bedrock, like, do you think you would hit fewer restrictions in Bedrock? Uh, that has been my experience, um, that, um, it's kind of hard to tell, tell sometimes from the error message that you get. Um, I think that in this case, uh, the same thing works with Sonnet or like with what I was doing. So I've, I've actually hit the, um, two things that happened to me yesterday. I was trying to do a merge request on offensive tooling and, uh, um, and get lab through clause. Like, Hey, go ahead and do, you know, the request. Right. And that's what triggered it. Okay. So there must've been something in the code in the like, I guess. Right. And then I tried to go back, you know, hit escape twice and go back and redo it and still triggered the same thing. And then another thing, um, I was having it try and clean up a disc and download nuclei results, which, I mean, come on, that's pretty benign. I'm just, and, and I don't know, maybe data exfiltration, but I was downloading it. Like, I don't know. It just, these were things that I had been doing a lot and they stopped working with Opus 4.8. So I think that if I had to give percentages, I'd say 60% in the harness and the API, um, and 40% in the model. Hmm. And when you say harness, you're, you're talking primarily like the backend harness, like, uh, yeah, no, I meant like, you know, like there were, I think in terms of like the protections all around, I think quad code, like built in the system prompt or certain protections. And it was actually, you know, late last fall where they actually put in, uh, pen authorized pen testing and, uh, CTFs and to call it system prompt as an authorized behavior. Right. And then I think in, inside of that, I know, after that, there are probably protections or there should be protections on the API, right. Looking at your prompts and the input. Um, and as a side note, I bet their telemetry is amazing. Like what's coming in from everywhere. Um, talk about a big beta problem, same with open API or open AI, but then I also, there are also guardrails in the model. So I think there is like three, at least three, like, um, yeah, defense in depth points, I guess. Well, it's, it's interesting that they would go to the trouble of putting that on the doc, the client side, the harness side, because it, it's kind of the same thing you get with, um, web application security. Like you can put all the logic and permission checks you want in the client side JavaScript. None of it matters if you're attacking it. So you have to enforce it on the server side to make it actually effective. To your point, like, or to your question, you know, like, uh, you know, do, you know, do you think that, you know, where is it? And what Bronwyn said earlier, like I've been doing a lot of agentic work. I've been using Bedrock and I found it like Bedrock is like less restrictive in terms of like, it's just like raw API calls to a model and Amazon doesn't seem to be putting much in the way. So, uh. Well, if they did, malicious actors wouldn't be launching as many attacks from AWS systems. I, maybe I gotta say that those attackers need to have some deep pockets because that crap ain't cheap. Right, John? No, it ain't. Um. So how do we, how do we make it cheaper, Derek? I think you have the answer sitting on your desk, don't you? Yeah. It's actually under my desk. But, uh, so two things. Uh, one, like, again, I'm not knocking the frontier providers. Uh, this is all uncharted territory. And, and, you know, I think you had mentioned also again in chat earlier that you thought maybe this was tightening of guardrails to try and. And for them to figure out how to release mythos level capabilities out, uh, into, uh, you know, the world. You know, and I think they've got to do something. And now, you know, like if there's a, uh, you know, executive order where the government, uh, is wanting to, um, review models for 30 days, uh, before they're released to the public. Which, uh, I mean. That's crazy. I hadn't heard that. Oh, yeah. Yeah. It's a new executive order. Wow. I don't, I don't. We originally wanted 90 days and they got talked down to 30 days. But. And the frontier providers were like, we don't want to do any of that. So no one's happy with 30 days. Right? Yeah. What good is it? I mean, it's like that, that meme of the guy, you know, doing this to, you know, not even touching people when they're supposed to be patting them down. Like. What do you mean that there's probably nobody in the government that's qualified to actually do like a, an effective, like evaluation of the capabilities? I mean, I wouldn't say, I wouldn't say that, but like. I bet they don't work in the executive branch. Yeah. Well, maybe they do. I don't know. There probably are. But I, so to me, like, I like your analogy. Mine's also, was it the old, like the Dutch boy putting the finger in the, the dike? That's what I think of. I was like, there is no way at this point, the cat is so far out of the bag that the, that's. It seems like most everything that we're going to do is just going to stifle us innovation. Like we're in the race. But yeah. One of the, speaking of the race, like most of the open models that get released are coming from China. And one of our points here, spoiler is like, Hey, to have true data sovereignty, you run a local model. You can only do that with open weight models. And one of the key differentiators of like open weight versus frontier is like the frontier models are better. They're more like bigger for, for sure. Like, unless you have access to the data center yourself. But, um, that, that, but the argument, the counter argument is like the open weight models are just as good as the frontier models were like several months ago. So it's, it's true. Well, there's a little nuance to it, right? Because like what I can run here locally on my side, right now I have a Quinn three six, 27 billion parameter model running with a 200 K context window. Cause that's the context window size. I think not two 56 either way. There's a portion of it. That's the system prompt for the harness and it takes up X amount. And then you get, I think it's something like I was getting like a 80 K or a hundred K like for the, um, the agent. And I mean, I mean, it's not the math maybe, but so I've got, you know, a smaller like window and also the model is smaller, but also there's open weight models that you can run in the cloud. And so, uh, and since we just sold this, I guess it's like an offering. We now offer, uh, agentic AI pen tests at Black Hills. Apparently I do. Um, and, uh, and I'm really excited about the platform and the, and so, um, basically it's a, a, a custom built, uh, you know, harness that runs using AWS bedrock. And I use GLM five one for the heavy lifting. And the reason why is because it was specifically designed again, a Chinese model for a long range agentic tasks. And in testing, when I was using, it was Opus four, six at the time when I was doing development of it. And, uh, Opus four, six is really expensive and I racked up a pretty big AWS bill, but it was worth it because the platform like early in testing actually discovered a critical for a customer that our humans had missed. Right. And that's happened now, like a couple of times, but there is also the converse that they, the, the AI is sometimes wrong and the human has to correct it. So, but, but anyway, when I started. Well, good hits and bad misses. We, we deal with false positives all the time. Oh yeah, exactly. And so, but anyway, GLM five one, let's just say that it's 85% as capable as Opus four, six. And then that I was getting like basically the same results and testing, like the same findings consistently between the models. And it was, and, and, and GLM five one's like 90% cheaper. Like we went from it being like a thousand dollars a run to like a hundred dollars a run. And I also then use Opus to come through and now like re rate and re valve re validate and then do reporting. Cause it's a larger context window. So I'm using an ensemble of, of different models and there's so many out there. Like now I'm in a position where we could start testing and we could try it like a new chemi model just came out. But to your point. Yes. The Chinese definitely have the open weight model. Like they're, they're doing better. And I, you know, I want to believe in Gemma four that came out recently, but I'm having issues with it. And the GPT stuff that came out, I never really got to do anything useful with, but I will say that the Quinn. Um, so I've been using Hermes with the spark and I had it do a explain the DJ spark quick. Oh yeah. That's a good, uh, the DJ spark. DJ. So Nvidia basically has a small form factor computer that has a GPU in it. Um, or an AI chip. I mean, it's not, you're not playing games on the thing. It's a, uh, a black world chip. I think, uh, that has 128 gig of Ram, um, two terabyte disc. I forgot what CPU is in it. Um, it's basically built to be like a little mini desktop, like supercomputer. Right. And 128 gig of Ram's pretty decent for running, you know, these edge, like, uh, you know, um, you know, open weight models. And so I put, uh, VLLM instead of OLUM on it because it's faster and I'm getting pretty decent results. I just started kind of down the road. Our idea is, as inference costs start to rise, can we effectively do like, you know, everyday pen tester tasks from the DJ at spark? And the answer might be yes. Um, but, uh, but then, you know, that GLM model for comparison, that GLM five, one model, I believe is a 400 and some billion parameter model to run it. I actually just, I'm writing up stuff now because now the question's coming from management at Black Hills is, well, what if we wanted to run the, that whole AI platform, not in bedrock, but here locally, what would we need? And the answer is 16 H200s and, and, and, and two like clustered systems, eight in each system. Um, and, uh, cause we need about 1.5 terabytes of, uh, GPU. You realize DRock is going to kill you because now he's got to upgrade your conditioning again. DRock's asked me, who asked me to size it. So, uh, it isn't me asking for it. I'm on the fence of what should we do it. And that's with me, like getting to be, you know, part of like making it, but there's a couple of things you have to keep in mind. It's not just the, um, it's not just the model itself, right? It's also the context window or the KV cache, right? If you talk about like the attention mechanism that has to be in GPU too. So as you're going through this, you know, auto-aggressive loop and the LLM predicting the next token, you want to cache what you've computed already in, in, uh, in the attention mechanism. So you have to do it every single time, right? And so that takes up GPU too. And so to run it at the full context window, uh, and then also run up to 10 tests, so like 10, like engagements a week. So we have anywhere from five to 10 externals. The calculation ended up being 1.5 terabytes of GPU, which is like 16 H200s. So, so I think you've covered several different levels here that we can tie back into our, our different sovereignties. Um, so if we talk about like data sovereignty, uh, what one level is, Hey, anthropic or open AI, like you're sending all your stuff to this AI company, which you have agreements with, but maybe you don't trust them. Or maybe you just, you know, don't want to have to trust them. So the next level would be, okay, take an open weight model, which you can host yourself. Give you, if you have the hardware, if, um, but if you don't have the hardware, you can host it on Bedrock or what's the, let's see a jury one. Like foundation or something. Foundry. And there are other providers too, right? Right. Like NVIDIA. Yeah. Someone else's data centers. Someone else's GPUs. Yeah, but you're still hosting and sending data to a third party. Yeah. You're shifting who you trust. You're shifting from the open AI, the anthropic, the frontier models to whoever's hosting your data center. And then the third level is you bring your own hardware and the DGX spark is a way to, to do that, to start doing real work. Um, but I think, so data sovereignty, I mean, you're obviously shifting, but then also like the, the terms exposure or the, the limits that the rate limits and stuff. So with the frontier models, you've got your rate limits and there's really no way around that. Especially if you're on a subscription, you've got your five hour blocks and your, your week blocks. If you start hosting that yourself in like the AWS data, um, like Bedrock, that, that all goes away. Right. Derek, there's no like rate limits to speak of. Um, so I think there are some rate limits, but I, I spawned like, you know, eight concurrent agents at a time, 16 concurrent agents. And I haven't really hit it. I think there are some rate limits like per model, but so far, no, I have not had an issue. Okay. And which by the way, uh, uh, Amazon Bedrock does have a pretty clear and concise privacy policy where they basically say, we're not saving any of that data for anything. Yeah. They're just, I mean, it's basically an API call to them. Right. Yeah. Um, so I imagine they might have some restrictions on terms of use because AWS in general just, you know, has terms of use. So if you're doing, obviously if you're doing completely illegal things, you're on a silk road or whatever, and, or you're spitting up, um, hacking campaigns. Oh, I promise your point earlier. Maybe people are doing that anyway and getting away with it, but presumably it's against Amazon's terms of use. And so to get around, like if you're doing something shady or maybe contextually, okay, morally, arguably, okay, but companies don't want to be a part of it, then you're kind of forced to shift it even further to your own hardware. Yeah. I mean, even with Amazon, there's nothing, you know, with it, like it's just a company with a terms of use that's theoretically possible for them to change their terms and be able to now take a cornerstone of like, what's going to be like our future business away. Now, are they going to do it? I doubt it. We're not doing anything illegal. Right. So, but I mean, the possibility exists again, like with looking at risk, there is, the risk is not non-zero, just like the risk of using a Chinese, like open weight money. Model with a coding, you know, a coding harness, like open code or call and code or codex or whatever, you're using a Chinese open weight model. The possibility exists that it could locally, like code a back door and do something. Now, is that going to happen? You know, practically speaking, no, but it's not, the risk is not non-zero. I mean, but I. Same risk would be present in frontier models as well. That's just, which, which country do you trust more? Yeah. Which companies and where they, which country they reside in, I guess, do you trust more? But I mean, that's going down. Oh, sorry. Go ahead, Bronwyn. No. Well, there's also the fact that I know in the reports I've seen more and more supply chain attacks seem to be happening. I mean, supply chain has always been an issue. But now it's like, okay, if I'm using Anthropic, well, Anthropic is a big target. They've already had source code leaked. I'm sorry. I'm just, I'm babbling here. But the whole idea is that even if we're trusting the third party frontier models, their targets, they might be attacked directly. We go with a hosting service is the same kind of thing. And so for the truly paranoid, local may be the only way to go. The problem is that the more you're doing yourself, the more maintenance and upkeep. How much work are you doing, Derek, to develop these new tools that are running locally on that Spark? Wait, are you asking me like how much like work have I done like to get the Spark up and running? Well, I mean, it's not just getting it up and running. It's the fine tuning. It's the torquing. It's the testing. Yeah. So the heaviest lift was because I chose to go with VLLM. If I would have went with a llama or llama CPP as an inference engine, because if you're going to get into open weight models, the first thing is choosing an inference engine. Probably the most popular is a llama and it's it's pretty easy to get installed. But if you're going to go down the road. Very easy. Yeah, exactly. So most of this is pretty easy to get up and running until you start saying, you know, start using things like VLLM where it's more of like kind of like a production ready. I wanted to learn that because I was being there were rumblings of being asked of, hey, what if we want to like, you know, host all this ourselves? Like, oh, yeah, that's a good question. What if we want to do that? And so VLLM is what, you know, the, you know, the kids on the street use to host like production models. Right. And so I went with that. That was a little bit of a heavy lift. I had Claude helping me honestly figure it out. Now I have, you know, a Docker project that I can just switch between models. So it's pretty, pretty easy getting that. I actually wrote a setup guide. I could actually publish it on my GitHub or something. I mean, let's polish it up a little bit. And so now the Spark, I got VLLM running a couple of models on there. And then I have a Hermes agent and a Kali VM just because I wanted to put in a VM instead of, because I wanted to be able to use the Spark when I wasn't at home. Right. Because that's, you know, cloud versus local. And so I put a tail scale network on the Spark and then on the Kali VM because I didn't want to put tail scale on my production computer. I didn't know if, you know, systems would be happy with me if I did that. So I just did it in a VM and it works pretty well. Like I was up at the, you know, my daughter's swim practice at the pool yesterday, 30 miles from my house. And I could, you know, access the Spark like I was sitting next to it. So everything like, but the, I'd say everything was pretty easy standard open source install stuff except for getting VLLM up and running. But now that it's working, I only have a couple of complaints. One complaint I have about the Spark is it's not easy to encrypt the hard drive. I haven't gone down that road yet. Brian Furman did. It was an exercise in frustration, apparently. I think he's now, there's actually a cheaper one from Asus. I can't remember the model name that it's about $3,500. Because, you know, like Ethan was saying, there's kind of different levels to this, right? You know, the first level is, you know, cheaper inference bedrock open weight models. It's kind of, you know, it's, you know, not free, but it's pretty cheap depending on your, you're doing usage paying. It's $3,500 to $5,000 for local inference of 128 gig of RAM. And then, yeah, if you're going to host the big models, well, you're looking at, what, a quarter million dollars? Interesting. Something in that neighborhood. Maybe more. I think this entire conversation seems very similar to, like, we could be hosting a self-hosted podcast right now. Like, it's the same argument of why you would want to get away from big companies, like software as a service, and host it yourself. So, are we all using OBS and VPNs? Like, I don't even know how that works. Yeah. But, well, I mean, so I think that the whole idea here is that every company and every individual has different needs and different risk postures. Absolutely. And so, it depends on what you want to do. If you're listening to this podcast and you're just getting into AI and cybersecurity, what I would do would be to install, like, Hermes Agent on a VPS or VM and go to NVIDIA and try their free inference to get started. Because then, you know, and setting up Hermes is pretty easy. They have, like, a, if you, similar to OpenClaw in the sense that they have a, like, kind of like a walkthrough, like, configure it this way. So, you really just need to go to NVIDIA, create an account, and get an API key. You could use Anthropic API keys or OpenAI API keys. They have a lot of different, like, model providers or options. And then just start using it. I actually had the thought earlier that maybe I should take a step back away from CloudCode because that's what I've been using for so long and just kind of, you know, test the waters with some other stuff. I haven't even really used Codex that much yet. And so, you know, maybe it's time to kind of to branch out a little bit. So, but I guess what I'm getting at is, like, pick a harness and go and kind of learn it and get to use it and do it as cheaply as you can. I know one of the things that I've been seeing more and more and I've been also seeing in my own experimentation is that it's easier to go deep with a single set of whatever tools. Whether, you know, regardless of the model you're using or the harness you're using, pick one, go deep with those. After you've gotten a deeper understanding, then go and do, like, what you're going to do. Branch out, see what the other kids are doing, and expand from there. And what I've seen is a lot of the knowledge is transferable from one set to the other, but it's not until you get into the weeds that you get into the nuances. Yeah, and if you're sitting here thinking, man, I don't even know, like, what to even start to do, like, for as a project. I mean, that's okay. It's sometimes hard to, like, get started. What I would do is pick something that you do at work. Like, if you're a blue teamer and you're looking at log files, take a set of log files and have, you know, your harness. Hermes, OpenCode, you know, CloudCode, Codex. Gemini has one. Gemini name theirs the same that they name their models. It's called Gemini. Why do you do that, Google? Well, just also like Google, they have already canceled or, like, deprecated that. Oh, at all? And now they've rolled it into AGY, their anti-gravity CLI. Good lord. I can't even keep up. And I have a, like, that's the, I pay for Google Pro, like, $20 a month just to, like, be part of their ecosystem. Oh, God. And I still can't keep up. I'm reading through alerts and articles and updates at least two to three hours a day. Over and above, you know, just, and I'm a speed reader. And I still can't keep up. Yeah. But that's the challenge for anyone listening this week is pick a harness and the cheapest inference that you can if you don't have it through work. And try and run some real tasks through it. You don't have to be creative. You don't have to have it, you know, create you, you know, your own, you know, personal GitHub or something, right? Like, you don't have to have it recode a platform or something. Just, like, get started with, like, everyday kind of tasks. That's where I think personal, like, local open weight, open harness agents are going to shine is helping you do things like, you know, draft, you know, reports. I use Claude and I'm like, I should probably start doing this with, you know, Hermes instead. I have it help me come up with my weekly, I do a weekly status report just so I know, even if no one's reading them, like, what I've done for the weeks, right? Because I have teenagers and I forget. And so, you know, just the things that you do that you have to do all the time that are repetitive and see if it'll help you out. So we talked about, like, hosting your own models, right? And you don't have to go out and buy a DGX Spark. You don't have to go buy the most fancy MacBook Pro or whatever. So there's, I don't know if we have show notes or not, but there's a GitHub project, Alexis Jones. It's called LLM Fit. LLM Fit. Yeah. And if you download that, run it, it will give you a list of open-weight models that will work and, like, how well it predicts they'll work on your hardware. Yeah. And there's actually a website, too, but I can't, I don't remember what that one was. That's one of the features of Misty that I like a lot. Okay. I mean, Misty is primarily a GUI, but you do have the ability to download and run multiple models locally or via API key. And you can do side-by-side chat. So you're sharing the same prompt. You're sharing the same data stack. You're sharing all of this stuff across, you know, I think three is the maximum number of chats you can run simultaneously. But it's nice because when you're going to connect to a llama or hugging face, it shows you right away whether or not a specific model will perform well or poorly. And that is another nice feature. And, you know, if you're focusing on learning the differences between the models as opposed to the harness, Misty, M-S-T-Y, is a nice tool. Well, I think that's probably a good place to leave it. You know, we're 36 minutes in, and so I think, yeah, I'm sure that we'll have more to say over the next coming weeks to months about open-weight models and open-source harnesses. And this seems like it's a roller coaster ride that's going into a dark tunnel. But maybe there's light at the end of that tunnel. So we kind of focused on different risk and data sovereignty and whatever, the three different pillars here. One type of risk that cuts across all three of them that we didn't cover, and maybe we can do in a future episode, is local prompt injection. Like, what risk does your harness pose? Prompt injection or even just like hallucination and your model goes rogue and starts deleting stuff on your system. Like, it doesn't matter where you're hosting or what model you're using. The risk is the same. I actually have slides on that in my class, so we can just take from that and do it. Now we know what next week's episode will be. Very cool. All right. Stay tuned. And with that, yeah, keep on prompting. Stay tuned.