← Back to search

Stop Building AI Agents: Build Harnesses Instead | Hamza Tahir (ZenML / Kitaru)

Domesticating AI · 2026-06-06 · 43 min
relevance 51 7423 words Episode page ↗ Audio ↗
Show full episode description
Everyone is building AI agents. OpenAI SDKs, Claude Code, Deep Agent systems, custom workflows, and orchestration frameworks all promise more autonomous AI. But as these systems become more capable, they start running into familiar engineering problems: retries state management orchestration context control durable execution This week we're joined by Hamza Tahir, CTO and co-founder of ZenML and creator of Kitaru, to discuss what happens when agents stop being simple chat interfaces and start behaving like long-running distributed systems. We explore: what an agent harness actually is durable execution and why it matters orchestration vs business logic state management for long-running agents retries, checkpoints, and human-in-the-loop workflows context management and token costs open vs closed agent frameworks why everyone seems to be rebuilding the same layer of infrastructure One of the biggest questions we kept coming back to: What is a meta harness? If you have an answer, let us know in the comments. Kitaru https://github.com/zenml-io/kitaru ZenML https://www.zenml.io Hamza Tahir https://www.linkedin.com/in/hamzatahir/ Pedro Agentware https://github.com/Soypete/pedro-agentware OpenAI Agents SDK https://platform.openai.com/docs/guides/agents Temporal https://temporal.io DBOS https://www.dbos.dev Apache Airflow https://airflow.apache.org Prefect https://www.prefect.io Domesticating AI is a bi-weekly podcast about practical AI for developers. We help you brace the feral open-source AI landscape — so you can tame it instead of getting dragged by it. Subscribe on YouTube, follow on Spotify or Apple Podcasts, and support the show on Patreon. Keep your AI on a leash. Links
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Should you build custom AI agents or harnesses, and why durable execution matters for long-running agents.
Benefits
  • Clear definitions: base LLM + harness produces an agent
  • Open source brings more eyes, security, societal trust
  • Smaller distilled models cut runaway inference costs
  • Durable execution enables background, long-running agents
  • Infrastructure-agnostic stacks decouple infra from business logic
Use cases
  • Categorization tasks hit the biggest premier model, '100% more expensive' than a bag-of-words approach
  • Agent autonomy time grew from a few hours to six, then eight hours per model release
  • ZenML stacks separate orchestration, artifact storage, container registry from business logic
  • Kitaru built as a durable-execution / agent-orchestration layer for dynamic DAGs
KPIs / results
  • Premier model ~100% more expensive for simple categorization
  • Autonomous agent runtime: hours → 6h → 8h per release
Tools / build
0:00 / 0:00
Domesticating AI We have our guest today. He does not have camera problems. Hamza, you're the CTO of ZenML, right? I didn't get your title wrong. I pulled it from LinkedIn, I think. Sweet! Go founder, CTO. Whatever you want to say. We'll put some more connection information for him in the show notes, but just a really quick intro question for you, Hamza, is why do you think open source is so important to the AI landscape right now? Well, open source is important for AI the same way open source is important for software in general. In general, more eyes on a problem in a community where it tends to bring better products out into the world, more secure products. And I feel like especially in the AI from a societal perspective, we don't want decisions being made behind closed doors in a walled garden that will have a very, in a very foundational technology that will have a big effect on normal people's lives. So, I feel like it's, it almost lends itself to being open source first. And that's largely how AI has been until the- Until the past few years, we have had large language models that are closed first. Yeah. Exactly. Although, although I have a, I have a huge doubt that this is anything but a transient moment and we'd go back to open source and open models pretty soon. Cool. I mean, there still are a lot of amazing open source models out there. So it's, yeah, I think it's a good wager. Yeah. It's like, like the way I liked it for him. I heard it from Sebastian Rechka. If you know him, he's, he's talking a lot about MLOps as well. And he's, he was saying that it's about where you, you want to spend your tokens. So you want to spend them on drain time or do you want to spend them on inference time? And I feel like that's a purely economical question, right? Like right now the inference time is just cheaper, but you can already see this year, you know, with, with more and more agent and agentic loops, you have more and more. And with more tokens being spent, the frontier models are no longer affordable for many use cases. And that will naturally drive us towards training smaller, more economical models. And, you know, distillation is quite easy in many ways. So I feel like that's just, you know, where we're going to get there. Yeah, no, definitely. I mean, you look at a lot of the use cases people are using LLMs for and like, honestly, I, I see it a lot of just like people using them to do basic categorization. That is honestly what I want to use them for. And you're like, train your own freaking, train your own. But I don't want to. So yeah. I mean, oftentimes it's, it's like silly categorization too. It's, it's stuff like you could have used a bag of words model for, you know, like, yeah, I want to categorize, you know, like newspaper clippings, which has been a traditional, like data science project, like beginner data science projects forever. But it's just like, oh, well, engineers didn't have to learn the data science so they can just hit an, an API. And of course they always hit the biggest premier model they need to. It's 100% more expensive. I know. That's why I'm saying like my use cases, I want to use it to categorize and I should be using bag of words. And I'm like, I don't care. But that is just me being lazy. Like it's like not the cost equation in that situation. Cool. So I have a good question to start. I think a lot of the times people have associated agents with, with the commonplace agent harnesses, right? You have cloud code, codex, OpenClaw. We've talked about OpenClaw a lot. And now you have Nemo claw and whatever. That is how people think of agents. But like, say you want to build an agent and not just use, use, use one of those, those existing harnesses. You want to build your own custom thing. Should you look at building a harness? Should you look at building a workflow? What on earth is scaffolding? People talk about that. People talk about turns and you're like, what's a turn? I think like what, foundationally, where do you start? You define what words mean. I feel like that's a good start. So, I mean, in a previous episode, you've already talked about what is an agent. I think the word harness is going through a similar, like, like I heard it often on Twitter, like the biggest rebranding win for AI was when we renamed wrapper, like tin wrapper for harness. A chat GPT wrapper is just a harness now? Oh, great. Yeah, it is literally, it is literally everything. So like mentally, the way I think about it is you have the base LLM model. The LLM model takes input and gets output tokens. And everything you put around that model for it to do things is the harness. So that's the operating system. That's the thing that actually makes things happen in the real world. And what gets produced as a result of it is an agent. So CloudCode is a harness because it allows you to, amongst other things, read files, edit files, store memories, call out to the internet. And one instantiation of that, like when you run a coding agent through CloudCode, that is the result of that process being called in that while loop. How do you guys think about that? No, no, I agree with that. That's exactly what I feel like harnesses are. Yeah, I think that's a really good definition. And I would encourage listeners to go back and re-listen because we're probably going to be talking about harnesses a lot in this episode. And how much Hamza hates the word harness as well. I just hate that we keep changing the, because what I just said used to be an agent last year. So now it's like the outcome of the thing is the agent. Now we have meta harnesses. Like I'm getting quite sick and tired of changing. Like it seems the world is changing beneath our feet. But at the end, I've realized it doesn't matter. I need to calm down. And as long as you understand what I'm saying and the listeners understand what we're saying, we're good. We can talk about what actually means, which was your original great question is, should we build our own harnesses? Should we be using CloudCode? Should we be doing custom agents? Like those are good questions, I feel like. I mean, for sure. And I think a lot of times, I guess I've seen this push, particularly in my job. It's like, okay, great. Now we've built our chat bot. Now we need to build our agent. I think there's a lot of overhead with that. It's like, okay, great. That's more than just like, originally we said, okay, the first thing you need to do with your agent is figure out the task you're trying to automate. But now we're just like, okay, great. How much freedom do you give it? Right? I think that's part of like the scaffolding question. Are you letting it determine its own steps? Are you giving it a hard-coded step? I think like N8N was more like, okay, great. This is the step, the outcome, move on to the next one, right? That kind of a more static workflow, which is really popular. But then some of these SDK agents are just like, okay, great. Give us a tool call and a set of turns and we'll give you something. Like there's a lot of options out there that I think beg a certain amount of agentic architecture discussion that I, we have, I don't know. I feel like the open AI SDK just like abstracts that away and makes it be not as much of a building software question as just a tooling question. And that to me, I think is what's scary in the longterm is are we doing it with an architecture and a software goal engineering practice in mind? Or are we just like throwing tools at it and hope that it behaves? Oh, you know, the answer is the latter. That's what most people are doing. I mean, that's what we did in envelopes, right? So I think it's a great way to think about it. Like these abstractions, like we're still trying to find the right abstractions, right? So when you talk about what the open AI agent SDK does or what the anthropic agent SDK does or what something like PyDantic AI does, it feels like we're currently, we're like in sort of a messy state of the world where all the concerns for all the use cases are collapsing into these frameworks. So for example, the open AI agent SDK also does things like deploy a sandbox for code execution, which would be classically more of an orchestration job on the operational layer of the platform. And like if you abstract out from one turn or one harness and you have swarms of agents, that's also more of a platform level question because you need to handle things like what happens if you preempt a subagent or if you steer like a subagent, if it's going in the wrong direction. And how do you attribute that to the parent ID? And those things also, the more deterministic you make them, they start looking like DAGs. So I come from the world of machine learning DAGs, right? Like I co-founded ZenML and ZenML is all about deterministic DAGs for machine learning pipelines. And I guess the more autonomy you give things, the more structured the DAG gets. And again, the orchestration of the DAG is different from the definition of the DAG. So it's a very interesting place to be in right now. I think that's a good time to kind of transition. Let's talk about like ZenML is an amazing analogs tool, but you guys have a new tool out there, Kataru, which is really more around, like it feels like you guys took all the best practices from, you know, building ZenML and MLOps and like now applying it to the new world agents. Like how, like what were some of the big lessons that you took over and like what were, what like didn't apply? What was changing? So this is something we've been thinking about for the last two to three years. And to be honest, ever since like agents and chat GPT or whatever, you know, all that craze for the last few years, we intentionally wanted to not be on the hype train. Like I guess just naturally as, as, as a team, like we were focused on MLOps and those sorts of problems. And we have big enterprise customers that were obviously demanding things, but I think something has changed in the last, last half a year to a year, maybe that made us think that we had an opinion that could count for something that we see the AI agentic stack going towards, which, I mean, again, another rebrand or not orchestration, but durable execution is, is something that I, I feel like we thought internally that, you know, ZenML had a lot of those answers from, from the world of MLOps that we'd solve very well. So things like tracking artifacts or orchestrating things in a infrastructure agnostic way, which was one of the big concepts of ZenML, right? We had the concept of stacks, which is you separate your orchestration and your artifact storage and your container registry, experiment tracking, things like that from the business logic of the core. So your data scientists back in the day and now AI engineers don't need to essentially take the infrastructure level decisions, like, like have a coupling effect in the enterprise with the business logic of the agent. So those things we sort of carried over and we thought we solved them well and at scale. So that led to the creation of like Hitaru because we, what we couldn't really convince people, I think just from a purely marketing perspective is to use a tool like ZenML, which is very well known for pipelines, for machine learning pipelines to, you know, craft a durable execution or agent orchestration layer to run dynamic tags, dynamic pipelines that are orchestrated, you know, in your, in your, in your infrastructure. On top of that, we also felt like the, like purely from an integration perspective, the way we wanted to create integrations into harnesses, for example, we would build somewhat differently in a different product than if we were to do it in ZenML. So we didn't want to mishmash and make ZenML like a Frankenstein. So early this year, we decided to just create a new brand and a new product around it. And, and it's been really fun. I mean, it's great to be back after five years creating product again from grounds up. I mean, you know, as products get mature, you get more customers and the demands change, but now it feels like we're back in 2020, 2021, when we were first thinking about these problems. And yeah, it's, it's been, it's been a great start so far. And if I can dig into that a little bit, I think you said a rebrand into durable execution. And if you don't do DAGs, which I feel like people crumbling from like credit backgrounds or even like sometimes event sourcing, right? So I come from data engineering, so we're like airflow all the way. Why does that state management of durable execution become so important to agents? I mean, I have my opinions, but I want to see why you think that that has become a benefit to building agents. So it's only recently been a benefit, if I'm being honest. Like if you take something like cloud code and you ask, think, ask about durable execution or orchestration, people would be like, yeah, why do I need a threaded while loop in my process that's running on my computer to have durable execution? I can just do cloud dash dash resume or codex resume and start from the transcript, right? So it's a little bit like people have thought a lot about agents as coding agents and coding agents are by default, the harnesses that are built around them are local. And they're not things that run on the background, right? They don't run outside of your laptop if you shut your machine down, even if you have token anxiety or whatever, or if you're token maxing. It's sort of hard to do it in your laptop. So now what's happened is that, I mean, if you see the meter curve, like how far can autonomous agents go? Like it used to be like a few hours, then it's like six hours, eight hours. So every model release, we've started to really see value, I guess, an enterprise of more long running autonomous agents that run in the background rather than in your laptop and then also coordinate between each other. And when you do that sort of system architecture, that's materially different from having a single process, single box, having the harness inside your machine, that's editing your file system. There's no market volumes on Kubernetes pods. There's no shared state. There's no complex context engineering and sub-agent delegation that's running on a different pod in Kubernetes, maybe with a different network access and permissions. So those sorts of things are typically where I think you need a new layer of tooling. And I think that's why we can carry a lot of what we have already solved with machine learning, where we did stuff like that, right? When you ran a machine learning job, which did a hyperparameter tuning sweep over thousands of parameters, and those needed to run on disparate resources with different permissions, maybe with different data flowing in, downloading Docker images that are eight gigabits big, so you can't just do that like thousands of times over. Like those sorts of problems, like those nitty-gritty platform problems, I think we've already solved. So why do we need to reinvent the wheel? Yeah, I think that highlights a really big reason of why like MLOps exists in the first place, is like training these models can take days or weeks. And most people were just used to training an XGBoost model that took like five minutes max, usually like less than a couple seconds, you know, depending on their data set. And you don't, a lot of the MLOps harnesses that make things nice, like it don't really matter when you can just retrain the whole model in seconds. But when it takes a long time, it matters. And you're right, like as we have agents that are doing more complex tasks and going off, and you're like, hey, just go build this entire theme, and you just let it run for hours. Yeah, it's going to make sense that we want a lot of that same tooling and reliability built into the system. So I think from the software perspective, when you do CRUD apps, you think a lot about state management, and you're like, okay, great. If this fails, what is my retry mode? The first time I dealt with DAGs was the example of like a checkout, right? If somebody's on a website, and they add things to their cart, and they leave the website, they come back. Is that a continuation of the process or a new process? In DAG orchestration, that's a continuation of the process, and you have to save that state for a retry. Agents do the same thing, right? When you're having these autonomous agents try to do something, there's a failure mode. So where do you restart? Do you restart from the beginning, or do you restart from the last step? And I think that is, in distributed systems, you can be like, oh, we want it to be event-driven, so we care about isolation. I'm like, I think with the way context works, that's not quite possible yet. And so we need to think about, okay, great, what is our outcome? What are our failure modes? What are our retry modes? And I think that's where DAGs become really powerful, because you can add isolation inside the steps, but you can also have retry for failure modes. So I really found a lot of value with that. My first foray into agents was actually background agents, and I liked label DAGs as well as DAGs. Label DAGs just giving the agents a little bit more ability to say, actually, we can skip this step and go to the next step. That's experimental on my end. I don't think you should always have that option. But I do like it as the idea of, okay, great, something goes wrong, what's next? And there's a reason why there were a specific set of tooling running on infrastructure at scale to solve exactly that problem, right? The thing is that you don't want your developers at scale writing defensible code, meaning just writing code to guard against every type of failure mode embedded in your business logic. So you want that abstracted, almost that the developer doesn't need to worry about it. So, for example, in your use case, if somebody goes away from a cart on Amazon and an e-commerce website, then you don't want necessarily in your business logic and your actual backend logic to store the state and do that. If that can happen inside a DAG process that you can define once, and then it'll complete. So agents in a lot of ways are, as you said, they're doing things like having a human in the loop or having another agent in the loop. And they're doing delegation of agents, so subagents. If you have 10 swarms of agents that are running and the parent agent is waiting, things like what happens if one of the subagent dies? Or maybe you don't want to steer the subagent towards another direction. Or maybe the parent dies. What happens? How do you recover context? Is that necessary to wire into your business logic all the time? Probably not. And that's why I think those kind of new tools exist. For sure. Sweet. We're experimenting with a break where I say, hey, guys, if you're enjoying this, go and follow us on Patreon, where we have fun behind the scene clips, where we go on weird rants about different things. We haven't ranted too much this episode, so please get a little bit more pedantic, Kamza, so we can have some good behind-the-scenes clips, please. You have to push me. You have to push me. You're being very kind. Chris is always so good with the hot takes, and he is unavailable this week. Well, sorry. If nothing else, in the comments, tell us what you think a harness is. Yes, that's a good thing. In the comments, tell us what you think a harness is. Because when you watch this video in six months, it'll be different. And please tell us what a meta-harness is. I mean, that's really interesting, right? Anthropic's talking about meta-harnesses now with their managed agent. I mean, it's a managed product, and it's supposed to be a meta-harness. What is a meta-harness? I don't know. What is a meta-harness? Please, please explain that to me in the comments. We all want to know what a meta-harness is. Good luck. Good luck. How expensive the agent is, right? Like, it's a one-to-one. That's why I'm really passionate about it. I'm like, is there a way to limit context size but still have the same amount of information? That feels very data engineering to me. What are some strategies behind caveman speak? Because I think that's ridiculous. But for... I personally love it. I don't know what you're talking about. Okay. I know that that is how you Claude code now. Now, how do you... With agents, right? Agent pipelines, agent harnesses, or platforms like Kitaru. What are some strategies for managing token spend? I'm sure you've run into this all the time, Matt. I'm sure that's your only job right now as a consultant is telling people how to save money on tokens. Hamza, you can jump in. But yeah, I mean, we've actually... At my company, we've conned and we've... A team went out and they created a long document of like, hey, here's some basic things everyone should be doing to try to like conserve the context. And, you know, like it really does add up. And honestly, I think one of the smartest things you can do right now is just picking the right model for the job. A lot of people just, you know, they'll leave it on Opus 4.7 or... Yeah, I was going to say, Opus 4.7 isn't the magic model for every single job? Yeah, I mean... Even text classification? It can do a lot of things. It's just like... But do you need that extra power? Do you need that extra spend for something basic? And the thing I find is that a lot of people are sending Opus 4.7 to do the exact same work that they were doing with Sonnet 4.5. You know, like... It's ridiculous. Absolutely ridiculous. The 4.5 worked for that and it's a lot cheaper. Why don't you just, you know, like... But isn't that exactly the most interesting question right now? Like, I mean, like Maria, you said that you're like coming from this world of context engineering. And I feel like anybody who's ever done software, all they're doing is context engineering. I mean, abstracted away, what are you doing? If you're making a credit application, it's just... How can you take data and use it in your business logic? Yeah, that is what people have... People forget that software engineering is literally just moving bits around. Like, that is at the end of the day what we've been doing for 70 years, 80 years. And I think now that it's not... We don't think of it as bits. We think of it as gigabytes of context that we have to have in order to make something function. And I don't like that. I do think... I do think, actually, Matt brought up a big point, which I think is choosing the right model for the job. One thing I don't see people doing very often is switching models inside their agent workflows. And I think it's because people don't know they can. Like, we can say, okay, great. In this step, use this model. And then in the next step, use the other model. And then maybe our final step, we want to use a fancier model or something. Or, I don't know. I think it's people... Use software automation and then pass it through ChatGPT at the end. It's just a lot harder because what's happening is... And this is something we should have touched on when we were talking about harnesses. Is that foundational model labs are just RLing their models on their own harnesses, right? And the harness is consuming more and more of the turns. So when you ask cloud something and it's thinking three minutes, you're not in control of that anymore. Like, the harness knows how to edit a file, read a file, and stores things and makes decisions in a very closed, walled, garden way. So it's a bit like having a Mac and not like Linux, right? So you don't really know what's happening behind the scenes. So it's not like you can just, in the next turn, use the codecs and it will have the same performance. It will probably not work as good. Not with current harnesses, right? Because I do think they have that. And I'm interested, Matt. I know you do a lot of education thing. I have a conspiracy theory that they've started to optimize harnesses to increase turns in order to increase token spend. That's true. That's not even a conspiracy. That's 100%. Like, I'm just like sitting here and I'm like, okay, great. You used to one-shot this and now you have to make 15 calls for that my API-based spend can be higher. That's fantastic. They're so tool-calling hungry nowadays. Like, they just want to call a tool. I'm like, you don't need five tools to read this file. Well, I do think that we have leaned really heavily into if we just have tool calls and let the agents manage their own context, it'll be smarter. Because Anthropic wrote this paper one time that's like, if we just let the agent iterate on it forever, it works. And then you turn around and Matt's like, just tell it what you want to do the first time. And it one-shots it. And I'm like, oh, wait, what? I can do that? I don't have to just let it iterate forever? I can give it 15 tool calls in order for it to manage its own context. It's like, I guess I'm not saying that we should switch models every time. I'm just saying in software, in microservice architecture, sorry, I've just gone on five tangents. But I'm saying I like to think of what comes next with agents. And I think of microservices, right? Where we say, okay, great, this is a domain that does this one thing. And I'm like, will agents end up becoming like that? Like, this is the model. This is the domain. This is the set of information for that task. And that then goes to the next task. I think that's where it will end up because we all just repeat ourselves over and over again. But I don't think we think about the side effects, which is like context, token costs, right? And token, the other thing, token spend is one-to-one at this point scaling-wise with compute, right? The more tokens you pass, the more context you have, the more compute you need in order to be able to get your inference. So then you have a scaling that has to come in. And it's just there's a whole lot there that it's just like, who's thinking about it? I mean, I'm thinking about it. I'm just not building it. I only care about context. Yeah. There's a lot there. And I can't necessarily respond to all of it. But what it brings to mind is, I mean, there's an example where someone had the same problem. They fed it one to like JLM5 and said, hey, do this. And JLM5 just one-shotted it. And then they gave it also to like Cloud Code and Codex. And those, they had to like create a plan and then execute on the plan. And all of these came up with the exact same solution. But like one immediately cut there. One took extra steps. And they came up with the exact same solution. But like it's just, there's definitely a matter of incentives, right? Like these big AI shops are, you know, like they're being paid per token. So they're not incentivized to reduce your token spend. And a lot of the systems that they're building increase how many tokens you use, right? Like you do not need to create a plan for everything. However, obviously creating a plan is really, really useful because we forget to add context. And that's the whole point of the planning phase is, you know, like if we could give the model everything it needed in one go, then it can go off of it. And so, yeah, going back, it's not necessarily a conspiracy theory, but it's just like they're not incentivized to reduce our token spend. Like if anything, they're going to increase it. Like it has to be, isn't that more an argument for open harnesses? Like because it feels, it feels to me that a lot of the things that you guys are speaking about is more about we don't have enough visibility over the reasoning process. And why certain agents make decisions? And we continue to outsource that to bigger labs. And I think that there's a set of problems where that's totally okay. So, I mean, maybe they solve a certain type of task management or coding or something very well. But do you want to just live in a world where all those decisions are made for you without you knowing? And then you're eventually, at the end of the day, these are real humans running one company, right? So you can't have one company solve every type of use case in an efficient way. And you're going to obviously end up using a Canon when you might as well have used a knife. And it's just very, like for me, that sounds to me that we're headed, if I was to put on my prediction at, that it feels we need, we've overly abstracted harnesses. And we need to sort of reveal what's happening with context behind the scenes. So I don't think we know enough about why a certain thing, like how many tool calls were made, what was the most inefficient one, what was expensive, what would have happened if I'd loaded that skill? Or how would you even evaluate a skill file? Or like maybe it would have been different if something else had happened. And this whole process is for me what I do when I develop a software program, like for the last 15 years, right? So, and that sounds to me like we're going to have an open base, which is neutral from the model perspective. I mean, despite what the bigger labs want us to do. So that's why open clause is fine. Open code is great. Pi is great. Like these things matter, in my opinion, materially, because we need to push against this. Plot code is the only way to do it. And you can't put like open AI inside it, it won't work as good because it's not reinforcement learning on how to edit a particular file. Because God only knows how to edit a file now. That's ridiculous. Come on, guys. So I think we're headed towards more open harnesses. And then the whole stack is going to, you know, obviously we're going to need some sort of swarm orchestration layer when you have multiple agents at scale. So, yeah, it's quite interesting how it develops in the next few months. In building Kicharu, or I guess, Matt, you've also done a lot with like Pydantic and stuff. What are your strategies to kind of start managing cost and context? And I don't know if like I want to say give visibility, but I do want to say like monitoring that. I think observability into that cost and context management is super important to the development and optimization processes later on. We've thought a lot about this. Like I think that's one of the foundational goals of any orchestration system is to be more efficient. And like cost is one side of it. But as I was saying, there's other failure modes that you want to catch right on a system level. So, so things like, like maybe one particular tool call is having a certain type of error or maybe like the agent keeps calling 10 times the same tool. And, you know, it could have just called it one time if you change the parameters. That is what happens in the agents I have built myself all the time. I'm like, if you failed once, my prompt is bad because I was literally saying read file. Exactly. I mean, and there's, it's so underexplored. I mean, imagine how little we know. I mean, when I'm developing agents, this is not necessarily at the forefront of my mind. The forefront of my mind is how do I get this thing actually solving the problem as quickly as possible? Because I'm probably under deadline pressures. Efficiency. Efficiency and you don't want to, you know, you don't want to prematurely optimize. But I think eventually at scale, you have to think about those things. And that's why the first thing that we are thinking about when we built Kitaru is that how do we surface those insights better and put it in front of the developer in a way that doesn't seem intrusive from the very beginning of agent development. So our answer to that is to integrate very deeply into the harnesses that are popular. So if you want to bring cloud code or if you want to bring bad NDK, fantastic harnesses. But you want to exfiltrate them and you want to understand what's happening behind the scenes. So if you plug in Kitaru, you get way more observability around orchestration and running different scenarios when things could have been different. Surfacing costs, expensive tool calls, expensive turns. Surfacing costs, expensive turns. Those sorts of things are very top of mind when I talk to my customers, early customers of Kitaru. Totally. I think, like I said, I like the idea is in software engineering, you do the make it work, make it right and then make it fast. So, like the first time when you make it work, yeah, you're not going to be doing cost optimization. But, like, I think MLOps was created because data science was too expensive. So we're like, okay, great. We've made it work. Now let's make it right and then we'll make it fast, right? We'll optimize that pipeline. Same thing with data pipelines. That's SRE's whole job is just cost savings anymore. I think I will wait for somebody to start having an AIOps job and I'll be like, guys, that's the same thing. It's just do you have observability and can you use that observability to make optimizations, right? Like, tell me that that's not been your job for the past 14 years, Matt. It's just optimizing things. Oh, yeah. I mean, that's the job. And I guarantee you, yeah, the job title is going to change. Like, that's the thing I've learned going back to one of the things Hamza said earlier. Like, these names, you know, first we were calling it agent. Now it's a harness. Now the harness creates agents. And, like, we're just always coming up. Like, the data field is addicted to, like, new titles and new names to describe the exact same things. How much of that is, do you guys are driven by this Bay Area VC? Oh, it's all marketing driven. I don't think it's just, like, Bay Area VC, but I do think it is marketing, right? Like, if we make a new name, we can make a new brand and then we can make a SaaS company around it. And yeah, it's a VC thing, right? Like, but yeah, it's money and we are cool and our stock price will go up because we adopted this new thing. When it's all at the same day, like, the same fundamental practice of do you have metrics? Are you monitoring? Can you optimize? Is this a platform versus a one-off, right? I think platform engineering was the thing. It's like, oh, now we have people building platforms. It's like, no, that's what we've been doing for 20 years. We've always been building platforms, right? We just called them, like, I don't know, operations. Like, so. But, like, it's so weird because we live in a world where words mean a lot more than even they used to before because we're literally making decisions based on output tokens, which are words, right? So, like, I remember when I was, like, a year ago when I was learning about durable execution. Like, I felt like such an idiot because I'm, like, I've been doing orchestration for 10 years and I don't know what durable execution really means. And I argued with an LLM for about one hour trying to understand what is the difference between having a DAG, which versions artifacts and does caching, like XenML or Prefect or Airflow, versus something like Temporal or DbOS, which is doing state checkpointing and has worker queues and all those things. What is the material difference? And it kept saying, okay, one is machine learning orchestration and the other is durable execution. And I said, no, can we go beyond that and try to understand the actual semantic? Like, it feels like these two things just developed independently, but they are the same thing, right? Temporal tries to say they're a DAG, though. Like, if you go to their docs, they're like, we're a DAG. And I'm like, I can't use you as a DAG, though. Like, you can't make the choice. It's just, it's hard, right? I mean, like, we're in this probably day-to-day, both of you as well. But imagine somebody just arriving at the seed. They're like, they wouldn't know what to do. They wouldn't know what to search for. So that's what makes it hard. Well, and like I said, you think of agents as agent harnesses that kind of just like, well, if you let it chug on the problem long enough, it'll figure it out. And like, well, that's not the point. The point is we've been building, not necessarily pipelines, but I think execution flows for years. And that's the same thing an agent is. It's just now we're letting it do things. We're letting it be more flexible with the kinds of things it can do, I guess. I don't even know if that's the right word. It's just a new flavor of automation. I really enjoy that aspect of the fact that, yeah, like we often create words that mean the exact same thing. But then we come up with very strange, minute semantic reasons of why these two things are different. But we also often do the opposite where two things are created independently and we use different words because someone used it in field A and someone used it in field B. And then you realize, oh, they're the exact same. And I mean, this has been happening forever. Like, you know, even in like subdomains where like statisticians will use one word to describe something and mathematicians will use a different word and they mean the exact same thing. And then oftentimes you get mixing between all these different fields and industries. And then in the data world, we get stuck with all the baggage, right? Like all of these different words. We're 20 years behind everybody until now. Now AI is first. So we get to invent the words and we don't know how to do it. Well, we... Everything's an agent. Everything's an agent. That's easy. I think you can explain it. Everything isn't. That's my point, though. I'm just saying, just like how everything was AI, I'm like, don't build a new agent. Just say the thing you were already doing is the agent. And then you can just sell it as an agent. Why start over, guys? Do you want to do it? Do you want to tell people where they can find Kitaru? Yeah. So we're like completely open source. It's sort of funded by ZenML in a way. So that's making money. It has a commercial product. Kitaru is open source free where we're really just looking for if you want to contribute or if you want to give us a spin. We have great integrations to existing harnesses like PytantiKi or Anthropic, Cloud Agents SDK. So if you want to give that a whirl, leave an issue, give us a start on GitHub. That would be the best. Support open source. And yeah. And I'm going to be making a Go harness for it eventually because Hamza challenged me to do that. Make a Go harness. Make a Rust. Make a Rust CLI. We're welcoming everything which isn't super vibe coded. I mean, you can leave us a spec. Maybe we can work on that. Super vibe coded. I don't write code. I do make sure it works though. I own my code. I own all of my code. There you go. As long as that's true, we're going to approve that. Nice. Thank you. Check it out. Check out Kitaru. Amazing project for sure. We're at time. And we stayed on track, guys. What on earth are we doing providing high quality content? Today we talked about agents, durability, context management, architecture, the fact that everybody has a different word for the same thing. We want to thank Hamza for coming. It was great. I really enjoyed having you. I'll have to check out Kitaru. I should have asked if anybody. I know Chris has tried the Hermes agent. I was wondering if either of you have messed with it yet to see if it was cool. But that can be. We'll do that in a minute. Let's not. Let's do this. So, like I said, we have podcasts on all major podcast providers. You can follow and get new episodes every two weeks. You can do Patreon for behind the scenes. We want you to tell us in your comment what you think a meta agent is because we're trying to figure that out. And then we can have a podcast on it if you tell us. And the most important thing is our outro. So, are you ready? I prepped you for this, Hamza. What was it again? I forgot. Keep your AI on a leash. Okay. Keep your AI on a leash. Keep your AI on a leash. Yeah. Like we domesticate AI. So, it's like our pet. All right. Okay. Ready? One. Two. Three. Keep your AI on a leash. See? Look, we're so good. I don't know if that worked. It was close enough. It was close enough.