← Back to search

Building AI Agents That Support Enterprise Teams

The Enterprise AI Show · 2026-09-16 · 39 min
relevance 42 7815 words Episode page ↗ Audio ↗
Show full episode description
Aaron interviews Diego Oppenheimer , general partner at the AIT Fund (the venture fund behind the AI Tinkerers builder community), about what it takes to move AI agents from personal experimentation into enterprise-grade deployment. Drawing on his own agent fleet (a chief-of-staff agent, a content strategist, and Ashley, an always-on AI community manager serving AI Tinkerers' 127,000+ members) as a proving ground, Oppenheimer lays out an enterprise security model built on treating agents as digital employees with their own delegated identities and credentials, scoped access rather than blanket access to accounts like email and calendar, and continuous audit logging (via the open-source tool Hyperware) in place of trust based on prior behavior. He argues enterprises get burned by chasing a single omnipresent agent instead of narrowing agents to specialized, restricted workflows, the same specialization principle that has driven results throughout machine learning, and flags scheduling as a deceptively hard example where the "10% edge cases" eat most of the effort. He also describes deliberately red-teaming his own agents for weeks (trying to break out of containers, extract credentials, and social-engineer them) before granting them any real access, and points to NanoClaw's containerized, minimal-component design as a model worth enterprise attention. The conversation closes on identity, responsibility, and permissioning as the unresolved internals enterprises must solve before scaling agent autonomy, and on Oppenheimer's prediction that many enterprise roles will shift toward "exception handling" as teams of AI coworkers absorb routine work and escalate only what needs human judgment. SHOW: 1063 SHOW TRANSCRIPT: The Enterprise AI Show #1063 Transcript SHOW VIDEO: https://youtu.be/vL8btTMcInE SHOW SPONSORS NordLayer - Use ENTERPRISE10 for 10% off Nasuni - Activate your data for AI and request a demo SHOW LINKS Diego Oppenheimer homepage Diego on LinkedIn AI Tinkerers Guardrails AI NanoClaw on GitHub The New Stack, "OpenClaw vs. Hermes Agent: The race to build AI assistants that never forget" The New Stack, "NanoClaw and Docker team up to isolate AI agents inside MicroVM sandboxes" GUEST BIO Diego Oppenheimer is working full-time at AIT Fund (the fund behind AI Tinkerers) and previously was a partner at Factory and served as CEO-in-residence at Factory. He founded Algorithmia, an enterprise MLOps platform acquired by DataRobot, co-founded Guardrails AI, and earlier in his career led teams at Microsoft shipping Excel, SQL Server, and Power BI. He is currently running AI agents inside his own team to take on real, sustained work, including one named Ashley, and documenting what he's learned in a new video series with Joe Heitzeberg. FEEDBACK? Email: show @ the enterprise ai show dot com Bluesky: @TheEntAIShow.bsky.social Twitter/X: @TheEntAIShow Instagram: @TheEntAIShow
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How enterprises can turn agent harnesses like OpenClaw, Hermes, and Nanoclaw into reliable AI workers instead of risky experiments.
Benefits
  • Insider comparison of OpenClaw, Hermes, and Nanoclaw design philosophies
  • Blueprint for replicating a $200-250k chief of staff with agents
  • Model for a tiny team operating a massive global community with AI
  • Framework: automate operations, keep human interaction as the core work
Use cases
  • Diego Oppenheimer built an AI chief of staff targeting $200-250k/year human-equivalent quality
  • Built a content strategist agent that queries his thoughts and flags outdated takes with evidence
  • 'Ashley' AI worker supports AI Tinkerers: 127,000 members across 250+ cities run by ~3 people
  • AI screens event applications at scale where 1 minute per applicant is impossible
KPIs / results
  • AI Tinkerers: 250+ cities, 127,000+ members, an event roughly every 6 hours
  • Community operated by essentially 3 people, one full-time
  • Chief of staff benchmark: $200,000-$250,000/year human equivalent
Tools / build
  • OpenClaw
  • Hermes
  • Nanoclaw
  • Ashley (community AI worker)
  • Howie scheduling agents
0:00 / 0:00
[SPEAKER_02] Good morning, good evening, OVR, and welcome back to The Enterprise AI Show. This is your host, Aaron, and today's interview is the second in a two-part series on agents in the enterprise. Today, we're talking to Diego Oppenheimer. He breaks down the agent harness landscape, including OpenClaw, Hermes, and Nanoclaw, and what he's learned building agents into a genuine force multiplier for enterprise teams. With that, let's jump right into the interview right after this quick break. [SPEAKER_02] What does the outside world already know about your company? Cybercriminals could be seeing leaked credentials, compromised session cookies, exposed infrastructure, and even impersonations of your brand and executives. NordLayer Intelligence by NordStellar gives security teams visibility into those threats across the deep and dark web. Data breaches, attack surfaces, and brand impersonation all in one platform. [SPEAKER_02] Find out what attackers know before they can use it. Visit nordlayer.com slash intelligence slash enterprise AI and use code ENTERPRISE10 for 10% off your NordStellar plan. [SPEAKER_01] Today's show is sponsored by Nasuni. There's a growing gap in AI right now between what's possible in theory and what successfully works at scale inside an enterprise. The difference comes down to unstructured file data. Many AI initiatives struggle because the file data they depend on is scattered, unstructured, and disconnected from where and how work actually happens. Nasuni changes that. [SPEAKER_01] It brings your unstructured file data into a single, secure foundation so AI, both generative and agentic, can access it with the context, governance, and performance it needs in production. Bring AI to where your unstructured data lives. See what it takes to activate your data for AI and request a demo at nasuni.com slash AI. [SPEAKER_02] And we're back and for our topic today. Actually, this is kind of a, I'm going to use this as a part two, if you will, of our previous topic on NanoCo and NanoClaw and working with agents in the enterprise. And for that, we're going to be talking about building AI agents for enterprise teams and enterprise use cases and have a gentleman today. [SPEAKER_02] I hope I can call him a gentleman that ran into at a conference here recently, Diego Oppenheimer. And Diego works full time at the AIT fund. And I'm going to have him explain a little bit more about what that means here in a second. But like I said, this is kind of a part two, if you will. The first one was what is agents and especially some of these agent harnesses and going into that in depth. [SPEAKER_02] But now we want to talk about a little bit more today about actual use cases and what it's like, you know, and somebody that's actually doing pretty cutting edge stuff with that. And so, Diego, first of all, welcome to the show. And first of all, give everyone a quick introduction and a little bit about your background, please. [SPEAKER_00] Sure. So Diego Oppenheimer, currently a general partner and founding partner at the AIT fund, which is a inception stage venture fund that invests in kind of like AI forward companies, backing technical founders when they just start their companies. The genesis of the fund is we help run the largest AI builder community in the world called AI Tinkerers. There's probably one in your city, wherever you are. [SPEAKER_00] We're in over 250 cities, over 127,000 people worldwide. There's usually an event about every six hours somewhere in the world. And the idea behind that was being able to build a place where people could share messy experimentation of what they were building with AI. So that's kind of like the genesis of it. Previous to that, I was the co-founder of two different companies that have exited in the AI space. I built products at Microsoft. I helped teach applied AI at Stanford. [SPEAKER_00] It's been working in the AI space for the last 15 years, essentially. [SPEAKER_02] Very nice. Very nice. Yes, absolutely. So let's kind of start at the harness level because we were, you know, like I said, we wanted to talk about that. And I think that'll kind of set up everything that we're doing here later. I'm going to refer to it almost like an arms race right now. Like Open Claw came onto the scene and Hermes, however we want to say it. And then Nanoclaw, you know, comes up in conversations as well more and more here recently. [SPEAKER_02] I think each of them have a bit of a different design philosophy and certainly different strengths and weaknesses. But tell me a little bit about, first of all, like what do you think about or how do you evaluate these? And then maybe we'll go into pros and cons of each other. [SPEAKER_00] Yeah, so I mean, I started out, so I use them all. Some have had more success than others. Mostly the style of agent that I've been building. My goal was to experiment with as many of these as possible and build with them and understand it. I think the, you know, if you go look at Open Claw was probably the first one that became popular. It evolved very much like, and it's changed quite a bit and it's good right now. [SPEAKER_00] But like at the time, it's popularity and kind of like, let's just call it like it's breath, right? It had connectors to everything. But it was also like had a lot of crud in it, so it became very hard to pursue stability with it. And that's changed now. But again, I was experimenting with these before they came out to the scene. [SPEAKER_00] Actually, funny enough that Peter, Pete, you know, one of the first meetups where he talked about Open Claw ever was at NAI Tinkers in London before it had its meteor to rise. I think each of them has taken a little bit of a different approach. And it's interesting because if you were on the internet in this last week, there's kind of like a next wave, right? Because the three that you mentioned are the, hey, here's the tooling, you know, bring the brain, here's the exoskeleton, put them together, make them do something. [SPEAKER_00] And over this last week, you know, we've now seen the, oh, it's all in a box for you. You don't even have to worry about that, right? But then like the launches of like Instinct and Muse from Meta and like, you know, so we're kind of seeing that's more on the personal side. It's very consumery. But if you kind of think of it from evolution, right, we had the kind of like we have a model. We're not really sure how to get it to do. Oh, hey, we're going to kind of give it a personality. It's going to work a little bit on its own. [SPEAKER_00] And we went to the harnesses, which are like, here's ways of activating, you know, which again, for those who have you like dug under the scenes, like the initial harness or what it was like. It's a model with data access and a bunch of cron jobs initially, like they're more complicated now, but like that's what it was, right? It's like now it's much more proactive. So anyway, so I think what we're seeing live is the evolution of the space, right? And at mock speed. [SPEAKER_02] Yeah. Yeah. And so let's do this. I want to explore like for those that are out there listening, like practical use cases and ways to get started. And I think it's better to start with a little bit of like what you've seen and what do you do? Like when you say you've used them all, like tell everyone a little bit about some of the use cases you do and like some of the personalities you have your agents running. Because like you and I were in a session together and you were doing some stuff that everyone else was like, really? Yeah. [SPEAKER_02] So like tell everyone a little bit about how far you're pushing this when it comes to agents and agents, you know, doing isolated jobs or jobs on their own. [SPEAKER_00] Yeah. So, you know, one of the things to think through is like, you know, as kind of the different companies and then the exits have happened, like, you know, at some point I had 400 people reporting to me, chief of staff, EA, kind of like the whole shebang. And then one thing to know is that I'm like completely calendar dyslexic. Like I cannot get like things on a calendar in the proper way without messing it up. And so a big part of my initial kind of like go into is like, well, you know, right now I run a small fund. [SPEAKER_00] I kind of am on my own. Like it doesn't really justify having a human helping out with these things, but I absolutely needed the help. So I went obviously in the chase of AIs, right, to help me. And I was an early investor in a company called Howie, which builds kind of like scheduling agents that, you know, figure things out there. And I also wanted to kind of try to replicate the chief of staff, right? [SPEAKER_00] And so one of my ethos here was when I'm going to go when I'm going to go chase these things, I'm looking to say if I can build a human work like worker, right. And that has like a lot of value, at least a lot of value to me. Right. And so if I think about a chief of staff to me, you know, chief of staff for me in the real world would be somebody that'd be earning 200, $250,000 a year. [SPEAKER_00] And so that's the quality level that I'm trying to achieve in terms of like ability to do work and be preemptive and stuff like that. So that chief of staff, which is like kind of like the use case that everybody starts with because everybody's like, who wouldn't want a chief of staff? By the way, they're great. Fantastic. If you ever had a great chief of staff, they changed your life. Totally agree. Right. And so there's the chase of that. But then it kind of went down those lines. It's like, well, it's just not that. Maybe it's also the, you know, you see customer support. I don't do customer support, but you see customer support come up showing up in this. For me, it was also around content management. [SPEAKER_00] I write a lot and I've considered myself a pretty bad writer in general. You know, it's English is a second language. You know, there's a bunch of reasons, you know, like technical stuff. Yes, I've written like specs, you know, but things that people want to enjoy reading. And I find writing to be how I kind of get my thoughts out. Right. I can explore a space. I can understand and I can get a little bit more technical about it. And so one of the things that I wanted to do is also build a content strategist. [SPEAKER_00] And so, you know, if you were a CEO of a Fortune 500 company and you wanted to, you probably have somebody who's helping with brand and somebody who's helping you with PR and somebody who's actually helping you with content, maybe even a ghostwriter. So I didn't want a ghostwriter because I didn't want the AI to write for me. But what I wanted is somebody to help me query my thoughts and take care of my thoughts and understand like the topics I'm interested in and then help me evolve different think pieces. And then also go back and be like, new information came in and be like, hey, remember when you said this thing? Well, that was totally wrong. [SPEAKER_00] And here's evidence of it. You want to write about it again? And so I went to build that as well. Right. So there's like kind of constantly on thought partner around that. Then for the community to be able to operate, you know, so there's three people who operate this community. There's volunteers in every city. But, you know, essentially three people, you know, and only essentially one of them who's full time operate this mega large community worldwide. Right. [SPEAKER_00] And so the only way to be able to do that is we also needed AI to help you answer questions. And so that's the idea of, you know, you mentioned before of Ashley came out as this kind of like global AI worker who's always on and always present and always trying to be helpful across the entire community. Gets to know all 127,000 of our members at a personal level, which we would never be able to do. Like just, you know, you know, to give you an idea of the scale, we were doing the math. [SPEAKER_00] So we have to approve any person to get into any of these events because we want builders only. Right. This is meant for builders who are deep into it. It's not about selling and pitching. It's about showing what you're building, showing what you want to do. [SPEAKER_00] So if you spent, you know, you know, you, you know, you had, you're going to have a hundred people show up to, you know, a meetup, you know, and you spent one minute looking at each person's application or, you know, that's impossible. Now try to do that at 127,000 people. Like it's just not possible. Right. And so AI has to help us. And so, so that's another place where we've, we've done it. [SPEAKER_00] So in a lot of cases, our ethos has been, what are some of the operational work that AI can do really, really well, allow us to continue doing the parts that we think are core to the business for us, which is interacting with humans and, and automate everything else. So operationally automate everything else. [SPEAKER_02] Yes. Now let me ask you this though, because I think one of the big things, especially when you see like, you know, like open and open AI, you know, models getting out and attacking, hugging face. And like, there's almost this like, oh my gosh, the, you know, the, the, the, the evil robots are taking over the world. Skynet thing all over again. [SPEAKER_02] But, but I'm going to add it this way of like, okay, everyone seems to want agents because they will do stuff, but no one, everyone is afraid the agents are going to go off the rail. And so like, how do you balance that of like, okay, especially if it's like a chief of staff role, like let's, let's use that as a good example. A chief of staff role is going to have access to your entire life. And so, or if it's good, you know, if it's good at its job. Right. [SPEAKER_02] And so like, how do you start to think about, okay, I want it to have autonomy, but also how do I keep it from going off the rails? Let's like talk about that. And I want, I kind of want to get into memory and long-term tasks and all the other things after that. But like, let's start fundamentally at like, how do you do that? Yeah. [SPEAKER_00] So, you know, this is actually something that I spent a lot of time thinking about because I have this like weird duality, which is one where I don't believe, I believe in privacy. I just don't believe it exists anymore. Fair. Right. [SPEAKER_02] I tend to agree with you actually, but go ahead. Go ahead. [SPEAKER_00] And, you know, I've given, you know, my life away to Google already multiple times over. And I am pretty, if I can get the results that I want, if I get like the outcome and I would get the, you know, the profit from my, my personal profit, I'm more than happy. You know, I'll share my medical, you know, I'll share my medical history. I'll share my private stuff. Like I don't have a problem, but I expect there to be ROI on that personal ROI. Right. I always joke around and it's like, I've already given you my, all my data. [SPEAKER_00] If you recommend me another grill, even though I already have that same grill, like that's on you, not me. Like that's a problem from, for your system. And I'm not happy about it because I, you should know that I want these other seven things and be recommending it perfectly to me because I've already given you all my data. At the same time, I was also very nervous and I still am. And I kind of, I look at it, which is there's this really interesting combo where it's like your email plus access to your SIM card are essentially like how you can. [SPEAKER_00] That's the, that's really the only two things you need. Like if I have your email and your SIM card, like you're done. Right. Like that's like, you know, you lose everything. And this isn't like kind of the more personal. And so that was a, got me thinking about how to think about security and what I wanted with these things, which is like, well, I already hire employees. Right. And what do I do when I hire employees? I give them emails. I give them GitHub accounts. I give them their own identities and then you can like hire fire them inside that ecosystem. [SPEAKER_00] And so that was been my approach since day one on building AI agents. I didn't like add an agent to my stuff. I said, you are your own thing and I will delegate things to you. Right. And so the same way I did with an EA or same thing I did with a chief of staff, I can say, Hey, you have delicate access to my email. You have delicate access to my calendar, but it's all as this kind of like unique persona around that. [SPEAKER_00] And so that kind of access we already did today, you know, obviously you trusted, you know, you had to trust this human you were hiring. Right. So, well, why did you do that? Well, you look at the things that they have control. You put it, you know, I can't go ask for a resume from the model and say, you know, like, was there an honest, you know, were they an honest model before? And like, you know, would you trust them again? [SPEAKER_00] But what I can do is understand like what scopes they're going to have access to and audit that regularly and understand what they're touching, what they're not touching and anything that's touching that shouldn't be touched, get alerts for it. And so I thought about it a lot more like that. How do you inject credentials into the system? How not to store credentials in terms of systems, you know, their computer systems at the end of the day. And like my background's all in fairly deep enterprise systems. So like, I think that way. [SPEAKER_00] And so that's that was the the genesis of how to build out these employees, these digital employees who had just the same shape as a human employee, except they work 24 seven. And they were smarter and dumber at the same time than any of my previous employees. And yes. And, you know, kind of like you kind of dealt with that. [SPEAKER_00] But so a lot of it is like you do this kind of like inline auditing and guard railing and building out kind of things around that versus the kind of like historical trust based on, you know, what they did before. Yeah. That's kind of like the ethos. [SPEAKER_02] But I so I think that that's a key thing right there. What you said, this whole idea of treated like an employee and basically instead of providing it access, provided delegate access is a key thing. I think a lot of folks aren't talking about and they're not necessarily thinking about it that way. They're always thinking about it of like, like, I'm going to I don't know, you know, like Claude today and I'm going to give it, you know, my Gmail and and calendar and all these other things. [SPEAKER_02] And it's going to do things as me as opposed to an employee of me. So I think that that's a super important aspect of it, especially for folks out there that are like looking to do more of this. Go do it that way. And also the credentials is super important. But let me I'm going to take it one step further. But I do want to talk about failure modes because I think failure modes is such a good topic here. But before we get to that, let's take it one step further. Like, OK, that's how they get access. [SPEAKER_02] How do we do memory and self-improvement and long term tasks and everything else over time? Like that that gets you day one or day zero access. How do you keep this thing on the rails over time? [SPEAKER_00] Well, so so you you constantly audit. So I have I've built I've adopted some logging systems that I use across my all my agents. So I use this open source. Actually, I didn't even realize I got their T-shirt on. I use this open source tooling called Hyperware. And the what it has is it like it ships all the logs to a central location of what the agents do at any time. [SPEAKER_00] And then I have constant reviews over that. Now, this is only like allows you to kind of like, oops, something happened. Let me go figure it out. And you can do some proactive stuff like, hey, let's go look at things and make sure. Honestly, it's a lot less nefarious and more like, why did you spend three days trying to recreate that CLI and spend so much tokens? Like it's much more like along those lines. So you see a lot of efficiency gains from looking at your logs more than catching something completely nefarious. The other one is like how you set yourself up. Right. [SPEAKER_00] So I think in line, I was really particular about just because we talked about it earlier and it sounds like he was on your show. Like I really like NanoClaw because of how and I'm not an investor or anything like that. I just use their product. They took this like really smart approach to minimalizing components and containerizing everything. And so, you know, I probably spent and this again is a little bit more my nature than anything else. [SPEAKER_00] Like before I made my agent do anything, I spent probably a couple of weeks just trying to break in, make it do bad stuff without like anything useful at all. Right. And so I was trying to do like, you know, where are you storing your credentials? How do I do credential injection? How do I break out of the containers? How do I kind of get privilege access? I'm sending other agents to try to break it. I was putting my agent in a chat room with a bunch of my like friends who I knew would try to break that agent and try to get into it. [SPEAKER_00] And so like, you know, I spend a lot of time kind of hardening the real time components of these agents that I use. But that was a lot of my curiosity, to be honest. Like, I think the space is so new and exciting that, you know, you have to do these kind of explorations to understand. Like anybody who's coming in and saying, this is the standard, this is how we're doing things. [SPEAKER_00] Like this is, you know, this one is just like inaccurate at this stage because it's rotating at like such high speed. That doesn't mean it shouldn't be adopted. Right. You know, like hopefully your listeners are not listening and being like, oh, technology too new. Like we should like, you know, look the other way until it's more mature. Like because you can't do that in this space at all. But that's kind of like, you know, we're still in the very let's figure things out. Let's see where things are going. Let's, you know, around that. [SPEAKER_02] Yeah. Yeah. And I like that idea of hardening them before you even go any further. Because yeah, here's the thing, like I've been, I've been playing around with those and I've even been something as simple as like, hey, do I want to do it in the cloud or do I want to do it in a local box, you know, behind a VPN kind of thing under my desk? You know, because again, enterprise background, I'm trying to approach this from most secure way possible and, and, you know, self-serving for this podcast. [SPEAKER_02] Like, hey, there's lots of things I can automate with this podcast, but I haven't gone to that level yet because I'm doing exactly the same thing you are of like, okay, what's it going to look like? How is it going to act? And not just that, because it's the non-deterministic nature. How is it going to act the majority of the time versus the minority of the time? And also if it's going to fail, how spectacular is it going to fail? [SPEAKER_02] And I think that's maybe like something we should talk about because you, you've kind of made a hobby out of, you know, agents failing. And so tell everyone a little bit about that as well, because I think it's super interesting. And I think it's, you know, for everyone considering doing these, I think considering how bad it can fail is worth looking into. So tell everyone a little bit about that. [SPEAKER_00] Well, if you'll allow me, you know, to plug a little bit of a show on another show, it's like really short. Please. So we have this series called Oh, Ashley. And so Ashley was the main AI employee behind AI Tinkerers created by Joe and both the show and Ashley herself. And, you know, we, we started, you know, because we're like ultimate tinkerers and we're absolute dog fooders of our own stuff. Like we went all in, you know, we're like, we're like, so AI pilled. [SPEAKER_00] Like we, we like, there's the, let me figure everything out and slowly adopt. Or there's like, let me adopt all at once. No stuff is going to break. It's going to be slightly embarrassing, but like, I'd rather fix it at the edge and learn than like not be all in. And so we're both in that camp for sure. And, but over the last, I want to say a couple of months, the failures have been hilarious and kind of entertaining. Some serious, some not. [SPEAKER_00] We have a friend in common, Brian and, and, and, and, and, you know, he was, there's a thread and we were planning a dinner in Paris. And when we say, Hey, you know, we have a bunch of AI people coming to this dinner in Paris. You should come. I think it'd be really interesting. He's like, I would love to, blah, blah, blah. But I can't, cause I'm leaving the day before. Can't remember exactly the thing, but I'd love to get coffee sometime. So Brian was telling me they'd love to get coffee. We both live in Seattle. And I was like, yeah, perfect. [SPEAKER_00] Like, let's go do that before I can respond. Ashley being very proactive, jumped in and said, Diego would love to get coffee with you, but not bright, not representing Brian, representing herself. Right. Being an AI. And I was like, that's funny. And so then I responded by saying, cause I knew like this was an AI problem. And I was like, actually, I don't want to do coffee anymore. Now I want to go do dinner at this really expensive restaurant in Seattle. And by the way, Joe's going to pay. [SPEAKER_00] And it was actually in all fairness, like he had these like inline checks, right. That were like looking and be like, Hey, this might be fishing. This might be somebody trying to take advantage of the system. And it actually, after it tried creating a in real life coffee between an AI and myself, as soon as I went in to try to like abuse the system, it was able to qualify that abuse. As potential abuse and it shut it down. And so I did not get my expensive dinner. But anyway, to all this point is we have this series called Oh Ashley that we make available. [SPEAKER_00] It's like 10 minutes, 10 minute little micro podcast where we interview people who have had agent failures and they talk about their agent failures. And they, uh, how they worked and what they got to do and like, you know, what they're looking at to do better. So it's been pretty fun. So we think everybody talks about everything that works all the time. And, you know, if you go online, like agents are doing everything and like fixing everything for everyone and the world's automated. And a lot of that's true. And that's like the dream. [SPEAKER_00] But, um, the agent failures are just as hilarious. And the more we talk about them, the more we can solve them. So we've just figured that we put a different view on the here is to absolutely AI pilled people who run their businesses purely on with AI agents. Showing you that we just got as many failures, if not more and how we think about them. [SPEAKER_02] Yeah. Yeah. And I love that because I think for a lot of our listeners that are out there, the idea of running agents in these agent harnesses, it tends to be the side hobby weekend thing. Maybe even, you know, a second career kind of thing at best. [SPEAKER_02] I don't know that many folks that are out there doing that as part of their day job today. But I was actually going to ask you that, you know, as somebody who speaks to a lot of folks and talks to a lot of folks that are running these things and are on the cutting edge. Where do you think the state of all of this is today? There is obviously there's your side of it where you're all in. Then there's other folks that are, you know, hobbyist and failing. [SPEAKER_02] But like tell everyone like a legitimate current state, you know, based off of where the technology is today. Where do you think this fits and where should it not fit? [SPEAKER_00] I think agents doing. Restricted workflows or dedicated tasks. Right. And those dedicated doesn't, I'm not talking about just like things that you would automate with one cron, but actually kind of like maybe like dedicated to, you know, hey, onboarding or in my case, kind of like content management or scheduling. Scheduling is actually extremely hard. I think people underestimate. [SPEAKER_00] Everybody wants to start with like, oh, I'm just going to have an AI that schedules for me. Scheduling has, you know, it's an operations research problem. It's wildly hard to do. And every single agent gets it wrong. Again, I'll plug for my folks over at Howie that has actually like solved a lot of this. But if you're going to build from scratch, don't start with scheduling. Like you're just going to be very frustrated. Like the 90% case is super easy. And then like the 10% case is like virtually impossible. [SPEAKER_00] And so you spend your entire life in corner cases. Anyway, that's a different. But like my point is like the adoption here is like find these small workflows that you can get really well. Because you can use memory. You can use context engineering. You can use restrictions of tooling. And the reality is that you can get there. Like today, the, you know, hands down, you know, it's not even state of the art. I think like we can adopt agents to do certain workflows. [SPEAKER_00] And now, of course, in enterprise, which is your audience, like you have to figure out identity. You have to figure out responsibility. You have to figure out who, you know, who's doing what and when and why do they do it? And when do they have permission to do it? So there's a whole bunch of kind of like internals that need to be figured out before you can let loose and kind of say do it. But I think the major failure point that I see, and particularly in enterprise, is this like omnipresent agent that does everything, right? [SPEAKER_00] Like still today, even though the models are, you know, have more breadth than humans, they still are not good enough at like kind of narrowing themselves down to certain tasks and getting those things right. Right. And so it's that balancing act where like success to me is, and by the way, this has been true in machine learning since the dawn of time, which is, you know, the more specialist you can get something, the higher results you get today. And that will change over time and you can expand from there. [SPEAKER_00] But like specialization has been the way that I've gotten things truly working and in production. [SPEAKER_02] Yeah. Yeah. That makes perfect sense. And let me ask you this then too, like for those that are out there today, what is the best way to get started with something like this? Like, because obviously there's picking a harness, but then each harness has, you know, different strengths and weaknesses. And I'm not necessarily looking for that. [SPEAKER_02] I'm more looking for like, okay, if you want to get started in this space, how do you even begin to get your head wrapped around all of this? Of like, what do you even want to evaluate? Yeah. Like, how do you start thinking about that first step? [SPEAKER_00] So you have to kind of split on terms of like, I want to get started, but I really want to just like have a thing that does stuff for me where the you're more focused on the get me to the outcome as quickly as possible. And I don't actually care as much as like the internals or to kind of evolve it over time. And I would argue your best case there is to go get one of these kind of prepackaged agents, right? [SPEAKER_00] Go to the Zappias of the world or the Muses of the world or whatever it is that you're going to go use and, you know, download that on your phone and get yourself up and run. Right. And, but the, the, that, will you learn how to build these things? No. Will you learn how to control them? Will you make them malleable? No. Like, will they do tasks for you? Absolutely. It'll be amazing. You'll love it. Right. So that's like, Hey, if you really want to get in deep and you want to say, look, I need to understand. I want to go build coworkers. I want to go do it. [SPEAKER_00] Then I would argue that the best way to do it is to kind of like get started with some of these harnesses plus the models and experiment into like what I call getting very good at micro task completion. So pick like the one, two, three things that you wanted to do. In some cases, you know, like, um, a lot of the use cases I really like are, so you do this. Cause like there's go to market. [SPEAKER_00] There's a lot of kind of like repeatable, repeated operational stuff outside of the creative and kind of like interactions you need to do. Right. So like once you have a piece of content preparing it for 17 different audiences and how to do that. And so these are all things that are great to do these. So like, what I would say is if you want to experiment with, uh, don't overcomplicate things, start with the beauty of this is also you can use models to teach you how to use these things. Right. So like there's literally no excuse. [SPEAKER_00] I tell people, um, you know, so I teach MBAs applied AI. Right. And so initially this is like a year ago. It was like, Oh, you know, I'm not technical. I was like that excuse went away a long time ago. Like that's not a thing anymore. Yeah. Like if you can talk and you can, you speak because if you can talk and you can speak, you can learn. [SPEAKER_02] Yeah. I mean, here's, here's the, the, the, like the things I do, like, okay. You know, automation projects for the podcast and I will do it and I'll get some level of, of success. And then I will literally keep asking the AI, okay, we did this thing. How do I write this prompt better in the future? Or how do you write this skill better in the future? Or how do you write the instructions for this project better in the future? [SPEAKER_02] And I actually get them to write all of the instructions for me. Like a little bit like your point. I'm like, okay, I do this thing. And I just, I admit, like, I kind of iterate on them over time, but I'm always like, okay, it did these things great. I didn't see this case and okay, now format this more better in the future or do other things like that. I know enough now to tell it to make itself better. Yeah. [SPEAKER_00] Totally. But you know, I, I, I say like, I think like the learning here is experimenting is, you know, I'm biased. I'm a tinker at heart. We run this thing called AI Tinkers. Go play with it. Like, it's like, it's the, it's one of the greatest moments in my opinion, to be alive right now. Everything gets to be played on. You can get to build things. You get to see, you can make things like if you can dream it, you can make it come to life. And so building these things is super fun. Right. [SPEAKER_00] And I would say if you're, if you're focused on usefulness and that learning on the usefulness is just, you know, go get one of these harnesses. You don't need to go buy a Mac mini. You don't need to go do like, you know, pay billions of dollars on model. I mean, it gets expensive once you're, if you don't have good inference deal on models, but the, go talk to Aaron about that. But the, I think that's the, it's fascinating and it's, it can be done, right? Like you can build it. [SPEAKER_00] Like my mom has her own agent now and, you know, she's, you know, 70 years old and calls it her best friend, but she like helps her plan like, you know, little vacations and recipes and stuff like that. And so it's really accessible. [SPEAKER_02] Yeah. What is your final closing question here? What is your thoughts on, I mean, we talked about it once or twice here, this idea of, I won't call them the second generation of agents, but this whole idea of like, okay, it's less about the framework and building it yourself in tinkering and more about a bundled agent like Muse or some of these other ones where it's, maybe it's all in one thing and it's free, maybe it's something that, you know, you, you paying per agent per month [SPEAKER_02] kind of thing. Like, tell me a little bit about where you think this is all going. [SPEAKER_00] I think we're going to have, I think people are going to be able to have AI assistants that are going to extend like how they do work. And some of them are going to be coworkers in a enterprise environment. And otherwise you're going to have a personal AI or multiple personal AIs that help you with like your daily life and you'll have control over them. They'll understand you. [SPEAKER_00] Memory and context will become key in terms of both access, but also kind of in how do you manicure these things? How do they learn about you? How do they kind of provide for you in that special way? And I think it's going to create a explosion of personal productivity. And like, uh, I legitimately feel that today I can time plex myself. Uh, it's the first time in my entire life that I've ever, I'm usually a pretty type [SPEAKER_00] very persona who like is somewhat efficient outside of not being able to use a calendar. And so I've never felt more productive in my life. And so I think that's where these things go is like, we're just going to be normal. And in an enterprise environment, I've thought a lot about this because I think, you know, there's a lot of fear around people are like, am I going to get replaced by AI and stuff like that? But I think what we'll find is that a lot of our jobs will be, which is what I do today [SPEAKER_00] a lot is AI still need humans for exception handling in a lot of cases. And, you know, I joke around a lot of my manager friends don't love how I frame this, but I'm like, what is a manager other than an exception handler? Right. If like everybody's doing things fine, they don't need you. Right. But sometimes there's exceptions and it gets bubbled up to you. And then you make a decision on that exception. And then you kind of like, you know, motivate the team and tell them, give it a direction. And then you kind of like deal with exceptions. Great. [SPEAKER_00] So now imagine that you're going to just do that, except everybody under you is just going to be an AI agent. They're going to bubble up the exceptions to you. You're going to like process it. And so if you go look at how software is being built today with like software factories, that's essentially it. Right. You're doing this, like, you know, specking in the beginning, sending it into the factory, dealing with the exceptions, software is coming out the other way. And so that's kind of the future I see, at least today. What I tell people is given this space, I fully reserve the right to change my mind as new information comes in. But that's where we're at. [SPEAKER_02] Yeah, no, I like that. And I'll add to that too. I agree with everything you're saying, but I also too, like, I can't remember where it was, but there was a job description I saw recently. And it was, it was actually for somebody. It was marketing for an AI company. And what was, what stood out to me was to your exact point of, okay, it wasn't necessarily, hey, we're replacing the marketing team with agents. [SPEAKER_02] But if this was a head of marketing job overseeing a team of people, but then also it was like, you need to be able to write agents and manage agents and be hands-on with agents and have done a lot and be at pretty cutting edge with the agents. And that's something I see going forward as well is like, people are going to be asking you a little bit of like, okay, not just what have you done with your career, but like, what [SPEAKER_02] have you done with agents? Yeah. [SPEAKER_00] I mean, I think, you know, I think it's important to be up to speed with these technologies and understand them. Right. I don't, I find it hard to anybody who is a technologist to be ignoring this entirely. I mean, it's everywhere. I mean, how do you, I feel like it'd be harder to ignore than to adopt at this point. [SPEAKER_02] Agreed. Agreed. And I think that's actually a great spot for us to call it a day here. Diego, thank you very much for your time. And anyone that is out there that wants to kind of get started, reach out to you, follows, you know, some of the things we mentioned, you know, where can everyone get started? [SPEAKER_00] I think there's a, so probably one of my favorite, there's a, depending on kind of like what your area is and what you're working on. There's a couple of different areas. I used to read a lot of, there's a couple of newsletters like Ben's Bites. That's really interesting. I don't know if he's still doing it. And then there's like Lenny's podcast. There's another person who's like actually working on a lot of stuff that I think are particularly interesting. But it's a, there's no better way in my opinion than if you're already a builder and you're [SPEAKER_00] already a tinkerer, you know, go find an A&A Tinkers event in your city, sign up, go show up, go look at the demos. I promise you, if you're not inspired, I owe you a beer. [SPEAKER_02] There you go. I love it. Fantastic. Guaranteed success by Diego. You heard it here. All right, everyone. Thank you very much for listening this week. And on behalf of Brian and myself, Diego, thank you very much for your time. If you enjoy the podcast and you can leave us a review wherever you get your podcasts. And of course, we're always welcome to feedback and ideas for shows and guests. For Brian and myself, that's it for this week. And we will talk to everyone next week. [SPEAKER_01] Thanks for listening. Check us out at theenterpriseaishow.com for past shows, newsletters, and all things Enterprise AI.