← Back to search
Agent Client Protocol
AI Infrastructure · 2026-08-16 · 46 min
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
How AWS's open-source Kiro Crew gateway, built on the Agent Client Protocol, compares to a self-hosted
Hermes Agent setup — and how to swap in your own agent backend.
Benefits
- Full-featured web dashboard vs interacting with Hermes only through Discord
- Fork Kiro Crew to run codex on a $20 ChatGPT subscription
- Out-of-the-box memory system with local embedding model, no manual setup
- One command (kiro crew cloud) launches a remote crew on EC2
- Everything self-owned: your box, memories, S3 publishing — Apache 2 open source
Use cases
- Forked Kiro Crew to replace Kiro CLI with codex ACP, run on a $20/mo ChatGPT Plus subscription — only 8% of weekly limit used
- Ran a private EC2 remote crew; NAT gateway drove costs to ~$50/month, cut via EventBridge scheduled scaling
- Claude found the Singapore NAT gateway cost $40/mo; moving to us-east cuts the bill by $10
- Generated sequence diagrams and interactive HTML artifacts during a call, publishable to S3/CloudFront
- Diagnosed Hermes memory failures: Markdown file size limits blocked memory writes for a week
KPIs / results
- ~$50/month for a private EC2 crew instance; NAT gateway alone $40/month in Singapore
- Nightly scale-down saves ~$30/month; region switch to us-east saves $10 more
- 8% of weekly ChatGPT Plus limit used running the whole crew
- Embedding vectors have ~1024 dimensions for semantic memory lookups
Tools / build
- Kiro Crew (AWS, Apache 2)
- Agent Client Protocol (ACP)
- Kiro Crew fork running codex ACP on ChatGPT Plus
- Hermes Agent VM with hindsight.ai memory + Postgres vector embeddings
- fck-nat on T4G nano to replace NAT Gateway
📑 Chapters — tap a time to jump there
0:00
Holiday, kids, and startup life
- Staycation with kids ages 10 and 8; cousins moving next door
- Paul Graham: balancing family life with startup culture remains unsolved
3:58
The pressure to stay relevant with AI
- Underlying pressure to stay relevant and ahead of the pack in AI
- Five years of remote-work lock-in without early retirement
5:06
Solar farms vs food security
- Tennis-court debate: England solar farm buildout vs food security objections
5:49
Screen share: what is Kiro Crew?
- Kiro Crew: AWS's open-source OpenClaw/Hermes-style enterprise agent solution
- macOS app bundles a gateway hooking into local authenticated Kiro CLI
8:06
Agent Client Protocol is not MCP
- Agent Client Protocol connects everything — distinct from Model Context Protocol
- Kiro Crew is a gateway like OpenClaw was for Telegram/WhatsApp
10:23
Where does the gateway actually run?
- Gateway runs inside the macOS DMG: React/Vite SPA plus Python gateway
- Talks to Kiro CLI on your machine via ACP
11:24
Remote crew: one command to launch on EC2
- kiro crew cloud bootstraps CloudFormation and spins up an EC2 gateway instance
- Remote pane served over SSH/SSM private tunnel; slow first load, fast after
- Nicer UI than interacting with a Hermes VM through Discord threads
15:53
Memory systems: Hermes vs Kiro out of the box
- Hermes memory choice: hindsight.ai, Nous cloud, or Markdown files
- Markdown hit file-size ceilings; memory writes failed silently for a week
- Kiro ships a full memory stack with local embedding model out of the box
18:19
What an embedding model actually does
- Embedding models convert text to ~1024-dimension vectors for semantic similarity
- Cat, dog, house pet land close together in vector space
20:26
Maintaining memories: markdown vs vector DB
- Markdown is portable (SCP to new machine); vector DBs need cleanup jobs
- Memory systems should self-maintain stale memories in the background
21:53
The multiplayer question
- Gateway lacks multi-tenant users; artifact comments are the multiplayer feature
23:46
The NAT gateway tax and fck-nat
- Private-subnet instance needs NAT gateway: ~$50/month for single user
- EventBridge scheduled scaling stops it nightly; fck-nat on T4G nano is cheaper
- Claude: Singapore NAT costs $40; us-east switch saves $10 alone
27:41
Using it for real: diagrams and artifacts
- Asked about AgentCore; got mermaid diagrams, then a proper HTML artifact
- Artifacts gallery keeps all previously created diagrams locally
29:19
Forking Kiro Crew onto a ChatGPT subscription
- Public fork swaps Kiro CLI for codex on a $20 ChatGPT Plus subscription
- Usage dashboard shows GPT Plus stats: 8% of weekly limit used
31:25
Elicitation and ACP capability mismatches
- Gateway wrongly advertised elicitation support; codex requests got 500/reject errors
- codex implements elicitation, Kiro doesn't — ACP capability mismatch
33:44
Faking usage reporting with a facade
- Built a facade exposing codex usage through Kiro's usage-reporting format
- Fable pinpointed that codex sends usage with every message
35:56
Why a Claude Code ACP gets your account banned
- AWS runs Claude Code ACP internally but stripped it: subscription use gets accounts banned
- OpenAI permits codex on subscription; fork may merge upstream
38:47
Option chips beat AskUserQuestion
- Without elicitation, codex returns structured options rendered as clickable chips
- Chips beat AskUserQuestion: click multiple, edit before sending
41:05
Local auto-classifier and trust prompts
- Kiro Crew security: local auto-classifier with trust-this-action or trust-all prompts
42:40
Why AWS rewrote Kiro onto a single harness
- AWS rewrote Kiro IDE, Kiro CLI (ex-Q CLI), and cloud Kiro onto one ACP harness
43:43
The artifact reveal and publishing to S3
- Interactive ACP-explainer artifact with comments, reminiscent of Claude desktop
- On EC2 with permissions it publishes straight to S3/CloudFront
45:31
Take a break
- Still to explore: extensions, scheduler, task runner, GitHub identity
- Happy with Hermes Agent but Kiro Crew may replace the Hermes box
How are you? Yeah, good. I've been on holiday this week. But I've just been at home. So in the end of the day, I just been trying to spend time with my kids and not to sure I'm doing the greatest job. I just end up arranging activities and then watching them play from afar. How old are they? Ten and eight. Oh, that's nice. They're close to each other. Yeah. And they can entertain themselves there at an age. Yeah, it's quite good. Yeah. And then the big news is that my sister and her husband's moving in here and they're going to live next to us so that they got two cousins to play with all the time. Oh, that's great. So I guess I don't need to be around anymore. I can just fully devote myself to AI. Yeah. Oh my God. That sounds so, I mean, not at all what I'm doing though. No. No, no, no. You've got the balance. You've got the balance just right. Yeah. Yeah. My son is 17 and my daughter is 12, no 10. Uh, she's six years younger, so she must be 11. Yeah. She's going to be 11. Yeah. And so it's a bigger gap and there's a bit more, a different power dynamic. Like he acts like, you know, cause he, he feels like I spoil her because she's my little princess. So he's there to set things right. You know, to be a bit more disciplined when she's, you know, taking all this for a walk. I don't know if you can see this ring like in the reflection. Oh, I can actually, I can. Is this your, your light or something? It's this feature on when I got this Mac book and I joined the meeting, I wanted, I was in a dark room and it says, said the light is bad. Do you want to turn on the cap, the ring camera? And I thought it was a Microsoft teams thing. And I was like, I'll figure out how to turn it off at some point, but it's everywhere. And it's annoying because it's, and I haven't figured out how to turn it off yet. I should ask Claude to turn it off for me. Yeah. So you're in a different room now. Yes. So my, I had to give my son's room back so he could spend some private time. He can get some private time. Why you say it like that? Well, yeah, it's been, it's been good to disconnect for a little bit, but I still like read Hacker News in the morning, like a lunatic. Funny enough, the Paul Graham was, was talking about that one problem that he felt that was not solved with the Y carbonators is how to balance family life with, with startup culture. Family life startup culture. Yeah. It doesn't seem to be any way of making those two mix very, very well. I mean, I personally know at least, at least a couple of people that have, have thrown themselves deep into tech. And of course they don't have a family. I would say, which I think it's a bit of a shame. I don't know if you have got any of those examples to hand, but I know, yeah, I definitely know a couple of people who went deep. Yeah. I, aside from myself, I have a family, but just cause they're quite accommodating to, to, to me being so like, honestly, it was supposed to be temporary, right? After COVID, you know, it was supposed to be like, you know, just, you know, lock in for a while. It's been five years now, um, to, to prepare for the future because, you know, COVID made remote work possible. And exactly. It allowed me to try and put away as much as I can and retire early, except I've been doing that for five years and I'm not retiring early. So yeah. I, yeah. Like with AI and the investments I'm making, I feel like I, I'm going to be working to the day I, well, until the, into retirement age at the very least. But yeah, I, I feel the pressure with AI. I mean, I, I do enjoy AI. It's just that of course there is, there is like an under, an undertone and underlying pressure to, to stay relevant and, and, and stay ahead of the pack. And it's, yeah, it is a bit, it is a bit much. It is a bit much, but like, just gotta keep on doing it and try and make it fun. So hopefully you have some fun things to tell me, Vincent. Otherwise I'm just gonna, I'm gonna sign off now and spend some time with my children. Let's not stand still and contemplate or decide, like, make sense of what we're doing and let's just dive in and focus on what, what, what other fun stuff we read about. Yeah. Never stand still, never look back. Gotta keep it fun. Yeah. Don't ask, uh, you know, ethical questions. Yeah. Oh, I got, I got, I got into, I got into a silly debate yesterday with one of the people I played tennis with. Uh, I'm a proponent of having, of building solar farms. I don't know what it's like in Vietnam, but like this, there's been a bit of a build out of, of zero, my God, what's it called? Zero, whatever. Um, zero footprint energy, whatever it's called. So basically solar farms in, in England. And there's like a real opposition to it because like, Oh, what about our food security? But like, I can't have argue like, isn't energy security slightly more important. Yeah. And saving the world. Anyway. Okay. Those distractions aside, Vincent, you look at a bit drained and distracted. Come on. Yeah, no, let's, let me share the, the screen, entire screen. I'm going to do it on the screen that has the camera. Ooh, massive, um, inception right there. Okay. So you're back on the Kiro stuff. Yeah. So I'm actually running like a local build, um, where I did a couple of really fun things. So my version might not actually work properly, uh, because I went quite invasive and I mean, Kiro crew is kind of AWS open claw slash harmless solution for enterprise. Right. And since they made it public, they now are also sharing about how quickly it evolves similar to open claw or Hermes, you know, having agents on their own contribute the features. And then they have a whole testing harness and, um, review harness. So this is actually very interesting to, to see. And also to, to, to link up with the AWS bench that we were talking about last time. Uh, maybe because I do believe they, they, they do benchmark heroes capabilities. Um, specifically AWS type of, um, you know, problems. Do you want to focus on that first or like, let's just. No, no, no, no, no, no. This is called Kiro Crew. I kinda. Yes. So there's a big difference here. Maybe you see it at the top. I am currently running the Mac OS app and that comes bundled with a gateway that then hooks into your local Mac book, um, where you will have Kiro CLI authenticated. And the Kiro CLI and the Kiro IDE, which is the original IDE. They launched respect driven development last year has all been rewritten to work on a single harness, a single bank ends. And it uses the agent client protocol to connect everything. So you could, you get all of these different experiences. Um, and with, with like one, like if I go into the builder article I wrote. I noticed that AWS was contributing to the whole MCP thing and making it easier to. So this is different. This is agent client protocol, not, not MCP. Agent client protocol. Oh, okay. It's not the model. It's not the model context protocol. So what, what Kiro Crew is here? Um, it's a, it's a gateway and, um, very much like how, how, um, you know, open cloud originally was a gateway to telegram or WhatsApp, uh, for you to talk to your, um, agent, um, wherever it ran. Like maybe it's your, um, you know, codec CLI or your cloud code, you know, originally it was open claw claw, like cloud, cloud bot. It was related to cloud code, right? Yeah. And then codec, uh, you know, open AI hired Peter. And anyway, this is the same thing, right? It starts as a gateway and the gateway connects to MCP tools. And this is, uh, the dashboard app. Um, this is the actual dashboard app. You can also have it in your browser. So this is the dashboard app running in the browser. Maybe I was going to ask, is there a mobile app? Um, potentially, I'm not, I'm not sure. I haven't looked into that. Uh, but anyway, so, so this is the whole component layout, which I think is a little bit murky when you first install it and you don't know, like when you first download or build the dashboard app on your machine, it's actually running on your, on your laptop. Um, this, this macOS app right here comes with sessions to, for chat schedule. Um, everything is, is, is there and it runs on your laptop, which I did not like. Like, I don't want this on my laptop. I want it. Yeah. It does feel a little bit like an anti pattern. You know, when people install open claw on the, the Mac minis and like, yeah, take over everything. Yeah. So, so, so that's where I got confused and that's where I really liked this diagram because I didn't know what the components were within. And, uh, but the most important part is that the dashboard app comes with like a bundled gateway, uh, and then it connects to your Kiro CLI on your machine. And then if you authenticate Kiro CLI or you use Kiro before it will inherit everything you already have there. So, um, then it's, it's a little bit like herder, what you mentioned, you know, it gives you an overview of all this, of all the sessions. Where does this gateway run? Is this an AWS hosted thing? So for example, here, um, let me just, this is, you get like all your chats of your, your local session. Right. So when you do the gateway in the Mac, Mac OS app, it's, it's basically just, it talks to your, uh, CLI on your machine using ACP, right? Agent client protocol. Okay, cool. So, um, oops. Yeah. So, so that's the first thing. What does the gateway run on your machine? Yes. I'm going to say it the third time now for you. Okay. I, this thing runs the gateway inside the Mac OS DMG. If you build it, it bundles the single page app, like a react VIT and the gateway, which is a Python, uh, gateway. And that connect with the protocol to your Kiro on your machine. Okay. So this, this Mac OS app runs the gateway. Got it. But then, and now let me, let me, let me tell you one more thing. One more thing here is that this particular gateway has a remote crew capability here. Ah, yeah. So with this remote crew, you can, um, if you use the Kiro crew CLI, you can, you have a cloud sub command. So actually that's, um, the next part of this, of this article here, uh, is that you can set up remote crew and you can run Kiro crew cloud, um, to manage remote instances. And, and that basically bootstrap a cloud formation templates into your AWS account and spins up an instance with the gateway as well. So that's a separate remote gateway that you didn't, that you then connect to. Okay. So that's a, that's a separate machine. You may run the Kiro CLI on it. Um, then you need to authenticate your Kiro CLI and, um, with the device mode authentication on your subscription. And then you, you get, um, you know, Kiro access and all the models that Kiro support. Um, and you, you, you get the same thing here. Um, sorry, you get the, where was I? Uh, where is the app? Oh, here it is. So, so you get, you get, you can connect. So let's, let's say you spin up one and then you can connect here and then it will pop up at the top. And now you have, uh, all of the chat sessions you have with that instance over there and it can run, um, you know, you can set up tasks to run on, on, on a schedule. So the thing runs on its own. And, um, I mean, I guess this is like what you would get if you run the Hermes app, which I never ran. Um, so, um, I have a Hermes VM and right now the only way I interact with it is through the discord channel. And it's not a very nice experience, right? Every interaction is either through a discord thread. Um, and, and in this case, actually I get a nice UI and I think, I do think Hermes has this as well, but this is the AWS version. I, I've shown you this before, but I, I use this thing called command and control, which has a mobile app and a webpage where I, I, I connect to all my, my different, uh. Okay. So, so this, this, this is confusing because I have another sessions here, settings here. And then here I don't have the remote crew, uh, option because this is a remote gateway already. Um, and what's very interesting is when you're developing and you're making changes to the gateway, and then it goes and deploys the gateway into that remote instance. You have to like reload this whole pain because this whole pain is served over SSH or over SSM via private tunnel. So, so this actually is a separate instance of a gateway running separately that you, that you get over, um, over a tunnel, like an SSH tunnel. It's very interesting. Yeah. So that was interesting because I was building, I was re uh, changing things and, and, and reloading the app is actually a little bit unintuitive because I was like, do I need to rebuild my, my, my local, uh, Mac OS? So, so this app and, and, uh, cloud was like, no, uh, the whole app loads over SSH remotely. So once the, the whole, um, asset, all of the assets have been downloaded over SSH, then it runs very quickly. Um, but the first time it can take a while if you run a crew somewhere in Singapore and you're in Vietnam and your connection is not great. Um, then, then the first time loading that is a bit slow. Uh, but once it's loaded, it's, it's pretty fast. I like the look of this whole ecosystem. I guess, I guess this is like remote control. Control, Claude remote control on steroids. I guess we're gonna see more of these sort of build outs, right? Where you, you, you are able to see your, your, a number of agents in essentially one place, right? Yeah. Uh, and there's also the ability here. Okay. I, I wanna really go into the features that this thing has because it has a lot more like my experience setting up Hermes VM and then coming into this, um, is like, holy, you know, S word. This is them, um, packed with features. Um, you know, I haven't tried this yet, but right now this remote, uh, instance is actually a virtual machine on my desktop at home, only accessible over LAN. Uh, I'm very excited to enable like tail net, uh, or, or like tail scale and then being able to connect to it remotely. Um, ideally over the control. Yeah. I think when I'm using OpenClaw, I have that kind of up the box. Uh, yeah, sure. Now let's go and talk about some of these cool things here, like knowledge. So I don't know if you've set up, like, I don't know about OpenClaw, but Hermes, when you set it up, one of the first thing you have to do is like, what is your memory system? It's flexible, right? You can choose which one you want. You can use hindsight.ai. Uh, you can use another, like you ideally knows, knows, knows, uh, the, the company behind Hermes wants you to use their control, their cloud to use their services. Right. But you are free to use. I don't remember. I don't remember choosing one. Yeah. Yeah. So I, yeah, it also can use Markdown files. So on this kit, we'll have a sold.md and maybe some other, uh, Markdown files. But what I found was that it very quickly hit a ceiling on like the restrictions on the Markdown file size. And also when it tries to save memories, it writes down to Markdown and then it hit that size limit. I basically used Hermes without really knowing what, what it was doing for a week. And then I asked Soné, what are the common failures? What are the errors that you see in the logs? And one of the first things Soné, uh, highlighted or cloud code highlighted was that, well, it's having a lot of issues writing memories. Like it's hitting Markdown file, C and things like that. Oh, really? So I wanted to figure out what does a memory system provide and like, how can I use it? And I started to learn a little bit more about, um, like on the memory side of things. There's, there's information about how do these memories, uh, link together. Um, you know, semantically, you know, recall rate, like when were the last recall? Are they still fresh? There's a lot of, uh, aspects, uh, about a memory system that you don't really want to bother with. Right. But, but I had to learn it because I had to like understand a little bit better about Hermes and setting up the memory system. So I w I'm running on the desktop, a complete full stack of, uh, containers, which, uh, uses Postgres vector embedding. I have a embedding model that I had to set up with my GPU. And then you come to Kiro, you install it and everything is there. Right. I'm trying to find now, but in terms of like, if I look at the memory configuration, it comes with a local embedding model. Where is the memory? What does, what does embedding mean here? Yeah. So embedding models are basically converting texts into token, into, uh, token arrays, right? So, um, a vector. So, so it's a bit more. So tokens becomes a vector that gets stored into a vector database for semantic lookups. What, what's the benefit of that? If I say cat and, um, it can, it can relate to house pet, even though they're not lexically close, the concept of cat. Dog is, is, is semantically close to house pet. So converting a word like cat or dog to tokens and then calculating, uh, using an embedding model. So you have to use the same embedding model that you, that you, um, store the data with the vectors with, because it consistently converts all of this. It's, it's a small, open, uh, weights model that, um, gets text and then converts them into vectors. Mm-hmm . And then those vectors, they have like similarity. So cat, dog and house pet all have very close semantic meanings. So their vectors will be very close to each other. So if you do, you can imagine, like if you, if you imagine a two dimensional vectors, you can calculate the distance between them easily. Right. But these are vectors with like 1024 dimensions. It's very hard to imagine how, uh, you know, but you can calculate the similarity between them. So, so that's what an embedding model does. Cool. Um, you know, lookups, you need to run an embedding model. Uh, when you get text stored into a vector database, and then when you want to query the database, you need to, um, embed that query, send it, um, to the vector store to do a similarity lookup. And then it's going to return facts about that are related to that query. So embedding is very important and running embedded locally on your machine is kind of, it's happening. You have a lot of small little, um, models that are running on your machines already. Cloud code ships with the local, uh, model. It, it, it, it runs a local model to do the auto mode classifier. I'm pretty sure. Oh, actually no, actually no, because, because when you, sometimes you get an error saying the classifier hits an API or something. So it is, it is remote. Yeah, it is remote, but the, the, on the topic of memories. So how do you maintain your memories? Cause like Markdown is cool because you can just like SCP your, your, your Markdown into a new machine. How do you maintain your, your embeddings database here? It doesn't sound really portable. So ideally you don't, right? Like, so if ideally that's where, where these, these memory systems have all kinds of like capabilities and cleanup jobs and, and, and, and background functionality to maintain the memory. Right. So I don't like with hindsight, I don't even know if there's like a, a cron running to like identify any stale memories and re, re, reorganize the memories to see what's, what's relevant. And I honestly don't really want to worry too much about it. Um, but you see how like that you miss a lot of that when you have simple Markdown on disk, right? Yeah. Yeah. This is, this is why I kind of, I always kind of gravitate back down to Markdown cause I can understand it in my human monkey brain. Yeah. So here it is. Embedding is using Lama C C C C P P. I want to find because it actually lists out the local models that, that Kiro ships with and it runs. So I can't find the model. I think it's in security. If I go to settings and security, I think the auto mode classifier, um, is somewhere in the settings. So it's, it's very interesting to go through and, and, and figure that out. Uh, but I think it runs at least two or three. But the next burning question in my mind is that what's the, the multiplayer aspect? Is it just through, just, are you supposed to, is one person supposed to admin this and then your team is on Slack or something? Or was there some other usage paradigm that's supposed to happen here? I do know that Kiro, um, when you use it, it has this artifacts capability. And I know if you look at the artifact, you, you can leave comments on things. Um, and I do think that this comments thing, I feel like that's really something that is a multiplayer capability. Um, but. So that means every, every one in your team runs. I don't know if you can like, I guess as far as I can see right now, the gateway doesn't have like a multi tenant or multi user, um, capability. When I look at these and I want to click publish to share it with the team, it's going to use a skill to. If this is a box running in my AWS, it's gonna try and publish it through like cloud front or S3. So the cool thing is that everything is yours, right? If you use cloud code and you create an artifact and then you, it gets like a cloud hosted URL. Publishing it is all controlled by Anthropic. This is all your box, your memories, your, your S3, your account. This is the power of AWS. They can sort of connect the dots, can't they? Yeah. With your, with this, with this capability to just run one command, uh, Kiro, crew cloud, and then you can just launch. And then you, you consider, do you want light balance, but you do realize that you needed a significant amount of CPU and memory because you're running all these local embedding models, right? Well, you know, and you're running a whole memory system and more. So, so the light one is, it's not cheap. And I guess it's not serverless. The thing is running all the time, even if you don't use it, right? Even worse, I, I went for the full private instance. That means it was running on a private subnet, not public. And that means I had to have a net gateway. That means, uh, I got a whole bunch of extra costs, which is already $50 a month for like a hobby or a single user. That doesn't make sense. Right? So, yeah. Yeah. So, so I did set up scheduled scaling, um, which basically with EventBridge, uh, stops it at night in the evening and then scales it up again in the morning. But, um, but still it says the net gateway is really the cause of your costs. Oh, this is your own blog. Is it? Yeah. This was the blog post I posted. I have a, I want to post four parts. The first one was this initial cloud based exploration. I did the same with Hermes, right? I just, um, you know, spun up a Hermes instance, looked at how it felt, looked at what I could do with it, realize it's going to cost me 50 or more dollars a month. Definitely if I, I wanted to have a memory system, you know, I wanted to have like the hindsight and all that, uh, fire crawl for, for web search. Wait, but why does it cost you 50 dollars? You don't have a, you don't have a pie or something? You don't have a VPS? What do you mean? Um, yeah, this is a different question. Okay. I'm talking about exploring and getting started quickly. Um, I didn't have, like, I, I, I didn't have like pie lying around or anything like that. Um, so I wanted to quickly, you know, try, try this out. And, and my quick start to try anything out is, you know, spin up an EC2 instance, do a bunch of things. If I don't like it, tear everything down. Uh, if I keep, if I want to keep it running, it's going to cost me at least $50 a month. Um, but of course, if I run it for an hour or two, it's going to cost me maybe $2, $3 for an experiment. I don't have to go and buy a, a pie for that. Okay. So, um, NatGateway and also funny, um, because again, the, the, the information for this blog post came mostly from cloud, but I actually rewrote it and wrote it all manually. I mean, large parts of it, most of it. Well, why did you have to rewrite it manually just because it? Cause I wanted to make sure it wasn't to AI slop and it was actually, um, you know, You tightened it up. Yeah. I tightened it up. Um, but the funny thing is that cloud is like, yeah, but you're running it in, in, in Singapore and Singapore, the NatGateway is costing you $40. You, the savings that you have on a nightly, um, if you scale it down night, nightly is only three, like $30. But if you would switch everything to the U S East, you would, uh, already cut the bill by $10 without even the up down schedule, which is a significant difference. Right. If you, if you put it in proportion that way, I thought that was pretty interesting finding. Yeah. Yeah. I I'm, I'm always surprised how expensive the NatGateway is and it's just to be in a meme for a decade now, hasn't it? Yeah. So there's two things I want to follow up on this article. It's to use FCK net, right? Which is the, um, Oh yeah. You told me about that. You told me about that. Yeah. I use it in, in dev environments at work. It's really cheap. Um, it's a T4G nano because ultimately if you don't need a high available multi-instance auto recovering high, high, high throughput NatGateway, uh, you can run on a single Nat and set up, um, what's this called? Nat is really just a couple of, uh, IP tables. It's nothing. Yeah, exactly. You can just do it with IP tables. If AWS announced that NatGateway was being folded into something and you wouldn't be in charge for it, it would be so great. But anyway, yeah. So great thing about Riverside is it connects, it records everything locally. It's just gonna be weird on the edit. Anyway, back, back to career crew. So already like, this is great. Like I told you about all of the capabilities comes with out of the box, but how does it feel to actually use it? So I, I tried some things out. For example, I was, uh, I was on a call. Let's, uh, close the sessions. Um, and I wanted to ask a question about, um, agent core, some capabilities there. So I was asking here, it gave me some diagrams, but then I asked it to, to create. I hope that it would automatically create, like I wanted the sequence diagram because it shows these things and I wanted to as a sequence diagram. I was hoping it would create an artifact, but it's put out a bunch of mermaid. I said, now please just create a proper artifact. And, um, what's a proper artifact? HTML basically, I suppose. Yeah. The same, like if you talk in cloud code and you say create an artifact and then it opens up the browser with that HTML, uh, uh, I don't really like mermaid diagrams, but oh well. That could be just. Yeah. Well, even cloud code is doing most of it. Uh, when it creates the artifact, it's using a mermaid or other diagram tools under the hood. Um, but anyway, so while I was talking, it created these diagrams. Um, and then you can ideally easily share them with your team. Right? So this was pretty cool. Like I found that a very nice experience. Uh, it's all locally, um, on my machine. Those are good. Those are good. I click on artifacts. I can see all of the previous artifacts it has created. Um, for example, here. All can do better in this regard. Well. Yeah. But remember that here. Okay. This is one thing I haven't told you yet. I'm, I'm running this all off. Um, my local fork, which uses my chat CPT $20 subscription. It doesn't like, it's not cloud code. I'm not paying a lot of money here. Uh, basically my, my family is using the, the chat CPT app on their phone and they're very happy with it. Uh, and I'm using it to generate, you know, on my, my machine. So that's what you did. You, you basically forked Kiro so that I can work with you, with your own personal. Cloud chat CPT subscription. So if I go here in settings and I click on a usage overview here, I can see here my, uh, it's my GPT plus and weekly reset is in six days. And I've only used 8% of my limit. Um, how many chats I have. So all of these statistics are, are my chat CPT subscription. And, and I think this is the greatest thing about Kiro crew. Um, if you go back to that diagram here, where is it? It's the whole thing's open source. Is it? It's fully open source Apache two. No. And it's using this ACP protocol or agent client protocol, which means they ship it with Kiro, but you, um, can contribute your own. Um, so what I did, I replaced the Kiro CLI with codex. And as a result, my whole Kiro crew is using codex under the hood instead of Kiro. Ah, now I understand why you've been excited to tell me about this. Yeah, that does sound quite cool. Yeah. So I've, I've, I've showed it to the team because when they announced it, they said that, um, this is ACP. So you can use any, uh, compliant and, and somebody asked, I think like, can we contribute our own other than Kiro? And they said, yeah. And I was like, okay, you, you said we can, it wasn't me. Somebody else said, so I said, uh, to fable, what would it take to rewrite this thing and not use Kiro, uh, but instead use, use codex and use this ACP. So I wanted to know, and it's actually very interesting because this agent client protocol, um, is a way of like, the client here is the gateway, right? Gateway is the client to the agent. So, so the gateway advertises capabilities of what it can show to the user. Um, for example, elicitation, when I ask a question, um, it may ask, like when the user asks a question through the gateway, the model might decide that it's better to get like, um, you know, feedback before it answers the question. So it gives you a couple of options, right? Um, that's, that's the ask user question tool in, in, in, um, in, um, in cloud code. Um, so this elicitation is a defined capability in the agent client protocol. So there's a difference between the capabilities of Kiro and codex. So codex implements elicitation. Kiro does not yet implement it. Okay. So what does that mean is that the gateway, uh, currently advertises that it can display. It doesn't actually eliminate elicitation as a capability. So codex assumes that the gateway can, can display. That means that codex will come back to the gateway and say, show the user a cart with three options. And the gateway, um, didn't say that it doesn't support it. So, um, it actually right now just responds with a 500 like internal server error. Um, it actually responds with a reject code, which keep, which codex input interprets as it was a security rejection and it cancels the request. So that was kind of like the unexpected, um, you know, kind of problems you run into because of, um, that was one thing. Also the whole usage like dashboard that I just showed you, it didn't work because the way that Kiro ACP, um, sends back usage information is different from the way that codex does. So I had to like fable highlighted exactly like where the problem lies, like the, the, the codex backend is sending the usage with every message. And I said, like, how about we create like, um, um, what do you call this a facade or like an interface on top of it? So I'm, I'm using codex ACP under the hood, but I'm exposing it as with the same capabilities as Kiro. Like codex says, I can't show you usage. So when you go to usage, the dashboard is empty, but I said, Hey, codex, there is usage information. Just present yourself as the Kiro usage and, and under the hood, it actually, you know, when codex sends the message, it captures the usage information. And exposes it, uh, as a side channel. So it still works. Yeah. So that's the kind of cool stuff that I learned doing this, right? Cause yeah, obviously I don't want to maintain a fork. Like I was actually thinking I set up this Kiro instance to, to automatically maintain its own fork and every time rebases on top of the upstream, uh, and, and, and things like that. But, uh, thankfully the, the AWS team is actually living up to their promise. And, um, and they're like, you know, evaluating because they're also saying, why don't we just, if you do the codex authentication, you get a bearer token. What if we just like use the Kiro ACP? Um, but you use the codex like bearer token that you get through the auth flow. Um, and that might be another thing that you could do instead of just replacing the whole ACP backend. I just did it for my, like my own exploration. Like what is HCP? What is the impact? How does that, you know, manifest itself? It was very cool. Yeah, this, this sounds really interesting. Uh, yeah, thanks for introducing me to agent decline protocol, but wait, but wait a minute. Um, sorry, Vincent, I was just daydreaming there and about using it myself. So you said you didn't want to maintain a fork. So, so going forward here, if I wanted to get started using my own, um, codex or codex, I'm going to use the codex. Um, you know, where, where the codex ACP protocol and the Kiro ACP protocol did not match up. And the ways that I worked around it so that I make things work. And, um, my fork currently is public. So anyone can just build a copy of Kiro Crew and run it with, with codex, uh, if they want. Um, I would not recommend building a code, a cloud code ACP. Um, but they, in fact, there is, there is a bunch of codes stripped out. So internally within AWS, they run Kiro Crew with cloud code ACP, I believe. Um, but they've stripped it out because if you do that on a subscription, you will get your account banned. Um, so, but, but open AI tells you, you can do it. So if you want to run Kiro Crew with codex ACP today, you can use my fork, build your own instance and use it. Hopefully in the next few weeks, if, if they, if they do believe this is a good addition, it might get merged upstream, upstream, right? And then you just have a single Kiro Crew and then you can choose either codex or Kiro. There's definitely a gap for the market. I mean, I was always kind of assuming that open code or some, somebody else was kind of filling it up. But like if, if AWS is, is pushing, is, is publishing good stuff with the Apache two license. I can't help. I think that it's, it's on, it's, it's a very good game plan. Very good. I want to show you this elicitation thing, um, because the fix was very interesting actually. And you can see it at work here rather. Yeah. So instead of showing a card here that, that this is the real elicitation. So normally codex would reply with like a, um, a schema of what should be in the form. It should show you the options and then the user can click on one of those options. And then the option gets fed back into the gateway gets, goes back to codex, but it's not supported. So what's, what's happening right now? I smell something weird like smoke. I hope I'm not on fire. Um, anyway, I'll go check in a minute. It's like, it's like this meme where the guy's on fire and he's like, fire, fire. It's always a good idea to nip the fire in the bud. Just go and have a look. Christ. Maybe it's the car that was on the car alarm. Maybe it was on fire. Maybe that's why the car alarm was going off. It's fine. Yeah. Anyway, so it's fine. What, what, what's what happened? The fix that, that fable suggested was just for the gateway. You know, remember the gateway is the client here to just not clearly advertise that it does not support elicitation. And when codex knows that it does not support elicitation, but it wants to get feedback. It actually sends back a structured response, um, with just a list of options. And that's what happened here. So I was like, I want to build ask user question. Like I want to have that nice interactivity of options. Yeah. And, uh, and then what I realized is I didn't have to build it because the gateway is able to show these options right here as little chips above my chat box. Can you click in and add notes? Oh, I see. So when you, when you click it, you can even click both of them and then it goes, they go inside here and then you can just modify it. That's an interesting implementation. It's a very nice, yeah. And I thought like, this is better than, than, than ask user questions, right? I was trying to build something I knew, um, but there was a better pattern that already worked. So that's kind of also something I learned, right? Sometimes you got to really use the system and understand how it works before you try to make it work exactly the way that you know. Yeah. Yeah. I mean, it is strange to think, I mean, we've talked about ask question tool a number of times, but it's one of those things that you can do. It's one of those things that we, we will keep on revisiting, I suppose, because the way you use it has so many nuances. And, uh, I don't think we've even touched on the multiplayer, um, aspect really, because ultimately the, the ask user tool probably is like one of the major decision points in your decision tree about how you get to your end product or somehow. It is quite fascinating to see where those, those turns were made. Yeah. So here's the interaction and maybe you can say, oh, this is like what the chat GPT app gives me. Right. But again, it runs on your box, your memory system, all your data. The second thing is you have the context window, you know, tracer here. You can choose the models. And because I built codecs, I have all of the chat GPT models. Can you show the tracer? I'm just curious because I've often find the tracing implementation is sometimes not what I expect it to be. What is the tracing? Oh, like when you see the, the actual calls and the, the thinking and things like that. Or maybe it's not there. Ah, you mean the thinking, the thinking steps that it did. Okay. Yeah. Sorry. I thought you wanted to see like telemetry. So is there some thinking? Okay. There's no, there's been no thinking on this chat, but let's go. What should the agent elicit from the user, uh, buying a car options? And let's stop with the questions and I have it build an artifact. Should the walkthrough expose the protocol? Enough questions. Build the artifact. Oh my God. So yeah. So here's the thinking, right? So you can see the thinking box at the top here. Oh yeah. And then this is the, this is one of the main features, uh, from Kiro Crew as well, uh, which is like security. There is this embedding model and auto classifier. Sorry, not embedding. It's a local auto classifier. So right now it's, it's, um, it says you can say trust this particular action or trust all. I'm just gonna go with the most secure way. We just trust everything. Yeah. It's really interesting. Yeah. But like if you only managed, like if you only used, um, Hermes through discord, like I have, this is a breath of fresh air. And I haven't even touched, and because of all of the capabilities that it comes with, right? With the task runner, the, the knowledge base, the schedule. And then you have the whole apps ecosystem. So they have a whole SDK and I haven't tried any of these, but you can have, um, like there's extensions where you have more than one agent. And so you can have a whole, um, animation and a world where you see like, oh, suddenly you review agent woke up because a new issue has been raised or a new pull request has been raised. And you see that agent moving towards a desk to start reviewing the things. Like those are the things that people show. And I think this is a bit stupid, but it looks cute. Right. And it's great, great visuals and it's good advertisement. Right. Um, but I haven't enabled any of that. Like not yet. Not at this point. Yeah. Um, I've got, oh, this is, man. Yeah. This is, this is a great little development. So the title of this podcast should definitely just be agent context, agent client protocol. Sorry. Yeah. Agent client protocol. It's a thing. And it's really interesting. Well, it's a thing that powers Kiro and it's a major rewrite that AWS did on how they built Kiro because originally they had three surfaces. One was the Kiro IDE, which is like a VS code IDE plus spectrum development. Then they had the Kiro CLI, which is a renamed Q CLI. And then they had the cloud Kiro, um, offering like a, a cloud service. And each one of them had to, had a separate code base and had to be redeveloped. So what they did is they rewrote all of it into a single harness, um, that supports the, uh, agent client protocol. And then they build different interfaces on top of it. So they build this, um, Kiro CRU gateway on top of it. They build the Kiro CLI as one of the, you know, interfaces on top of the harness. And they're also rewriting the, um, you know, other capabilities there. So I do hope that's all public information. I do think because they announced this, uh, that they made this public and they're advertising. Uh, no, I've definitely seen the LinkedIn posts. Okay. Well, let's end the podcast there. Cause I, I need to get on with some other. Wait, you need to see the artifact. It has created the artifact for you. And then, and then of course I gotta put my security hat on and work out. It worked so hard to create an ACP explanation for you using what the multiplayer, what the multiplayer story is. So I really love this interactivity here. Like when you select it, you can just go quote, um, then the artifact has been built. You click on it. It opens on the side. You can make it full screen. You can leave some comments on it. It reminds me of Claude, Claude code. I mean, not Claude code. Claude desktop. Yeah. Yeah. Maybe, maybe, maybe some. Do you not use Claude desktop? Not so much. No, no. I sometimes like I ask it to like fill out a PDF for me. Like, I don't know anything. Well, there's still, there's still, there's definitely value in an open source Claude code in my opinion. Definitely. This is pretty cool. Look at this. It created this, like users start to turn. Oh, it's like an interactive artifact. Oh, I haven't seen. Actually, no, you can do that with Claude desktop. It's really cool. Yeah, that's cool. I can see why you're excited about it. And then ideally I can publish this. Uh, if I, because right now this runs on a, on a VM on my, my, um, desktop at home and it doesn't have AWS credentials, but if it runs on an EC2 box and that box has the right permissions, it will publish this. And makes this available through whatever, like, do you have a blog and do you, you know, host it on S3 with cloud front? Then you can just directly publish your, you know, visualizations and sharing with other people right there. There's definitely, um, some, some value, a lot of value in having it open and this is quite exciting. Well, thanks. Thanks for letting, showing me and I'm going to get back to my holiday now. Are you, are you not taking any holiday in August or September? At the end of August. And this is only one part of it, right? I only showed you the artifacts. I haven't even started with all the extensions capabilities or the schedule or the task runner. I haven't even given it a GitHub identity yet. Like this is supposed to replace my Hermes box, right? And there's so much more that I have to do, uh, but I've had so much fun already with it. Uh, but anyway, yeah. Um, read, I'm going to post more blogs and, and share them with you. It's going to break down. Um, yeah, I'm pretty, I'm pretty happy with Hermes, Hermes agent. So, but, um, but I definitely will play around with this. Definitely. Cool. All right. Um, thanks for listening this far. If you got this far, like the video, do all that stuff. Otherwise I wish you a good summer holiday and take a break. Like I'm trying to do now. Take a break. So I guess I won't see it. Will I see you next week? Never stand still. Never look back. Don't look back in anger. Anyway. Okay. Take a break. See you, man. Bye.