← Back to search

Build Your Own Siri (Locally)

Rubber Duck Radio · 2026-07-24 · 32 min
relevance 66 4637 words Episode page ↗ Audio ↗
Show full episode description
Tim Williams unveils two open-source tools that turn your Mac into a fully local, privacy-first AI assistant, indexing your iCloud documents and iMessage history so you can finally ask real questions about your digital life. Along the way, he and Paul Mason dissect why Apple hasn't shipped this yet, call out a shady AI wrapper built on open-source hero Hermes, and share a hot tip for developers: SQLite is the unsung foundation of local AI. If you're tired of the hype and want an assistant you control, this episode is your blueprint.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Apple hasn't shipped a Siri that can reason over your documents and messages, so the host built fully local, private AI tools that do it.
Benefits
  • Fully local indexing — no data leaves your machine
  • Ask natural questions about iCloud documents with cited answers
  • Search a decade of iMessage history privately
  • MCP-native: works with Claude Desktop, Claude Code, Codex, LM Studio
  • Phone access over private Tailscale network with TLS
Use cases
  • Asked agent for auto insurance renewal date — answered with citation: car_insurance.pdf page 1, March 15th
  • Queried net proceeds from a 2022 house sale from indexed closing documents
  • iMessage queries like 'catch me up on the reno group chat' with contact names resolved locally
  • Mobile access to home Mac's index via Tailscale Serve while at the grocery store
KPIs / results
  • Apple has ~2 billion active devices constraining risk tolerance
  • M4 Mac with 32GB unified memory runs Gemma 12B comfortably
  • One developer built both tools in Rust over a few weeks
Tools / build
  • iiCloud — local iCloud Drive indexing daemon (Rust, MCP)
  • AI iMessage — local iMessage search over MCP
  • whisper.cpp transcription pipeline
  • SQLite FTS5 + vector hybrid search with reciprocal rank fusion
  • Tailscale-fronted private mobile access
0:00 / 0:00
and two, the competitive pressure is becoming existential. Google is shipping Gemini-powered features on Pixel phones that can reason about your photos and messages. Samsung has Galaxy AI. If you can buy a $500 Android phone that has a genuinely useful AI assistant that knows your documents and messages, and your $1,200 iPhone can't do any of it, the privacy argument stops being enough. At some point, privacy can't be the excuse for not shipping features. Hey there, listeners. Coming at you with episode 24 of the Rubber Duck Radio, your favorite nostalgic debugger. Speaking of nostalgic debuggers, how are you, Paul Mason? I'm good. I hear you've been cooking up some good stuff this week. What have you been up to? So I think there's a lot of justified AI hate out there right now. Your average user sees nothing but slop. They see jobs being threatened. They see these enormous data centers going up that are going to suck up all the water and drive up electricity costs. These are valid concerns. I'm not here to dismiss any of that. But here's the thing. I think there's actually a lot to be optimistic about in the AI space. And I don't mean the big splashy demos from the major labs. I don't mean the fundraising rounds or the press releases. I mean, what's happening at the local level? Models you can actually run on a mid-tier Mac laptop that are better than the multi-billion dollar state-of-the-art models from just a few years ago. These models can do things like what you wanted Siri to do when it first came out. Understand your documents. Search your messages. Answer questions about your actual life. Not just set a timer. Which begs a pretty obvious question. How come Apple hasn't done this yet? Like, genuinely, they've had 15 years. Exactly. How come you can't just ask Siri what somebody texted you or what's in your iCloud documents about your insurance policy? And there are reasons, real structural reasons we should talk about. But instead of waiting around for a trillion dollar company to figure it out, I've been building these tools myself to see what's actually possible. And, Paul, the results have been kind of incredible. Yeah, this is the stuff you've been teasing on the side for weeks now. Let's get into it. What did you actually build and how does it work? All right, so two projects, both open source, both in Rust. The first one is called iiCloud. It's a background daemon that indexes everything in your iCloud drive. PDFs, scan documents, images, spreadsheets, even voice memos and videos. Runs OCR, transcription, document understanding, and embeddings locally. And then serves the whole thing to any AI agent over MCP. Fully local. Nothing goes to a data center. Not even a telemetry pane. So it's like you point an AI at your entire digital filing cabinet and it can actually answer questions about what's in there. Not just search. Actually reason about the content. That's exactly it. Think about everything you would have put in a filing cabinet 10 years ago. Closing statements, insurance policies, tax forms, warranties, medical records. That one important email you printed to PDF. Today it's all scattered across iCloud drive in whatever folder you happen to be in when you saved it. And when a question comes up, When does my auto insurance renew? You're digging through folders, squinting at scans on your phone, hoping you name the file something useful. Now you just ask your agent. It searches the index, finds the policy document, reads the renewal date out of the extracted facts, and answers with the exact figure and the file that proves it. Car insurance pot PDF. Page 1. Renewal date. March 15th. Same for what were my net proceeds when I sold the house or find the parking pass for the game. Every answer comes with a citation. The proof is one click away instead of one drawer down. That's actually useful. Like, actually, actually useful. Not AI wrote me a poem about my insurance, which, I mean, I guess that's fine, but it doesn't help me when I'm standing at the body shop and they need my policy number. Right. It's not a parlor trick. And the questions get genuinely interesting. What were my net proceeds when I sold the house in 2022? What's my HOA contact information? Show me every document related to the kitchen renovation. This is the kind of thing Siri should have been doing five years ago. So walk me through the architecture because I know you and I know you didn't just slap an API wrapper on something. It's a two-pass design. First pass is pure extraction. A background demon watches your iCloud folder with F's events, pulls text out of PDFs using a Swift sidecar that calls PDFKit per page, OCR scans and images with Apple Vision, transcribes audio and video with whisper.cpp, all local. Nothing touches the network. Then the second pass enriches everything. A local LLM reads each document along with rendered page images and produces a structured summary document type, key value facts and tags. All JSON schema constrained so the output is predictable. Then it chunks everything, embeds it into a SQLite database with FTS5 for full text search and raw float 32 vectors for semantic search. Fused with reciprocal rank fusion so you get hybrid retrieval. Keyword plus meaning in a single query. The whole thing runs as a launched background service. You install it once with Homebrew, run the setup wizard, and it just keeps your digital life searchable forever. And the MCP tools it exposes, what can an agent actually do with it? Seven tools. Search documents. Get document with page full text plus facts and tags. Get chunk for neighbor context. List documents with folder and type filtering. Search facts for structured dates and amounts and parties. Sync status. And this is the cool one. Ask, which is a server-side research loop where the local model searches, reads multiple documents, and synthesizes an answer with cited file paths. All in one MCP call. So it's not just search, it's an actual research assistant that lives on your machine. Exactly. And it's MCP native. So any agent that speaks MCP, Cloud Desktop, Cloud Code, Codex, LM Studio can use it. You run AI Cloud Connect, and it prints ready-to-paste config for every client. Even prints a tail scale setup so you can access it from your phone over your own private tail net with TLS. Full privacy. Your network, your hardware. Wait, so I can be on my phone at the grocery store, ask my local AI a question about a document sitting on my Mac at home, and it just answers over my own private network? Yep. It's like having a private cloud that only you can access. The server binds to loopback only 127.0.0.1. Every request requires a bearer token. If you want mobile access, you front it with Tailscale Serve, and it adds TLS. Nothing leaves your machine unless you explicitly opt in. There's a privacy config key, allow Ramon endpoints, and it defaults to false. That's huge. And you said there's a second project? Yeah. AI iMessage. Same architecture, same philosophy, but over Apple messages. It indexes your entire iMessage history, every conversation, every group chat, going back years into a private SQLite database on your Mac, and serves read-only search tools over MCP. Your message's history is the most personal data set you own. Family conversations, money discussions, health updates, travel plans, every relationship. It's all in that chat.db file. And it's also exactly what makes an AI assistant genuinely useful. When did I last talk to Joe, and what did we decide? What did Marissa and I settle on for the trip dates? Catch me up on the Reno group chat. What am I actually on the hook for? I'm going to be honest. The idea of uploading my entire message history to some cloud AI makes my skin crawl. That's a decade of my most private conversations. And that's the whole point. AI Message refuses that trade entirely. It indexes locally. Embeddings run on device via ONAX. No API calls to an embedding service. The message's database is opened with SIL ET, Open Recob plus Pragma query only, enforced at the SQL level and backed by tests. No telemetry whatsoever. The only network access in the default configuration is a one-time download of the public embedding model weights. And that contains none of your data. You can even skip that with ETL mash no embed. And it resolves contact names from the local macOS contact store, right? So search results say Alice Smith instead of plus 1916. Yeah. Read only access to contacts. Names never leave the machine. You can disable it entirely if you want. And the tool surfaces deliberately minimal. Search messages with hybrid keyword plus semantic retrieval. Get recent messages for the chronological tail. Get conversation to expand a search hit with surrounding context. And list chats by recency. Four, read only tools. The server can never write. It can only ever see the local index it built. So you basically built the thing that Apple should have built years ago. Two things actually. Right. And that's what I want to dig into. Why hasn't Apple done this? Because on the surface, it's baffling. They control the hardware, the OS, the messages database, the iCloud infrastructure, the entire vertical stack. They've been shipping neural engines since the A11 Bionic in 2017. They have on-device ML accelerators that are genuinely impressive. They've had 15 years of Siri data to learn from. So what's the holdup? Is it just Apple being Apple, slow, controlling, can't ship software anymore? Here's my theory. And it's actually more nuanced than Apple is slow. Apple has roughly 2 billion active devices. 2 billion. If they ship a Siri that reads your iMessages and your iCloud documents, and it hallucinates, and it will hallucinate because every model does, they have a PR catastrophe on a scale that makes AntennaGate look like a rounding error. Imagine Siri confidently telling someone the wrong medical information from their documents or misreading a legal contract. Or, and this is the nightmare scenario, fabricating a text message that never existed and attributing it to a real person. Siri told me my husband said X and he never did. Apple cannot afford that at the scale of 2 billion devices. So it's fundamentally a liability problem, not a capability problem. They could build it. They're scared to ship it. That's part one. Part two. Most of those 2 billion devices simply can't run this stack locally. You need a Mac with enough unified memory to run a capable model. The iPhone in your pocket. Even the latest Pro. It's getting close, but running OCR transcription embeddings and a capable LLM simultaneously while maintaining reasonable battery life? That's genuinely hard. Phones are thermally constrained in ways laptops aren't. So they'd have to do at least some of it in the cloud, which breaks the privacy story they've been selling for over a decade. Exactly. And Apple's privacy positioning is arguably their most valuable brand asset. They've spent billions marketing it. What happens on your iPhone stays on your iPhone. They put it on billboards. If they suddenly say, actually, we're going to need to send your messages and documents to our servers so Siri can read them. That's not just a product problem. That's a brand destroying reversal. The same people who pay the Apple premium specifically for privacy would revolt. And even their private cloud compute thing, which is genuinely impressive from a security engineering standpoint, still means your data leaves your device. The cryptographic attestation is cool, but most users don't understand or trust it. Right. PCC is a real technical achievement. Anonymous stateless computation on trusted hardware with verifiable transparency logs. It's good engineering, but it's still cloud. And for the kind of deeply personal data we're talking about, your messages, your financial documents, your medical records, cloud is a fundamentally different trust model than local only. So they're stuck in a double bind. Ship local and it doesn't work on most devices. Ship cloud and betray the privacy promise. Ship nothing and fall behind while Android and third party tools eat their lunch. Yep. But here's the optimistic part. I think they're going to do it anyway and relatively soon. What makes you say that? Because I've been burned by Apple AI promises before. Two things. One, Apple's whole AI strategy is already pushing relentlessly toward on-device. They've been quietly building the infrastructure for years. The neural engine getting more capable every generation. MLX as a first class ML framework for Apple Silicon. The on-device foundation models they've published papers about. The hardware is catching up fast. An M4 Mac with 32 gigs of unified memory can run Gemma 412B comfortably. That's a genuinely capable model that can do document understanding and structured extraction. And two, the competitive pressure is becoming existential. Google is shipping Gemini powered features on Pixel phones that can reason about your photos and messages. Samsung has Galaxy AI. If you can buy a $500 Android phone that has a genuinely useful AI assistant that knows your documents and messages and your $1,200 iPhone can't do any of it, the privacy argument stops being enough. At some point, privacy can't be the excuse for not shipping features. Yeah, I think you're right about the pressure. And honestly, your projects kind of prove the point. If one developer can build this in Rust over a few weeks, Apple can definitely do it with their resources. The building blocks are all off the shelf at this point. Well, to be fair, one developer is doing a lot of heavy lifting there. And I'm standing on the shoulders of open source giants, whisper.cpp, own NX, the entire Rust ecosystem. But the point stands. MCP, local models, vector search, OCR, transcription. They're all mature enough to stitch together. The hard part isn't the technology anymore. The hard part is the organizational will to ship something that might hallucinate in ways that make headlines. And the liability framework. I don't envy the Apple lawyer who has to sign off on Siri can now read your legal documents and answer questions about them. That's the meeting where someone puts their head in their hands. No, that's a rough meeting. But here's the thing. They don't have to launch it as Siri knows everything about your life. They could start with a much narrower surface. Siri can search your documents for dates, names, and amounts. Factual retrieval with citations, not generative reasoning. Narrow the blast radius. Build trust incrementally the way they did with Apple Pay. Start small, prove it works. Expand the surface over years. Yeah, the crawl, walk, run approach. All right, let's switch gears. You mentioned another topic you wanted to hit. Something about Jarvis and you seemed less than thrilled about it. Oh, man. Paul, have you seen this obnoxious bullshit that's been making the rounds? Which obnoxious bullshit? You're going to have to be more specific. It's been a rich week for obnoxious bullshit in AI. Fair. I'm talking about Jarvis. Use Jarvis.dev. This thing positioning itself as a revolutionary AI agent that operates software alongside humans with a little animated pebble cursor. It's been all over social media with very slick demo videos. Book a flight, edit a video, sort your desktop. All very impressive in a 30-second clip. Oh, yeah. I've seen the demo. The pebble cursor thing. Claims it has two cursors on screen at the same time. Yours and the AI's. Very Iron Man. What's the actual story under the hood? I looked into it so you don't have to. Here's what it actually is. Jarvis is essentially a UI skin and a text-to-speech layer. Kokoro, by the way, wrapped around Hermes Agent, which is a genuinely good open-source agentic harness from news research. Hermes does the actual work, the agent loop, the memory system, the skills architecture, the tool dispatch. Jarvis adds a cursor animation, a voice, and a subscription fee. Wait, so they're charging $9.99 a month to put an animated cursor on top of an open-source project? One that you can install yourself for free? Yes. And here's what really gets me. They market it like they invented the entire concept. They have a manifesto, a landing page that frames it as the native path versus the MCP path, and positions screen control as some kind of philosophical breakthrough. But the agent core, Hermes. The persistent memory, Hermes. The skill system, Hermes. The multi-platform gateway, Hermes. They didn't build the engine. They wrapped the engine in a pretty UI and wrote a manifesto about it. So it's the classic AI wrapper grift. Take something open-source, add a pretty face, charge a subscription, write a manifesto, and hope nobody checks the dependencies. And the thing is, Hermes is genuinely good. It's from Noose Research. It's fully open-source. It has persistent memory that actually works. A skill system where you define capabilities as markdown files. Messaging gateway support across Discord and Telegram and WhatsApp. It can connect to email, write code, manage your calendar. If you want an actual agentic harness that functions as a personal assistant, just use Hermes directly. It's free. Forget the Jarvis wrapper with its $9.99 cursor animation. What I genuinely don't understand is why does this keep working? Why do people keep falling for wrappers? We've seen this pattern a hundred times now. Because most people don't want to configure a terminal. They don't want to edit YAML config files. They don't want to understand MCP servers and model endpoints and bearer tokens. They want a thing that looks polished and just works. And there's a completely legitimate market for that. I'm not opposed to charging for convenience. The problem is when the wrapper pretends to be the innovation. When the marketing obscures the actual open-source project doing all the real work. Right. Build a better UX on top of Hermes. Charge for the convenience. I have zero problem with that. But be honest about what you are. Put Powered by Hermes on the landing page. Don't write a manifesto like you've discovered a new paradigm for human-computer interaction. Don't write a manifesto. That should genuinely be a t-shirt. But there's something else that bothers me about Jarvis specifically. The screen control approach they're positioning as superior to MCP. They frame it as the native path versus the MCP path like structured APIs or somehow legacy. But screen control is fragile as hell. It relies on pixel coordinates and ARIA labels. It breaks on every UI refresh. It's slow. It's the kind of thing that looks amazing in a polished demo and falls apart the moment a website updates its CSS. Yeah, I've actually tried screen control agents. Multiple ones. They're incredible for about five minutes. You feel like Tony Stark. Then you realize you can't use your computer while it's working because it's hijacking your cursor. And half the time it clicks the wrong button because a div moved three pixels. Or the page loaded differently this time. Or there's a cookie banner it didn't expect. Exactly. And Hermes, the actual project underneath, is MCP native. It speaks the protocol that every serious AI tool is converging on. It has structured reliable APIs for tool calls. Why would you take something that has a clean, programmatic interface and wrap it in a screen-scraping cursor that breaks constantly? It's like putting a horse in front of a car and calling it a transportation innovation. The car already works. The horse is not an upgrade. The horse is not an upgrade. I'm stealing that. All right, I think we've adequately slammed Jarvis. What's the takeaway for someone listening who's actually interested in this agent space? If you want an agentic harness that's actually useful, one that can connect to your email, write code, manage your calendar, operate as a genuine personal assistant, look at Hermes directly. It's open source. It's from Noose Research. It has a real community-building skills and contributions. Or look at OpenClaw if you prefer a different architecture. Both are genuine projects with actual engineering behind them. Don't pay 10 euros a month for a cursor animation and a manifesto. And for what it's worth, both Hermes and OpenClaw have MCP support, so you can plug Tim's iCloud and iMessage tools right into them. The ecosystem actually connects. Yeah. And that's the bigger point. These tools are designed to work together. AI Cloud and AI Message expose your data over MCP. Hermes or OpenClaw provides the agent layer. A local model does the reasoning. Skewilite holds the index. The whole stack runs on hardware you own. You don't need a cloud subscription for any of it. Speaking of AceQ Lite, you had a hot tip for AI builders, right? Something you wanted to get on the record? Yes. This is my actual hot tip. And I want developers to hear this. If you're building local AI utilities, proofs of concept, MCP servers, or anything that needs persistent storage, reach for Skewilite, not JSON files, not Markdown files. Not some trendy vector database with venture capital pricing, Skewilite. Skewilite, the most boring, most reliable, most battle-tested technology in the entire stack. It's been around since 2000, and it's probably going to outlive both of us. And that's exactly why it's perfect for this. It's been in continuous development for over 25 years. The test suite is the stuff of legend. The AceQualite team maintains something like 600 times more test code than implementation code. It runs on literally everything. It's public domain. And now, critically for AI work, it has vector support through extensions like Skewilite Vect. Yeah, the vector extension changes the game. You get semantic search without standing up a separate vector database. No pine cone, no Weave8. No cloud service that charges per query and owns your embeddings. Right, and for local RAG applications, that's the killer feature. Let me give you a concrete use case. Say you want to build a local RAG MCP server. You want an agent on your computer to be able to reason about PDFs and images and documents without sending anything to the cloud. No open AI embeddings API. No pine cone. Nothing. You build a SQLite database. You vectorize your PDF documents and image descriptions using a local embedding model, something like BGE small n is B1.5 running through fast embed or on an X. You store the file metadata right alongside the vectors, file name, path, date, modified tags in the same database. Then you plug that into any AI agent over MCP. With very minimal tokens, you can answer questions like, when was the last picture I took on vacation and where was I? Or show me every receipt from Costco in the last six months. You have some extremely specific Costco tracking needs, my friend. Look, spaghetti inventory is a real problem and I will not apologize for wanting to solve it. But the broader point is, SQLite is the perfect foundation for local AI. It's self-contained in a single file. It's infinitely portable. It supports FTS5 for full text keyword search. It supports vectors for semantic search. And you get hybrid retrieval, keyword plus meaning fused with reciprocal rank fusion, all in one query against one database file on your disk. And honestly, it's so much better than what most developers reach for first, dumping everything into JSON or markdown files and then trying to grip through them or load them all into memory. I've done that. It works for about a week. We've all done that. I'll just store it in JSON. It's simple. It's readable. What could go wrong? And then six months later, you have 400 JSON files, no indexing, no querying. And you're writing increasingly elaborate node scripts just to find something. Skewilite from day one. Future you will thank, present you. Future you will thank, present you. I feel like I say that every episode, but it's genuinely true here. And the vector thing is what really makes it interesting for AI work. You can build a fully local semantic search engine in like 50 lines of code. Embed your documents, store the vectors, query by similarity. It's not complicated anymore. And that's exactly what both iCloud and IMSage do under the hood. They use Eskulite with FTS5 for keyword search plus raw float 32 vectors for semantic search, fused with reciprocal rank fusion for hybrid retrieval. It's fast, it's reliable, and it's completely local. No pine cone, no WeV8, no cloud vector database that charges you per query and stores your embeddings on someone else's hardware. Just a single bot skeet file in your application support directory with owner-only permissions. You know what I really love about this whole approach? It brings the entire AI stack back to something a single developer can understand and control. You don't need a distributed system to do semantic search. You don't need a Kubernetes cluster to run an agent. A capable Mac, some open source tools, and SQLite. That's the whole stack. You can hold the whole thing in your head. That's the moral of this whole episode, honestly. You don't need to wait for Apple to figure out their Siri strategy. You don't need to pay 10 euros a month for a wrapper that adds nothing but animation. You don't need a cloud vector database or an embeddings API. A capable Mac, some thoughtfully designed open source tools, and SQLite. You can build the AI assistant you actually want right now, today. The infrastructure is ready. And the privacy story is real. It's not marketing copy. It's not a promise in a terms of service that changes next quarter. When everything runs on your machine, you can verify, actually verify, that nothing is being sent anywhere. That's not a trust relationship with a corporation. That's an architectural property. If the server only binds to loop back, it physically cannot send data to the internet. So to bring it all together, iCloud and IMSG are available right now, open source, dual license, MIT and Apache 2.0. Brew install Tim Williams slash iCloud and you're up and running. The readmes are thorough. There are over 140 tests in AI Cloud alone. There's a setup wizard that walks you through everything. What to index, which model to use, your privacy boundaries. It's real, production quality software, not a demo. I built these because I wanted them to exist. And now they do. And the ecosystem is actually coming together around MCP. You can use these with Claude Desktop, with Codex, with LM Studio running a local model, with Hermes or OpenClau as the agent layer. It's not a bunch of isolated tools anymore. They actually compose. Yeah. And that's the thing I'm genuinely optimistic about. Not the hype cycles. Not the fundraising announcements. Not the manifestos. The quiet infrastructure being built by people who just want their computers to be more useful and their data to stay their own. That's the stuff that's going to last. That's the stuff that's going to matter. Well said. And if anyone builds that spaghetti tracker with CQLite, you know where to find us. We want to see the repo. I'm absolutely not kidding about the spaghetti. Someone build it. All right. That's our show. Links to AI Cloud, AI Message, Hermes, and everything we talked about are in the show notes. Paul, always a pleasure. Same here. Thanks for having me. See you next time. Here's looking at you. Thank you. Thank you. Thank you.