← Back to search

Local AI Gets Serious on the M5 Ultra Mac Studio

AppStories · 2026-09-21 · 39 min
relevance 84 5952 words Episode page ↗ Audio ↗
Show full episode description
This week on AppStories, John interviews Federico about his review of the M5 Ultra Mac Studio and it local AI capabilities. On AppStories+, Federico takes listeners behind the scenes on how he tested the M5 Ultra Mac Studio. Also available on YouTube here . Sponsored By: Workbench : Remote desktop for AI agents and headless Mac minis. Free to use for 20 minutes/day. Claude – For problems worth solving — get started with Claude today. Links and Show Notes Running Local AI Models on Apple's Latest Macs M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents Apple Hardware Mac Studio Mac mini Studio Display XDR Local Models Qwen3.8-Flash-Next Qwen3.8-27B DeepSeek-V4-Flash GLM-5.3-Flash Cohere Transcribe Qwen3-TTS Qwen-Image-2.1 Local Inference and Clustering oMLX exo Agent Apps and Tools Hermes Agent Hermes Agent on GitHub Open Minis Cadu OpenAI Brings Computer Use to Codex App on Windows Frontier Models and Workflow GPT-6 Astra Claude Fable 5.1 Claude Projects Artificial Analysis Intelligence Index Elyx NVIDIA PC Comparison GeForce RTX 5090 Lian Li A3-mATX NVIDIA DGX Spark AppStories+ Post-Show: A preview of how Federico used local AI models to support his iOS and iPadOS 27 review. Leave Feedback for John and Federico AppStories Feedback Form Follow us on Social Media Federico is @viticci everywhere John is @johnvoorhees everywhere Affiliate Linking Policy
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Running always-on local AI agents on Mac hardware without burning API costs, subscriptions, or rate limits.
Benefits
  • Roughly 2x prompt-processing speed on M5 Pro vs M4 Pro
  • 50% higher memory bandwidth on M5 Ultra (1.2 TB/s vs 819 GB/s)
  • Neural accelerator in each of the 80 GPU cores boosts token pre-fill
  • Local models avoid API costs, rate limits, and keep documents private
  • Agents run continuously, limited only by electricity and compute
Use cases
  • Ran DeepSeek V4 Flash locally on M3 Ultra Mac Studio powering iOS 27 review agents for 99 days straight
  • Agents fired every five minutes to ingest clipped research, reconcile with a Notion database, and cross-reference review chapters
  • Benchmarked Qwen3 8/27B: 60 tok/s read on M4 Pro vs 150 tok/s on M5 Pro; 16K prompts read at 91 vs 400 tok/s
  • M4 Pro Mac Mini used as outboard processor for video/audio and AI, offloading an M1 Max Mac Studio
KPIs / results
  • Prompt read speed: 60 vs 150 tok/s (short), 91 vs 400 tok/s (16K prompts)
  • Memory bandwidth: 819 GB/s (M3 Ultra) to 1.2 TB/s (M5 Ultra), +50%
  • Agents ran continuously for 99 days on M3 Ultra
  • M5 Ultra: 80-core GPU, 256 GB RAM review unit; 512 GB model due late October
Tools / build
  • DeepSeek V4 Flash local agent pipeline on Mac Studio
  • Qwen3 27B local benchmarks
  • Notion research database with agent cross-referencing
  • Astropad Workbench remote desktop for headless Mac minis
0:00 / 0:00
[SPEAKER_00] Hello and welcome back to another special episode of AppStories. Today's episode is brought to you by Astropad Workbench and Claude. I'm John Voorhees and this is Federico Viticci who just has another big review out. How are you doing, Teach? [SPEAKER_01] Yeah, hi. [SPEAKER_00] You just know there's no sleep for you, man. [SPEAKER_01] Nope. Yeah, as we'll see in this episode, I jumped straight from the iOS and iPadOS 27 review onto another one. I had a one day break in between them. All right. Okay, so let's recap. What are we talking about? [SPEAKER_00] We're talking about two pieces of hardware. We're going to talk about the latest Apple Mac Mini and Mac Studio. All right, we're going to talk about those things during the main show and then during AppStories Plus for Club Mac Stories Premier members and AppStories Plus subscribers. We're going to talk a little bit about some behind the scenes stuff, which we'll get into during the post show. [SPEAKER_00] But, you know, I think of these two computers having read your review, the one that obviously is the showcase here because it's capable of the most is the Mac Studio. But I don't want to give the Mac Mini short shrift, Federico. I want to hear a little bit about the Mac Mini first because, I mean, it looks the same, but it does do more, right? [SPEAKER_01] Yeah, so to recap, Apple sent me an M5 Pro Mac Mini with 64 gigs of RAM and obviously the start of the show, the M5 Ultra Mac Studio with 256 gigs of RAM since the 512 gig version is not available yet. It'll come out in late October. I don't know actually if anybody else got review units for the 512, but I didn't have it. I only had the 256. [SPEAKER_01] Now, my article is about the, so there's an article on Mac Stories about the M5 Ultra Mac Studio. Most of my testing was focused on the M5 Ultra Mac Studio. I did run some tests on the M5 Pro Mac Mini and I do plan on following up on my coverage of the Mac Mini. I actually do plan on rethinking my desk around this M5 Pro Mac Mini review unit. I just wanted to focus on the Mac Studio. But to give you some numbers, I compared the Mac Mini. [SPEAKER_01] Obviously, both of all of my tests and the article on the site were focused on local AI. I think that's why Apple sent me these review units. I'm not a video editor. I'm not an app developer, but I do tinker a lot with local AI models. [SPEAKER_00] You're no Dr. Dre out there producing records. [SPEAKER_01] No, I'm not doing Final Cut. I'm not doing video processing pipelines, that kind of stuff. My coverage here and on the site is focused on local AI performance. So this M5 Pro Mac Mini, I compared it to my M4 Pro MacBook Pro that I'm using right now here with 48 gigs of RAM. That was the only M4 Pro machine I had. So it's less RAM, but it's the same chipset. [SPEAKER_01] To give you some numbers, I tested QN 3.8 27 billion. So very popular local model from this year on both the M5 Pro and the M4 Pro. And on the M5, we're basically looking at a double of prompt processing. [SPEAKER_01] So how quickly can the chip read a prompt before answering? That is called token pre-fill. Like how quickly can the machine ingest the entire token before decoding? So before actually generating a response token by token. [SPEAKER_01] And on the M4 Pro, I ran a bunch of tests with short pros, 4K token prompts, as well as 16K token prompts. To give you some numbers, I got, for example, 60 token per second read on the M4 Pro versus 150 token per second read on the M5 Pro. [SPEAKER_01] On 16K, so 16,000 tokens in the context window, my M4 Pro MacBook Pro read the prompt at 90, 91 token per second. The M5 Pro Mac Mini read it at 400. So like, it's an incredible jump from a portable M4 Pro. Obviously, like this point of comparison is not, it's not like the different machines. [SPEAKER_01] But realistically, we're looking at quite a bit of a, quite a bit of a bump. [SPEAKER_00] Yeah, I can add a little context to this too, because, because I've been running an M4 Pro Mac Mini, which is pretty, it sounds like the same specs as your laptop, because it has 48 gigs of RAM too. And these Pro Minis are really, really capable machines. What I've done, and you know, you say you're thinking about resetting up your desk. [SPEAKER_00] I actually have the M4 Pro Mac Mini piggybacking on my old M1 Macs Mac Studio, because it has all the, you know, encoding, decoding chipsets inside the Mini. And now the Mini has become my outboard place for processing a lot of video and audio that I work on, as well as doing some AI stuff myself, [SPEAKER_00] because it takes the load off the Mac Studio, which has become kind of like the front end that doesn't do as much of the heavy lifting anymore, and instead outsources it to the Mac Mini, which is kind of remarkable, given that an M1 Macs at the time was, you know, other than the Ultra, was the top of the line Mac Studio that you could get in 2022, I guess it was now, or 2023. [SPEAKER_00] This special episode of App Stories is brought to you by Astropad Workbench, the remote desktop for AI agents and headless Mac Minis. Workbench is the first remote desktop app built specifically for AI agents. If you have an agent working on your Mac, Workbench lets you connect remotely to see how they're doing, spot issues, make changes, and restart automated jobs without being tethered to your desk. That's super useful if you're running agents on a headless Mac Mini. [SPEAKER_00] Workbench handles remote access, external connectivity, and display in a single app. There's no port forwarding, no VPN, and no router configuration. Just install Workbench on your Mac and on the device you want to connect from, and you're ready to go. My favorite feature of Workbench is the unified virtual display. This can combine multiple Mac displays into a single screen that matches the resolution of your remote device. [SPEAKER_00] For a headless Mac Mini, your iPhone, iPad, or another Mac becomes its display. That makes it easy to check on a running agent from your iPhone while you're out, see the actual desktop, and send prompts to keep things moving. Voice input lets you dictate to your iPhone or iPad to fill text fields trigger inputs on your Mac so you don't have to type everything on a small screen. [SPEAKER_00] Workbench also manages your Mac sleep settings with an intelligent sleep feature, keeping your Mac awake and reachable when you need it. Workbench is a native app on all three platforms, and the streaming technology behind Workbench is Astropad's proprietary multi-codec engine Liquid, the same engine that powers Lunar Display and Astropad Studio. It delivers what Astropad describes as perceptually lossless video with retina support. [SPEAKER_00] Compared to VNCs and screen sharing, Astropad emphasizes crisp text, accurate color, low latency, even over a cellular connection. And Astropad has shared a preview with us. Terminal support is coming later this year, and AppStories listeners are the first to hear about it. You'll be able to open a terminal window on your remote Mac and run commands directly from your iPhone or iPad. [SPEAKER_00] That means checking logs, tweaking a script, or restarting a process without switching to a separate SSH app. You can use Workbench free for 20 minutes per day or upgrade to unlimited access for $79 a year. Just visit astropad.com slash app stories. To get started, that's astropad.com slash app stories. Our thanks to Astropad Workbench for the support of this special episode of AppStories. [SPEAKER_00] Okay, Federico. Well, I think we need to start talking about the Mac Studio because that's really, I think, the machine that is a step change, at least from the perspective of somebody who wants to do local AI from reading your story. What did you, I guess, recap the specs and then tell me a little bit about what you tested? [SPEAKER_01] Right. So you're looking at a brand new GPU with, it's an 80-core GPU in this machine, M5 Ultra, 256 gigs of RAM, 4 terabytes of storage in this unit. And you're looking at this new GPU with 80 cores. And inside each core, there's a neural accelerator with a new architecture. So this is something that Apple started doing, as we saw before earlier in the M5 generation, putting a neural accelerator inside each GPU core. [SPEAKER_01] So you got 80 of those in the M5 Ultra in addition to the new GPU, which, you know, on average, Apple says, should perform like up to four times, you know, better in terms of compute than the M3 Ultra, like nine times better or something than the M1 Ultra. So I think sort of the angle of my story was I've been using the Mac Studio as a local server for all kinds of things for the past year. [SPEAKER_01] But especially this summer, I went all in with local AI on the Mac Studio because I had this big project, the iOS 27 review that I structured in such a way that required an agent to be running at all times, meaning every five minutes, an agent would actually fire up and do some processing. Multiple of those agents with different schedules. So effectively, it was always running something. [SPEAKER_01] In that case, the M3 Ultra Mac Studio was running DeepSeq V4 Flash. And I structured it in a way that I had to be using a local model. First, because I didn't want to burn through my subscriptions. I didn't want to run these API costs. [SPEAKER_01] And I just preferred the thought of having these documents and especially the chapters of my iOS review being fully self-contained in the Mac Studio. Because, you know, you never know. And so I saw with the iOS 27 review this summer, the benefit of being able to have this agent with, you know, [SPEAKER_01] with the kind of performance that would have been considered frontier performance like six to nine months ago, running entirely locally on a Mac. [SPEAKER_00] I'm going to say, I think one thing that also is a factor here that is worth mentioning when you talk about burning through your tokens and subscriptions is that because you're running it locally, you were able to do things like have your local agents monitoring your research on like practically a minute by minute basis. So that when you added. Right. So that when you added new research, it was immediately ingested and organized and categorized and all that stuff. [SPEAKER_00] Because you do that with a subscription, you're going to burn, you know, 24 hours a day. You're going to burn through it pretty darn fast. [SPEAKER_01] Yeah. And even if you use GPT Luna, which, you know, it's very cheap model. It's very cheap, yeah. But the moment you start running it every minute, you know, I just didn't like the thought of that. And, you know, you get API rate limited, like all those kinds of things. Whereas with a local model, you're just limited by your own electricity and your own compute. But on the M3 Ultra, it was possible and it worked for 99 days. This occurred almost 100 days. [SPEAKER_01] And I have the screenshots in the story. 99 days, these agents ran essentially all the time on the Mac Studio. But they were slow. They worked. But if I clipped something, I knew that the agent would wake up and find it, right? [SPEAKER_01] For example, if I saved an article from, you know, about a certain feature of iOS 27 that needed to be reconciled with my Notion database and also cross-referenced with my chapters and my existing list of features. That would wake, the agent would wake up instantly. But it would take like 10 minutes to categorize it, to run the whole thing. [SPEAKER_01] And that was because on the M3 Ultra, you were limited by token pre-fill speeds and text generation speeds. So like that's, and those two elements, right? How quickly can an agent process a prompt? And how quickly can an agent start generating a response? And especially when you're dealing with agents, right, in an agentic loop. So there's multi-turn tasks. [SPEAKER_01] And the agent goes back and forth in the loop and needs to reread, you know, the thinking steps from before. And so the conversation grows longer. You start getting that performance decay, right? Where the conversation grows slower over time and the context window gets larger. And so you are limited by those two elements, the nature of the GPU, so the GPU itself, and the memory bandwidth of the GPU. [SPEAKER_01] So with the M5 Ultra, this is changing fundamentally because of the better GPU and because of the increased memory bandwidth. Now, I want to make sure that I get my numbers correct. Because in terms of memory bandwidth, we are jumping from 819 gigabit per second to 1.2 terabyte per second. [SPEAKER_01] So it's a 50% higher memory bandwidth in the M5 Ultra compared to the M3 Ultra. That means, you know, you get those, you get essentially the, when it comes to generation, the model is able to, you know, pull out from its context window. And from, you know, when it's writing token by token, it's able to go much faster. [SPEAKER_01] But the new architecture of the GPU also means that I saw some incredible jumps in the token pre-fill stage. So how quickly can the model start, how quickly can the model read a prompt and go from essentially zero to the first token that it generates? [SPEAKER_01] And so in this review, by the way, you will find, in addition to my classic article, you will find this section, this long section with tons of animated visualizations as well as static charts. I really wanted to find a way to convey these things. [SPEAKER_01] And you will find all of these examples in interactive animations and charts that you can click through, pick different models, pick different prompts, and see exactly the comparison between the M3 Ultra, the M5 Ultra, as well as my NVIDIA 1590, which is also part of the benchmarks. But we can talk about the benchmarks and the testing environments later. [SPEAKER_00] Okay. [SPEAKER_01] Anyway, the token pre-fill on a 4K prompt, for example, the M3 Ultra could read at 1,100 tokens per second. [SPEAKER_00] Okay. [SPEAKER_01] So 1,100 tokens per second. The M5 Ultra, so 1,100, we're jumping to 2,700. Oh, wow. On a 16K prompt, okay? So imagine like big system prompt, lots of tools, lots of skills. Maybe you're giving it, I don't know, a chapter, a section of a review that you're writing, right? We're jumping from 1,100 again. So these numbers, they stay pretty stable. [SPEAKER_01] But with the GPU of the M5 Ultra, you do actually get a performance bump in the pre-fill stage when the prompt goes longer. And so we go from 1,100 to 2,800. Okay. Okay. So, and mind you, that's the pre-fill. When it comes to generation at that size, the M3 Ultra was writing at 37 tokens per second. The M5 is writing at 52 tokens per second. Okay. Very nice. [SPEAKER_01] And so, long story short, and there's plenty of numbers and plenty of things that you can read on Mac Stories. I highly, highly recommend, even if you don't want to read my stuff because maybe you don't like what I say, but the numbers, they speak for themselves. They do. And so I urge you to scroll down and take a look at the numbers. [SPEAKER_00] Yeah, and I think it's worth looking at these visualizations because they're not just all bar charts. There are animated bar charts and things like that, but they're also, you have some really neat animations that kind of give you a better sense for what it's like if you were running these things yourself, which I think is a good way to look at it. [SPEAKER_01] Yeah, I wanted to find a way to visualize like the time to first token or what it actually means to write at, I don't know, 50 tokens per second versus 110. The thing for me is that this machine lets me run QAM 3.8 Flash Next, which is a preview of QAM 4. [SPEAKER_01] It's based on a new architecture where it's a hybrid architecture where the model has 125 billion parameters. So it's a medium model. It's not a small model. It's a medium to large model. But in addition to that, it has 51 extra billion parameters stored in a so-called NGRAM table. And that's a very complex conversation that's actually beyond my knowledge. I read a lot about it. [SPEAKER_01] The idea is that it's like a lookup table with additional parameters that the model can use to look up when and where to activate certain parameters at runtime. And so the combination of these parameters that are in the actual model plus the extra lookup table, the model is able to activate only six or seven billion parameters depending on your request. [SPEAKER_01] And so it can go really fast, but it requires storage and it requires RAM. With this machine, I have gone all in with Flash Next as my local AI model. Interesting. It's very nice to use in things like Hermes Agent or Open Minis on iOS. I even set it up as a native local sub-agent in Codex because OpenAI doesn't care. [SPEAKER_01] They let you add a local model provider as a native model in Codex. You actually see it in the model picker in the Codex app on the Mac. And when I set reasoning to like even extra high and the model is loaded in OMLX, which is the backend that I've been using, I get like a time to first token of like three seconds. [SPEAKER_00] That's amazing. [SPEAKER_01] The model can start obviously at 100, 110 TPS generation at the very beginning of a long session. But then to give you an example, John, last night I worked on this local web app exclusively using Open Minis on iOS, talking to the M5 Ultra for two hours. It put together this 3D voxel coliseum example. [SPEAKER_00] Oh, yeah, yeah. [SPEAKER_01] And even when the context started getting to the limit of the model, like 200,000 tokens, it was still going at like 80 tokens per second. [SPEAKER_00] That's amazing. That's amazing. What's the quantization on the... [SPEAKER_01] So interestingly enough, most of my tests I ran at 4-bit. It's based on this new hybrid quantization technique for OMLX called OQ. But actually, the M5 Ultra with 256 gigs of RAM can also hold entirely in RAM the 5-bit version. I also tested for completionist's sake the 6-bit and the 8-bit versions. Okay. [SPEAKER_00] How'd that go? [SPEAKER_01] Now, the M5 Ultra can hold both 6-bit and 8-bit entirely in RAM because of the 5-12. You got a lot more. [SPEAKER_00] You got double the RAM on that machine. [SPEAKER_01] But thanks to the new Ngram table architecture of FlashNext, and thanks to OMLX's support for offloading Ngram tables to an SSD, I can actually use 6-bit and 8-bit on the M5 Ultra that I have right now if I want to. That's amazing. Just that part of the model is loaded in RAM and the rest goes offloaded to SSD. [SPEAKER_00] Right. [SPEAKER_01] But I think as a daily driver, I'm just going to use the 5-bit. Yeah. [SPEAKER_00] Is it faster to do the 5-bit then, do you think, than to offload some of the stuff to the SSD? [SPEAKER_01] Yeah. I'm in practice. I just need to be very careful because, like, I'm used to loading, like, a couple of models in memory on the M3 Ultra. With this model, the 256, I need to measure my performance just right. Because, for example, when I load FlashNext and it takes 110 gigs of RAM. [SPEAKER_00] Right. [SPEAKER_01] If I also want, which is something that I don't usually do, but for the sake of this review, I also wanted to test local image generation, right? Oh, sure. When Image 2.1 launched last night and I wanted to test it on the M5 Ultra. And when you generate high-resolution images, like 2.5K image resolution, that also becomes, like, 70 gigs of RAM. [SPEAKER_01] And so 110 plus 70 to 80, you're approaching 200 and you're right there at the limit because the macOS is going to kill your process. Because it's going to say, well, I need the rest to run the OS and your open apps. Right. So that will need more testing whether I want to go with 4-bit or 5-bit or 6-bit offloaded to the SSD in the future. [SPEAKER_00] Yeah. It'll be interesting to see what people are able to do with the 512 gigabyte model because that opens up a lot more room to do. You'd easily do the 8-bit and then be able to run another model at the same time, I would think. Yeah. [SPEAKER_01] I really hope that Apple allows me to test the 512 version because I have some ideas. Like, I didn't even have time to set up ExoLabs so you can daisy chain multiple Thunderbolt 5 computers, right? And use the RDMA, which is an Apple framework. I've been looking at that a lot lately. Essentially split, distribute inference across multiple Macs over Thunderbolt 5 because of the crazy bandwidth that Thunderbolt 5 has. [SPEAKER_01] And in theory, like, I could imagine a scenario where combining an M3 Ultra, an M5 Ultra with 256 and an M5 Ultra with 512. So you're allocating one and a half terabyte, essentially, of RAM to running something like, I don't know, Kimi K3 quantized to 4-bit or Qen 3.8 Max also quantized to 4-bit. Like, with that amount of RAM, those are like 1.2 trillion parameter models, right? So they're giant models. [SPEAKER_01] But with one and a half terabyte of RAM over Thunderbolt 5 and the right quant, it probably becomes possible. So that this conversation doesn't get too technical. I mean, it already is quite technical. In practice, in practice, I have switched my Hermes agent using the excellent upcoming Cadu iOS app. That's a very nice app. Which I highly recommend. And Open Minis, which is my beloved iOS agent. [SPEAKER_01] Both of those apps I am exclusively using with a local model running on a local server on the M5 Ultra. Flashnext, extra high in both Open Minis and Hermes agent. They feel very nice to use. [SPEAKER_00] Yeah, just in terms of like the speed and the quality of the... [SPEAKER_01] Speed, the quality of the model, right? I mean, it's a great model. It's rated on a 40 intelligence index on artificial analysis. For context, Astra and Fable are like 51 and 53. So a 40 artificial analysis index value was like Frontier in January. Right. To give you some context. And this is a model that you can run on your desk. I mean, it's literally running on my desk. [SPEAKER_00] Like maybe comparable to Opus 4.6 or something like that. Yeah. [SPEAKER_01] It's like you have Opus 4.6 level on your desk. Yeah, right. And it's speedy. That was the thing that always killed me with local agents before. They were fine for the iOS 27 review because I knew they were slow, but they would just work in the background. [SPEAKER_00] And you didn't need that stuff right away, right? It's a different kind of use case. [SPEAKER_01] As long as it happened in the background, I was not looking at the model, right? Right. When I'm using Open Minis, when I'm using Hermes Agent, I am looking at my phone and I'm waiting for an answer, right? And so speed becomes important. And now with this computer, that kind of personal agentic setup is possible. I should also mention, I have been loading other models on the M5 Ultra. [SPEAKER_00] Okay. [SPEAKER_01] Cohere Transcribe, excellent transcription as a voice to speech to text model. Running locally transcribes text in like five seconds. It's like super fast. I have been running QAN 3 TTS. So this is a voice model. I've actually been doing some very early tests because this beta of KADU came out last night. But it is possible to use a text model. [SPEAKER_01] So Flash Next, Inner Miss Agent with a live voice model on top. And so I tested it. It's possible. I can use the M5 Ultra with Flash Next as the underlying model doing the work and generating responses. And QAN 3 TTS with a custom voice also running on the M5 Ultra speaking the responses in real time. Right. [SPEAKER_00] What's the delay like there? Can you tell from looking at the... [SPEAKER_01] There's a bit of a delay. It's not like obviously... And it's also not a bi-directional voice model. Right. So it's not like GPT Live. You cannot interrupt the model. You're just talking to it. But it's a turn by turn conversation. [SPEAKER_00] That's interesting. That's interesting. This episode of App Stories is brought to you by Claude from Anthropic. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. [SPEAKER_00] I've been using Claude for months to help me rethink my podcast production system from the ground up. It started by building one-off scripts and web apps for individual parts of my workflow. Over time, though, Claude has helped me bring those pieces together into a comprehensive app that tracks the entire production process from start to finish. It saves time because I always know where things are, and I can concentrate on editing instead of finding things. [SPEAKER_00] Claude is built by Anthropic, a public benefit corporation founded on a hard question. How do we make sure AI turns out well for people? They take on hard questions about AI in the open. Jobs, kids, trust, because a future worth hoping for can't be built behind closed doors. I'm a big fan of Claude Design, too. It's the creative collaborator for design work. Describe what you want, a website, a prototype, a slide deck, [SPEAKER_00] attach your own design system, and it builds with your real components, so everything comes back to you in your own design language. Edit right on the canvas, then export clean PDFs and PowerPoints when you're done. Then there's also Claude Fable, which is the newest and smartest model from Anthropic, kicking off the Claude 5 family, and it really shines on long, tricky work. The tougher the task, the bigger its edge gets. [SPEAKER_00] Hand it a job you'd normally split into 10 separate requests, and it'll take on the whole thing at once. I'm also a fan of deep research. It takes things much further than standard web search, delivering deep, trustworthy analysis backed by real sources, turning what would take hours into minutes. It can surface links between 50-plus sources that you'd likely miss on your own. Why do people trust Claude? [SPEAKER_00] Claude comes from Anthropic, a public benefit corporation. There are no ads and no tricks designed to keep you scrolling. Anthropic also puts out genuine research on tough topics like what AI means to jobs rather than avoiding them. Companies including Stripe, Shopify, Pfizer, Rakuten, and Notion trust Anthropic to help guide how they bring AI into their businesses. For problems worth solving, get started with Claude today at claude.ai slash app stories. [SPEAKER_00] That's C-L-A-U-D-E dot A-I slash app stories. And check out Claude Pro, which includes access to all the features mentioned in today's episode. Claude.ai slash app stories. Our thanks to Claude for their support of the show. Federico, I want to steer this to your 5090 because I think it's interesting to compare and contrast [SPEAKER_00] the different architectures between a GPU like that and what kind of things that can accomplish compared to what Apple's architecture can do. Yeah. [SPEAKER_01] So obviously you have a dedicated GPU with, I forgot how many Tensor cores are in the GPU. [SPEAKER_00] And I think for context, this 5090 is sitting in your gaming PC. And what is it? It's 5090. Is it a mini tower? [SPEAKER_01] It's an ATX build with a Lian Li A3 case. So this is a gaming PC. It's not a PC optimized for local AI inference. [SPEAKER_00] And it's not small either. It's not small. [SPEAKER_01] But it is compact, right? And so there are bigger, quieter builds that you can do with a 5090. Mine is not quiet because it needed to be small, right? This is not a machine that I use for local AI. I use it for gaming, but it's a 5090. So you can use it for AI. It's faster than a 95 Ultra, especially when it comes to shorter context windows. [SPEAKER_01] But the gap is getting much, much narrower than before. So to give you some numbers, I mean, first off, a couple of things worth keeping in mind. Tensor cores are... NVIDIA has been doing excellent work with their Tensor cores designed specifically for the matrix multiplications that LLMs do. Second, you have faster bandwidth, right? Memory bandwidth. You got 1.... [SPEAKER_01] Almost 1.8. 1.79. 1.79 terabyte per second. Terabit per second. On a 5090? [SPEAKER_00] Compared to what was at 1.2? [SPEAKER_01] Compared to 1.2. Right? So you get, what, a 30% difference there? The other difference, on this build, I don't have a unified memory setup. The fast memory at that speed, the 1.8 terabyte per second, is only the VRAM that's inside the GPU. So it's 32 gigs of VRAM. It cannot access... See, that's the key advantage here. [SPEAKER_01] You got a really fast GPU that beats, you know, by a good margin, the M5 Ultra still on local AI performance, but it's limited to 32 gigs. And unless you get yourself one of those DGX Sparks, you know, unified... NVIDIA is now doing unified memory things. [SPEAKER_01] Or unless you want to offload it to the much, much slower system RAM running over PCIe, like, you are going to be limited by the 32 gigs of VRAM in the GPU. Apple, with a unified memory approach, that is, it's like a shared pool across the CPU and the GPU, you got 256 gigs. So, like, you can run bigger models, you can run, you know, bigger prompts on a single computer. And also, the size comparison is not even fair. Like, this... The size, the heat, the sound. [SPEAKER_01] The heat, the sound. Like, this PC is like five times the size of a Mac Studio. And it's incredibly loud and it gets incredibly warm. When I walked into my office the other day and I was doing benchmarks on the 5090, it was hot in here. And you could tell the difference. When I'm running, like, right now, in this very moment, Flash Next 4-bit is loaded on my Mac Studio. It doesn't make a single noise. I can barely hear the fan if I actually place my ear on top of the computer. [SPEAKER_01] Like, physically touching the computer with my ear. I can hear the fan. Otherwise, I cannot. I can hear the fan when I'm generating images at 4K. In that case, because it's very GPU-bound, in that case, a 4K image will spin up the fans. But it does so for three minutes and then it's done. So, 5090, I think, it's not an unfair comparison. I think Apple has been moving toward that target, right? And I think now, if you look at... [SPEAKER_01] Let me take a look at the numbers again. If you take a look at these numbers that I'm getting with an M5 Ultra and the 5090, the 5090, you know, I had to use different models, right? To test... I wanted to test the same model across the M3 Ultra, M5 Ultra, and the 5090. So, I couldn't use that build of Flash Next. So, it's a different model with different numbers. But the... But the... Sort of the scale holds up. [SPEAKER_01] Because the same model was reading, right? So, processing the prompt at 400 tokens per second on an M3 Ultra. Jumped to 1700 on an M5 Ultra. So, 400 to 1700. But it was going at 3000 on a 5090, right? So, I think, though, at this point, it's not the rumor of Apple, you know, [SPEAKER_01] moving ahead with the M7 family for AI and an M7 Ultra. [SPEAKER_00] And building servers and other rumors like that. [SPEAKER_01] I think their M7 Ultra will match or perhaps even exceed a 5090. That's going to be a 2028 product. [SPEAKER_00] Yeah, I'm very, very excited about the M7 having heard the rumors. [SPEAKER_01] So, I... Long story short, I wanted to be objective about it. A 5090 is still better. If you're just one shot in a prompt and loading up, you know, I couldn't test Flash next. I tested QAM 3.8 27B. It's faster than an M5 Ultra. But the M5 Ultra doesn't do as bad as the M3 Ultra did compared to the 5090. But I just prefer using a Mac. [SPEAKER_01] I just prefer using something that is a pretty good operating system that doesn't suck. I'm sorry, I hate Windows 11 with a passion. As a pretty great app ecosystem, it's more elegant, smaller, quieter. Doesn't run as hot or as loud. And so, yeah, more power to you if you want to use a 5090 for local AI on your desk. Not for me, but objectively speaking, it is true. It's better. [SPEAKER_00] Yeah. Look, I think there's a lot to be said for you want to have a setup that is your everyday computer that sits on your desk, isn't super hot, isn't super noisy. And that's, you know, and that Mac Studio is just a jack of all trades in that way. It's not as maybe specialized as having a custom PC with a 5090 in it, but it can handle a lot of that stuff itself too. So, I mean, I think that's the solution I would go with if I were buying a new Mac Studio today. [SPEAKER_01] Yeah. Do you want to talk about the behind the scenes and the making of in App Stories Plus? [SPEAKER_00] Yeah, let's do that. Let's move on to App Stories Plus. But first, you know, Federico, I want a special thanks to your review sponsor. That's AstroPad Workbench. It's a really great tool for looking in on your Macs. [SPEAKER_01] Which I use on the Mac Studio. Yeah, exactly. [SPEAKER_00] Looking in on your Macs and pushing your agents forward. And it's designed specifically for dealing with headless Macs and things like that. So thanks to them. Thanks also to Claude from Anthropic for sponsoring this episode. We will be back in another week. In the meantime, go read Federico's review. It's up on the website. And you can find the two of us on social media where Federico is at Viticci. That's V-I-T-I-C-C-I. And I'm at John Voorhees. [SPEAKER_00] J-O-H-N-V-O-R-H-E-S. Talk to you next week, Federico. [SPEAKER_01] Ciao, John.