← Back to search

Every AI Model Explained In 20 Minutes (Update)

Tina Huang · 2026-08-27 · 21 min
relevance 66 4669 words Episode page ↗ Audio ↗
Show full episode description
Sponsored by Viktor, the AI employee that lives in Slack and Microsoft Teams and connects to 3,200+ tools. Hire Viktor for your team: https://ref.viktor.com/thuang-ytPromo code TINA100. $100 of his work free, no credit card, running in minutes.@getviktor_com 🧠 Signup for my live teachable workshop (it's free!) 👉 https://bit.ly/4wSHn4oIn this video I cover every AI model! 📑 Video resource mentioned in video (prompts etc.) 👉 https://www.lonelyoctopus.com/download-resource-every-ai-model🐙 Free 28-Day AI Sprint Roadmap - pick your goal to get a clear, day-by-day path forward 👉 https://www.lonelyoctopus.com/ai-sprint-roadmap🤖 Want to get ahead in your career using AI? Join the waitlist for my AI Agent Bootcamp: https://www.lonelyoctopus.com/ai-agent-bootcamp🤝 Business Inquiries: https://tally.so/r/mRDV99🖱️Links mentioned in video========================🔗Affiliates========================My SQL for data science interviews course (10 full interviews):https://365datascience.com/learn-sql-for-data-science-interviews/ 365 Data Science: https://365datascience.pxf.io/WD0za3 (link for 57% discount for their complete data science training)Check out StrataScratch for data science interview prep: https://stratascratch.com/?via=tina🎥 My filming setup ========================📷 camera: https://amzn.to/3LHbi7N🎤 mic: https://amzn.to/3LqoFJb🔭 tripod: https://amzn.to/3DkjGHe💡 lights: https://amzn.to/3LmOhqk⏰Timestamps========================
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
The AI model landscape shifts fast; this episode maps which current models to use for which tasks across flagship, mid-tier, and light categories.
Benefits
  • Clear tiering of flagship vs workhorse vs light models
  • Cheap open-source alternatives to Claude and GPT identified
  • Model-to-task matching for coding, agents, and multimodal work
  • Highlights near-flagship open-weight models like Kimi K3 and GLM 5.2
Use cases
  • Victor AI employee auto-pulls YouTube analytics and posts Slack summaries, saving 4 hours every Monday
  • Minimax M3 powers the host's Hermes agent cheaply with vision and 1M context
  • DeepSeek V4 Pro used as dirt-cheap coding fallback when Claude/OpenAI credits run out
  • GPT 5.6 Sol used to build ADHD-coping desktop apps with accurate visual designs
  • Hunyuan 3 (HY3) tried free on Nous portal for Hermes for reasoning and fact verification
KPIs / results
  • Victor saves ~4 hours of Monday marketing work and connects to 3,200+ tools
  • GPT 5.6 Sol offers Fable-level coding at half the price
  • Mimo V 2.5 Pro: near-flagship quality at ~1/8 of Opus pricing
  • Minimax M3: 1 million token context window
Tools / build
  • Hermes agent powered by Minimax M3
  • Victor AI employee in Slack
  • ADHD-coping desktop apps built with GPT 5.6 Sol
  • Nous portal free trial of Hunyuan 3 for Hermes
0:00 / 0:00
📑 Chapters — tap a time to jump there
00:00
Intro
  • Snappy refresh of the shifting AI model landscape
  • Categories: flagship, workhorse mid-tier, light, plus specialized models
01:01
Flagship Models
  • Claude Fable 5 smartest but pricey; GPT 5.6 Sol best all-rounder
  • Opus 5, Gemini 3.1 Pro multimodal king, Grok 4.5 with X data
  • Kimi K3: open-source flagship beating all but Fable and Sol
04:18
Viktor: AI employee overview05:19 Work Horse/Mid-tier Models11:36 Light Models
  • Victor AI employee pulls YouTube analytics, posts Slack summaries
  • Replaces 4 hours of Monday work; connects to 3,200+ tools
Wo gibt's eigentlich die weißen Rosen aus Athen? Und wo regnet's für mich nochmal rote Rosen? Natürlich im Europa-Rosarium Sangerhausen mit der größten Rosensammlung der Welt. Viele tausend Rosenarten warten auf Sie mit prachtvollen Blüten und herrlichem Duft. www.rosarium.de Die richtige Business-Idee ist da, aber es fehlt an Zeit- und Programmierkenntnissen. Warum? Der IONOS KI App und Zeitbilder setzt Ihre Idee schnell und einfach per Chat-Eingabe um. Ob App, Website oder digitales Tool. Die KI baut es und mit nur einem Klick geht es live. Domain und sicheres Hosting auf europäischen Servern inklusive. Jetzt schon ab 9 Euro pro Monat starten. Auf IONOS.de Slash App IONOS. Digital an ihrer Seite. This is an updated video on every AI Model. You see, the AI Model Landscape is shifting. So I wanted to do like a snappy little refresh to make sure that you're not missing out on some amazing new models and new use cases out there. Especially with so many powerful open source AI Models available now, I have personally been updating the way that I use different models for different things. So maybe you can too after watching this video. Sounds good? Alright, let's go. A portion of this video is sponsored by Victory. Okay, so here are the categories of models that I will be covering today. So first category is the flagship models. These are like the frontier highest performance models. Then we have the Workhorse Mid-Tier Class. They're the most balanced in terms of capability, price, and speed. And accounts for like 80% of the work that you would normally do on a day-to-day basis. Then there are the light models. Small, fast, and cheap models that are the most suitable for automations and bulk processing. And I also want to point out some more specialized models too. For example, those that are best at coding, different types of multi-modality like music, video, audio. Those are best for customization. Like if you want to fine-tune and build on top of them, etc. So let's start off with the flagship models. Just as a note, a company may define a model as their flagship model, but I may not include it in this category because compared to other flagship models out there, maybe it's just not quite there yet. I want to keep this list of flagship models like legitimately the best of the best. So starting off, we have Claude Fable 5. This currently is the smartest model that you can buy. In fact, it's so smart. As most of you know, the US government literally took it offline for like three weeks. So yes, very capable for sure. But it's also expensive, slow, and has restrictions. For me personally, I like using Fable for my most complex planning and brainstorming sessions. Also for like very complicated software projects. Then there's GPT 5.6 Sol by OpenAI. This is the current flagship all-rounder. It is at Fable level coding, but only half its price. It's also the best in class when it comes to terminal and web-based browsing and doesn't ask for permissions on things like continuously. Maybe because it has less restrictions. This one, I don't know. Not verified, but I feel like it. And it's often bundled with image generation capabilities, which Fable does not have. And because I do a lot of coding-related tasks, I especially like using this model for building things when I need like very accurate aesthetic visual designs. Like these desktop apps that I made, which I use every day to cope with my ADHD. Next, we have Cloud Opus 5. Opus used to be Anthropix's flagship model, but now it's been demoted to the reliable flagship model. Really great to use when your Fable quota runs out. And I personally like how it's more direct and honest about what it knows and doesn't know compared to Fable. Next, we have Gemini 3.1 Pro. Come on, Google. You promised us Gemini 3.5 Pro like months ago. But despite that, you still belong in the flagship category because you are the multimodal king. Gemini Pro is the only frontier model that actually can watch a video. But when it comes to raw capabilities, especially coding abilities, it is lagging behind other flagships. And because of how much work that Google has put in to integrate Gemini across its mass ecosystem of Google products. You probably are using this model or one of its siblings a lot more than you think. Like through Gmail, Notebook LM, Google Docs, Sheets, YouTube. This is just not a model that I would really directly reach for. And especially when it comes to coding. But maybe with 3.5 Pro, whenever it drops, things would be different. Do let me know in the comments by the time you're watching this video, has 3.5 Pro dropped and what you think about it. Next up, we have Grok 4.5. But by the time that you are watching this video, it might be Grok 4.6. But regardless, it still belongs in this flagship category. With this model, you get near opus level coding abilities, but at a fraction of the cost. And it does have X data that's built into it, which is very helpful if you're someone who's interested in having real-time social news pulse. Do let me know in the comments as well if by the time you see your video, it is in fact 4.6 and any first impressions that you have. All right, time for the last model of this category. And it is a very special model. And that is Kimi K3. Kimi K3 is a very special model because it is a flagship model that is also open source. Well, technically open way. But basically, you can go and actually download this model yourself and run it on your machine. Technically speaking, like most of us don't have machines that are big enough to run this model. But technically, you could. Capability-wise, it beats out all models out there other than Fable and GPT Sol. And when it came out, there was so much hype surrounding it. Justified. Having used them myself can confirm as well. This is a really big deal because the gap between open source and closed source models is so, so, so small. So we shall see what happens in the next few months. And on that dramatic note, let's go on to our next category. Now, all of these models themselves represent only raw potential. What actually matters is where you point that potential at. Every Monday, our marketing team used to spend four hours manually pulling YouTube analytics and marketing numbers across platforms before we can even have a strategic conversation about what to do next for the week. But now, we hired Victor for this job. Now, every Monday morning before anybody even opens their laptop, Victor has already pulled our YouTube analytics, cross-referenced our content performance across platforms, flag well worked and didn't work, and posted a clean summary onto our Slack channel. He also drafts content recommendations based upon how we perform the best that week and queues them for our approval. So yeah, as a team collectively, that used to be four hours of work every single Monday. But now, it just appears. Victor is an AI employee that lives in our Slack, connects to 3,200 plus tools, and does the work instead of just answering questions about it. The whole team shares him, and he works through weekends, building things when nobody else is at the desk. So, if you also want to hire Victor for your team, my link and promo code is in the description. $100 of his work free, no credit card, running in minutes. Thank you so much, Victor, for sponsoring this portion of the video. Now, back to the video. The Workhorse Mid-Tier Models, aka your daily drivers. Models that are still very capable, but also faster and more reasonably priced. Most suitable for probably 80% of the tasks that you do on a day-to-day basis. Now, it is really this mid-tier category of models that has exploded in the past few months with some very exciting new players. But let's start off with the usual suspects first. Claude Sonnet 5, the best value in the Claude family. It's my go-to model if I'm just chatting with Claude, using it with Co-Work, or doing some casual coding. And next, followed up, of course, my GPT 5.6 Terra. Now, this model was last year's flagship model, but only half the price. The relationship between Sonnet and Terra is a very similar dynamic between Fable and Sol from the flagship category. Terra is similar in capability to Sonnet, but cheaper. And it's a default option in the chat GPT ecosystem. Next up, Gemini 3.6 Flash. This one is kind of embarrassing, Google, because how come your mid-tier model is like low-key better than your flagship model in some things? I don't know about that. But basically, this is Google's mid-tier model, really great model. It's also fast and cheap, especially when it comes to multi-modality. And if you're using any Google product, this is probably the default model. Great. We are not done with our usual suspects, our closed-source models. Let's now move on to some very interesting open-source models that has recently been really making waves. And you should probably try out one of these models yourself, if you haven't already. Let's start off with GLM 5.2. This is a Chinese open-source model developed by Z.ai. And it's technically the best open-weight model out there, actually beating out Sonnet 5 when it comes to terminal programming. Apparently, 5.3 is supposed to be coming out soon as well. So I wonder if GLM 5.3 might end up being in the flagship territory too. In any case, this is a really cool model trial if you want to start dabbling in the open-source space. Although it is way too large for you to actually download yourself, probably, unless you have hardware that is worth hundreds and thousands of dollars. So you probably want to try it out through some hosted version. I'm not going to go into too much detail about how to actually run and use open-source models, because I want to keep this video nice and snappy. And I do have another video, which I'll link over here, that goes into much more details about that. So please do check it out. And that applies to all of the open-source models that I will be talking about. And you know what? I'm also going to link in the free guide in description, a little quick start guide for how to try out open-source models, if you wish. I promise you, it's a lot easier than you think. Okay, so next up is Minimax M3. This is also an open-weight model that is characterized as a triple threat because it has almost flagship level coding, long 1 million context window, and native vision multimodality. So you can do things like give it screenshots of stuff, ask it to turn it into code, make UI automations, and also use it to power multimodal agents. Personally, this is one of my favorite models to use on my Hermes agent, because it genuinely is really good and also extremely cheap. Moving on to what I like to call a side gig model, the Mimo V 2.5 Pro by Xiaomi. So I call it a side gig model because Xiaomi is primarily a phone company, although they do dabble in a lot of different products too. But I just think it's kind of funny because you got like a phone company that creates this open source model that is also extremely, extremely good. It has almost flagship level quality at only like one eighth of Opus pricing and is really climbing up the ranks on open router. Truthfully, I haven't personally used this model that much. I'm using its little brother a lot more, which belongs in the flash category, which I'll talk about later. But I do want to experiment more with this model because it's supposed to be an alternative to Sonic 5 and especially good at powering agents. Okay, so we've currently covered three open source Chinese models that are mid-tier level and also really cheap. So you might already be thinking, wow, that's some great options already out there. But we're actually only halfway there. I have three more Chinese open source mid-tier models that you can choose from. We have the Hu Ren 3 HY3 model by Tencent. This is not considered the smartest open source mid-tier level, but I gotta say personal experience wise, it is really, really good. They had a free trial of this model on Nu's portal for Hermes and it was really, really good. People like to use it specifically for reasoning and fact verification because it is a model that overthinks everything. It is like the overthinker of the AI world. Plus, it does follow instructions really well too. Another model we have is the Quen 3 model. The Quen family has the most diverse variety of models out there to pick from. So I'm going to talk a lot more about the Quen family models a little bit later in the video, but it very much has also earned its place in the mid-tier model category too. As a little fun fact, whenever I try out like a new agent or harness or software that requires some type of open source model, I usually go for one of the Quen family. And finally, we have my personal favorite open source mid-tier model. That is the DeepSeq V4 Pro. It performs better than most mid-tier models, especially when it comes to math, coding, and deep reasoning. Yet, it is dirt cheap. So it is my favorite go-to model when it comes to coding tasks, especially when I run out of both Claude and OpenAI credits. Great budget option. So without a doubt, Chinese open source models is absolutely dominating, especially in this mid-tier category. But there are a couple of other models in this category that is not Chinese too. So let's start off with Inkling by Thinking Machines. This is America's answers to open source models. It is deliberately not the smartest model out there, but is the most customizable, calibrated, and least censored Western model. It also has a very impressive multi-modality with native audio, vision, and text processing. I must admit, I have personally not given this model its fair chance yet, especially when it comes to customization. So that is on my to-do list. But let me know in the comments if you have tried out this model and what you think about it. Now time for the other competitive non-Chinese open source mid-tier model. Europe's Pride, Mistral Large 3. This is Europe's biggest open source model. For people who don't want to be using US or Chinese models, it does have some specific EU slash GDPR regulations that are built into it, which I believe is very important if you are from Europe because you guys have a lot of laws. Again, I must admit, this is not a model that I really played around with much because I just never really had that incentive because I'm not from Europe. But if you are, please leave me in the comments on what you think about this model. Okay, so before I go on to the next category, I do have to round out this category with what we call the ghost of the open models. That would be Llama 4. This is the fading formal king of open source models developed by Meta. Unfortunately, it has been frozen in time because Meta has decided to go closed source now with their new Muse pivot. Speaking of Muse, I'm sorry, I lied. I do have to include Muse 1.1 Spark into this category of mid-tier models. Since this is a video about every AI model, this is the first closed source model by Meta. And honestly, I find this model quite disappointing. It is supposed to be a generalist model that's especially good at computers and multi-agent orchestration. But I did use it for computer use and multi-agent orchestration and was still quite disappointed by it. I mean, like, it's not bad, right? It's not like a bad model, but it's just like nothing special. Especially since it's a closed source model, I don't see why I would reach for this model compared to all the other open source models that are out there. So low-key a flop in my opinion. But let me know in the comments what you think. With that, let's move on to the final category of models, the light models. Models that are fast, cheap, and quick. Best to be used for high volume bulk processing and automations. For the open source models in this category, you can actually download them and run many of these models on your local machine, which is incredibly cool. I'll be talking more specifically about which models are the best to run locally on your machine a little bit later in this video. But let's get through this final category of models first. First, we have Quad Haiku 4.5. This is an anthropic model that is up to four to five times faster than Sonnet and is meant to power sub-agents under a smarter general model like Sonnet, Opus, or Fable. It is the go-to high volume processing model if you want to stay in the anthropic ecosystem. Then we have GPT 5.6 Luna. Luna is the smallest, cheapest, and fastest model of the GPT 5.6 family. So this is the model to choose for your high volume chats, bulk processing, and for powering your sub-agents if you want to stay in the open AI ecosystem. Our team does have some open AI-specific products, so we do use the Luna model quite often. Next up is the Gemini 3.5 Flashlight. This is actually pretty impressive, Google, because it is the fastest model ever measured, allegedly. It's great to use if you need something real-time, like real-time autocomplete search functionalities, instant document triaging, high-speed throughput. It's its defining trait, and it's very impressive. Good job! The other two closed-source options in this category is the Quen 3 Flash, which, as its name suggests, is extremely fast and also has a 1 million-plus context window and extremely cheap. There's, of course, also the Grok 4.5. This is XAI's small, fast model that is especially good at live news routing, taking advantage of the XAI news feed and ecosystem. Hello, Tina from the future here. Just wanted to pop in to let you know that I have a free live session on August 17th with Teachable AI Academy on my AI COO setup. I have offloaded a lot of ops stuff now to AI, so I'm really excited to show you guys. Plus, there will be a live Q&A as well. Hope to see you guys there. Link is in the description. Now let's talk about some open models in this category, starting off with DeepSeq V4 Flash. This is the little brother of my favorite open-source model, the DeepSeq V4 Pro. And just like it, it is really, really cheap. In fact, at more than 55 times cheaper than GPT 5.5, it is practically free, I would say. GirlMath. And it's still a legitimately really good model. It especially excels at doing classifications at large scales. Unfortunately for most people, you probably still need to use this as a hosted version through something like Open Router because the hardware requirements is just out of reach for most people. Because you do need hardware that's like over 128 gigs, which at least I don't have. And that is also the case for the next model, dubbed the Volume King, the Mimo V2.5 from Xiaomi. This is the little brother of the Mimo V2.5 Pro, which we talked about earlier. And it is actually the number one most used model on the entire Open Router marketplace because you get near mid-tier level performance for absolute dirt cheap. So people use it for high volume tasks, especially high volume code assistance. Two other models to pay attention to that are just out of reach for most consumer hardware is Ling 3.0 Flash from Ant Group. Personally, I haven't used this model, but it is growing very fast in popularity. So maybe I should. Second is Mistral Small 4. This is three models in one. It has reasoning, vision, and coding, all merged together into a tiny little package at a very low cost. Now, very exciting. To round up this category, I have three other models that I want to highlight, which you can in fact run on consumer hardware. Like whatever laptop you're using can probably run this. Even your phone or Raspberry Pi. So the first one are the smaller models of the Quinn family. Like I said earlier, the Quinn family is massive and they have all sorts of sizes of different models, ranging from models that are only 0.5B all the way up to a trillion parameters. So I'm going to put on screen some of my favorite Quinn models that you can run on your consumer devices. And just to put it out there, my personal favorite is the Quinn 3.635B model. It does require over 32 gigs of VRAM, so slightly on the higher end side, but this is the model that basically runs my Hermes agent. So gotta give a shout out here. The Gemma 4 family model from Google is also amazing for running on low-code devices. It is considered the best model model, especially for phones. And of course, you can run it on practically any laptop that you have. And finally, there is the 5-4 family from Microsoft. This is Microsoft's take on a small model. It is especially good at math and logic, considering how tiny it is. This is my personal favorite model to use on a Raspberry Pi. Amazing! Wow, we have covered a lot of models. I don't know about you, but my head is already spinning. But since this video is called Every AI Model, I would be amiss to not also cover the specialized models out there, models that are trained specifically for a certain task. So the good news is that we have actually covered most of these models already. So for each of these specialist categories, I'm only going to briefly touch on the models that I haven't already covered earlier and put the rest of them on screen. All right, ready? Lightning round. Coding specialist. Take a screenshot. Or check out the free guide in description. We have pretty much covered all of the coding models actually, except for Kimi K 2.7 Code, which is a great model for cheap daily coding and swarms. The Quintry Coder Next, also open source model, great as a coding agent. And I can actually fit on your machine for local and private coding. Mistral Devstral, which is Europe's agentic coder with EU compliant automations and Longcat 2.0. First of all, I think this is a hilarious name, Longcat. And it's developed by Meituan, which is a Chinese food delivery company as a side hustle. But despite that, it's actually really good. Next category, multi-model reasoning. Take a screenshot. We have actually covered all of these models already previously. Image generation. The absolute best image generation model right now is GPT Image 2 from OpenAI. Fun fact, when you use ChatGPT, the flagship models don't actually have image generation capabilities. They're actually secretly bundling together GPT Image 2. This is followed closely by the Nano Banana family from Google and MidJourney. C Dream is one of my personal favorites for image generation and Irogram is considered the most reliable when it comes to complex image generation. And of course, we do need to mention Muse Image from Meta. Close source model is part of the Muse family. And the last model that I do have to mention by name is Flux 2. It is widely considered the best open source image model that you can actually download and generate images yourself with. Now for the rest of them, take a screenshot because we're moving on to video generation. Video 3.1 from Google is the current king of video generation, especially since OpenAI's Sora has died. This is followed by RunwayGen 4.5 that people really like because it also has an editor suite where you can edit and manipulate videos very easily. Cling Omni is also very popular, especially because you can lip sync in five different languages. And Seed Dance is currently having a viral moment with top raw quality and the ability of mixing all different types of assets together. Here are some examples of things that people have created using Seed Dance. Really amazing. And for the rest of them, take a screenshot. Now next category is music, voice, and audio. The top two closed source competitors is Suno and Udio. Suno is considered the best for full songs and vocals, while Udio is considered to have the best raw audio fidelity for instrumentals. It's also considered legally cleaner because Udio has fully licensed UMG plus Warner platforms. Now when it comes to voice cloning across 70 different languages, the clear winner is Eleven Labs with its sister Eleven Labs Music that produces licensed data music and that is safe to monetize. So very popular with content creators. Liria 3 is Google's API first composer with real-time streaming. And GPT Live is OpenAI's new real-time voice tier. Now on the open source side, Voxroad TTS is for open-weight speech which you can also self-host. And Stitch Audio for open source free for AI text-to-speech and free voice cloning. Take a screenshot. Now when it comes to customization, I mean like fine-tuning these models and building on top of these models. Inking, which I already mentioned previously, was really built just for this. As well as NVIDIA's NemoTron 3 family. And of course, we can't forget the Kuen family. It is the widest model family on earth. 201 languages and every size model. It is the favorite base model to self-host and fine-tune for today. Next category, Enterprise and Retrieval. We already mentioned NVIDIA's NemoTron 3 which is also the standout of this category. There's also Cohere which is nicknamed the Rag Plumbing Specialist. Great for enterprise search and retrieval pipelines. There's also Amazon Nova 2 plus Nova Act. The all-AWS stack for enterprises that are already in the Amazon ecosystem. Finally, there's Perplexity Sonar specifically trained for search functionalities. Last major category, the local slash self-hosted category. I've already covered most of these models. The Kuen 3.6 35B is my favorite local model which I run on my Mac Studio. And because I am a big fan of Hermes, the Noose Hermes 4 family is also very good and optimized for running your Hermes. And the only other family I haven't mentioned yet is the GPT-OSS family. This is OpenAI's open source models that ranges in different sizes if you do like the OpenAI family of models. And finally, the last model that I do have to mention because I really am serious about covering every AI model that I can is the Step 3.7 Flash model. This is a very specialized model and people like to call it the Big Mac Special because it is a model that is specifically great when it comes to running on Apple products. Although you probably can't run it yourself because it requires over 128 gigs of RAM. Alas. All right, that's it. Oh my goodness. I think I have now covered every single AI model. At least I tried my best. Do let me know in the comments what you think about this new model landscape and if there's any models that you are curious to try out now. Plus, what's your current favorite model? And I'll see you guys in the next video or live stream.