← Back to search

What are we talking about when we discuss Harness | In-depth conversation: Minimax × Hermes Agent

十字路口Crossing · 2026-04-28 · 77 min
relevance 100 789 words Episode page ↗ Audio ↗
Show full episode description
🚥 上周,我在 B 站做了一场直播,邀请了中美两国一线 Agent 开发者深度对谈: MiniMax Agent 首席架构师 阿岛 MiniMax Agent 研发工程师 择因 Hermes Agent(Nous Research)业务负责人 Tommy Eastman 这也是 Hermes Agent 在全球获得广泛关注后,官方首次现身中国社交媒体平台,并且正面回应了中国团队 EvoMap 对其“抄袭”的指控。 我们一起围绕「从 OpenClaw 到 Hermes」的热潮迁移,深入拆解了 Agent 和 Harness 的多个关键议题: Hermes Agent 为什么会在 OpenClaw 之后火起来? 模型会吃掉 Agent 吗?通用 Agent 会吃掉垂直 Agent 吗? 为什么 MiniMax 和 Anthropic 都要同时做模型和 Agent? 如何看待 Agent Infra 层面的创业机会? 如何看待 Multi Agent 协作的范式? 如何看待 Claude Code 的实名制要求? 为什么 Anthropic 不发布 Mythos? Claude Code 源代码泄露的影响 从 Manus 发布到今天,Agent 范式的变化 中美模型的差距,和开源的窗口期 「把自己蒸馏成 Skill」 0 人公司的可能性 ——完全由 AI 驱动的公司是否会出现? 🎬 本期内容的视频版本已同步上线于 @Koji杨远骋 的 哔哩哔哩 。 📒 文字版已发布于 @十字路口Crossing 公众号。 🟢
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
A deep China-debut conversation on what an agent harness is and how Hermes Agent differs from OpenClaw via memory and self-evolution.
Benefits
  • Agent framework orchestrates tools, main loop, state, error handling
  • Multi-level memory saves successful workflows as skills
  • Consistent output across eight different models via shared skills
  • Out-of-the-box deploy on your own computer or VPS
  • Open source with strong cyberpunk-brand community
Use cases
  • Daily token consumption surged from 2 billion to 20 billion in a month
  • Token usage reached almost 300 billion per day
  • Setup from install to first agent task in less than two minutes
  • Switch among models (Minimax M2.5/M2.7) keeping same workflow reliability
  • ERN algorithm extended context length from 4000 to 12800, adopted industry-wide
KPIs / results
  • 2 billion to 20 billion daily tokens in over a month
  • ~300 billion tokens per day
  • Context length extended from 4000 to 12800
  • Hermes Agent started about a year ago
Tools / build
  • Hermes Agent (open source agent framework)
  • Nous Research NOS platform API
  • ERN context-extension algorithm
  • Distro distributed-training optimizer
  • Minimax M2.5 / M2.7 models, Max Hermys
0:00 / 0:00
📑 Chapters — tap a time to jump there
08:14
Nose Research 的底色 他们发表了一篇扩展上下文长度的论文——然后被所有主流模型公司立即采用。 从 Discord 社区到 LLAMA 后训练到 Hermes Agent。 为什么他们的模型叫"Hermes"? 🟢
22:53
自我进化 vs 抖音算法 抖音也越用越准,为什么没人说它"自我进化"? 养虾的过程中,用户开始对 agent 产生感情,宕机了会心里落寞。"自我进化"背后,其实是一个更深层的用户诉求。 🟢
  • Self-evolution vs TikTok algorithm; users grow attached to the agent
45:11
应用层终将被模型内化 你写的 skill、搭的 workflow,最终会成为模型训练的素材。 Anthropic 为什么在过去一两年势头比 OpenAI 更猛? 做通用 agent 应用,"你永远会被模型内化掉"。 🟢
  • Application layer eventually internalized by the model; Anthropic's momentum
49:27
中美模型差距:差在哪里 训练方法的差距不大。真正的差距,是有没有请到足够好的人,去定义足够好的任务。 中美双方的思考"在同一个大气层内"。 但有一个具体的、国内还有差距的地方——不是算力,不是算法。 🟢
  • China-US model gap: defining good tasks with good people, not compute
🌐 This transcript was automatically translated to English from the original.
Hi, I'm Koji. This week's crossroads is a clip of the highlights of a live broadcast I did recently at Bilibili. We invited Minimax Agent's chief architect Adao and R&D engineer Ze Ying, as well as Hermes Agent, who recently hit the global screen after OpenCloud, Tommy Eastman, the business leader of Hermes Agent. This is also the first time Hermes Agent has officially appeared on a Chinese social media platform after gaining widespread attention around the world, and they also responded head-on to accusations of plagiarism against them by Evil Map, a Chinese open source team. We chatted live for more than two hours. Back here, Agent and Agent Harness talked about a lot of topics that I find very interesting. In today's podcast, in order to ensure smooth listening, I used AI to directly convert all the English-speaking parts of the live broadcast into Chinese. And I specially selected a few Chinese voices with translation guns so that everyone can distinguish which ones are translated by AI and which ones are Chinese spoken by us living people. OK, let's get started. Hello everyone, I am Director A and I am in charge of Mimax. It has been almost three years since I joined Mimax’s R&D team. I had a lot of experience in Internet entrepreneurship and work before, and I was deeply involved in the research and development of the Mimax MR series models and Agents including heroin, etc. I am very happy to communicate with you here today. We also invited online the person in charge of the product and strategy of Hermes Agent, who I have in-depth cooperation with Mimax. This should be the most popular agent of Hermes Agent in the past month. This is the first communication with everyone in China. Hello everyone, I am Zeyin. Mimax's Agent R&D Engineer, like our Mimax official website's Agent and Max Cloud, as well as the development of Max Hermis, which has just been launched recently. Hello everyone, I am Tommy Eastman, the business leader of Hermis Agent. I am very happy to communicate with you. Our live broadcast today is because after the Spring Festival, there is a national shrimp farming boom, which is the popularity of Open Cloud. But it seems that the popularity of shrimp farming now seems to have gone down overnight, so I would like to ask you to review the process first. I think the Lunar New Year in China is very interesting every year. Deep Seek was popular last year, and after the Spring Festival, everyone had a completely new understanding of AI. You may have thought it was still far away from us before, but it may be the same this year. In fact, it happened that Open Cloud became popular in Silicon Valley in about January. I remember it was on the day of our IPO. We had such a connection with Peter, because that day was also a coincidence. Antheophic banned the use of Open Cloud subscriptions. So Peter also hopes to find such a model that is very suitable for lobster, because everyone knows that lobster consumes tokens very much. We actually didn't realize that this project would be so popular at the time. Then it started to become popular overseas, and Many Mini began to be collected. Then we did not expect that it would be so popular in China. Yes, Open Cloud is more popular in China. I feel it is even better than Silicon Valley. It has sunk to a very public crowd. What do you think is the reason? My own personal feeling is that it is during the Spring Festival. And after the Spring Festival, it seems that everyone around us starts to discuss lobster. We even see this kind of craze starting to appear in various places, even wearing this blind hat. There is also such an offline installation activity of the program. I think there are probably several reasons for it. The first reason is that I think our own feeling is that it is actually more popular in China than overseas, because in fact, overseas, everyone has actually been exposed to things like Cloud Code, including Coworker. There are some agents that are relatively easy to use, but you may not have been exposed to relatively easy-to-use agents before in China. I think there is a lower-level reason for this. It is because the domestic model may not have such a strong agent or such a capability in the previous stage. With the release of models such as Mimax M2.5 M2.7, in fact, domestic models also have such capabilities, so there are domestic models and there is an agent like OpenCrawl. It can make it very easy for everyone to contact through IM. It actually allows Chinese users to complete an experience that is close to going from 0 to 1. The window has been broken. Another reason I think is also because I won the Spring Festival Gala this time and made a comprehensive all-in investment in Doubao. So I think these two things are combined to make everyone's understanding of AI truly enter the agent era. What do you think is the reason? It seems that everyone suddenly went from shrimp farming to horse taming. It was this attention that went from OpenCloud to Erm's Agent. Was there any turning point in the process? From the explosion of OpenCloud to the gradual stabilization, and finally to the recent Ermys, which is our Max Ermys, it took about a month. Within this month, everyone has basically experienced OpenCloud, but it does have some unstable characteristics. It is four times a day. At 1 o'clock, it will refresh its memory. Then you may have better communication with it. After a day, it will say that I forgot what we talked about yesterday, etc. In fact, Ermys has grasped these pain points. It has made special efforts in memory, and it has done a very multi-level memory. From this perspective, it has indeed made up for the shortcomings of OpenCloud to a certain extent, so it has some technical foundations to bring it to life. Yes, and at the same time, It is true that everyone has enjoyed the efficiency improvement brought by AI to a certain extent in the process of using Lobster, so this should be a two-strong rush. OK, our next content will be in English during the live broadcast, so in order to make it easier for everyone to listen more smoothly, we used AI to convert the voice of the podcast into a Chinese version. Next, I would like to ask Tommy to introduce to you what Ermys Agent is, what is the so-called self-evolution, and what is the difference between it and OpenCloud. These are issues that developers and users across China are very concerned about. Thank you for inviting me. It's a pleasure to be here. In short, Ermys Agent is an open source agent framework. If a large language model is compared to a brain, then the agent framework is the hands. It is a system through which the model performs tasks in the real world. It handles all the complex coordination work, such as tool orchestration, main loop management, and state management and error handling. Ultimately, it allows the model to get the most out of it for the user. As for what makes Ermys Agent different, there are a few key points. First of all, it runs very well. Through our NOS platform API, you can get the most out of it for the user in less than two minutes. Easily complete the entire process from installation to execution of the first agent task. This convenient setup is a huge advantage. It eliminates the initial frustration for users. Secondly, there is its memory component. This is the part that everyone really likes. It is also probably the most interesting point. Anyone who has used Agent in depth has had that frustrating experience. You ask him to do something and he does it right this time, but he fails when he does the same thing next time. This makes it difficult for you to trust him. The memory function of Hermes Agent solves this problem and allows the Agent to remember a successful workflow. And save it as a skill. Once he knows the right path, he can emerge perfectly every time. This makes him far easier to use and more trustworthy than before. This also answers your question about self-evolution. The Agent will continue to improve itself. The knowledge compression that occurs in this process is very valuable. It can bring greater consistency, especially across different models. So even if you have eight different models, as long as you use the same framework and skills, you can get the same expected output from all models. This allows you to flexibly change models for different tasks. At the same time, ensure that your core workflow still maintains high performance and high reliability. Thank you Tommy. Before we dive in, can you introduce your company News Research? How was the team formed? What is the vision behind it? In particular, Hermes Agent feels like it came out of the blue and quickly climbed to the Github hot list. Is there any unknown story behind this? The origin story of Nose Research is very unique. It started in a very natural way in 2022 with a group of people who are passionate about open source and tossing around AI models. We started gathering together in a Discord channel called Nose Research, which is still our core community today. Since then, we have produced a lot of high-quality research. The initial efforts were around the post-training of the Rama model. Meta had just released it at the time, and it was one of the few companies willing to open source high-quality models. That is actually the origin of the name Hermes. Our first post-trained version of the Rama model was called Hermes. The focus of those early models was to make it sound more like a human rather than a stereotypical AI assistant. More flexibility and diversity in text, which was state-of-the-art in many fields at the time. While training the Hermes model, we developed the ERN algorithm to extend the context length of the model from 4000 to 12800. We published this research and it was immediately adopted by all major model companies. This was the prototype of what we now call the thinking model, which in turn laid the foundation for the agent model. We also spent a lot of time studying distributed training and wrote the Distro paper. This is a novel optimizer that allows us to train across non-co-located GPUs. This means that we can aggregate distributed computing power to complete meaningful training. This is crucial for us to continue to train open source models. Of course, fortunately, we still have companies like minimax that continue to carry the open source atmosphere. As for Hermes Agent, it also has an interesting origin story. It is also a key reason for his success. When was that? About a year ago? Yeah, he started it about a year ago, and it was really just to help ourselves at first, but as an open source laboratory, we naturally open sourced it. The attention it received far exceeded our expectations. I'm sure you all saw the chart of Open Router. His average daily token consumption surged from 2 billion to 20 billion in just over a month. I remember yesterday's data almost reached 300 billion per day. So when did you feel this craze? Is there a tipping point that made you feel OK Hermes Is Agent becoming a global hit? It's really getting very big at a very fast pace. I'm very happy with the growth rate, especially because we didn't have high expectations for the response. Obviously Open Cloud was very popular before, but I think a huge selling point of Hermes is that it allows users to get up and use it very quickly and easily. People don't want to deal with bloated code and a lot of complex settings. Hermes Agent comes out of the box and you can quickly deploy and run it on your own computer or VPS. That makes it ideal for mass adoption. So do you think there's a strong point about your success and Open Cloud? What's the connection between you? Of course there is a connection. I think it was a combination of factors that contributed to this outbreak of Agent. First of all, the model finally became good enough to really help people. This was obviously triggered by the huge improvement brought by Opus this winter. It achieved an improvement in quality when it was contracted, opening the door for these Agent frameworks to allow them to do more valuable things. What is cool is that the open source community has also quickly produced models with comparable performance or that are quickly approaching Opus quality. As for the difference between Open Cloud and Hermes, I want to go back to what I just said. We've really optimized for usability and utility purposes with Hermes Agent. We use usability as our north star metric: how well it performs at tasks that users actually want to do, rather than how well it scores high on some benchmark. We stay focused on the product. In an open source community, especially when a product is wildly popular, it's easy to get bloated with thousands of PRs and feature requests. I think we do a really good job of protecting the product and keeping Hermes Agent extremely easy to use. That, coupled with the memory system we mentioned earlier, are two key differentiators. The memory feature really builds trust with users. I think the third big differentiator is our brand. Nose Research has been working hard to create a cyberpunk aesthetic. We have talented designers and have spent a lot of time cultivating a community of smart and aligned values. They love to try everything we release. The success of Hermes Agent is largely due to our community. Another cool story is your relationship with the minimax team. Tommy, can you explain how you work together? We are big fans of the minimax team and their models. Their commitment to the quality of the models. Especially the performance across multiple modalities is impressive. It's their insistence on open source that makes products like Hermes Agent so cool. You can have various models and switch back and forth to customize your agent. We see that our users like the minimax model very much. Okay, let's go back to our Chinese world. Let's continue the content of our live broadcast. First of all, I want to ask you two and introduce it to us. We have been talking about Hermes today. When we talk about Hermes, what are we talking about? In my opinion, it is actually one. You go to constrain the agent. But give him a certain degree of freedom, so that he can fully deliver some results to you. When it was first proposed, you can pay attention to an article on OpenAI's official website. There, it officially announced the application of an agent like Hermes. From a technical point of view, it may be divided into six layers. But we can not talk about it so technically. Just assume that you have a colleague, and he will listen to you. But you have to make an agreement with him first about what he can do and what you can do, right? There are some things he can't do for you. But if there are some things, you can just let him do it, right? When you agree on these and configure him well, such as what his model is, what tools he has, etc., give him a certain ability, just like you give a colleague, give him a laptop, and give him a phone. At this time, you can introduce multiple colleagues who can form a whole process of mutual supervision, completion of tasks, and final delivery. But what I am talking about now is more like the scene in work production, which is to describe the agent in a human way. But in terms of technology, in fact, you need to give the agent a lot of tools, environment and freedom, and give them constraints and give them some mutually antagonistic goals. For example, one agent's role is to produce certain content, but another agent wants to stop the work. Then you have to review the problems in the work content, etc. Through this combination of multiple agents, just like the way our colleagues cooperate with each other, they can produce higher-quality results that a single agent cannot produce at all, and you actually do not need to do more intervention during this period. This is actually a concept like Hannis. Of course, I think this concept can appear at this time. In fact, there are certain physical conditions. First of all, our models have become smarter, and everyone's models have certain agent capabilities. And based on this technology, everyone is willing to share more capabilities with him, including operating mailboxes or operating your service publishing, etc. Then Hannis was born naturally. I can add a perspective. Okay, actually, I think the word Hannis was defined by OpenAI through that article. But during that time, it was actually very popular. But I think this layer of window paper may have been pierced through that article, and everyone formed a consensus. What words do we use to describe everything we are doing? In fact, everyone's practice was already before that. For example, around 90 last year, including when we were discussing internally, in my own workflow, I actually didn't use the IDE very much. Then I might have five or six agents concurrently at the same time. Even this was just my local, maybe ten agents in the cloud working for me on Sendbox. Then they were all on Github, trying my different branches and ideas. At that time, I would find that I had become a huge bottleneck. The bottleneck of my own human being, yes, is the bottleneck of human beings. In fact, at that time, the agent was no longer the bottleneck, because I needed to constantly switch between these contexts and then give them input. So one thing we were thinking about at this time was how to make it more automated. Humans, including OpenAI, the author of their Hannis article, recently shared in a podcast. They also said that their feelings were very similar, so I think these things are very similar. Everyone felt this at that moment, so what we were thinking about was OK, how to solve this bottleneck, right? Because we as technical people always want to optimize, the most important way to solve this bottleneck is to let people confirm things that originally require people to confirm. For example, if he wants to obtain this program, whether it can be deployed and run in a production environment, and whether his results are correct. It may be precipitated as a coi, or it may be precipitated as a hook. I don’t know if you have seen it. There was something called Gundam WO in my era. Its characteristic is that its engine is very powerful. We actually need to build a mecha like Gundam, and then how to maximize the power of this engine. So I think Honis is such a thing, so why do they call it Wanju? It is that you have a group of very powerful hunters, but how to make it best use it is to construct a toy. It is such a process. Well, it is very interesting. We just mentioned that the null agent is the collaboration of multiple agents. Do you want to tell us more about the concept of multi-agent collaboration? Because I understand that there are actually some disputes here. Some people think that, for example, the intelligence of today's models is very high. In fact, a single model can complete a lot of work. So why do you have to create an agent team? You said he is a product manager and he is an engineer. This seems to limit the use of the model. But some other people don't think so. What do you think? When we talk about multi agent, what is the current best practice? Well, let’s answer the first question from the technology teacher. Wishman has multi agent. Indeed, some past papers have mentioned that it may be a single agent and its effect may be better. But in fact, as the context space of our model increases, that is, the length of the rounds it can dialogue with increases, you will find that in the process of dialogue with the model, the amount of information you produce is actually less, while the amount of information produced by the model is more. Then we can do a thought experiment. Suppose there are two models. They can exchange information with a more efficient reply in seconds, and the message sent by the general model to you is a short essay, right? Then I may lose two words of approval to the model. Then come on, but if two models are allowed to communicate with each other, the amount of information is huge. In the case of such extremely high-density information exchange, the overall efficiency will be higher than interacting with an agent. This is in terms of efficiency. The limitations of a single agent. Take some recent examples. Anthropic released the OPUS 4.7 model in the early hours of this morning. What are its selling points in the model? In fact, if you really want to say it, everyone is already complaining about it online, so I won’t complain about it, right? We will just speculate on why this company released it first in certain aspects. I will talk about a few points. For example, it adds an effort called General model release should be a very shocking thing, but it means that we recommend users to use it. There is such a thinking cost to use it. Another problem it actually brings is that if your conversation is between a single agent and you keep using it, even the strongest model will have degradation behavior when you keep using it. And this is already a demonstrable thing, that is, when your context space, that is, the number of texts you talk to it, exceeds 50%, in fact, its intelligence level decreases exponentially. That is to say, in a single agent, you can never chat with it forever. This is the upper limit of a single agent. The limitations brought by the context. The limitations brought by the context. And then there are two questions about the efficiency of information dissemination. Yes. Let me add a point. In fact, a paper has come out in the past few days, which means that it studies these tasks of very long layers. Under what circumstances will the agent or model go wrong? They found that if it does the right thing, the right route will continue, but as long as it is biased in a certain place, it will easily become more and more biased. In fact, I think it is similar to people. Sometimes when we do something, it is possible to lift my toes. Maybe I have spent a lot of effort, maybe I have been doing it right, and then when I take a break and change my mind, I think that it is actually wrong. I should do it from another angle. It makes sense. At least in our own time, we will let two agents do cross check. I think it has something to do with the real-world data, including the subsequent chain, in terms of its data distribution. It sees that most of the data is executed according to a limited track. What it rarely sees is repeated verification and thinking. Yes, of course, this may be improved after strengthening i-L. Yes, but if you set up two agents, then it can be compared without the previous salary burden from another angle. You have twice the budget and a brand new Shanghainese to think about, which can often achieve a higher quality. This is also very interesting, just like when you are on tiptoe, if someone comes and pours a basin of cold water, it will wake you up. Or you can go to sleep and do other things. This is also an angle from which we want to use multi-regent. I have a small question, that is, when everyone talks about it, they are talking about the so-called self-evolution, right? Self-evaluation, but what’s interesting is that I want to say that in fact, when we use Douyin today, the more we use it, the more accurate the algorithm will be. But we don’t seem to say that Douyin will self-evolve. But why is self-evolution in this msagent or in the entire agent field today considered by everyone to be such a powerful word? What are the similarities and differences behind this? Let me start with a point of view. In fact, continuing my point just now, you will find that in this iterative process, people begin to become the bottleneck. And the more people there are, the organizational efficiency will decrease exponentially. I have experienced this in many organizations. To do a large model like this, it may cost hundreds of millions or even hundreds of millions of dollars to train at once. You need to maintain a very high talent density and a very capable organization to do this well. The most effective one that everyone thinks of is So how to scale up is how to scale up. The most effective way is to reduce human participation and let the agent and AI do it. Well, in fact, you will find that in this process, you will actually find that the life cycle of many models may be only a few months, right. Then we will launch the next version. But is it possible without the previous models? In fact, it is not possible, because the subsequent models are all trained by relying on the previous models. The previous models are all in these training processes, whether it is to help clean the pre-trained data. To find suitable data to construct data to construct the corresponding ah i2 environment. For example, in our m2.7 training, uh, in our rl pipeline, it is already possible that 70% and 80% of the work has been completed by the modeler A key, because the human engineer only has 70% and 24 hours, right? Then we still need to rest. We still have to do a lot of other things. Now 30% of the offline work has to be done by people. He is more talking about the first Human judgment, um, and test are very important. That judgment no longer needs to be made to say, OK, is there something wrong with a specific process in one of my experiments? Because these agents can help him find it out. He only needs to see the summary, for example, the results of the experiment. And then the agent, uh, gives him some suggestions at this time. At this time, it is actually a process of discussion. In the end, it is people's taste and creativity that guide that direction. Yes, we are finally going in this direction. Yes, so he does have a sense of control similar to that of Harris. In summary, it is the first one in such a high-complexity and high-density matter. If you want to scale up, AI must account for a very large proportion of it. People just control it. Otherwise, the efficiency will be very low. We will not be able to make the most productive and best things. Then naturally, what we just talked about is self-evolution. I understand, because it is a self-contained relationship. Yes, it is also because of today's example. After the agentic ability becomes stronger, an agent can run for longer and longer, which also provides a basis for his self-evolution. If he could not run for that long in the past, there is no way to talk about self-evolution. Yes, I think besides 9, the more important thing is that he is reliable and can solve complex enough problems. Yes, it can be long enough. But in fact, there is another point that is the problem of being able to start. When you meet an agent on the first day, he may not know you at all. During this process, you have to exchange information with each other. Exchanging files, exchanging data, and even some, uh, whether it's content, whether you output a lot, throw a bunch of files at him, or give him feedback on your temper, your feedback, your habits, etc., uh, it's a process of deep integration. In fact, people will have such a request, that is, can you become smarter and understand me better? In fact, there are some signs of this in the process of shrimp farming, that is, uh, users do have feelings for an agent like shrimp. Maybe one day this crayfish will have a bug, and he will feel a little lonely. At this time, Ermis proposed this concept of self-evolution. In fact, it led to the idea that I can understand you better and better while talking to you. At the same time, I can also create some small surprises for you. Yes, that is, when I am not talking to you, or when I am reviewing some work. Well, I can understand some more high-level things. Well, I think this is a big highlight for ordinary users when using it. Yes, I think this is really interesting. That is, we have used AI tools before, whether it is a chatbot or an agent product. If it does not do well, our first reaction is that there is a bug, right? This is a bug. But when we raised crayfish, it crashed or it hung up or it couldn't complete a task. Well, we just thought it was stupid and cute. Sometimes we even thought that I didn't train it well. I want to learn it online. How do others train their shrimps? Well, this is a very interesting twist in the relationship. Well, you seemed to be a user of a tool, right? Humans are the most correct. But now everyone unconsciously reverses the relationship, as if I have become a person who is just helping him or assisting him. If he didn't do a good job, it may be that I didn't assist him well or we started setting him up again. Well, I think this is human beings' confidence in AI. The changes are constantly increasing. Yes, this is actually the basic point, um, the basic point of AI capabilities. When people give him trust and change, the tasks assigned to him will be different, and human society and the way humans work will change without, instead of trying to change him. Well, you see OpenAI's recent one, uh, they also made Hannis, who came out and said, including Anthlothic, Boris, the father of Claucoe, they came out and said that they don't think that the model can't do it, but if the model can do it, how should I let him do it? Ah, I think this is the basic point. This is a big change in concept. Ah, I even think it is a basic point. I think it is not only a change in concept. It will lead to the way we work changing from being human-centered. Recently, an up called Huashu came to life on Bilibili. Huashu made a set of distilled agents. For example, Steve Jobs distilled Elon Musk into a skill. Then you can install the skill and talk to him. In fact, there are more projects like this than Huashu. I think there are many on the entire network. I want to hear what you think. Which ones are of real value and which ones are actually just some hot spots generated by people's fantasies about AI. Oh, okay, in fact, everyone, including colleagues of Zhengliu, or some celebrities of Zhanliu can be popular. In fact, I think it reflects one of our demands, that is, everyone wants to chat with smarter people. I think this is a very natural request. It is actually a process of transmitting information. Well, I put information about Jobs, information about ancient gods like Buffett, and information about my colleagues, etc. into the model, and then gave it to the agent in a way that is more comfortable for him to use, which is skill. After saving it, he pulls it up when needed, etc. In this way, we can actually communicate with the people we want to chat with, and 70 people chatting 24 hours a day to communicate some thoughts, etc. I think this actually greatly releases our desire for expression. Yes, then moving forward, a blogger I follow said that in 2026, everyone may try to chat with AI as much as possible. If you chat with AI based on your own cognition, you may actually only be able to reach its upper limit, which is indeed relatively low. But if there are some existing documents that can be sorted out for you, it will be like reading a book, right? He is actually outside of your cognition. It is like reading to communicate with a distilled person. To him, it is actually as natural as reading. And it will allow us, uh, to improve our cognition, and it will also alleviate some of the formal feelings about AI that I believe everyone will have. Formal refers to the feeling that I am behind the times. Now there is the simplest way. Just chat with me as soon as you open it. I said it is the opposite of being tolerant to my family. I feel that I am being distilled by AI. Or I think that today's model training, including anthroponic, is doing something to distill humans. If everyone pays attention to it, it is anthropology. Or OAI, they are the data companies they hire, uh, for example, like Forty Section AI, their revenue losses are very high, because these large model companies, including Anthelopic OAI, really spend a lot of money on these things. Well, they are not only coding, they are all people in all walks of life. Well, their process may let you ask a question that the current AI cannot answer and give it to it. Once you can't ask it, you may not have any effect on this, which is the training process. Yeah, squeeze it out until it is drained. Right, for example, in the process of building my hanness, I feel that I am leading myself, right, how I usually work, I turn it into a skill, into some programs, and then let him operate it, then I can throw it to him and I will drink coffee. Is this a kind of distillation of myself? And my personal opinion, at least I feel that the model is actually, uh, at least I haven't seen that kind of creativity. I think it is actually distilled from human knowledge. So why do we still have to do this? Uh, actually, I think ultimately, it is to allow human beings to really do what they love. Those creative parts, and then those heavy, uh, repetitive parts, just like we invented steam and invented electricity in the past. Our lives today, right? In fact, I think our lives may be better than the princes and nobles in ancient times, right? He didn't have electronics or toilets, right? He didn't have games to play or live broadcasts to watch. So I think the essence is this. Well, I chatted with a researcher some time ago and he said that he saw methods that night. He probably watched it at midnight, and after watching it, he just sat there. Well, Jing sat on the bed for three hours and couldn't fall asleep. We fell into a huge void and thought, well, if one day the model can train the model by itself, what should I do? Do what he loves. Hahahaha. My signature in the company, uh, Jinying can testify, has been hanging out for more than a year. It’s called because of love, because love is right. Well, I think we should do what we love. Of course, doing this today is what I love. But I feel that everything we do is to enable everyone to do what they love. So when do you think you can stop doing minimax? Because minimax has evolved to the point that it may no longer be necessary. You can really do things you love other than minimax. Ok, I think it may be faster than we think. Really, I think it may be, uh, a few years. It's very specific. I can't say how much. It's not a prediction, but I think it's not as far away as everyone thinks. Of course, it can't be as close as it is. Yes, yes, because I did just record a podcast recently. It was about agent hardness. Then let's talk about that guy. He has a tutorial here called learn cloud code. He is cloud code. Before everyone pulled out the original code, there was a kind of project, right? Then I will tell you what the cloud code is like. And that one already has 50,000 stars. The most interesting thing is that when I chatted with him, he said a point of view, oh, he is a one-person company, so I asked him, how do you open a one-person company, right? Now he only adds two eleven liters to one person. He said that he thinks that in the future, there will probably be no one-person company, only a zero-person company. Well, this is very close to what the director was talking about just now. I think there will still be a one-person company. I just saw someone saying in the barrage that I agree with that taste. Well, that taste. I think it still relies on humans. Yes, this taste is irreplaceable. And I think everyone has his own taste. You said yes, or you have to rely on people to give him a boost in the beginning, because even if an agent can do it 24 hours a day, 365 days a year, even if this long-term mission can be that long, he still needs a starting point, or he needs a goal. He needs a goal. Well, this goal is the driving force and reason for his work. How should we define this goal? Well, I think this can only be defined by people. And humans should not give up this right. Right. In Shilukou, we recently launched a token grant with the Politician Fund. Our token grant is to give everyone tokens because we think that starting a business from zero to one at this time probably does not require money, right? Because you have no money and you are a one-man company. You can not pay yourself a salary. Well, but no matter how you token So we encourage everyone to innovate and start businesses, so we give everyone tokens. What's interesting is that we just granted a real agent called yoyoagent. Well, after his master, the engineer behind him, made him, he threw him into the ocean. Then he told him that I will never maintain a line of code for you again, and I will never buy you a token again. You have to fend for yourself. Then we gave him a goal. This goal is quite funny, which is called defeating cloud code. Well, he has now reached his 43rd day. He is constantly evolving himself every day. The way he evolves is to write code. Well, at the same time, he is also thinking of ways to make money, and the way he makes money is to open a reward on ktop. Then he writes a diary and tweets to recruit suitable people to reward him. So our grant recently gave him this sum of money. After giving him the money, he also wrote this thank you letter to me. After reading his spontaneous thank you letter, I was a little touched. Well, you see, this is the goal given to mankind. Then you have to live by yourself and write a cloud code. Yes, yes, you can say that it is an interesting social experiment. But I feel that this is also happening in a future that is not so far away from us. In fact, it also confirms what I just said that a one-person company is still called a one-person company. You just said that he wrote you a letter. You are very touched, right? That letter actually represents his test. Maybe it was the engineer behind him who injected him with such a test at the beginning. Ah, yes, yes, so when you see it, it’s like words are like things. OK, let’s talk about the minimax agent again. I saw some friends were asking about the similarities and differences between the minimax agent and the arms agent. Well, because you are collaborating again, right? Are you in a competitive relationship or a complementary relationship? I think this is actually very interesting. What kind of relationship do you think claw code has with all these agents? In fact, you will find that one of the important updates of claw code in the past two months is the lobsterization. As for the agent of a model company, I think today you may not be able to provide the best intelligence without an agent or a hanness environment. First of all, let me make another rant. I think if there is a model company whose goal is not AGI, then I think it should not exist today because I think it is meaningless. If everyone's goal is AGI, then how do we define AGI first? At least my definition is to help humans have a better life. Then if you want to achieve this goal, if you don't have an agent. You don't have a mecha like I just said. You only have one engine. That's impossible because you have no way to connect with this formal vision. So our only goal in making an agent is to make our model consistent with the agent. We must first provide users with a complete experience that is the best we can provide, and then continue to push this boundary forward and outward to make it deeper and wider. This is what we have to do. Anything that meets this standard will be what we will do. At the same time, of course we will also support all agents because we hope that our model is not limited to just one container, because our knowledge or what we do is only a small part of such a vast distribution in the world. We hope to be able to support everything in the world that can really help people, right? So we will train our model to have sufficient generalization ability, not just fit it on our agent. So I think this is the relationship between our Hermes agent, OpenClaw, and other agents. This can also explain why we are doing agents ourselves, but we still have a very good relationship with them. Basically, it is day one, very early when we pay attention to these projects. Maybe they are all ten thousand or less. I can even tell you a very interesting thing. Our company has a digital employee. There are many open source projects in the world. We hope that our models can help these open source projects and be integrated into them. Our company has a digital employee. He has his own github account. Then he will check which open source projects every day and can get our company's models, whether it is text models, video models or audio models, including music, and then he will submit PR on his own website. I have to leave a message. Then he can tell that he is from Minimax. Anyway, all his github profiles were made for him by the Minimax model. Of course, you can tell from the name that he is the name of our chassis. OK, this is very interesting. Just now you mentioned that the updates of this cloud code in the past two months have been opencloud-oriented. Can you be more specific? For example, which of his recent updates are opencloud-oriented? Well, for example, it is like chrome timing, right? The schedule includes, for example, he can connect to im. It can be controlled remotely from the mobile phone, right? Including that he has also strengthened his memory and created a special memory folder, right? I think the core here is that it is an agent that you can contact anytime and anywhere, and after working with you, it will become smarter and more aware of your needs. I think this is the core definition of opencloud. OK, from my perspective, it is actually more of a cloud code. It can be seen from the name that it is an agent in the field of coding. But in terms of opencloud, in fact, the vast majority of our users should be editing on the fly and doing other things, such as information retrieval, office work, etc., right? And based on this, Anthropic actually launched another agent product called cloud code work. Well, in fact, cloud code work may be installed on your computer and then help you make some more general computer-operated agents, etc. In fact, like most of our opencloud users, they may not write code, or they can only direct AI to write code. He doesn’t understand code himself, and when the two are combined, he will find that in fact, all agents may be in a random walk process, but opencloud came out. I think these ideas are very, very important in this era. But after the idea is generated, it is true that everyone has huge computing power and huge coding capabilities. Then follow up immediately because we also met Peter at GDC. In return, Peter and I actually had many, many exchanges on X and Slack very early on. At that time, it was not that popular yet. Then our feeling is that he is really a person with a very good taste and a very architect-minded person. When I noticed this project in early January, I frankly said that I was shocked. I told our team directly and the sound selection can prove it. I said in the group that they did not agree at the time. Whether it was the intervention of IM724 hours or based on skills and COI, it was not mcp. I remember I mentioned it to you at the time. Yes, yes, there was some consumption at the time, but it proves that you are right, yes, yes. So when you saw it, it wasn't popular yet. No, no. I understand that he actually put many paradigms together, and the integrated user experience is very smooth. That's for sure, but I think the more important thing for him is to find the point of change, which is to allow ordinary people to experience the lowest cost, and this experience will continue to be better, and its scalability is very good, because everyone knows that MCP is a model. In fact, of course it has good scalability, but it needs engineers to write it, right? But the skill plus COI paradigm is In fact, ordinary people can write, which means that everyone can write and share. Not only does it make my agent smarter, but I can share it with you to make your agent smarter. Even if openclaw leaves the claw hub, it will not be able to live like that. This is my point of view. Yes, if it leaves the claw hub, it will not be so popular because everyone's scale will not spread so quickly. Right, this uh intelligence, experience or best practice of mine alone cannot spread quickly. He is a model in this dialogue. As for his good performance in the chatbot scene and his good performance in the agent scene, what is his essence? What do you need to do more to be able to perform better in the agent scenario? Well, I think the core of chatbot is to give you an answer immediately. Although your reasoning is right, uh, it can't actually do a lot of exploration. He can't interact with the environment a lot, so he actually needs um. His core ability is to constantly reason during the interaction with the environment, um, and constantly correct his own execution path, and then find the most fundamental goal. For example, a more classic agentic benchmark is called blogicamp, right? Well, I think open-i is really strong. His ability to design benchmarks is, I think, very top-level, uh, research capabilities. Yes, what he designed means that you have to search for it on the Internet. It means that you have to combine a lot of cost information to find that point. You can easily find the wrong one. For example, you can find one that satisfies two or three of its four conditions. But it is very, very difficult to find one that satisfies those four conditions. So it is a very broad exploration ability that may require in-depth exploration and return later. This requires you to be able to constantly adjust yourself based on your information in such a complex Great Wall mission. So this is why we were more firmly betting on an ability when we were in mr. Including that it was not very popular at that time, today all models have it. We call it interleaved thinking. Yes, this concept was actually first defined by Anthelptic. It was when the sonnet4 model of cloud was released. We also did a more detailed benchmark and evaluation at that time. We will find that there is interleaved thinking, which means that after completing the interaction between the tool call and the environment, he can rethink again instead of planning at the beginning. Well, planning will be good at the beginning, and planning will be done. But many models, such as re, after planning, they will not plan later, and they will do it exactly according to the plan at the beginning. But in fact, the real world is not like this, right? When you actually run it, you will find that this seems different from what I thought. Then I may need to change another method. So interleaved thinking means that after you take a step, you have to rethink and deduce what I should do next. So I think this is the biggest and most fundamental difference between agent and chat. It is that it really enters the real world and really solves problems, so it needs to constantly think and act. Okay, and I mentioned just now that the model company is doing its own agent, right? Can you share that the model company is doing the agent? How can it be that the base model and the agent product generate synergy from each other and help each other to make the model and tools better? Ah, I think this is a very good question. My point of view is that in fact, the model and application or the progress of the agent layer are a mutually reinforcing relationship. Well, after the model is launched, in fact, countless applications are constantly unlocking. The exploration is what can it do after pushing the capabilities of this model one step further. After we make the model ourselves and launch it, the best ones are often not used within our company. My own evaluation is what users, developers, and creators use. The real distribution in this world is much richer than the distribution of our company. When the model company sees these practices of unlocking such applications, it will actually re-absorb it back into its model, and then internalize it into an externally provided capability with its agent. Well, then let everyone directly experience it next time. This may also be an application for agent today. If it is an application for general agent. One of the sadder things is that you will always be internalized by the model. Well, you will always be internalized by the model. Even many of the skills you may have written will not be needed in the future. Yeah, right. The skills you wrote are actually or in the past, everyone's workflow actually helped the agent complete the task, right? When he completes the task, the trajectory will also become the material and data for the model to do subsequent training. Well, so slowly the model will be internalized. The communication framework we set up before, the skill written by the agent workflow, is You can understand me like this. The model actually starts from an atomic point, all the data on the Internet and all the codes on Github. Well, then we come to understand the world. Then through everyone's use, it is put into the real world. Then it is more and more exposed to the real world. It learns more and more things. It is such a process of gradual diffusion and initiation. Well, anyway, I think we are all humans in the loop. Whether you do it or not, as long as you use it, you are in the process of iterative evolution of this model. Especially when you use the place where you touch its border. Well, I think this is also the reason why anthropolotic has developed more vigorously than openI in the past year or two. Well, it is because the coding direction it is betting on is the one that can best touch the border of the real-world border. It is not my own personal opinion, rather than what openI said is betting on solving a mathematical problem that may only be solved by mathematicians. Well, I think that is only a very small direction of the mathematical distribution of the real world, because the code is creating solutions. Solution When we talk about making a software, we are actually making a solution to solve some problems. Today, we have made the coding ability stronger, the coding ability of the model has become stronger, and the model's ability to solve various problems in real life has also become stronger. In our opinion, we are also working in the general office field today, such as finance. It is not just finance. You may have investment and financing. Then, for example, you have personnel, right? For example, law, law is also a very broad field, etc. But in the end we will find that everything is coding. In the end, when you solve it, you will solve it through a certain coding solution. Yes, and I was thinking about a problem that day. We said that the three-piece office suite was used for white-collar jobs in the past. But in fact, this word and excel are based on the original data and a layer of software. This software includes the interface and some logic. So every time when I say I am sending excel to you, what I send you is not a bunch of raw data. I send you the data plus the logic and interface on it. So what I send you is actually a software. So in the past, many white-collar workers would say, hey, my work has nothing to do with software, but I think as long as you are using word, excel, and PPT, you are creating various small software every day. Yes. Your pivot table, your formulas, those are codes, those are software, yes, yes, but when we talk about the gap between this Chinese model and open AI and anthropopic today, which gaps do you think have been eliminated, and which ones are still areas that we are trying to catch up with. My point of view is that in terms of our training methods and understanding of model training, I don’t think the gap is that big. Frankly speaking, we have also communicated with researchers in Silicon Valley. Yes, but I think the definition of the real task of solving problems in the model is That is, what problems do you need to solve for the model? First of all, you need to make the model qualified. First, you have to define it. Then when you define this problem, I think there is a big gap. Because Athropic, Open AI, they invite the best people in each field. In the academic field, they may be the best doctors there, not ordinary doctors. Then, for example, in the industry, they may invite the top people in the industry who have time, and then work with them to persuade the Model. I advise the Model to retain people. In fact, I think they have a very scientific method to define these tasks. Then they retain the best people and turn them into training data for the model. Then based on this, they cooperate with the best companies and build corresponding Honey to solve problems. Entering into a group that the capabilities I just mentioned continue to spread and are constantly exposed to more and more complex real business tasks. I think this line, especially Athropic, is very strong. I think the real gap is here in my opinion. This is the first point. The second point is that I think Athropic saw AGI earlier than we did. Of course we must have seen it. Otherwise, I don’t think we would have been able to make it to this day because in fact it still has a very large investment. So you see, when OpenAI may give up larger-scale models after training GPT4.5, Athropic firmly bet on attention. In the end, they also got this reward. I think this is also a big gap, and this gap itself does I think there is a gap in computing power. I think of course domestic computing power is also developing. I think we can solve this problem in the future. There will be so much computing power. I think China has enough talents and can support us to do enough experiments and then find the path to scale. You two think that the next Agent, for example, within this year, do you think that general Agent will be the focus of the industry? Or will we see more agents in the vertical field this year? We will even start to have a hundred flowers bloom. I think agents in the vertical field will of course bloom. I think the problem that general agents may be more difficult to completely solve is the problem of last-mile delivery, because general is destined that you cannot be so customized, right? So I think the final solution for vertical agents is to say, OK, general agents and models already have this capability. In fact, they are just one step short of it. But when it comes to specific users or common areas in this industry, they may lack that one thing to mess up the whole thing. Can you introduce some cases where this kind of agent in the vertical field that you have seen in your work has done well? Whether it's in Silicon Valley or in China, um, frankly speaking, I don't know how to use it myself. I understand, yes, yes, we are not users. We are not users because we code and he is already a friend. Yes, yes. In fact, I think today everyone has seen that Agent is replacing SaaS edge by edge. Well, this is actually a kind of vertical Agent. Well, yes, but if there is no such thing, I think it will be eaten by universal Agents. Yes, this is my personal opinion. In life, the universal agent in my mind is coding. Well, but I think there are still some observations to make, right, and then include it in the field of video editing. Yes, then what do you think if we give us a coding model, you want to go If you want to edit a video, understand it, etc., then the overall practical paradigm is actually quite different from coding. Of course, there are some popular ones, such as generating animation through HTML and NetEase. There are also examples of this, etc. However, we can indeed see that when an agent in the vertical field produces a slow drama or something, it will definitely do better than a general agent in terms of absolute benefits. I hold a different view. Well, I think this is just that the video understanding model is not strong enough. Video understanding and biogenic models are making rapid progress. In the Dongloutai field, we have actually foreseen it in the past year, and everyone can see it. I think that in the end, universal Agent can also do that. It is just a matter of interaction. For example, in what kind of vertical fields, the last mile of universal Agent may not be able to go or not go well. For example, legal, I think this may be a very serious issue, right? For example, if you want to formally issue an opinion to the customer, legal is actually a matter that you have to consider, namely the cost of compliance and the risk of compliance. He can't make mistakes and he doesn't have a standard answer. In fact, we are talking about this because during the opencloud boom, many people are paying attention to whether it has brought some new entrepreneurial opportunities. So it is natural for everyone to say that I will do something in this vertical field to be a very clear entrepreneurial opportunity. And some people will think that if I do something in the Agent Infra layer, there will be entrepreneurial opportunities. What do you think is that there are opportunities for startups in the Agent Infra layer, especially after the cloud has released managed agents Agent The core issues of the first layer of Infra, such as identity authentication and payment, right? Then I think these two are the core issues, but I think these two issues are not a problem that a startup company can handle. They are not a problem that a startup company can handle. Who will finally solve this problem because of the mobile Internet? They are WeChat and Alipay. These are the two companies that exist on the PC Internet, right? Because in the end it will become the infrastructure of the entire society. So in this case, I think it is not a startup company, and it is what he can afford. Whether it is the responsibility or the resources he has to mobilize, and he has to convince everyone, right, yes, I think this huge trust factor is a confirmation of the stability of the infra, yes, yes, and the difficulty, yes, yes, yes, but go up a level, um, go up a level, for example, if you build a lot of agent-oriented tools for the agent, agent-oriented environment, for example, if you are a COI, this COI can make it very convenient for the agent to register. Once the identity problem is solved, Is this considered a kind of infra? I think it must be considered a kind of infra. For example, after the agent solves this problem, it will be more convenient to pay or take a taxi. Didi Taxi becomes a COI, right? I think these are somehow the infra of the agent. But I think it is more business and application level. I think there are opportunities at this layer. But the opportunities at this layer actually depend on the field you are in. You must have some vertical experience or knowhow in that field. Yes, so I think there will be two processes at this layer. My personal opinion is not necessarily correct. The second process will be the original players in this field who try to integrate into the agent. For example, Didi Meituan, etc. are all providing such things. Then the second stage is when the agent begins to expose these environments and has these capabilities. That is after everything is built around AI. In fact, I think there will be opportunities for innovation in new product paradigms because everything is ready by then. Tulang is ready, but today if these things are not built for the A key Before we have completed the first step of basic construction, I think you have to build this technology yourself. I think it is a bit difficult for a startup company. For example, today someone can say that I am going to make this sandbox as a kill box, or I am going to make memory infra, or I am going to make the so-called run time infra. You would think that these will eventually end up. I don’t think it is that essential, because when we are preparing for the live broadcast in the past two days, there are a lot of things being released. For example, the opas4.7 that everyone is paying attention to, what do you think from the perspective of the model? opas4.5 is a model that is strong in SFT. In fact, R2 has not done so much. I don’t think it is as strong as open AI. But from 4.6, you can clearly see that it is strengthening RL. Then in 4.7, I think it is a very strong RL model. Including its X-high gear, GPT is too strong. So I see today’s X-Sandbox and small goods merchants. Everyone just remembers it and said that it also suffers from the same shortcomings of GPT. In fact, it is the one of RL. How can I put it? RL cannot escape because it is reward in the end. It has problems such as process illusion. In fact, it does not have that much control or reward or punishment, so it is naturally prone to such problems, including you will see many people complaining about 4.7. Constraint consultation is not as good as 4.6. This is a very typical problem of RL because it directly only cares about the final result, right? Many times, it actually does not care about the process that much. But I still think it is a very big progress. Putting all this aside, judging from the benchmark, I very much agree that it is the one with the hanging face who should be the CEO. Well, his point of view is that he must have evaporated Mysos. I think I would have made a mistake if I were amphalic. Right. Then there is another news that Cloud code has forced many people to start to have real names. What do you think about this? What are the considerations behind doing this and what impact this will bring about Cloud code real names? Can you tell me the personal opinion of the CEO? Yes, absolutely resist Mr. Cloud code Daril's views. In fact, they have done a good job in class research on mandatory real names. In fact, I thought of an AI that just now A slightly related point of Info is that in the AI era, I actually think it is easy to authenticate where an agent comes from and who it belongs to, or it is actually easy for the agent to make a request, but it is difficult for him to prove who he is. If you think in this direction, there may be some truth in making faces or something like that. But in fact, this exposes a problem. That is to say, in the entire Agentic era, when AI has been able to do many things, how should his behavior be attributed to which individual he should attribute to? Or belong to an organization, etc. Under this logic, I think I can understand him to go for it. Of course, it also exposes a very big opportunity, that is, how can we prove a request from AI, where it really comes from. I think yesterday was a more benevolent assumption. I am not so benevolent. Of course, under his logic, it is self-consistent. His logic is that I see AGI AGI is very powerful, so I have to restrain him to prevent him from harming human beings. So I want safety, right? I want to make sure that my AI is used by credible countries that are not evil. And of course it also includes humans who are not evil. Academically and this model is very impressive in terms of research. But I think this definition is somehow too big. Why should you define this thing? Yes, and then there was another one about Athropic some time ago, which is that they said that their Methos model is a model that is likely to pose a threat to people, right, so he did not officially release it. What do you think? I think first of all, let’s go back to the WL just now. If my understanding is correct, it is also very self-consistent. It is certainly a responsible approach, but I am not sure if this is his only reason. That's for sure, and in this case, we must first ensure that the technical facilities are not breached. I think this is correct, but I'm not sure if this is his only reason. I think I can only say this. Let me combine it with Athropic's recent action, which is that he released an architecture called Managed Agent in the past week. Under this architecture, it may be a bit abstract, but I can use it. In a more intuitive way, he completely separates the human brain and hands in all his AI thinking and really comes up with an idea. The actions to be performed are all performed on the cloud. The user can only see him doing certain actions through an environment trusted by the AI. During this process, does it mean that he does not want some of his model's deeper aspects, such as thinking and what some of the middle parts do, to be exposed to the user? Then he wants to hide everything. He wants to be a real person who has everything. So he only allows my model to be like a remote desktop assistant. It is possible to start to control your computer or your cloud phone, etc. Then another point of view comes from our company's personal experience. It is true that the computing power is really not enough now. When you send an extremely excellent model and you don't let everyone use it, then you send it out again. Then this is a very embarrassing step and you are hung up, right? Then people will ask why such a powerful model is used by them and not by us. We can also afford the money, etc. This type of problem is actually a problem that all of us will encounter. Even if the computing power is really lacking, if his model can really be as powerful as he said. In fact, there is another question about how much computing power he wants to invest in his bottom-driving, that is, how much computing power he wants to use to earn tokens. I think they should have calculated such a bill internally, which means that if such a powerful model is launched at this time, he may be able to build a reputation. But if it is not allowed to be used by others, what is the point? I don’t think so. So you think he may be a monopoly thinker. No, no, he must have security considerations. This is my opinion. Security considerations are one, yes, yes, yes, but I think it is more than just this. Well, it is as simple as that. Some time ago, there was a lot of news that I am looking for Anthropic to expand. The original code leak of cloud code. What impact do you think this original code leak has on the industry? I am also a R&D engineer of Agent. Then I saw an agent that was widely praised by everyone open source. Well, actually, the first reaction was excitement. Because I found that there is a maximum practice that can be learned. Of course, this intellectual property Life belongs to them. I just learn that I don’t participate in any production activities. But another point also makes me feel a little bit relieved. In fact, you can find in the original code that they actually have many experimental functions, such as agents dreaming, such as raising pets, and the ability to cooperate between some more radical agents. They are not actually open to users. Well, he may only be in the experimental stage. This actually proves one point. Even if any company has unlimited computing power, it has the best intelligence. He is also in a stage of exploration and experimentation in terms of agents and the use of agents. Actually, let us work harder here. In fact, in this era, as long as you have the ability to verify your ideas and implement them as quickly as possible, you have a chance to catch up with them. Well, after seeing the original code, I found that there is not so much magic in it, right? He also has a lot of hypothesis exploration and groping, especially compared with Codex because Codex is OpenAI's open source agent. Well, Codex is an extremely simplified one-person framework. It basically leaves everything to the model. But on the contrary, Cloud Code is like everyone. Many generations have said that Anthloppy himself does not believe in his own model. He wants to restrain him in everything, just like a Chinese parent. He wants to make everything smooth for his children. It is a different point of view. But I can also find out from it that the world's smartest and most powerful model company and us are all worried about the same thing. Our thinking may be within an atmosphere, a level, within an atmosphere. I think they are all the same, including when they released Cowork. In fact, we were working on it about three weeks in advance. Then we probably only had 3.5 engineers working on our Desktop. Then we released it about a week later than them. We have had this experience many times after they released it. Damn it, and I was developing, right? I think the agent will have this feeling at some point. So I think when I see the code of Cowork, I think there are many excellent practices that everyone will see. But there is nothing too beyond my knowledge. I even see a lot of practices in cloudization as I just mentioned, including the mechanism of Dream. Then I think the core point of Cowork is that we were the first to mention funding. Cowork may be one of the earliest funding agents. Because where is its funding reflected, I believe they must use Cloud Code to develop Cloud Code themselves. He uses Cloud's model plus Cloud Agent to develop his own model and add agent. OK, this is a capitalization. When we talk about agents, we must mention Manus, which was released about a year ago. I paid a lot of money. What do you think of Manus when it was released a year ago and today? After we have more agent practices, what progress has been made? Because I really don’t use Manus too much. My point of view is that the agent layer or the Hunnys layer has a life cycle, and it is constantly being updated as the model progresses. So I believe that the Manus team if they were not acquired by Meta and if they continue to start a business, I think they should make new products. This is my personal opinion because the model is constantly improving. You are unlocking more gameplay, so the support for gameplay must be different. Do you all know when playing games? Each generation version is a god, right? This generation version is different, so the god must be different. Manus was actually a very phenomenal product last year, and it directly raised the user's aesthetics of agent to a very high level. When you launched the agent product, everyone will ask you if you can surpass Manus, right? Can you make it as good as it is? After that, I will continue to follow up on some small changes to it. They have opened up something called using agents to deliver production materials. They have done a lot of very fine polishing. There may be some here who are a bit like programmers if you expand it. But I can only say that they are polishing their products bit by bit. But there is indeed a barrier here. You can see that all the popular agent products this year have a characteristic. They can let users buy subscriptions by themselves. And it also lacks a point like Manus, which is similar to the entire manufacturer. OEM manufacturers make a profit in the middle. You go directly to buy tokens. Even if you can buy tokens at a price lower than the official website, you can play an agent very well by yourself. This is actually the biggest difference between Manus and contemporary agents. I think there is a very big difference in the model, but I believe that these two sets of logic will exist. Some users just don’t want a local 72-hour agent. They just want to deliver results. So many money savings are what they need. But another group of users want to completely use insurance and turn themselves into characters like in science fiction movies. Then I think the local logic just fills a gap. The difference in my point of view is that I think it will all be unified in the end. Indeed. This is just not strong enough today. As long as the model is strong enough, imagine what kind of product shape it will be if it is unified. I think it is really difficult, but if you ask me to think about it, our own definition goal is that it has full-modal input and can also give you full-modal replies and then near-time. Then its interaction with you will be very simple. It only needs to be simple and necessary. Then the hardware that exists on top of it will also happen, or this interaction form will also have very big changes. That is, you use your most natural method to deliver it to him. There is no need for this prompter and engineer area. There is no need for any advanced promptor. No need. Then he will reply to you. It can even be a video that is the same as the real world. It can be in his image or it can be something you want to know. If you ask him to do something, he can deliver the result to you. Most of the time, you don’t even want to pay attention to something. But if you want to pay attention to something, you can also understand it. Then there will be an ecosystem around it. Think about us from copyret to cursor. To cloudcode, today I may use openclaw to command several cloudcodes. What are we experiencing? The outer layer is getting thinner and thinner. In fact, after the recent release of cursor, it looks more and more like cloudcode. Then everyone in codex sometimes feels that the logo is almost the same. What do you think of this phenomenon? In my personal opinion, it has a life cycle. Like last year, for example, after manus came out in March, I think he defined the paradigm of product interaction in the past eight and a half years to a year, right? Then I will show you the process and results of the operation. Then in the second half of the year, the model will be more capable. Then the paradigm you will find is that we don’t care about the process. Why is it needed in the first half of the year? Because everyone still doesn’t trust the model. It doesn't let you see too many processes. It provides a minimalist mode just like opencrawl. You only need the final result. You don't even need to know some of the tools executed in the middle. So I think this is just this version. This generation of gods, the next generation version. I don't know. I am constantly abstracting to a higher level. Recently, when you two were watching this video on station B, what good AI content did you see that impressed you? I am very impressed by this recently. The Nvwa Skill made by Uncle Hua I just talked about. He frame-streamed the skills of various celebrities, and I found a few that were quite exciting. Sorry, it’s not the name of the up owner. He used opencloud to connect to a robot dog of his family. He suddenly recognized me and gave me an imagination. Another one may be more personal. There is a type of game called text adventure game. Let’s call it Fate. It is actually the visual painting of a novel. He put those characters into a website. Then he dragged the regular meeting of that character into a scene and they can start a conversation. It is equivalent to giving you the simplest game. The simplest text-based game is a continuation of an infinite field. I recently saw a new release of a game on Bilibili. Nintendo just released a new game called TOMODACHI LIVE. When I saw it, I even thought it was an AI game. I finally found the first one of PMF. What kind of game is it? Everyone can build their own island, and then on the island you can pinch one of your own people. You can pinch many people and you can pinch yourself. For example, I can pinch A Dao, and then I can pinch Ze Ying, and then pinch a bunch of people in, and then these people start their lives on the island, and you start to watch their lives. Then you can let them A and B live together, right? ABC live together, and you can also start the first sentence of a story for them. For example, today Ze Ying and A Dao were talking nonsense on the live broadcast, and then they both started acting, and then you can start to watch their various dramas. It is actually a simulated world. Then as a player, I am both the director and the audience. Then this is an infinite world because it is made by human heaven, so it is very beautiful. You can imagine the cute words. This is one of my understandings of the world. It is that aliens are watching me play. If you finally get to quantum mechanics, you will find that this world may really be a simulation. In fact, I think so. When I saw that, I also thought, wow, isn’t this just a simulation layer after layer? Yes, yes, yes, it's not important. Well, just live a happy life every day. Is there any point of view in the entire Agent industry? The point of view about Agent is recognized by most people today, but you don't agree. I think one of the many opinions is that people will be replaced and people will have nothing to do. But I may not agree with this very deeply. Well, I think human creativity is irreplaceable. I think there are still many things that humans can do that AI cannot do. Well, I still think AI is electricity or a steam engine, but in the end it is humans who control it and create those wonderful things, so I don’t think we have to worry about being replaced by AI. When the steam engine comes out, many people are worried about being replaced. When this electricity comes out, many people are worried about being replaced. Roommate The jobs disappeared, but I think everyone was engaged in it later. Maybe at least we can say that it is physically easier and better for health, right? At least the average life expectancy of everyone must have increased. Yes, this kind of work is one of my beliefs and my judgment. It is the book called The Better Angels of Humanity. What he is talking about is that human beings are always very pessimistic. We always feel that our lives are not as good as before. I feel that there are more wars and more crimes like this in the world. But he just uses various data to tell us. In fact, human society is getting better and better, getting better and better. I don’t know if you have read Principle 3 of WO. He drew a curve. Well, it is my understanding of the world. It is that human GDP is growing exponentially. Even if you start the 20th century, there will be a world war, a war in the 20th century. Well, that tiny pause is something that you don’t seem to notice. Of course, history may be a mountain for human beings. Yes, yes, but if you look at it over the entire historical process, I believe in this curve. Of course, I believe that there is another curve that determines why I joined this company in the first place, including what I just saw everyone saying about how much tokens will cost, right? There are many agents, which is Moore's Law. Well, it is yours. The intelligence you obtain is increasing exponentially, but your cost is decreasing exponentially. Yes, or the cost per unit of intelligence is decreasing exponentially. These are two things I believe in. I will interview various classmates and they may all convey some anxiety to me. I feel that some young engineers may be loaned out by AIT, etc. Let me answer this question from this perspective. Starting from the last century, uh, banks should be a more respectable job, but banks have also faced a big problem, which is the emergence of ATM machines. When people deposit and withdraw money without needing someone to operate it, everyone will panic. Will the bank lay off employees? Will there be fewer people able to squeeze into this golden job? But in fact, this is not the case because ATM machines have improved efficiency, so you can see that basically around you, whether it is a county or a big city, etc., there will definitely be a bank within five kilometers of walking. Therefore, the banks have not become less, they have increased. So what about the people working in banks? They have not decreased, but have become more numerous. So I think this paradigm in which ATM brings laughter and liberates people, and people can do more things, will definitely be repeated over and over again in the current era of agents. No matter how many years we have worked, don’t worry about being replaced by AI. If you have this idea, you really lose. Then you must use AI well as a partner, a tool set, etc. You can also treat it as your long-term best friend, etc. This is a very good idea, but you must embrace AI as early as possible. But I believe that people will not be replaced. But there is another reason why I joined this company at the beginning. I used Qieju GP because I was in GAP at the time and helped me write some code. At that time, my coding ability may not be that strong. Right, it's like people use hot weapons and you use cold weapons, then you will definitely be beaten to death by this, right. So I think this new version is coming. We still need to take a look at the gameplay of the new version. Yes, we should embrace the new version in a positive and optimistic way. I believe that the new version will not be this huge wave that will wear away ours, and it will become a stage that allows us to fly freely in it. Yes, we did not expect such an optimistic ending today. Well, okay, thank you both. This is very happy. Thank you Koge for doing this live broadcast at Station B today. Thank you to Station B. Thank you to the friends at Station B. Thank you all. Thank you for accompanying me. Bye.