🌐 This transcript was automatically translated to English from the original.
からあげ帝国放送局始まります。 この配信ではAIの会社で働きながら作家として本を書いたり、個人でメーカーとしてものづくりを楽しむ私からあげが、技術の話であったり個人のスモールビジネス、その他雑多のことをお話ししていく配信です。 今日もですね、いつものように近所の公園を散歩しながら収録しています。 曇りでですね、すごい散歩しやすい天気ですね。 ちょっとこの後天気が崩れるみたいなので、明日とかもですね、散歩できるかちょっと心配ですかね。 最初に少しお知らせというかですね、私は個人で AIアシスタントツールのZanjiというですね、ソフトを開発しているんですけれども、 いわゆるオープンクロー的なソフトですね。 これがですね、Telegramというですね、チャットツールに対応する変更を機能してリリースしましたというところですね。 この変更はですね、豊さんという、多分読み方合っていると思うんですけれども、 豊さんがプルリクエストという変更のですね、リクエストをしてくれて実現した機能になります。 豊さんはですね、先週のNT金沢というイベントにですね、来てくださって、 お話もいろいろできてですね、その時にTelegramの話もしていてですね、 プルリクエストしていいですかみたいな話があったので、 ぜひぜひという形で実現したものになります。 で、このTelegramもですね、私全然あんまり知らなかったんですけれど、 なんかLINEっぽいですね、チャットツールで。 で、実はですね、オープンクローとか、エルメスエージェントとかですね、 そういった今の主流のエージェントって、 結構第一にこのTelegramに対応しているので、 気にはなっていたんですよね。 で、使ってみたらですね、LINEっぽいインターフェースの チャットツールで調べてみたらですね、 なんか中国系だと勝手に思ってたんですけど、 どうも違うみたいで、 もともとロシアの技術者の方がですね、作っていたもので、 なんかその国の正常的にちょっと開発がしづらいので、 今はですね、拠点を点々と落として、 ドバイで開発しているみたいなものみたいですね。 で、結構、なんていうんですかね、仮想通貨のコミュニティとかで 使われていたりしたようなツールだったりして、 世界的にはですね、かなりメジャーな、 だからDiscordとかよりワールドワイドで見ると使われている チャットツールみたいですね。 で、私も少しですね、動作確認のために使ってみたんですけれど、 今のところですね、ちょっと良さがまだ分からないっていうのが正直なところですね。 多分ですね、そもそもチャットツールなので、 周りの人がたくさん使っているかどうかっていうところが、 重要なファクターにはなると思うんですけれど、 周りでですね、テレグラム使っている人、 一人も今のところですね、少なくとも一緒に使おうぜっていう人はいないので、 正直良さは分からないっていうところですかね。 なので、もしですね、テレグラム詳しい方とか、 こうやって使うと便利だよみたいなことを持っている方がいたらですね、 情報をいただけると嬉しく思いますというところですね。 で、今日の本題に入っていこうかなと思うんですけれど、 今日はですね、AIによる行き過ぎた自動化の果てにあるものについて話そうかなと思います。 この話のきっかけはですね、ゆる言語学ラジオのMCのですね、堀本さんですね。 で、その堀本さんがやっているですね、 ゆるコンピューター学ラジオという番組があるんですけれど、 いわゆるコンピューターサイエンスのいろんな話を取り扱ったような、 ポッドキャストとかYouTubeで配信しているコンテンツですね。 で、ここで個人情報は捨てましょうみたいなタイトルでですね、 クロードコードですかね、を使ったですね、 様々な作業を自動化している話をですね、 している回が最近ありました。 また概要欄からリンクは貼っておこうと思うんですけれど、 これがですね、すごい共感というかですね、 まさにですね、私もクロードコードではなくて自作のですね、 ザンギというソフトでやっていることとですね、 重なることが非常に多かったですね。 で、堀本さんはですね、個人情報とかは一切、 保護を捨ててですね、 全部自分のことをAIに教え込ませて、 さまざまな作業の自動化とかもこなしているみたいな話をしてですね、 育てるのにも300時間って言ってたかな、 それだけ時間かけて調教したから、 すごい便利に使えてますよって話がありましたね。 で、まさに私もですね、 ずっと同じようなことをやっていてですね、 同じワークスペースをずっと使い続けて、 そこに自分のあらゆる情報ですね。 However, in my case, I don't include my real name or any work information, but rather my activities as Karaage. Regarding that, we input all the information from blogs, past Twitter archives, etc., and automate various tasks by turning them into skills. However, what they were doing was pretty similar, so I felt a great deal of sympathy for them. However, to be honest, the amount of time I spend working on it is quite a bit longer. As for the AI tools themselves, I think the difference is that we create them ourselves and develop them while improving them little by little. However, to be honest, this is just a hobby of mine, and since I work in the field of AI, I'm also learning about it, so in terms of cost performance, I can't say it's good at all, so for regular use, I'd say something like Claude Code or Codex.I'd be happy to have someone else use it, but I'm still using the tools I like. I think it's good value for money to switch to something better when it comes out, but if you're interested in AI or like to develop your own software, I think it would be interesting to create and develop your own AI agents, assistants, and even the software itself, so I personally recommend it. So, we've been automating a lot of work, and finally, yesterday, we achieved complete automation of the audio distribution of this podcast, from recording to distribution.We've achieved this, so I'd like to report it here, and I'm proud of it. Well, this was also triggered by the fact that the branch manager of my audio distribution team told me from a manufacturing perspective, on a program, that he was automating his own audio distribution to a large extent. He also talked about wanting to use Claude's code, so this is something I can't give up on. I also worked on automation by adding audio. So, the branch manager said that since they are using Voysie for the final registration, they have not been able to automate that part, so I am aiming to go beyond that and have achieved complete automation. So, to be more specific, the premise is that automation was already possible to some extent in the first place. It feels like semi-automation. So, how do you do it? Gemini, Google's Gemini Gem, well, it's a skill-like function. It's a certain amount of work, and there is a function to automate it by summarizing it in prompts, and I created two gems to automate the work. So, one is Gem, which transcribes the transcription and summarizes the contents of the transcription. So, first of all, when you upload an audio file, it will be automatically transcribed and you will get an overview of what kind of distribution it will be. This will be used as an overview when subscribing to the podcast, so if you can get it. They also suggest titles. And after that, the other gem is the artwork. We are making gems that can be used to create cover images. So, this is my character sheet, and I've already loaded the color game teacher's icon into Gem. So, the prompt is a square image like this, and it's a gem that asks you to write according to the content, so I just asked you to paste the transcribed summary and generate an image. It's nanobana. We asked the image generation AI to draw the artwork image for us. Now that you have the images and the outline of their distribution, all you have to do is log in to Spotify in your browser, paste the outline and cover image, and click here, because I use Spotify to distribute to various platforms. Enter the episode number and press the broadcast button. Also, this is a personal work, but as a backup, I'm also saving the transcribed text data, audio data, and artwork in Notion. This series of tasks is also highly automated, so the total work time is not that long. I would wait for a while, do a few clicks, wait a while, and then the image would appear, and if I didn't like it, I would edit it, and if there was no problem, I would just click the same button and upload it.Well, of course, I had some work to do here and there, so if I didn't have time, I would forget about it, or I wouldn't have time to switch things up, so I recorded it, but... It seems like I couldn't upload anything until midnight, and it's been happening quite often lately, and it was actually causing me stress. So, even if we try to automate it further, for example, with GEM, we have experimented with merging it into one, but as the prompt becomes larger, the performance decreases, and although it works well, problems such as unstable images occur. Also, things like posting on Spotify, etc. GEM doesn't allow you to do things like operate the browser at all, so it's inevitable that you won't be able to automate things like that. This workflow was also relatively stable, so I didn't automate this part, but now it's time to do it. Since this is a daily task, I thought I'd like to automate it, so I gave it a try. The software I used was Zanji, which I introduced at the beginning, which I made myself. However, I think it would be possible to do something similar with Claude Code or Codex. So, in the first place, I used GEM to transcribe the text and summarize it, and since I had already prepared prompts for things like creating artwork, I decided to just use them as skills and let them do the work. One of the bottlenecks is that I originally used it for image generation, but it was free to use like Nano Banana, but there was a problem that there were not many high quality products, but recently, Glock has become able to generate images, and I, X, am paying a fee to Twitter, so I can use this to a certain extent. So, this image generation of Glock can be called from something called Hermes Agent, so it's a bit roundabout, but Zangi uses the skill to call Hermes Agent, and from there it generates an image of Glock, so image generation can also be automated. After that, you have to click on the browser and register it, which I thought would be quite difficult, but it's called browser use, and there's a technology, AI, that operates the browser on a click-by-click basis. So, with this Zangi, in addition to the computer I usually use, I use something like NVIDIA's small AI supercomputer called DGX Spark, and there, I have achieved automation to some extent by running the browser, clicking the button, and having people post. Well, I don't look at the screen at all, and to begin with, this DGX Spark is not normally connected to a display. Well, if you ask them to access it using a browser and register, it will automatically launch the browser on their computer and start posting. I wondered if it was actually possible, but it was posted properly, so I decided not to do it anymore. So, by turning this browser operation part into a skill, there will be a skill to generate an overview, a skill to generate images, and a skill to post to the browser. So, by creating something like an integrated skill that calls each of these three, and having them work in sequence, it's like automating posting. Now that we've achieved complete automation, all I have to do is pass the audio to the AI, and the AI will handle everything, including the distribution. Well, it does take some time, but the good thing is that as long as you give them the audio file, you can leave it alone. During that time, I can do other tasks without having to worry about anything, so up until now, while I was working, I had to occasionally stop and switch things up, do some work, and then go back to the original work.With this, I often ended up forgetting things and leaving them unattended, but personally, I'm glad that this is no longer the case. The amount of time required to do the work is very small, so in a sense there is a sense of self-satisfaction on the part of the individual, and when you try to do it at work, it's hard to show the effects because it's not very efficient when you look at the time saved, but I think it's a very satisfying job for individuals to do. Also, I didn't put any background music on the handles or anything, so I think it was easy for me to do it because there wasn't a lot of background music. I think it would be more difficult for people who need to add background music, but I don't think it's impossible, so I also think it would be possible to have a stream that automatically inserts fancy background music if you feel like it.The question I have for this stream is how many people are looking for cool background music, so I guess it's just a matter of feeling like it. If anyone has any requests, I would be happy to hear your comments. In this way, we can automate more and more things, but there is always the question of what will happen if we automate too much. Actually, as I mentioned before, I used to write a diary every day, but lately I've been inputting a lot of things I did and ate into the AI, and the AI writes the diary on its own. So, at first, I was planning to leave the automated diary to the AI, but I would write in a diary every day, handwriting the other parts that couldn't be automated, but lately, I've become somewhat satisfied with the AI diary, and I've started to think that this is fine, so I'm actually scared. In the past, I used to write things like this on my blog that helped me organize my thoughts and lead to new ideas, but now I'm leaving everything to AI, and I wonder where all the things I said back then have gone. I'm also worried that I might lose something important. Also, you know, once you start automating things like taking a walk and recording audio, you'll really be done with it. I wonder how I can automate it. If you create a robot or something, have it walk around, talk while walking, record it, and distribute it as a podcast, that's really the end of it. With physical AI, everything has been taken away from me, and I feel like I'll just be crying like a baby at home, but if I let my guard down, if I can really do that, I'm afraid that I'll be satisfied with just that, so I'm wondering what I'm going to continue doing, and I wonder what it's okay to automate. If we don't let our guard down, everything will become automatic, so these days I feel like I need to reconsider the rules, or put a stop to it. Anyway, today's post is getting a bit long, so I'd like to end it without replying to comments. We look forward to hearing from you, so please keep sending them in. See you soon.