← Back to search
Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas — 2026-08-18
Impact Vector: AI Tools · 2026-08-18 · 3 min
Show full episode description
## Short Segments ByteDance Seed and Tsinghua AIR have unveiled CUDA Agent, a reinforcement learning system that optimizes GPU kernel generation. This system trains a large language model to write faster CUDA kernels, outperforming traditional compilers. On the KernelBench benchmark, CUDA Agent achieves a 98.8% pass rate and a 96.8% success rate in generating faster kernels than the torch.compile method. While the trained agent isn't publicly available, the system's components, such as the CUDA-Agent-Ops-6K dataset, are accessible for mid-size teams to integrate into their workflows. This development is significant for teams looking to enhance computational efficiency in deep learning infrastructure. Meet SAM, the Sovereign Agent Mesh, a zero-config, zero-trust P2P network for AI agents. This Apache-2.0 project allows autonomous AI agents to share tools securely without exposing internal scripts or APIs to the public internet. SAM operates like a private VPN, enabling agent-to-agent tool sharing over the Model Context Protocol. While still in beta, SAM offers Go binaries, Docker images, and a Kubernetes deployment guide, making it suitable for mid-market and enterprise engineering organizations. This innovation is crucial for teams managing agents across multiple network boundaries, enhancing security and efficiency. Nous Research introduces Bot Mode for Hermes Agent, transforming agent profiles into a roster of named bots. This feature allows each bot to have its own chat, memory, skills, and pinned model, facilitating communication through a persistent Agent Inbox. Bot Mode is now bundled and default-on in Hermes Desktop, available at no license cost. It's ideal for solo builders, startups, and small-to-mid engineering teams, offering a flexible tool for managing multi-model agent workflows. Enterprises, however, should consider it a workstation tool due to the lack of centralized management features. ## Feature Story Cartesia's Sonic-3.6 text-to-speech model now leads both Artificial Analysis speech arenas, setting a new standard in real-time TTS technology. Released just three months after Sonic-3.5, Sonic-3.6 achieves top scores on both the Provider Voice and Controlled Voice leaderboards, with the latter being particularly noteworthy as it isolates the synthesis engine from the voice catalog. This advancement is attributed to its state space model architecture, which delivers sub-90ms time-to-first-audio, enhancing naturalness and responsiveness. Available in beta as a hosted API, Sonic-3.6 is not open-source, requiring users to rent the service rather than self-hosting. Its deployment spans various industries, including financial services, healthcare, and e-commerce, catering to solo developers, startups, and large enterprises alike. As Sonic-3.6 sets a new benchmark in TTS performance, it highlights the growing importance of natural and efficient speech synthesis in diverse applications, from customer service to content creation. Looking ahead, the focus will likely be on further refining the model's capabilities and expanding its accessibility to a broader range of users and industries.
✨ Episode Outline — click any point to jump to it in the episode
Problem solved
Covering new agent infrastructure and Cartesia's Sonic-3.6 leading both speech synthesis leaderboards.
Benefits
- Sonic-3.6 leads both Artificial Analysis speech arenas
- Sub-90ms time-to-first-audio via state-space architecture
- SAM enables zero-trust P2P agent tool sharing
- Hermes bot mode gives each bot own chat, memory, skills
- CUDA Agent writes faster kernels than Torch.compile
Use cases
- CUDA Agent: 98.8% pass rate, 96.8% faster-kernel success on KernelBench
- Sonic-3.6 deployed across financial services, healthcare, e-commerce
- SAM ships Go binaries, Docker images, Kubernetes deployment guide
- Hermes bot mode bundled and default-on in desktop at no license cost
- Sonic-3.6 released three months after Sonic 3.5
KPIs / results
- 98.8% KernelBench pass rate, 96.8% faster-kernel success
- Sub-90ms time-to-first-audio
- Top scores on provider and controlled voice leaderboards
Tools / build
- Cartesia Sonic-3.6 TTS API
- CUDA Agent + CUDA Agent Ops 6K dataset
- SAM (Sovereign Agent Mesh)
- Hermes Agent bot mode
ByteDance Seed and Xinhua AIR have unveiled CUDA Agent, a reinforcement learning system that optimizes GPU kernel generation. This system trains a large language model to write faster CUDA kernels, outperforming traditional compilers. On the KernelBench benchmark, CUDA Agent achieves a 98.8% pass rate and a 96.8% success rate in generating faster kernels than the Torch.compile method. While the trained agent isn't publicly available, the system's components, such as the CUDA Agent Ops 6K dataset, are accessible for mid-sized teams to integrate into their workflows. This development is significant for teams looking to enhance computational efficiency in deep learning infrastructure. Meet SAM, the Sovereign Agent Mesh, a zero-config, zero-trust P2P network for AI agents. This Apache 2.0 project allows autonomous AI agents to share tools securely without exposing internal scripts or APIs to the public internet. SAM operates like a private VPN, enabling agent-to-agent tool sharing over the model context protocol. While still in beta, SAM offers Go binaries, Docker images, and a Kubernetes deployment guide, making it suitable for mid-market and enterprise engineering organizations. This innovation is crucial for teams managing agents across multiple network boundaries, enhancing security and efficiency. Noose Research introduces bot mode for Hermes Agent, transforming agent profiles into a roster of named bots. This feature allows each bot to have its own chat, memory, skills, and pinned model, facilitating communication through a persistent agent inbox. Bot mode is now bundled and default on in Hermes desktop, available at no license cost. It's ideal for solo builders, startups, and small-to-mid engineering teams, offering a flexible tool for managing multi-model agent workflows. Enterprises, however, should consider it a workstation tool due to the lack of centralized management features. Cartesia's Sonic 3.6 text-to-speech model now leads both artificial analysis speech arenas, setting a new standard in real-time TTS technology. Released just three months after Sonic 3.5, Sonic 3.6 achieves top scores on both the provider voice and controlled voice leaderboards, with the latter being particularly noteworthy as it isolates the synthesis engine from the voice catalog. This advancement is attributed to its state-space model architecture, which delivers sub-90 milliseconds time-to-first audio, enhancing naturalness and responsiveness. Available in beta as a hosted API, Sonic 3.6 is not open source, requiring users to rent the service rather than self-hosting. Its deployment spans various industries, including financial services, healthcare, and e-commerce, catering to solo developers, startups, and large enterprises alike. As Sonic 3.6 sets a new benchmark in TTS performance, it highlights the growing importance of natural and efficient speech synthesis in diverse applications, from customer service to content creation. Looking ahead, the focus will likely be on further refining the model's capabilities and expanding its accessibility to a broader range of users and industries. Let's see. We'll see you next time. We'll see you next time.