Cartesia completes its voice stack with Sonic 3.5 (TTS) and Ink 2 (STT)

Cartesia launches Sonic 3.5 for text-to-speech and Ink 2 for speech-to-text, topping Artificial Analysis rankings with low-latency voice agent models.

With Sonic 3.5 and Ink 2, Cartesia covers both ends of the real-time voice agent chain: speech synthesis (TTS) on one side, transcription (STT) on the other. Founded by Stanford researchers who originated State Space Models, the company claims the top spot for both on Artificial Analysis' streaming rankings, and positions itself as the only provider to simultaneously hold the top position for both speaking and listening.

Sonic 3.5, presented as the most natural voice model, boasts a latency of under 90 ms to the first sound and native support for 42 languages. It correctly reads numbers, codes, emails, and phone numbers without prior setup, and resolves English heteronyms based on context. Ink 2, on its part, focuses on accuracy and end-of-turn detection: it semantically identifies when a person has finished speaking, without a separate voice activity detector, and emits a complete cycle of turn-taking events. It also natively handles structured data like numbers and dates. At this stage, however, Ink 2 only processes English, whereas Sonic 3.5 covers dozens of languages.

Both models rely on Cartesia's SSM architecture, which the company contrasts with transformers for its low streaming latency. Sonic 3.5 and Ink 2 are accessible via Cartesia's API and playground.