Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
MarkTechPost
Read Full Article at MarkTechPost →Ad Slot — In-Article (728x90)
Cartesia has released Sonic-3. 6, a streaming text-to-speech model built on state space models rather than transformers.
It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to isolate the synthesis engine. Cartesia states sub-90ms time-to-first-audio.
This is a summary. For the full story, read the original article at MarkTechPost.
Original source: MarkTechPost