Blog / Product

Introducing Sonic-3.6

Cartesia TTS Research 
A hand-drawn outlined sheet of paper overlapping a green watercolour speech bubble marked 3.6

At Cartesia, we’re deeply invested in making foundational advances to AI architectures and algorithms, because that’s what drives step changes in model capabilities.

Sonic-3.6 is the culmination of our latest advances in architectures that learn efficiently from multilingual audio across pre- and post-training. Sonic-3.6 makes huge strides compared to Sonic-3.5 in naturalness and quality, with listeners preferring it in up to 93% of blind head-to-head tests across fifteen locales.

At the start of this year, we made a bet that instead of incrementally adjusting the existing paradigm, the way to do the best research in the world was to rethink everything from first principles. Since then, we’ve rebuilt our data, model architecture, training and evaluation from the ground up with novel ideas, and we’re seeing those efforts pay off.

Sonic-3.6 lands just two months after Sonic-3.5, and it once again takes #1 on the Artificial Analysis leaderboard across both the controlled and provider voice boards. On the controlled board, the incumbent model we beat is Sonic-3.5. Overtaking our own model while no other provider has closed the gap is a testament to our accelerating research velocity.

Artificial Analysis controlled voice leaderboard, fifteen text-to-speech models ranked by Elo. Cartesia Sonic 3.6 is first on 1120 and Sonic 3.5 second on 1095, ahead of ElevenLabs Eleven v3 on 1066
Our lead over the nearest competitor, at #3, is wider than the gap from #3 down to #12.

Closing the gap to natural speech

For anyone running voice at scale, naturalness decides whether a customer stays on the line. A voice that sounds even slightly synthetic makes callers lose trust and ask for a human. We built Sonic-3.6 to sound natural and native enough to keep the conversation moving across 44 languages.

In US English, listeners chose Sonic-3.6 over Eleven v3 92% of the time in blind head-to-head tests, thanks to better intonation, higher voice quality, and context-driven emotionality. Sonic-3.6 reads a line the way it’s meant to be heard and adapts delivery to the conversation’s context the way a person does.

Text: I understand things are difficult right now. We can start with a smaller amount, around $35 a month, and adjust it later. There's no pressure to commit to more than you can manage.

Eleven v3

Sonic-3.6

¿Hablas español?

We’ve always viewed English as table stakes, but sounding native in every other language or accent is the real frontier that decides whether voice AI stays a mostly-English technology or becomes something the whole world can actually talk to. Most models are intelligible abroad, but few are convincing enough to get the accent, tone, and local pronunciation of a name or a place right.

We tested Sonic-3.6 the way a customer would judge it: by ear, in each language, against leading competitors. Native-speaker panels preferred Sonic-3.6 head-to-head in every core language we tested.

Native-speaker preference

Blind head-to-head against the leading competitor in each locale.

Competitor
Sonic-3.6
English vs Eleven v34%92%
English vs Eleven v313%72%
English vs Eleven v33%77%
Spanish vs Eleven v316%70%
Spanish vs Eleven v36%91%
Portuguese vs Eleven v327%67%
French vs Eleven v312%86%
Hindi vs Gemini 3.125%55%

Methodology note: For each language we used the leading competitor and matched its voice to a comparable Sonic-3.6 voice, generated audio from identical transcripts, then asked native-speaker panels to blind-rate each pair. Vote share = a model's wins ÷ total valid votes

What these improvements unlock across languages:

  • Broader reach: 61 locales (11 of them Indic) and 500+ preset voices now live, plus two new languages: Odia and Urdu, both launching above 96% transcript accuracy.
  • Hinglish support: Code-switch between Hindi and English in a single generation using transcripts in Devanagari, Latin script, or a mix.
  • Improved accent adherence for instant voice clones: Sonic-3.6 holds onto a speaker’s accent far better, including less widely spoken ones.
  • Locale-awareness: Set the optional locale field so that dates, times, and numbers come out the way a local would say them.

Seamlessly code-switches between Hindi and English:

Text: हमारी टीम ने पिछले quarter में customer feedback के आधार पर 12/08/2026 को नया mobile app launch किया, जिससे revenue में लगभग पंद्रह प्रतिशत की growth देखी गई।

Reads the date the way an Australian would:

Text: Right, so I've confirmed your Qantas booking, QF11, departing Sydney on 21/07/2026. You're in seat 32A, you've got one checked bag at twenty-three kilos, and check-in opens three hours before the 9:50am departure.

Building for production

Naturalness may win in the demo, but we know the live call is the real test. Built for enterprise workloads at scale, Sonic-3.6:

  • Replies under 90ms and generates nearly 2x faster than v3 Conversational (132 vs 68 characters/sec).
  • Runs at scale across cloud, on-prem, and localized endpoints, backed by a 99.9% uptime SLA.
  • Powers millions of production calls today with the reliability that volume demands.
  • Deploys in your own cloud through Baseten, Amazon SageMaker, or Together AI, so audio never has to leave your infrastructure.
  • Meets enterprise compliance, with SOC 2, PCI, HIPAA, and GDPR covered.
  • Fits your brand with hands-on help from our team to develop and find your custom voice.

Hear it for yourself

Sonic-3.6 is now generally available in the playground or via the API.

Try it soon, though. At the rate we’re going, we might have a new one out before you finish integrating.

Hear Sonic-3.6 for yourself

Try Cartesia

Architecting AI that learns and interacts like humans.

Status