Customer stories

Speko's voice AI benchmark recommends Sonic 3.5 most often

Speko is a router for voice AI. Developers build and run voice agents through one Speko API, and Speko picks the speech-to-text, LLM, and text-to-speech models for each call based on its language and use case. Those picks come from a benchmark Speko re-runs every two to three days across roughly a dozen providers. On the text-to-speech side, the model that comes out on top most often is Cartesia’s Sonic 3.5.

We are building an open router of voice AI basically. We are trying to be this neutral layer where we help people to pick the right model stack.

Beknazar AbdikamalovFounder, Speko

The challenge: three model decisions and nobody neutral to ask

Roughly 95% of voice AI runs cascaded: speech-to-text, then an LLM, then text-to-speech. That is three separate model decisions per agent, and most of the providers with an opinion on them have a stake in the answer. New providers keep arriving, so there are more options every month without more clarity about which ones hold up.

It gets harder outside English. A model that reads well in English can drift or flatten in another language, and there is very little public data on which ones do. Companies large enough to run their own evaluations already know which stacks work where. Everyone else picks from marketing pages.

The solution: a benchmark that refreshes every two to three days

Speko evaluates in phases, so a new model gets cheap automated checks before anyone spends real money on it. Agents watch X and Hugging Face for promising releases, and anything they surface goes through those checks first. Models that do well move on to human evaluation and higher-budget testing. Roughly a dozen providers are tracked at any given time.

Re-running the benchmark every two to three days keeps the numbers on current model versions rather than on whatever shipped last quarter. Each category has its own criteria. Text-to-speech has three:

  • Intelligibility: does the model read the text in a context-aware, sensible, production-usable way, beyond getting the words right?
  • Drift: does the voice stay consistent across a generation, or do quality, tone, and timing wander?
  • Naturalness: does the voice sound human, or are there robotic qualities that break the experience?

Scoring the three separately is what makes a recommendation specific to the use case. Abdikamalov puts most customers in one of two groups. Customer support agents care about latency and measure success by fast, accurate completion. AI coaches and therapy products need voices that feel human. Most providers are good at one or the other.

Why Speko chose Cartesia

Sonic 3.5 leads Speko’s overall text-to-speech ranking, with the strongest combination of speed and naturalness, and it holds up on price against the other models Speko tracks. Those two axes cover most of what Speko’s customers build: latency for support, naturalness for coaching and therapy. Being strong on both is rare, which is why Sonic 3.5 is Speko’s most frequent text-to-speech recommendation.

Default here means most frequently recommended, not automatically selected. When a customer’s requirements point elsewhere, Speko routes them elsewhere. The recommendation follows the use case; Sonic happens to win a lot of those cases on the criteria that matter.

Looking ahead: making model selection invisible

Speko is working on two things at once. The first is making model selection something customers never have to think about. Abdikamalov does not want the right stack to depend on how well a customer understands the models underneath it.

To make people ignorant in terms of model selection so that, I think, ideally they shouldn’t even know what’s under the hood.

Beknazar AbdikamalovFounder, Speko

The second is becoming the benchmark the field trusts. Plenty of leaderboards rank voice models without publishing enough methodology to check the results, and others lean on narrow demos that say little about production behavior. Speko wants a benchmark that refreshes on a schedule, publishes how it works, and tests the things customers actually build.

Abdikamalov estimates that about 1% of companies use voice AI today. As that number grows, so will the number of models, and the hard part stays the same: knowing which ones to trust, and where.

Architecting AI that learns and interacts like humans.

Status