Learn

Best speech to text models compared (2026)

Rene 
Best speech to text models compared (2026)

You are comparing speech-to-text models because something downstream broke: an agent booked the wrong day, a transcript mangled a phone number, or a batch job cost more than the product it powers. This page sorts the models that matter in 2026 by the job you are hiring them for, and says where each one falls short.

The short version: if you are transcribing live English audio for a voice agent, Ink 2 is the strongest starting point. It ranked first on Artificial Analysis’s streaming speech-to-text leaderboard for word error rate as of June 2026, and it handles turn detection without a separate model. For multilingual streaming today, Deepgram Nova-3 and AssemblyAI’s Universal models cover more languages. For transcribing recorded files on hardware you control, OpenAI’s Whisper large-v3 is the open-source default. Rankings move; check the current leaderboards rather than trusting any single page, including this one.

Overview: streaming and batch are different jobs

The first split is not between vendors, it is between two kinds of work:

  • Streaming (real-time) transcription returns text while the speaker is still talking. Voice agents, live captioning, and call routing need it. The model commits to words it has only partially heard, so it is judged on accuracy and on latency and on whether it can tell when the speaker finished a thought.
  • Batch transcription takes a complete recording and returns a transcript. Podcast post-production, meeting notes, and compliance review live here. Latency matters much less; accuracy and cost per hour matter more.

A model built for one job rarely shines at the other. Whisper is the clearest case: excellent batch accuracy for an open model, no native streaming at all. Buying on a single blended leaderboard misses this, which is why the comparisons below are grouped by job.

How to judge a speech-to-text model

Three axes decide most real-time deployments:

  1. Accuracy, especially on the strings your product cares about. Average word error rate (WER) hides the errors that hurt. A model can post a respectable overall WER and still garble confirmation codes, mangle email addresses, or silently drop a spoken number — the exact failures that break an agent’s booking flow. We covered the deeper problem in our guide to word error rate: WER is a diagnostic, not a verdict, because it depends on the reference transcript and says nothing about which errors a listener forgives.
  2. Turn detection. In a conversation, someone has to decide when the user stopped talking. Many stacks bolt on a voice activity detector and a silence timeout; that combination misfires on pauses, and every extra model in the pipeline adds latency. Models with end-of-turn prediction built in remove the bolt-on entirely.
  3. Latency under real conditions. Vendor pages quote median times on clean audio. Your audio has accents, crosstalk, and background noise, and your p95 is what callers experience. Measure time-to-final-transcript on recordings that resemble your traffic.

Price is the fourth axis, and it is the one where the honest answer is boring: per-minute rates differ by a few tenths of a cent, change often, and matter less than picking a model that does not mishear your customers. Check the pricing pages current when you read this.

The models worth comparing

These are the models a builder shortlisting speech-to-text in 2026 will run into. Claims are sourced to each vendor’s own materials or independent leaderboards, labeled as such.

Cartesia Ink 2 — streaming English for voice agents

Ink 2 is a streaming model built for real-time agents, and it is the model we build. What it is good at, per the launch post and Artificial Analysis’s streaming leaderboard, where it ranked first for word error rate as of June 2026:

  • Accuracy on production audio: line recordings, accented speech, noisy conditions, and earnings calls, outperforming Deepgram Flux, Soniox RT-V4, AssemblyAI’s realtime models, and ElevenLabs Scribe-2-realtime in those tests. It handles structured entities — phone numbers, emails, UUIDs, dates — by waiting for the full sequence before committing it.
  • Turn detection built in: end-of-turn prediction is part of the model, so you do not run Silero or a separate turn-taking service. In a Pipecat pipeline, Ink’s eager turn-end predictions cut roughly half a second off response time by letting the agent start its reply before the user’s turn formally ends.
  • Noise robustness without a preprocessing filter, which removes a Krisp-style stage (and its cost and latency) from the stack.

The honest limits: Ink 2 transcribes English only today, with a multilingual model in preview. If your agents take calls in Spanish or Hindi today, pick from the multilingual options below and revisit. Cartesia offers cloud, on-prem, and air-gapped deployment, and you can try Ink free in the playground.

OpenAI Whisper large-v3 — the open-source batch default

Whisper is the model to beat for batch transcription on your own GPUs: open weights, about 99 languages, decent accuracy on clean audio, no vendor lock-in. Teams self-host it for privacy-sensitive or cost-sensitive workloads.

Where it falls short is the job in this page’s title. Whisper is batch-only: no native streaming, no turn detection, and a known habit of hallucinating phrases during silence or non-speech audio — the failure mode AssemblyAI’s writeup calls the catch, and every streaming wrapper around Whisper inherits it. Chunked streaming setups land around 500 milliseconds of added latency and own the endpointing problem themselves. Use Whisper for files. Do not use it as the ear of a real-time agent.

Deepgram Nova-3 — multilingual streaming at scale

Deepgram’s Nova-3 is the established streaming API for multilingual deployments: sub-300 millisecond streaming latency (vendor-reported), real-time code-switching across ten languages, and domain customization. Nova-3’s launch post reports 6.84% median WER on streaming audio, a vendor figure — independent indexes put it higher, which is a good reminder to read vendor numbers as marketing until you reproduce them. Deepgram also ships Flux, a model that adds end-of-turn detection for voice agents, an acknowledgment that classic STT alone does not solve turn-taking.

Pick Nova-3 when you need many languages in production today and can accept a separate turn-detection layer or Flux alongside it.

AssemblyAI Universal — streaming with bundled speech intelligence

AssemblyAI’s Universal line pairs streaming transcription with summarization, entity detection, sentiment, and PII redaction in one API, and its realtime models compete directly in voice-agent benchmarks — the company’s own writeups report 5.19% WER for Universal-3.6 Pro Realtime on its English voice-agent benchmark, ahead of Scribe v2 and Nova-3 on the same test. Treat that as a vendor-run benchmark, but the positioning is clear: one vendor for transcription plus audio intelligence. Strong pick when the transcript feeds analytics as well as an agent.

Google Chirp 3 — for teams already on Google Cloud

Google’s Chirp 3 rounds out the managed options, with wide language coverage and the operational simplicity of staying inside a cloud provider many teams already run. It rarely tops independent accuracy leaderboards for streaming English, but procurement reality — existing contracts, compliance postures, one invoice — keeps it in real deployments.

Ink vs Whisper: when the open-source default loses

The most common comparison we see is Cartesia Ink against OpenAI Whisper, so it deserves its own answer. The question sounds like “which model is better,” and the real question is “which job do you have.”

Whisper wins when you have files, many languages, and your own GPUs: batch transcripts of recorded calls, podcasts, or archives, transcribed cheaply at scale with no per-minute vendor bill.

Ink wins when you have a live conversation: it streams as the user speaks, commits to words faster, predicts when the turn ends without a separate endpointing model, and holds accuracy on numbers, IDs, and accented speech — the errors that turn into wrong bookings. Running Whisper behind a streaming wrapper means you own chunking, endpointing, and silence hallucinations, and you pay for it in round-trip latency.

A useful tiebreaker: if your pipeline already includes Silero, Krisp, and a silence-timeout state machine held together with retries, that machinery exists to compensate for a batch model in a streaming job. A streaming-native model replaces it.

Which speech-to-text model should you pick?

Match the model to the job:

Your jobStart with
Live English voice agentInk 2
Live multilingual transcriptionDeepgram Nova-3, AssemblyAI Universal realtime
Batch files, self-hostedWhisper large-v3
Transcription plus analytics in one APIAssemblyAI Universal
Staying inside an existing cloud contractGoogle Chirp 3

Then test on your own audio before committing. Every figure on this page came from someone else’s recordings; the ranking that matters is the one your traffic produces. Play a few of your hardest calls — accents, crosstalk, a caller spelling out a confirmation code — through the two or three finalists and compare what comes back, including the errors your product cannot afford.

If English voice agents are the job, try Ink 2 in the playground with no setup, or read how to build a voice agent to see where transcription fits in the larger pipeline.

FAQs

What is the best speech-to-text model?

It depends on the job. For streaming English in a voice agent, Ink 2 topped Artificial Analysis's streaming leaderboard for word error rate as of June 2026 and has turn detection built in. For multilingual streaming, Deepgram Nova-3 and AssemblyAI's Universal models cover more languages today. For batch transcription on your own hardware, Whisper large-v3 is the open-source default. Check the current leaderboards rather than trusting any single page, including this one.

Can Whisper transcribe in real time?

Not natively. Whisper is a batch model: it expects a complete audio segment before it returns text. Teams build streaming wrappers around it by chunking audio and handling endpointing themselves, which adds latency and failure modes. If you need live transcription, a streaming-native model such as Cartesia Ink, Deepgram Nova-3, or AssemblyAI's realtime models is the simpler path.

What is word error rate, and is a lower WER always better?

Word error rate (WER) is the share of words a model transcribes incorrectly against a reference transcript. It is a useful diagnostic, not a complete verdict: it penalizes formatting choices, it depends on which ASR model scored the reference, and two deployments with the same WER can feel very different if one drops numbers and IDs. Evaluate on your own audio, and weight the errors your application cannot tolerate.

Which speech-to-text model handles numbers and IDs best?

Look for entity-level evaluation, not just average WER. Phone numbers, email addresses, alphanumerics, and dates fail differently from ordinary words, and some models handle mid-sequence entities better than others. Ink 2, for example, waits for a complete sequence before committing it. Whatever you pick, test with the strings your product actually hears: order IDs, confirmation codes, street addresses.

How much does Ink 2 cost?

Ink 2 is priced per minute of audio, with volume discounts and custom pricing on enterprise contracts. Current rates are on the Cartesia pricing page, and you can test the model free in the Cartesia playground.