You are comparing speech-to-text models because something downstream broke: an agent booked the wrong day, a transcript mangled a phone number, or a batch job cost more than the product it powers. This page sorts the models that matter in 2026 by the job you are hiring them for, and says where each one falls short.
The short version: if you are transcribing live English audio for a voice agent, Ink 2 is the strongest starting point. It ranked first on Artificial Analysis’s streaming speech-to-text leaderboard for word error rate as of June 2026, and it handles turn detection without a separate model. For multilingual streaming today, Deepgram Nova-3 and AssemblyAI’s Universal models cover more languages. For transcribing recorded files on hardware you control, OpenAI’s Whisper large-v3 is the open-source default. Rankings move; check the current leaderboards rather than trusting any single page, including this one.
Overview: streaming and batch are different jobs
The first split is not between vendors, it is between two kinds of work:
- Streaming (real-time) transcription returns text while the speaker is still talking. Voice agents, live captioning, and call routing need it. The model commits to words it has only partially heard, so it is judged on accuracy and on latency and on whether it can tell when the speaker finished a thought.
- Batch transcription takes a complete recording and returns a transcript. Podcast post-production, meeting notes, and compliance review live here. Latency matters much less; accuracy and cost per hour matter more.
A model built for one job rarely shines at the other. Whisper is the clearest case: excellent batch accuracy for an open model, no native streaming at all. Buying on a single blended leaderboard misses this, which is why the comparisons below are grouped by job.
How to judge a speech-to-text model
Three axes decide most real-time deployments:
- Accuracy, especially on the strings your product cares about. Average word error rate (WER) hides the errors that hurt. A model can post a respectable overall WER and still garble confirmation codes, mangle email addresses, or silently drop a spoken number — the exact failures that break an agent’s booking flow. We covered the deeper problem in our guide to word error rate: WER is a diagnostic, not a verdict, because it depends on the reference transcript and says nothing about which errors a listener forgives.
- Turn detection. In a conversation, someone has to decide when the user stopped talking. Many stacks bolt on a voice activity detector and a silence timeout; that combination misfires on pauses, and every extra model in the pipeline adds latency. Models with end-of-turn prediction built in remove the bolt-on entirely.
- Latency under real conditions. Vendor pages quote median times on clean audio. Your audio has accents, crosstalk, and background noise, and your p95 is what callers experience. Measure time-to-final-transcript on recordings that resemble your traffic.
Price is the fourth axis, and it is the one where the honest answer is boring: per-minute rates differ by a few tenths of a cent, change often, and matter less than picking a model that does not mishear your customers. Check the pricing pages current when you read this.
The models worth comparing
These are the models a builder shortlisting speech-to-text in 2026 will run into. Claims are sourced to each vendor’s own materials or independent leaderboards, labeled as such.
Cartesia Ink 2 — streaming English for voice agents
Ink 2 is a streaming model built for real-time agents, and it is the model we build. What it is good at, per the launch post and Artificial Analysis’s streaming leaderboard, where it ranked first for word error rate as of June 2026:
- Accuracy on production audio: line recordings, accented speech, noisy conditions, and earnings calls, outperforming Deepgram Flux, Soniox RT-V4, AssemblyAI’s realtime models, and ElevenLabs Scribe-2-realtime in those tests. It handles structured entities — phone numbers, emails, UUIDs, dates — by waiting for the full sequence before committing it.
- Turn detection built in: end-of-turn prediction is part of the model, so you do not run Silero or a separate turn-taking service. In a Pipecat pipeline, Ink’s eager turn-end predictions cut roughly half a second off response time by letting the agent start its reply before the user’s turn formally ends.
- Noise robustness without a preprocessing filter, which removes a Krisp-style stage (and its cost and latency) from the stack.
The honest limits: Ink 2 transcribes English only today, with a multilingual model in preview. If your agents take calls in Spanish or Hindi today, pick from the multilingual options below and revisit. Cartesia offers cloud, on-prem, and air-gapped deployment, and you can try Ink free in the playground.
OpenAI Whisper large-v3 — the open-source batch default
Whisper is the model to beat for batch transcription on your own GPUs: open weights, about 99 languages, decent accuracy on clean audio, no vendor lock-in. Teams self-host it for privacy-sensitive or cost-sensitive workloads.
Where it falls short is the job in this page’s title. Whisper is batch-only: no native streaming, no turn detection, and a known habit of hallucinating phrases during silence or non-speech audio — the failure mode AssemblyAI’s writeup calls the catch, and every streaming wrapper around Whisper inherits it. Chunked streaming setups land around 500 milliseconds of added latency and own the endpointing problem themselves. Use Whisper for files. Do not use it as the ear of a real-time agent.
Deepgram Nova-3 — multilingual streaming at scale
Deepgram’s Nova-3 is the established streaming API for multilingual deployments: sub-300 millisecond streaming latency (vendor-reported), real-time code-switching across ten languages, and domain customization. Nova-3’s launch post reports 6.84% median WER on streaming audio, a vendor figure — independent indexes put it higher, which is a good reminder to read vendor numbers as marketing until you reproduce them. Deepgram also ships Flux, a model that adds end-of-turn detection for voice agents, an acknowledgment that classic STT alone does not solve turn-taking.
Pick Nova-3 when you need many languages in production today and can accept a separate turn-detection layer or Flux alongside it.
AssemblyAI Universal — streaming with bundled speech intelligence
AssemblyAI’s Universal line pairs streaming transcription with summarization, entity detection, sentiment, and PII redaction in one API, and its realtime models compete directly in voice-agent benchmarks — the company’s own writeups report 5.19% WER for Universal-3.6 Pro Realtime on its English voice-agent benchmark, ahead of Scribe v2 and Nova-3 on the same test. Treat that as a vendor-run benchmark, but the positioning is clear: one vendor for transcription plus audio intelligence. Strong pick when the transcript feeds analytics as well as an agent.
Google Chirp 3 — for teams already on Google Cloud
Google’s Chirp 3 rounds out the managed options, with wide language coverage and the operational simplicity of staying inside a cloud provider many teams already run. It rarely tops independent accuracy leaderboards for streaming English, but procurement reality — existing contracts, compliance postures, one invoice — keeps it in real deployments.
Ink vs Whisper: when the open-source default loses
The most common comparison we see is Cartesia Ink against OpenAI Whisper, so it deserves its own answer. The question sounds like “which model is better,” and the real question is “which job do you have.”
Whisper wins when you have files, many languages, and your own GPUs: batch transcripts of recorded calls, podcasts, or archives, transcribed cheaply at scale with no per-minute vendor bill.
Ink wins when you have a live conversation: it streams as the user speaks, commits to words faster, predicts when the turn ends without a separate endpointing model, and holds accuracy on numbers, IDs, and accented speech — the errors that turn into wrong bookings. Running Whisper behind a streaming wrapper means you own chunking, endpointing, and silence hallucinations, and you pay for it in round-trip latency.
A useful tiebreaker: if your pipeline already includes Silero, Krisp, and a silence-timeout state machine held together with retries, that machinery exists to compensate for a batch model in a streaming job. A streaming-native model replaces it.
Which speech-to-text model should you pick?
Match the model to the job:
| Your job | Start with |
|---|---|
| Live English voice agent | Ink 2 |
| Live multilingual transcription | Deepgram Nova-3, AssemblyAI Universal realtime |
| Batch files, self-hosted | Whisper large-v3 |
| Transcription plus analytics in one API | AssemblyAI Universal |
| Staying inside an existing cloud contract | Google Chirp 3 |
Then test on your own audio before committing. Every figure on this page came from someone else’s recordings; the ranking that matters is the one your traffic produces. Play a few of your hardest calls — accents, crosstalk, a caller spelling out a confirmation code — through the two or three finalists and compare what comes back, including the errors your product cannot afford.
If English voice agents are the job, try Ink 2 in the playground with no setup, or read how to build a voice agent to see where transcription fits in the larger pipeline.
Related documentation
- Ink 2 launch post — the accuracy, turn-detection, and latency results behind the ranking claim
- Word error rate, explained — why WER is a diagnostic, not a verdict
- Voice activity detection for voice agents — why silence-based endpointing misfires
- Build a voice agent with Pipecat and Cartesia — a working streaming stack with Ink’s eager turn ends
- Artificial Analysis streaming speech-to-text leaderboard — current standings across providers