A mean opinion score (MOS) is the average of the ratings a group of listeners gives a piece of audio, usually on a scale from 1 (bad) to 5 (excellent). It started in telephony as a way to put a number on call quality, and it is now the default quality number in text-to-speech papers and vendor comparisons.
If you are choosing a voice model or tuning a voice agent, MOS is worth understanding mostly so you can tell when it’s being misused. Here’s what goes into the number, what a good score looks like for calls and for synthetic speech, and the questions to ask before you trust one.
The MOS scale
Most MOS tests use the Absolute Category Rating (ACR) scale defined in ITU-T Recommendation P.800. Listeners hear one sample at a time and pick a label:
| Score | Label |
|---|---|
| 5 | Excellent |
| 4 | Good |
| 3 | Fair |
| 2 | Poor |
| 1 | Bad |
The MOS is the arithmetic mean of those ratings. If ten people rate a clip 5, 4, 4, 5, 3, 4, 4, 5, 4, 3, the MOS is 4.1.
P.800 dates from 1996 and was written for telephone networks. It specifies the test conditions down to the room, which must be quiet, with reverberation under 500 ms and background noise below 30 dBA, as Wikipedia’s summary of the standard notes. A companion standard, ITU-T P.800.1, defines labels for where a score came from. That matters, because “MOS” on a product page can mean any of several different measurements.
One acronym, three numbers
When someone quotes a MOS, find out which kind it is.
| What it is | How it’s produced | Where you’ll see it |
|---|---|---|
| Subjective MOS | People listen and rate | TTS papers, model launches, codec listening tests |
| Predicted MOS (speech) | A model compares or listens to audio: PESQ, POLQA, DNSMOS, NISQA, UTMOS | Codec and network testing, TTS training dashboards |
| Estimated MOS (calls) | Computed from network stats such as jitter, latency, and packet loss | VoIP and contact-center dashboards |
Subjective MOS is the original. It’s slow and costs money, because you have to recruit and pay listeners.
Predicted MOS replaces the listeners with a model trained on their past ratings. PESQ (ITU-T P.862) compares a degraded call recording against a clean reference. ITU withdrew it in January 2024 in favor of POLQA (ITU-T P.863). For synthetic speech there is no clean reference to compare against, so researchers use reference-free predictors such as DNSMOS, NISQA, and UTMOS.
Estimated MOS on a call-quality dashboard usually isn’t anyone’s opinion. Network-planning models such as the ITU-T G.107 E-model map delay, jitter, and packet loss to a score on the MOS scale. Dialpad’s glossary describes its MOS this way, as a number that “accounts for jitter, latency, and packet loss.”
All three are reported on the same 1–5 scale. They don’t measure the same thing.
What is a good MOS score?
It depends on which number you’re looking at.
For phone and VoIP calls, there are working conventions. Twilio’s glossary calls roughly 4.3 to 4.5 an excellent quality target and says calls become unacceptable below about 3.5. Ratings rarely reach 5, because people avoid giving perfect scores. If your contact-center dashboard sits above 4, the network probably isn’t your problem.
For text-to-speech, there’s no universal threshold, and anyone who gives you one is skipping a step. A TTS MOS describes a system relative to whatever else was in the same test. Put a strong model next to weak baselines and it scores well. Put the same model next to recorded human speech and its score drops, even though the audio didn’t change.
So read a TTS MOS as a ranking within one test. The useful question isn’t “is 4.3 good?” It’s “4.3 compared with what, rated by whom?”
Why two MOS numbers usually don’t compare
This is the part most MOS write-ups skip, and it’s where most bad decisions come from.
Listeners use the whole scale they’re given. Raters spread their scores across 1 to 5 within a test, a pattern called range-equalization bias. A system rated among clearly worse samples gets pushed up, and the same system rated among better ones gets pushed down. ITU-T P.800.2, the standard for reporting MOS, says scores from separate experiments shouldn’t be compared directly unless the experiments were designed to be compared.
The setup changes the number. Instructions, the sentences used, playback loudness, headphones versus a laptop speaker, and how tired the raters are all move scores. Our post Is this TTS model good? shows four listeners giving the same waveform four different scores depending on their setup.
People want different things. A warm, slow voice that scores well for a meditation app can sound wrong on a pharmacy refill line. MOS gives you an average, and an average can hide two groups of listeners who disagree.
Small samples are noisy. For the ten ratings above, the 95% confidence interval is about ±0.5. You can’t tell a 4.1 from a 4.5 with that sample. With 100 ratings and the same spread, the interval narrows to about ±0.15. If a published MOS has no rating count or interval next to it, assume it’s less precise than it looks.
The practical rule: only compare MOS numbers that came out of the same test. A model launch post that quotes 4.4 against a competitor’s 4.2 from the competitor’s own paper tells you very little.
Predicted MOS: useful, with blind spots
Running human listening tests on every training checkpoint isn’t realistic, so teams use predictors. They’re good at what they were trained on: noise, distortion, and codec artifacts.
They’re much worse at things like the wrong emphasis, the wrong emotion, or an accent that drifts. A 2026 Interspeech paper, Investigating Human-Model Discrepancies in Speech Quality Assessment, added increasing amounts of wrong Japanese pitch accent to synthetic speech. Human ratings fell by 1.84 MOS points. UTMOS, UTMOSv2, NISQA, and DNSMOS all moved by less than 0.1, and some scored the most corrupted speech slightly higher. The authors found all of the models “insensitive to prosodic errors despite large subjective score drops.”
That doesn’t make predictors useless. If UTMOS drops 0.3 on a new checkpoint, something probably broke. If it holds steady, all you’ve learned is that nothing broke in a way UTMOS can hear. Use predictors to catch regressions, and use people to make decisions.
How to run a MOS test that tells you something
If you’re comparing voice models for a real product, a small, well-designed test beats a big, generic one.
- Write down the use case first. “Reads order numbers to a caller on a phone line” is a test you can design. “Sounds natural” isn’t. Our guide to choosing voice AI models has a framework for this.
- Use your own text. Include the hard parts: names, addresses, dates, prices, codes, and the jargon in your domain. Clean read speech hides most failures.
- Play audio the way users will hear it. For a phone agent, that means 8 kHz μ-law, not a studio WAV. Match loudness across systems so the louder clip doesn’t win by default.
- Rate every system in the same session. Mix the order, hide which system is which, and include a few repeated or reference clips as anchors so you can spot careless raters.
- Use listeners who match your users. For a Spanish-language support line in Mexico, that means native Mexican Spanish speakers. ITU’s P.808 covers running these tests through crowdsourcing.
- Collect enough ratings and report the interval. Aim for dozens of ratings per system, not a handful, and publish the confidence interval with the mean.
- Ask a second question. A side-by-side preference test (“which of these two would you rather hear?”) often separates close systems better than two absolute scores do.
MOS for voice agents: necessary, not sufficient
MOS rates a clip in isolation. A voice agent is a conversation, and many of the failures that end calls never show up in a clip. The agent waits too long to answer, talks over the caller, reads a confirmation code too fast, or says “2026” in a way nobody expected.
Our research team’s view, laid out in Is this TTS model good?, is that listening scores are one input among several. The others are word-level accuracy, how names and numbers are pronounced, whether the voice stays consistent across a long call, and how it holds up on the real streaming path. The word error rate guide covers the accuracy side.
If you’re comparing models, the fastest honest test is your own script, read by each model and played over the channel your users will hear it on. You can try Sonic in the playground with your own text, or run the same script through the API with pcm_mulaw at 8000 Hz to hear it as a caller would.