What is a mean opinion score (MOS)?

Rene, Kabir Goel 
What is a mean opinion score (MOS)?

A mean opinion score (MOS) is the average of the ratings a group of listeners gives a piece of audio, usually on a scale from 1 (bad) to 5 (excellent). It started in telephony as a way to put a number on call quality, and it is now the default quality number in text-to-speech papers and vendor comparisons.

If you are choosing a voice model or tuning a voice agent, MOS is worth understanding mostly so you can tell when it’s being misused. Here’s what goes into the number, what a good score looks like for calls and for synthetic speech, and the questions to ask before you trust one.

The MOS scale

Most MOS tests use the Absolute Category Rating (ACR) scale defined in ITU-T Recommendation P.800. Listeners hear one sample at a time and pick a label:

ScoreLabel
5Excellent
4Good
3Fair
2Poor
1Bad

The MOS is the arithmetic mean of those ratings. If ten people rate a clip 5, 4, 4, 5, 3, 4, 4, 5, 4, 3, the MOS is 4.1.

P.800 dates from 1996 and was written for telephone networks. It specifies the test conditions down to the room, which must be quiet, with reverberation under 500 ms and background noise below 30 dBA, as Wikipedia’s summary of the standard notes. A companion standard, ITU-T P.800.1, defines labels for where a score came from. That matters, because “MOS” on a product page can mean any of several different measurements.

One acronym, three numbers

When someone quotes a MOS, find out which kind it is.

What it isHow it’s producedWhere you’ll see it
Subjective MOSPeople listen and rateTTS papers, model launches, codec listening tests
Predicted MOS (speech)A model compares or listens to audio: PESQ, POLQA, DNSMOS, NISQA, UTMOSCodec and network testing, TTS training dashboards
Estimated MOS (calls)Computed from network stats such as jitter, latency, and packet lossVoIP and contact-center dashboards

Subjective MOS is the original. It’s slow and costs money, because you have to recruit and pay listeners.

Predicted MOS replaces the listeners with a model trained on their past ratings. PESQ (ITU-T P.862) compares a degraded call recording against a clean reference. ITU withdrew it in January 2024 in favor of POLQA (ITU-T P.863). For synthetic speech there is no clean reference to compare against, so researchers use reference-free predictors such as DNSMOS, NISQA, and UTMOS.

Estimated MOS on a call-quality dashboard usually isn’t anyone’s opinion. Network-planning models such as the ITU-T G.107 E-model map delay, jitter, and packet loss to a score on the MOS scale. Dialpad’s glossary describes its MOS this way, as a number that “accounts for jitter, latency, and packet loss.”

All three are reported on the same 1–5 scale. They don’t measure the same thing.

What is a good MOS score?

It depends on which number you’re looking at.

For phone and VoIP calls, there are working conventions. Twilio’s glossary calls roughly 4.3 to 4.5 an excellent quality target and says calls become unacceptable below about 3.5. Ratings rarely reach 5, because people avoid giving perfect scores. If your contact-center dashboard sits above 4, the network probably isn’t your problem.

For text-to-speech, there’s no universal threshold, and anyone who gives you one is skipping a step. A TTS MOS describes a system relative to whatever else was in the same test. Put a strong model next to weak baselines and it scores well. Put the same model next to recorded human speech and its score drops, even though the audio didn’t change.

So read a TTS MOS as a ranking within one test. The useful question isn’t “is 4.3 good?” It’s “4.3 compared with what, rated by whom?”

Why two MOS numbers usually don’t compare

This is the part most MOS write-ups skip, and it’s where most bad decisions come from.

Listeners use the whole scale they’re given. Raters spread their scores across 1 to 5 within a test, a pattern called range-equalization bias. A system rated among clearly worse samples gets pushed up, and the same system rated among better ones gets pushed down. ITU-T P.800.2, the standard for reporting MOS, says scores from separate experiments shouldn’t be compared directly unless the experiments were designed to be compared.

The setup changes the number. Instructions, the sentences used, playback loudness, headphones versus a laptop speaker, and how tired the raters are all move scores. Our post Is this TTS model good? shows four listeners giving the same waveform four different scores depending on their setup.

People want different things. A warm, slow voice that scores well for a meditation app can sound wrong on a pharmacy refill line. MOS gives you an average, and an average can hide two groups of listeners who disagree.

Small samples are noisy. For the ten ratings above, the 95% confidence interval is about ±0.5. You can’t tell a 4.1 from a 4.5 with that sample. With 100 ratings and the same spread, the interval narrows to about ±0.15. If a published MOS has no rating count or interval next to it, assume it’s less precise than it looks.

The practical rule: only compare MOS numbers that came out of the same test. A model launch post that quotes 4.4 against a competitor’s 4.2 from the competitor’s own paper tells you very little.

Predicted MOS: useful, with blind spots

Running human listening tests on every training checkpoint isn’t realistic, so teams use predictors. They’re good at what they were trained on: noise, distortion, and codec artifacts.

They’re much worse at things like the wrong emphasis, the wrong emotion, or an accent that drifts. A 2026 Interspeech paper, Investigating Human-Model Discrepancies in Speech Quality Assessment, added increasing amounts of wrong Japanese pitch accent to synthetic speech. Human ratings fell by 1.84 MOS points. UTMOS, UTMOSv2, NISQA, and DNSMOS all moved by less than 0.1, and some scored the most corrupted speech slightly higher. The authors found all of the models “insensitive to prosodic errors despite large subjective score drops.”

That doesn’t make predictors useless. If UTMOS drops 0.3 on a new checkpoint, something probably broke. If it holds steady, all you’ve learned is that nothing broke in a way UTMOS can hear. Use predictors to catch regressions, and use people to make decisions.

How to run a MOS test that tells you something

If you’re comparing voice models for a real product, a small, well-designed test beats a big, generic one.

  1. Write down the use case first. “Reads order numbers to a caller on a phone line” is a test you can design. “Sounds natural” isn’t. Our guide to choosing voice AI models has a framework for this.
  2. Use your own text. Include the hard parts: names, addresses, dates, prices, codes, and the jargon in your domain. Clean read speech hides most failures.
  3. Play audio the way users will hear it. For a phone agent, that means 8 kHz μ-law, not a studio WAV. Match loudness across systems so the louder clip doesn’t win by default.
  4. Rate every system in the same session. Mix the order, hide which system is which, and include a few repeated or reference clips as anchors so you can spot careless raters.
  5. Use listeners who match your users. For a Spanish-language support line in Mexico, that means native Mexican Spanish speakers. ITU’s P.808 covers running these tests through crowdsourcing.
  6. Collect enough ratings and report the interval. Aim for dozens of ratings per system, not a handful, and publish the confidence interval with the mean.
  7. Ask a second question. A side-by-side preference test (“which of these two would you rather hear?”) often separates close systems better than two absolute scores do.

MOS for voice agents: necessary, not sufficient

MOS rates a clip in isolation. A voice agent is a conversation, and many of the failures that end calls never show up in a clip. The agent waits too long to answer, talks over the caller, reads a confirmation code too fast, or says “2026” in a way nobody expected.

Our research team’s view, laid out in Is this TTS model good?, is that listening scores are one input among several. The others are word-level accuracy, how names and numbers are pronounced, whether the voice stays consistent across a long call, and how it holds up on the real streaming path. The word error rate guide covers the accuracy side.

If you’re comparing models, the fastest honest test is your own script, read by each model and played over the channel your users will hear it on. You can try Sonic in the playground with your own text, or run the same script through the API with pcm_mulaw at 8000 Hz to hear it as a caller would.

FAQs

What is a mean opinion score (MOS)?

A mean opinion score is the average of quality ratings that listeners give on a fixed scale, most often the five-point Absolute Category Rating scale from ITU-T P.800: 5 Excellent, 4 Good, 3 Fair, 2 Poor, 1 Bad. Telecom teams use it for call quality, and speech researchers use it for text-to-speech and voice conversion. Many tools also predict MOS with a model instead of asking people, and that predicted number should be labeled as an estimate.

What is a good MOS score?

For phone and VoIP calls, Twilio describes about 4.3 to 4.5 as an excellent target and below roughly 3.5 as unacceptable. For text-to-speech, there is no fixed good number. A TTS MOS only means something next to the other systems rated in the same test, by the same listeners, under the same instructions. A 4.2 from one paper and a 4.4 from another can't be ranked.

How is MOS calculated?

Add up every listener's rating for a clip or system and divide by the number of ratings. Report the number of ratings and a 95% confidence interval with it. With 10 ratings and a typical spread, the interval is about plus or minus half a point. With 100 ratings it shrinks to around plus or minus 0.15.

Can you compare MOS scores from different tests?

Not directly. ITU-T P.800.2 says MOS values from separate experiments shouldn't be compared unless the experiments were designed for it. Listeners use the whole scale within each test, so a system's score depends on what else it was rated alongside. Compare systems inside one test, or include shared anchor samples across tests.

What is the difference between MOS and predicted MOS?

MOS comes from people listening. Predicted MOS comes from a model trained to imitate those ratings: PESQ and POLQA for telephone networks, and DNSMOS, NISQA, and UTMOS for speech enhancement and synthesis. Predictors are cheap and useful for catching regressions across thousands of clips, but they miss failures they weren't trained on, such as wrong emphasis or the wrong pitch accent.