What Is MOS, and How Do You Measure Call Quality?
MOS (Mean Opinion Score) is a call quality score on a scale of 1 to 5: 5 is excellent, 1 is poor. In the past the score was set by people who listened to calls; today, systems calculate it from network data: latency, jitter and packet loss. It gives a common language for the question "how good was the call."
Definition and origin
MOS stands for Mean Opinion Score. The method comes from the world of classic telephony: you seat a group of people, play them recordings of calls, and each one gives a score. The average is the MOS. The method is described in a recommendation of the International Telecommunication Union, ITU-T P.800.
Obviously, you can't seat a group of listeners for every call in the office. So methods were developed that calculate an estimated score, which should be close to what people would have given. When a phone system displays "MOS 4.1," it is almost always a calculated score of this kind.
A simple analogy: a restaurant's rating on a review site. No rating tells the whole story, but a restaurant with 4.5 and a restaurant with 2.8 are probably not the same.
The scale
| Score | Original description | What it means in practice |
|---|---|---|
| 5 | Excellent | Like face to face. On the phone you almost never reach this |
| 4 | Good | A clear call, no effort |
| 3 | Fair | Understandable, but you notice interruptions or distortion |
| 2 | Poor | You have to repeat things; frustrating |
| 1 | Bad | Almost impossible to talk |
In telephony, a score of about 4 and above is generally considered good, a score between about 3.6 and 4 is fair, and a score below about 3.5 is a sign of a problem worth checking. These are rules of thumb, not a law.
Note: a score of 5 almost never appears in phone calls. Even in perfect conditions, the codec itself limits the score. A regular G.711 call on a clean network reaches about 4.4 by the standard calculation.
Why an average? Because two people who hear the same call don't always agree. One is sensitive to delay, the other to noise. The average of many listeners gives a stable picture that doesn't depend on one person's taste.
How a score is calculated without listeners
There are two families of methods:
- Calculation from network data: the system doesn't listen to the voice, but looks at the numbers: how much latency, how much jitter, how many packets were lost, and which codec was used. From these it calculates a score by a formula. The common method is called the E-model (ITU-T recommendation G.107), and it first calculates a score called the R-Factor and then converts it to MOS.
- Voice comparison: a known recording is transmitted and what arrived is compared to the original. Methods like PESQ and POLQA work this way. They are more accurate, but they require test equipment and are used mainly in special tests.
Most phones and PBXs use the first method, because it can be applied to every call, in real time, without listening to the content.
There is also an "estimated score in advance": before setting up an office, you can enter the expected latency and loss into the formula and get an estimate of call quality. This helps you decide, for example, whether a certain connection is sufficient.
What lowers the score
| Factor | How it sounds | Effect on the score |
|---|---|---|
| Packet loss | Gaps, swallowed words | Very high: even a few percent is noticeable |
| Latency | Delay, people talk over each other | Moderate up to about 150 milliseconds, then grows quickly |
| Jitter | Choppy, robotic voice | Depends on the phone's jitter buffer |
| Compressed codec | A "flatter" voice | Lowers the ceiling from the start |
| Echo | You hear yourself late | Very noticeable when latency is high |
These factors add up. Moderate latency alone isn't a problem, and small loss alone isn't a problem, but both together can bring a call down from "good" to "fair."
So when you fix things, it's worth starting with the factor that has the most impact, usually packet loss, and not the factor that is easiest to change.
Where to see the score
- On the phone itself: many phones show, during or after a call, a "Status" or "Quality" screen with the codec, jitter, loss and sometimes MOS too.
- In the PBX: many systems keep a quality report for each call, built from the voice channel data (RTCP, see RTP).
- In monitoring tools: large organizations run tools that track MOS averages over time and alert when they drop.
An important note: every call has two directions, and each direction has its own score. A call can be 4.3 from the office to the customer and 3.1 from the customer to the office. That is exactly the information that helps you understand where the problem is.
When you ask support to check a call, it's worth saying from which side it sounded bad: whether you heard it badly, or the other side complained. This points right away in the right direction.
How to use MOS for diagnosis
MOS alone doesn't tell you what the problem is. It tells you there is a problem, and it helps you ask the right questions:
- Which direction is the score low in? If it's only the direction coming into the office, look at the download side of the connection, or at the router. If it's only outgoing, look at the upload.
- All extensions or just one? One: a cable, Wi-Fi or the device. All of them: the internet connection.
- All hours or specific hours? A score that drops every day at a fixed time hints at load: computer backups, updates, a video class.
- Which component is highest? Loss, jitter or delay: each one points to a different fix.
This is how a single score becomes a checklist. And once you make a fix, you can test again and see whether the score really went up.
A high score, but still a complaint
Sometimes the numbers are good, and the user still says "it sounds bad." There are a few explanations:
- The problem isn't in the network: a faulty headset, a microphone far from the mouth, low volume, noise in the room. An MOS calculated from network data doesn't see any of these.
- The problem is on the other end: the customer is speaking from a cell phone with weak reception. The score measures only the part of the call that passes through your system.
- A short glitch: the average of a whole call can hide ten bad seconds in the middle.
- Echo: sometimes caused by the device or the speaker, and not always included in the calculation.
That's why MOS is a good aid, but not a substitute for a simple question: "What exactly did you hear, and when?"
R-Factor: the score behind the score
In the E-model calculation, the system first calculates a score called R-Factor, on a scale of 0 to 100, and only then converts it to MOS. In professional tools you sometimes run into both numbers.
| Approx. R-Factor | Approx. MOS | Estimated satisfaction |
|---|---|---|
| 90 and above | 4.3 and above | Most users are very satisfied |
| 80–90 | 4.0–4.3 | Satisfied |
| 70–80 | 3.6–4.0 | Some users are dissatisfied |
| 60–70 | 3.1–3.6 | Many are dissatisfied |
| Below 60 | Below 3.1 | Most users are dissatisfied |
The table is based on what is common in the standard, and the numbers are rounded. The goal isn't to be accurate to the tenth, but to understand which range the call falls in.
How to improve the score
- A cable instead of Wi-Fi: for desk phones, and especially at stations that do a lot of talking.
- Voice prioritization on the network: set up QoS on the router and switch, so voice goes ahead of downloads and backups.
- Enough bandwidth: especially for upload, which on many connections is narrower. See Bandwidth.
- A suitable router: SIP ALG off, and reasonable connection hold times.
- A suitable codec: when there is enough bandwidth, a less compressed codec gives a higher ceiling.
- Separating traffic: in large organizations, a separate network for phones (VLAN).
An example from the field
At the Greenfeld family's bookkeeping office (a fictitious name), they complained that calls "sound robotic" at around four in the afternoon. Checking the call data showed that the score drops at exactly those hours, in the outgoing direction only, and that the prominent component was packet loss.
The cause: a daily backup of the computers to the cloud, which started at four and filled the upload line. After the backup was moved to nighttime hours and voice prioritization was set up on the router, the score returned to the good range. No one would have found this without seeing that the problem was in one direction and at a fixed hour.
MOS over time: look for trends
The score of a single call is interesting. The scores of a hundred calls over a week are much more. When you look at averages, patterns appear that you can't see in a single call:
- A drop at fixed hours: daily load on the network.
- A drop at one branch: that branch's connection or router.
- A gradual decline over weeks: equipment wearing out, or more and more devices on the same connection.
- A sudden drop from a certain day: a change that was made: a new router, an update, a different provider.
The question "What changed on the day the complaints began?" solves quite a few problems, even before you check anything technical.
What MOS doesn't measure
It's worth remembering the limits of this measure, so you don't rely on it too much:
- It doesn't measure whether the call was answered on time, or how long people waited in the queue.
- It doesn't measure whether the agent was polite or answered correctly.
- It doesn't measure problems before the call began, such as an extension that doesn't ring.
- In most cases it doesn't see what happens outside the network being measured, for example on the customer's cell phone.
In other words: MOS answers one question, "how well did the voice come through." Other questions have other measures, such as wait times and abandoned calls.
How it works with us at Kesher
In Kesher's phone system, the call log shows every call, including unanswered ones. When a call that sounded bad is reported, the time and the extension make it possible to find it quickly and check it.
When there is a quality problem, every question is answered by a person, not an automated answering system or a queue waiting for an agent. In most cases the check starts at the office network: the router, the connection and Wi-Fi.
Our standard router recommendation, SIP ALG off and a UDP connection hold time of 180 seconds, prevents some of the problems that lower call quality.
And when setting up a communications cabinet, with cables, a patch panel, a PoE switch, a router and a UPS, all neatly labeled, the phones get a stable wired infrastructure, which is the basis for a good call.
FAQ
What is MOS, in simple terms?
A call quality score from 1 to 5. In telephony, about 4 and above is considered good.
Why does it never reach 5?
Because even under perfect conditions, the codec limits the quality. A regular G.711 call on a clean network reaches about 4.4.
Is the score calculated from what was said in the call?
No. In most systems the score is calculated from network data only (delay, jitter and loss), without listening to the content.
Is a score of 3.5 bad?
It's a sign that it's worth checking. The call is still understandable, but many people will already notice breaks or distortion.
Why are there two scores for one call?
Because each direction has its own score. The difference between them helps you understand whether the problem is in the download or the upload.
The score is good and the customer says they can't hear. How come?
The problem can be in the headset, the microphone, noise in the room, or on the customer's own end. An MOS calculated from the network doesn't see these.
What is R-Factor?
A score on a scale of 0 to 100 from which MOS is calculated in the common method. 80 and above is considered good.
What affects the score the most?
Usually packet loss. Even a few percent is noticeable, especially when it comes in a row.
How do you improve a low score?
Start with the network: a cable instead of Wi-Fi, voice prioritization on the router, and a check that nothing is filling the upload line.
Does a higher score mean HD voice?
Not necessarily, but a wideband codec allows a higher ceiling. An HD call on a good network usually gets a higher score than the same call in narrowband.