Mira

MOS evaluation

MOS evaluation asks trained listeners to score speech under conditions that are actually controlled. The panels are calibrated against anchor samples before they touch your audio, so a four from one listener means the same thing as a four from another. We scope every panel to the languages and domains you ship in — a listener who has never heard the accent should not be grading it. Without that scoping a MOS number is a poll, not a measurement.

The workflow begins with the panel, not the audio. Listeners are selected for the languages and domains you ship in, screened for hearing, and trained on anchor samples that span the full quality range. Before a panel touches your clips, every listener is calibrated against those anchors, and calibration is checked across the run so a listener who drifts mid-batch is caught. Each clip is scored for MOS, comparison MOS against a reference, and intelligibility, with a short note when a clip is marked low — a four that means clean-but-robotic is not the same four as one that means noisy-but-natural.

Where listeners disagree, the clip is not averaged into a quiet mean. It goes to a senior listener who re-scores it blind and settles the score with a written reason. The output is more than a mean opinion score: we report MOS, comparison MOS and intelligibility per utterance, with confidence intervals, per-listener spread, and the audio that produced each judgement. The point of MOS is not the mean; it is the measurement you can defend clip by clip.