Two tools exist to score AI-generated speech automatically, without a human ever listening: UTMOS and NISQA-TTS, both “no-reference” quality predictors that output a 1-5 score. We ran them against real human listening across two separate tests, and found something worth knowing before you trust either one to pick a text-to-speech provider: they don’t always agree with what a person actually prefers — and sometimes they don’t even agree with each other.
What these scores actually measure
Both tools are trained to predict how natural a clip would sound to a listener — specifically, how close it is to an unremarkable, neutral human recording. That’s a narrower question than “does this sound emotionally alive.” A flat, neutral reading can score well on naturalness while sounding lifeless; a more expressive, dynamic performance can score worse on the same scale, precisely because it’s further from “neutral,” even when it’s the better result for what you actually wanted.
The finding: two providers, two listening tests, one consistent surprise
We compared three TTS providers — ElevenLabs Turbo v2.5, ElevenLabs Eleven v3, and Gemini 3.1 Flash TTS — on two different tests: five single lines spanning a real emotional range, and a ~2-3 minute two-person dialogue scene. Both times, a human listening pass was run alongside the automatic scores.
Both times, the human listener picked Gemini 3.1 Flash TTS as the most emotionally convincing of the three — despite it scoring lowest on both automatic tools, at both lengths. The falling automatic scores over the course of the dialogue scene tracked the character’s delivery going quieter and more broken, exactly as the mood called for — a naturalness predictor reads that as “worse,” a person hears it as “sadder.” Two independent listening passes, at very different clip lengths, landing on the same reversal is a real, repeatable pattern — not a one-off fluke.
A fresh pair, scored honestly — listen for yourself
Rather than just cite our own past results, we generated a new pair of clips specifically for this piece — a single line, same text, one provider each — and ran both automatic scorers on them just now. We’re not going to tell you which one we think sounds better, because we can’t hear it ourselves. Listen to both and judge for yourself.
| Provider | UTMOS | NISQA-TTS |
|---|---|---|
| ElevenLabs Turbo v2.5 | 4.30 | 4.11 |
| Gemini 3.1 Flash TTS | 4.21 | 4.67 |
Notice the two tools don’t even agree with each other here: UTMOS scores ElevenLabs slightly higher, NISQA-TTS scores Gemini clearly higher, on the exact same pair of clips. That’s the same disagreement pattern we saw in the larger study — and it’s a useful reminder on its own: if two respected automatic scorers can land on opposite rankings for the same audio, neither one should be treated as a final verdict.
The practical takeaway
Use automatic naturalness scores for what they’re actually good at: catching genuinely broken output. In our testing, both tools correctly flagged a real defect — a clip that had rendered as an accidental whisper scored severely lower than everything else, and once we fixed it, both scores jumped back to normal. That kind of outlier-catching is genuinely useful.
What they’re not reliable for is picking a provider based on which one sounds more expressive or emotionally alive — that’s a different question than the one these tools are trained to answer, and our results point the opposite direction from the automatic scores on that question, twice. If emotional delivery is what you actually care about, do a real listening pass before you decide — don’t let a naturalness score make the call for you.
Both providers in this piece are reachable through the same fal.ai account, alongside the image and video models we’ve covered in our earlier parameter study on AI portrait realism.
