English
EN
← Research

How Do You Measure the Quality of an AI-Dubbed Video?

WER, BLEU, MOS, and lip-sync scores each measure one slice of AI dubbing. Production quality is a vector—translation, speaker identity, prosody, emotion, pace, and visual sync—plus the failures viewers actually notice.

Why WER, BLEU, MOS, and lip-sync scores are not enough

AI dubbing has made an impressive transition from research demos to production systems. A video can now be translated, voiced, and lip-synced in minutes.

But this creates a less obvious problem:

How do we know whether the result is actually good?

It is tempting to evaluate a dubbing system using the metrics we already know:

  • WER for speech recognition
  • BLEU or COMET for translation
  • MOS for speech quality
  • speaker similarity for voice cloning
  • SyncNet scores for lip synchronization

All of these metrics are useful.

None of them, individually, tells us whether a viewer will think the dubbed video is convincing.

The reason is simple:

AI dubbing is not a single task. It is the reconstruction of a performance.

A successful dub has to preserve meaning, identity, timing, emotion, delivery, and audiovisual coherence—while producing natural speech in another language.

And these objectives can conflict with one another.

A translation can be extremely accurate but too long to fit the scene.

A voice can be highly similar to the original speaker but deliver the sentence with the wrong emotion.

A video can achieve excellent lip-sync while sounding completely unnatural.

A 60-minute video can have excellent average scores while containing two catastrophic mistakes that viewers immediately notice.

So perhaps the right question isn’t:

“What is the quality score of this dub?”

It is:

“What dimensions of quality matter for this particular video, and how reliably can we measure each one?”


There Is No Single Dubbing Quality Score

Consider this short scene:

Original:
“You did what?!

Suppose the translated version is:

“你做了什么?”

The translation may be perfectly correct.

The cloned voice may have 0.90 speaker similarity.

The audio may have excellent signal quality.

The lip-sync model may report a strong synchronization score.

But if the generated voice says it calmly instead of shouting in disbelief, the dub still feels wrong.

Now consider another example:

Original:
“I can’t believe you actually did it.”

The translated sentence is excellent, the emotion is correct, and the speaker sounds right—but the generated speech lasts 2.8 seconds while the original shot only gives the speaker 1.9 seconds.

The result may overlap with the next speaker.

Again, the individual components can all look good while the video-level result fails.

This is why dubbing quality is inherently multidimensional.

A useful conceptual model is:

A useful conceptual model is

And there is another dimension above all of these:

Consistency across the entire video.


What Should We Actually Measure?

I would divide the evaluation problem into roughly seven dimensions.

DimensionWhat we want to knowExamples of measurements
TranslationDid we preserve the meaning?COMET, human evaluation, terminology accuracy
VoiceDoes the speaker still sound like themselves?Speaker embedding similarity, human preference
NaturalnessDoes the speech sound human?MOS, neural speech-quality metrics
PerformanceIs it delivered correctly?Prosody, emotion, pitch, rhythm, pauses
TimingDoes it fit the original performance?Duration ratio, speech overlap, pause alignment
VisualDoes it fit the video?Lip-sync / AV synchrony metrics
ConsistencyDoes the whole video remain coherent?Speaker, terminology, voice, style, narrative consistency

This is already much closer to how humans actually judge dubbed content.

Interestingly, this isn’t just an engineering intuition.

A large-scale study of professional human dubbing analyzed 319.57 hours of video, 674 episodes, 54 shows, and 9,215 speakers. The study found that vocal naturalness and translation quality were particularly important, while also showing that source-side audio characteristics beyond the words—including emphasis and emotion—matter to human dubbing.

In other words:

Professional dubbing itself is multidimensional.

Why should AI dubbing be evaluated differently?


Translation Accuracy Is Necessary—but Not Sufficient

The most obvious metric is linguistic quality.

We can ask:

  • Is the meaning correct?
  • Are names correct?
  • Are numbers correct?
  • Are technical terms correct?
  • Is anything missing?
  • Is the translation natural in the target language?
  • Are cultural references handled appropriately?
  • Is terminology consistent throughout the video?

Traditional machine translation metrics such as BLEU can be useful for benchmarking, while newer metrics such as COMET attempt to correlate better with human translation judgments.

But even a perfect translation metric cannot tell us whether the sentence works as spoken dialogue.

Consider:

“Are you kidding me?”

There are dozens of reasonable translations depending on context.

And even the best textual translation doesn’t tell us whether the actor should sound:

  • angry,
  • amused,
  • shocked,
  • sarcastic,
  • disappointed,
  • or genuinely confused.

A 2024 discussion of evaluating end-to-end speech-to-speech translation for dubbing made essentially this point: conventional text-centric metrics do not capture important properties such as naturalness, target-voice similarity, accent, and rhythm.

So:

Translation quality measures what was said. Dubbing quality also has to measure how it was said.


Speaker Similarity: Does the Character Still Sound Like the Same Person?

Voice cloning introduces another dimension.

For each speaker, we can ask:

Does the generated voice preserve the identity of the original speaker?

Speaker embeddings give us a useful automated signal.

For example:

For example

But speaker similarity has an important limitation.

Similarity is not the same as consistency.

Imagine a 30-minute documentary with one narrator.

The generated voice might have:

Segment 1    similarity = 0.91
Segment 2    similarity = 0.89
Segment 3    similarity = 0.92
...
Segment 57   similarity = 0.90
Segment 58   similarity = 0.64

The average may still look excellent.

But segment 58 may sound like a completely different person.

This is why we need to distinguish:

Speaker similarity

Does this generated voice resemble the original?

from:

Speaker consistency

Does this generated voice remain recognizable as the same speaker throughout the entire video?

This distinction is especially important for long-form content.

And recent research suggests that even sophisticated audio-language models are not yet reliable judges for this problem. A 2026 ACL benchmark called SpeakerSleuth evaluated 1,818 human-verified instances across four datasets and 12 large audio-language models. The models struggled to reliably detect acoustic speaker inconsistencies; providing textual context could make performance dramatically worse because models tended to prioritize textual coherence over acoustic evidence.

That’s a fascinating result.

It means:

Even an AI model that understands the conversation may not be a reliable judge of whether the voice is actually consistent.


Naturalness Is More Than “Does the Voice Sound Good?”

Traditional speech synthesis evaluation often uses MOS—Mean Opinion Score.

Humans listen to generated speech and rate its quality.

For example:

1 ── Very poor
2 ── Poor
3 ── Fair
4 ── Good
5 ── Excellent

MOS remains useful because ultimately people are the audience.

But it has a fundamental limitation:

A listener may give a voice a 4.5/5 for naturalness while still giving the overall dub a 2/5.

Why?

Because naturalness is only one component.

A beautifully synthesized voice can still:

  • use the wrong emotion,
  • speak too quickly,
  • have the wrong pause,
  • sound like the wrong character,
  • overlap the next speaker,
  • contradict the video,
  • or pronounce a character’s name incorrectly.

This suggests an important principle:

Speech quality and dubbing quality are different evaluation problems.


Prosody: The Missing Layer Between Words and Voice

This is one of the dimensions I would emphasize heavily in the article.

Prosody includes things such as:

  • pitch
  • stress
  • rhythm
  • intonation
  • emphasis
  • pauses
  • phrasing
  • speaking rate

Consider:

“You really did that.”

The words are identical.

But:

YOU really did that?

sounds different from:

You REALLY did that?

which sounds different from:

You really did THAT?

The text is unchanged.

The performance is completely different.

Research on automatic dubbing has explicitly identified duration, lip movements, gestures, timbre, emotion, prosody, and environmental characteristics as relevant properties for making translated speech sound natural in the original video context.

So a useful evaluation framework should measure not only what was spoken, but also how it was performed.


Speaking Pace Is Its Own Metric

Speaking rate deserves to be separated from general timing.

Suppose the original speaker says:

"I don't think this is going to work."

at 3.8 words/second.

The translated version might be generated at:

2.4 words/second

or:

5.1 words/second

Both can be problematic.

We could measure:

Speech-rate ratio

Rpace=PdubPoriginal

where P is words/second, syllables/second, or phonemes/second.

But even this is not enough.

A good dub does not necessarily need to have the same speaking rate.

Different languages have different information densities and phonetic structures.

The real objective is closer to:

Does the target-language performance feel like a natural rendition of the original performance while fitting the available time?

This distinction becomes particularly important for languages with very different speaking-rate characteristics.

Recent research on duration-aware multilingual dubbing demonstrates how important this problem is: a 2025 EMNLP system that explicitly constrained translation duration achieved up to 24% relative improvement in speech overlap compared with translations without explicit length constraints, while maintaining competitive COMET translation scores.

That result is important because it demonstrates something fundamental:

Translation quality and timing quality are related—but they are not the same objective.


Emotion: The Sentence Can Be Correct and Still Be Wrong

Emotion is particularly difficult to evaluate automatically.

We might classify the original as:

Anger:       0.82
Sadness:     0.03
Joy:         0.05
Neutral:     0.10

and compare the generated performance.

But emotion isn’t just a classification label.

A human actor expresses emotion through:

  • pitch
  • loudness
  • speaking rate
  • pauses
  • voice quality
  • stress
  • articulation
  • breath
  • timing

Therefore, a system that simply predicts:

“This is an angry sentence”

still has to solve the much harder problem:

“How should anger be performed in the target language by this particular speaker?”

This is one reason why dubbing is closer to performance reconstruction than ordinary speech synthesis.


Volume and Audio Continuity Matter More Than We Think

There are also seemingly mundane dimensions that become very obvious when they fail.

For example:

For example

The generated voice might suddenly be 4 dB louder than the surrounding dialogue.

Or it might have:

  • different reverberation,
  • different microphone characteristics,
  • different background noise,
  • clipping,
  • compression artifacts,
  • unnatural transitions.

These errors may not be detected by translation or speaker-similarity metrics.

Yet they can immediately make the video feel artificial.

The 2026 research on human/AI dubbing evaluation also identifies sound design alongside synchronization, voice consistency/expressiveness, narrative cohesion, and translation as areas where human judgment remains important.


Lip Sync Is Important—but Lip Sync Is Not Dubbing Quality

Lip synchronization is probably the most visible AI-dubbing metric.

Systems such as SyncNet provide automated measures such as:

  • LSE-D — lip-sync error distance
  • LSE-C — lip-sync confidence

Generally, lower LSE-D and higher LSE-C indicate better audio-visual synchronization. These metrics are widely used in talking-head and lip-sync research.

This is extremely useful.

But consider two versions:

Version A

Translation:       95%
Voice similarity:  92%
Naturalness:       91%
Emotion:           88%
Lip sync:          70%

Version B

Translation:       95%
Voice similarity:  65%
Naturalness:       62%
Emotion:           40%
Lip sync:          98%

Which one is better?

There isn’t an obvious answer.

For a close-up dramatic scene, Version A might feel much more convincing.

For a talking-head YouTube video where the speaker’s mouth is visible throughout, lip sync might deserve much greater weight.

This illustrates an important principle:

A metric measures a property. It does not determine the importance of that property.


Video Category Changes the Definition of “Good”

This is where I think the evaluation framework can become particularly interesting.

Not every video needs the same quality profile.

TutorialPodcastDramaComedyMusic
Translation●●●●●●●●●●●●●●
Speaker●●●●●●●●●●●●●●
Emotion●●●●●●●●●●●
Pace●●●●●●●●●●●●●●
Lip Sync●●●●●●●
Rhythm●●●●●●●●

A dubbing system that performs extremely well on tutorials may not be suitable for songs.

Likewise, a system optimized for movies may unnecessarily spend computational resources on lip synchronization for a podcast where the speaker is never visible.

So I would introduce the idea of a Video Category Feasibility Profile

Instead of asking:

“Can this system dub videos?”

we should ask:

“For which categories of video can this system reliably produce acceptable results?”


The Worst Errors Are Not Always the Most Frequent Errors

This is another important production insight.

Suppose a 60-minute video contains:

  • 20 minor pronunciation issues
  • 10 slightly unnatural pauses
  • 5 small timing deviations
  • 1 wrong speaker assignment
  • 1 completely incorrect translation of a product name

A simple average metric might conclude:

98.5% of the content is good.

But imagine that the wrong product name appears in the most important scene.

The perceived quality of the video may be much worse than 98.5%.

This suggests that a production system needs severity-aware evaluation.

For example:

ErrorFrequencySeverityActionComedyMusic
Slight pronunciation deviationHighLowPass●●●●●
Slight pace mismatchMediumLowPass●●●●●●
Unnatural pauseMediumMediumPass●●●●●●
Wrong terminologyLowHighReview●●●●●●
Wrong speakerLowVery highPass●●●●
Missing dialogueLowVery highReview●●●●●●
Wrong emotion in critical sceneLowHighReview
Severe lip-sync failureLowHighPass
Timing89HighPass
Volume97HighPass
Lip sync91HighPass
Audio quality95HighPass

A useful conceptual model is therefore:

Qvideomean(Qsegments)

Instead:

Qvideo=f(average quality,consistency,severity,context)

This is why production QA cannot simply calculate one average score.


Segment-Level Quality vs. Video-Level Quality

This leads to another distinction.

A production system should evaluate at multiple granularities.

Segment-level QA

Can tell us:

“Segment 183 has a speaker mismatch.”

Scene-level QA

Can tell us:

“The emotional tone changes unnaturally between these three shots.”

Video-level QA

Can tell us:

“This character’s voice gradually becomes inconsistent over the course of the episode.”

This hierarchy is important because different errors exist at different scales.


Can an LLM Judge AI Dubbing?

This is becoming an obvious direction.

A multimodal model can potentially inspect:

  • the original video,
  • the dubbed video,
  • the transcript,
  • the translation,
  • speaker information,
  • timestamps,

and produce a structured evaluation.

For example:

Translation:       9.2 / 10
Speaker identity:  8.8 / 10
Naturalness:       8.5 / 10
Emotion:           7.1 / 10
Timing:            8.0 / 10
Lip sync:          7.6 / 10

Critical issues:
- Segment 183: wrong speaker
- Segment 227: emotional mismatch
- Segment 311: target speech overlaps next speaker

This is extremely attractive for large-scale production.

But we should be careful.

The recent SpeakerSleuth results are a good warning: current large audio-language models can struggle with acoustic speaker-consistency judgments, and textual context can actually bias them toward the wrong answer.

So the answer probably isn’t:

“Let an LLM judge everything.”

Instead:

Use specialized objective metrics for measurable properties, multimodal models for contextual reasoning, and humans for high-impact subjective decisions.


A Better Evaluation Architecture

This suggests a three-layer evaluation system.

This is fundamentally different from:

Generate → Watch → Looks OK

At production scale, we need:

Generate → Measure → Diagnose → Repair → Re-measure


The Quality Scorecard

A practical system might therefore produce something like:

CategoryScoreConfidenceAction
Translation96HighPass
Terminology98HighPass
Speaker similarity93HighPass
Speaker consistency87MediumReview
Naturalness91MediumPass
Prosody84MediumReview
Emotion78LowReview
Pace94HighPass
Timing89HighPass
Volume97HighPass
Lip sync91HighPass
Audio quality95HighPass

And importantly, we shouldn’t necessarily collapse these numbers into one scalar.

Instead, we might produce:

OVERALL
──────────────
GOOD

⚠ 3 segments require review

Segment 183
  Speaker consistency    FAIL

Segment 227
  Emotion                WARNING

Segment 311
  Timing                 FAIL

This is much more actionable.


What VMEG Can Learn From This

At VMEG, we process localization tasks at a scale where manually watching every finished video is not practical.

That changes the engineering problem.

The objective isn’t simply:

Build a better dubbing model.

It is:

Build a system that can continuously determine whether its own output is trustworthy.

This is where the architecture from the previous article becomes important.

End-to-end intelligence. Pipeline-level control.

The AI model can reason about the video globally.

But the production system needs to expose intermediate decisions:

Video
 ↓
Transcript
 ↓
Speaker map
 ↓
Translation
 ↓
Segmentation
 ↓
Timing
 ↓
Voice assignment
 ↓
Speech generation
 ↓
Audio mixing
 ↓
Lip sync
 ↓
Quality evaluation

Each stage creates opportunities for measurement.

And when something fails, we don’t necessarily want to regenerate the entire video.

If the problem is:

Segment 183
Speaker = wrong

we should fix speaker assignment.

If the problem is:

Segment 227
Duration = 3.4s
Target = 2.5s

we should fix timing or translation.

If the problem is:

Segment 311
Emotion = neutral
Expected = angry

we should regenerate the performance.

This is what makes a dubbing system debuggable.


From “Quality Score” to “Quality Engineering”

This is perhaps the bigger lesson.

The next generation of AI localization systems will not be judged only by the quality of their underlying models.

They will be judged by how reliably they can answer:

What went wrong?

And:

Can we fix only the part that went wrong?

And:

How do we know the fix actually improved the result?

That turns evaluation into a feedback loop:

This is the essence of production-grade AI.


The Future: A Dubbing Quality Vector, Not a Single Number

Perhaps the most useful abstraction is to represent every dubbed video as a quality vector:

Q=[T,S,N,P,E,R,V,C]

where:

  • T = translation accuracy
  • S = speaker similarity/consistency
  • N = naturalness
  • P = prosody
  • E = emotion
  • R = rhythm/pace/timing
  • V = visual synchronization
  • C = contextual consistency

The weights shouldn’t necessarily be fixed.

For a tutorial:

Qtutorial=f(T,R,N,pronunciation)

For a drama:

Qdrama=f(T,S,N,P,E,R,V)

For a podcast:

Qpodcast=f(T,S,N,P,E)

For a music video:

Qmusic=f(lyrics,rhythm,melody,voice,synchronization)

This is a much more realistic representation of quality than a single “AI dubbing score.”


Conclusion: Good Dubbing Is a System Property

The biggest mistake we can make when evaluating AI dubbing is to evaluate its components independently.

Translation can be excellent.

TTS can be excellent.

Voice cloning can be excellent.

Lip-sync can be excellent.

And the final video can still be bad.

Because the viewer doesn’t experience these components separately.

They experience one performance.

This is why AI dubbing needs a broader evaluation framework—one that measures:

meaning, voice identity, consistency, naturalness, prosody, emotion, pace, timing, audio quality, visual synchronization, and contextual coherence.

And perhaps most importantly, it needs to understand which of these dimensions matter most for the particular video being localized.

The future of AI dubbing evaluation is therefore unlikely to be a single benchmark.

It will be a quality system:

Measure every important dimension. Detect the failures that matter. Repair them locally. And continuously verify the result.

Because ultimately:

Model quality is not video quality.

And:

The quality of an AI-dubbed video is the quality of the reconstructed performance.


References

  1. Brannon, Virkar & Thompson, “Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing,” TACL 2023. The study covers 319.57 hours, 54 professionally produced titles, 674 episodes and 9,215 speakers. ⁠ACL Anthology
  2. Federico et al., “Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing,” Interspeech 2020. ⁠ISCA Archive
  3. Bane, “Evaluating End-to-End Speech-to-Speech Translation for Dubbing: Challenges and New Metrics,” AMTA 2024. ⁠ACL Anthology
  4. Won et al., “End-to-End Multilingual Automatic Dubbing via Duration-based Translation with Large Language Models,” EMNLP 2025. ⁠ACL Anthology
  5. Lee et al., “SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?” ACL 2026. ⁠ACL Anthology
  6. Spiteri Miggiani, “Quality in interlingual AI dubbing: Exploring machine and human intersections,” Paralleles, 2026. ⁠University of Geneva / Paralleles
  7. For a broader technical review of multilingual video dubbing and evaluation metrics, see the 2023 review in Frontiers in Signal Processing. ⁠Frontiers