English
EN
← Research

Dubbed pace: what we can measure, and how we keep it natural

Absolute syllable-rate meters break down on noisy, cut, or short source audio. Relative perceived pace tracks the ear better, so we expand neighboring slots first, rewrite when needed, and change playback speed last.

If a dubbed line sounds rushed or dragging, can you fail it with syllables per second or how often the loudness bounces? Not in this pipeline. Source audio is often noisy, cut, or unclear, and timestamps drift. Both envelope-based and forced-alignment-based syllable rates can disagree with perceived pace. A relative pace score (how much the dub is stretched or squeezed versus the original, then adjusted for the language pair) matches the ear on the 14 lines below.

Human dubbers mostly follow translation length, not speaking rate; when the two conflict, they would rather loosen timing than force an unnatural pace (Brannon et al., 2023). Tight timing windows make synthesized speech too fast, too slow, or uneven across neighboring lines; a little boundary slack helps (Federico et al., 2020). Stretching synthesized audio afterwards reduces intelligibility and naturalness (Sahipjohn et al., 2024; Choi et al., 2025). Our order is therefore: expand neighboring time first, rewrite the translation when needed, and change playback speed last. The material below comes from a live report of uneven pacing: 14 Hindi lines dubbed into Indian English.

What phonetics already knows

Cross-language comparisons rarely use words per minute. They count syllables or phones. That can explain why Japanese often sounds faster than Mandarin. Dubbing asks a different question: does this translated line sound like a person talking.

Pellegrino et al. (2011), seven languages: denser languages tend to speak syllables more slowly. On their materials, Japanese is about 7.84 syl/s, English 6.19, Mandarin 5.18. The tradeoff does not make information rate identical (Japanese relative to Vietnamese about 0.74, English about 1.08). Coupé et al. (2019), 17 languages: information rate averages about 39 bit/s (SD about 5); syllable rate about 6.63 syl/s. That is a tendency across many speakers, not a pass/fail line for one dubbed sentence. We did not estimate bits per second here.

Two algorithms we actually used.

Envelope methods ignore text. They use energy peaks as an approximation of syllable nuclei and estimate the number of syllables per unit time, then divide by duration (Mermelstein, 1975; Morgan & Fosler-Lussier, 1998; Wang & Narayanan, 2007; De Jong & Wempe, 2009). No transcript needed. They were mostly checked inside one language against hand-marked syllable nuclei. A hard cut or a jump in the noise floor makes a bump that is not a syllable.

Forced-alignment syllable rate. The written transcript is force-aligned to the audio, after which we count the syllables within each time window. We use Montreal Forced Aligner (McAuliffe et al., 2017; GitHub). Count source and dub, then divide: above 1 looks like a faster dub. Short lines, unclear speech, and Hindi versus English all magnify alignment error. Syllable rate and phone rate capture different aspects of speech rate, so they are not interchangeable (Trouvain & Möbius, 2014).

Perceived pace is not either count alone. Pfitzinger (1998, 1999) found listeners track a mix of syllable and phone rate (about 0.88–0.91). Foreign speech also often sounds faster than it is (Pfitzinger & Tamashima, 2006; Bosker & Reinisch, 2017). A dub is heard by target-language listeners against a picture, not against a lab-average syllable rate.

Four cases

Listen to the finished dub in each case. Material: 14 Hindi lines dubbed into Indian English.

Case 1: Why envelope rate can fail

Listen to the finished dub: it does not drag; if anything, it is slightly fast. “Madam, No, Saturn with Rahu can go to hell, sorry. Every child brings their own karma.” Envelope rate scores it as slower than the source, while perceived pace does not. The source contains hard edits and jumps in the noise floor. Those jumps create extra energy peaks, inflate the estimated source rate, and make the dub look slow.

Dub:

Source (listen for hard cuts and jumps in the noise floor):

Case 2: Why forced-alignment syllable rate can fail

Listen to the finished dub: it does not sound slow. “She needs training, she doesn't need remedies.” This is the shortest of the 14 lines. Forced alignment gives it a much lower syllable density than the source, while perceived pace is normal. In such a short window, a small alignment error, magnified by comparing two languages, can label a normal line as slow.

Dub:

Source:

Case 3: Perceived pace catches a dub that sounds unnatural

Listen to the three-line finished dub. The middle line is: “Child, the area around your home is opening up, because she can read energy.” Perceived pace and envelope rate both call it slow, matching the ear. Phone and syllable rates relative to the source are above 1, as if the dub were faster. But the source itself is slow, and this kind of consultation is not meant to rush. A direct cross-language syllable-rate comparison cannot separate a language's inherent rhythm from a pacing mismatch between the dub and the picture.

Dub (three lines):

Source (three lines):

Case 4: Small neighbor-boundary adjustments can smooth the pace

The same three lines: one neighbor rushes while another drags. We are experimenting with moving their boundaries slightly, guided by perceived pace. This is not the production default. Listen to the current dub first, then the adjusted version.

Current dub (three lines):

The same three lines after neighbor-boundary adjustment guided by perceived pace (experimental):

What alignment and rewrite actually do

Chaume (2012) splits lip sync into three jobs: when sound starts and stops, whether the mouth looks right, and whether the body acting still makes sense. Matching duration is only the first. It does not require the target language to speak at the source’s syllable rate.

Brannon et al. (2023), 54 titles, 300+ hours: translation length tracks dubbed duration at about 0.52 (0 = unrelated, 1 = lockstep); speaking rate tracks duration at about 0.16. Spanish and German dubbers vary rate less than the English source. Forced to choose, people break timing rather than warp rate.

Alignment still needs the line near the original mouth. Endpoints can slacken a little so TTS is not crushed (Federico et al., 2020). Some synthesizers hit the duration at generation time instead of stretching afterwards; the latter hurts intelligibility (Sahipjohn et al., 2024; Choi et al., 2025). Rewriting the translation is the same idea as putting duration constraints into MT (Saboo & Baumann, 2019). We still keep playback-speed change. We just put it last.

So the current order is: expand neighboring blocks, rewrite when needed, and change playback speed last.

Perceived pace compares the dub's timing with the source, then adjusts for the usual pacing of that language pair. It asks whether the dub feels fast or slow once placed on the timeline, not which of two original-speed TTS versions is faster. On these 14 lines it agrees with the ear; envelope rate, phone density, and syllable density each have a counterexample. We use it to explain problems and in the Case 4 neighbor-boundary experiment, not as a pass/fail rule for every line.

Overly fast lines are still worth limiting. A mandatory lower bound for slow lines is more costly: in a sample of 50 production lines that already sounded slow, most timing windows could not be shortened further without making the voice end while the mouth was still moving.

Expanding neighboring blocks reduces the need for speed changes. Rewriting changes how much must be said, so perceived pace and lip sync can improve together, consistent with Saboo and Baumann's use of duration constraints in translation.

Takeaway

Absolute speech-rate measures can describe speech, but they cannot by themselves decide whether a dubbed line sounds natural. Noise, edits, short windows, alignment errors, and cross-language rhythm can all make envelope or syllable rate disagree with the ear. Perceived pace fits these 14 lines better, but it is evidence, not a new universal threshold.

The engineering order matters more than another pass/fail number: give neighboring lines more room first, rewrite how much must be said when needed, and change playback speed last. The goal is not to make one metric pass. It is to keep voice, meaning, and visible articulation natural together.

This conclusion comes from 14 Hindi-to-Indian-English lines. Neighbor-boundary tuning remains experimental, and perceived pace still needs counterexamples from more language pairs.

References

  1. Pellegrino, F., Coupé, C., & Marsico, E. (2011). A cross-language perspective on speech information rate. Language, 87(3), 539–558. https://doi.org/10.1353/lan.2011.0057
  2. Coupé, C., Oh, Y. M., Dediu, D., & Pellegrino, F. (2019). Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche. Science Advances, 5(9), eaaw2594. https://doi.org/10.1126/sciadv.aaw2594
  3. Trouvain, J., & Möbius, B. (2014). Sources of variation of articulation rate in native and non-native speech: Comparisons of French and German. Speech Prosody 2014, 423–427. PDF
  4. Dellwo, V., & Wagner, P. (2003). Relations between language rhythm and speech rate. Proceedings of ICPhS 15, 471–474. PDF
  5. Pfitzinger, H. R. (1998). Local speech rate as a combination of syllable and phone rate. ICSLP 1998. https://doi.org/10.21437/ICSLP.1998-545
  6. Pfitzinger, H. R. (1999). Local speech rate perception in German speech. ICPhS 14, 893–896. PDF
  7. Pfitzinger, H. R., & Tamashima, M. (2006). Comparing perceptual local speech rate of German and Japanese speech. Speech Prosody 2006. https://doi.org/10.21437/SpeechProsody.2006-30
  8. Bosker, H. R., & Reinisch, E. (2017). Foreign languages sound fast: Evidence from implicit rate normalization. Frontiers in Psychology, 8, 1063. https://doi.org/10.3389/fpsyg.2017.01063
  9. Koreman, J. (2006). Perceived speech rate: The effects of articulation rate and speaking style in spontaneous speech. Journal of the Acoustical Society of America, 119(1), 582–596. https://doi.org/10.1121/1.2133436
  10. Plug, I., Lennon, R., & Smith, R. (2023). Testing for canonical form orientation in speech tempo perception. Quarterly Journal of Experimental Psychology. https://doi.org/10.1177/17470218231198344
  11. Mermelstein, P. (1975). Automatic segmentation of speech into syllabic units. Journal of the Acoustical Society of America, 58(4), 880–883. https://doi.org/10.1121/1.380738
  12. Morgan, N., & Fosler-Lussier, E. (1998). Combining multiple estimators of speaking rate. ICASSP 1998. https://doi.org/10.1109/ICASSP.1998.675368
  13. Wang, D., & Narayanan, S. (2007). Robust speech rate estimation for spontaneous speech. IEEE Transactions on Audio, Speech, and Language Processing, 15(8), 2190–2201. https://doi.org/10.1109/TASL.2007.905178
  14. De Jong, N. H., & Wempe, T. (2009). Praat script to detect syllable nuclei and measure speech rate automatically. Behavior Research Methods, 41(2), 385–390. https://doi.org/10.3758/BRM.41.2.385
  15. McAuliffe, M., Socolof, M., Mihuc, S., Wagner, M., & Sonderegger, M. (2017). Montreal Forced Aligner: Trainable text-speech alignment using Kaldi. Interspeech 2017, 498–502. https://doi.org/10.21437/Interspeech.2017-1386
  16. Montreal Forced Aligner GitHub repository: https://github.com/MontrealCorpusTools/Montreal-Forced-Aligner
  17. McAuliffe, M., et al. (2026). Montreal Forced Aligner and the state of speech-to-text alignment in 2026. https://arxiv.org/abs/2606.18466
  18. Brannon, W., Virkar, Y., & Thompson, B. (2023). Dubbing in practice: A large-scale study of human localization with insights for automatic dubbing. Transactions of the Association for Computational Linguistics, 11, 419–435. https://doi.org/10.1162/tacl_a_00551
  19. Chaume, F. (2012). Audiovisual Translation: Dubbing. St. Jerome Publishing / Routledge. ISBN 978-1-905763-91-7. Routledge
  20. Federico, M., Virkar, Y., Enyedi, R., & Barra-Chicote, R. (2020). Evaluating and optimizing prosodic alignment for automatic dubbing. Interspeech 2020, 1481–1485. https://doi.org/10.21437/Interspeech.2020-2983
  21. Saboo, A., & Baumann, T. (2019). Integration of dubbing constraints into machine translation. Proceedings of the 4th Workshop on Discourse in Machine Translation (DiscoMT 2019), 102–107. https://aclanthology.org/D19-5614/
  22. Sahipjohn, N., et al. (2024). DubWise: Video-guided speech duration control in multimodal LLM-based text-to-speech for dubbing. Interspeech 2024. https://doi.org/10.21437/Interspeech.2024-399
  23. Choi, J., Kim, J., & Chung, J. S. (2025). Dub-S2ST: Textless speech-to-speech translation for seamless dubbing. Findings of EMNLP 2025, 9871–9881. https://doi.org/10.18653/v1/2025.findings-emnlp.524