Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Making Things With AI

Images, video, voice and music — how they work, where they break, who owns them.

Lesson 55 of 848 min

The dubbing chain, and where it breaks

Four stages, four failure modes

Automatic dubbing chains together:

  1. Recognition — speech to text in the source language.
  2. Translation — source text to target text.
  3. Synthesis — target text to speech, often in the original speaker's cloned voice.
  4. Alignment — fitting the new audio to the timing, sometimes with lip adjustment.

Each stage takes the previous stage's output as truth. A name misheard at stage one is translated confidently at stage two, spoken clearly at stage three, and lip-synced convincingly at stage four. The pipeline's output is fluent, well-timed and wrong, with nothing anywhere signalling a problem.

This compounding is the central thing to understand. If each stage is 95% right, the chain is not 95% right.

The specific breakages

Recognition errors on names and terms propagate all the way through, and they are precisely the words a viewer will notice.

Translation without register. Machine translation chooses a register and often chooses wrong: formal where the speaker was casual, or the wrong second-person form in languages that distinguish them. In several languages this makes a speaker sound rude or absurdly deferential, which is a serious failure that no automatic check catches.

Length mismatch. Languages differ in how long the same content takes to say — a translation can be 30% longer than the original. The pipeline then either speeds up the speech, which sounds rushed, or lets it run over the shot, which breaks the edit. Good human dubbing solves this by rewriting for length, which is a craft skill with a name — adaptation — and it is the part automatic systems handle worst.

Cloned prosody carried across languages. A voice print captures the speaker's identity; the delivery patterns of the source language do not necessarily suit the target. The result can sound like a foreigner reading the target language even though the accent is correct.

Cultural content. Idioms, jokes, references and units. A joke translated literally is not a joke, and nothing in the chain knows a joke was there.

Where automatic dubbing is genuinely good

Being fair about this matters, because the failure list makes it sound useless and it is not.

It works well for informational content with a controlled script: training material, product explanations, lectures, documentation. The register is neutral, the terms can be supplied in advance, the timing is flexible, and there are no jokes.

It works well as a draft for a human adapter, who then fixes register, length and the twenty lines that matter. This is the shape most professional workflows have converged on, and it makes translation affordable for material that would never have been dubbed at all.

It works well for accessibility, where a rough version now beats a perfect version never.

The controls that make it safe

Check the transcript before translating. Ten minutes of review at stage one removes most of the compounding. This is the single highest-value intervention in the chain.

Have a speaker of the target language review the translation. Not a checker of grammar, a checker of register and meaning.

Listen to the final audio against the original. Not read it — listen.

Disclose it. A viewer watching a dubbed video in a cloned voice should know. Several jurisdictions now require this in some form, which the final module covers, and the ethical case does not depend on the legal one.

Get consent for the voice, from the speaker, specifically for synthetic reproduction in other languages. A performer's agreement to be filmed is not an agreement to speak Portuguese.

There is a cheaper alternative that deserves more consideration than it gets: subtitles. They are far more reliable, they cost a fraction as much, they preserve the original performance entirely, and a large part of the world's audience already prefers them. Dubbing is the right answer for young children, for people who cannot read the subtitles comfortably, and for material watched while doing something else. For everything else, the choice between a good subtitle track and a mediocre automatic dub is not close, and the dub is frequently chosen because it sounds more impressive rather than because it serves the viewer better.

The honest limitation of the whole area: this is a chain of statistical systems, none of which knows what it is talking about, producing output that is confident at every stage. It is enormously useful and it is not a fire-and-forget pipeline. Every organisation that has deployed it at scale has learned this by publishing something embarrassing, and the ones who put a human at stage one learned it more cheaply.

The one thing to keep

A dubbing pipeline is four error-prone models in series, so errors compound and the only reliable control is a human check at each stage rather than at the end.

Before you move on

Why is checking the transcript the highest-value intervention in an automatic dubbing chain?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly