Prosody, and why it still sounds slightly wrong
The information that is not in the words
Say "I never said she took the money" seven times, stressing a different word each time. Seven different meanings, one sentence. The stress carries information the text does not contain.
A text-to-speech system receives only the text. Everything about delivery — which word carries the emphasis, where the phrase breaks, whether the sentence rises at the end, how fast, how much energy — has to be inferred. The model does this from what usually happens with sentences of that shape, which is right on average and wrong on the specific sentence often enough to notice.
This is why synthetic narration passes at the sentence level and fails at the paragraph level. Each line is plausible; the sequence has no argument running through it. Human readers place emphasis according to what they are doing with the passage — contrasting, conceding, listing, concluding — and no such intention exists in the model.
The specific failures to listen for
Once you know the list, you will hear them:
- Contrastive stress missed. "It was not the red one, it was the blue one" delivered with even weight, so the contrast disappears.
- Lists flattened. Every item read identically, with no rise-rise-fall shape, so the listener cannot tell when the list ends.
- Questions that are not questions. Rising intonation on a statement, or a flat delivery on a genuine question.
- Wrong phrase breaks. A pause inside a noun phrase rather than at the clause boundary. This is the one that most reliably reveals a synthetic reader.
- Emotional monotony over length. Three minutes at exactly the same energy, which no person produces.
- Homographs. "Read", "lead", "live", "bow", "record", "tear". Systems guess from context and get it wrong on the harder cases, and in some languages the equivalent problem is far worse.
What actually improves it
Reference audio for style. The strongest control available. Supply a short clip in the delivery you want — energetic, conversational, sombre — and the model conditions on the style as well as the identity. This works far better than any adjective, for the reasons the reference-image lesson gave: you are supplying the thing rather than a word for it.
Rewrite the text for the ear. Shorter sentences. Explicit connectives. Punctuation that marks the breath. Spelling out numbers and abbreviations the way you want them read. This is real work and it improves human narration too. Anyone who has written for radio already knows it.
Markup where it is supported. Some systems accept tags for emphasis, breaks and rate. The dialect varies; where it exists it is precise and worth learning for the twenty lines in a script that matter.
Split and re-roll. Generate line by line, keep the deliveries that land, regenerate the ones that do not. Assembling in Audacity is free and this is how professional-sounding synthetic narration is actually made. Nobody generates ten minutes in one pass and ships it.
Direct it like a performance. Where a system accepts an instruction alongside the text — "read this as if explaining to a friend" — use it, and be specific about intent rather than emotion. "Sceptical, then convinced by the end of the paragraph" outperforms "happy".
The part that remains
There is a category of prosodic meaning that no current system produces reliably: the delivery that depends on understanding what the passage is arguing. Irony, dramatic timing, a pause that lands because of what came three sentences earlier, the shift in register when a speaker changes their mind.
This is not a data problem in an obvious way, and it is a reasonable place to expect slow progress. It is also the reason that synthetic narration is now entirely adequate for instructional material, announcements, audio descriptions and drafts, and still noticeably inferior for anything where the reading is an interpretation — audiobook fiction, poetry, advertising with a joke in it, drama.
Knowing which of those two categories your job is in saves an enormous amount of argument. For the first, use synthesis and put the effort into the script. For the second, hire someone, or record it yourself.
One more practical note, because it is the fix most people never try. If a line will not come out right after four attempts, the problem is almost always the writing rather than the model. Read the line aloud yourself. If you have to think about how to deliver it, so does the system, and it has less to go on. Splitting the sentence, moving the important word to the end, or replacing a subordinate clause with a full stop will usually produce a clean read on the first attempt. Writing for the ear is a real skill, it is learnable in a week, and it makes both synthetic and human narration better.
The one thing to keep
Prosody carries meaning that the text does not, so a system given only text has to guess the interpretation, and its guesses are wrong in ways that are subtle and cumulative.
Before you move on
Why does synthetic narration often pass at the sentence level and fail over a paragraph?
Pick the one you would defend. Nobody sees your answer.