Fonti.Studio

HomeGuides

Loudness, true-peak, LRA: the three numbers your AI dub vendor isn't measuring

An AI dub can clone the cast, fit every line to the cue, and still fail delivery on a number nobody on the project was watching.

The number is almost never the voice. It is the mix. A dub master is graded by a loudness meter before a human ever listens to it, and the meter reports three values that most AI dubbing pipelines do not optimise for and frequently do not measure at all: integrated loudness, true-peak, and loudness range. A file that sounds correct playing out of a laptop can sit two units hot, clip on a sibilant the browser smoothed over, and swing wide between a whisper and a punch. All three are delivery failures. None of them are audible as "wrong" on casual playback.

This is the most under-served stage of the AI dub stack. Cloning gets the attention because it is the visibly hard part. The mix is where the file actually passes or fails QC, and it is where the AI vendors we see in the field are weakest, because they treat mastering as a normalisation afterthought instead of a constraint the whole render has to meet.

The three numbers

Integrated loudness (LUFS, or LKFS). This is the average perceived loudness of the whole programme, measured the way an ear weights it rather than the way a peak meter reads it. The two acronyms are the same unit. LKFS comes from the ITU specification; LUFS is the name the EBU gave the identical measurement. Since the loudness gate was added in the second revision of ITU-R BS.1770 in 2011, the algorithm ignores silences so they do not drag the average down, and the two labels are interchangeable. One unit is one decibel. This is the number a delivery spec sets as its target, and the one a too-loud or too-quiet master fails.

True-peak (dBTP). Not the sample peak, the true peak. When a digital signal is reconstructed into analogue on its way to a speaker, the waveform between samples can rise above the highest sample value. A file that reads 0 dBFS on a sample meter can overshoot above zero on reconstruction and clip on the listener's hardware. True-peak, defined in the same BS.1770 family, oversamples the signal to estimate that inter-sample maximum. Specs cap it below zero precisely to leave room for the overshoot.

Loudness range (LRA). This is the spread between the quiet parts and the loud parts of the programme, measured statistically. EBU Tech 3342 defines LRA as the difference between the 10th and 95th percentiles of the loudness distribution, which is what keeps a single gunshot or a fade-out from dominating the figure. A low LRA is flat and fatiguing. A high LRA means the viewer rides the volume knob: dialogue too quiet to follow, then an action beat that wakes the house. LRA is the number that quietly governs whether a mix is comfortable, and it is the one AI pipelines ignore most completely.

The specs that set them

There is no single target, and that is the first thing an AI vendor has to get right. The number depends on where the file is going.

EBU R128 is the European broadcast standard. It sets integrated loudness at -23 LUFS, permits a tolerance of up to ±1 LU on less predictable material, and caps true-peak at -1 dBTP. The recommendation itself is short; the discipline is in hitting it across a feature without crushing the dynamics.

ATSC A/85 is the US broadcast standard, given legal force by the CALM Act. It targets -24 LKFS. The recommended practice uses the same BS.1770 measurement as Europe, set one unit lower, which is enough on its own to bounce a master tuned for the European number.

Netflix measures differently again. Its loudness specification asks for -27 LKFS (±3 LKFS) measured dialog-gated using ITU-R BS.1770-1, so the average is taken over the speech, not the explosions. If there is not enough dialogue to gate on, it falls back to -24 LKFS measured over the full programme. True-peaks must not exceed -1 dBTP. A dub vendor who masters to -23 and ships it to the streaming spec is four units hot before anyone opens the file.

Three destinations, three numbers, two measurement methods. The dub mix has to be authored to one specific target, not normalised to a vague "broadcast-ish" level and hoped through. A streaming platform may apply its own playback normalisation on top, but that adjusts the listener's volume; it does not repair a master that arrives out of spec on true-peak or LRA.

What goes wrong on AI dubs specifically

A traditionally recorded dub is mixed by an engineer on a calibrated stage who watches all three meters in real time. An AI dub is assembled from hundreds of independently synthesised cues, and that assembly introduces failure modes a human mix never has.

Per-cue level inconsistency. Each line comes back from the text-to-speech engine at its own loudness. One cue lands a few units quieter than the next, depending on phrasing and the synthesiser's internal gain. Concatenate a few hundred of those and the integrated figure can land anywhere, while the line-to-line jumps inflate LRA into a mix that lurches. A pipeline that normalises only the final stereo file fixes the average and leaves the lurch.

Sibilance and plosive true-peak spikes. Synthetic speech produces sharp, narrow transients on s, t and k sounds. These read fine on a sample meter and overshoot on true-peak reconstruction. A master that looks safe at -1 dBFS sample peak can still cross zero dBTP on a single hissed consonant, and a true-peak-aware QC tool flags it.

Silence-floor mismatch. The synthesised dialogue sits on a digital silence floor. The music-and-effects bed, preserved from the original mix, carries room tone and ambience. Drop the clean dialogue onto the textured bed without matching floors and every line entrance has an audible seam: the bed ducks to dead silence under the voice, then returns. The ear reads it as an edit even when no individual cue is wrong.

None of these are voice problems. They are mix problems, and they are exactly the problems a vendor who stops at "clone the actor" never sees.

The Fonti Studio mastering loop

We treat mastering as an optimisation, not a normalisation. Our dub_master stage runs a parameter sweep over the background level, the ducking depth, and the bed's high-pass, and scores each candidate against a composite that weighs several things at once:

The sweep picks the parameter set that maximises the composite rather than the one that hits a single number. On the reference run, that distinction was the whole point: re-tuning the background level and the ducking depth lifted the composite mastering score from 83.2 to 90.0. The dialogue got less isolated and sat over a more present music bed, and the mix got better, because a dub that buries the music to chase a dialogue number is not a finished mix. It is a voiceover.

The point of scoring against the composite is that you cannot game it by winning one term. A master that is dead-on its loudness target but clips on true-peak loses. A master that is pristine on peaks but flat on LRA loses. The optimiser is forced to satisfy all three loudness numbers and keep the dialogue intelligible, which is the same set of constraints the delivery QC tool is about to apply, run before the file leaves the building instead of after it bounces.

The gate downstream

Mastering sets the levels. It does not catch a voice that went wrong. Downstream of the mix sits a separate, mandatory check: a pitch QA gate that scans every cue's fundamental frequency against the character's expected range, on the median and on the 90th percentile, because a cloned male voice can break into a brief falsetto on a single line while the whole-track average hides it completely. The gate flags the cues that contradict the character, and each one is re-synthesised at high stability until it holds. Loudness mastering and pitch QA are orthogonal: one grades the mix, the other grades the performance, and a delivery has to clear both.

Ask the vendor the three questions

Most AI dub vendors will tell you their output is "broadcast loudness." That phrase is doing a lot of hiding. Ask which spec the master is tuned to. Ask for the true-peak figure. Ask for the LRA. A vendor who can answer all three is measuring all three; a vendor who answers "it's normalised" is measuring one and hoping on the other two.

If you have a feature heading into a platform delivery and you would rather not find out from the platform which number it failed on, send us the master and we will tell you which of the three it is failing on first.


Fonti Studio is an AI-native subtitle and dubbing service for film distributors and sales agents. Broadcast-grade output validated on feature-length deliveries. €3,500 per language per dub. Flat. EUR. Free 5-minute preview before you pay.

Heading into a delivery window?

Send a film master, trailer, and CCSL or original-language SRT. We send back a free preview in your target languages and tell you up front whether the master will pass.

Email us a brief