Fonti Studio

HomeGuides

Subtitle translation: why translating an SRT file cue by cue gives you a file that looks right and reads wrong

In short: translating an SRT cue by cue forces the target language into English word order and past the reading-speed ceiling, and it lets names drift. A professional subtitle translation pass reconstructs whole sentences first, translates them with a locked glossary, re-segments onto the original timing, condenses to reading speed with a meaning review, and verifies alignment, flow and naturalness before the file ships.

Search for "translate SRT file" and every result is a tool: upload, pick a language, download in a minute. The result parses, the cue numbers match, the timecodes are untouched, and if you do not read the target language yourself it is very easy to ship.

Then a native speaker watches the film and tells you, politely, that it reads like a translation. Not wrong, exactly. Wrong where the lines break, wrong in how fast it goes by, wrong in the name that changes spelling forty minutes in. This is why that happens, what a proper subtitle translation pass does instead, and how to tell the difference before you pay.

An SRT is not a text, it is a timing grid

The mistake is built into the file. An SRT looks like a list of short sentences with timestamps, so the obvious move is to translate each one. But a professionally spotted subtitle (an SRT or any other format) is not a list of sentences. It is sentences sliced into fragments, each timed to the frame against the actor's delivery.

"...and that's why, after everything, I left." arrives as three cues: "...and that's why," then "after everything," then "I left." Each sits exactly where those words are spoken, and in English the fragments still read as English.

Translate them fragment by fragment and you have forced Dutch, French or German into English word order. Those languages move clauses: French may want "I left" at the front and the reason at the end; Dutch puts the verb elsewhere in a subordinate clause. A cue-by-cue tool never sees the sentence, only the fragment, so it produces three fragments that are each a fair translation of their English piece and, read in sequence, are not a sentence anyone would write. The breaks land in the wrong place, the verb turns up in the wrong cue, and the viewer reassembles it in their head at 24 frames per second.

Broadcasters wrote this down long ago. The BBC Subtitle Guidelines ask for subtitles that contain single sentences, broken at natural points with "priority given to linguistic considerations", and Netflix's Timed Text Style Guide lists what a break must not separate: a noun from its article, a first name from a last name, a verb from its subject pronoun. A tool translating inside the fragment boundaries cannot honour any of that, because the boundaries were drawn for a different language.

What the fix looks like

Three steps, in this order.

First, reconstruct whole sentences. Before anything is translated, the fragments are joined back into the sentences they came from, and short cues are combined with their neighbours into units that carry a complete thought. It is the step every one-click tool skips.

Second, translate the sentence with its context, so it is rendered the way the target language actually says it.

Third, re-segment the translation back onto the original cue grid: the same number of cues, split where the target language breaks naturally, each inheriting its timing from the original spotting. Timing is never invented. A cue may be extended into the idle gap before the next one to bring its reading speed down, and a cue start always tracks the moment the words are spoken. Nothing else moves. The output is still an SRT on the same clock, but each cue is now a piece of a target-language sentence, not a translated piece of an English one.

Reading speed: translation and condensation are one job

The second thing a tool gets wrong is invisible in the file and obvious on screen. In our deliveries, Dutch and French run roughly 10 to 25 percent longer than the English they translate. The time on screen does not grow with them, because the time is the actor's.

The limits are fixed. Netflix's language guides cap reading speed at 20 characters per second for adult programmes and 17 for children's, with 42 characters per line, two lines per event, and a duration between five sixths of a second and seven seconds (the full set, and where deliveries fail it, is in our Netflix subtitle requirements piece). A cue that read comfortably at 17 characters per second in English is over the ceiling in French the moment it grows by a fifth. On a fast, dialogue-dense film we have seen around a hundred cues per language land above 20 after a faithful first translation.

So the translator condenses while translating: drops a redundant adverb, picks the shorter synonym, restructures to fit the seconds available. Ordinary craft, but condensing changes meaning if nobody checks: we have watched a shortening pass turn "general subjects" into "the classics" in a school scene, a fluent line that no longer said what the character said. Condensation has to be followed by a faithfulness review of the condensed cues, not run as the last step before export.

One caveat: some cues cannot be brought under the limit without damage, because the English source was already reading at 25 or 30 characters per second. There the right call is to match the source's pace, and a good vendor tells you which cues those are rather than quietly trimming meaning to hit a number.

Names and terms: the glossary, and the connotation trap

Names drift. A character is Shauna in reel one and Shana in reel five because translation happens in chunks that decide independently. The fix is a locked glossary, built and reviewed before a line is translated and enforced by a validator afterwards, so an unapproved spelling is an error, not a choice.

There is a subtler version: the term the source keeps in English. A tool copies an anglicism or a technical term through because that looks safe, but the term can mean something else in the target language. A military rank that is the correct dictionary equivalent can read on screen as "the commanding officer" rather than a rank; a place-phrase that is neutral in English can carry a specific historical association in French. No spell-check catches these. A reviewer has to ask what the word connotes in the target language, not only what it denotes, so a well-run pass flags every capitalised term carried over verbatim and asks that question once per term.

Verification a tool does not do

This is what separates a subtitle translation service from a button. After translation and condensation, the file goes through checks independent of the model that produced it.

A faithfulness review by a second, independent reader or model, comparing each target cue with its source: meaning changes, dropped clauses, wrong register (the "tu" that should have been "vous"), profanity a polite model softened.

A semantic alignment check: does each cue mean its own source line? One bad batch can shift a region of forty cues by one position, so every cue carries its neighbour's translation: right length, right reading speed, grammatical, entirely wrong. An embedding-based comparison of each cue against its own source and its neighbours finds it in seconds; a human spot-check of a ninety-minute file usually does not.

A cross-cue flow read: a sentence split across cues must read as one sentence when the cues are concatenated. No capital opening a continuation, no full stop closing a clause that continues. Reviewers read the cues in order, as a script.

A naturalness pass by a native speaker, because a line can be accurate, the right length and grammatical and still be something no one says: a bare infinitive opening a sentence, a calque of an English idiom. This is what clients notice first, usually on the trailer.

And the format checks, the cheap part: reading speed, line length, line count, overlaps, minimum duration, gaps, stray sound descriptors from the source SDH. You can run those on any file in the free subtitle validator. Passing them is necessary and nowhere near sufficient.

What the full chain produces: The Last Kumite (2024), Dutch and French: 0 errors against the Netflix Branded Delivery subtitle spec across 1,212 cues per language.

What to ask a vendor, and what it costs

Four questions sort vendors quickly. Do you translate whole sentences or cues? When is the glossary built, and do I see it? Who reviews the condensed cues for meaning? Will a native speaker read the file before I do?

On price, the traditional market quotes per runtime minute and adds QC on top, so a feature into one language lands in the low four figures; our localization cost breakdown does the arithmetic. Fonti Studio prices subtitle translation flat: €400 per language for a feature, €50 per language for a trailer, with a free preview in your target languages within 48 hours of a producer approving the brief, and the full masters within 48 hours of payment. The service, including per-reel and bilingual delivery, is on the subtitling page.

Common questions

Can I just translate an SRT file with AI?

You can, and the file will parse; the problem is what it reads like. Cue-by-cue translation forces the target language into the source language's word order, so line breaks land in the wrong place, reading speed climbs past the platform ceiling as the text expands, and names drift. Whole-sentence reconstruction, condensation with review, a locked glossary and a native read separate a usable track from a parsed one.

How much does subtitle translation cost?

At Fonti Studio, €400 per language for a feature and €50 per language for a trailer, flat, with a free preview before you pay. Per-minute vendors typically land in the low four figures per language for a feature once QC is included; tool subscriptions are cheap until you add the human pass that makes the output shippable.

How long does subtitle translation take?

At Fonti Studio, a free preview in your target languages comes back within 48 hours of a producer approving the brief, and the full masters within 48 hours of payment, for features up to 180 minutes. That window covers translation, condensation, the review passes and the format gates.

What is the best way to translate subtitles?

Reconstruct the sentences before translating, translate with context and a locked glossary, re-segment onto the original timing, condense to reading speed with a meaning review, then verify: faithfulness, alignment, cross-cue flow, naturalness, format. More steps than a button, which is why the button's output reads the way it does.

The short version

Translating a subtitle file is re-authoring the same timing grid in a language with different word order and more characters, then proving every cue still means what the actor said. Fonti Studio does exactly that: glossary-locked, reviewed subtitle translation into Dutch, French and other target languages, flat per language, with a free preview before you pay. Send the master or the trailer to hello@fonti.studio.

Heading into a delivery window?

Send a film master, trailer, and CCSL or original-language SRT. We send back a free preview in your target languages and tell you up front whether the master will pass.

Send a brief

Prefer email? hello@fonti.studio