Subtitle converter

Runs in your browser. Your subtitle file is never uploaded.

Or paste subtitles
Output

A positive shift shows subtitles later, a negative one earlier. The frame-rate change is applied first, then the shift. Two-point sync, if both times are filled in, fixes lag and drift together: it moves the first cue to the first time and the last cue to the last time and stretches everything between. Play your video to find when the first and last lines are actually spoken.

Line length and clean-up

What each format can hold

The converter reads every format into the same list of cues (start, end, text) and writes that list out again. Anything a format cannot hold is dropped on the way, and the report says so.

SubRip (.srt) WebVTT (.vtt) ASS / SSA (.ass, .ssa) YouTube SBV (.sbv)
Timestamp 00:01:02,500 00:01:02.500 (hours optional) 0:01:02.50 (hundredths) 0:01:02.500
Precision 1 ms 1 ms 10 ms 1 ms
Cue numbers Required, ignored on reading Optional identifiers, kept None None
Italic, bold, underline <i> <b> <u> <i> <b> <u> {\i1} {\b1} {\u1} No formatting
Position, colour, fonts <font> only (not standard) Cue settings, STYLE blocks Full styling and positioning None
Speaker names No <v Name> voice tags Name field No
Used by Most players and editors HTML5 video, YouTube, Vimeo Anime fansubs, Aegisub YouTube Studio

Conversions keep italic, bold and underline between SRT, WebVTT and ASS. WebVTT voice tags become "Name: " at the start of the line, because SRT and SBV have no speaker field. WebVTT cue settings (position, alignment) are kept when the output is WebVTT and dropped otherwise. ASS styling beyond italic, bold and underline is dropped when converting to another format; converting ASS to ASS (to shift times, for example) keeps the original script info, styles, fields and override tags of every cue whose text you did not change. ASS stores hundredths of a second, so converting to ASS rounds times to the nearest 10 ms.

Fixing subtitles that drift: frame rates

If subtitles start in sync and slowly drift further off, they were timed for a video at a different frame rate. The usual case is a film at 23.976 frames per second released in a 25 fps (PAL) country: the video plays about 4% faster, and an hour of film lasts 57.5 minutes. The fix multiplies every time by the ratio of the two rates:

new time = old time x (frame rate the subtitles were made for / frame rate of your video)

From 23.976 to 25 fps, the factor is (24000/1001) / 25 = 0.959041, so a cue at 1:00:00.000 moves to 57:32.547. The converter uses the exact NTSC rates (23.976 is really 24000/1001 and 29.97 is 30000/1001). The rounded figures would be off by about 4 ms per hour: small, but there is no reason to add it.

If subtitles are off by the same amount all the way through, it is not a frame-rate problem. Use the shift instead: a positive number of milliseconds shows every cue later, a negative one earlier. Cues pushed before 0:00 are trimmed to start at 0:00, or dropped if they would end before it. When both are set, the frame-rate change is applied first and the shift second, so the shift is measured on your video's timeline.

Two-point sync fixes both at once. Play the video, note when the first and last lines are actually spoken, and enter those two times. The converter moves the first cue to the first time and the last cue to the last, and stretches everything between in proportion (new time = old time x rate + offset, solved from the two points). It reads the start of the first and last cue, so pick lines that start on a spoken word. It runs after the frame-rate change and the shift; leave those at zero when you use it.

Line length and cue splitting

Subtitle style guides set a limit on characters per line so a line can be read at a glance. Netflix's English guide allows 42 characters per line and two lines per cue, and many broadcasters use 37 to 42. The converter's default is 42 and two lines; change both to match your platform.

With Re-wrap long lines on, a cue that breaks the limits is wrapped again at word boundaries into lines of similar length (when two lines are needed, the split that makes the longer line shortest wins). A cue whose text cannot fit in the allowed lines is split into several cues. Its time is shared in proportion to the characters in each part, and italic or bold that runs across the split is closed and reopened, so every cue stays valid. Cues that already fit are left exactly as they were. A single word longer than the limit gets a line of its own.

Merge flash cues joins a cue shorter than 1.2 seconds to its neighbour when the gap between them is at most half a second, the result lasts no more than 7 seconds, and the joined text still fits the line limits.

What the validation report checks

Check Level Why it matters
Bad timestamp (unparseable, minutes or seconds of 60 or more) Error, cue skipped Players reject the file or skip the cue.
Cue ends before it starts Error The cue never shows; some players reject the whole file.
Starts and ends at the same time Warning The cue never shows.
Overlaps the previous cue Warning Two cues on screen at once, or one hidden, depending on the player. Fix overlaps sorts the cues and trims each one to end where the next begins.
Starts before the cue above it Warning Some players and YouTube skip cues that are out of order.
Empty text Warning A blank flash on screen.
Missing blank line, junk between cues, cue number that is not a number, missing WEBVTT header Warning The file is malformed; the converter reads it anyway and writes a clean one.
Line longer than the limit, more lines than the limit Note Hard to read, especially on phones.
More than 20 characters per second Note Viewers cannot finish reading. Netflix's adult English limit is 20.
Shown for less than 0.7 seconds Note Too short to read.

Every issue names the cue number and the line in your file, so you can find it in an editor.

Encodings

Files are read as UTF-8 (with or without a byte-order mark) or UTF-16 with a byte-order mark. A file that is not valid UTF-8, typical of older .srt files from Windows, is read as Windows-1252, and the result names the encoding it used. Output is always UTF-8. Choose Windows line endings if an old player needs them.

Privacy

The file is read by your browser and converted in the same tab. It is never uploaded, and the page keeps working with the network switched off. See the privacy page.

If your subtitles came from a speech-to-text tool, the WER calculator shows how many words it got wrong, and the benchmark shows how open-source Whisper models do on recordings like yours.

Also available as Markdown.