How to Make an SRT File: By Hand, From a Video or From Audio
How to make an SRT file: type it in Notepad or TextEdit, time it in a subtitle editor or generate it from the audio, then save it as UTF-8 so players read it.
To make an SRT file, write numbered subtitle blocks — a number, a start and end time, the text, a blank line — in a plain-text editor, and save it as UTF-8 with an .srt extension. Typing works for a short clip. For anything longer, generate the file from the audio and correct it.
Below: four ways to make one, the save-dialog traps on Windows and Mac, how to check a file before you upload it, and how to make one per language.
What goes inside an SRT file?
An SRT file is plain text, and two blocks show the whole format:
1
00:00:01,200 --> 00:00:04,000
Welcome back to the channel.
2
00:00:04,400 --> 00:00:07,850
Today we're fixing the audio
before anything else.
Four rules cover almost every file that works:
- Blocks are numbered from 1, in order. The number sits on its own line.
- Times are hours, minutes, seconds and milliseconds, with a comma.
00:00:04,400— the comma is SRT's, the period is WebVTT's. Start and end are joined by-->with a space on each side. - The text takes one or two lines. Around 42 characters per line is the usual ceiling for Latin scripts, and much less for Japanese or Korean; the reading-speed guide has the numbers per language.
- A blank line ends the block. Without it, the next block runs into this one.
Most styling does not survive in SRT. YouTube's help page is blunt about SRT: only basic versions are supported, no style markup is recognised, and the file must be plain UTF-8. If you need fonts, colours or positions, that is a different format — SRT vs VTT vs ASS covers which one carries what.
Four ways to make an SRT file
- By hand in a text editor — free and fine for a short clip, slow for anything long.
- In a subtitle editor — free, with a waveform to set the timings precisely.
- Generated from the audio — speech recognition writes the text and the timings, and you correct them.
- Exported from your video editor — when the captions already exist in Premiere Pro or DaVinci Resolve.
Method 1: Type it in Notepad or TextEdit
Any plain-text editor will do. What goes wrong is the save.
On Windows, with Notepad:
- Type the blocks in the format above.
- Choose File → Save as.
- Set Save as type to All files and name the file
subtitles.srt. - Set Encoding to UTF-8, then save.
Left on Text documents, Notepad can save subtitles.srt.txt instead. File Explorer hides known extensions by default, so the file looks right and still fails to load. Turn on file name extensions once in File Explorer's View menu, and that mistake stays visible for good.
On a Mac, with TextEdit:
- Open a new document and choose Format → Make Plain Text before typing. A default TextEdit document is rich text, and a rich-text file renamed to
.srtis full of formatting codes no player reads. - Type the blocks.
- Choose File → Save, name the file
subtitles.srt, and set the encoding to Unicode (UTF-8). If TextEdit asks whether to use.srtor.txt, keep.srt. - In Finder, check that the name really ends in
.srt. Finder → Settings → Advanced → Show all filename extensions makes that visible.
Typing is fine for a 30-second clip or a quick fix. For a ten-minute interview it means hundreds of timestamps read off a player, and the next two methods exist for exactly that.
Method 2: Time it against a waveform in Subtitle Edit
Subtitle Edit is free and open source, runs on Windows, macOS and Linux, and is built for this job. Subtitle Edit's download page has every version. Open the video and the audio appears as a waveform under the player: type a line, mark where it starts and ends against the visible speech, move on. Saving as SubRip writes a valid .srt, numbered and comma-formatted, so the format rules above stop being your problem.
It is still manual timing, only far faster than reading times off a clock. Subtitle Edit can also generate the text with Whisper from its speech-to-text menu; the offline subtitle generator comparison covers that side of it.
Method 3: Generate it from the video or audio
Speech recognition writes the text and the timings in one pass, and your job becomes correcting rather than typing. Whisper does this on your own computer, so nothing is uploaded, and it takes video and audio files alike. Online generators do the same job after you upload the file.
With the Whisper command-line tool (free, needs Python and FFmpeg):
whisper interview.mp4 --model turbo --output_format srt
That writes interview.srt into the current folder. Add --language es, or the code for whatever language is spoken, if detection guesses wrong.
With a desktop app: in Sablate, drop the video or audio file in, transcribe it, correct the text in the editor — which flags lines that read too fast — and choose Export → Subtitle file (.srt). The free plan takes files up to 10 minutes, three a month, with the Fast and Balanced Whisper models, and adds one short "Made with Sablate" line at the end of each exported .srt. Pro is $29 once: any length, every model, no credit line.
Either way, read the file before you ship it. Automatic transcripts go wrong in predictable places — names, numbers, punctuation — and Whisper has two stranger habits: writing a phrase such as "Thank you for watching" over silence, and repeating one line many times. Why Whisper repeats lines and invents text covers how to prevent both.
Method 4: Export it from your video editor
If the captions already exist in your editor, export them rather than retyping:
- Premiere Pro: with a captions track on the timeline, choose File → Export → Captions and pick SRT. The captions panel's menu also has Export to SRT file.
- DaVinci Resolve: choose File → Export Subtitle and pick
.srtor.vtt, or right-click the subtitle track header and choose Export Subtitle.
Creating the captions inside the editor is a separate question. Premiere's speech-to-text comes with its subscription; Resolve's is Studio-only, and the free-version route imports an .srt made somewhere else.
Why won't my SRT file load?
Usually for one of five reasons, and all five show up in a plain-text editor:
- It is really
subtitles.srt.txt. Windows and macOS hide known extensions by default. Turn extensions on and rename the file. - It was saved as rich text. Opened in a plain editor, a rich-text file begins with RTF control codes instead of the number 1. Start again from a plain-text document.
- It is not UTF-8. Accented letters turn into garbage, or the upload is refused; YouTube requires plain UTF-8. Re-save the file with UTF-8 selected.
- A timestamp uses a period.
00:00:04.400is WebVTT syntax. Some players forgive it, others skip the block or reject the file. - A blank line is missing. The next block's number and timecode get read as subtitle text, and that block disappears.
On macOS and Linux, two commands catch the problems you cannot see:
file subtitles.srt
grep -nE "[0-9]{2}:[0-9]{2}:[0-9]{2}\.[0-9]{3}" subtitles.srt
file names the encoding: UTF-8 or plain ASCII is fine, and anything that mentions ISO-8859 or extended-ASCII needs a re-save. The grep line prints every timestamp written with a period. On Windows, Notepad shows the encoding in its status bar. For a translated file, also count its blocks against the original — the SRT translation guide has that one-liner.
Making SRT files in more than one language
Every language is its own file with the same timings. There is no multi-language SRT: YouTube takes one upload per language, and media servers read the language from the file name. Plex, for one, documents the rule:
Interview (2026).mp4
Interview (2026).en.srt
Interview (2026).es.srt
Interview (2026).de.srt
The code before .srt is a two-letter ISO 639-1 code or a three-letter ISO 639-2/B one.
Keep one set of timings for every language. Correct the first transcript, then translate its text on those same timings instead of generating each language from the audio separately — separate runs produce cue boundaries that drift apart, and every sync fix has to be made once per language. The multilingual workflow goes through the order. In Sablate a translation is a layer on the original timings, and Pro's All languages · Subtitles export writes each language as its own .srt into one folder.
Which method should you use?
| By hand | Subtitle Edit | Whisper CLI | Sablate | Editor export | |
|---|---|---|---|---|---|
| Price | Free | Free, open source | Free, open source | Free tier; Pro $29 one-time | Included with the editor |
| Writes the timings for you | No | With its Whisper engines | Yes | Yes | Only if the editor transcribed |
| Video stays on your computer | Yes | Yes | Yes | Yes | Depends on the editor |
| Correction editor | Your text editor | Yes | No | Yes | Yes |
| Other languages | Type them yourself | Machine translation built in | Into English only | Translation layers on the same timings | Depends on the editor |
| Platforms | Any | Windows, Mac, Linux | Windows, Mac, Linux | Windows and Mac | Depends on the editor |
"It's a 20-second clip." Type it. Five minutes in Notepad or TextEdit, and the four rules above are all you need.
"I want every cue timed by hand, exactly." Subtitle Edit and its waveform.
"I have an hour of interview." Generate it: Whisper's command line if you are comfortable there, Sablate if you want an editor and a reading-speed check before export. Then correct it.
"The captions already exist in my editor." Export them. Never retype what a program has already timed.
"I need it in five languages." Generate and correct one, then translate on the same timings, one file per language.
An SRT file is the simplest subtitle format there is, which is why it breaks for simple reasons: a hidden extension, rich text, the wrong encoding, a period where a comma belongs. Get those right and any of the four methods produces a file every platform reads. If you would rather correct a transcript than type one, download Sablate for Windows or Mac and make the file from the audio.