How to Add Word-by-Word Karaoke Captions to Shorts and Reels
Add word-by-word karaoke captions to Shorts, Reels and TikTok: where the word timing comes from, three ways to burn it in, and the edits that break the sync.
Word-by-word captions need one thing an ordinary subtitle file does not have: a timestamp for every word. Get those from a speech recogniser such as Whisper, choose how the spoken word stands out — a colour change, a box, a fill or one word at a time — and burn the captions into the video.
Burning in is not optional here. Shorts, Reels and TikTok accept no subtitle files, so the highlight has to live in the pixels. Below: where the word timing comes from and why it drifts, three ways to add the effect — one of them free — and the edits that quietly break the sync.
Why word-by-word captions need word timestamps
A normal subtitle cue has two timestamps: when the line appears and when it goes. That is enough for a line that sits still. A karaoke highlight needs a start and an end for every word inside the line, and a standard .srt has nowhere to store them.
Whisper produces word timings as a second pass. Once it has settled on the text, it aligns each word with the stretch of audio that most likely produced it, and the openai-whisper command line still labels the option experimental. The timings are good enough to drive a highlight and not good enough to trust blindly — and they come with a quirk that decides how the highlight behaves.
Adjacent words often share a boundary exactly. In the clip we examined for our Whisper model comparison, 172 of 177 gaps between words were exactly zero seconds. Where real pauses do exist, a tool has to decide what the highlight does during them:
- Light only the word being spoken. The highlight blinks off in every pause, and a short word such as "a" or "it" flashes for a frame or two. Whisper's own highlighter works this way: it writes a plain, unhighlighted copy of the line into every gap.
- Keep each word lit until the next one starts. The highlight never drops, and every word stays on screen long enough to register. This is how most caption apps behave, and it is what viewers of short-form video are used to.
The second approach also survives editing better, because the word windows cover the whole cue from its first word to its last, with nothing left over.
The four karaoke caption styles
"Karaoke" covers four different looks, and it helps to know which one you want before choosing a tool:
| Style | What the viewer sees | Often called |
|---|---|---|
| Colour | The spoken word changes colour; the rest of the line stays white | Hormozi-style captions |
| Box | A coloured box sits behind the spoken word | Box or pill highlight |
| Fill | Words change colour as they are spoken and stay changed | Classic karaoke sweep |
| One word | Only the spoken word is on screen, large and centred | Single-word captions |
Colour and box read best on talking-head footage. Fill is the lyric-video look. One word is the most insistent of the four: strong for a hook, tiring over a full minute.
Three ways to add word-by-word captions
- Whisper and FFmpeg — free, local, command line. A plain highlight, full control, some setup.
- A video editor's caption template — CapCut or DaVinci Resolve Studio. One click once you are inside the editor; the terms are what differ.
- A desktop caption app — Sablate. Presets, vertical framing and burn-in in one window, on your own machine.
Method 1: Whisper and FFmpeg (free, fully local)
This needs Python and FFmpeg installed and on your PATH. Everything runs on your machine, and the Whisper model downloads once, on first use.
- Install Whisper:
pip install -U openai-whisper
- Transcribe with word timestamps and highlighting switched on, three words per cue:
whisper clip.mp4 --model turbo --word_timestamps True --highlight_words True --max_words_per_line 3 --output_format srt
That writes clip.srt, in which each word gets its own cue and the spoken word is wrapped in <u> tags.
- Optionally, turn the underline into a colour. FFmpeg reads
<font color>tags in SRT files, so a find-and-replace is enough (macOS and Linux, or Git Bash on Windows):
sed -e 's/<u>/<font color="#FFE900">/g' -e 's#</u>#</font>#g' clip.srt > karaoke.srt
- Burn it in:
ffmpeg -i clip.mp4 -vf "subtitles=karaoke.srt:force_style='FontName=Arial,FontSize=16,Outline=2,MarginV=60'" -c:a copy clip-captioned.mp4
One detail explains those numbers. When FFmpeg burns an SRT, it first lays the text out on a virtual 384×288 canvas, so FontSize and MarginV are measured against those 288 units rather than your video's pixels. MarginV=60 lifts the line roughly a fifth of the way up the frame, clear of the Shorts and Reels interface. If FFmpeg says it cannot find the subtitle file on Windows, the path needs escaping — the burn-in guide covers the fix.
The honest limits: you get an underline or a flat colour, with no box and no pop, and the highlight blinks off in pauses. Editing is the awkward part. Every word is its own cue carrying the whole line, so a misspelled name in a three-word line is misspelled in at least three cues. Correct the transcript before you generate the SRT, not after.
Method 2: A video editor's caption template
If you already edit in one of these, the word highlight is a template away.
CapCut. Generate auto captions, then apply a caption template with a word highlight, on desktop or on a phone. Two conditions come with it. The audio is processed on CapCut's side: its privacy policy says it may collect content "through pre-uploading at the time of creation, import, or upload", for example to generate captions. And CapCut's help centre names auto captions as a feature that can include Pro-only functionality; a project that uses Pro features needs a subscription to export.
DaVinci Resolve Studio. Timeline → AI Tools → Create Subtitles from Audio transcribes on your own machine. Since Resolve 20, the Animated titles include a Word Highlight template: drag it onto the subtitle track header, then set the colour, font and speed in the Inspector. Automatic subtitles are a Studio feature — $295 one-time — and transcription covers 14 languages out of the box, with more through an extra download. If you already own Studio, this is the cleanest route on this page.
Method 3: A desktop caption app
A desktop caption app does the whole job in one window on your own machine: transcribe, correct, pick a style, frame for vertical, render.
In Sablate, six word-highlight presets — Hormozi, Beast, Reels, One Word, Karaoke and Pop — cover the four looks above, and the Word-by-word highlight switch adds the effect to any other style, with Color, Box, Fill or Word as the mode. The caption fonts behind the presets (Montserrat ExtraBold, Anton, Bangers, Bebas Neue and Archivo Black) ship inside the app and inside its renderer, so the burn-in draws the same face the preview shows. A missing font is one of the usual reasons exported captions look different from the editor.
The timing follows the rules from the top of this page. Each word stays lit until the next one begins, so the highlight never drops in a pause. Fixing a typo keeps the recorded timings as long as the word count stays the same; add or remove a word and only that line falls back to estimated timing. The burn-in is drawn by libass, the subtitle renderer media players use, rather than stamped on as a flat overlay, and it renders at 9:16, 1:1 or 16:9 by cropping or by padding with a blurred copy of the source.
Where it goes further than the other methods is languages. Each translation layer carries its own style and its own script font, so the same Short can go out with karaoke captions in English, Spanish and Korean from one transcription, as the multilingual workflow describes. The timing on a translated layer is estimated rather than measured, because nobody spoke those words, and the app says so next to the switch.
Word-by-word styling is part of Pro, a single $29 payment. The free plan previews every preset and burns in the other styles with a small corner mark.
Which method should you use?
| Whisper and FFmpeg | CapCut | DaVinci Resolve Studio | Sablate | |
|---|---|---|---|---|
| Price | Free | Free tier, Pro subscription | $295 one-time | $29 one-time |
| Audio leaves your machine | No | Yes, for auto captions | No | No |
| Highlight looks | Underline or colour | Template library | Word Highlight template | Colour, box, fill, one word |
| Full video editor | No | Yes | Yes | No |
| Runs on a phone | No | Yes | No | No |
| Vertical 9:16 output | With an extra filter | Yes | Yes | Yes, crop or blur |
Two of those rows go against Sablate: it is not a video editor, and it does not run on a phone. If you cut, add music and publish from your phone, a desktop caption app is an extra step rather than a shortcut.
"One clip, and I'm comfortable in a terminal." Whisper and FFmpeg. Free, local, and good enough for a colour-change highlight.
"I already edit in CapCut and post every day." Stay there. The templates are the point, and if you pay for Pro anyway, the cost argument is gone.
"I edit in DaVinci Resolve Studio." Use Word Highlight. You own the licence and the transcription stays local.
"The footage cannot leave my machine." Anything except CapCut. Offline subtitle generators compares the local options in more depth.
"The same Short goes out in three languages." Sablate: one transcription and a karaoke style per language, with estimated timing on the translations.
"It's a music video, and I want syllable-level timing." Aegisub, which is free. Its karaoke mode times ASS \k tags syllable by syllable, by hand — slow, and still the best result for lyrics. The ASS format is the only common subtitle format that stores karaoke timing at all.
How to keep word-by-word captions in sync
The highlight is only as good as the word timings underneath it, and a few habits keep them intact:
- Correct before you style. Timings belong to the words that were recognised. A same-length fix, "there" for "their", keeps them; rewriting a line throws them away.
- Give Whisper the names in advance. Whisper accepts a short prompt of names and terms that biases the transcription: the
--initial_promptoption, or Defaults → Custom vocabulary in Sablate. A name spelled right the first time never needs a correction that disturbs the timing. It is a bias, not a guarantee, so still read the names. - Split long cues rather than shrinking the font. On vertical video, one line of three to five words can be read while the highlight moves; two full lines cannot.
- Replay after retiming. Dragging a cue's edge usually leaves the word timings inside it where they were, so the words nearest the new edge get squeezed. Nudge, then watch the cue once.
- Check the render, not only the preview. Watch the first cue, the last cue and the fastest passage before you upload.
Word-by-word captions are a timing problem wearing a style. Get clean word timestamps, keep your corrections the same length, and choose the look after that. If you want the presets, vertical framing and several languages in one place, download Sablate for Windows or Mac and run one Short through it.