
Add captions by starting from a word-level timed transcript instead of the auto-captions your export tool generates on the fly, clean the text before you style anything, sync it in one or two short lines at a time, and design for a muted phone screen first. Burn captions in for social clips, ship a separate SRT file for long-form YouTube.
About 85 percent of Facebook video is watched with the sound off, a figure widely cited from the platform's own advertiser research and repeated across industry coverage including Digiday's reporting on muted mobile video. Forbes separately covered Verizon Media research putting the figure at 69 percent across devices. Whatever the exact number on a given platform, the pattern holds everywhere: if a clip has no captions, most of the audience is watching a silent video with no idea what is being said.
Muted autoplay gets most of the attention, but captions do real accessibility work too. They are the difference between a video being usable or unusable for deaf and hard-of-hearing viewers, and they carry a video across a language barrier for anyone watching in a second language, who often read faster than they can parse fast or accented speech. The Web Content Accessibility Guidelines, maintained by the W3C, list synchronized captions for audio content as a baseline requirement for accessible media, not an optional extra. Treating captions as only a growth hack misses that a meaningful share of any audience needs them just to follow along at all.
Most editors caption backwards, they finish the cut, then run an auto-caption feature on the export and hope the timing lands. It is faster to start from a transcript with word-level timecodes attached to the original footage. Every word already knows exactly when it was spoken, so captions built from that transcript sync automatically instead of drifting the way generic speech-to-text on a compressed export often does, especially with overlapping speakers or background noise.
Auto-transcription reliably mangles proper nouns, brand names, and industry terms. Fix those first, in the text, before you spend any time on font or color. Decide how you are handling filler words here too, um and uh usually come out, but see removing filler words for where to draw that line so captions do not read as choppier than the speaker actually sounded.
Captions should appear as the words are spoken and disappear before the next thought starts, not sit on screen a beat behind. Keep each caption block to one or two short lines, five to seven words is a reasonable ceiling per line on a vertical video. Longer blocks force viewers to read ahead of the audio, which breaks the sync feeling even when the timing is technically correct.
Every caption block competes with the pace of the speech underneath it. A dense line that stays on screen for a fraction of a second forces a choice between reading and watching, and most viewers give up on the caption entirely rather than pause a scrolling feed to finish it. The practical fix is simple: if a line needs more than a beat or two to read comfortably, split it into two caption blocks instead of cramming it into one. This matters more on fast, energetic speech, interviews and sales pitches, where the transcript is naturally denser per second than a slower, more deliberate narration.
Watching how captions actually get built and timed inside an editor makes the workflow concrete.
Design for a small screen with the sound off and one thumb scrolling past. Use a bold sans-serif with either a stroke or a solid background box behind the text so it reads over any footage, light or dark. Avoid thin fonts and low-contrast color pairings that look fine on a desktop preview and disappear against bright footage on a phone in direct sunlight. Keep text inside the safe zone, roughly the middle 60 percent of a 9:16 frame, so a platform's own UI elements never cover a word.
TikTok, Instagram Reels, and YouTube Shorts each place their own UI in slightly different spots. TikTok stacks a caption field, username, and sound title along the bottom third and a like/comment/share column down the right edge. Instagram Reels puts a similar action column on the right but reserves less space at the very bottom. YouTube Shorts keeps controls tucked closer to the edges but still covers a strip at the bottom on first load. A caption style built for one platform and exported identically to the other two will get clipped somewhere; check the actual app, not just an export preview, before publishing to each one.
The right captioning tool depends mostly on whether captions are the whole job or one step inside a bigger edit. A dedicated captioning tool wins on speed and style variety for a quick social pass; an editor that keeps captions inside the same text-based editing transcript as the cut wins when the same word-level timing needs to drive both the edit and the caption without a second pass.
| Tool | Best for | Starting price | Notes |
|---|---|---|---|
| CapCut | Mobile captioning with templates | Free; Standard $10/mo | Auto-captions built in, huge style preset library |
| Submagic | Animated, high-retention caption styles | Starter $19/mo ($12/mo billed annually) | Word-by-word animation, hook titles, B-roll suggestions |
| Kapwing | Browser-based team captioning | Free; Pro $24/mo ($16/mo billed annually) | Auto-transcribe plus manual timeline correction |
| Descript | Captioning as part of a full edit | Free; Hobbyist $24/mo ($16/mo billed annually) | Caption styling lives inside the same transcript editor as the cut |
| ScriptCut | Word-level timed transcript and SRT export | Free; Starter $19/mo ($15/mo billed annually) | Captions inherit the same timecodes used for the edit, no re-syncing |
Several captioning tools, Submagic among them, now offer one-click caption translation into other languages. That is worth using deliberately rather than as an afterthought: a clip that performs well in English is a reasonable candidate for a Spanish or Portuguese caption pass if a meaningful chunk of the audience data shows viewers from those regions. Translated captions still need the same word-level timing discipline as the original, a machine-translated line that runs long tends to blow past the sync window a short caption block was designed for, since translated sentences rarely map word-for-word to the source length.
TikTok, Instagram, and YouTube all now offer a built-in auto-caption button, and it is tempting to just tap it and move on. The problem is consistency and control: platform auto-captions vary in accuracy from clip to clip, cannot be styled beyond a couple of presets, and have to be regenerated separately on every platform since none of them share caption data with each other. A clip captioned from a word-level transcript gets styled once, with a known-accurate source, and the same caption data can be reused across every platform export instead of trusting four different auto-caption engines to each get the same line right.
Burned-in captions are permanent text baked into the video file, the right choice for TikTok, Reels, and Shorts, where the caption is part of the design and the platform cannot be trusted to render an uploaded caption file consistently. For long-form YouTube, upload a separate SRT file instead. It stays editable, lets viewers toggle it off, and YouTube uses it to generate searchable transcripts and translated captions automatically.
Pull a 45-second clip from a transcript that already has word-level timing. Export the matching caption data directly, no re-transcribing the exported clip from scratch. Scan for any misheard names in that short window, fix them. Apply a bold two-line style with a background box, confirm nothing sits in the bottom safe zone where a platform's UI usually lives, then export burned-in for the vertical post and keep the SRT on hand in case the same clip goes up on YouTube later.
Good captions start with accurate word-level timing, not a caption app run on a finished export. Get the transcript and timing right first, clean the text, then style for a muted phone screen, and pick burned-in or SRT based on where the video is actually going to live.
Burn captions in for TikTok, Reels, and Shorts, where the caption is part of the on-screen design. Use a separate SRT file for long-form YouTube, where it stays editable and helps with search.
Most social video is watched with the sound off, and captions also make video usable for deaf and hard-of-hearing viewers and anyone watching in a second language.
Start from a transcript that already has word-level timecodes attached to the footage, so caption timing is inherited automatically instead of re-synced by hand.
Roughly the middle 60 percent of a 9:16 frame on any of the three, but the exact UI placement differs slightly per app, so check the caption inside the actual platform before publishing.
No. Clean obvious filler like um and uh out of the text before styling, or the captions make a speaker look choppier on screen than they sounded out loud.
Burned-in captions are permanently part of the video image and cannot be turned off. Closed captions, delivered as a file like SRT, can be toggled on or off by the viewer and are editable after the fact.
It is worth doing selectively for clips that already perform well and show viewership from non-English-speaking regions, since translated captions can meaningfully widen the audience for that specific clip.