Editorial illustration of a vertical phone frame with caption bars
Podcast video

How to Add Captions to Your Video Clips

The ScriptCut Team
/
June 9, 2026
/
8 min read

Add captions by starting from a word-level timed transcript instead of the auto-captions your export tool generates on the fly, clean the text before you style anything, sync it in one or two short lines at a time, and design for a muted phone screen first. Burn captions in for social clips, ship a separate SRT file for long-form YouTube.

Why Captions

Most Viewers Never Turn the Sound On

About 85 percent of Facebook video is watched with the sound off, a figure widely cited from the platform's own advertiser research and repeated across industry coverage including Digiday's reporting on muted mobile video. Forbes separately covered Verizon Media research putting the figure at 69 percent across devices. Whatever the exact number on a given platform, the pattern holds everywhere: if a clip has no captions, most of the audience is watching a silent video with no idea what is being said.

Accessibility

Captions Are Not Just a Social Media Trick

Muted autoplay gets most of the attention, but captions do real accessibility work too. They are the difference between a video being usable or unusable for deaf and hard-of-hearing viewers, and they carry a video across a language barrier for anyone watching in a second language, who often read faster than they can parse fast or accented speech. The Web Content Accessibility Guidelines, maintained by the W3C, list synchronized captions for audio content as a baseline requirement for accessible media, not an optional extra. Treating captions as only a growth hack misses that a meaningful share of any audience needs them just to follow along at all.

Transcript First

Start From Word-Level Timing, Not Auto-Captions on Export

Most editors caption backwards, they finish the cut, then run an auto-caption feature on the export and hope the timing lands. It is faster to start from a transcript with word-level timecodes attached to the original footage. Every word already knows exactly when it was spoken, so captions built from that transcript sync automatically instead of drifting the way generic speech-to-text on a compressed export often does, especially with overlapping speakers or background noise.

Cleanup

Fix Names and Terms Before You Touch Styling

Auto-transcription reliably mangles proper nouns, brand names, and industry terms. Fix those first, in the text, before you spend any time on font or color. Decide how you are handling filler words here too, um and uh usually come out, but see removing filler words for where to draw that line so captions do not read as choppier than the speaker actually sounded.

Sync

One or Two Short Lines, Timed to the Word

Captions should appear as the words are spoken and disappear before the next thought starts, not sit on screen a beat behind. Keep each caption block to one or two short lines, five to seven words is a reasonable ceiling per line on a vertical video. Longer blocks force viewers to read ahead of the audio, which breaks the sync feeling even when the timing is technically correct.

Reading Speed

A Caption Someone Cannot Finish Reading Is Worse Than No Caption

Every caption block competes with the pace of the speech underneath it. A dense line that stays on screen for a fraction of a second forces a choice between reading and watching, and most viewers give up on the caption entirely rather than pause a scrolling feed to finish it. The practical fix is simple: if a line needs more than a beat or two to read comfortably, split it into two caption blocks instead of cramming it into one. This matters more on fast, energetic speech, interviews and sales pitches, where the transcript is naturally denser per second than a slower, more deliberate narration.

Watching how captions actually get built and timed inside an editor makes the workflow concrete.

Design

Bold, High-Contrast, Inside the Safe Zone

Design for a small screen with the sound off and one thumb scrolling past. Use a bold sans-serif with either a stroke or a solid background box behind the text so it reads over any footage, light or dark. Avoid thin fonts and low-contrast color pairings that look fine on a desktop preview and disappear against bright footage on a phone in direct sunlight. Keep text inside the safe zone, roughly the middle 60 percent of a 9:16 frame, so a platform's own UI elements never cover a word.

Per Platform

The Safe Zone Is Not Identical Everywhere

TikTok, Instagram Reels, and YouTube Shorts each place their own UI in slightly different spots. TikTok stacks a caption field, username, and sound title along the bottom third and a like/comment/share column down the right edge. Instagram Reels puts a similar action column on the right but reserves less space at the very bottom. YouTube Shorts keeps controls tucked closer to the edges but still covers a strip at the bottom on first load. A caption style built for one platform and exported identically to the other two will get clipped somewhere; check the actual app, not just an export preview, before publishing to each one.

Tools

Where Each Captioning Tool Actually Helps (as of August 2026)

The right captioning tool depends mostly on whether captions are the whole job or one step inside a bigger edit. A dedicated captioning tool wins on speed and style variety for a quick social pass; an editor that keeps captions inside the same text-based editing transcript as the cut wins when the same word-level timing needs to drive both the edit and the caption without a second pass.

ToolBest forStarting priceNotes
CapCutMobile captioning with templatesFree; Standard $10/moAuto-captions built in, huge style preset library
SubmagicAnimated, high-retention caption stylesStarter $19/mo ($12/mo billed annually)Word-by-word animation, hook titles, B-roll suggestions
KapwingBrowser-based team captioningFree; Pro $24/mo ($16/mo billed annually)Auto-transcribe plus manual timeline correction
DescriptCaptioning as part of a full editFree; Hobbyist $24/mo ($16/mo billed annually)Caption styling lives inside the same transcript editor as the cut
ScriptCutWord-level timed transcript and SRT exportFree; Starter $19/mo ($15/mo billed annually)Captions inherit the same timecodes used for the edit, no re-syncing

Multi-Language

Translated Captions Widen the Audience Further

Several captioning tools, Submagic among them, now offer one-click caption translation into other languages. That is worth using deliberately rather than as an afterthought: a clip that performs well in English is a reasonable candidate for a Spanish or Portuguese caption pass if a meaningful chunk of the audience data shows viewers from those regions. Translated captions still need the same word-level timing discipline as the original, a machine-translated line that runs long tends to blow past the sync window a short caption block was designed for, since translated sentences rarely map word-for-word to the source length.

Native Auto-Captions

Why the Platform's Own Caption Button Is Not Enough

TikTok, Instagram, and YouTube all now offer a built-in auto-caption button, and it is tempting to just tap it and move on. The problem is consistency and control: platform auto-captions vary in accuracy from clip to clip, cannot be styled beyond a couple of presets, and have to be regenerated separately on every platform since none of them share caption data with each other. A clip captioned from a word-level transcript gets styled once, with a known-accurate source, and the same caption data can be reused across every platform export instead of trusting four different auto-caption engines to each get the same line right.

Format

Burned-In for Social, SRT for Everything Else

Burned-in captions are permanent text baked into the video file, the right choice for TikTok, Reels, and Shorts, where the caption is part of the design and the platform cannot be trusted to render an uploaded caption file consistently. For long-form YouTube, upload a separate SRT file instead. It stays editable, lets viewers toggle it off, and YouTube uses it to generate searchable transcripts and translated captions automatically.

Worked Example

A 45-Second Clip, Captioned in Under Five Minutes

Pull a 45-second clip from a transcript that already has word-level timing. Export the matching caption data directly, no re-transcribing the exported clip from scratch. Scan for any misheard names in that short window, fix them. Apply a bold two-line style with a background box, confirm nothing sits in the bottom safe zone where a platform's UI usually lives, then export burned-in for the vertical post and keep the SRT on hand in case the same clip goes up on YouTube later.

Mistakes

Where Caption Workflows Go Wrong

  • Auto-captioning the final export instead of the source. You lose word-level accuracy and often have to manually retime lines that drifted.
  • Captioning every filler sound. A caption for every um and uh makes a speaker look worse on screen than they sounded out loud; clean the text first.
  • Cramming too much text into one line. A dense caption block forces viewers to choose between reading and watching, and most give up on the caption entirely.
  • Text sitting under platform UI. Always preview inside the actual app, not just the export, before publishing.
  • Using one caption style for every platform. The safe zone shifts enough between TikTok, Reels, and Shorts that a single export can get clipped on at least one of them.
  • Skipping the SRT for long-form. Burned-in only locks out viewers who want captions off and hurts searchability on YouTube.

The Takeaway

Captions Are a Timing Problem Before They Are a Design Problem

Good captions start with accurate word-level timing, not a caption app run on a finished export. Get the transcript and timing right first, clean the text, then style for a muted phone screen, and pick burned-in or SRT based on where the video is actually going to live.

Sources

frequently asked questions

How to Add Captions to Your Video Clips FAQs

Should I caption a clip with burned-in text or an SRT file?

Burn captions in for TikTok, Reels, and Shorts, where the caption is part of the on-screen design. Use a separate SRT file for long-form YouTube, where it stays editable and helps with search.

Why do clips need captions if they already have audio?

Most social video is watched with the sound off, and captions also make video usable for deaf and hard-of-hearing viewers and anyone watching in a second language.

How do I get accurate caption timing without retyping the transcript?

Start from a transcript that already has word-level timecodes attached to the footage, so caption timing is inherited automatically instead of re-synced by hand.

What is the safe zone for captions on TikTok, Reels, or Shorts?

Roughly the middle 60 percent of a 9:16 frame on any of the three, but the exact UI placement differs slightly per app, so check the caption inside the actual platform before publishing.

Should every filler word be captioned?

No. Clean obvious filler like um and uh out of the text before styling, or the captions make a speaker look choppier on screen than they sounded out loud.

What is the difference between burned-in captions and closed captions?

Burned-in captions are permanently part of the video image and cannot be turned off. Closed captions, delivered as a file like SRT, can be toggled on or off by the viewer and are editable after the fact.

Should I translate captions into other languages?

It is worth doing selectively for clips that already perform well and show viewership from non-English-speaking regions, since translated captions can meaningfully widen the audience for that specific clip.

Get the ScriptCut newsletter
Editing tips and product news. No spam, unsubscribe anytime.
Stop scrubbing. Start selecting.