How to add captions to a short
ViewMade's pricing documentation lists 7 caption styles as of 22 August 2026, and every one of them burns the captions into the image rather than
Last checked
ViewMade's pricing documentation lists 7 caption styles as of 22 August 2026, and every one of them burns the captions into the image rather than delivering them as a separate subtitle track. This page walks through the five steps of captioning a vertical short correctly: choosing the delivery method, sourcing the text, framing the safe area, pacing the words, and verifying against the narration.
The short answer
If you have the script, caption from the script instead of running the audio through automatic recognition. Burn the captions into the video for platforms where you cannot upload a subtitle file, and keep them inside the safe area above the bottom interface strip. One or two words per beat, timed to the voiceover. Then watch the final render once at full speed with sound on and fix any word that does not match what is spoken.
The steps
1. Decide between burned-in captions and platform subtitles
Burned-in captions are rendered into the pixels of the video. Platform subtitles are a separate file, such as an SRT, that YouTube stores alongside the video and displays when the viewer enables them. They do different jobs. Burned-in captions survive reuploads, edits, reposts to other platforms, and screenshots; they always look exactly how you designed them. Platform subtitles can be toggled off, translated by the platform, indexed for search, and read on devices where the video plays muted by default.
For most shorts you want both. The burned-in layer carries the styled, word-by-word display that keeps viewers watching without sound. The subtitle file gives accessibility, search indexing, and translation for free. If your tool only produces one, produce the burned-in version first, then export an SRT from the same script so the two never diverge.
Output: a decision recorded before any styling work starts.
2. Caption from the script, not from the audio, when you have the script
Automatic speech recognition listens to audio and guesses words. It makes predictable errors: it misspells proper names, transcribes numbers inconsistently, and drops or invents small words at fast speech rates. When you already have the script, none of that guessing is necessary. Every word is known before the first frame renders.
The workflow is simple. Paste the script into the captioning step, split it into beats manually or with the tool's splitter, and time each beat to the corresponding stretch of narration. Tools that generate the video from a script can do this automatically because the timing comes from the synthesis engine, not from a transcription pass. ViewMade works this way: its word-by-word captions come from the script that also drives the voiceover, so the caption track and the audio track share one source.
Output: a caption track whose text matches the script exactly.
3. Set the safe area before you set the style
The bottom strip of a short is occupied territory. On YouTube Shorts, the title, channel name, like and comment buttons, and the description overlay all sit over the lower portion of the frame, and they cover roughly the lower third depending on device. A caption placed there gets hidden behind the interface, and no amount of styling fixes that.
Set your safe area margins first: keep captions out of the bottom third and away from the extreme left and right edges, where rounded screen corners clip content on phones. Then choose the style within that box. Pick a style with enough contrast against your footage, add a shadow or outline if the background varies, and check the position on both a phone-shaped preview and a desktop player, because the overlays differ between them.
Output: a caption block positioned where no interface element covers it.
4. Keep one or two words per beat
Word-by-word captions exist to pace reading. A beat carrying six words forces the viewer to either skim past it or fall behind the narration, and both outcomes cost retention. One or two words per beat matches the natural rhythm of spoken English at short-form speed and keeps each flash readable in under half a second.
Split longer sentences at phrase boundaries, not mid-clause. "The bridge collapsed" reads cleanly as two beats; splitting it as "the bridge" and "collapsed" loses nothing, but "the bridge col" and "lapsed" destroys comprehension. Let punctuation guide the cuts: commas and periods mark natural break points. If a number appears in the sentence, give it its own beat, since numerals take longer to read than words of the same character count.
Output: a caption rhythm that tracks the narration beat for beat.
5. Check the caption against the narration one last time, at speed
Watch the finished render once, at normal playback speed, with sound on. Do not scrub. Scrubbing hides sync drift, because each beat looks correct in isolation even when the whole track has slipped half a second behind the voice. At speed, drift shows up immediately as captions landing before or after their words.
Fix three classes of error during this pass. First, mismatched words: anything the caption says that the voice does not. Second, mistimed beats: captions that flash early or linger late. Third, style failures: low contrast against a bright background frame, or a line that collides with the platform overlay you identified in step 3. Export again after fixing, and repeat this single-speed pass once more.
Output: a verified render where caption, audio, and frame agree.
What this will not fix
Captioning from a script removes transcription errors, but it does not remove every class of problem. Automatic recognition still fails in specific ways when a script is not available: proper names get spelled phonetically, numbers get written out or digitized inconsistently, and homophones get swapped. This page does not publish an error rate for those failures, because we have not measured one, and a percentage we did not measure would be decoration.
Script-based captioning also does not fix bad source material. A script with a wrong date produces a caption with a wrong date, perfectly synced. A mumbled or overlapping section of narration still produces awkward beat splits even when the text is correct. And no captioning method fixes a video whose footage contradicts its own narration; that is a research failure, not a captioning failure. Check facts at the script stage, before any of these five steps begin.
Where to go next
If you want the vocabulary for step 1 in detail, our glossary entry on burned-in captions explains how rendering text into pixels differs from sidecar subtitle files and when each survives a repost. For step 2 to produce good results, the script itself has to be written for synthetic narration, which is a separate skill covered in how to write for a synthetic voice; a script built for a human reader often breaks the beat pacing described in step 4.
<!-- faq -->Frequently asked questions
Should I use the same caption style on YouTube Shorts and TikTok?
Not necessarily. TikTok's interface occupies different parts of the frame than YouTube Shorts, so a position that clears YouTube's overlay may collide with TikTok's buttons. The safest approach is to keep captions in the middle band of the frame, above the bottom third, which clears both interfaces. Style can stay consistent across platforms; position should be checked per platform against that platform's overlay layout.
Do burned-in captions hurt reach?
No measurable penalty exists for having text in the frame. What matters is whether the captions help or hinder comprehension. Well-paced captions raise watch time on muted playback, which is common for shorts discovered in feeds. Poorly paced captions, with too many words per beat, push viewers to skip. The format itself is neutral; execution decides the outcome.
Can I edit captions after publishing?
If the captions are burned in, no. The text is part of the image, so correcting a typo means re-rendering and re-uploading the video. This is why the verification pass in step 5 matters: catching errors after publish costs a full re-render. Platform subtitles, by contrast, can be edited or replaced in YouTube Studio without touching the video file, which is an argument for uploading an SRT alongside the burned-in version.
How many words per minute is right for short captions?
There is no fixed number worth quoting here, because the constraint is per-beat readability, not aggregate speed. Two words flashing every half second and six words held for a second and a half can carry the same words-per-minute while feeling completely different. Aim for one or two words per beat and let the total rate follow from your narration speed. If a beat feels unreadable on the single-speed review, split it.
What about translating captions into other languages?
A burned-in caption cannot be translated without a new render, because the text lives in the pixels. An SRT file can be translated and uploaded as an additional subtitle track at no cost to the original video. If you plan to serve multiple languages, keep the SRT as the translation surface and leave the burned-in layer in the primary language. ViewMade generates voiceovers in 5 languages as of 22 August 2026, but the caption language follows the narration language per render.