Guides / Shorts
Where captions go on a Short: measure it
Tested on: Carousel Studios & Lab caption style, measured with ffmpeg 7
TL;DR
- Phrase captions and one-word cards are different objects. They want different positions.
- Do not compute where a caption lands. Burn it over black and read the box with ffmpeg, outline painted white so it shows.
- The famous caption statistics say nothing about placement. I leave them out.
Step 1: Name the kind of caption you have
We run two kinds. Phrase captions carry a sentence at a time. The Section 508 captioning page says to center captions in the lower third, except where they block important text, like signs or person identifiers. On a 1080x1920 frame that third starts at row 1280, but the app draws over the bottom of the frame (Step 4). Our August 2026 notes put phrase captions over a talking head just below the chin, roughly rows 1000 to 1250. That is a reasoned pick, not a tested result.
One-word cards are another object: one large word, swapped about every third of a second at speaking pace, so the eye parks at one point. On Carousel Studios & Lab we put them at dead center, row 960, by choice. On shots with a face mid-frame the word covers it, and I accepted that.
So the rule in our August 2026 notes, never dead center, suits a sentence over a talking head, not one word at a fixed point. I found no published test of position against watch time then, and did not repeat the search. Center is a taste call, not a result.
Step 2: Burn the style over black
Burn your style onto a black frame and ask ffmpeg where the bright pixels are. Black removes the footage from the question, and anything black from the answer.
Save this as card.ass: our one-word card style as it runs on the channel. Arial Black, size 120, bold, uppercase, an 8 px black outline, bottom-center, the bottom margin the only dial. Each card pops from 116% to 100% in 70 ms.
[Script Info]
ScriptType: v4.00+
PlayResX: 1080
PlayResY: 1920
WrapStyle: 0
ScaledBorderAndShadow: yes
[V4+ Styles]
Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding
Style: Card,Arial Black,120,&H00FFFFFF,&H00FFFFFF,&H00000000,&H80000000,-1,0,0,0,100,100,0,0,1,8,0,2,60,60,902,1
[Events]
Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text
Dialogue: 0,0:00:00.00,0:00:02.00,Card,,0,0,0,,{\fscx116\fscy116\t(0,70,\fscx100\fscy100)}HOUSTON
Put the font in a fonts folder beside it, then run this.
ffmpeg -f lavfi -i color=c=black:s=1080x1920:r=30:d=1 -vf "subtitles=card.ass:fontsdir=fonts,trim=start_frame=15:end_frame=16,bbox=min_val=40" -f null -
ffmpeg 7.0.1 printed x1:307 x2:772 y1:929 y2:992 on Oct 2, 2026: the first and last columns and rows holding a pixel brighter than min_val. With no caption on screen it prints no coordinates.
Keep min_val at or above the black level. Our frame is luma 16, which is also ffmpeg’s default, so min_val=1 boxed the whole frame.
fontsdiradds your font file but does not lock it in. If the style’s font name does not match, libass quietly picks another font: a one-letter typo moved our box from rows 929 to 992 to rows 917 to 996. Check thefontselectline in the log.trimkeeps one frame, half a second in. Frames 3 to 6 gave identical boxes, frame 2 was off by up to 2 pixels, and frame 0 caught the pop at its starting 116% scale.
Step 3: Read the rectangle, not the font size
Same style, Oct 2, 2026. Only the cue, margin, size and frame change between rows. The last column paints the stroke white.
| Case | Bottom margin | Fill rows | Rows with stroke showing |
|---|---|---|---|
| HOUSTON, size 120, center card | 902 | 929 to 992 | 921 to 1000 |
| Same card, earlier position | 700 | 1131 to 1194 | 1123 to 1202 |
| Lowercase “piggy jumps” | 902 | 930 to 1009 | 922 to 1017 |
| Same card, frame 0 (pop starts at 116%) | 902 | 915 to 988 | 907 to 996 |
| 48-character phrase, size 64 | 902 | 907 to 1013 | 899 to 1021 |
The margin is measured to the line box, not the ink. At margin 902 the line box ends at row 1018, but the fill ends at row 992, because the box reserves room for descenders.
The font size is not the glyph size. A size of 120 drew capitals only 64 pixels tall, because libass scales a font so its Windows ascent plus descent equals the size.
Long cues wrap and climb. The phrase “Houston, we have a problem with the long caption” broke into two lines. The bottom stayed near row 1013 and the top rose to row 907. Measure the longest cue you will ship.
Step 4: Check the app overlay yourself
Captions leave the bottom because the app draws over it. How far up is the part people quote with false confidence. What works: play your own upload on your phone, screenshot it, and lay it over your export. It beats any table.
YouTube’s help page on enhancing Shorts says the editor shows white lines for a “non-safe” area near the edge and icons where viewer elements might sit. It gives no pixel or percentage figures.
Google Ads Help does publish one image, for vertical video ads, which can also run on Shorts. On a 1080x1920 frame it keeps important elements 288 pixels from the top, 672 from the bottom, 48 from the left and 192 from the right. That is an ad spec, not the organic Shorts player, so I publish no Shorts figure.
Step 5: Leave the famous statistics out
Two numbers appear in most caption pitches.
“85% of video is watched without sound.” I traced it to a Digiday piece from May 17, 2016. The figures came from publishers’ own data, not Facebook, and one publisher gave a range of 50 to 80 percent. It describes one platform, in 2016.
“Captions raise completion by 80%.” The source is a Verizon Media and Publicis Media survey from April 2019 of 5,616 US adults. Eighty percent said they were more likely to watch an entire video with captions. That is a stated preference, not a completion rate. I read it through 3Play Media’s summary, not the original release.
I repeat neither as fact. If a caption position comes with a percentage, ask for the test.
What did not work: moving it on taste, untested
Run it yourself
- Save
card.asstwice, with your shortest cue and your longest, and run the command on each. - Check that the
fontselectline in the log names your font. - Run it again with
OutlineColourandBackColourset to&H00FFFFFF. - Read
y1andy2and work out the optical center. - Change the margin only, re-measure, and date every number.
Sources
All links re-checked Oct 2, 2026.
- YouTube Help, Enhance your Shorts, Android version. Supports: white lines for a non-safe area, icons, no figures.
- Google Ads Help, Use square and vertical video, with its safe-zone image. Supports: vertical ads can show on Shorts, and the pixel figures.
- Section508.gov, Captions and transcripts. Supports: center of the lower third except where that blocks important text, and white text on a black translucent background by default.
- Digiday, May 17, 2016. Supports: the figures came from publishers’ own data, with one range of 50 to 80 percent.
- 3Play Media, summary of the Verizon Media and Publicis Media study. Supports: April 2019, 5,616 US adults, self-reported.
- FFmpeg 7.0.1 filters doc, FFmpeg 7.0.1 muxers doc, libass 0.17.2 source, FreeType sizing docs. Supports, in that order: bbox default of 16 and fontsdir,
-update 1(image2 muxer), font size scaled to the Windows ascent plus descent. - Our own run: every command and measurement above, ffmpeg 7.0.1 on Windows, with the style our channel used on Oct 2, 2026. Other filter behavior is what I observed.
- Not re-checked: our August 2026 search for position tests. Instagram’s caption help page returned no text, so this guide says nothing about TikTok or Instagram captions.
Start with a channel teardown.
Send me a channel. I write up what I find, with specific fixes, caption placement included.
Request a teardown