Skip to content
All guides

How to measure cut rhythm

A 2026-08-19 measurement of four pages on storyshort.ai found the same server-rendered body on each one: 11 words.

Last checked

A 2026-08-19 measurement of four pages on storyshort.ai found the same server-rendered body on each one: 11 words. That tells you what most tools in this category publish about rhythm, which is nothing. This page shows you how to measure cut rhythm yourself, on your own videos, without buying anything.

The short answer

Count every cut in a finished video, divide by its runtime in seconds to get cuts per second, then break that number down per section: opening, middle, close. Compare sections against each other, not against another creator's average. If your cuts land at fixed intervals, you will hear it as a metronome; move them to sentence boundaries instead. There is no universal target number, because rhythm depends on topic and pacing of speech.

The steps

1. Count your own cuts before you copy anyone's

Open one finished video and mark every cut: scene change, B-roll switch, zoom punch, caption style change that resets attention. Count them all. Divide by total runtime in seconds. Write the result down as cuts per second, and also note the raw count and the runtime so you can recheck later. Do this for three videos before drawing any conclusion, because a single video can be an outlier. Your own average is the denominator for every comparison you make afterwards; without it, any advice about rhythm is a guess borrowed from someone else's footage.

Output: one cuts-per-second figure per video, plus the raw counts behind it.

2. Measure per section, not per video

Split the same video into three sections: the first ten seconds, everything between the hook and the final beat, and the last ten seconds. Count cuts inside each section separately. In most vertical documentaries the opening carries a faster rate than the middle, and that difference is intentional rather than sloppy. When you average across the whole video, the fast opening hides the slow middle and you lose the only signal worth acting on. Record the three numbers side by side. If all three are identical, your edit is probably running on a template rather than on attention.

Output: three cuts-per-second figures for one video, one per section.

3. Cut on the sentence, not on a fixed interval

Play the video back and check whether cuts land at even intervals, for example every two seconds regardless of what is being said. Fixed-interval cutting reads as mechanical within a few seconds of watching. Instead, align each cut with a sentence boundary or a completed thought in the voiceover. To verify this, mute the audio and watch: if the visual changes feel arbitrary, they are not tied to language. Then unmute and confirm each visual change coincides with a spoken sentence ending. This single adjustment usually does more than changing the cut count itself.

Output: a pass or fail verdict on whether your cuts track sentences instead of a timer.

4. Watch where you got bored yourself

Replay your own video and note the exact second where your attention drops, where you reach for the skip button even though you made the thing. That timestamp is the most honest data point in this whole method, because it comes before rationalisation starts. Check the cut density around that second using the numbers from step 2. A bored moment often sits right after a long stretch with no cuts, but it can also sit where cuts are frequent yet say nothing new. Either way, write down the timestamp and what is happening on screen at that moment.

Output: a list of timestamps where attention drops, matched against local cut density.

5. Change one thing and re-measure

Take the weakest section identified in step 4 and change exactly one variable: either cut frequency or cut placement, never both. Re-cut that section, render, and run the full count again using the same method from steps 1 and 2. Compare the new per-section figure against the old one. If you changed two things at once, you cannot tell which one moved the number, and you have learned nothing transferable. Keep a simple log of what you changed and what the count did. Two or three cycles of this produce more usable knowledge than any generic rule about pacing.

Output: a before-and-after pair of cuts-per-second figures for one section, with the single change documented.

What this will not fix

This page gives no answer to the question "how many cuts per second should I use", because we have no measured basis for such a number and no evidence bank entry supports one. Rhythm targets vary by topic: a war-history short and a bedtime-story short carry different natural rates, and forcing one number across both would damage both. Measuring cut rhythm also does not fix weak research, flat narration, or a hook that never lands. It diagnoses pacing only. If retention problems persist after your rhythm numbers look healthy, the cause sits upstream in script or topic selection, not in the timeline.

Where to go next

Once you have per-section rhythm numbers, /guides/how-to-measure-whether-a-format-is-working shows how to tie those numbers to outcomes like completion instead of treating pace as an end in itself. If the diagnosis points back to structure rather than timing, /guides/how-to-make-a-documentary-short covers how a sourced documentary short is assembled from script outward, which is where most rhythm problems actually originate.

<!-- faq -->

Frequently asked questions

Do I need editing software with analytics to do this?

No. The whole method runs on a counter, a stopwatch, and playback. Any editor's timeline shows cut positions, and the phone stopwatch handles section boundaries. Analytics platforms report aggregate retention, which is useful later, but they cannot tell you whether a cut landed on a sentence or on a fixed interval. That distinction requires eyes and ears on the specific edit, which no dashboard provides.

How many videos should I measure before acting on the results?

Three finished videos is the minimum for a personal baseline, because a single video reflects one topic and one mood. Measure the same three sections in each. Once you have nine per-section figures, patterns become visible: openings clustering around one rate, middles around another. Acting on fewer measurements risks tuning your edit toward an outlier rather than toward your actual working range.

Does faster cutting always improve retention?

No, and claiming otherwise would require retention data we do not have. What the method above shows is that mismatched rhythm hurts: a metronomic cut pattern reads as artificial, and dead stretches read as slow. Both extremes cost attention. The goal is variation that tracks speech, not maximising the cut count. Some of the strongest documentary shorts hold a single shot far longer than any formula would allow.

Should captions count as cuts?

Count caption style changes that reset attention, such as a colour shift or a layout change, but not routine word-by-word subtitle updates. Word-level subtitles fire dozens of times per video and counting them inflates your figure until it measures nothing. The test is simple: if a viewer consciously registers the change as a new visual event, it counts. If it blends into continuous reading, it does not.

How does this apply when the visuals are archive footage?

The same way, with one addition: archive clips have their own internal motion, so a static shot of a map and a moving piece of historical film carry different energy at the same cut rate. When measuring, note the source type next to each cut. Over time you will see which clip types tolerate longer holds in your edits, and the medya künyesi attached to each rendered clip makes that mapping explicit rather than remembered.