How to measure whether a format is working
On 19 August 2026, four pages of a topic-to-video tool each returned an 11-word server-rendered body when fetched without JavaScript.
Last checked
On 19 August 2026, four pages of a topic-to-video tool each returned an 11-word server-rendered body when fetched without JavaScript. That single number told more about how that site works than any claim on its homepage. This page gives you the same kind of discipline for your own channel: a fixed method for deciding whether a video format is working, using numbers you set before you publish, not feelings you have after.
The short answer
A format is working when it clears a threshold you wrote down before the first video went out, judged on the median of a fixed batch of videos rather than the best one. Define the format in one sentence, commit to a batch size, track median retention and the exact second where viewers leave, then kill or keep the format against your pre-set number. If you cannot state the kill threshold in advance, you are not measuring a format, you are rationalizing one after the fact.
The steps
1. Define the format in one sentence
Write down what the format is before you measure anything: "60-second vertical documentaries about aviation disasters, narrated, with archive footage and burned-in captions." One sentence, no adjectives, no ambition statements. If the definition takes two sentences or contains the word "and" twice, you have two formats, and neither will produce readable data. A vague definition produces a vague result: when the numbers come back flat, you will not know which part of the format failed, because you never specified what the format contained. The sentence also fixes the comparison set. Two videos only count as the same format if they match the sentence.
The output of this step is one written sentence that names the format, its length, its subject area and its visual method.
2. Give it a fixed number of videos, decided in advance
Decide the batch size before publishing the first video, and write it next to the definition. Five is the practical minimum for a vertical format; ten gives a cleaner median but costs more production time. The point of fixing the number is that it removes the two ways this test usually gets corrupted. First, stopping early: one video underperforms, the format dies on day three, and you never learn whether the third video was the problem or the topic was. Second, extending forever: the format limps along for months because quitting feels like waste. A fixed batch turns an open-ended hope into a bounded experiment with an end date. When ViewMade renders a batch of documentary shorts from one topic, the render count per plan is fixed in advance, 50 per month on the Starter plan as of 22 August 2026, and that cap is what makes per-video cost calculable at $0.58. The same logic applies to your batch.
The output of this step is a written batch size and a date by which every video in the batch will be published.
3. Track the median, not the best one
After the batch is published, sort the videos by views and take the middle value, not the top one. The best video measures luck: one video getting picked up by the algorithm tells you nothing about the format, because you cannot reproduce whatever caused the spike. The median measures the format, because it describes what a typical execution of that format does. Record three medians for the batch: median views, median average view duration, and median likes per thousand views. Write them next to the definition and the batch size. If the spread between your best and worst video is enormous while the median sits low, the format may be fine and your topics or hooks are inconsistent, which is a different problem with a different fix.
The output of this step is three median values attached to the format definition.
4. Look at where viewers leave, in the same second across videos
Open the retention graph for every video in the batch and find the timestamp where the steepest drop happens. Then compare across videos. If viewers leave at wildly different points, the exits are topic-driven and the format itself is neutral. If they leave at roughly the same second in most of the batch, you have found a structural defect: a segment length, a caption style, a music cue, a pacing pattern that the format repeats and viewers reject. This is the step most channels skip, and it is the one that actually improves the format, because it converts a vague verdict ("the format is not working") into a specific edit ("viewers leave during the second archive clip"). Fix that segment, keep everything else, and rerun a small batch.
The output of this step is either "exits are scattered across timestamps" or "a shared exit window exists at approximately [timestamp], present in most of the batch."
5. Kill or keep on the number you set in advance
Go back to the threshold you wrote in step 2 and apply it mechanically. If the median cleared the bar, keep the format and scale it. If it did not, kill it or change one variable and run a new batch, but do not quietly extend the old one. The reason for mechanical application is that post-hoc judgment always finds a reason to continue: the thumbnail was bad, the upload time was wrong, the algorithm was cold. Some of those excuses will even be true. But a threshold you can renegotiate after seeing the results is not a threshold, it is a preference. Write the verdict down with the date, the median values and the decision, so the next format test starts from a record instead of a memory.
The output of this step is a dated verdict: kept, killed, or revised with a named change and a new batch size.
What this will not fix
This method cannot rescue a decision made on a tiny sample. On a small channel, a five-video batch can produce medians that swing widely from pure variance, and no amount of retention analysis changes that. The honest position is that format decisions made below a few thousand views per video are statistically weak, and this page deliberately does not give you a minimum view count that makes them strong, because such a number depends on your niche, your audience size and your posting cadence, none of which a general guide can set for you. What the method does give you is consistency: the same definition, the same batch size and the same thresholds applied every time, so that over several batches your decisions get better even when no single batch is conclusive. It also will not fix a bad topic choice dressed up as a format problem, which is why step 4 separates exit timing from topic effects before you blame the format.
Where to go next
To read the raw numbers this method depends on, see /guides/how-to-read-shorts-analytics, which explains what each metric in Shorts analytics means and which ones move together. To act on the exit-window finding from step 4, see /guides/how-to-measure-cut-rhythm, which covers how cut frequency relates to where viewers stop watching.
<!-- faq -->Frequently asked questions
How many videos do I need before I can judge a format at all?
Five is the workable floor for a first verdict, ten for a confident one. Below five, the median is dominated by single-video swings and tells you almost nothing. What matters more than the exact count is that you fix it before publishing and do not extend it after seeing early results, because an adjustable batch size silently converts a test into a hope.
What if one video in the batch massively outperforms the rest?
Exclude it from the median judgment and treat it as a separate question. A large outlier usually reflects the topic, the hook or an external share rather than the format, since the format was identical across the whole batch. Note what was different about that video, because it may be worth testing as its own variant, but do not let it raise your verdict for the format.
Should I judge formats on views or on watch time?
Use both, but weight watch time more heavily for format decisions. Views measure distribution, which depends heavily on factors outside the format, while average view duration and retention shape reflect how the video itself holds attention. A format with modest views but strong median retention is often worth keeping, because distribution improves with consistency faster than structure does.
How long should I wait after the last video before applying the threshold?
Wait until each video has had comparable exposure time, typically two to four weeks for Shorts, since early impressions accumulate unevenly. Judging a video published Monday against one published three weeks earlier corrupts the median. Set the evaluation date at the same time as the batch size in step 2, so the end of the experiment is known before it starts.
Can I change the format mid-batch if the first videos clearly fail?
No. Changing mid-batch destroys the comparison, because the later videos are no longer the same format and the median stops meaning anything. If the early evidence is overwhelming, kill the batch, write the verdict, define a revised format as a new entry, and start a fresh batch with a new threshold. The record of the killed attempt is worth more than a salvaged one.