Shorts · 8 min read
Why your AI Shorts look like stock footage, and the four things that fix it
The four measurable reasons an automated vertical video reads as generic, checked against a frame-accurate reference set of ten high-retention Shorts and one full production manifest.
Published
"AI slop" is a useful insult and a useless diagnosis. It describes the feeling of watching something assembled rather than made, but it does not tell you which decision produced the feeling, so it cannot tell you what to change.
There are four things doing most of the work, and all four are measurable. A shot that has nothing developing inside it. A cut rate following a stopwatch instead of a sentence. Every clip arriving from the same source. And the two small production faults — black frames and silence — that a human editor removes without thinking and an automated pipeline emits by default.
The numbers below come from a frame-accurate reference set of ten high-retention vertical documentaries, seven English and three Turkish, measured shot by shot. Ten videos is a reference band, not a law of physics. It is enough to kill some confident advice, and it is not enough to replace it with a formula.
1. The shot has nothing developing inside it
Median shot length across the reference set ran from 2.03 to 6.00 seconds.
The six-second figure is the interesting one, because the received wisdom says a static vertical shot dies after about two. It does not. It dies when nothing is developing inside it.
Those are different failures. A six-second shot of a crowd moving toward the camera is not the same object as a six-second shot of a still photograph, even though a stopwatch cannot tell them apart. The first one is still delivering information at second five. The second one delivered everything it had in the first quarter-second and spent the rest asking the viewer to wait.
This is why the "cut faster" advice both works and misleads. Cutting faster does fix the symptom — you are no longer sitting on a dead frame — but it fixes it by hiding it. The underlying problem is that the clip was chosen for its subject and not for its movement, and a library of clips chosen that way will read as stock no matter how quickly you cut between them.
What to look for in a candidate clip: is anything in it changing between its first frame and its last? Camera movement counts. Subject movement counts. A slow push on a still image is a substitute for it, and viewers have learned to recognise the substitute.
2. The cut rate is following a stopwatch, not a sentence
Cuts per minute across the reference set ranged from 8.2 to 25.1.
Look at how wide that is. The bottom of the range is a cut roughly every seven seconds; the top is one every two and a half. Both were in the set, and both held attention. That single fact kills the most repeated piece of Shorts advice in one line: there is no correct cutting speed. There is a correct cutting speed for the thing being said.
The pattern underneath is easier to state than the number. The cut rate follows the density of the narration, not the clock. A sentence that names three places in eight seconds wants three shots. A sentence that lands one hard fact wants one shot that stays there while the fact registers.
An automated pipeline that cuts on a fixed interval gets this wrong in a specific, recognisable way: the visuals stop agreeing with the sentence. A new image arrives in the middle of a clause and the viewer's attention is pulled off the words and onto the change. Do it forty times in a minute and the result is exhausting rather than fast, which is the actual complaint behind most "this feels like AI" reactions.
3. Every clip came from the same place
One real render, measured: job b2e4b3eb, a 53.47-second piece about Nadia Comăneci at the 1976 Olympics. Fourteen cuts, 15.7 cuts per minute, median shot 3.2 seconds — and footage drawn from five separate archives.
Source count is the variable nobody talks about and the one most visible to a viewer who cannot articulate why something looks cheap. Material from a single library shares a grade, a grain, a framing convention and often a decade. Cut six clips from one stock provider together and the result has a uniform surface, which reads as a catalogue rather than as a record of something that happened.
This is also the reason so many faceless channels look like each other rather than like themselves. They are drawing from the same handful of libraries, so their videos inherit the same visual accent.
4. Black frames and silence
Three things were true in all ten reference videos, with no exceptions.
No black frames. Zero. Not between scenes, not as a fade, nothing. A single dark frame in a vertical feed reads as the video ending, and the thumb is already moving before the next shot arrives.
No silence. Also zero. Every gap in the narration is carrying something, even if that something is room tone.
A narrow loudness band, measured at -14.8 to -11.5 dBFS integrated. Not because a number is magic, but because a video that is quieter than the feed around it gets skipped before its first sentence finishes.
These are the faults that separate an automated pipeline from an edited video most reliably, because they are invisible in a storyboard and obvious in playback. A human editor removes them without noticing they made a decision. A renderer emits them unless someone decided it should not.
Where the footage comes from
Most of the four faults above trace back to one upstream choice, so it is worth naming the options honestly rather than pretending one wins everywhere.
| Footage source | What it is | Repeats across channels | Can be credited | Movement inside the frame |
|---|---|---|---|---|
| Generated (video model) | Invented from a prompt | Low — every clip is unique | No original to credit | High, but often physically wrong |
| Stock library | Licensed catalogue clips | High — same libraries, same clips | Licence, not authorship | Mixed; much of it is B-roll shot to be generic |
| Real archive | Recordings of the actual event | Low, if sourcing is broad | Yes, with a named source | Whatever really happened |
| Stills with motion | Photographs, pushed or panned | Medium | Yes | Simulated |
None of these is disqualifying. Generated footage is the strongest option when the subject never existed in front of a camera. Stock is defensible for abstract sequences. Archive is the only one that can carry a claim about a real event, and it is the one whose weakness is availability rather than credibility.
The failure is not choosing any particular source. It is choosing one and using it for everything.
Checking your own video
Five things you can measure on a Short you already published, in about ten minutes:
- Median shot length. If it is under two seconds throughout, you are probably hiding dead clips rather than pacing a story.
- Cut rate against the narration. Play it with the audio only, then with the video only. If the cuts land mid-clause, the edit is following a timer.
- Source count. How many distinct places did the footage come from? One is a warning.
- Black frames. Step through every transition. There should be none.
- Silence. Any gap longer than a breath is a place the viewer can leave.
Every number in this piece describes a reference set of ten videos and one production manifest. Treat it as a band to check yourself against, not a specification to hit — a Short at 26 cuts per minute is not automatically wrong, it is outside the range we measured and worth a second look.
<!-- faq -->Questions people ask
Is AI-generated footage always worse than archive footage?
No. Generated footage is the strongest option when the subject never existed in front of a camera, and the weakest when it did. The problem is not that it is generated, it is that a video built entirely from one source — generated, stock or archive — inherits that source's uniform surface. Our reference render used five separate archives for a 53-second piece.
How long can a single shot be in a Short?
Longer than the common advice suggests. Median shot length in our reference set ranged from 2.03 to 6.00 seconds, and the videos at the top of that range held attention. What matters is whether something is developing inside the shot, not how many seconds it occupies.
Does cutting faster improve retention?
Not on its own. Cuts per minute ranged from 8.2 to 25.1 across the reference set and both extremes worked. The cut rate that holds attention is the one that follows the density of what is being said, so a fixed interval is the wrong tool regardless of which interval you pick.
Why do so many faceless channels look identical?
Because they draw from the same few stock libraries, and material from one library shares a grade, a grain and a framing convention. Broadening the source count is the cheapest change available: it costs sourcing effort rather than production budget.
What loudness should a Short be?
Our reference set sat between -14.8 and -11.5 dBFS integrated. The precise figure matters less than the relative one: a video noticeably quieter than the feed around it loses viewers before the first sentence lands.
Does using archive footage affect monetisation?
Sourcing and monetisation are separate questions, and the second one turns on originality and attribution rather than on whether a machine was involved. It is worth reading YouTube's policy language directly rather than a summary of it.
What this is based on
- Frame-accurate reference set of 10 high-retention vertical documentary Shorts (7 English
- 3 Turkish) measured by the ViewMade engine
- Per-job production manifest for render b2e4b3eb showing 14 cuts across 5 separate archives with zero black frames and zero silence
- Measured loudness band of -14.8 to -11.5 dBFS across the reference set
Shorts
Related reading
- Shorts · 6 min read
How to publish YouTube Shorts daily when you cannot edit video
Where "no editing required" is true, where it quietly stops being true, and the six checks that let someone who has never opened a timeline judge whether a finished video is any good.
- Shorts · 4 min read
We measured the cut rhythm of Shorts that work. Here are the numbers.
Ten high-retention vertical documentaries, measured shot by shot: how long they run, how often they cut, and how long a single shot is allowed to sit before attention leaves.
- Faceless channels · 7 min read
The best AI video generator for a faceless YouTube channel in 2026, judged by output
Most comparisons rank these tools by feature list. This one sorts them by what they actually produce, using a frame-accurate reference band from ten high-retention Shorts as the measuring stick.
Stop researching. Start uploading.
ViewMade finds the topic, writes the script, sources real footage and delivers a finished video.
See what it makes