Scripts · 6 min read
Writing for a synthetic voice
ViewMade offers narration in 5 languages, a figure measured from its own output documentation on 22 August 2026.
Published
ViewMade offers narration in 5 languages, a figure measured from its own output documentation on 22 August 2026. That number is small enough to make the point of this article plainly: writing for a synthetic voice is not the same job as writing for a human one. This article covers what changes in the script when the reader is software, and how to structure sentences so they survive it.
The short answer
Writing for a synthetic voice means removing everything the voice has to interpret rather than read. Synthetic voices handle short declarative sentences, explicit punctuation, numbers written out with their units, and one idea per sentence. They stumble on long clauses, ambiguous abbreviations, rhetorical flourishes, and implied emphasis that a human narrator would supply through tone. The practical method is simple: write the script as if a careful but literal reader will perform it, then cut every sentence that depends on performance rather than text.
Why synthetic voices fail on certain sentences
A synthetic voice reads characters. It does not understand intent behind them, so every ambiguity in the script becomes audible. Three failure patterns cover most of the damage.
First, dependent clauses stacked inside one sentence. A line that hangs several qualifications off one subject forces the voice to carry multiple levels of nesting without a breath. Splitting it into three short sentences fixes it.
Second, abbreviation and shorthand. Approximations, century markers, and rounded counts all render badly or inconsistently. Write the full word, write the exact figure, spell it out.
Third, rhetorical devices. A dash used for dramatic pause, an ellipsis implying hesitation, all caps implying shouting: none of these survive. The voice applies its own pacing regardless, so the device adds nothing except risk.
The general rule: if a sentence needs a human to sound right, rewrite it until it does not.
Sentence length and rhythm
Short sentences are the default. A brief sentence gives the speech engine clean boundaries and gives the listener a natural pause point between ideas. Longer sentences work only when they are structurally flat: one subject, one verb chain, no nesting.
Varying rhythm matters less than you would expect. Human narrators build tension through pacing changes; synthetic voices have limited control there, and most viewers of short-form documentary content do not register the absence. What registers instead is clarity. Two consecutive sentences at similar length read as steady, which suits narration.
Practical checks before you record:
- Read each sentence aloud yourself. If you run out of breath, split it.
- Count commas per sentence. More than two usually means nesting.
- Check that every number appears with its unit attached.
- Confirm no sentence relies on a pause mark for meaning.
Punctuation, pauses, and numbers
Punctuation is your only real directing tool with a synthetic voice. Periods create hard stops. Commas create shorter breaks. Colons often produce a pause plus a slight tonal shift into what follows. Use these deliberately: end a paragraph's final sentence with a period where you want silence, use a comma mid-sentence where you want flow.
Numbers deserve special handling. Write digits for years, spelled-out words for small counts where the surrounding prose is narrative, and always attach units: "eleven words", "fifty renders per month". Never leave a bare numeral floating, because the voice may read it in isolation and the viewer loses context. Dates should be unambiguous: write the day, month name, and year in full, since slash formats are read inconsistently across engines.
Avoid semicolons in narration scripts. They signal written prose, and voices treat them unpredictably. Replace with a period and a new sentence.
Structuring a script the voice can carry
Structure the script around facts, not around a presenter's personality. A working pattern for a vertical documentary script:
- Open with the measured fact. One sentence, one number, dated if the number could change.
- State what the video will show. One sentence.
- Move through the body as a sequence of single-idea sentences, grouped into beats of roughly three to five sentences each.
- Close with a plain statement, not a call to emotion.
Each beat should be self-contained enough that a viewer joining mid-video follows it. This is also why scripts for synthetic narration tend to read slightly repetitive when printed: repetition of framing phrases such as dated references is doing orientation work that a human narrator would do with tone.
Keep visual cues out of the spoken track. If your production pipeline uses markers for archive footage, put them in a separate column or file. Mixing directions into the script risks a stray direction being voiced.
Where ViewMade fits in this workflow
ViewMade generates the full package from a topic: research, script, narration, real archive footage, burned-in captions, and a media credit file naming the source of every clip. Its narration layer supports 5 languages, measured on 22 August 2026, and the script it produces follows the pattern above: declarative sentences, numbers with dates, no rhetorical devices. If you write scripts by hand for a synthetic voice, the same constraints apply whether the tooling is yours or outsourced; the voice does not care who wrote the text.
What this does not tell you
This article describes structural rules, not quality measurements. We have not measured how specific speech engines render specific punctuation, so the guidance here reflects common behavior across text-to-speech systems rather than benchmarked results per engine. We have not compared listener retention between hand-written and machine-generated scripts, and we do not claim any retention effect from following these rules. The claims about sentence failure modes are drawn from editing practice, not from controlled tests. If you need engine-specific direction, such as exact pause timing on a particular platform, test it against your own output; this article will not answer that question.
<!-- faq -->Frequently asked questions
Should I write the script first or pick the voice first? Write the script first under the constraints above. Voice selection changes timbre, not tolerance for bad structure. A well-formed script works across voices; a poorly formed one fails on all of them. Choosing a voice first tends to lock you into compensating through retakes instead of fixing the text.
How do I handle statistics in a narration script? Attach three things to every statistic: the value, the unit, and the measurement date. "Nineteen words per page, measured 19 August 2026" survives narration; an undated approximation does not. If you cannot date the figure, say so in the script itself or drop the sentence. Undated numbers invite the viewer to distrust the whole video.
Do captions change how I should write? Yes, mildly. Word-for-word captions mean every filler phrase appears on screen. Cut fillers, false starts, and repeated framing words that exist only for spoken rhythm. Keep sentences short enough that a caption line fits comfortably on a vertical frame; long lines force wrapping that competes with the footage.
Can I use humor or personality in a synthetic-voice script? Limited amounts. Deadpan statements land better than jokes, because jokes depend on timing the voice cannot deliver. If you want personality, put it in sentence choice and subject matter rather than delivery cues. A dry, precise script performed flatly often reads as intentional style; a joke performed flatly reads as an error.
What this is based on
- ViewMade render output specification
- read from the product on 2026-08-22
- sample 1
- captionStyles 7
- narrationLanguages 5
- interfaceLanguages 7
Scripts
Related reading
- Scripts · 6 min read
AI script writer vs writing it yourself
On 22 August 2026 we pulled the server-side HTML of a competitor's pricing page without running JavaScript and counted 4 words of body text; our own homepage
- Scripts · 7 min read
Avoiding the tells of AI writing
On 19 August 2026, a direct HTTP pull of storyshort.ai without running JavaScript returned 11 words of body text on every page tested, while the same method
- Scripts · 7 min read
Fact checking a documentary Short
A documentary Short built by ViewMade ships with a media credits file that names the source of every clip in the render, alongside a Starter plan priced at $29
Stop researching. Start uploading.
ViewMade finds the topic, writes the script, sources real footage and delivers a finished video.
See what it makes