Craft Guide

How to Make an AI Documentary Video — A Craft Guide for Beginners

Script-first workflow, three-tier visual planning, reconstruction ethics, EU AI Act disclosure rules, documentary prompting techniques, and a complete worked outline. The guide the tools don't teach.

There's an uncomfortable truth at the center of this subject: AI does not help you make a documentary. It helps you illustrate one.

The actual work — finding out what happened, deciding what it means, and building an argument a viewer will trust — is unchanged. What generative tools have genuinely transformed is access to footage of things you could never afford to film, or that no longer exist to be filmed. That's a real unlock. It's also the part most likely to get you into trouble if you don't think carefully about it.

Part 1: Decide What Kind of Documentary You're Making

The label covers wildly different forms, and they place very different demands on generated imagery. Pick one — beginners who try to blend an explainer with an investigation usually end up with neither.

🟡 The Historical Narrative — Use with Care

You're reconstructing events before or beyond the reach of a camera. Generated imagery does something genuinely valuable here — but the honesty burden is high, because viewers instinctively read documentary imagery as evidence. Reconstruction must be clearly marked.

🟡 The Profile — Hardest Category

A person, living or dead. This is the hardest category because the central ethical constraint — don't synthesize a real person's face or voice — sits right on top of the material you most want. Proceed with legal advice for public figures.

🔴 The Investigative Piece — Almost No Place for Generated Footage

You're making factual claims about real people and real events. Generated imagery has almost no place here beyond abstract or diagrammatic sequences. If your film's power depends on the viewer believing what they see, don't show them things that didn't happen.

Part 2: The Rules That Have to Come First

Most guides put ethics at the end, as a disclaimer. That's backwards — this decision shapes your entire shot list. The reconstruction line is where the craft begins.

Three lines worth holding

Never generate something that purports to be a record of a real, specific event. A generated clip of a real protest, a real crash, a real speech is a fabricated document — regardless of intent.
Never synthesize a real person's face or voice without explicit permission — and for the deceased, without the estate's. This is both an ethical floor and, increasingly, a legal one.
Always mark reconstruction visually. Give generated sequences a consistent treatment — a grade, a grain, a slight stylization, a subtle frame — so the viewer learns within the first minute which register they're in.
⚖️ Legal Requirements — 2026

EU AI Act Article 50 (applicable from 2 August 2026)

Requires clear disclosure of AI-generated or manipulated audio, video, image, and text, with a reinforced obligation on deepfakes. Applies to deployers whose content reaches the EU market — which, for anything on the open internet, is most people. Penalties are substantial.

YouTube (May 2026)

Moved from voluntary disclosure to automatic detection and labeling of significant photorealistic AI content. Undisclosed synthetic footage may get labeled whether or not you check the box.

California SB 942

Comparable disclosure requirements for AI-generated content. C2PA provenance metadata is emerging as the technical standard underneath all of this.

The practical takeaway:

  • Check the platform's synthetic content disclosure at upload. Always.
  • Put a plain-language line in your description: which parts are generated, which are real.
  • Include a disclosure card in the film itself before or after generated sequences.
  • Keep your generated assets' provenance metadata intact rather than stripping it.

Disclosure costs you nothing with an audience already assuming AI involvement. Getting caught not disclosing costs you the entire film's credibility.

Part 3: Research Before Script, Always

Your documentary is only as good as what you actually know. Before writing a word of narration, build three things:

DocumentWhat it isWhy it matters
Source documentEvery fact, with a link or citation next to itSomething you can hand to someone — not a browser history
Claim ledgerTwo columns: the claim you're making, and the evidence for itEvery narration sentence should trace to a row in this ledger. If it doesn't, you're speculating
The thesisOne sentence: your argument"The 1889 flood was bad" is a topic. "The 1889 flood was an engineering decision, not a weather event" is a documentary

Prefer primary sources — court records, transcripts, papers, archives, direct interviews. A documentary built on other people's YouTube videos is a remix, and viewers can feel it.

Part 4: Write the Script First, Then Find Pictures

This is the single biggest workflow difference between documentary and fiction. In narrative work you often start with images. In documentary you start with words, and pictures serve them.

👂

Write for the ear

Short sentences. One idea each. Concrete nouns. Read every paragraph aloud — if you stumble, the sentence is wrong. Prose that looks elegant on a page frequently sounds pompous in a voiceover.

🔢

Do the timing math

Comfortable narration runs 140–160 words per minute. A ten-minute documentary is roughly 1,300–1,500 words of narration — less than beginners expect, because you need stretches with no narration at all.

📋

Build the paper edit

Add a second column to your script showing what the viewer sees during each line. This is your shot list. Build it before generating anything or you'll end up with forty beautiful clips and no idea what they're for.

🗂️

Structure in sequences

A sequence is 60–120 seconds that does one job. A ten-minute film is six to eight sequences. Name each on an index card and order them before writing any prose. Each sequence should be reducible to one sentence.

Part 5: Voice Is the Film

Viewers forgive a lot of visual roughness. They will not forgive a narrator they don't want to listen to. This is the production decision that carries the most weight and gets the least attention.

Use your own voice if you possibly can. Even a mediocre human read with a $60 microphone and a duvet over your head usually beats an excellent synthetic one, because listeners respond to the sense of a person who cares about the material. Record in a small room with soft furnishings, close to the mic, speak more slowly than feels natural.

If you use synthetic narration — direct it

Generated voice is not a text box you paste into. These are the craft controls:

ProblemFix
Rhythm and pacingPunctuation is timing. Commas, full stops, and paragraph breaks are your primary controls. Break long sentences into short ones to force pauses where you want emphasis.
Inconsistent energyRecord in chunks — generate section by section, not the whole script at once. You get more consistent energy and can re-roll a bad line without regenerating everything.
PronunciationProper nouns, place names, and technical terms are where synthetic voices break the spell. Use phonetic overrides or spelling hacks. Check every name.
The liltSynthetic narration defaults to an even, slightly upward, uniformly enthusiastic cadence. Cure it in the writing: vary sentence length aggressively, use fragments, let some lines be flat statements.
Over-tighteningDon't remove every pause. Human narration has air in it. Leave the gaps.

Pick one voice and never change it. Consistency of narrator is more important than quality of narrator. And disclose synthetic voice — it falls squarely inside current disclosure rules.

Part 6: The Three-Tier Visual Plan

Before generating anything, sort every item in your paper edit into three tiers. The discipline is that Tier 3 fills gaps Tiers 1 and 2 can't — not the other way around.

Tier 1 — Use First, Always

Real Material

Archival photographs, public-domain footage, documents, maps, your own filmed footage, screen recordings, charts built from real data. A grainy real photograph of the thing is worth more than a flawless generated impression of it — both evidentially and emotionally. If this exists, use it.

Tier 2 — Use When Tier 1 Doesn't Cover It

Licensed Stock

For generic connective imagery — traffic, weather, crowds, landscapes, hands typing. Stock is cheap, real, and consistent. Many beginners generate footage they could have licensed for a few dollars in higher quality. Check this tier before opening a generator.

Tier 3 — Last Resort, Clearly Marked

Generated Footage

For what cannot be obtained from Tiers 1 and 2: a room that no longer exists, a process inside a body, a scale or place that can't be filmed, an abstract concept made visual, a reconstruction of an undocumented moment. Documentaries that generate everything look like nothing — uniform, weightless, unanchored. The texture contrast between real archival material and generated sequences is itself expressive.

Part 7: Prompting for Documentary Imagery

Documentary prompting is nearly the opposite of trailer or narrative prompting. There, you push toward the cinematic. Here, cinematic is a liability — it reads as made, which is exactly the wrong signal for a form that depends on the viewer's trust.

✗ Narrative/trailer prompt
Epic wide shot of an oil rig at dawn, cinematic, golden hour, 4K, stunning visuals, anamorphic lens
✓ Documentary prompt
Handheld observational shot, available light only, slightly underexposed. An oil rig on flat grey water. Overcast, no visible sun. 16mm grain, slight camera shake, imperfect framing, mixed color temperature. Subject partially obscured by sea haze.

Documentary prompting vocabulary

Shoot style

observational, handheld, fly-on-the-wall, available light only, no fill light, shot from a distance, subject partially obscured

Imperfection

soft focus, dust on lens, motion blur, blown highlights in the window, flat contrast, faded colors, minor scratches, slight underexposed

Period accuracy

Name specific clothing styles, signage typography, vehicle models, architectural details. Research before prompting. Hunt every frame for anachronisms.

Avoid faces

Hands, backs of heads, crowds from behind, rooms after everyone left, objects on tables, feet. Faces are least convincing and most likely to imply a specific real person.

Period accuracy is where beginners fall apart. If you're depicting 1911, every element in frame is a factual claim — clothing, signage, typography, vehicles, architecture, hairstyles, materials. Models will happily give you a 1970s streetlamp in a 1911 street. Research the period specifically, name the details in the prompt, then audit every generated frame for anachronisms.

Other visual tools that aren't generated video

ToolWhen to use it
Stills with motion (Ken Burns)Slow push or drift across a real archival photograph. Cheap, uses real material, and the slowness reads as seriousness. Keep moves slow, in one direction, ending before the cut.
Documents on screenShow the actual page. Highlight the actual line. Nothing builds trust faster than letting the viewer read the evidence themselves.
Maps and dataAnimated maps orient viewers instantly. Simple animated charts from real data are far more persuasive than any imagery. Learn basic motion graphics for these.
Lower thirds and typographyOne typeface, one weight for names, one for titles. Names visible for at least three seconds. Restraint reads as authority.

Part 8: Assembly and Pacing

🎙️

Narration first

Put the full VO on the timeline before a single picture. The narration is the spine; the pictures hang on it. This is the opposite of how most beginners work.

📸

Tier 1 before Tier 3

Place real footage and stills first. Only then fill remaining holes with generated and stock material. This ordering keeps generated footage in a supporting role by construction.

🌬️

Give the film air

Every 90 seconds, stop talking. Let a shot run with only ambience for 4–5 seconds. Beginners narrate continuously and the result is exhausting. Pauses are where a documentary breathes.

✂️

Cut on meaning

This is the opposite of trailer editing. Change the picture when the idea changes. Evidence shots hold long. Transitional shots move fast. Vary shot length by function, not by instinct.

Part 9: Sound

Ambience under everything. Every generated shot should have room tone, wind, a street hum — something. Silent picture reads as fake more than any visual artifact does. Adding ambience is the highest-value hour you'll spend in post.
ElementGuidance
AmbienceUnder every shot, always. Room tone, wind, hum. Even three seconds of quiet room tone under a silent shot transforms it.
ScoreNearly invisible — sparse, low, often absent. If the music is telling the audience how to feel, the script isn't doing its job. Cut score entirely under your most important lines.
MixNarration sits on top; everything else ducks under it. If a viewer strains to hear the voice, nothing else matters.
LicensingMusic and archival material both. A documentary taken down for a music claim is a wasted year. License everything before publishing.

Part 10: Fact-Check, Then Credit

Before publishing, take the claim ledger from Part 3 and walk the finished film against it line by line. Every factual sentence gets checked against its source in the form it actually appears in the cut — narration often drifts toward overstatement during editing.

Then build a sources card or description block listing your major sources, archival credits, and AI disclosure. Documentaries that show their work get taken seriously. It takes twenty minutes.

Part 11: The Beginner Mistake Checklist

Before you export, check all eleven.

  • Generating footage before writing the script
  • Generated imagery presented as a record of real events
  • No disclosure at upload or in the description
  • A synthetic narrator that sounds like a synthetic narrator (the lilt, the uniform enthusiasm)
  • Continuous narration with no silence anywhere
  • Everything generated, nothing archival — uniform, weightless, unanchored
  • Anachronisms in period reconstruction (the 1970s streetlamp in the 1911 street)
  • Silent picture with no ambience under it
  • Music too loud and too constant — tells the audience how to feel
  • Facts without sources — no claim ledger, no sources card
  • Thirty minutes long on a ten-minute idea

Part 12: A Worked Outline — Start to Finish

Thesis: The 1889 Johnstown flood was an engineering decision, not a weather event.
Length: Ten minutes · Narration: ~1,200 words · Generated footage: One sequence only
SEQ 1 The Place · 90 seconds

Footage: Archival photographs of the valley with slow Ken Burns pushes. One generated wide of the reservoir at dusk, graded to match the archival stock, marked with a reconstruction card.
Job: Establishes the town, the dam above it, and the sense of scale.
Generated: One shot. Clearly marked.
SEQ 2 The Decision · 120 seconds

Footage: Documents on screen — the actual engineering correspondence, key lines highlighted.
Job: Delivers the evidence. This is the argument's foundation.
Generated: None. This is evidence and it should look like evidence.
SEQ 3 The Warning · 90 seconds

Footage: Animated map showing the valley's geography and dam position. Narration walks the chain of ignored reports.
Job: Shows why the outcome was avoidable.
Generated: None. Maps are more persuasive than any imagery here.
SEQ 4 The Failure · 60 seconds

Footage: Generated reconstruction, stylized and clearly marked: water, structure, scale. No people, no faces. Ambience only, no score.
Job: Shows what could not be filmed.
Generated: All of it — this is the one sequence where nothing real exists.
SEQ 5 After · 90 seconds

Footage: Real archival photographs of the aftermath with slow moves. Very little narration. Let the images sit.
Job: Emotional weight through restraint.
Generated: None. Real photographs do this work better than anything else.
SEQ 6 The Reckoning · 120 seconds

Footage: Court records on screen, the outcome, the thesis restated. Sources card at close.
Job: Completes the argument and shows the work.
Generated: None.
Ten minutes, roughly 1,200 words of narration, and generated footage in exactly one sequence — where nothing real exists. That ratio is the discipline. Documentaries that generate everything look like nothing.

Start with an Explainer

If this is your first documentary, don't start with the historical reconstruction. Make a six-minute video essay about something you already understand well, where the visuals are frankly illustrative and the honesty burden is low. Learn the script-first workflow, the voice, the pacing, and the sound on a project where you can't do much harm.

The tools will keep improving. The obligation to know what you're talking about won't change at all.

Tools for Making AI Documentaries

Revid AIAI video generation for documentary B-roll — code CLIPVERDICT for 20% off
ElevenLabsAI narration with voice cloning — read our review
DaVinci ResolveFree editor for color grading, audio mixing, motion graphics
Licensed musicEpidemic Sound, Artlist, or Musicbed. License before publishing.
Try Revid AI Free → AI Movie Trailer Guide →

* Affiliate link. Full disclosure →

Frequently Asked Questions

Is it legal to use AI-generated footage in a documentary?
Yes, with required disclosure. EU AI Act Article 50 became applicable on 2 August 2026 and requires clear disclosure of AI-generated audio, video, image, and text. YouTube moved to automatic detection and labeling of significant photorealistic AI content in May 2026. California's SB 942 imposes comparable requirements. Disclose at upload, in your description, and in the film itself. It costs you nothing with an audience already assuming AI involvement.
Can I use AI to generate footage of real historical events?
With strict conditions. Never generate something that purports to be a record of a real, specific event — a generated clip of a real protest or speech is a fabricated document regardless of intent. Historical reconstruction is acceptable when clearly marked as reconstruction, visually distinguished from archival material, and disclosed as AI-generated. Research the period specifically — models will place 1970s streetlamps in 1911 streets, and every anachronism is a factual error.
Should I use my own voice or AI narration?
Your own voice if you possibly can. Even a mediocre human read with basic equipment usually beats a synthetic one, because listeners respond to a person who cares about the material. If you use synthetic narration, direct it carefully — punctuation controls rhythm, record in chunks, fix proper nouns, and fight the uniform upward lilt synthetic voices default to. Disclose it regardless.
How long should an AI documentary be?
Start with six minutes. Comfortable narration runs 140–160 words per minute, so a six-minute documentary is roughly 700–800 words of actual narration — less than beginners expect, because you also need stretches with no narration. A ten-minute film is six to eight sequences of 60–120 seconds each. Beginners consistently write thirty minutes of material for a ten-minute idea. Write to the count.