Craft Guide

How to Write AI Video Prompts That Actually Work (2026)

Stop writing vibe prompts. The shot brief framework — 10 components, Revid-specific syntax, platform tips for TikTok and Shorts, ready-to-use templates, and a failure/fix table.

Most AI video prompts fail for one reason: they describe a vibe, not a video.

"Make a cinematic video about productivity" isn't a prompt — it's a wish. The model doesn't know what platform you're publishing on, what the viewer sees in the first frame, what moves, where the camera is, or what the video is supposed to accomplish. So it guesses. And that guess looks like every other generic AI video.

The Shot Brief Mental Model

An AI video model is not a genie. It's a camera crew waiting for direction. When a director says "make it cinematic," the crew asks: which lens? Where's the camera? What's the subject doing? What's the lighting source?

A prompt that works is a mini production brief. Not a longer prompt — a smarter one. The formula:

Goal + Platform + Shot Type + Subject + Action + Setting + Camera + Style + Audio + Constraints
✗ Vibe description — doesn't work
Make a viral productivity video.
✓ Shot brief — works
9:16 TikTok opening. Close-up product shot.
Subject: messy desk — laptop, notebook, coffee.
Action: phone lights up with notifications, then notebook slides in with "3-minute focus reset."
Camera: slow push-in from slightly above, shallow DOF.
Style: clean creator aesthetic, soft daylight, warm neutrals.
Audio: notification pings, then calm upbeat music.
Constraint: readable text space at top for caption.

How AI Video Prompting Changed in 2026

Modern AI video tools have separated two distinct layers. Mixing them makes results worse.

Layer 1 — The Container

Aspect ratio, duration, platform, output quality, export format.
Set this in the tool's settings — not in your prompt.

Layer 2 — The Scene

Subject, action, camera, setting, lighting, style, audio, pacing.
This is what your prompt controls.

Writing "make this exactly 9:16" inside a prose prompt often doesn't work if the tool has a separate aspect ratio setting. In Revid AI, set format controls in the tool settings; use your script to control story, visuals, pacing, and tone.

The 10 Components of a Working AI Video Prompt

1Goal — what job is the video doing?

Tell the model the job. A video meant to explain a concept gets different shots than one meant to drive immediate action. "Create a product demo that makes freelancers understand the benefit in under 10 seconds" gives the model a purpose. "Create a cool product video" gives it nothing.

2Platform — where is this video going?

A TikTok hook video is not a YouTube explainer. Platform affects pacing expectations, safe zone placement, and caption style. Specify: "9:16 TikTok video with a visual hook in the first second" vs. "16:9 cinematic website hero background, no text, no dialogue."

3Shot type — how should the camera frame the scene?

Useful shot types: close-up, medium shot, wide shot, overhead shot, POV shot, tracking shot, macro shot, establishing shot, screen-recording style, talking-head, product hero shot. Name the shot type at the start of every prompt — it anchors everything else.

4Subject — who or what is in the frame?

Weak: "A person working." Strong: "A tired freelance designer in a black hoodie sits at a small desk covered with sticky notes, a laptop, and a half-empty coffee cup." Be specific enough to anchor the visual. Only include details that need to stay visible — overloading the subject description confuses composition.

5Action — what is actually happening?

This is where most prompts improve fastest. The AI needs a verb, not a mood. Weak: "The founder is successful." Strong: "The founder refreshes the dashboard, sees the revenue graph spike, freezes for a second, then laughs in disbelief." The more physical the verb, the better the result.

6Camera — how does the shot move?

Camera instructions shape emotional experience. Useful movements: slow push-in, handheld documentary feel, locked-off tripod, low-angle tracking, overhead flat lay, fast whip pan, gentle dolly backward, rack focus. Use one camera move per prompt — conflicting instructions produce weird motion.

7Lighting and color — what does the scene look like?

Weak: "cinematic lighting." Stronger: "soft morning window light, warm beige shadows, muted blue-gray background, realistic skin tones." Describe the quality of light and choose a palette of three to five colors for consistent visual direction.

8Style — what is the aesthetic direction?

Style should support content, not bury it. Useful labels: UGC ad, documentary, clean SaaS demo, faceless educational, high-energy TikTok edit, cinematic B-roll, retro VHS, anime-inspired. Don't stack contradictory styles — "cinematic anime Pixar cyberpunk UGC" gives the model no coherent direction. Pick one.

9Audio — what does the viewer hear?

Audio belongs in the prompt, not as an afterthought. Include voiceover tone, dialogue, music style, sound effects, ambient sound. Example: "Audio: quiet room tone, soft keyboard typing, one notification ping, calm voiceover with confident pacing." A prompt with no audio instruction is a silent film you didn't mean to make.

10Constraints — what should the AI avoid or preserve?

Constraints prevent the most common failures: "No extra people. Keep text inside center safe zone. Do not change the character's outfit. One continuous shot. No distorted hands." Important: rephrase negatives positively — "empty hallway" beats "no people"; "open natural landscape" beats "no buildings." Positive descriptions give a clearer rendering target.

The 2 Rules That Make the Biggest Difference

Rule 1: Prompt for motion, not appearance

Video is not an image with extra seconds. A static description of how something looks tells the AI nothing useful about how the shot should feel in motion.

✗ Appearance prompt
A sleek black electric bike in a futuristic city, cinematic lighting.
✓ Motion prompt
A sleek black electric bike glides through a rain-soaked city street at night. Camera tracks low beside the front wheel as neon reflections ripple across wet pavement. Steam rises from a street vent as the bike passes. Slow push-in on the glowing dashboard.

Strong motion words: slides, turns, blinks, leans, floats, sprints, rotates, pours, snaps open, drifts, reveals, zooms, tracks, pushes in, pulls back.

Rule 2: One scene per generation

Most AI video models struggle with multiple scene changes in one clip. Generate one scene at a time, then edit them together.

✗ Too many scenes
Show a founder waking up, checking analytics, filming a TikTok, meeting investors, launching a product, and celebrating — cinematic montage.
✓ One scene (generate the rest separately)
A founder sits alone at a kitchen table before sunrise, laptop open, blue analytics dashboard glowing on their face. They refresh the page, pause, then smile as the graph spikes. Slow handheld push-in, quiet room, warm practical lamp, documentary style.

Writing Prompts for TikTok, Reels, and Shorts

For short-form platforms, the prompt needs to address platform behavior, not just scene content. Start with the hook.

Write the hook first

Four questions before anything else:

  • What does the viewer see in the first frame?
  • What text appears in the first second?
  • What is the first spoken line?
  • Why would someone keep watching?
✗ Weak opening
Create a video about saving time with AI.
✓ Strong opening — retention strategy built in
Opening frame: a creator staring at a 12-tab browser with a stressed expression. Large caption at the top: "You're not slow. Your workflow is broken." In the first second, the cursor rapidly jumps between tabs while notification sounds overlap.

Platform-specific notes

PlatformKey prompt considerations
TikTok9:16, hook in first 3 seconds, clear text overlays, under 60s for organic discovery. Avoid bottom-right and right-edge safe zones where UI overlaps.
YouTube ShortsVertical up to 3 minutes, but shorter performs better for new channels. Videos meeting vertical criteria are automatically categorized as Shorts after upload.
Instagram ReelsReels over 3 minutes are not recommended to new audiences — get to the point fast when discovery is the goal.

Safe zones

Add to every short-form prompt: "Keep all important text in the center third of the frame. Avoid placing key text near the bottom, right edge, or top UI area. Leave clean negative space behind captions."

And fewer words on screen. "Here are the three most important things you need to know" is a terrible caption. Three short versions are better: "Your prompt is too vague" / "Video needs motion" / "Use this formula."

Revid AI Prompt Syntax — Four Tools Most Users Miss

Revid AI isn't a raw video model — it's a complete short-form video workflow: your script goes in, and a video with visuals, voiceover, captions, music, and edits comes out. That changes how you write.

SyntaxWhat it doesExample
Line breaksEach new paragraph triggers a different visual sceneFirst spoken sentence.

Second sentence. Visual just changed.
[Bracket notes]Visual instructions that don't get spoken[Visual: messy timeline] Most people don't have a content problem.
<break time="Xs" />Controlled pauses in voiceover for emphasisMost AI prompts fail for one reason. <break time="0.5s" /> They describe a vibe.
Short sentencesCleaner TTS delivery — punctuation controls prosodySubject. Action. Camera. Timing. (not one long clause)

A complete production-ready Revid script

[Opening visual: a slot machine spinning with the words "cinematic," "viral," and "epic."] On-screen caption: "This is why your AI videos look random." Most people are not prompting. They are gambling. <break time="0.4s" /> [Visual: the slot machine stops on three vague words: cool, epic, cinematic.] A vague prompt gives the AI too many choices. [Visual: the vague words transform into a checklist: Subject, Action, Camera, Timing.] A good video prompt works like a shot brief. It tells the AI what the viewer sees. What moves. Where the camera is. [Visual: split screen. Left: "Make a viral productivity video." Right: "Close-up of a phone lighting up with notifications while the camera slowly pushes in."] This prompt is a wish. This prompt is a scene. [Ending visual: a vertical video draft with captions, voiceover, and a publish button.] Stop describing the vibe. Start directing the shot.

Ready-to-Use Prompt Templates

Universal template

Create a [duration] [aspect ratio] video for [platform]. Goal: Help [audience] understand/want/do [specific outcome]. Opening frame: [describe first visual in detail.] On-screen text: "[hook caption]" Main action: [Subject] [specific physical action] in [setting]. [Secondary movement or environmental detail.] Camera: [Shot type], [movement], [framing], [focus]. Look: [Lighting], [color palette], [style], [mood]. Audio: [Voiceover / music / SFX / ambient]. Editing: [Pacing, cuts, captions, CTA]. Constraints: [No extra characters, safe-zone text, consistent outfit, etc.]

TikTok / Reels hook template

Create a 9:16 short-form video. Target viewer: [specific audience]. First frame: [unexpected visual]. First caption: "[pattern interrupt]" First spoken line: "[strong claim or observation]" Beat 1: [problem visual]. Beat 2: [simple explanation]. Beat 3: [example]. Beat 4: [payoff]. CTA: [save / comment / follow / try].

Revid faceless video template

[Opening visual: describe a strong first frame.] On-screen caption: "[hook]" [Visual: describe scene 1.] First spoken sentence. <break time="0.4s" /> [Visual: describe scene 2.] Second spoken sentence. [Visual: describe scene 3.] Third spoken sentence with a concrete example. [Visual: describe final payoff or CTA.] Final spoken sentence.

Cinematic B-roll template

[Shot type] of [subject] in [setting]. Action: [Specific slow physical movement]. Camera: [Movement and framing]. Lighting: [Time of day, source, color palette]. Texture: [Steam, dust, rain, reflections, fabric, glass, metal]. Audio: [Ambient sound or subtle SFX]. Style: [Realistic cinematic / documentary / luxury / moody / lifestyle].

Image-to-video template

When using image-to-video, your prompt should focus on motion — not restating what's already visible in the image.

The [subject] slowly [specific motion action]. [Secondary environmental motion — steam, wind, light shift]. Camera: [one camera movement]. [What changes or is revealed at the end of the shot]. Style: [brief aesthetic note only if it differs from the image].

Why Prompts Fail — and How to Fix Them

ProblemLikely causeFix
Looks good but doesn't match the ideaPrompt described style, not actionAdd specific subject, action, setting, and camera
Character changes halfway throughNo reference or consistency constraintUse a reference image; describe what must not change
Output feels randomToo many scenes in one promptSplit into one scene per generation
Camera movement is weirdConflicting camera instructionsUse one camera move only
Visually busyToo many subjects or objectsOne main subject, one main action
Text unreadableNot designed for mobile or safe zonesFewer words, centered placement
Motion is weakPrompt only described appearanceAdd physical verbs and environmental movement
Style is inconsistentToo many style referencesOne style, one color palette
Negative prompt made it worseModel handled negatives poorlyRephrase positively: "empty street" not "no people"
Burning credits without improvementChanging too many things at onceChange one variable per iteration

How to iterate the right way

The fastest improvement is methodical iteration — not longer prompts.

  1. Generate the simplest version of your prompt
  2. Fix only the action in the next iteration
  3. Fix only the platform hook
  4. Fix only the edit or constraints
  5. Four focused iterations beats one heavily revised prompt

How Prompting Changes by Video Type

Video typePrompt directionKey constraint
Trailer / narrativePush toward cinematic — lens, grain, anamorphic flare, dramatic lightingOne action per shot; accelerate toward the drop
Documentary B-rollPush away from cinematic — observational, available light, imperfect framing, 16mm grainAvoid faces; mark reconstructions; match period accurately
Faceless TikTok / ShortsHook-first, fast cuts, bold captions, clear CTASafe zones; no faces needed; one scene per beat
Product demo / SaaSScreen-recording style, smooth zooms, cursor highlights, clean UI aestheticKeep UI text readable on mobile
UGC-style adHandheld front-facing, natural lighting, conversational tone, no beauty filterNo medical claims, no exaggerated transformations, no visible logos
Music videoVisual world matched to song mood; beat sync on downbeats and chorusNo lip-sync unless specified

Recommended Tool: Revid AI

Script-to-videoPaste script with bracket notes and break tags → complete video with captions, voiceover, and music
Article to VideoPaste a URL → Revid extracts, scripts, and generates automatically
Talking AvatarSolves character consistency entirely — the avatar is the character
DiscountCode CLIPVERDICT = 20% off any plan
Try Revid AI Free → Read Full Revid AI Review →

* Affiliate link. Full disclosure →

Frequently Asked Questions

Why do my AI video prompts keep failing?
The most common cause is describing a style instead of a scene. "Cinematic and viral" gives the model nothing actionable — it needs to know what the viewer sees in the first frame, what moves, where the camera is, and what the pacing should be. The second most common cause is too many scene changes in a single prompt. Generate one scene per clip.
How long should an AI video prompt be?
As long as it needs to remove ambiguity, and no longer. A simple motion clip needs 1–3 sentences. A short-form social scene needs 5–10 lines. A full Revid AI script works best at 100–250 words with bracketed visual instructions. Very long prompts with competing instructions often produce worse results than focused shorter ones.
Do negative prompts work for AI video?
Sometimes, but not consistently. Most official guides recommend clear positive descriptions instead. "Empty hallway" beats "no people." "Open natural landscape" beats "no buildings." Positive descriptions give the model a clearer rendering target.
How do I keep the same character consistent across AI video shots?
Use reference images when the tool supports them. Describe stable traits explicitly: age range, hair, outfit, posture, accessories. Add a constraint: "Do not change the character's outfit, age, or hairstyle." Describe characters by silhouette and wardrobe rather than face — wardrobe stays consistent where faces drift. For presenter-style content, Revid AI's Talking Avatar solves this entirely.