How to Write AI Video Prompts
A video prompt describes one shot, not a story: who is on screen, the one thing they do, where the camera stands and how it moves, the light and the look. This guide covers that order, the film words models understand, how Sora 2, Veo 3, Kling and Runway each read a prompt, and why length and frame shape belong in the app.
One shot, one action
Today's video models make clips of 4 to 15 seconds. That is one shot, and a shot holds one action well. "A red fox pounces into a snowdrift" fits in eight seconds; "a fox wakes up, hunts, eats and goes back to its den" does not, and the model will rush or skip the middle. When you need a sequence, write one prompt per shot and cut them together. OpenAI's Sora 2 prompting guide gives the same advice: keep motion simple and describe it in beats.
Subject, action, camera, light, style
Build the prompt in that order. The subject is who or what is on screen, concrete enough to picture: "a red fox in fresh snow", not "an animal". The action is a verb you can see: pounces, turns, spills, lifts off. Then the camera, then the light, then the look. A model fills every gap you leave with something generic, so each part you name is a decision you keep.
Direct the camera with film words
Models were trained on footage described in film language, so use it. A shot size says how close: wide shot, medium shot, close-up, extreme close-up. An angle says from where: low angle, high angle, over the shoulder, aerial. A move says what the camera does: static, slow push in, pull back, pan, tilt, tracking, orbit, crane up. Pick one move per clip. Two moves in eight seconds usually read as a wobble.
Each model reads the layout differently
- Sora 2 works well with a short scene line followed by a cinematography block: camera shot, lighting, style. OpenAI's guide frames a prompt as a storyboard panel briefed to a cinematographer.
- Veo 3 reads one shot-list sentence well: camera first, then subject, action, setting and look. Veo also makes sound, so a line of audio, such as wind or footsteps in snow, is worth adding. Google's Veo documentation covers the options.
- Kling follows the subject and its movement first, then the scene, the camera, the lighting and the atmosphere.
- Runway is directed by motion: lead with how the camera and the subject move. Runway's Gen-4.5 help also warns against negative phrasing: write "locked camera on a tripod", not "no camera movement", because "no X" can bring X in.
The video prompt generator lays your words out in each of these four orders, with nothing added but the labels.
Length and frame shape belong in the app
Every app sets clip length and frame shape in its own settings, and the setting wins over the prompt. Naming them in the prompt still helps the model pace the motion, but set them in the app too: 9:16 for a phone feed, 16:9 for a screen, 1:1 for a square post. Shorter clips hold an action more cleanly; use the longest setting only when the action needs it.
Change one thing at a time
Treat the first clip as a draft. If the motion is right and the light is wrong, change the light and nothing else, so you learn what each word does. The same habit works for stills: the guide to AI image prompts covers it from the image side, and a strong still is often the best first frame for an image-to-video clip.