This entry covers the skeleton of a video prompt: what to write for each of the five elements—subject, scene, motion, camera, style—and in what order. It suits creators transitioning from image prompts to video, and can also serve as a pre-writing checklist.
The Five Elements and Recommended Order
Write in the order “subject → scene → motion → camera → style,” moving from concrete to abstract, with 1–2 sentences per element:
- Subject: Who/what is in the frame. Give 2–4 visual anchors (age, clothing, color, material) so the subject stays recognizable and consistent throughout the shot.
- Scene: Where, when, and what atmosphere. Location + time + light/weather, in 1–2 short phrases—don’t turn it into a novel.
- Motion: What the subject is doing. This is the core difference between video and images—mandatory. Verbs should be specific (“slowly turns” beats “is moving”); 1–2 coherent actions per prompt is ideal.
- Camera: Shot size + camera position + camera movement. Use standard terms from the camera movement glossary (e.g., push in, pan, tracking shot)—don’t invent your own phrasing.
- Style: Visual style/lighting/texture/color tone. End with 3–5 words placed last to avoid diluting the concrete details before them.
A full prompt is recommended at 60–120 Chinese characters (40–80 English words); under 30 characters lacks information, and beyond 150 characters the later descriptions tend to get diluted. Effective length limits per model: unverified (to be filled in after real-world testing).
Complete Example (Annotated by Element)
Chinese:
一位穿驼色长风衣的银发老妇人〔主体〕,黄昏的空旷海边栈桥,海雾弥漫〔场景〕,她缓缓转身面向镜头,风衣下摆被海风掀起〔运动〕,中景,平视机位,镜头缓慢向前推进〔镜头〕,电影感,冷蓝色调,柔和逆光,35mm 胶片颗粒〔风格〕
English:
A silver-haired old woman in a long camel trench coat,〔subject〕 on an empty seaside pier at foggy dusk,〔scene〕 she slowly turns to face the camera as the sea wind lifts the hem of her coat,〔motion〕 medium shot, eye level, slow push in,〔camera〕 cinematic, cool blue tones, soft backlight, 35mm film grain〔style〕
Degradation Behavior When Elements Are Missing
- Missing motion → the frame is nearly static, with only slight “breathing” drift; the result looks like an animated GIF rather than a video
- Missing camera → the model decides framing and movement on its own, often producing random push-ins, pans, or even mid-shot cuts
- Missing subject anchors → the subject’s appearance drifts frame by frame, with face or clothing changes
- Missing scene → the background is randomly generated and feels disconnected from the subject’s character
- Missing style → the output falls into the model’s default texture—“correct but bland”
These are general tendencies; specific degradation behavior per model: unverified (to be filled in after real-world testing).
Common Mistakes
- Mixing all five elements into one muddle (style words inserted mid-action, subject details scattered throughout) → the model can’t grasp the focus, and key information gets swallowed
- Stacking only style words with no motion (“cinematic, epic, 8K, masterful lighting…”) → you get a slightly wobbling image
- Inconsistent subject descriptions (says “red dress” in the first half, “blue coat” in the second) → the subject transforms mid-shot or blurs into a hybrid of both
- Cramming 3+ actions into one prompt → actions get skipped or compressed into twitching; better to split into multiple segments and generate separately