This entry explains how to state the three format parameters—aspect ratio, duration, and frame rate (fps)—in the very first sentence of a prompt. They are global constraints; if placed at the end, they risk being overlooked, and getting them wrong means regenerating the entire video.
Why Put Them in the First Sentence
Aspect ratio, duration, and frame rate are not visual content but the “container specifications” of the entire video. They must be determined before generation begins and cannot be changed mid-way. When placed at the end of a prompt, the dozens of preceding words describing content dilute their presence, and the format declaration can easily be treated as a minor modifier (whether different models actually ignore it: unverified, pending real-world testing; but placing global constraints up front is always the safer approach). Build the habit: the first sentence covers format only; visual content starts from the second sentence.
Standard Sentence Patterns
- Chinese: “15 秒竖版 9:16 视频,24fps。画面:……”
- English:
a 10-second 16:9 video at 24fps, ...followed by the content description - All three elements present with clear units: seconds, ratio, fps; if one is missing, the platform’s default fills in, and the default may not be what you want
What Each of the Three Parameters Controls
Aspect ratio: Determines the composition approach and must be decided before writing content. 9:16 vertical (portrait) suits single subjects, vertical depth (corridors, waterfalls, full-body standing poses); 16:9 horizontal (landscape) suits wide scenes and multi-subject lateral staging; 1:1 suits feed content, with the subject centered being the safest. The same content requires different composition wording when the aspect ratio changes—you cannot reuse the text verbatim.
Duration: Constrained by the model’s single-generation limit; a common tier is 5–10 seconds, with some models supporting 15–60 seconds (refer to the current documentation of the model in use). Declaring beyond the limit typically results in truncation to the cap, or the entire motion being compressed into an awkward rhythm. For a 30-second final piece, the correct approach is to split it into 3–6 segments generated separately and then edited, rather than forcing 30 seconds into a single prompt.
Frame rate: Affects the feel of motion. 24fps is the cinematic baseline; 30fps is closer to TV and direct phone output; 60fps is smooth, suited to motion and game-like visuals, but can easily show a “too-smooth cheapness.” When in doubt, write 24fps.
Aspect Ratio and Platform Correspondence
- 9:16 vertical → mobile short-video feeds, full-screen immersion, final pieces typically 15–60 seconds
- 16:9 horizontal → long-form video platforms, TV, projection, the mainstream for narrative and multi-shot editing
- 1:1 / 4:5 → GIF-style short videos within image-text feeds
- The order is: decide the distribution platform first, then the aspect ratio, then write content—never the reverse
Does Declaring It Mean the Model Will Obey?
Not necessarily. In most products, aspect ratio is specified separately via a parameter panel or button; writing it in the prompt is only a fallback. The degree to which each model follows frame rate and duration: unverified (pending real-world testing). After each generation, check three things: the actual ratio, the actual seconds, and whether motion shows obvious frame drops. Trust the generated result, not the declaration.
Common Mistakes
- Format written mid-prompt or at the end → content is already composed for the default horizontal ratio before 9:16 is “mentioned as an afterthought,” leading to cropped subjects and broken composition
- Vague units (“shorter,” “vertical-ish”) → the model can only guess; if you need 8 seconds, write 8 seconds; if you need vertical, write 9:16
- Aspect ratio mismatched with content → a wide group scene forced into 9:16, with subjects squeezed into a single column; or vertical space unused in a portrait video, with empty margins top and bottom
- Declaring a duration beyond the model’s limit → truncated or rhythm compressed; long-form content should be generated in segments and then edited