vezo0 cr
← ALL TUTS

// GUIDE — SEEDANCE 2.5

How to prompt Seedance 2.5

Seedance 2.5 rewards structure. This is the working grammar — how to name a subject, bind a reference, stage thirty seconds, direct a camera, and write audio — with templates and examples you can paste straight into a render.

PUBLISHED
21 aug 2026
UPDATED
21 aug 2026
READ
18 min
SECTIONS
15
  • PROMPTING
  • SEEDANCE
  • BYTEDANCE
  • LONG-FORM
01

WHAT PROMPTING ACTUALLY CONTROLS

Prompt structure moves three things: how closely the model follows your instructions, how consistently it uses the material you give it, and how much of the result you can steer on purpose rather than by luck.

It does not move the model's ceiling. Image quality, human realism, physics, complex camera moves and clean cuts come from the model itself, from the material you feed it, and from the randomness in any given render. Everything below reduces ambiguity — which raises your hit rate. Nothing below is a guarantee.

Practical consequence: when a render misses, first ask whether the prompt was ambiguous. If it wasn't, re-roll instead of rewriting. Re-rolling a clear prompt beats endlessly editing a clear prompt.

02

THE CORE FORMULA

Every Seedance prompt is some combination of six slots. Write them in this order and the model reads them in the order it needs them.

  • Subject + action — who or what does what. This is the only mandatory slot.
  • Scene — place, time of day, weather, spatial layout, background state.
  • Look — light, colour, materials, texture, overall mood.
  • Camera — shot size, angle, movement, what it holds focus on, where it cuts.
  • Audio — dialogue, voice character, ambience, effects, music.
  • References — which uploaded material defines which of the above.
TEMPLATE — BASIC
<subject> does <primary action or event> in <scene>.
The look is <light, colour, texture, mood>.
Shoot it <shot size, angle, movement, cuts>.
Audio: <dialogue, ambience, effects, music>.
EXAMPLE — BASIC
A ceramicist finishes a pale blue cup at dawn, lifts it off the wheel,
and sets it in the middle of a wooden shelf.
Soft morning light through one window; wet clay still has a sheen; the
bench is tidy.
Open on a medium shot of the wheel, push in slowly on the cup's surface,
then cut to a frontal view of the shelf.
Audio: the low hum of the wheel, clay friction, quiet room tone.

Drop any slot you don't care about — an unfilled slot is the model's to choose, and it usually chooses reasonably. Do not put duration, aspect ratio or resolution in the prompt text; those are settings on the render form, and writing them into the prompt only adds noise.

Write in the positive

Describe what should be in frame, not what shouldn't. "An empty street at night" works; "a street with no people" invites people. The one exception is exclusions attached to a reference image, which are covered next — there, a negative is about material use, not about the frame.

03

REFERENCE MATERIALS AND THEIR ROLES

An uploaded image carries everything in it: the person, their clothes, the background, the composition, the lighting. If you only say "use this image," you have asked for all of it. Name the attributes you want and rule out the ones you don't.

TEMPLATE — REFERENCE ROLES
@Image 1 defines <subject>'s <face, clothing, structure, material>.
@Video 1 defines <motion, camera move, pacing>.
@Audio 1 defines <who or what>'s <voice, line, ambience, music>.

<subject> does <primary action> in <scene>.
The look is <style>, shot <camera treatment>.
EXAMPLE — REFERENCE ROLES
@Image 1 defines the ceramicist's face, hair and dark green apron.
Do not use its background.
@Image 2 defines the studio's bench, window position and morning light.
Do not use the people in it.
@Video 1 defines the pacing of throwing, lifting and setting down.
Do not use its performer, clothing or location.

The ceramicist finishes a pale blue cup at dawn and sets it on the shelf.
Open medium on the wheel, then push in on the cup's surface.
Audio: wheel hum, clay friction, room tone.

Bindings live in the prompt, not in the pictures. Text labels burned into an image are unreliable, and asking the model to infer which upload is which character is how two people end up wearing each other's coat.

How much material to give it

MATERIALHARD LIMITWORKS BEST AT
Images30, each up to 4K1–8 distinct subjects
Video10 clips, 30s combined1–5 subjects, 5–10s each
Audio10 clips, 30s combinedonly voices and ambience the shot needs
Editing1 source video + imagessource under 20s, 1–5 reference images

Fifty materials is the ceiling across all types. The right-hand column is where results stay stable — you can push past it (9–12 subjects from images, 6–8 references for an edit) at a falling success rate. Trim before you add: a reference that isn't doing a job is a reference that can leak.

Several views of one thing

When a product or face needs multiple angles, upload one view per image and say so explicitly. Separate images beat a single contact-sheet collage, which the model may read as several different objects.

EXAMPLE — MULTI-VIEW
@Image 1 is the front of the folding desk lamp.
@Image 2 is its left side.
@Image 3 is its right side.
@Image 4 is its back.
All four show one lamp. Exactly one lamp appears in the output.

If a reference video already carries the motion, camera and order you want, inherit it by name and stop describing the action in words. Re-describing a move the video already shows sets your sentence against your footage, and the model has to pick a winner.

04

AUDIO AND TEXT SYNTAX

Plain prose works. When a shot has music and effects and a line of dialogue at once, the brackets keep them apart.

FORUSEEXAMPLE
Music( )(soft rhythmic piano under the scene)
Sound effect< ><a bell rings somewhere far off>
Dialogue{ }{Hello — welcome back.}
On-screen text【 】【Chapter One: Departure】

Getting the language right

Non-English dialogue needs the language named before the line, or the model will read your text in whatever accent it defaults to. Same fix when English text comes back spoken in another language, or when you need a specific regional voice.

FORMULA — SPOKEN LANGUAGE
<language> + <region or accent> + <delivery> + <speaker> + {line}

Dialogue language: American English. The girl says it plainly, in
conversational American English: {I thought you weren't coming.}

Dialogue language: Japanese. She says it softly: {もう大丈夫です}

Keep lines short. A ten-second shot holds roughly two sentences of natural speech; pack in more and the delivery speeds up to fit, which reads as rushed rather than urgent.

05

WRITING FOR THIRTY SECONDS

Thirty seconds is the reason to reach for this model, and it is also the thing most prompts get wrong. A 30-second prompt written as one continuous description gives the model no idea what should have happened by second twelve. Split the clip into consecutive stages. Give each stage exactly one change of state, and finish each stage with what a viewer would see at that moment.

The end state is the load-bearing part. It is what the next stage continues from, and it is what stops a prop from teleporting between hands.

TEMPLATE — STAGED CLIP
[Goal]
A <type of video>. Subject: <subject>. Story: <one-line summary>.

[Stage 1]
Start state: <where the people, props and scene are>.
Event: <one change>.
End state: <what is visible when this stage lands>.

[Stage 2]
Carried over: <what must not change>.
Event: <one change>.
End state: <what is visible>.

[Stage 3]
Event: <closing change>.
End state: <final visible frame>.

[Hold constant]
<identities, headcount, clothing, who holds what, screen direction, audio>.
EXAMPLE — STAGED CLIP
[Goal]
A how-it's-made clip: a florist and a shop assistant build, wrap and
hand off one bouquet.

[Stage 1]
Start state: the florist stands behind the bench. Loose stems, shears
and wrapping paper on the top.
Event: she sorts the stems and trims them to length.
End state: the bouquet is in her left hand; the shears are back on the
right of the bench.

[Stage 2]
Carried over: both people keep their faces and clothes; the bouquet
stays in her left hand.
Event: the assistant unfolds the paper, she lays the bouquet in and
ties it with green ribbon.
End state: the wrapped bouquet lies flat and centred on the bench, bow
facing camera.

[Stage 3]
Event: the assistant lifts the bouquet onto the pickup shelf.
End state: the bouquet sits centred on the shelf; both stand behind the
bench looking at it.

[Hold constant]
Both identities and outfits, bench orientation, shears position, and who
is holding the bouquet at each point.

When to reach for timestamps

Stages are the default. Use clock times only when a specific beat has to land at a specific moment — a handoff, an entrance, a transition, a cut on a sound.

PATTERNWRITE IT AS
Range0–3s … 3–7s … 7–12s …
PointAt 5s the camera whip-pans left and the scene changes.
RelativeThree seconds after he hits the switch, the lights fade.
EXAMPLE — TIMED SEQUENCE
0–5s: an empty wooden table. A hand sets down a white ceramic plate.
End state: the hand is out of frame; only the plate remains, centred.
5–10s: the plate is removed and a clear glass is set down.
End state: only the glass remains, centred.
10–15s: the glass is removed and a green vase is set down.
End state: only the vase remains, centred.

Ranges must run consecutively and must not overlap. Treat a range as a budget rather than an edit point — an action may land slightly either side of a boundary. Underfill a range and the model invents filler; overfill it and you get frantic cutting or dropped events. Never ask for a rate, as in "three actions in one second."

06

CASTING WITH MANY REFERENCES

With up to fifty materials in play, the job stops being description and becomes casting. Work in four passes: bind each subject one at a time, group the bindings by type, give recurring characters a profile, then choose materials scene by scene.

Pass 1 — one binding per line

TEMPLATE — BINDINGS
<Character A> is @Image 1. Use its face, hair and clothing only.
<Character B> is @Image 2. Use its face, hair and clothing only.
<Prop A> is @Image 3. Use its structure, material and colour only.
<Scene A> is @Image 4. Use its layout, architecture and light only.
Do not use the people in it.

Never write "@Images 1–4 define the four characters." That sentence contains no mapping — it tells the model there are four of each and leaves the pairing to chance.

Pass 2 — group by type

TEMPLATE — GROUPED CAST
[Cast]
<Conservator> is @Image 1 — face, hair, clothing.
<Registrar> is @Image 2 — face, hair, clothing.
<Installer> is @Image 3 — face, hair, clothing.
<Guide> is @Image 4 — face, hair, clothing.
Never swap their faces, clothes, actions, positions or lines.

[Props]
<Sample case> is @Image 5 and belongs to <Conservator> only.
<Record board> is @Image 6 and belongs to <Registrar> only.

[Locations]
<Lab> is @Image 7 — space, materials, light.
<Gallery> is @Image 8 — space, materials, light.

[Motion and sound]
@Video 1 is how <Conservator> opens <Sample case>. Not its performer or
location.
@Audio 1 is <Guide>'s voice and lines.

Pass 3 — a profile for anyone who recurs

TEMPLATE — SUBJECT PROFILE
[Profile — Conservator]
Look: @Image 1.
Always carries: <Sample case> (@Image 5).
Appears in: <Lab>, <Gallery>.
Moves like: the case-opening in @Video 1, the placement in @Video 2.
Never: wears another character's clothes; holds <Record board>.

Pass 4 — pick per scene

EXAMPLE — SCENE SELECTION
Scene 1 | Lab inspection
Uses: <Conservator>, <Sample case>, <Lab>, the opening motion from @Video 1.
Event: she opens <Sample case> at the bench and inspects what's inside.
End state: she stays on the inner side of the bench; the case sits by her
right hand, frame left.

Scene 2 | Gallery registration
Uses: <Registrar>, <Record board>, <Gallery>.
Event: he checks the number on <Record board> beside the display case.
End state: he still holds the board in both hands; nobody else enters the
display-case area.

Written this way, an unmentioned material sits out that scene instead of crowding into it. That is the whole point — you are telling the model what to leave in the drawer.

07

EDITING AN EXISTING VIDEO

An edit is not a new generation with extra steps. Declare the source clip as the single authority for everything, then carve out the smallest possible exception. Four blocks: goal, master, scope, preserve.

TEMPLATE — GENERAL EDIT
[Goal]
Edit @Video 1. In <whole clip | time range>, <add | remove | replace |
adjust> <object, region or sound>.

[Master]
@Video 1 is the sole master. It defines the people, scene, action,
framing, camera, occlusion, audio and event order.

[Target]
@Image 1 / @Audio 1 defines <the specific attributes of the new thing>.

[Scope]
Change only <object, region, time range or audio category>.

[Preserve]
Everything else from @Video 1 — <the specifics that must not move>.
EXAMPLE — RELIGHT ONE WALL
[Goal]
Edit @Video 1. Between 4s and 7s only, change the cool blue light on the
right-hand wall to warm orange.

[Master]
@Video 1 is the sole master — character, room, action, framing, camera,
audio and event order.

[Scope]
Only the wall's light colour and the area it falls on. Let her skin tone
respond naturally to the new light.

[Preserve]
Her identity, clothes, expression, position and movement; the room's
structure; the camera move; the dialogue and room tone.

Swapping a subject

A replacement has to inherit a timeline, not just a shape. State the count explicitly — "exactly one lamp, for the whole clip" — because an unstated count is how you get two.

EXAMPLE — OBJECT SWAP
[Goal]
Edit @Video 1. Replace only the yellow folding desk lamp with the white
one in @Image 1.

[Master]
@Video 1 is the sole master — desk, books, hands, camera position and
move, occlusion, event order.

[Target]
@Image 1 gives the white lamp's shape, structure and material only. Not
its background, framing or other objects.

[Scope]
Exactly one white lamp throughout. Replace only the yellow lamp. Leave
the books, desk, hands and background alone.

[Timeline]
The white lamp inherits every entrance, arm rotation, hand occlusion and
exit of the yellow one — same timings, same path, same speed changes.
Apart from that object, nothing else in @Video 1 changes.

Swapping a background

EXAMPLE — BACKGROUND SWAP
@Video 1 is the sole master — people, action, framing, camera, order.
@Image 1 supplies only a daylit glass greenhouse: its layout, depth,
ambient colour and light direction. Not the people in it.

Replace only the light grey background outside her silhouette in
@Video 1 with that greenhouse.
Keep her identity, features, hair, clothing, expression, position, scale
and the arm-raising movement exactly as they are.
Apart from that region, nothing else in @Video 1 changes.

Editing sound alone

Dialogue, language, voice, score and effects are separately addressable. Name the category, the change, and what stays.

EXAMPLE — AUDIO EDITS
Edit @Video 1. Remove only the background music. Keep the dialogue and
its lip sync, the room tone and the action effects, and keep the picture,
camera and cutting rhythm as they are.

Edit @Video 1. Change the presenter's spoken language to natural American
English, keeping the same words and the same speaking times. Every other
voice, the music, the ambience and the picture stay as they are.
08

EXTENDING A VIDEO

An extension has one hard problem: the join. Describe the boundary frame before you describe anything new, and describe it as a state — pose, facing, prop positions, framing, light, direction of travel — not as "continue from the end."

TEMPLATE — EXTEND FORWARD
@Video 1 is the clip to extend forward.

Extend @Video 1. Its first new frame continues straight from @Video 1's
last frame: keep <pose and facing>, <prop positions>, <background and
spatial relations>, <camera position and framing>, <light>, and
<direction of movement>.

Then: <the new action, event, camera move or sound>.

Across the extension keep <identity and clothing>, <key props>,
<background layout> and <screen direction> continuous. Each subject stays
one continuous instance — never duplicated, split, or changing part count.
EXAMPLE — EXTEND FORWARD
@Video 1 is the clip to extend forward.

Extend @Video 1. Its first new frame continues straight from @Video 1's
last frame: same locked-off medium shot, the orange paper plane in the
same place and attitude, the classroom window behind, afternoon light,
still travelling frame right.

Then the plane glides on to the right and out of frame while the white
curtain by the window sways once. Camera and classroom stay exactly as
the last frame left them.

Adding fresh references to an extension is fine — a face, an outfit, a prop — as long as you say plainly that the source video still owns the opening frame. Otherwise a new reference image will try to restage the join.

Extending backwards

Going the other way, the source video's first frame becomes your target end state, and it has to be written out in full. "Then it connects to the original" is the sentence that drags later characters into the opening seconds, or lets the picture drift after it has already arrived.

EXAMPLE — EXTEND BACKWARD
@Video 1 is the clip to extend backward.

Before @Video 1 begins: the same glass greenhouse, empty. Morning mist
drifts near the floor, the overhead shade rises slowly, nobody is in yet.

The extension's last frame lands on @Video 1's first frame: the central
aisle, planting tables both sides, glass frame, soft morning light, the
same locked-off wide. By then the shade is fully up, the aisle is empty,
and the leaves still move slightly.

Anything that only appears after @Video 1 starts must not appear early
here.

A seam that reads as continuous is the target; identical pixels are not. Review both sides of the join and then watch the whole extension — drift usually shows up a second or two past the boundary, not at it.

09

KEYFRAMES, STORYBOARDS AND BLOCKOUTS

First and last frame

Declare the anchors on their own lines — one role per image. Extra references can still define a face or a material, as long as you say they must not restage the anchors. The output takes its aspect ratio from the first image, so give both anchors the same ratio or the last frame gets stretched to fit.

EXAMPLE — FIRST + LAST
@Image 1 is the first frame: opening framing, positions, poses, tabletop
state, workshop, camera direction.
@Image 2 is the last frame: closing framing, positions, poses, tabletop
state, workshop, camera direction.
@Image 3 is the perfumer's face, hair and dark green apron. It must not
change the framing set by @Image 1 or @Image 2.
@Image 4 is the glass bottle's shape, material and label position. Same
restriction.

From the opening pose, the perfumer picks up a dropper and the bottle,
drips amber oil in, swirls it, seats the stopper, sets the finished bottle
in the middle of the table, and arrives naturally at @Image 2.
Throughout: her identity and clothes, the bottle count and shape, the
table layout, the warm side light and the camera direction all hold.

A run of keyframes

More than two anchors works the same way: state the order once, then say what each image is the state of. Keyframes govern stage order and the look of each landing point — they do not reproduce themselves frame for frame.

EXAMPLE — KEYFRAME RUN
Use @Image 1 through @Image 4 as keyframes, in that order.

@Image 1 is the first frame: an orange paper plane on the left of a
classroom desk, nose pointing right, locked-off medium.
@Image 2 is the second state: a hand lifts that same plane off the desk,
attitude unchanged.
@Image 3 is the third state: the same plane passes the window as the
curtain shifts slightly right.
@Image 4 is the last frame: the same plane at rest on the middle shelf of
the bookcase, still pointing right.

Move through those four states in order, with flight direction and speed
continuous between them. Hold the plane's orange paper, size and folds,
the classroom layout, the afternoon side light and the camera axis.

Storyboard grids

A grid communicates order and rough framing — not detail. Keep it under about fifteen panels, keep the drawing to clean line art, and keep text out of it. Then say how to read the grid, what to ignore about its style, and what each panel actually contains.

EXAMPLE — STORYBOARD GRID
@Image 1 is a four-panel storyboard for shot order and rough framing.
Read it left to right, top to bottom. Ignore its line-art style and any
labels in it.
@Image 2 is the potter's face, short hair and dark grey apron.
@Image 3 is the blue-glazed cup's proportions, glaze and curved handle.

Shot 1: wide — a quiet studio, the potter seated at the wheel.
Shot 2: side medium — both hands shaping the turning clay, the wall rising.
Shot 3: close — fingers refining the rim and the handle join, slip moving
over the fingertips.
Shot 4: medium close — the fired cup on a wooden shelf as both hands
withdraw.

Finish it as realistic documentary. Audio: wheel rotation, wet clay, the
room.

Blockouts

Decide first which kind you have. A coarse blockout is grey geometry that carries timing and space; a fine blockout is finished modelling that needs skin. They take opposite instructions.

COARSEFINE
Carriespaths, blocking, camera, cuts, light changes, rhythmfull structure, action, layout, camera
You supplywho and what each shape becomesmaterials, colour, character, scene, style
Material must besimple shapes with clear relationsclean model — no path lines, axes or camera frustums
Prompt doesmap every shape, name what to inheritpreserve structure and motion, define the re-render
EXAMPLE — COARSE BLOCKOUT
@Video 1 is a coarse blockout. Take from it only the walking path, the
cart's direction, the locked-off camera, the one push-in and the two
cuts. Not its grey geometry or empty room.
The tall cylinder is <Guide>.
The rectangular block is <Display cart>.
@Image 1 is <Guide>'s face, blue uniform and name badge.
@Image 2 is the cart's white metal frame and clear cover.
@Image 3 is the gallery: curved walls, grey floor, overhead strip lights.

<Guide> pushes <Display cart> along the curved wall, stops at the central
display and opens the cover. Keep the path, the blocking, the push-in
direction and the cut points from @Video 1.
Bright realistic museum documentary. Audio: footsteps, castors, room tone.
EXAMPLE — FINE BLOCKOUT
@Video 1 is a fine blockout. Preserve the sculpture's full structure, the
three-ring rotation, the plinth position, the orbiting camera and the
cuts. Not its grey material or empty background.
@Image 1 is the outer ring's brushed brass.
@Image 2 is the inner blades' translucent blue glass.
@Image 3 is a contemporary gallery: white curved walls, dark grey floor,
soft overhead light.

Re-render the rings as a brass and blue-glass kinetic sculpture, and the
room as that gallery. Keep the structure, the rotation rhythm, the orbit
and the cuts. Audio: a low mechanical turning, quiet interior.
10

ASSEMBLY AND TRANSITIONS

One-click video from stills

"Make these into a video" is not an instruction. Five things need saying: what each image is for, what order they run in, how much each one moves, how it's cut and finished, and what it sounds like.

EXAMPLE — STILLS TO FILM
[Roles]
@Image 1 — night-market entrance, the opening.
@Image 2 — the traveller walking the street.
@Image 3 — the lantern stall and its detail.
@Image 4 — three friends eating.
@Image 5 — the river at night, reflections.
@Image 6 — the closing group photo on the bridge.
@Video 1 — its cutting rhythm, sticker treatment and transitions only.
Not its people or places.

[Order]
Run @Image 1 to @Image 6 in that order: arrive, wander, eat, walk the
river, take the photo. The three friends keep their faces and clothes
throughout; never mix them up.

[Motion]
Slow push-ins and slight parallax on the environments. On the people,
only blinking, a head turn, a raised glass, a little cloth movement.
Stall structure, table position and bridge railing stay put.

[Finish]
Upbeat travel-cut rhythm, scenes joined on occlusion and matching colour,
hand-drawn stickers kept to the edges of frame.

[Audio]
Market chatter, light crockery, river wind, and unobtrusive upbeat
instrumental.

Seamless transitions between two clips

A bridge shot needs seven things: which clip is before, which is after, what triggers the change, what the camera does, what visually morphs into what, where it arrives, and how the sound crosses over.

EXAMPLE — BRIDGE SHOT
@Video 1 is the outgoing clip — take its rainy street, red umbrella,
slow push-in and rain.
@Video 2 is the incoming clip — take its round gallery skylight, rising
camera and quiet interior reverb.
Both clips keep their people, spaces, structures and main actions.

At the end of @Video 1 the red umbrella comes toward camera and fills the
frame — that starts the transition.
The camera keeps moving forward. The umbrella's circular edge becomes the
skylight's metal ring; the red fabric becomes white daylight through the
glass.
It lands on @Video 2's opening low-angle framing, the move changing from
forward travel into a rise.
The rain thins into footsteps echoing inside the gallery.
TRANSITIONWHAT TO SPELL OUT
Dive / reversecamera direction, speed change, when the next space begins
Subject turnpose, rotation direction, how clothing and backdrop change with it
Foreground wipewhen the object fills frame, and the framing waiting behind it
Object morphwhich shapes correspond, which materials, how the change happens
Push or rackthe move, the focus target, the spatial link across the cut
11

DIRECTING PERFORMANCE

"Tense," "warm" and "oppressive" set a direction and leave the acting open. To close the gap, give observable cues — eyes, brows, mouth, breath, where they look, what the hands do. Two to four cues carry a single emotional turn; you do not need a full anatomy of the face.

TEMPLATE — ONE TURN
The feeling moves from <start> to <end>.
After <trigger>, <subject> first <the immediate, involuntary reaction>.
Then <eyes, brows, mouth, breath, gaze or hands> gradually <change>.
Finally <subject> shows <target feeling> by <what they outwardly do>.
EXAMPLE — MULTI-STAGE
Applause reaches him from behind the stage. His fingers stop on the
programme, his eyes turn slowly toward the curtain, his shoulders stay
locked.
Once it's clear the curtain call is over, he lets out a breath. The
shoulders drop, a held-back smile arrives, his eyes slowly fill — and he
still doesn't turn to leave.

For several changes in one clip, hang each on its trigger — heard, seen, or confirmed — instead of listing feelings in sequence. An emotion with a cause reads; an emotion on a schedule doesn't.

12

CAMERA LANGUAGE

Ordinary film vocabulary goes in as-is. The model knows shot sizes, moves and angles.

TERMS THAT LAND
Shot sizeextreme wide, wide, medium, close-up, extreme close-up
Movementpush in, pull out, pan, track, follow, orbit, dive, dolly out, tilt up, handheld
Anglelow angle, overhead, first person
Techniquesone-take, dolly zoom, aerial, FPV, bullet time, handheld, speed ramp

With more than one subject in frame, every move still needs an object: which subject the camera follows or orbits, where the move starts and where it stops. A dolly zoom needs to know whose scale to hold and whether the background should rush in or away. Bullet time needs to know what freezes and which way the camera travels around it.

Terms the model may not know

For niche or contested vocabulary, keep the term and immediately restate it as a visible change — subject, what changes, foreground against background, direction or speed.

EXAMPLES — CAMERA
Rack focus: pull focus from the leaves in front to the woman behind.
The leaves soften as her face comes sharp.

Shallow portrait: keep the pastry chef's eyes and face sharp while the
jars and lights behind her fall into soft round bokeh.

Tracking: move sideways at the skater's speed, holding him sharp while
the roadside wall smears into horizontal blur, right to left.

Golden hour: warm low sun from behind and left of the hiker, throwing
long shadows across the ridge.

Natural vignette: darken the four corners gradually, keeping the centre
brightness and her skin tone natural — no hard black border.

Whip pan: at 5s, snap the camera left; cut when the foreground shelving
covers frame, and keep moving left at the same speed in the next scene.

Aperture, focal length and shutter values are allowed, but the visible result you want is almost always the clearer instruction. "f/1.4" is a hope; "background falls into soft round bokeh" is a description.

13

LOCKED PARAMETERS

Some tasks take their parameters from the material rather than from the form. Knowing which saves you from writing a setting the model was never going to honour.

TASKASPECT RATIODURATION
Text or image to videoyours to setyours to set
Video editlocked to the sourcelocked to the source, ±0.3s
First / first + last framelocked to the first imageyours to set
Extensionlocked to the sourceyours to set

A locked value cannot be overridden on the form or through the API. Two consequences worth remembering: mismatched first and last images stretch the last frame, and an extended segment can sit at a slightly different loudness from its source.

14

BEFORE YOU RENDER

  • The subject and the main action are stated in one readable sentence.
  • Every reference says what to take from it — and what to ignore.
  • Every distinct person, product and prop is named and bound to exactly one material.
  • Materials are chosen per scene, not all dumped into every scene.
  • Each stage carries one change and ends on a visible state.
  • Headcount, clothing, who-holds-what and screen direction are pinned across stages.
  • Time ranges run consecutively, don't overlap, and don't demand a rate.
  • Emotions and unusual camera terms are paired with something you could actually see or hear.
  • First and last images share an aspect ratio; each anchor image has one role.
  • For an edit: one master, an explicit scope, a stated count, and a preserve list.
  • For an extension: the boundary state is written out before any new action.
  • For assembly: roles, order, amount of motion, finish and audio are all specified.
  • Nothing in the prompt tries to set a parameter the task has already locked.
15

WHERE PROMPTING STOPS

  • Timestamps budget time for events. They are not frame-accurate edit points.
  • An edit prompt raises the odds that key events line up with the source. It cannot promise frame-for-frame overlap.
  • Multi-reference work is about choosing and combining the right material — not about making everything appear at once.
  • Subtitles, formulas, signage, spec numbers and exact frame timing still want a real edit afterwards. Generate the shot, finish the text in post.
  • A generated transition is a bridge, not a splice; it will not preserve both source clips pixel for pixel.
  • Visual quality, human realism and physical plausibility come from the model. Structure raises your hit rate; it does not raise the ceiling.

Everything here is technique, not warranty. The same prompt renders differently on different days — that is the medium, and the reason to draft short, re-roll cheaply, and only commit the full thirty seconds to a structure that already reads.