// GUIDE — SEEDANCE 2.5
How to prompt Seedance 2.5
Seedance 2.5 rewards structure. This is the working grammar — how to name a subject, bind a reference, stage thirty seconds, direct a camera, and write audio — with templates and examples you can paste straight into a render.
- PUBLISHED
- 21 aug 2026
- UPDATED
- 21 aug 2026
- READ
- 18 min
- SECTIONS
- 15
- PROMPTING
- SEEDANCE
- BYTEDANCE
- LONG-FORM
CONTENTS
— jump to a section- 01WHAT PROMPTING ACTUALLY CONTROLS
- 02THE CORE FORMULA
- 03REFERENCE MATERIALS AND THEIR ROLES
- 04AUDIO AND TEXT SYNTAX
- 05WRITING FOR THIRTY SECONDS
- 06CASTING WITH MANY REFERENCES
- 07EDITING AN EXISTING VIDEO
- 08EXTENDING A VIDEO
- 09KEYFRAMES, STORYBOARDS AND BLOCKOUTS
- 10ASSEMBLY AND TRANSITIONS
- 11DIRECTING PERFORMANCE
- 12CAMERA LANGUAGE
- 13LOCKED PARAMETERS
- 14BEFORE YOU RENDER
- 15WHERE PROMPTING STOPS
WHAT PROMPTING ACTUALLY CONTROLS
— set expectations before you write a wordPrompt structure moves three things: how closely the model follows your instructions, how consistently it uses the material you give it, and how much of the result you can steer on purpose rather than by luck.
It does not move the model's ceiling. Image quality, human realism, physics, complex camera moves and clean cuts come from the model itself, from the material you feed it, and from the randomness in any given render. Everything below reduces ambiguity — which raises your hit rate. Nothing below is a guarantee.
Practical consequence: when a render misses, first ask whether the prompt was ambiguous. If it wasn't, re-roll instead of rewriting. Re-rolling a clear prompt beats endlessly editing a clear prompt.
THE CORE FORMULA
— six slots, all but the first two optionalEvery Seedance prompt is some combination of six slots. Write them in this order and the model reads them in the order it needs them.
- Subject + action — who or what does what. This is the only mandatory slot.
- Scene — place, time of day, weather, spatial layout, background state.
- Look — light, colour, materials, texture, overall mood.
- Camera — shot size, angle, movement, what it holds focus on, where it cuts.
- Audio — dialogue, voice character, ambience, effects, music.
- References — which uploaded material defines which of the above.
<subject> does <primary action or event> in <scene>.
The look is <light, colour, texture, mood>.
Shoot it <shot size, angle, movement, cuts>.
Audio: <dialogue, ambience, effects, music>.A ceramicist finishes a pale blue cup at dawn, lifts it off the wheel,
and sets it in the middle of a wooden shelf.
Soft morning light through one window; wet clay still has a sheen; the
bench is tidy.
Open on a medium shot of the wheel, push in slowly on the cup's surface,
then cut to a frontal view of the shelf.
Audio: the low hum of the wheel, clay friction, quiet room tone.Drop any slot you don't care about — an unfilled slot is the model's to choose, and it usually chooses reasonably. Do not put duration, aspect ratio or resolution in the prompt text; those are settings on the render form, and writing them into the prompt only adds noise.
Write in the positive
Describe what should be in frame, not what shouldn't. "An empty street at night" works; "a street with no people" invites people. The one exception is exclusions attached to a reference image, which are covered next — there, a negative is about material use, not about the frame.
REFERENCE MATERIALS AND THEIR ROLES
— say what each upload is for, and what it is not forAn uploaded image carries everything in it: the person, their clothes, the background, the composition, the lighting. If you only say "use this image," you have asked for all of it. Name the attributes you want and rule out the ones you don't.
@Image 1 defines <subject>'s <face, clothing, structure, material>.
@Video 1 defines <motion, camera move, pacing>.
@Audio 1 defines <who or what>'s <voice, line, ambience, music>.
<subject> does <primary action> in <scene>.
The look is <style>, shot <camera treatment>.@Image 1 defines the ceramicist's face, hair and dark green apron.
Do not use its background.
@Image 2 defines the studio's bench, window position and morning light.
Do not use the people in it.
@Video 1 defines the pacing of throwing, lifting and setting down.
Do not use its performer, clothing or location.
The ceramicist finishes a pale blue cup at dawn and sets it on the shelf.
Open medium on the wheel, then push in on the cup's surface.
Audio: wheel hum, clay friction, room tone.Bindings live in the prompt, not in the pictures. Text labels burned into an image are unreliable, and asking the model to infer which upload is which character is how two people end up wearing each other's coat.
How much material to give it
| MATERIAL | HARD LIMIT | WORKS BEST AT |
|---|---|---|
| Images | 30, each up to 4K | 1–8 distinct subjects |
| Video | 10 clips, 30s combined | 1–5 subjects, 5–10s each |
| Audio | 10 clips, 30s combined | only voices and ambience the shot needs |
| Editing | 1 source video + images | source under 20s, 1–5 reference images |
Fifty materials is the ceiling across all types. The right-hand column is where results stay stable — you can push past it (9–12 subjects from images, 6–8 references for an edit) at a falling success rate. Trim before you add: a reference that isn't doing a job is a reference that can leak.
Several views of one thing
When a product or face needs multiple angles, upload one view per image and say so explicitly. Separate images beat a single contact-sheet collage, which the model may read as several different objects.
@Image 1 is the front of the folding desk lamp.
@Image 2 is its left side.
@Image 3 is its right side.
@Image 4 is its back.
All four show one lamp. Exactly one lamp appears in the output.If a reference video already carries the motion, camera and order you want, inherit it by name and stop describing the action in words. Re-describing a move the video already shows sets your sentence against your footage, and the model has to pick a winner.
AUDIO AND TEXT SYNTAX
— four brackets that stop audio from blurring togetherPlain prose works. When a shot has music and effects and a line of dialogue at once, the brackets keep them apart.
| FOR | USE | EXAMPLE |
|---|---|---|
| Music | ( ) | (soft rhythmic piano under the scene) |
| Sound effect | < > | <a bell rings somewhere far off> |
| Dialogue | { } | {Hello — welcome back.} |
| On-screen text | 【 】 | 【Chapter One: Departure】 |
Getting the language right
Non-English dialogue needs the language named before the line, or the model will read your text in whatever accent it defaults to. Same fix when English text comes back spoken in another language, or when you need a specific regional voice.
<language> + <region or accent> + <delivery> + <speaker> + {line}
Dialogue language: American English. The girl says it plainly, in
conversational American English: {I thought you weren't coming.}
Dialogue language: Japanese. She says it softly: {もう大丈夫です}Keep lines short. A ten-second shot holds roughly two sentences of natural speech; pack in more and the delivery speeds up to fit, which reads as rushed rather than urgent.
WRITING FOR THIRTY SECONDS
— stages and end states, not one long paragraphThirty seconds is the reason to reach for this model, and it is also the thing most prompts get wrong. A 30-second prompt written as one continuous description gives the model no idea what should have happened by second twelve. Split the clip into consecutive stages. Give each stage exactly one change of state, and finish each stage with what a viewer would see at that moment.
The end state is the load-bearing part. It is what the next stage continues from, and it is what stops a prop from teleporting between hands.
[Goal]
A <type of video>. Subject: <subject>. Story: <one-line summary>.
[Stage 1]
Start state: <where the people, props and scene are>.
Event: <one change>.
End state: <what is visible when this stage lands>.
[Stage 2]
Carried over: <what must not change>.
Event: <one change>.
End state: <what is visible>.
[Stage 3]
Event: <closing change>.
End state: <final visible frame>.
[Hold constant]
<identities, headcount, clothing, who holds what, screen direction, audio>.[Goal]
A how-it's-made clip: a florist and a shop assistant build, wrap and
hand off one bouquet.
[Stage 1]
Start state: the florist stands behind the bench. Loose stems, shears
and wrapping paper on the top.
Event: she sorts the stems and trims them to length.
End state: the bouquet is in her left hand; the shears are back on the
right of the bench.
[Stage 2]
Carried over: both people keep their faces and clothes; the bouquet
stays in her left hand.
Event: the assistant unfolds the paper, she lays the bouquet in and
ties it with green ribbon.
End state: the wrapped bouquet lies flat and centred on the bench, bow
facing camera.
[Stage 3]
Event: the assistant lifts the bouquet onto the pickup shelf.
End state: the bouquet sits centred on the shelf; both stand behind the
bench looking at it.
[Hold constant]
Both identities and outfits, bench orientation, shears position, and who
is holding the bouquet at each point.When to reach for timestamps
Stages are the default. Use clock times only when a specific beat has to land at a specific moment — a handoff, an entrance, a transition, a cut on a sound.
| PATTERN | WRITE IT AS |
|---|---|
| Range | 0–3s … 3–7s … 7–12s … |
| Point | At 5s the camera whip-pans left and the scene changes. |
| Relative | Three seconds after he hits the switch, the lights fade. |
0–5s: an empty wooden table. A hand sets down a white ceramic plate.
End state: the hand is out of frame; only the plate remains, centred.
5–10s: the plate is removed and a clear glass is set down.
End state: only the glass remains, centred.
10–15s: the glass is removed and a green vase is set down.
End state: only the vase remains, centred.Ranges must run consecutively and must not overlap. Treat a range as a budget rather than an edit point — an action may land slightly either side of a boundary. Underfill a range and the model invents filler; overfill it and you get frantic cutting or dropped events. Never ask for a rate, as in "three actions in one second."
CASTING WITH MANY REFERENCES
— the goal is the right material per scene, not all of it at onceWith up to fifty materials in play, the job stops being description and becomes casting. Work in four passes: bind each subject one at a time, group the bindings by type, give recurring characters a profile, then choose materials scene by scene.
Pass 1 — one binding per line
<Character A> is @Image 1. Use its face, hair and clothing only.
<Character B> is @Image 2. Use its face, hair and clothing only.
<Prop A> is @Image 3. Use its structure, material and colour only.
<Scene A> is @Image 4. Use its layout, architecture and light only.
Do not use the people in it.Never write "@Images 1–4 define the four characters." That sentence contains no mapping — it tells the model there are four of each and leaves the pairing to chance.
Pass 2 — group by type
[Cast]
<Conservator> is @Image 1 — face, hair, clothing.
<Registrar> is @Image 2 — face, hair, clothing.
<Installer> is @Image 3 — face, hair, clothing.
<Guide> is @Image 4 — face, hair, clothing.
Never swap their faces, clothes, actions, positions or lines.
[Props]
<Sample case> is @Image 5 and belongs to <Conservator> only.
<Record board> is @Image 6 and belongs to <Registrar> only.
[Locations]
<Lab> is @Image 7 — space, materials, light.
<Gallery> is @Image 8 — space, materials, light.
[Motion and sound]
@Video 1 is how <Conservator> opens <Sample case>. Not its performer or
location.
@Audio 1 is <Guide>'s voice and lines.Pass 3 — a profile for anyone who recurs
[Profile — Conservator]
Look: @Image 1.
Always carries: <Sample case> (@Image 5).
Appears in: <Lab>, <Gallery>.
Moves like: the case-opening in @Video 1, the placement in @Video 2.
Never: wears another character's clothes; holds <Record board>.Pass 4 — pick per scene
Scene 1 | Lab inspection
Uses: <Conservator>, <Sample case>, <Lab>, the opening motion from @Video 1.
Event: she opens <Sample case> at the bench and inspects what's inside.
End state: she stays on the inner side of the bench; the case sits by her
right hand, frame left.
Scene 2 | Gallery registration
Uses: <Registrar>, <Record board>, <Gallery>.
Event: he checks the number on <Record board> beside the display case.
End state: he still holds the board in both hands; nobody else enters the
display-case area.Written this way, an unmentioned material sits out that scene instead of crowding into it. That is the whole point — you are telling the model what to leave in the drawer.
EDITING AN EXISTING VIDEO
— name the master, the scope, and everything that must surviveAn edit is not a new generation with extra steps. Declare the source clip as the single authority for everything, then carve out the smallest possible exception. Four blocks: goal, master, scope, preserve.
[Goal]
Edit @Video 1. In <whole clip | time range>, <add | remove | replace |
adjust> <object, region or sound>.
[Master]
@Video 1 is the sole master. It defines the people, scene, action,
framing, camera, occlusion, audio and event order.
[Target]
@Image 1 / @Audio 1 defines <the specific attributes of the new thing>.
[Scope]
Change only <object, region, time range or audio category>.
[Preserve]
Everything else from @Video 1 — <the specifics that must not move>.[Goal]
Edit @Video 1. Between 4s and 7s only, change the cool blue light on the
right-hand wall to warm orange.
[Master]
@Video 1 is the sole master — character, room, action, framing, camera,
audio and event order.
[Scope]
Only the wall's light colour and the area it falls on. Let her skin tone
respond naturally to the new light.
[Preserve]
Her identity, clothes, expression, position and movement; the room's
structure; the camera move; the dialogue and room tone.Swapping a subject
A replacement has to inherit a timeline, not just a shape. State the count explicitly — "exactly one lamp, for the whole clip" — because an unstated count is how you get two.
[Goal]
Edit @Video 1. Replace only the yellow folding desk lamp with the white
one in @Image 1.
[Master]
@Video 1 is the sole master — desk, books, hands, camera position and
move, occlusion, event order.
[Target]
@Image 1 gives the white lamp's shape, structure and material only. Not
its background, framing or other objects.
[Scope]
Exactly one white lamp throughout. Replace only the yellow lamp. Leave
the books, desk, hands and background alone.
[Timeline]
The white lamp inherits every entrance, arm rotation, hand occlusion and
exit of the yellow one — same timings, same path, same speed changes.
Apart from that object, nothing else in @Video 1 changes.Swapping a background
@Video 1 is the sole master — people, action, framing, camera, order.
@Image 1 supplies only a daylit glass greenhouse: its layout, depth,
ambient colour and light direction. Not the people in it.
Replace only the light grey background outside her silhouette in
@Video 1 with that greenhouse.
Keep her identity, features, hair, clothing, expression, position, scale
and the arm-raising movement exactly as they are.
Apart from that region, nothing else in @Video 1 changes.Editing sound alone
Dialogue, language, voice, score and effects are separately addressable. Name the category, the change, and what stays.
Edit @Video 1. Remove only the background music. Keep the dialogue and
its lip sync, the room tone and the action effects, and keep the picture,
camera and cutting rhythm as they are.
Edit @Video 1. Change the presenter's spoken language to natural American
English, keeping the same words and the same speaking times. Every other
voice, the music, the ambience and the picture stay as they are.EXTENDING A VIDEO
— align the seam first, then describe the new partAn extension has one hard problem: the join. Describe the boundary frame before you describe anything new, and describe it as a state — pose, facing, prop positions, framing, light, direction of travel — not as "continue from the end."
@Video 1 is the clip to extend forward.
Extend @Video 1. Its first new frame continues straight from @Video 1's
last frame: keep <pose and facing>, <prop positions>, <background and
spatial relations>, <camera position and framing>, <light>, and
<direction of movement>.
Then: <the new action, event, camera move or sound>.
Across the extension keep <identity and clothing>, <key props>,
<background layout> and <screen direction> continuous. Each subject stays
one continuous instance — never duplicated, split, or changing part count.@Video 1 is the clip to extend forward.
Extend @Video 1. Its first new frame continues straight from @Video 1's
last frame: same locked-off medium shot, the orange paper plane in the
same place and attitude, the classroom window behind, afternoon light,
still travelling frame right.
Then the plane glides on to the right and out of frame while the white
curtain by the window sways once. Camera and classroom stay exactly as
the last frame left them.Adding fresh references to an extension is fine — a face, an outfit, a prop — as long as you say plainly that the source video still owns the opening frame. Otherwise a new reference image will try to restage the join.
Extending backwards
Going the other way, the source video's first frame becomes your target end state, and it has to be written out in full. "Then it connects to the original" is the sentence that drags later characters into the opening seconds, or lets the picture drift after it has already arrived.
@Video 1 is the clip to extend backward.
Before @Video 1 begins: the same glass greenhouse, empty. Morning mist
drifts near the floor, the overhead shade rises slowly, nobody is in yet.
The extension's last frame lands on @Video 1's first frame: the central
aisle, planting tables both sides, glass frame, soft morning light, the
same locked-off wide. By then the shade is fully up, the aisle is empty,
and the leaves still move slightly.
Anything that only appears after @Video 1 starts must not appear early
here.A seam that reads as continuous is the target; identical pixels are not. Review both sides of the join and then watch the whole extension — drift usually shows up a second or two past the boundary, not at it.
KEYFRAMES, STORYBOARDS AND BLOCKOUTS
— when a picture specifies the shot better than a sentenceFirst and last frame
Declare the anchors on their own lines — one role per image. Extra references can still define a face or a material, as long as you say they must not restage the anchors. The output takes its aspect ratio from the first image, so give both anchors the same ratio or the last frame gets stretched to fit.
@Image 1 is the first frame: opening framing, positions, poses, tabletop
state, workshop, camera direction.
@Image 2 is the last frame: closing framing, positions, poses, tabletop
state, workshop, camera direction.
@Image 3 is the perfumer's face, hair and dark green apron. It must not
change the framing set by @Image 1 or @Image 2.
@Image 4 is the glass bottle's shape, material and label position. Same
restriction.
From the opening pose, the perfumer picks up a dropper and the bottle,
drips amber oil in, swirls it, seats the stopper, sets the finished bottle
in the middle of the table, and arrives naturally at @Image 2.
Throughout: her identity and clothes, the bottle count and shape, the
table layout, the warm side light and the camera direction all hold.A run of keyframes
More than two anchors works the same way: state the order once, then say what each image is the state of. Keyframes govern stage order and the look of each landing point — they do not reproduce themselves frame for frame.
Use @Image 1 through @Image 4 as keyframes, in that order.
@Image 1 is the first frame: an orange paper plane on the left of a
classroom desk, nose pointing right, locked-off medium.
@Image 2 is the second state: a hand lifts that same plane off the desk,
attitude unchanged.
@Image 3 is the third state: the same plane passes the window as the
curtain shifts slightly right.
@Image 4 is the last frame: the same plane at rest on the middle shelf of
the bookcase, still pointing right.
Move through those four states in order, with flight direction and speed
continuous between them. Hold the plane's orange paper, size and folds,
the classroom layout, the afternoon side light and the camera axis.Storyboard grids
A grid communicates order and rough framing — not detail. Keep it under about fifteen panels, keep the drawing to clean line art, and keep text out of it. Then say how to read the grid, what to ignore about its style, and what each panel actually contains.
@Image 1 is a four-panel storyboard for shot order and rough framing.
Read it left to right, top to bottom. Ignore its line-art style and any
labels in it.
@Image 2 is the potter's face, short hair and dark grey apron.
@Image 3 is the blue-glazed cup's proportions, glaze and curved handle.
Shot 1: wide — a quiet studio, the potter seated at the wheel.
Shot 2: side medium — both hands shaping the turning clay, the wall rising.
Shot 3: close — fingers refining the rim and the handle join, slip moving
over the fingertips.
Shot 4: medium close — the fired cup on a wooden shelf as both hands
withdraw.
Finish it as realistic documentary. Audio: wheel rotation, wet clay, the
room.Blockouts
Decide first which kind you have. A coarse blockout is grey geometry that carries timing and space; a fine blockout is finished modelling that needs skin. They take opposite instructions.
| COARSE | FINE | |
|---|---|---|
| Carries | paths, blocking, camera, cuts, light changes, rhythm | full structure, action, layout, camera |
| You supply | who and what each shape becomes | materials, colour, character, scene, style |
| Material must be | simple shapes with clear relations | clean model — no path lines, axes or camera frustums |
| Prompt does | map every shape, name what to inherit | preserve structure and motion, define the re-render |
@Video 1 is a coarse blockout. Take from it only the walking path, the
cart's direction, the locked-off camera, the one push-in and the two
cuts. Not its grey geometry or empty room.
The tall cylinder is <Guide>.
The rectangular block is <Display cart>.
@Image 1 is <Guide>'s face, blue uniform and name badge.
@Image 2 is the cart's white metal frame and clear cover.
@Image 3 is the gallery: curved walls, grey floor, overhead strip lights.
<Guide> pushes <Display cart> along the curved wall, stops at the central
display and opens the cover. Keep the path, the blocking, the push-in
direction and the cut points from @Video 1.
Bright realistic museum documentary. Audio: footsteps, castors, room tone.@Video 1 is a fine blockout. Preserve the sculpture's full structure, the
three-ring rotation, the plinth position, the orbiting camera and the
cuts. Not its grey material or empty background.
@Image 1 is the outer ring's brushed brass.
@Image 2 is the inner blades' translucent blue glass.
@Image 3 is a contemporary gallery: white curved walls, dark grey floor,
soft overhead light.
Re-render the rings as a brass and blue-glass kinetic sculpture, and the
room as that gallery. Keep the structure, the rotation rhythm, the orbit
and the cuts. Audio: a low mechanical turning, quiet interior.ASSEMBLY AND TRANSITIONS
— turning a pile of material into one pieceOne-click video from stills
"Make these into a video" is not an instruction. Five things need saying: what each image is for, what order they run in, how much each one moves, how it's cut and finished, and what it sounds like.
[Roles]
@Image 1 — night-market entrance, the opening.
@Image 2 — the traveller walking the street.
@Image 3 — the lantern stall and its detail.
@Image 4 — three friends eating.
@Image 5 — the river at night, reflections.
@Image 6 — the closing group photo on the bridge.
@Video 1 — its cutting rhythm, sticker treatment and transitions only.
Not its people or places.
[Order]
Run @Image 1 to @Image 6 in that order: arrive, wander, eat, walk the
river, take the photo. The three friends keep their faces and clothes
throughout; never mix them up.
[Motion]
Slow push-ins and slight parallax on the environments. On the people,
only blinking, a head turn, a raised glass, a little cloth movement.
Stall structure, table position and bridge railing stay put.
[Finish]
Upbeat travel-cut rhythm, scenes joined on occlusion and matching colour,
hand-drawn stickers kept to the edges of frame.
[Audio]
Market chatter, light crockery, river wind, and unobtrusive upbeat
instrumental.Seamless transitions between two clips
A bridge shot needs seven things: which clip is before, which is after, what triggers the change, what the camera does, what visually morphs into what, where it arrives, and how the sound crosses over.
@Video 1 is the outgoing clip — take its rainy street, red umbrella,
slow push-in and rain.
@Video 2 is the incoming clip — take its round gallery skylight, rising
camera and quiet interior reverb.
Both clips keep their people, spaces, structures and main actions.
At the end of @Video 1 the red umbrella comes toward camera and fills the
frame — that starts the transition.
The camera keeps moving forward. The umbrella's circular edge becomes the
skylight's metal ring; the red fabric becomes white daylight through the
glass.
It lands on @Video 2's opening low-angle framing, the move changing from
forward travel into a rise.
The rain thins into footsteps echoing inside the gallery.| TRANSITION | WHAT TO SPELL OUT |
|---|---|
| Dive / reverse | camera direction, speed change, when the next space begins |
| Subject turn | pose, rotation direction, how clothing and backdrop change with it |
| Foreground wipe | when the object fills frame, and the framing waiting behind it |
| Object morph | which shapes correspond, which materials, how the change happens |
| Push or rack | the move, the focus target, the spatial link across the cut |
DIRECTING PERFORMANCE
— convert the adjective into something a camera can see"Tense," "warm" and "oppressive" set a direction and leave the acting open. To close the gap, give observable cues — eyes, brows, mouth, breath, where they look, what the hands do. Two to four cues carry a single emotional turn; you do not need a full anatomy of the face.
The feeling moves from <start> to <end>.
After <trigger>, <subject> first <the immediate, involuntary reaction>.
Then <eyes, brows, mouth, breath, gaze or hands> gradually <change>.
Finally <subject> shows <target feeling> by <what they outwardly do>.Applause reaches him from behind the stage. His fingers stop on the
programme, his eyes turn slowly toward the curtain, his shoulders stay
locked.
Once it's clear the curtain call is over, he lets out a breath. The
shoulders drop, a held-back smile arrives, his eyes slowly fill — and he
still doesn't turn to leave.For several changes in one clip, hang each on its trigger — heard, seen, or confirmed — instead of listing feelings in sequence. An emotion with a cause reads; an emotion on a schedule doesn't.
CAMERA LANGUAGE
— standard terms work; unusual ones need translatingOrdinary film vocabulary goes in as-is. The model knows shot sizes, moves and angles.
| TERMS THAT LAND | |
|---|---|
| Shot size | extreme wide, wide, medium, close-up, extreme close-up |
| Movement | push in, pull out, pan, track, follow, orbit, dive, dolly out, tilt up, handheld |
| Angle | low angle, overhead, first person |
| Techniques | one-take, dolly zoom, aerial, FPV, bullet time, handheld, speed ramp |
With more than one subject in frame, every move still needs an object: which subject the camera follows or orbits, where the move starts and where it stops. A dolly zoom needs to know whose scale to hold and whether the background should rush in or away. Bullet time needs to know what freezes and which way the camera travels around it.
Terms the model may not know
For niche or contested vocabulary, keep the term and immediately restate it as a visible change — subject, what changes, foreground against background, direction or speed.
Rack focus: pull focus from the leaves in front to the woman behind.
The leaves soften as her face comes sharp.
Shallow portrait: keep the pastry chef's eyes and face sharp while the
jars and lights behind her fall into soft round bokeh.
Tracking: move sideways at the skater's speed, holding him sharp while
the roadside wall smears into horizontal blur, right to left.
Golden hour: warm low sun from behind and left of the hiker, throwing
long shadows across the ridge.
Natural vignette: darken the four corners gradually, keeping the centre
brightness and her skin tone natural — no hard black border.
Whip pan: at 5s, snap the camera left; cut when the foreground shelving
covers frame, and keep moving left at the same speed in the next scene.Aperture, focal length and shutter values are allowed, but the visible result you want is almost always the clearer instruction. "f/1.4" is a hope; "background falls into soft round bokeh" is a description.
LOCKED PARAMETERS
— what the task type decides for youSome tasks take their parameters from the material rather than from the form. Knowing which saves you from writing a setting the model was never going to honour.
| TASK | ASPECT RATIO | DURATION |
|---|---|---|
| Text or image to video | yours to set | yours to set |
| Video edit | locked to the source | locked to the source, ±0.3s |
| First / first + last frame | locked to the first image | yours to set |
| Extension | locked to the source | yours to set |
A locked value cannot be overridden on the form or through the API. Two consequences worth remembering: mismatched first and last images stretch the last frame, and an extended segment can sit at a slightly different loudness from its source.
BEFORE YOU RENDER
— the pass that pays for itself at 270 credits- The subject and the main action are stated in one readable sentence.
- Every reference says what to take from it — and what to ignore.
- Every distinct person, product and prop is named and bound to exactly one material.
- Materials are chosen per scene, not all dumped into every scene.
- Each stage carries one change and ends on a visible state.
- Headcount, clothing, who-holds-what and screen direction are pinned across stages.
- Time ranges run consecutively, don't overlap, and don't demand a rate.
- Emotions and unusual camera terms are paired with something you could actually see or hear.
- First and last images share an aspect ratio; each anchor image has one role.
- For an edit: one master, an explicit scope, a stated count, and a preserve list.
- For an extension: the boundary state is written out before any new action.
- For assembly: roles, order, amount of motion, finish and audio are all specified.
- Nothing in the prompt tries to set a parameter the task has already locked.
WHERE PROMPTING STOPS
— the honest limits- Timestamps budget time for events. They are not frame-accurate edit points.
- An edit prompt raises the odds that key events line up with the source. It cannot promise frame-for-frame overlap.
- Multi-reference work is about choosing and combining the right material — not about making everything appear at once.
- Subtitles, formulas, signage, spec numbers and exact frame timing still want a real edit afterwards. Generate the shot, finish the text in post.
- A generated transition is a bridge, not a splice; it will not preserve both source clips pixel for pixel.
- Visual quality, human realism and physical plausibility come from the model. Structure raises your hit rate; it does not raise the ceiling.
Everything here is technique, not warranty. The same prompt renders differently on different days — that is the medium, and the reason to draft short, re-roll cheaply, and only commit the full thirty seconds to a structure that already reads.