PromptSama Model field notes
Static edition All models Download prompt skill ↓
Seedance 2.0 Multimodal field guide · Aug 2026

Direct the space.
Sequence the time.

Seedance 2.0 reads images, videos, audio, and text as one production brief. Assign each asset a bounded role, organize complex action as ordered shots, and state exactly what edits or extensions must preserve.

MULTIMODAL / DIRECTOR STACK01—05
  1. 01Choose the taskgenerate · edit · extend · join
  2. 02Assign asset rolesImage 1 · Video 1 · Audio 1
  3. 03Bind the subjectsidentity · wardrobe · props
  4. 04Sequence the shotscamera · action · space · audio
  5. 05Constrain driftcount · style · text · continuity
4–15sgenerated video
duration range
9reference images
per request
3reference videos
15 seconds combined
3reference audio clips
15 seconds combined

01 / OVERVIEW

Write for two layers at once

The prompt becomes easier to control when you separate what exists in the frame from what changes over time.

A

Spatial truth

Define subjects, identity, props, scene layout, lighting, style, and generated text. Images are strongest when their authority is specific.

B

Temporal change

Define actions, camera movement, cuts, effects, dialogue, and event order. Videos and audio can provide motion, rhythm, or voice evidence.

C

Bounded authority

An identity image need not control its background. A motion video need not control its actor, scene, color grade, or wardrobe.

D

Visible behavior

Replace “very anxious” with darting eyes, shallow breath, shifting weight, and fingers tapping a bag handle.

!
Version-specific syntax

Seedance 2.0 uses ordered labels such as Image 1, Video 1, and Audio 1. Do not copy Seedance 2.5’s @image1 syntax into a 2.0 prompt.

02 / TASK MODES

Choose the verb before the visuals

The wording determines whether an asset inspires a new video, modifies an existing one, or continues its timeline.

01

Reference

Extract identity, scene, style, action, camera, effects, voice, or rhythm from one or more assets to generate a new video.

02

Edit

Directly edit Video 1. Add, remove, or replace a bounded target and name every category that must remain unchanged.

03

Extend

Continue Video 1 forward or backward while inheriting its seam: pose, object state, composition, lighting, motion, and sound.

04

Complete tracks

Join two or three ordered videos with motivated transition descriptions. Use this for scene turns and higher-action sequences.

AVOID FOR EDITS

“Reference Video 1 and replace the package.”

This can be interpreted as a new generation inspired by the source.
WRITE

“Strictly edit Video 1. Replace only the package; preserve all motion and camera work.”

The edit target and source authority are explicit.

03 / REFERENCE LABELS

One ordered asset, one declared job

Labels identify the file. Role sentences tell the model which information to extract from it.

Image 1defines a character’s face, hair, build, and clothing—not the background.
Image 2defines a prop’s structure, material, color, and wear.
Image 3defines scene layout, architecture, light, and atmosphere.
Video 1defines only an action, movement path, effect, pacing, or camera behavior.
Audio 1defines a voice’s timbre, pace, emotional tone, or rhythmic character.

REFERENCE CONTRACT

Every role sentence answers:

  • Which subject or production dimension is this asset?
  • What evidence should be inherited?
  • What incidental content should be ignored?
AMBIGUOUS

“Use Image 1 and Image 2 for the two people respectively.”

The subject-to-image relationship is implied.
BOUND

“Define the woman in Image 1 as <Maya>. Define the man in Image 2 as <Noah>.”

The same stable labels can now be used in every shot.
Input structure that affects prompt designExpand +
Images

Up to 9 in multimodal reference mode.

Prefer a clean headshot and full-body image over a dense character collage.
Video

Up to 3; 15 seconds combined.

Name the exact motion, camera, effect, or sound dimension to extract.
Audio

Up to 3; 15 seconds combined.

Audio cannot be the only reference modality.
Boundary frames

First/Last Frame is a separate mode.

Use it when exact opening and closing frames matter.

04 / PROMPT ARCHITECTURE

Build from identity to constraint

ADVANCED FORMULA
Precise subject+ Action details+ Scene+ Light & color+ Camera+ Style & quality+ Constraints

Put asset roles before this formula when references define any part of the result.

Reference-first template

[Asset Roles]
Image 1 defines <subject>. Ignore <incidental content>.
Video 1 defines only <motion or camera>.

[Generation Goal]
Create <one-sentence video brief>.

Shot 1: <camera, action, space, audio, end state>.
Shot 2: Continue from Shot 1. <next coherent change>.

[Global Direction]
<lighting, color, style, quality, and sound bed>.

[Constraints]
<identity, count, text, logo, watermark, and style locks>.

Production order

  1. 01

    Asset roles — authority, subject, inherited traits, exclusions.

  2. 02

    Generation goal — one sentence that fixes the task.

  3. 03

    Shot order — camera, action, space, audio, and end state.

  4. 04

    Global direction — style, quality, light, color, and sound bed.

  5. 05

    Constraints — only the failure modes likely in this scene.

05 / SHOTS & MOTION

Order shots; don’t over-time them

For complex clips, ordered shot language is more stable than forcing several exact second ranges.

SHOT 1Establish and begin

Medium two-shot. The subjects start walking; one raises an open hand.

End: both still moving; the hand lowers.
SHOT 2Change one relationship

Cut to a tracked profile. The listener turns, tightens the jaw, and responds.

End: gaze returns forward; pace remains even.
SHOT 3Resolve visibly

Return to the two-shot. Shoulders release and both continue past camera.

End: stable identity, direction, and spacing.

Action detailName body part, direction, range, speed, and force.

Action continuityDescribe how momentum carries from one action into the next.

Camera disciplineUse one principal movement per shot; use cuts for a new camera idea.

EmotionDirect gaze, breath, posture, fingers, jaw, shoulders, and pacing—not abstract intensity.

06 / AUDIO & GENERATED TEXT

Mark the information type

( )
Music

(Muted electronic pulse beneath the scene.)

< >
Sound effect

<Escalator hum and distant footsteps.>

{ }
Dialogue

{Transparency has to be part of the product.}

【 】
Text / subtitle

【Design Review — Monday】

Keep dialogue in one language and identify it when ambiguity matters. Dialogue language: natural American English. <Maya> says with contained frustration: {That is not what I approved.}

07 / WORKFLOWS

Use the contract that matches the verb

09 / TROUBLESHOOTING

Repair the contract, not the adjectives

FailureLikely causeRewrite
Identity drifts

Face is too small or mixed into a collage.

Use a clean headshot first and repeat one stable role label.

Duplicate characters

Multi-view people or ambiguous counts.

Use separate single-person images and require exactly one of each role.

Unexpected subtitles

Dialogue or reference media encourages text.

Add “keep subtitle-free” and remove unnecessary text from references.

Style becomes realistic

Reference style conflicts with the target.

State the target style globally or preprocess the identity assets into it.

Extension seam jumps

Boundary pose and motion are vague.

Restate pose, objects, composition, lighting, motion direction, and sound.

Effect is structurally wrong

A complex effect was described only in text.

Supply a clean effect video and inherit only its formation path and timing.

Voice match is weak

Only timbre was named.

Add pitch, texture, pace, emotion, and sentence-ending behavior.

Too many actions vanish

One short clip carries too many events.

Reduce shot density or generate separate clips and join them.

Prompt copied