PromptSama Model field notes
Static edition All models Download prompt skill ↓
Seedance 2.5 Foundation guide · Aug 2026

Direct the reference.
Then direct the scene.

A practical guide to multimodal video prompting: assign every image, video, and audio clip a role; describe the visible state changes; and make every scene hand its continuity to the next.

OMNI / CONTROL STACK01—05
  1. 01Map each reference@image1 → character
  2. 02Lock identitiesappearance · clothing · props
  3. 03Stage the eventone primary change per beat
  4. 04Direct the camerashot · angle · motion · cut
  5. 05Declare end statewhat must remain visible
50combined reference
materials supported
30images per
OMNI request
30sstandard generation
and input pools
180sdedicated long-video
mode

01 / OVERVIEW

Think like a continuity director

The model can infer a lot. Your job is to remove the expensive ambiguities.

A

References define truth

Say what each asset contributes and what must be ignored. Never make the model guess which person or prop an image represents.

B

Events define motion

Write observable actions in chronological order. For longer clips, give each stage one primary state change.

C

End states define continuity

Finish each beat with positions, prop ownership, screen direction, and the visible state the next beat must inherit.

D

Exclusions prevent leakage

Explicitly reject incidental people, backgrounds, compositions, music, subtitles, or styles carried by a reference.

!
The most useful rule

Use more references only when they reduce uncertainty. Extra materials without clear roles usually reduce stability.

02 / PROMPT ANATOMY

Build from concrete to cinematic

CORE FORMULA
Subject+ Action / event+ Scene+ Visual treatment+ Camera+ Audio

Only subject and action are essential. Add the other layers when they materially constrain the result.

Short-form template

<Subject> performs <primary action> in <scene>.
The visuals feature <lighting, texture, and mood>.
Use <shot size, angle, movement, or cuts>.
Audio includes <dialogue, ambience, SFX, or music>.

Prompt order for OMNI

  1. 01

    Reference roles — what each asset defines and excludes.

  2. 02

    Generation goal — the video type and one-sentence story.

  3. 03

    Timeline — events, camera, sound, and visible end states.

  4. 04

    Global locks — identities, props, space, axis, and negatives.

03 / OMNI REFERENCE MAPPING

One material, one declared job

Reference syntax is less about tagging files and more about assigning authority.

@image1defines the character’s face, hairstyle, and clothing.
@image2defines the prop’s structure, material, and color.
@image3defines spatial layout and lighting. Ignore its people.
@video1defines motion path, pacing, and camera movement only.
@audio1defines the presenter’s voice and delivery.

REFERENCE CONTRACT

For every asset, answer three questions:

  • Who or what does this material correspond to?
  • Which attributes should the model inherit?
  • Which incidental details must it ignore?
AVOID

“@Images 1–4 define four characters respectively.”

The mapping is ambiguous.
WRITE

“<Conservator> corresponds to @image1. Use only the face, hair, and clothing.”

The subject and inherited attributes are explicit.
Official input limits and recommended working rangesExpand +
Images

Up to 30, each no larger than 4K.

Prefer 1–8 distinct subjects.
Video

Up to 10, 30 seconds combined.

Prefer 5–10 seconds per subject.
Audio

Up to 10, 30 seconds combined.

Keep only task-relevant sound.
Edit mode

One source video plus reference images.

Prefer source under 20 seconds and 1–5 images.

04 / AUDIO & TEXT SYNTAX

Punctuation as a production cue

( )
Music

(Soft rhythmic piano plays underneath.)

< >
Sound effect

<A bicycle bell rings in the distance.>

{ }
Dialogue

{I thought you weren’t coming.}

【 】
On-screen text

【Chapter One: Departure】

For non-Chinese dialogue, reinforce the language before the line. Dialogue language: natural American English. The presenter says calmly: {Let’s begin.}

05 / TIMING & CONTINUITY

Use stages by default; seconds for critical beats

Timestamps allocate attention. They are not frame-accurate edit decisions.

00—05Set the state

Empty table. A hand places a ceramic plate in the center.

End: hand exits; only plate remains.
05—10Change one thing

Remove the plate. Place a clear glass in the same position.

End: only glass remains.
10—15Resolve visibly

Remove the glass. Place a green vase, centered and still.

End: one vase; no hands in frame.

Time rangeUse for a beat’s time budget: “0–5 seconds…”

Exact pointReserve for a key event: “At 5 seconds, whip-pan left.”

Relative timingUse for causality: “Three seconds after the button press…”

06 / WORKFLOWS

Choose the contract that matches the task

08 / TROUBLESHOOTING

Fix ambiguity before adding adjectives

FailureLikely causeRewrite
Characters swap identities

Grouped or implied mappings.

Name every subject and bind it to one reference.

Reference background leaks in

No exclusions.

Say “use only appearance; do not use background or composition.”

Long clip skips actions

Too many state changes per range.

Give each stage one primary event and an observable end state.

Extension jumps at the seam

Boundary frame is underspecified.

Restate pose, prop position, camera, lighting, and motion direction.

Edit changes untouched areas

Preservation contract is vague.

Name the sole master video, exact edit scope, and everything to preserve.

Reference motion fights the prompt

The action was redundantly rewritten.

If the video defines motion well, state only which traits to inherit.

Prompt copied