PromptSama Model field notes
Static edition All models Download prompt skill ↓
MiniMax H3 Multimodal guide · Aug 2026

Assign authority.
Then stage the shot.

H3 treats text, images, video, and audio as one creative context. Reliable prompts tell it what each asset controls, lock the identities and screen geography, then describe a timed audiovisual scene.

H3 / AUTHORITY STACK01—06
  1. 01Reference useidentity · motion · camera · voice
  2. 02Continuity lockswardrobe · props · count · geography
  3. 03Scene intentlocation · conflict · emotional turn
  4. 04Timed shotsframing · action · dialogue · transition
  5. 05Image & soundlight · ambience · stereo · music
  6. 06Negativesno drift · no swaps · no extra cuts
4–15soutput
duration
12mixed files
maximum
2K1440px
short edge
24frames per second
with stereo audio

01 / OVERVIEW

One model, four kinds of evidence

H3 combines references instead of treating generation, transfer, and editing as isolated tasks.

A

Generate commercially useful shots

Direct film, brand, product, game, VFX, typography, and stylized content with native picture and sound.

B

Understand mixed references

Combine character identity, performance, camera, voice, environment, style, atmosphere, and edit rhythm.

C

Transfer performance precisely

Use a video for motion and timing while separate images control who appears and audio controls how they sound.

D

Edit without rebuilding everything

Target a character, object, background, light, dialogue, voice, effect, or rhythm while preserving untargeted content.

!
The core H3 rule

References are not a pile of inspiration. Each one needs a declared job and a boundary on what it must not control.

02 / SPECIFICATIONS

Design inside the input envelope

Working specification snapshot reviewed August 2, 2026.

ControlMiniMax H3
Output

4–15 seconds · 24 FPS · every generation includes native stereo audio.

Aspect ratios

21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Omni Reference also offers Auto; First/Last Frame follows the input image.

Resolution

2K mode uses a 1440px short edge from 16:9 through 9:16; wider formats target about 3.7 MP. The 768p mode is marked as coming soon.

First / Last Frame

0, 1, or 2 images. Image dimensions 256–5760px; aspect ratio 5:2 through 2:5. No images becomes text-to-video.

Omni Reference

Up to 9 images, 3 videos, and 3 audio clips; up to 12 mixed files total. Video and audio pools allow 15 seconds total each.

Prompt

Up to 7,000 characters.

Formats

Images: JPG, JPEG, PNG, WEBP, HEIC, HEIF. Video: H.264/AVC or H.265/HEVC with AAC/MP3 audio. Audio: WAV or MP3.

Per-file limits

Video 50 MB · image 30 MB · audio 15 MB · API request body 64 MB. URL media is recommended for API workflows.

03 / PROMPT ARCHITECTURE

Write a production brief, not a paragraph

Put stable truth first. Then time the scene. Finish with camera, sound, and a short negative list.

01
REFERENCE USE

Assign identity, wardrobe, motion, camera, scene, voice, sound, or style. State what to ignore.

02
IDENTITY / CONTINUITY LOCKS

Fix subject count, appearance, wardrobe, props, voice ownership, screen direction, and spatial relationships.

03
SCENE + DIALOGUE

State the location, dramatic goal, emotional shift, exact spoken lines, speakers, and delivery.

04
SCREEN GEOGRAPHY

Place people and objects in frame before asking them to move. Keep left/right relationships stable.

05
SHOT LIST

For each time range: framing, camera motion, visible action, dialogue, transition, and end state.

06
ACTING + LIGHT + SOUND + NEGATIVES

Separate performance, visual treatment, production audio, and only the failure modes that matter.

[REFERENCE USE]
Image1 defines <Character A>'s identity and wardrobe only. Ignore its background.
Image2 defines <Character B>'s identity and wardrobe only. Ignore its background.
Video1 defines only performance timing and body motion. Do not inherit its people or location.
Audio1 defines <Character A>'s voice timbre and delivery.

[IDENTITY / CONTINUITY LOCKS]
Keep exactly two people. Preserve faces, hair, clothing, body proportions, and voice ownership.
Keep screen direction and prop ownership stable. No identity or wardrobe swaps.

[SCENE]
<Location, time of day, dramatic goal, and emotional progression.>

[DIALOGUE]
<Character A> says: “<Exact line.>”
<Character B> replies: “<Exact line.>”

[SCREEN GEOGRAPHY]
<Character A> remains frame left; <Character B> remains frame right.

[SHOT LIST]
0–4s — <framing, camera, action, dialogue, transition>.
4–9s — <framing, camera, action, dialogue, transition>.
9–15s — <framing, camera, action, dialogue, final state>.

[ACTING]
<Posture, gaze, gesture, emotional progression.>

[LIGHT AND IMAGE]
<Lighting, palette, texture, lens character, depth of field.>

[CAMERA]
<Shot sizes, movement rules, axis, and transition language.>

[PRODUCTION SOUND]
Native stereo ambience, clear dialogue, motivated effects. <Music rule.>

[NEGATIVES]
No extra people, face drift, wardrobe changes, voice swaps, broken eyelines, or unmotivated cuts.

04 / REFERENCE MODES

Define what each asset is allowed to control

Use compact labels—Image1, Video1, and Audio1—with no @ symbol or space. Keep every authority narrow and truthful.

Image1

Identity

Face, hair, body, wardrobe, product geometry, or a specific object.

Exclude background, pose, and composition unless needed.
Video1

Motion

Body mechanics, path, gesture order, expression timing, or interaction rhythm.

Exclude the original actor, wardrobe, and location.
Video2

Camera / edit

Framing, dolly path, shot rhythm, transitions, or overall visual language.

Do not let it overwrite identity or scene content.
Audio1

Voice

Timbre, accent, cadence, emotion, singing, or another vocal performance.

Name the target speaker and exact line.
Image2

Scene / style

Environment, lighting, palette, material treatment, graphic effect, or atmosphere.

Say whether it controls layout, appearance, or both.
Video3

Target result

A consolidated example of style, soundscape, pacing, or final behavior.

Use only when its authority is clearer than separate references.
AMBIGUOUS

“Use all references to make the same scene.”

Every asset competes for identity, composition, motion, and style.
CONTROLLED

“Image1 defines identity only. Video1 defines motion timing only. Audio1 defines her voice.”

Authority is explicit and cross-contamination is bounded.

05 / SHOT TIMING

One visible beat per time range

H3 outputs 4–15 seconds, so every second needs a job. Use ranges as editorial budgets, not frame-accurate commands.

00—04Establish

Two-shot, geography, shared task, first line, subtle move in.

End: who is left/right and what each person holds.
04—09Escalate

Closer framing, interruption, reaction, prop or gaze transfer.

End: new emotional state and readable eyeline.
09—15Resolve

Final reply, decisive action, camera settle, audible scene tail.

End: stable tableau that confirms the outcome.

FramingDeclare shot size and who is visible before describing action.

ActionUse one primary physical event with a clear cause and result.

DialogueBind every exact line to one speaker and one delivery.

TransitionName a cut, pan, occlusion, or continuous move only when it matters.

End stateFinish with observable positions, gaze, props, and emotional state.

07 / PRECISE EDITING

Name the target and protect the rest

A reliable H3 edit has one source master, a bounded change list, and an equally explicit preservation list.

1

Declare the master

“Video1 is the sole source for timeline, subjects, camera, and audio.”

2

List the changes

Bind every replacement to an exact subject, object, region, line, or time range.

3

List what survives

Protect identity, movement, occlusion, lighting, camera, dialogue, ambience, and timing as needed.

Multi-element scene edit

Video1 is the sole editing master.

Replace the newspaper with one green hardcover book. Replace the chair with one red sofa. Remove the subject’s sunglasses and reveal the same face. Remove the burning-car effect and restore the same vehicle to normal. Replace the coat photograph with one small black notebook. Add one tree at frame left.

Preserve every actor’s identity, action timing, occlusion, eyeline, camera move, scene layout, dialogue, and ambience. Do not modify anything else.

Dialogue and performance edit

Video1 is the sole source for the woman, scene, camera, interaction, and duration.
Audio1 defines the replacement line, voice, timing, and emotional delivery.

Replace only the woman’s original line with: “Please don’t go. This time, let’s not let each other go.” Adjust her facial expression and breath subtly to support the new plea.

Preserve both identities, the other person’s performance, setting, camera, edit rhythm, ambience, and total duration. Maintain accurate lip sync.

08 / TROUBLESHOOTING

Repair authority, timing, and preservation

FailureLikely causeRewrite
Identity drifts

Appearance is implied or split across competing references.

Use one primary identity reference and lock face, hair, body, and wardrobe.

Wrong person gets the voice

Audio has no named target.

Assign Audio1 to one character and restate exact dialogue and voice ownership.

Motion reference changes the cast

The video’s authority is too broad.

Say it defines motion and timing only; explicitly reject its actors and location.

Complex action becomes chaos

Position changes are not ordered.

Write the path as a chronological state sequence with stable left/right geography.

Edit rebuilds the whole frame

No source master or preservation contract.

Name the sole master, exact targets, and every untouched category to preserve.

Dialogue is rushed

Too much text for the shot budget.

Shorten lines, reserve reaction time, and stage one conversational turn per range.

Prompt copied