Assign authority.
Then stage the shot.
H3 treats text, images, video, and audio as one creative context. Reliable prompts tell it what each asset controls, lock the identities and screen geography, then describe a timed audiovisual scene.
- 01Reference useidentity · motion · camera · voice
- 02Continuity lockswardrobe · props · count · geography
- 03Scene intentlocation · conflict · emotional turn
- 04Timed shotsframing · action · dialogue · transition
- 05Image & soundlight · ambience · stereo · music
- 06Negativesno drift · no swaps · no extra cuts
duration
maximum
short edge
with stereo audio
01 / OVERVIEW
One model, four kinds of evidence
H3 combines references instead of treating generation, transfer, and editing as isolated tasks.
Generate commercially useful shots
Direct film, brand, product, game, VFX, typography, and stylized content with native picture and sound.
Understand mixed references
Combine character identity, performance, camera, voice, environment, style, atmosphere, and edit rhythm.
Transfer performance precisely
Use a video for motion and timing while separate images control who appears and audio controls how they sound.
Edit without rebuilding everything
Target a character, object, background, light, dialogue, voice, effect, or rhythm while preserving untargeted content.
References are not a pile of inspiration. Each one needs a declared job and a boundary on what it must not control.
02 / SPECIFICATIONS
Design inside the input envelope
Working specification snapshot reviewed August 2, 2026.
4–15 seconds · 24 FPS · every generation includes native stereo audio.
21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Omni Reference also offers Auto; First/Last Frame follows the input image.
2K mode uses a 1440px short edge from 16:9 through 9:16; wider formats target about 3.7 MP. The 768p mode is marked as coming soon.
0, 1, or 2 images. Image dimensions 256–5760px; aspect ratio 5:2 through 2:5. No images becomes text-to-video.
Up to 9 images, 3 videos, and 3 audio clips; up to 12 mixed files total. Video and audio pools allow 15 seconds total each.
Up to 7,000 characters.
Images: JPG, JPEG, PNG, WEBP, HEIC, HEIF. Video: H.264/AVC or H.265/HEVC with AAC/MP3 audio. Audio: WAV or MP3.
Video 50 MB · image 30 MB · audio 15 MB · API request body 64 MB. URL media is recommended for API workflows.
03 / PROMPT ARCHITECTURE
Write a production brief, not a paragraph
Put stable truth first. Then time the scene. Finish with camera, sound, and a short negative list.
Assign identity, wardrobe, motion, camera, scene, voice, sound, or style. State what to ignore.
Fix subject count, appearance, wardrobe, props, voice ownership, screen direction, and spatial relationships.
State the location, dramatic goal, emotional shift, exact spoken lines, speakers, and delivery.
Place people and objects in frame before asking them to move. Keep left/right relationships stable.
For each time range: framing, camera motion, visible action, dialogue, transition, and end state.
Separate performance, visual treatment, production audio, and only the failure modes that matter.
[REFERENCE USE]
Image1 defines <Character A>'s identity and wardrobe only. Ignore its background.
Image2 defines <Character B>'s identity and wardrobe only. Ignore its background.
Video1 defines only performance timing and body motion. Do not inherit its people or location.
Audio1 defines <Character A>'s voice timbre and delivery.
[IDENTITY / CONTINUITY LOCKS]
Keep exactly two people. Preserve faces, hair, clothing, body proportions, and voice ownership.
Keep screen direction and prop ownership stable. No identity or wardrobe swaps.
[SCENE]
<Location, time of day, dramatic goal, and emotional progression.>
[DIALOGUE]
<Character A> says: “<Exact line.>”
<Character B> replies: “<Exact line.>”
[SCREEN GEOGRAPHY]
<Character A> remains frame left; <Character B> remains frame right.
[SHOT LIST]
0–4s — <framing, camera, action, dialogue, transition>.
4–9s — <framing, camera, action, dialogue, transition>.
9–15s — <framing, camera, action, dialogue, final state>.
[ACTING]
<Posture, gaze, gesture, emotional progression.>
[LIGHT AND IMAGE]
<Lighting, palette, texture, lens character, depth of field.>
[CAMERA]
<Shot sizes, movement rules, axis, and transition language.>
[PRODUCTION SOUND]
Native stereo ambience, clear dialogue, motivated effects. <Music rule.>
[NEGATIVES]
No extra people, face drift, wardrobe changes, voice swaps, broken eyelines, or unmotivated cuts.
04 / REFERENCE MODES
Define what each asset is allowed to control
Use compact labels—Image1, Video1, and Audio1—with no @ symbol or space. Keep every authority narrow and truthful.
Image1Identity
Face, hair, body, wardrobe, product geometry, or a specific object.
Exclude background, pose, and composition unless needed.Video1Motion
Body mechanics, path, gesture order, expression timing, or interaction rhythm.
Exclude the original actor, wardrobe, and location.Video2Camera / edit
Framing, dolly path, shot rhythm, transitions, or overall visual language.
Do not let it overwrite identity or scene content.Audio1Voice
Timbre, accent, cadence, emotion, singing, or another vocal performance.
Name the target speaker and exact line.Image2Scene / style
Environment, lighting, palette, material treatment, graphic effect, or atmosphere.
Say whether it controls layout, appearance, or both.Video3Target result
A consolidated example of style, soundscape, pacing, or final behavior.
Use only when its authority is clearer than separate references.“Use all references to make the same scene.”
Every asset competes for identity, composition, motion, and style.“Image1 defines identity only. Video1 defines motion timing only. Audio1 defines her voice.”
Authority is explicit and cross-contamination is bounded.05 / SHOT TIMING
One visible beat per time range
H3 outputs 4–15 seconds, so every second needs a job. Use ranges as editorial budgets, not frame-accurate commands.
Two-shot, geography, shared task, first line, subtle move in.
End: who is left/right and what each person holds.Closer framing, interruption, reaction, prop or gaze transfer.
End: new emotional state and readable eyeline.Final reply, decisive action, camera settle, audible scene tail.
End: stable tableau that confirms the outcome.FramingDeclare shot size and who is visible before describing action.
ActionUse one primary physical event with a clear cause and result.
DialogueBind every exact line to one speaker and one delivery.
TransitionName a cut, pan, occlusion, or continuous move only when it matters.
End stateFinish with observable positions, gaze, props, and emotional state.
06 / EXAMPLE GALLERY
Copy the control pattern
Examples are structured as reusable production patterns.
Cinematic asset remix
Separate asset content from editorial language.
[REFERENCE USE]
Image1–Image6 define only the available characters, objects, and locations. Do not copy their framing.
Video1 defines shot rhythm, transition language, and music only. Do not copy its subjects.
[SCENE]
Create one coherent cinematic sequence using the six image assets. Preserve each asset’s defining appearance and keep subject count stable.
[EDITING]
Match Video1’s pace and transition grammar without recreating its story. Keep geography legible and every transition motivated.
Coffee surface to desert
Use shared texture as the transition mechanism.
Image1 defines the coffee surface: milk foam, cocoa dust, bubbles, and dark liquid.
Image2 defines the destination desert and its wind-carved dunes.
One continuous photoreal macro shot. Push rapidly into Image1 until particles, foam contours, and ripples fill frame. At the moment those forms align with the dune ridges and airborne sand in Image2, transform seamlessly into the desert and continue pushing forward to reveal the full landscape.
Extremely shallow depth of field at the start; quiet restrained sound. No hard cut, black frame, tear, or visible compositing seam.
Two-person foam fight
Bind identity to an image and interaction timing to video.
Image1 defines the two characters’ identities and wardrobe. Ignore its background and pose.
Video1 defines body motion, expressions, interaction timing, and camera framing only. Do not use its people.
At the sink on frame right, the man hands a clean plate to the woman on frame left. He turns and flicks dish-soap foam toward her with his right hand. She recoils, then immediately retaliates. They laugh, dodge, and throw foam back and forth.
Keep both identities stable, preserve left/right geography, and make foam, wet surfaces, shadows, and reactions physically believable.
Capybara pyramid transfer
Describe ordered positions when a motion reference is complex.
Video1 is the sole reference for motion paths, timing, and the locked-off wide camera. Replace its three suited men with exactly three photoreal capybaras.
Preserve the action order: all three drop to the floor; the left capybara jumps to center; the center capybara rolls to far left; the new center capybara rolls to far right; the right capybara jumps to center; finally the center capybara jumps onto the other two to form a pyramid.
Keep the camera fixed. Integrate fur, weight, floor contact, lighting, and shadows naturally. No extra animals or position swaps.
Voice-cloned dialogue
Assign the recording to one speaker and write the exact line.
Video1 defines the character, scene, camera, and visible performance.
Audio1 defines only the character’s voice timbre, cadence, and emotional delivery.
The character says exactly: “Follow the wind, live free. Leave worries behind and enjoy the moment.”
Match Audio1’s voice while preserving clear lip sync, the original character identity, and native stereo ambience. No subtitles, extra speech, or voice drift.
Green-screen replacement
Declare the source master and preserve foreground truth.
[SOURCE MASTER]
Video1 defines the subject, performance, foreground timing, camera, and duration.
Video2 defines only the target fairy-tale environment, atmosphere, and lighting style.
[EDIT]
Remove the green background in Video1 and replace it with the environment from Video2. Make background motion and parallax respond correctly to the subject. Relight the subject to match the new scene.
[PRESERVE]
Keep identity, hair and fabric edges, gesture timing, scale, camera, and duration unchanged. No green spill or altered performance.
07 / PRECISE EDITING
Name the target and protect the rest
A reliable H3 edit has one source master, a bounded change list, and an equally explicit preservation list.
Declare the master
“Video1 is the sole source for timeline, subjects, camera, and audio.”
List the changes
Bind every replacement to an exact subject, object, region, line, or time range.
List what survives
Protect identity, movement, occlusion, lighting, camera, dialogue, ambience, and timing as needed.
Multi-element scene edit
Video1 is the sole editing master.
Replace the newspaper with one green hardcover book. Replace the chair with one red sofa. Remove the subject’s sunglasses and reveal the same face. Remove the burning-car effect and restore the same vehicle to normal. Replace the coat photograph with one small black notebook. Add one tree at frame left.
Preserve every actor’s identity, action timing, occlusion, eyeline, camera move, scene layout, dialogue, and ambience. Do not modify anything else.
Dialogue and performance edit
Video1 is the sole source for the woman, scene, camera, interaction, and duration.
Audio1 defines the replacement line, voice, timing, and emotional delivery.
Replace only the woman’s original line with: “Please don’t go. This time, let’s not let each other go.” Adjust her facial expression and breath subtly to support the new plea.
Preserve both identities, the other person’s performance, setting, camera, edit rhythm, ambience, and total duration. Maintain accurate lip sync.
08 / TROUBLESHOOTING
Repair authority, timing, and preservation
Appearance is implied or split across competing references.
Use one primary identity reference and lock face, hair, body, and wardrobe.
Audio has no named target.
Assign Audio1 to one character and restate exact dialogue and voice ownership.
The video’s authority is too broad.
Say it defines motion and timing only; explicitly reject its actors and location.
Position changes are not ordered.
Write the path as a chronological state sequence with stable left/right geography.
No source master or preservation contract.
Name the sole master, exact targets, and every untouched category to preserve.
Too much text for the shot budget.
Shorten lines, reserve reaction time, and stage one conversational turn per range.
