Seedance 2.5 Prompt Guide
Seedance 2.5 is ByteDance's next-generation AI video creation model. The most important change is not simply better image quality: the model is designed to handle longer narratives, larger multimodal reference sets, more precise edits, and more detailed creative direction.
This Seedance 2.5 prompt guide turns those capabilities into practical writing methods. You will learn the basic prompt formula, how to assign roles to reference files, how to structure a 30-second story, and how to write prompts for editing, extension, keyframes, storyboards, white-model animation, seamless transitions, emotion, camera movement, and sound.
Ready to explore copyable ideas? Browse our 80+ Seedance 2.5 prompts with videos for professional examples covering stories, editing, transitions, ads, and multimodal creation.
Seedance 2.5 introduces longer generation, richer multimodal references, targeted editing, and stronger multilingual prompt control.
Seedance 2.5 vs Seedance 2.0: What Has Improved?
Compared with Seedance 2.0, Seedance 2.5 introduces four major workflow upgrades:
| Upgrade | What Seedance 2.5 adds | Why it matters |
|---|---|---|
| Longer storytelling | Up to 30 seconds in a single generation, with stronger temporal consistency | A complete emotional beat or narrative sequence can fit into one clip with fewer stitched generations |
| More multimodal references | Up to 50 combined image, video, and audio references | Creators can provide a full cast, multiple locations, camera references, product assets, and brand audio in one task |
| Precise video editing | Local changes to a background, product, character, or other selected element while preserving the wider shot | One generation can become a reusable master for localized ads, ecommerce variants, and multi-channel brand content |
| Broader language and instruction control | Native support for more than 10 languages, plus stronger understanding of complex shots and emotional changes | Global creators can describe ideas in their preferred language and give more exact direction without an extra translation step |
The practical result is a shift from making isolated short clips to directing a complete scene. Seedance 2.5 is especially useful for advertising, short films, product stories, branded content, and other work where character, scene, camera, and sound continuity all matter.
The Basic Seedance 2.5 Prompt Formula
Start with this flexible formula:
Subject + action or event + scene and environment (optional)
+ visual style (optional) + camera or cuts (optional) + sound (optional)
Each part has a distinct job:
- Subject + action or event: Who or what is doing what? Summarize the main process, then add detail only to the actions that matter most.
- Scene and environment: Define the place, time, weather, spatial relationships, and background state.
- Visual style: Describe lighting, color, material, image texture, and overall mood.
- Camera or cuts: Specify shot size, camera position, movement, focus, and transitions.
- Sound: Define dialogue, voice, ambience, sound effects, and music.
Basic prompt template
[Subject] performs [main action or event] in [scene and environment].
The image has [visual style].
Use [shot size, camera angle, camera movement, or cuts].
Sound includes [dialogue, ambience, sound effects, or music].
Basic prompt example
A ceramic artist finishes a pale-blue cup in a studio at dawn, removes it
from the wheel, and places it in the center of a wooden shelf.
Soft morning light enters through the window. The wet clay has a delicate
sheen, and the workbench remains tidy.
Begin with a medium shot of the shaping process, slowly push in toward the
surface texture of the cup, then cut to a front view of the shelf.
Keep the low rotation of the wheel, the friction of wet clay, and subtle room tone.
Omit any part you do not need. Set adjustable generation options, such as duration and aspect ratio, in the generation interface or API rather than repeating them in the prompt.
How to Use Reference Images, Videos, and Audio
Seedance 2.5 can combine as many as 50 multimodal reference assets. Stability is usually higher when every file has one clearly stated responsibility.
| Reference type | Input limit | Recommended starting point |
|---|---|---|
| Images | Up to 30 images, each no larger than 4K | 1-8 main subjects |
| Videos | Up to 10 clips, with a combined duration of 30 seconds | 1-5 main subjects, with each clip around 5-10 seconds |
| Audio | Up to 10 clips, with a combined duration of 30 seconds | Keep only dialogue, voice, ambience, or music relevant to the task |
| Video editing | An original video plus reference images | An original under 20 seconds and 1-5 reference images |
You can go beyond the recommended ranges—for example, 9-12 image subjects, 6-10 video subjects, or 6-8 editing references—but stability may decline as the set becomes more complex. If one subject needs several viewpoints, separate front, side, and rear views into individual images rather than placing every angle in one collage.
Assign a role to every asset
Do not rely on labels printed inside an image, and do not ask the model to guess which file belongs to which character or object. Put the mapping directly in the prompt.
@Image 1 defines the potter's facial features, hairstyle, and dark-green apron.
Do not use its background.
@Image 2 defines the wooden workbench, window position, and morning light of
the pottery studio. Do not use the person shown in the image.
@Video 1 defines the rhythm of shaping, lifting, and placing the cup. Do not
use the person's identity, clothes, or setting from the video.
The potter finishes a pale-blue cup in the pottery studio, lifts it from the
wheel, and places it in the center of a wooden shelf. Begin with a medium shot,
then slowly push in toward the cup's texture. Keep the wheel, clay, and room sounds.
For multiple views of one product, explicitly say that they define the same object:
@Image 1 defines the front of the same folding desk lamp.
@Image 2 defines its left-side construction.
@Image 3 defines its right-side construction.
@Image 4 defines its rear construction.
All four images define one folding desk lamp. Only one lamp appears in the video.
If a reference video already contains the exact action, camera move, and sequence you want, describe what to inherit instead of restating every motion. A long written action list can conflict with the motion already present in the reference.
Sound, Dialogue, Music, and Subtitle Syntax
Natural language works well, but these optional symbols can separate different sound categories:
(soft rhythmic piano music) Music
<a bell rings in the distance> Sound effect
{Hello, welcome back.} Dialogue
【Chapter One: Departure】 Subtitle
You can also state what to retain or remove:
No background music. Keep only dialogue, room ambience, and action sounds.
No subtitles.
No sound.
For non-English dialogue, name the language before the line. If regional speech matters, include the locale, accent, delivery, speaker, and exact words:
Dialogue language: natural, conversational American English.
The young woman says softly: {I thought you weren't coming.}
Prompting With Many Reference Assets
With a large reference set, organize the prompt in this order:
Asset roles → subject mapping → groups by type → subject profiles → scene-by-scene use
Name each important character, product, prop, and location. Avoid vague mappings such as “Images 1-4 define four characters.” The model needs to know which file belongs to which identity.
【Characters】
The Conservator corresponds to @Image 1. Use only the face, hair, and clothing.
The Registrar corresponds to @Image 2. Use only the face, hair, and clothing.
Do not exchange their appearance, clothes, actions, positions, or dialogue.
【Props】
The Sample Box corresponds to @Image 3 and belongs only to the Conservator.
The Record Board corresponds to @Image 4 and belongs only to the Registrar.
【Locations】
The Conservation Room uses only the space, materials, and light from @Image 5.
The Gallery uses only the space, materials, and light from @Image 6.
【Action references】
@Video 1 defines how the Conservator opens the Sample Box. Do not use the
person or location from the reference video.
When a subject appears across several scenes, add a compact profile:
【Subject profile: Conservator】
Appearance and clothing: @Image 1.
Fixed prop: the Sample Box from @Image 3.
Locations: the Conservation Room and the Gallery.
Action references: the opening action from @Video 1.
Do not use another character's clothing. The Conservator never holds the Record Board.
Then call only the relevant assets in each scene. The goal is correct selection, not forcing all 50 references to appear at once.
How to Prompt a 30-Second Seedance 2.5 Video
For a longer story, divide the prompt into consecutive stages. Give each stage one main change and an observable end state. The end state creates a clean handoff to the next part.
【Generation goal】
Create a [video type]. The main subject is [subject], and the story is [summary].
【Stage 1】
Start: [initial character, prop, and scene state].
Main event: [one action or event].
End: [visible character position, prop ownership, or scene state].
【Stage 2】
Continue from the previous stage: [state that must remain consistent].
Main event: [one action or event].
End: [observable state].
【Stage 3】
Main event: [closing event].
End: [final visible state].
【Continuity】
Keep [identities, number of people, clothes, prop ownership, spatial direction,
and sound relationships] consistent.
Example: flower shop order
Create a flower-shop order-packing video. A Florist and an Assistant arrange,
wrap, and deliver one bouquet.
Stage 1
Start: The Florist stands behind the workbench. Loose stems, scissors, and
wrapping paper lie on the table.
Main event: The Florist arranges the stems and trims them.
End: The bouquet is in the Florist's left hand, and the scissors are back on
the right side of the workbench.
Stage 2
Continue: Both people keep the same identity and clothing. The Florist still
holds the bouquet.
Main event: The Assistant opens the wrapping paper. The Florist places the
bouquet inside and ties a green ribbon around it.
End: The wrapped bouquet lies at the center of the workbench with the bow facing camera.
Stage 3
Main event: The Assistant lifts the bouquet and places it on the pickup shelf.
End: The bouquet is centered on the pickup shelf. Both people stand behind
the workbench and inspect the result.
Continuity: Keep both identities, clothes, workbench orientation, scissor
position, and bouquet ownership consistent.
When to use timestamps
Use stages for normal narrative work. Use one-second time ranges only for important handoffs, entrances, exits, transitions, or beats:
0-5 seconds: Show an empty wooden display. A hand places a white ceramic plate
in the center. At the end, the hand has left the frame and only the plate remains.
5-10 seconds: Remove the plate, then place a clear glass in the center. At the
end, only the glass remains.
10-15 seconds: Remove the glass, then place a green ceramic vase in the center.
At the end, only the vase remains.
Keep time ranges continuous and non-overlapping. They are rhythm budgets, not frame-accurate edit points. Avoid impossible frequency instructions such as “perform three separate actions in one second.”
Parameter Rules for Editing, Keyframes, and Extension
Some task types automatically lock generation settings:
| Task | Aspect ratio | Duration |
|---|---|---|
| Video editing | Inherits the input video; cannot be changed | Approximately inherits the input; cannot be changed and may differ by up to about 0.3 seconds |
| First-frame or first-and-last-frame generation | Uses the first image's ratio | Can be set by the creator |
| Video extension | Inherits the input video; cannot be changed | Extension duration can be set by the creator |
Use the same aspect ratio for first and last frame images. If they differ, the last frame may be stretched.
Seedance 2.5 Video Editing Prompt Template
Treat the original video as the single master. State the target, exact scope, replacement asset, and everything that must remain unchanged.
【Edit goal】
Edit @Video 1. During [the full video or a precise time range], [add, remove,
replace, or adjust] [object, region, or sound category].
【Original video role】
@Video 1 is the only editing master. It defines the characters, setting,
actions, composition, camera, occlusion, sound, and event order.
【Target reference role】
@Image 1 or @Audio 1 defines [specific properties of the new object or sound].
【Edit scope】
Change only [object, region, time range, or sound category].
【Keep unchanged】
Preserve [all unaffected images, actions, sounds, and timing] from @Video 1.
Example: replace one product
Edit @Video 1. Replace only the yellow folding desk lamp with the white folding
desk lamp from @Image 1.
@Video 1 is the only editing master. It defines the desk, books, hand actions,
camera position, camera movement, occlusion, and event order.
@Image 1 defines only the white lamp's appearance, construction, and material.
Do not use its background, composition, or other objects.
There is exactly one white folding desk lamp throughout the video. Replace only
the yellow lamp. Do not change the books, desk, hands, or background.
The white lamp inherits every appearance, arm rotation, hand occlusion, exit,
path, and speed change of the original yellow lamp. Everything else remains unchanged.
The same structure works for background replacement. Change only the area outside the subject's silhouette and explicitly preserve the face, hair, clothing, expression, scale, position, action, foreground objects, camera, and timing.
For sound editing, separate dialogue, language, voice, music, ambience, and effects:
Edit @Video 1. Remove only the original background music. Keep the dialogue,
lip movements, ambience, and action sound effects. Keep the image, camera,
and editing rhythm unchanged.
Extending a Video Forward or Backward
Video extension creates material outside an existing boundary. A backward extension starts from the original last frame; a forward extension ends at the original first frame. In both directions, preserve identity, props, background, motion, camera axis, light, and sound.
Extend after the original ending
@Video 1 is the original video to extend backward.
Extend @Video 1 backward. The extension's first image directly continues its
last frame. Preserve [subject pose and direction], [prop position], [background
and spatial relationships], [camera and composition], [light], [sound state],
and [motion trend].
Then [describe the new action, event, camera move, or sound].
Keep identity, clothing, key props, background layout, camera axis, and the
original sound environment continuous. Each subject remains one continuous
object and never duplicates or splits.
Extend before the original beginning
@Video 1 is the original video to extend forward.
Extend @Video 1 forward. Before the original begins, [describe the preceding
action, event, camera move, or sound].
The extension's final image naturally joins the first frame of @Video 1:
[subject pose and direction], [prop position], and [background relationship]
match the original first frame. Preserve its camera, composition, light, sound,
and motion trend.
Keep identity, clothing, key props, background layout, camera axis, and sound
continuous. Do not introduce any subject that appears only after the original begins.
Aim for a natural visual match, not a pixel-identical boundary. Also check audio across the join: extension volume may vary slightly from the source.
First and Last Frames, Keyframes, and Storyboards
In multimodal reference mode, you can define the first and last frames directly in the prompt:
@Image 1 is the first frame. It defines the opening composition, subject
positions, poses, prop states, setting, and camera direction.
@Image 2 is the last frame. It defines the ending composition, subject
positions, poses, prop states, setting, and camera direction.
@Image 3 defines Subject A's appearance and clothing without changing the
composition defined by @Image 1 or @Image 2.
[Describe one continuous action or event.] Start naturally from @Image 1,
perform the continuous action, and arrive at @Image 2. Keep identities, prop
structure and ownership, scene layout, and camera direction continuous.
Define each anchor separately; do not combine them into “Images 1 and 2 are the first and last frames.”
For several independent keyframes, declare their order first:
Use @Image 1 through @Image 4, in order, as keyframes.
@Image 1 is the first frame: an orange paper airplane rests on the left side
of a classroom desk, pointing screen-right, in a fixed medium shot.
@Image 2 is the second keyframe: one hand lifts the same airplane without
changing its direction.
@Image 3 is the third keyframe: the same airplane passes the window as the
curtain moves slightly to the right.
@Image 4 is the last frame: the airplane rests on the middle shelf at screen-right,
still pointing right.
Move through all four states in order with continuous action. Preserve the
airplane's orange material, scale, folds, flight direction, room layout,
afternoon side light, and camera axis.
Separate keyframe images generally align more reliably than a single collage. If you use a storyboard grid, keep it to roughly 15 panels or fewer, state the reading order, and explain that the grid controls sequence and broad composition rather than the final drawing style or every tiny detail.
Using Rough and Detailed White-Model References
A rough white-model or animatic should provide motion and spatial timing: paths, positions, entrances, camera movement, cuts, lighting changes, and sound rhythm. Map every geometric placeholder to its final subject and use other images for appearance.
@Video 1 is a rough white-model reference. Use only its walking path, subject
positions, fixed camera, one push-in, and two cuts. Do not use its gray geometry
or empty environment.
The tall cylinder in @Video 1 corresponds to the Presenter.
The cuboid corresponds to the Mobile Display Cart.
@Image 1 defines the Presenter's face, blue uniform, and badge.
@Image 2 defines the cart's white metal frame and transparent cover.
@Image 3 defines the curved walls, gray floor, and overhead strip lights of a
technology exhibition hall.
The Presenter pushes the cart along the curved wall, stops at the center display,
and opens the transparent cover. Preserve the reference path, positions, push-in,
and cut points. Render the result as a bright, realistic exhibition documentary.
Keep footsteps, wheel sounds, and exhibition-hall ambience.
A detailed white model already contains finished structure. Use it for re-rendering material, color, character appearance, environment, or style while preserving geometry, motion, camera, and cuts. Remove production overlays such as trajectory lines, axes, controls, and camera cones before uploading when possible.
One-Click Video Assembly
When turning several stills into one complete video, specify:
Asset roles → image order → amount of motion → edit style → visual package → sound
Do not write only “make these images into a video.” Define what each image shows, whether upload order matters, how much animation is allowed, and what must remain stable.
【Asset roles】
@Image 1 is the night-market entrance and opening shot.
@Image 2 shows the Traveler walking along the street.
@Image 3 shows the lantern stall and craft details.
@Image 4 shows three friends eating together.
@Image 5 shows the river at night.
@Image 6 is the final group photo beside the bridge.
@Video 1 defines only the upbeat edit rhythm, transitions, and hand-drawn stickers.
Do not use its people or location.
【Sequence】
Show @Image 1 through @Image 6 in order: arrival, exploration, meal, walk,
and final group photo. Keep all three friends' faces and clothes distinct.
【Motion】
Use slow push-ins and mild parallax for environments. For people, add only
natural blinking, head turns, raised glasses, and gentle clothing movement.
Keep stalls, table positions, and the bridge railing stable.
【Style and sound】
Use a bright travel-film rhythm, natural foreground wipes, related colors,
and stickers only near the frame edges. Keep crowd ambience, tableware sounds,
river wind, and light instrumental music.
Seamless Video Transition Prompt
A transition prompt must define the outgoing clip, incoming clip, trigger, camera movement, transformation, arrival state, and sound bridge.
@Video 1 is the outgoing clip. Use its rainy night street, red umbrella,
slow forward push, and rain sound.
@Video 2 is the incoming clip. Use its circular gallery skylight, upward camera
movement, and quiet interior reverb. Preserve the original people, settings,
structures, and main actions in both clips.
At the end of @Video 1, the red umbrella moves toward the camera until it fills
the frame and triggers the transition. The camera continues forward. The circular
edge of the umbrella gradually becomes the metal ring of the gallery skylight,
while the red fabric becomes white daylight.
Finish at the opening composition of @Video 2. Turn the forward camera move
smoothly into an upward tilt. Fade the rain into gallery footsteps and room reverb.
Generative transitions aim for visual and audio continuity, not pixel-perfect preservation of both source clips.
Direct Emotion Through Visible Behavior
Words such as “tense,” “warm,” or “oppressive” establish a mood but leave performance choices open. For better control, translate emotion into visible or audible behavior: gaze, eyebrows, mouth, breathing, shoulders, hands, and delivery.
Use two to four clear signals for one emotional turn:
The overall emotion changes from [starting emotion] to [ending emotion].
After [trigger], [subject] first shows [immediate observable response].
Then [gaze, brow, mouth, breathing, or hand movement] gradually changes.
Finally, [subject] expresses [target emotion] through [restrained visible behavior].
Example:
Applause from the end of the performance is heard behind the stage. The young
actor's fingers suddenly stop on the program. Their gaze slowly turns toward
the curtain while the shoulders remain tense. After confirming the curtain call,
the actor exhales, the shoulders relax, a restrained smile appears, and the eyes
slowly fill with tears, but the actor does not turn away.
For several emotional changes, tie each new reaction to a specific trigger instead of listing many facial details at once.
Camera Language That Seedance 2.5 Can Follow
Common terms can be written directly:
- Shot size: extreme wide shot, wide shot, medium shot, close-up, extreme close-up
- Movement: push in, pull out, pan, truck, follow, orbit, dive, crane back, tilt up, subtle handheld motion
- Angle: low angle, overhead view, first-person view
- Popular techniques: one continuous take, dolly zoom, aerial view, FPV flight, bullet time, handheld tracking, speed ramp and rebound
If a term is unusual or ambiguous, translate it into an observable result:
Rack focus: shift focus smoothly from the leaves in the foreground to the
person in the background. The leaves become soft while the person's face
changes from blurred to sharp.
For exact transitions, define the trigger, occluding object, movement direction, cut, and post-transition motion:
At second 5, whip-pan quickly left. Cut when the foreground bookshelf covers
the entire frame. In the next scene, continue moving left at a similar speed.
Lens, aperture, and shutter values can help, but describing their visible effect is usually more reliable than providing a number alone.
Seedance 2.5 Prompt Checklist
Before generating, confirm that your prompt answers these questions:
- Is the main subject and action clear?
- Does every reference say what to use and, where needed, what to ignore?
- Is each person, product, and prop named and bound to the correct reference?
- Does each scene call only the assets it needs?
- Does each long-video stage contain one main change and a visible end state?
- Are character count, clothing, prop ownership, and spatial relationships stable?
- Does an edit define one master, the exact scope, target quantity, and unchanged content?
- Are abstract emotions and technical camera terms translated into observable behavior?
- Is each first, last, or intermediate keyframe defined separately?
- Do first and last images use the same aspect ratio?
- Does a storyboard state its reading order and a white model state what timing or structure to inherit?
- Does an extension preserve the boundary image, motion trend, and sound continuity?
- Does one-click assembly define roles, order, motion strength, editing style, and sound?
- Does a seamless transition define both source clips, its trigger, transformation, arrival, and audio bridge?
Important Limitations
Keep these boundaries in mind when evaluating a result:
- Timestamps allocate narrative rhythm; they are not frame-accurate edit points.
- Editing prompts improve alignment with the original, but cannot guarantee identical frames.
- Multimodal prompting is about selecting and combining the right assets, not displaying all assets simultaneously.
- For text, formulas, signage, product specifications, or exact frame-level events that must be perfect, use prepared source assets and finish the work in post-production.
- Video editing locks the input ratio and approximate duration; output can differ by up to about 0.3 seconds.
- First/last-frame generation locks the first image's aspect ratio. Mismatched frame ratios can stretch the final frame.
- Video extension locks the source aspect ratio, and the extension's volume may differ slightly.
- If image order or character mapping matters in one-click assembly, state it explicitly.
- A seamless transition aims for natural continuity, not pixel-identical preservation.
Final Seedance 2.5 Prompt Template
Use this all-purpose structure as your starting point:
【Goal】
Create [video type] about [main subject and story].
【Reference roles】
@Image 1 defines [subject and exact attributes]. Do not use [unwanted content].
@Video 1 defines [action, camera, timing, or spatial relationship]. Do not use
[unwanted identity, clothing, or environment].
@Audio 1 defines [speaker, voice, dialogue, ambience, or music].
【Subject and scene】
[Named subject] performs [main action] in [scene and environment].
[State important identity, quantity, prop ownership, and spatial rules].
【Sequence】
Stage 1: [start, one main event, visible end state].
Stage 2: [continued state, one main event, visible end state].
Stage 3: [closing event and final state].
【Look and camera】
[Lighting, color, materials, atmosphere, shot size, camera position, movement,
focus, and transitions].
【Sound】
[Dialogue language and delivery, ambience, sound effects, and music].
【Continuity】
Keep [identity, appearance, clothing, object count, prop ownership, scene layout,
camera axis, motion direction, and audio relationships] consistent.
The best Seedance 2.5 prompts are not necessarily the longest. They are the ones that give every subject and reference a clear role, break long events into observable stages, and state exactly what should remain consistent from the first frame to the last.
