Guides · Video models · Updated August 10, 2026

The Seedance 2.5 prompting guide

How to write prompts that actually get the shot — the structure, the dialogue rules, the audio brief, and the failure modes, from running Seedance 2.5 in production every day.

Seedance 2.5 is ByteDance's frontier video model, and the headline is genuinely unusual: up to 30 seconds of continuous video in a single generation, with dialogue, sound effects and music generated natively alongside the picture. Most models give you five to ten seconds and silence.

That changes how you write. A five-second prompt is a description. A thirty-second prompt is a shot brief — it has an opening, a middle and an end, and you have to direct all three. Below is what works, what breaks, and the fixes.

1. Open with the shot structure

Before anything else, tell the model what it's making. One line, at the very top:

Total: 30s / 1 continuous shot / 9:16

This is the single highest-leverage line in a long prompt. Without it the model decides for itself how many shots your 30 seconds contains, and the answer is usually "several, at random." With it, you've set the contract.

If you want cuts, say so — Total: 20s / 4 shots / 9:16 — and Seedance will hold character and style consistency across them. Three or four cuts inside one generation is well within its range, and holding identity across a cut is something 2.5 is specifically built for.

Then label your beats to match what you asked for, because the labelling itself steers the model. ByteDance's own examples number their shots — S1, S2, S3 — with an explicit hard cut to between them. That's the right shape when you want an edit. It's the wrong shape when you don't: numbered shots invite cuts, so for a continuous take use timecodes instead and say outright that each window is the same camera still moving.

You wantLabel beats asAnd add
A cut sequenceS1・S2・S3…hard cut to between shots
One continuous take[0-4s], [4-8s]"the SAME camera, still moving — do not cut"

2. Write a shot brief, not a mood board

The most common bad prompt is a pile of adjectives. Words like beautiful, cinematic, epic and stunning give the model nothing it can point a camera at. Verbs and consequences do.

Instead ofWrite
A stunning cinematic shot of a woman with coffeeShe sets the mug down on the counter, and it slides an inch and tips
Epic dramatic lightingHard side light from a window on the left, deep shadow on the far wall
She looks emotionalHer jaw tightens, she looks away from the lens, then back
A beautiful forest sceneShe steps through wet ferns; each step scatters leaves and drops of water

The rule underneath: name the physical consequence. "Beads of condensation sliding down the aluminium" is worth more than a paragraph of atmosphere, because the model can render it and it can't render "premium".

Labelled lines beat a paragraph. This is the shape our own ad engine writes in, because one element per line is easier for the model to parse and far easier for you to tweak — you can change the camera without rewriting the sentence the action lives in:

Subject: the presenter holding the bottle up toward camera
Action: slight squint from the sun; deadpan, direct-to-lens delivery
Camera: eye-level, slow push-in
Composition: medium close-up, portrait
Focus: shallow, on her face
Sound: soft room tone

Subject, Action, Camera and Composition earn their place every time; Focus, Ambiance and Sound only when they add something. Each value is a descriptive phrase, not a bare keyword.

For anything past ten seconds, go one level up: write in blocks. ByteDance's own published example prompts for 2.5 are long — several hundred words — and they aren't paragraphs. They're a stack of labelled blocks, and the beats come last:

【TITLE】【One-line logline of the whole film】

LOOK — genre, film stock, lens, grade, grain, texture.
COLOUR — the palette, as a 60:30:10 split.
SUBJECT — who or what, described once, in full.
PHYSICS — how mass, contact and materials should behave.
CAMERA — lens character, handheld vs crane, and the speed plan.
AUDIO — the sound design and music brief.

[beats, in order]

POSITIVE CONSTRAINTS — every requirement above, restated as positives.

Three of those are worth calling out, because most people never think to write them:

  • COLOUR, as a ratio. ByteDance's examples specify palette as 60:30:10 — 60% base environment tone, 30% the light source of the scene, 10% hot accent. It's a real art-direction convention, and giving the model a ratio instead of a list of colours is what stops all three fighting for the frame.
  • PHYSICS, named explicitly. Steadier physics is one of 2.5's headline upgrades, and you can address it directly: "real gravity throughout; his weight is palpable; every impact distinct; no floatiness." Saying it out loud measurably helps on anything involving mass, impact, cloth or water.
  • POSITIVE CONSTRAINTS, at the end. See §12 — this one is load-bearing and it looks redundant.

Not every prompt needs all of this. A six-second clip doesn't. A twenty-second one with an arc does, and the block form is also just far easier to edit: you change the grade without touching the choreography.

3. Direct the camera, every time

Seedance reads standard camera language cleanly, so use it. Left unspecified, the camera drifts — usually a slow push-in it decided on by itself.

One move per moment, from this vocabulary:

locked-off · slow push-in · pull-out · pan · tilt · dolly · tracking / follow · orbit / arc · crane · whip-pan · rack focus · handheld drift · low angle · eye-level · high angle · over-the-shoulder

State the framing alongside it — "medium, waist up, centered" — or you'll get a different shot size every time you re-roll.

And for a continuous take, say what the camera is not doing. This is the one place a negative genuinely helps:

ONE CONTINUOUS TAKE — no cuts, no fades, no zoom, one unbroken camera move.

4. Keep dialogue short, quoted, and characterised

Native dialogue is the best thing about this model and the easiest to break.

Put spoken lines in double quotation marks. That's what tells the model to voice them instead of rendering them as on-screen text.

Budget about 2.5 words per second. A 4-second beat holds ~10 words, 8 seconds holds ~20. Overflow doesn't get squeezed — it gets cut off mid-word.

Long unbroken speeches drift out of lip-sync. If a line won't fit, split it across two beats rather than stretching it.

Describe the voice once, concretely, then keep that description byte-identical everywhere it appears. Change the words and you re-cast the actor:

DEV — South Asian man, mid-thirties, MUSTARD-YELLOW work polo. Voice: warm,
easygoing American man, light slightly nasal mid register, unhurried, completely
unbothered — no announcer.

Emotion goes in the action, not the voice line — "delivery dry and grudging" — so the character stays the same person while the reading changes.

Then describe how the line is said, the way a novel would. A tag after the line is the most direct control you have over a performance, and it costs nothing:

Marcus, spoken aloud and lip-synced: "Is that a... dinosaur?" — he says it slowly
and flatly, more puzzled than alarmed, the pause landing like he is
double-checking his own eyes.

Two rules make this safe rather than dangerous:

  • The tag lives OUTSIDE the quotation marks. Always. The model voices what's inside the quotes, so a tag that slips in becomes something your character says out loud. Closing quote, then the dash, then the tag. If a render ever does speak one, move it up into that beat's action sentence instead — same information, no risk.
  • Tags carry the emotion; the character block carries the voice. Never write a tag that re-casts the speaker — a whisper becoming a shout, a new accent, a different register. Within one clip that's a different person, and the model will render it as one.

Then write the line like speech, because it's going to be spoken. Three rules do most of the work here, and the first one catches almost everybody:

  • Full sentences, not headlines. Verbless ad-copy fragments — "Twelve hours of wear. Zero touch-ups." — read as on-screen captions, and voiced aloud they land as bullet points. Give the line a subject and a verb the way a person would actually say it: "I put this on at seven and I still haven't touched it." Short is good; fragmentary isn't. The two aren't in tension — "It takes two days." is both.
  • Never say what the viewer can already see. The picture has the setting and the action covered. A line that re-describes them spends its whole moment saying nothing — spend it on what the frame can't show: the stake, the reason, the thing that happened offscreen.
  • Plain words, no jargon. A "jacket", not a "shell". "It keeps the rain out", not "a waterproof membrane". Precise material vocabulary still belongs in the action and the seed image, where models reward it — just not in someone's mouth.

5. Treat audio as a brief, not an afterthought

Audio is co-generated with the picture in a single pass — the same forward pass, not a dub — which is why lip-sync, footsteps and musical hits land on the right frames by default. It also means you can't fix it afterwards, so brief it up front.

Don't rely on the default in either direction. Depending on who you ask, an unspecified audio brief either gets you a generic orchestral bed like a car advert, or gets background music actively suppressed. Both behaviours are documented by different sources, which tells you all you need to know: say what you want. You have four levers.

Sound effects. Name them per moment. "Shattering glass, a low impact boom, ringing" beats "loud noises".

Silence, used as a sound. This is the lever nobody pulls, and ByteDance's own example prompts lean on it hard — they mark the audio state per shot and slam the extremes against each other: a screaming close-up marked (DEAFENING) cutting straight to a wide shot marked (COMPLETE SILENCE). Dropping out is a real instruction and the model honours it:

[3-6s] ... Sound: a massive scream, wind roar, hull groan — BLASTING loud.
[6-8s] ... The frame goes completely silent — no sound at all, total stillness.

Silence also has a physical excuse in a lot of shots, which makes it land harder: vacuum, underwater, a held breath, the beat after an explosion when hearing has gone. Ask for a crescendo and you get loudness; ask for a drop-out and you get dynamics.

Music. Either brief it or ban it, but never ignore it. To brief it, describe genre, tempo, instrumentation and arc:

MUSIC — exactly two states and ONE change.
1. [0-4s] Bland convenience-store muzak through a tinny ceiling speaker, low
   in the mix, under the dialogue.
2. [4-30s] 120 BPM electronic rock — reverb-soaked surf-guitar lead, four-on-
   the-floor drums, distorted synth bass. Starts lean and BUILDS in layers,
   full power by the final explosion. No vocals, no orchestra.

To ban it, say it once, in the header — not on every line:

AUDIO — diegetic sound only: dialogue and sound effects in the room. Do not
compose or perform any musical score, soundtrack or backing track.

"Once" means once per generation, not once per project. Inside a single clip the instruction has scope over the whole thing, and repeating it eleven times just puts eleven units of weight on the idea of music. But if you're building a video from several clips, each one is a separate generation that inherits nothing — so the exclusion has to be restated in every prompt, along with anything else that must hold across them (wardrobe colours, the voice description, a held prop). Same rule, opposite conclusion, depending on whether you're writing one prompt or many.

6. Timecode anything over ten seconds

Past about ten seconds, a single paragraph of description gets averaged into mush — the model spreads your whole idea evenly across the runtime instead of sequencing it. Break it into labelled windows:

[0-4s]  Normal speed, calm. Two men chat by the slushie machine...
[4-6s]  DETONATION. The glass doors blow inward...
[6-9s]  Tight, low, handheld at floor level. Both men crouched...

Each window gets its own action, its own camera, its own sound. Order inside the sentence is the choreography: "she sets the mug down, turns to the shelf, and lifts the box into the light."

Write timecodes as [0-3s], and nothing else. Square brackets, whole seconds, no colons, no minutes, no "s" on the first number. This isn't a style preference — Gemini Omni parses that exact shape as a native timing instruction, while 0:00-0:07 reads as prose and is silently ignored. Seedance and Kling have no timecode parser at all, but the same shape is the clearest phrasing for them too, so one format serves every model.

✅[0-3s] · [3-7s]
❌0:00-0:03 · 00:07 · 7s-12s · [0s-3s]

Every clip's timing starts at 0. The numbers are offsets into this clip, never positions in the finished video. If you're cutting a longer script into clips, re-base each one: a beat you wrote as 0:12-0:14 that becomes a two-second clip is [0-2s]. Carried through unchanged it tells the model to do nothing for twelve seconds of a two-second clip.

Only timecode ordered beats. A single continuous action needs no markers at all, and a [0-8s] header on an 8-second clip carries no information the clip doesn't already have.

7. Chaining clips: a clip must never end

If you're building something longer than one generation, the usual method is frame-chaining — take the last frame of clip 1, feed it as the start frame of clip 2, and the two read as one continuous take. It works well, and it fails in one specific, maddening way.

The model doesn't know it's making a middle. Generate a clip on its own and it will tack a cinematic ending onto it — a whip-pan, a crash zoom, motion blur, a drop into slow motion, a fade toward black. That final frame is then the literal first frame of your next clip, so the ending flourish smears straight into it and the seam shatters.

So for every clip except the last, direct the ending explicitly:

End the clip on a sharp, in-focus, steadily-framed moment with calm motion that
simply continues — no motion blur, no whip-pan, no zoom, no slow motion, no fade.

Only the final clip is allowed to settle or resolve.

And keep hero objects in frame the whole way. In a chained take, an object that drifts out of frame doesn't come back — the next clip is generated from a frame it isn't in, so it vanishes or morphs into something else and can't be recovered. If a product is the point of the video, it stays visible in every clip: held, worn, or clearly placed in shot.

Give the ear a breath at each join too. There's no hard cut to punctuate the seam, so if one clip's speech runs to its last frame and the next opens on speech, the two phrases butt together and it reads as a machine, not a person. Every join wants silence on one side — either the outgoing clip ends with a beat of quiet, or the incoming one opens on a small non-speaking action (a glance down, a step) before the line starts.

8. Lock your characters and props

Anything that must stay the same needs a distinct, named colour: "a mustard-yellow work polo, distinct from his navy-blue jacket." Vague clothing blends between characters as the take runs.

The same goes for state that evolves. If your subject gets progressively dirtier, say so as an arc rather than hoping:

Dust, then blue slush, then soot accumulate on both men steadily across the take
until they are filthy at the end.

9. On-screen text: colour it, place it, hold it

In-frame text is one of the places 2.5 genuinely pulled ahead: it renders titles, captions and subtitles inside the frame, across major languages including English, Spanish, Chinese, Arabic, Japanese and Korean. If you're cutting the same spot for several markets, that's a typesetting pass you no longer need.

It's still the least reliable thing in any video model, though — "better" isn't "solved". It works far more often when you give it four things: the exact words in quotes, the placement, the weight, and the colour contrast.

The words "50% OFF" in bold white block capitals on a solid red banner, centred
in frame and held legible for a beat before the wave hits it.

Three things earn their keep there. Quotes stop the words being spoken aloud. Explicit contrast (white on red) stops the model rendering low-contrast mush. And "held legible for a beat" matters more than people expect — if text is going to be destroyed, obscured or moved, say that it reads clearly first, or the model will start the effect on frame one and you'll never see the words at all.

Keep strings short. Two or three words render reliably; a sentence rarely does.

10. Photoreal creatures need an explicit guard

Ask for an animal and you'll often get something that looks rendered. If you want a real one, say what it isn't — the one other place a negative pays off:

The dinosaur and the shark are REAL ANIMALS, filmed practically — no 3D,
no cartoon, no VFX, no CGI look.

Pair it with a film-stock cue in the style line: 35mm film quality, ARRI ALEXA aesthetic, photorealistic, film grain, motion blur on all fast action.

11. The image-to-video aspect trap

This one catches everybody, and it's specific to 2.5.

On image-to-video, Seedance 2.5 takes its aspect ratio from your start frame. It does not accept 9:16 or 16:9 as an instruction — the frame decides. So if you want a vertical video, generate a vertical seed image. Feeding a 16:9 still and asking for 9:16 in the prompt gets you a 16:9 video.

Text-to-video and reference-to-video take the full aspect list normally. It's only the start-frame path that inherits.

While you're there, the seed frame is worth over-writing. It's the only place identity is described, and everything downstream inherits it — so a vague seed costs you every clip, not just the first. Four things pay for themselves:

  • Say "portrait" and name facial detail for a person — "warm three-quarter portrait, soft catchlights in the eyes, light freckles, relaxed half-smile". Image and video models both reward the explicit portrait cue, and it's what holds the same face across clips.
  • Describe a product's physical FORM, not its brand — "a soft squeezable tube with a flip cap", not "the moisturiser". The precise form is what stops the model restyling it into a different object two clips later.
  • Compose for motion to begin — clean eyeline, headroom, and not mid-blink or frozen in an action pose. It's a first frame, not a hero still.
  • Keep added text out of frame — captions, UI, watermarks. A product's own real label and logo are part of the product and should render faithfully; anything else you overlay later is something the model will fight you on.

12. What the model won't do

You might expectReality
A negative_prompt fieldDoesn't exist on Seedance. Fold exclusions into the prompt — and close with a positive-constraints block (below)
A multi-shot / beat arrayNot on 2.5. Use timecodes, or Cut to: for real cuts
1080p or 4K480p and 720p only on 2.5. Use Seedance 2.0 for 1080p
Slow-motion rampsSupported, and ByteDance documents the phrasing (below). Still the first thing to check on render one

The missing negative prompt has a real workaround, and it's the last block in ByteDance's own examples. After the beats, they restate every requirement — as positives:

POSITIVE CONSTRAINTS — one unbroken continuous camera for the full 15 seconds,
following one single ship the whole way down. The same ship, unchanged in design
and scale, in every beat. Live-action photographic film texture, real physics,
real weight. Sound effects and one rising music cue only; no dialogue, no
voice-over, no subtitles.

It reads redundant and it isn't. With no negative field, repeating a constraint at the end is the only way to put extra weight on it — and phrasing it as what you do want ("photographic film texture") outperforms the thing you don't ("not CGI"), because models are poor at negation but excellent at description. Reserve it for the two or three things that would ruin the clip if they drifted.

And speed ramps are better supported than they look. ByteDance's published examples use them explicitly, so copy their phrasing — a global plan in the camera block, then a per-beat marker:

CAMERA — speed: mostly real-time; only two brief slow-motion moments — the instant
the cable takes his weight, and the instant the window shatters.

[5-7s] ... Speed ramp: brief slow motion as the cable takes his weight → back to
real-time as he swings.

Name the instant the ramp starts and the instant it ends, and say the speed you're returning to. A bare "slow motion here" is what gets ignored. If a ramp still matters and still won't land, shoot the beat at normal speed and retime it on the timeline — frame-accurate, and free.

And never put dialogue inside a slow-motion window. A lip-synced line rendered in slow motion drags the voice into a groan. Keep speech in normal-speed beats.

13. Budget your re-rolls

Thirty seconds of 720p is a real generation. Three things follow.

Draft at 480p. Settle the prompt on the cheap tier and spend 720p once, on a version you've already chosen — see the note at the top. Four 480p drafts plus one 720p finish costs less than two 720p rolls of the dice.

Front-load the risk. Judge render one on the single hardest thing in your prompt — the whispered line, the rendered text, the two-character lip-sync — not on the overall vibe. If the hard part fails, everything else is noise.

Simplify rather than repeat. Three failed re-rolls of an over-stuffed prompt cost more than one clean run of a simpler one. If a 30-second take with four set-pieces keeps collapsing, split it at a natural seam into two 15-second clips and join them on the timeline. You keep the look and buy back control.

The skeleton

Copy this and fill it in.

Total: <N>s / <N> shots / <aspect>

<ONE CONTINUOUS TAKE — no cuts, no fades, no zoom.  |  or omit for cuts>

Style: <genre reference>. <film stock, grade, grain, lens character>.
Photorealistic.

COLOUR: <60 base tone : 30 light source : 10 hot accent>

SETTING: <location, time of day, light sources, two or three hero props>

<CHARACTER> — <look, age, LOCKED COLOUR wardrobe>. Voice: <age + gender read,
accent, register, cadence, energy, anti-pattern>.

<state arc that evolves across the take, if any>

PHYSICS — <how mass, contact and materials should behave>       <-- if anything
                                                     heavy, wet or breakable moves

AUDIO — <"diegetic only, no score" OR the music brief>
SPEED — <"real-time throughout" OR the ramp plan, naming each instant>

[0-Xs] <Speed.> <Action, in the order it happens.> <ONE camera move + framing.>
<Character>, spoken aloud and lip-synced: "<line, ≤2.5 words per second>" — <how
they say it, outside the quotes>
Sound: <named diegetic effects, and any deliberate silence>

[X-Ys] ...

POSITIVE CONSTRAINTS — <the two or three things that must not drift, restated as
positives>

Worked example

Here's the opening of a 30-second spot we generated with 2.5 — two men in a convenience store, a robbery, and an escalating disaster they refuse to acknowledge. Nine shots and eight hard cuts, all out of one generation.

Note the shot header, the locked wardrobe colours, the music brief stated once rather than per beat, the delivery tags sitting outside the quotes, and how short the spoken lines are.

Total: 30s / 9 shots / 9:16

An edited action sequence. Hard cut wherever it says CUT TO; one continuous camera
within each shot. Straight cuts only — no fades, no dissolves. Change the shot size
and the camera move at every cut so the edit has rhythm.

LOOK: a big-budget American action film shot on 35mm anamorphic glass. Teal-and-
orange grade, deep shadows, blue lens flares, film grain, sparks and floating
embers. Photoreal throughout.

SETTING: a neon-lit American convenience store at night. Magenta and cyan
fluorescent tubes, a wall of glowing drinks coolers, a chrome slushie machine,
glass front doors onto a black parking lot.

HECTOR — Latino man, mid-thirties, stocky, thick black moustache, a MUSTARD-YELLOW
work polo stretched tight, holding a tall cup of electric-blue slush. Voice: warm,
easygoing American man, light Latino-American accent, unhurried and totally
unfazed — no announcer.

MARCUS — lean Black man, early forties, neat beard, thin gold chain, a crisp open
NAVY-BLUE bomber jacket. Voice: dry, smooth American man, low chest register, slow
deliberate cadence, flat and skeptical — no announcer.

Dust, then blue slush, then soot build on them across the take until they are
filthy at the end.

AUDIO — dialogue, sound effects AND music are all generated natively in this take.
MUSIC — exactly two states and ONE change. [0-4s] bland store muzak through a
tinny ceiling speaker, low under the dialogue. [4-30s] the muzak is gone and a
120 BPM electronic rock track takes over — surf-guitar lead, driving drums, builds
in layers to full power. No vocals, no orchestra.

[0-4s] Normal speed, calm and mundane. The two men stand chatting by the slushie
machine. Hector takes a pull on his straw and speaks. The camera drifts in slowly.
Nothing is wrong.
Hector, spoken aloud and lip-synced: "So did you hear Seedance two-point-five's on
TurboClip now?" — he says it offhand, barely looking up from his straw, the way you
mention the weather.
Sound: dialogue, a humming cooler and the slushie machine churning.

CUT TO:
[4-6s] DETONATION — a hard cut to the front of the store. WIDE, low angle, locked
off on the glass doors as they blow inward in a wall of shattering glass and the
robber bursts through, a black silhouette against blinding white headlights. The
glass hangs in brief slow motion, then time snaps back and he lands hard.
The robber, screaming at the top of his lungs, spoken aloud and lip-synced:
"HANDS UP!" — he shrieks it, voice cracking on the second word, far more frightened
than frightening.
Sound: shouted dialogue, a huge glass explosion, a low impact boom, ringing.

The short version

  1. Open with Total: <duration> / <shots> / <aspect>.
  2. Verbs and physical consequences, never adjectives.
  3. One named camera move and framing per moment.
  4. Dialogue in quotes, ~2.5 words per second, voice described once and never reworded.
  5. Say how the line is said — in a tag outside the closing quote.
  6. Brief the audio, and brief the silence too. Never leave it to the default.
  7. Timecode anything over ten seconds.
  8. Lock wardrobe and props with distinct named colours.
  9. On-screen text: quoted, high-contrast, short, and held legible before anything happens to it.
  10. no 3D, no cartoon, no VFX for photoreal animals.
  11. Vertical video needs a vertical seed image — image-to-video inherits the aspect from your start frame.
  12. Timecodes are [0-3s], always counting from 0 within that clip.
  13. Numbered shots (S1, S2) invite cuts; timecodes read as one take. Label for the film you want.
  14. Close a long prompt with a positive-constraints block — there's no negative prompt to fall back on.
  15. Chaining clips? Every clip but the last must end on a steady, in-focus frame — and anything that leaves frame never comes back.
  16. Draft at 480p. Deliver at 720p.

Seedance 2.5 is live in TurboClip now — pick it in the model menu on any video or ad generation. Start a project or see what it costs.