Guides · Video models · Updated August 8, 2026

The Seedance 2.5 prompting guide

How to write prompts that actually get the shot — the structure, the dialogue rules, the audio brief, and the failure modes, from running Seedance 2.5 in production every day.

Seedance 2.5 is ByteDance's frontier video model, and the headline is genuinely unusual: up to 30 seconds of continuous video in a single generation, with dialogue, sound effects and music generated natively alongside the picture. Most models give you five to ten seconds and silence.

That changes how you write. A five-second prompt is a description. A thirty-second prompt is a shot brief — it has an opening, a middle and an end, and you have to direct all three. Below is what works, what breaks, and the fixes.

1. Open with the shot structure

Before anything else, tell the model what it's making. One line, at the very top:

Total: 30s / 1 continuous shot / 9:16

This is the single highest-leverage line in a long prompt. Without it the model decides for itself how many shots your 30 seconds contains, and the answer is usually "several, at random." With it, you've set the contract.

If you want cuts, say so — Total: 20s / 4 shots / 9:16 — and Seedance will hold character and style consistency across them. Three or four cuts inside one generation is well within its range.

2. Write a shot brief, not a mood board

The most common bad prompt is a pile of adjectives. Words like beautiful, cinematic, epic and stunning give the model nothing it can point a camera at. Verbs and consequences do.

Instead ofWrite
A stunning cinematic shot of a woman with coffeeShe sets the mug down on the counter, and it slides an inch and tips
Epic dramatic lightingHard side light from a window on the left, deep shadow on the far wall
She looks emotionalHer jaw tightens, she looks away from the lens, then back
A beautiful forest sceneShe steps through wet ferns; each step scatters leaves and drops of water

The rule underneath: name the physical consequence. "Beads of condensation sliding down the aluminium" is worth more than a paragraph of atmosphere, because the model can render it and it can't render "premium".

Labelled lines beat a paragraph. This is the shape our own ad engine writes in, because one element per line is easier for the model to parse and far easier for you to tweak — you can change the camera without rewriting the sentence the action lives in:

Subject: the presenter holding the bottle up toward camera
Action: slight squint from the sun; deadpan, direct-to-lens delivery
Camera: eye-level, slow push-in
Composition: medium close-up, portrait
Focus: shallow, on her face
Sound: soft room tone

Subject, Action, Camera and Composition earn their place every time; Focus, Ambiance and Sound only when they add something. Each value is a descriptive phrase, not a bare keyword.

3. Direct the camera, every time

Seedance reads standard camera language cleanly, so use it. Left unspecified, the camera drifts — usually a slow push-in it decided on by itself.

One move per moment, from this vocabulary:

locked-off · slow push-in · pull-out · pan · tilt · dolly · tracking / follow · orbit / arc · crane · whip-pan · rack focus · handheld drift · low angle · eye-level · high angle · over-the-shoulder

State the framing alongside it — "medium, waist up, centered" — or you'll get a different shot size every time you re-roll.

And for a continuous take, say what the camera is not doing. This is the one place a negative genuinely helps:

ONE CONTINUOUS TAKE — no cuts, no fades, no zoom, one unbroken camera move.

4. Keep dialogue short, quoted, and characterised

Native dialogue is the best thing about this model and the easiest to break.

Put spoken lines in double quotation marks. That's what tells the model to voice them instead of rendering them as on-screen text.

Budget about 2.5 words per second. A 4-second beat holds ~10 words, 8 seconds holds ~20. Overflow doesn't get squeezed — it gets cut off mid-word.

Long unbroken speeches drift out of lip-sync. If a line won't fit, split it across two beats rather than stretching it.

Describe the voice once, concretely, then keep that description byte-identical everywhere it appears. Change the words and you re-cast the actor:

DEV — South Asian man, mid-thirties, MUSTARD-YELLOW work polo. Voice: warm,
easygoing American man, light slightly nasal mid register, unhurried, completely
unbothered — no announcer.

Emotion goes in the action, not the voice line — "delivery dry and grudging" — so the character stays the same person while the reading changes.

Then write the line like speech, because it's going to be spoken. Three rules do most of the work here, and the first one catches almost everybody:

  • Full sentences, not headlines. Verbless ad-copy fragments — "Twelve hours of wear. Zero touch-ups." — read as on-screen captions, and voiced aloud they land as bullet points. Give the line a subject and a verb the way a person would actually say it: "I put this on at seven and I still haven't touched it." Short is good; fragmentary isn't. The two aren't in tension — "It takes two days." is both.
  • Never say what the viewer can already see. The picture has the setting and the action covered. A line that re-describes them spends its whole moment saying nothing — spend it on what the frame can't show: the stake, the reason, the thing that happened offscreen.
  • Plain words, no jargon. A "jacket", not a "shell". "It keeps the rain out", not "a waterproof membrane". Precise material vocabulary still belongs in the action and the seed image, where models reward it — just not in someone's mouth.

5. Treat audio as a brief, not an afterthought

Leave audio unspecified and Seedance scores your clip like a car advert — a generic orchestral bed, every time. You have three levers, and you should pull all three deliberately.

Sound effects. Name them per moment. "Shattering glass, a low impact boom, ringing" beats "loud noises".

Music. Either brief it or ban it, but never ignore it. To brief it, describe genre, tempo, instrumentation and arc:

MUSIC — exactly two states and ONE change.
1. [0-4s] Bland convenience-store muzak through a tinny ceiling speaker, low
   in the mix, under the dialogue.
2. [4-30s] 120 BPM electronic rock — reverb-soaked surf-guitar lead, four-on-
   the-floor drums, distorted synth bass. Starts lean and BUILDS in layers,
   full power by the final explosion. No vocals, no orchestra.

To ban it, say it once, in the header — not on every line:

AUDIO — diegetic sound only: dialogue and sound effects in the room. Do not
compose or perform any musical score, soundtrack or backing track.

"Once" means once per generation, not once per project. Inside a single clip the instruction has scope over the whole thing, and repeating it eleven times just puts eleven units of weight on the idea of music. But if you're building a video from several clips, each one is a separate generation that inherits nothing — so the exclusion has to be restated in every prompt, along with anything else that must hold across them (wardrobe colours, the voice description, a held prop). Same rule, opposite conclusion, depending on whether you're writing one prompt or many.

6. Timecode anything over ten seconds

Past about ten seconds, a single paragraph of description gets averaged into mush — the model spreads your whole idea evenly across the runtime instead of sequencing it. Break it into labelled windows:

[0-4s]  Normal speed, calm. Two men chat by the slushie machine...
[4-6s]  DETONATION. The glass doors blow inward...
[6-9s]  Tight, low, handheld at floor level. Both men crouched...

Each window gets its own action, its own camera, its own sound. Order inside the sentence is the choreography: "she sets the mug down, turns to the shelf, and lifts the box into the light."

Write timecodes as [0-3s], and nothing else. Square brackets, whole seconds, no colons, no minutes, no "s" on the first number. This isn't a style preference — Gemini Omni parses that exact shape as a native timing instruction, while 0:00-0:07 reads as prose and is silently ignored. Seedance and Kling have no timecode parser at all, but the same shape is the clearest phrasing for them too, so one format serves every model.

[0-3s] · [3-7s]
0:00-0:03 · 00:07 · 7s-12s · [0s-3s]

Every clip's timing starts at 0. The numbers are offsets into this clip, never positions in the finished video. If you're cutting a longer script into clips, re-base each one: a beat you wrote as 0:12-0:14 that becomes a two-second clip is [0-2s]. Carried through unchanged it tells the model to do nothing for twelve seconds of a two-second clip.

Only timecode ordered beats. A single continuous action needs no markers at all, and a [0-8s] header on an 8-second clip carries no information the clip doesn't already have.

7. Chaining clips: a clip must never end

If you're building something longer than one generation, the usual method is frame-chaining — take the last frame of clip 1, feed it as the start frame of clip 2, and the two read as one continuous take. It works well, and it fails in one specific, maddening way.

The model doesn't know it's making a middle. Generate a clip on its own and it will tack a cinematic ending onto it — a whip-pan, a crash zoom, motion blur, a drop into slow motion, a fade toward black. That final frame is then the literal first frame of your next clip, so the ending flourish smears straight into it and the seam shatters.

So for every clip except the last, direct the ending explicitly:

End the clip on a sharp, in-focus, steadily-framed moment with calm motion that
simply continues — no motion blur, no whip-pan, no zoom, no slow motion, no fade.

Only the final clip is allowed to settle or resolve.

And keep hero objects in frame the whole way. In a chained take, an object that drifts out of frame doesn't come back — the next clip is generated from a frame it isn't in, so it vanishes or morphs into something else and can't be recovered. If a product is the point of the video, it stays visible in every clip: held, worn, or clearly placed in shot.

Give the ear a breath at each join too. There's no hard cut to punctuate the seam, so if one clip's speech runs to its last frame and the next opens on speech, the two phrases butt together and it reads as a machine, not a person. Every join wants silence on one side — either the outgoing clip ends with a beat of quiet, or the incoming one opens on a small non-speaking action (a glance down, a step) before the line starts.

8. Lock your characters and props

Anything that must stay the same needs a distinct, named colour: "a mustard-yellow work polo, distinct from his navy-blue jacket." Vague clothing blends between characters as the take runs.

The same goes for state that evolves. If your subject gets progressively dirtier, say so as an arc rather than hoping:

Dust, then blue slush, then soot accumulate on both men steadily across the take
until they are filthy at the end.

9. On-screen text: colour it, place it, hold it

Rendered lettering is the least reliable thing in any video model, but it works far more often when you give it four things: the exact words in quotes, the placement, the weight, and the colour contrast.

The words "50% OFF" in bold white block capitals on a solid red banner, centred
in frame and held legible for a beat before the wave hits it.

Three things earn their keep there. Quotes stop the words being spoken aloud. Explicit contrast (white on red) stops the model rendering low-contrast mush. And "held legible for a beat" matters more than people expect — if text is going to be destroyed, obscured or moved, say that it reads clearly first, or the model will start the effect on frame one and you'll never see the words at all.

Keep strings short. Two or three words render reliably; a sentence rarely does.

10. Photoreal creatures need an explicit guard

Ask for an animal and you'll often get something that looks rendered. If you want a real one, say what it isn't — the one other place a negative pays off:

The dinosaur and the shark are REAL ANIMALS, filmed practically — no 3D,
no cartoon, no VFX, no CGI look.

Pair it with a film-stock cue in the style line: 35mm film quality, ARRI ALEXA aesthetic, photorealistic, film grain, motion blur on all fast action.

11. The image-to-video aspect trap

This one catches everybody, and it's specific to 2.5.

On image-to-video, Seedance 2.5 takes its aspect ratio from your start frame. It does not accept 9:16 or 16:9 as an instruction — the frame decides. So if you want a vertical video, generate a vertical seed image. Feeding a 16:9 still and asking for 9:16 in the prompt gets you a 16:9 video.

Text-to-video and reference-to-video take the full aspect list normally. It's only the start-frame path that inherits.

While you're there, the seed frame is worth over-writing. It's the only place identity is described, and everything downstream inherits it — so a vague seed costs you every clip, not just the first. Four things pay for themselves:

  • Say "portrait" and name facial detail for a person — "warm three-quarter portrait, soft catchlights in the eyes, light freckles, relaxed half-smile". Image and video models both reward the explicit portrait cue, and it's what holds the same face across clips.
  • Describe a product's physical FORM, not its brand"a soft squeezable tube with a flip cap", not "the moisturiser". The precise form is what stops the model restyling it into a different object two clips later.
  • Compose for motion to begin — clean eyeline, headroom, and not mid-blink or frozen in an action pose. It's a first frame, not a hero still.
  • Keep added text out of frame — captions, UI, watermarks. A product's own real label and logo are part of the product and should render faithfully; anything else you overlay later is something the model will fight you on.

12. What the model won't do

You might expectReality
A negative_prompt fieldDoesn't exist on Seedance. Fold exclusions into the prompt (see §5, §9)
A multi-shot / beat arrayNot on 2.5. Use timecodes, or Cut to: for real cuts
1080p or 4K720p only. Use Seedance 2.0 for 1080p
Reliable slow-motion rampsUndocumented and inconsistent. Worth trying; have a plan B

On that last one: speed ramps sometimes work beautifully and sometimes get ignored entirely. If a ramp matters, shoot the beat at normal speed and retime it on the timeline — you get frame-accurate control and it's free.

And never put dialogue inside a slow-motion window. A lip-synced line rendered in slow motion drags the voice into a groan. Keep speech in normal-speed beats.

13. Budget your re-rolls

Thirty seconds of 720p is a real generation. Two things follow.

Front-load the risk. Judge render one on the single hardest thing in your prompt — the whispered line, the rendered text, the two-character lip-sync — not on the overall vibe. If the hard part fails, everything else is noise.

Simplify rather than repeat. Three failed re-rolls of an over-stuffed prompt cost more than one clean run of a simpler one. If a 30-second take with four set-pieces keeps collapsing, split it at a natural seam into two 15-second clips and join them on the timeline. You keep the look and buy back control.

The skeleton

Copy this and fill it in.

Total: <N>s / <N> shots / <aspect>

<ONE CONTINUOUS TAKE — no cuts, no fades, no zoom.  |  or omit for cuts>

Style: <genre reference>. <film stock, grade, grain, lens character>.
Photorealistic.

SETTING: <location, time of day, light sources, two or three hero props>

<CHARACTER> — <look, age, LOCKED COLOUR wardrobe>. Voice: <age + gender read,
accent, register, cadence, energy, anti-pattern>.

<state arc that evolves across the take, if any>

AUDIO — <"diegetic only, no score" OR the music brief>

[0-Xs] <Speed.> <Action, in the order it happens.> <ONE camera move + framing.>
<Character>, spoken aloud and lip-synced: "<line, ≤2.5 words per second>"
Sound: <named diegetic effects>

[X-Ys] ...

Worked example

Here's the opening of a 30-second spot we generated with 2.5 — two men in a convenience store, a robbery, and an escalating disaster they refuse to acknowledge. Note the shot header, the locked wardrobe colours, the music brief stated once, and how short the lines are.

Total: 30s / 1 continuous shot / 9:16

ONE CONTINUOUS 30-SECOND TAKE — no cuts, no fades, no zoom, one unbroken camera
move that travels through the whole scene.

Style: a big-budget American action film. 35mm film quality, photorealistic.
Anamorphic lens flares, teal-and-orange grade, low hero angles, film grain,
motion blur on all fast action.

SETTING: a neon-lit American convenience store at night. Magenta and cyan
fluorescent tubes, a wall of glowing drinks coolers, a chrome slushie machine,
glass front doors onto a black parking lot.

DEV — South Asian man, mid-thirties, MUSTARD-YELLOW work polo, holding a tall cup
of electric-blue slush. Voice: warm, easygoing American man, light slightly nasal
mid register, unhurried, completely unbothered — no announcer.

MARCUS — Black man, early forties, open NAVY-BLUE mechanic's jacket over a white
tee. Voice: dry, gravelly American man, low chest register, slow deliberate
cadence, flat and skeptical — no announcer.

Dust, then soot accumulate on both men steadily across the take.

AUDIO — dialogue, sound effects AND music are all generated natively.
MUSIC — two states, ONE change. [0-4s] bland store muzak through a tinny ceiling
speaker, low under the dialogue. [4-30s] 120 BPM electronic rock, surf-guitar
lead, driving drums, builds in layers to full power. No vocals, no orchestra.

[0-4s] Normal speed, calm and mundane. The two men stand chatting by the slushie
machine. Dev takes a pull on his straw and speaks. The camera drifts in slowly.
Dev, spoken aloud and lip-synced: "So did you hear Seedance two-point-five's on
TurboClip now?"
Sound: dialogue, a humming cooler and the slushie machine churning.

[4-6s] DETONATION. The glass front doors blow inward in a wall of shattering
glass. The camera WHIP-PANS hard to the doors as a robber bursts through, a black
silhouette against blinding headlights. Both men drop straight DOWN out of frame
instead of raising their hands.
The robber, screaming, spoken aloud and lip-synced: "HANDS UP!"
Sound: shouted dialogue, a huge glass explosion, a low impact boom, ringing.

The short version

  1. Open with Total: <duration> / <shots> / <aspect>.
  2. Verbs and physical consequences, never adjectives.
  3. One named camera move and framing per moment.
  4. Dialogue in quotes, ~2.5 words per second, voice described once and never reworded.
  5. Brief the audio — or you'll get a stock orchestral score.
  6. Timecode anything over ten seconds.
  7. Lock wardrobe and props with distinct named colours.
  8. On-screen text: quoted, high-contrast, short, and held legible before anything happens to it.
  9. no 3D, no cartoon, no VFX for photoreal animals.
  10. Vertical video needs a vertical seed image — image-to-video inherits the aspect from your start frame.
  11. Timecodes are [0-3s], always counting from 0 within that clip.
  12. Chaining clips? Every clip but the last must end on a steady, in-focus frame — and anything that leaves frame never comes back.

Seedance 2.5 is live in TurboClip now — pick it in the model menu on any video or ad generation. Start a project or see what it costs.