Fifteen seconds is Seedance 2.0’s maximum single-pass length — the model runs 4 to 15 seconds per generation. So a 15-second reel isn’t a comfortable middle setting. It’s the ceiling, and it’s exactly where the model is most likely to front-load your action and let the last third drift. Every prompt below is built as three labelled beats to stop that.
Copy them, swap the variables, run them vertical.
Settings to lock before you paste anything
| Setting | Use | Why |
|---|---|---|
duration | 15 | Accepts 4 through 15, or auto. Pin it — auto will often hand you 8 seconds. |
aspect_ratio | 9:16 | Reels are built around 9:16 at 1080×1920. Anything else gets letterboxed or cropped. |
resolution | 720p to draft, highest available to ship | 720p vertical renders 720×1280 — under Instagram’s 1080-wide target. Upscale before publishing. |
generate_audio | true | Audio is generated in the same pass at no extra cost, so switching it off saves nothing. |
| Frame rate | 30fps at export | Instagram’s standard. Export H.264, MP4. |
One thing the settings panel won’t tell you: Instagram layers a username, caption and a button stack over your frame. The model has no idea. Keep your subject in the middle 4:5 of the vertical frame and prompt for headroom — “framed with space above the head, subject centred, lower third empty” — or your payoff sits under the like button.
The three-beat structure that fills all 15 seconds
Every prompt here follows the same skeleton, and it’s the reason they hold up at max length:
Shot 1: [camera + framing] [subject + action] [environment + light]
Shot 2: [closer framing] [the payoff action or the quoted line]
Closing: [named held frame] for the final two seconds, silence holds final 0.5s.
[Audio mix line]
- [Suppressor stack]
Four rules make that skeleton work.
Use shot labels, not timestamps. Write Shot 1: and Shot 2:, not 0-5s:. ByteDance documents precise-timing control as unstable, and testing by Ambience AI found second-ranges get read as literal text the model tries to honour, producing degraded output. You’ll see timestamp-bracket prompting recommended elsewhere; on 2.0 specifically, labels are the safer bet.
Budget your words. Around 20 spoken words total across the clip, one sentence per quoted line, roughly 10 words per line. In Ambience’s testing a single 21-word quote transcribed cleanly for about five seconds and then dissolved into phoneme soup — the same brief split into two short quotes transcribed word-for-word across the full runtime.
Name the ending, including the silence. Left undirected, the model invents its own final transition and both picture and audio degrade. Direct it.
Stack your suppressors. There’s no negative prompt field. You close the prompt with a dashed list instead. And “no music” alone isn’t enough — library beds and uninvited narrator voiceovers still get through. The full stack does the job.
Verbs are what the model animates. Spend your prompt on actions and consequences, not adjectives. “Stunning, cinematic, 8K, masterpiece” is four words that point a camera nowhere.
1. The talking-head UGC hook
The highest-volume reel format on paid social, and the one lip sync was built for.
Shot 1: Vertical 9:16 phone-shot UGC. A woman in her late twenties in a grey sweatshirt sits on a sunlit sofa, holding a small amber glass bottle, medium close-up, locked camera, subject centred with headroom. Warm window light, slightly oversaturated phone-camera look, no studio polish. Casual half-amused delivery, realistic lip articulation, no exaggerated mouth opening, no head turns while speaking. She says "I gave this four weeks before writing it off." Shot 2: She tilts the bottle toward the lens in a tighter close-up and says "Week three is when it clicked." Closing: She shrugs and holds a small smile in a static medium close-up for the final two seconds, silence holds final 0.5s. Dialogue clean and prominent, faint room tone subtle.
No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo
Why it works: 16 spoken words, split across two quotes with a physical beat between them. The lip-restraint clause stops the over-acting that pulls mouths out of sync, and the locked camera protects it further — lip sync degrades as the camera moves.
Swap: the product, the objection in line one, the turning point in line two.
2. The three-cut product hero
No dialogue. Sound design carries it. This is the fastest spec ad you can generate.
A vertical 9:16 spec ad for a matte black steel water bottle, built as three cuts in one take. Shot 1: Macro close-up on condensation beading and sliding down the brushed metal on a wet slate slab, hard rim light from the right, cool desaturated grade. Shot 2: Cut to a hand entering frame and twisting the cap free with a sharp click, a thin curl of cold vapour lifting off the mouth of the bottle. Closing: Cut to a locked centred product frame, bottle upright, vapour settling, held for the final two seconds, silence holds final 0.5s. Audio: the drip of condensation, the metallic click of the cap, one low bass note underneath.
No music, no library audio, no voiceover narration, no subtitles, no logo
Why it works: each cut resolves a physical consequence — beading, sliding, the click, the vapour settling. Physics needs something concrete to chase or the model produces a static frame with drift.
Swap: the product and its one satisfying physical moment. Every product has one.
3. The street interview
Two voices, cut between. Reads as documentary, which buys credibility no polished ad gets.
Shot 1: Vertical 9:16 handheld documentary. A man in his thirties in a denim jacket stands on a busy evening pavement, neon shopfronts blurred behind him, medium close-up, slight handheld drift, subject centred with headroom. An off-camera interviewer asks "What's the one thing you'd tell your younger self?" Shot 2: Cut to a tighter single on his face as he thinks, then answers, dry and a little tired, realistic lip articulation, no exaggerated mouth opening: "Stop waiting to feel ready." Closing: He glances off-camera and half-laughs in a held frame for the final two seconds, silence holds final 0.5s. Dialogue clean and prominent, street ambience and distant traffic subtle.
No music, no library audio, no on-screen text, no subtitles, no logo
Why it works: the pause between question and answer is scripted as a beat, so the model animates thinking instead of jumping straight to speech. That gap is most of the realism.
Swap: the question, the location, the answer. Keep the answer under eight words.
4. The voiceover explainer with burned-in subtitles
For educational and listicle reels, where most viewers watch muted.
Shot 1: Vertical 9:16 overhead shot of hands sorting three stacks of paper across a pale oak desk, soft north-facing daylight, clean minimal documentary grade. A low unhurried female voice narrates "Most people sort their inbox by date." Shot 2: Cut to a closer overhead as one hand sweeps two stacks aside, leaving a single stack centred. She continues "Sort it by who's waiting on you." Closing: A held overhead frame on the single remaining stack for the final two seconds, silence holds final 0.5s. Her narration runs as subtitles along the lower third, timed to the voice, clean sans-serif. Narration clean and prominent, paper rustle subtle.
No music, no library audio, no logo
Why it works: subtitles get requested explicitly and timed to the narration, so you skip the captioning pass. Note this is the one prompt where “no subtitles” comes out of the suppressor stack — leave it in and you’ll fight your own instruction.
Swap: the two-line insight. Line one states the default, line two overturns it.
5. The sound-led ASMR loop
Audio is the scene. These over-perform on completion rate because they’re rewatchable.
Shot 1: Vertical 9:16 macro. A knife blade presses into the skin of a ripe blood orange on a dark wooden board, juice welling along the cut, single hard side light, deep saturated grade, shallow focus. Shot 2: The two halves fall apart and rock to a stop, pulp glistening, camera pushing in slowly. Closing: A held macro frame on the cut face, one bead of juice sliding down, for the final two seconds, silence holds final 0.5s. Audio carries the scene: the split of the skin, the wet separation, a soft knock as the halves settle on wood, no music at all.
No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo
Why it works: an open prompt comes back scored like a car advert. Naming every diegetic sound and calling for silence on purpose is the only way to get a genuinely quiet clip.
Swap: any single tactile action with a clean sound signature.
6. The character-consistent series episode
Use this when the reel is episode four of a series and the face has to match. Requires reference support on your platform — Seedance 2.0 accepts up to 9 images, 3 videos and 3 audio clips per generation.
Shot 1: Vertical 9:16. The woman from @Image1 stands in the kitchen from @Image2, wearing the same green apron, medium close-up, locked camera, centred with headroom, warm morning light. Bright quick delivery, realistic lip articulation, no exaggerated mouth opening. She says "Day four and I'm already cheating." Shot 2: She lifts a chipped blue mug into frame in a tighter shot and says "This one doesn't count." Closing: She sips and raises an eyebrow at the lens in a held medium close-up for the final two seconds, silence holds final 0.5s. Dialogue clean and prominent, kitchen ambience subtle.
No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo
Why it works: every reference tag stays attached to a noun — “the woman from @Image1”, never a bare tag. Bare tags are where character drift creeps back in. If your platform doesn’t expose references, reuse the same first frame through an image-to-video workflow instead.
Swap: the running gag, the props, the day number.
7. The transformation reveal
One continuous move, no cut, no dialogue. The camera does the storytelling.
Vertical 9:16, one continuous take, no cuts. A cluttered home office at dusk: papers across the desk, cables tangled, blinds half-shut, cold flat overhead light. The camera pushes in slowly and low toward the desk as the light shifts warmer and the room resolves — the desk now clear, a single lamp on, blinds open to blue evening, one notebook squared to the edge. The transition happens in the light and the movement, not a cut. Closing: The camera settles on a static frame of the lamp and notebook for the final two seconds, silence holds final 0.5s. Audio: a low room hum easing into quiet, one soft click as the lamp comes on, nothing else.
No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo
Why it works: it gives the model a lighting arc to resolve across the runtime instead of an event to finish early. That’s the cleanest fix for last-third drift when there’s no dialogue to carry it.
Swap: any before-and-after where the change can read as a lighting and clutter shift.
What a 15-second render actually costs
Retries are the real budget line, and at max duration they hurt.
| Route | Per second | One 15s render |
|---|---|---|
| fal, 480p (derived from published token formula) | ~$0.134 | ~$2.02 |
| fal fast endpoint, 720p | $0.2419 | $3.63 |
| fal standard, 720p | $0.3034 | $4.55 |
| fal standard, 1080p | $0.682 | $10.23 |
fal’s real meter is token-based — output height × width × duration × 24, divided by 1,024, billed at $0.014 per 1,000 tokens — which is why the per-second rate scales with pixels. Aspect ratio barely moves the number, so a 9:16 clip costs about what a 16:9 one does at the same resolution.
The workflow that follows from this: draft at 5 seconds and 480p to test whether the beat structure and the dialogue land, then re-run the winner at 15 seconds and the highest resolution you have. A 5-second 480p draft runs well under a dollar. Six of those still cost less than one 1080p full-length render.
Credit-based platforms price differently — Novoads charges 3 credits for 5 seconds and 7 for a full 15-second take, roughly $4.83 on their Pro tier. Run the same maximum-length clip through every available door and it lands anywhere between $3.63 and $10.23.
The four ways 15-second clips fail
| Symptom | Cause | Fix |
|---|---|---|
| Speech smears halfway through | One long quoted line | Split into two quotes of ≤10 words with an action beat between |
| Uninvited stock music bed | No suppressor stack | Rank the mix, then append the full dashed list including “no library audio” |
| Speech cut off mid-word | Dialogue running too late | Move all quotes into the first two-thirds, script a silent closing beat |
| Subject stands idle after the action | Only one beat written | Three labelled beats, so every stretch has a job |
Should you be on Seedance 2.5 instead?
ByteDance launched Seedance 2.5 on 31 July 2026, with single-pass clips up to 30 seconds, a far larger reference budget, and timestamp-level editing. API access opened on aggregators through early August.
For reels, mostly no — not yet. Thirty seconds solves a problem you don’t have when your deliverable fits inside 15. Regional availability at launch excluded the US, resolution on shipped integrations is capped at 720p, and Seedance 2.0 has no announced deprecation. Its API is live, its documentation is mature, and the prompt structures above are tested against it.
Revisit when your reels stop fitting in 15 seconds. Until then, the ceiling isn’t the constraint — the beat structure is.
Start here
Take prompt 1. Change nothing except the product and the two lines. Run it at 5 seconds and 480p first, confirm the lip sync and the delivery, then push it to 15 and ship it.
The prompt that beats the one above is the one you’ve already run four cheap drafts of.