Continuous 30-second videos
Double Seedance 2.0. A complete story in one render, with no stitched segments or continuity breaks.


Unveiled at Volcano Engine's FORCE 2026 conference, Dreamina Seedance 2.5 generates continuous cinematic videos of up to 30 seconds, accepts up to 50 multimodal references and edits or extends existing videos with text alone. This guide condenses the official documentation — creation modes, prompt syntax, technical limits and DXBuilder pricing.
Seedance 2.5 is an AI video generation model focused on long-form storytelling, multimodal referencing and precise editing. In a single call it can generate a complete 30-second story with realistic physics, a cinematic look and synchronized sound — or take one of your videos and transform, extend or repair it. It is the flagship engine of DXBuilder's Seedance 2.5 studio. → Studio
Double Seedance 2.0. A complete story in one render, with no stitched segments or continuity breaks.
Up to 30 images + 10 videos + 10 audio clips in a single request. Keeps characters, products, logos and styles consistent across scenes.
Replace the protagonist, add or remove objects, repair parts of the frame or swap the soundtrack — with text alone.
Continue any video forward or backward, with natural transitions and coherent light and motion.
Precisely lip-synced voice, sound effects and music generated in parallel with the image — no post-production.
Prompts and spoken dialogue in English, Portuguese, Spanish, Japanese, Korean, Arabic, Thai and more — with lip sync.
| Capability | Seedance 2.5 | Seedance 2.0 |
|---|---|---|
| Maximum duration per render | 30 seconds | 15 seconds |
| Multimodal references | 50 (30 img + 10 video + 10 audio) | 15 (9 img + 3 video + 3 audio) |
| Audio-only reference (no image) | ✓ Supported | ✗ Requires image/video |
| Output format | MP4 and MOV (4:4:4 chroma) | MP4 only |
| Total duration of video/audio refs | 30 seconds | 15 seconds |
| Pre-generation planning (3D white-model) | ✓ | ✗ |
| Cinematic intent understanding | Far superior (framing, camera language) | Basic |
The engine classifies every request into one of six task types based on the uploaded references and the words in the prompt. The editing, extension and frame modes have mandatory aspect-ratio and duration constraints — respecting them avoids generation errors.
How to trigger: All you need is a prompt — no files at all.
Rules: No restrictions: freely choose aspect ratio (16:9 … 9:16) and duration (4–30s).
"An FPV drone dives between skyscrapers in the rain, neon lights reflected on the asphalt. <rain and engines> (synthwave)"
How to trigger: 1 image as the starting frame + a prompt describing the motion.
Rules: The aspect ratio automatically matches the image (adaptive). Free duration 4–30s.
"The person in the image comes to life: smiles as their hair moves in the breeze, camera orbiting 360°."
How to trigger: 2 images: the first and last frames of the video. The engine interpolates the transition.
Rules: Aspect ratio inherited from the first frame (adaptive). Free duration 4–30s.
"Cinematic transition from the first to the last frame with light particles and realistic physics."
How to trigger: Combine up to 30 images (@Image1…), 10 videos (@Video1…) and 10 audio clips (@Audio1…).
Rules: Free aspect ratio and duration. ⚠ Avoid words like "edit" or "continue" in the prompt — the engine may reclassify the task and return an error.
"Reference @Image1 for the character. The character walks through a night market, camera movement from @Video1."
How to trigger: Upload a video and use trigger words: edit, add, remove, replace, modify.
Rules: Aspect ratio: adaptive only (keeps the video's). Duration: automatic (≈ same as the original, 0.4s tolerance). The input video must be 4–30s. The uploaded video's duration counts toward billing.
"Video edit: remove everyone in @Video1 except the protagonist."
How to trigger: Upload a video and use trigger words: continue, extend forward/backward, continue the story.
Rules: Aspect ratio: adaptive only (keeps the video's). Duration configurable 4–30s or automatic. The uploaded video's duration counts toward billing.
"Extend @Video1. After the window opens, move into the art gallery shown in @Video2."
FIRST FRAME (1:1)
LAST FRAMEPrompt used: "The girl in the frame says 'cheese' to the camera, with a 360-degree orbiting camera shot". The output video automatically inherits the 1:1 aspect ratio of the first frame.
Subject + action/event + scene and environment + visual style + camera movement / cuts + sound
You may omit unnecessary parts, but keep this order — it's the structure the model was trained to follow.
(epic orchestral music, building)
<distant thunder> <waves crashing>
{Welcome to the future!} — state the language first if not obvious
【AVAILABLE NOW】
@Image1, @Video2, @Audio1 — state what each asset provides (appearance, motion, timbre)

A single 30s render in steampunk style, structured in three windows — [0-10s] a brass clock unfolding into gears, [10-20s] a flight through a zoetrope and copper cable car, [20-30s] a mechanical ship, a giant moon and a return to the clock, with the @Image1 logo appearing in the final second.
"A premium, highly cinematic 30-second 3D motion-graphics sequence in a refined steampunk style... [0-10s]: A macro close-up of an antique brass clock face unfolds into interlocking gear rings... [10-20s]: The camera glides into an ornate brass zoetrope... [20-30s]: ...In the final second, a logo appears, referring to @Image1."

A cookie ad where every reference has a role: @Image1 provides the product, @Video1 the opening composition, @Video2 the camera movement, @Video3 the slow-motion impact, @Video4-6 choreography, typography and the brand ending. This is how you keep brand consistency across an entire video.
"...The strawberry flavor refers to @Image1. The opening builds visual focus on the fruit, referring to the composition of @Video1... one cookie snaps in half and enters slow motion, referring to the impact of @Video3..."

A cinematic rap on a sunset beach where the vocalist sings "hello" in 8 languages — English, Chinese, Japanese, Korean, Portuguese, Thai, Spanish and Arabic — with precise lip sync, hard cuts on the beat and 8 shots described one by one with time windows. It demonstrates Seedance 2.5's native support for 11 languages.
"...Lyrics (precise lip sync): English: 'Hello' / Portuguese: 'Olá' / Japanese: 'こんにちは'... Shot 5 [0:10-0:13] - A musician by the shore; fast lateral dolly... Lyric line 5 (Portuguese 'Olá'). Hard cut."
| Aspect ratio | 480p | 720p |
|---|---|---|
| 16:9 | 854 × 480 | 1280 × 720 |
| 4:3 | 752 × 560 | 1112 × 834 |
| 1:1 | 640 × 640 | 960 × 960 |
| 3:4 | 560 × 752 | 834 × 1112 |
| 9:16 | 480 × 854 | 720 × 1280 |
| 21:9 | 992 × 432 | 1470 × 630 |
From 4 to 30 seconds, or automatic — the engine picks the ideal duration for the prompt. In video editing, the output duration is always ≈ equal to the original video (0.4s maximum difference).
MP4 — maximum compatibility, ideal for publishing on social networks and the web. MOV Pro — H.264 + yuv444p chroma + PCM audio: superior color fidelity for color grading, keying and compositing. Recommended as input and output in editing/extension workflows.
Voice, sound effects and music generated natively and synchronized with the image, in 11 languages: English, Portuguese, Spanish, Chinese, Japanese, Korean, Arabic, Thai, Vietnamese, Indonesian and Malay.
• Without an uploaded video: generated duration × base rate (45 cr/s at 720p · 21 cr/s at 480p)
• With an uploaded video (editing/extension): (upload duration + generated duration) × reduced rate (27 cr/s at 720p · 13 cr/s at 480p)
The duration of the video you upload always adds to the billed time — trim the video before uploading if you only need part of it.
Subject + action/event + scene and environment + visual style + camera movement/cuts + sound. Omit what you don't need — but keep the order.
For long videos, structure the prompt with [0-10s], [10-20s], [20-30s] markers. The engine respects each block's timing and the story flows.
State explicitly what each asset provides: "@Image1 gives the character's appearance; @Video1 gives the camera movement; @Audio1 gives the voice timbre".
"Edit", "remove", "continue" and "extend" change the TYPE of task. If you only want a style reference, avoid those verbs or the request may fail with a parameter error.
If you plan to do color grading, keying or compositing, generate in MOV (4:4:4 chroma + PCM). For publishing directly to social media, MP4 is enough and universal.
Adaptive aspect ratio + automatic duration let the engine pick the ideal format for the content. In editing/extension/frame modes it's actually mandatory.
The engine rejects images with unauthorized real people's faces and trademarked content. Use generic characters or your own licensed assets.
It is ByteDance's latest-generation AI video model, unveiled at Volcano Engine's FORCE 2026 conference. It generates continuous cinematic videos of up to 30 seconds with native synchronized audio, accepts up to 50 multimodal references (images, videos and audio) and performs prompt-based video editing and extension. On DXBuilder you get direct access with no API key required.
2.5 doubles the maximum duration (30s vs 15s), more than triples the references (50 vs 15), adds MOV output with 4:4:4 chroma, accepts audio-only references, and understands the cinematic language of references far better — framing, motion and creative intent.
Between 4 and 30 seconds per render. You can also let the engine automatically pick the ideal duration for your prompt. In extension modes you can continue an existing video multiple times to tell longer stories.
Yes — natively. Use () for music, <> for sound effects, {} for lip-synced spoken dialogue and 【】 for subtitles. It supports voice in 11 languages, including English and Portuguese.
480p and 720p, in 7 aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 and adaptive. Output in MP4 (universal) or professional MOV with H.264 + yuv444p + PCM for color grading and compositing.
You upload a video and write the change: "remove everyone except the protagonist", "replace the character in @Video1 with the one in @Image1", "remove the background music". The engine automatically keeps the original video's aspect ratio and duration.
Billing is per second: 45 credits/s at 720p and 21 credits/s at 480p. If you upload a reference video, its duration adds to the billed time, but the per-second rate drops (27 cr/s at 720p, 13 cr/s at 480p). Use the calculator on this page for the exact amount.
Only with authorization. The engine blocks faces of unauthorized real people and public figures. For human characters use AI-generated images, generic characters or your own licensed assets.
Up to 50: 30 images (@Image1–@Image30), 10 videos (@Video1–@Video10) and 10 audio clips (@Audio1–@Audio10). The total duration of reference videos cannot exceed 30 seconds, and the same applies to audio.
30-second product launch ads, e-commerce videos, cinematic trailers and short films, virtual influencers with consistent characters, multi-scene brand storytelling, 9:16 content for TikTok/Reels/Shorts and multilingual music videos.
No API key, no setup — write the prompt, pick the format and the engine does the rest. If you're out of ideas, the AI invents one for you.
Open the Seedance 2.5 Studio