DX Builder
DX Builder
OFFICIAL DOCUMENTATION · MODES · PROMPTS · PRICINGByteDance Seedance 2.5 — Complete Guide
← Back to the Seedance 2.5 Studio

Seedance 2.5: the complete guide to ByteDance's most advanced video model

Unveiled at Volcano Engine's FORCE 2026 conference, Dreamina Seedance 2.5 generates continuous cinematic videos of up to 30 seconds, accepts up to 50 multimodal references and edits or extends existing videos with text alone. This guide condenses the official documentation — creation modes, prompt syntax, technical limits and DXBuilder pricing.

30s per render480p · 720p7 aspect ratiosMP4 + MOV 4:4:450 references11 languagesNative audio

What is ByteDance Seedance 2.5?

Seedance 2.5 is an AI video generation model focused on long-form storytelling, multimodal referencing and precise editing. In a single call it can generate a complete 30-second story with realistic physics, a cinematic look and synchronized sound — or take one of your videos and transform, extend or repair it. It is the flagship engine of DXBuilder's Seedance 2.5 studio. → Studio

Continuous 30-second videos

Double Seedance 2.0. A complete story in one render, with no stitched segments or continuity breaks.

50 multimodal references

Up to 30 images + 10 videos + 10 audio clips in a single request. Keeps characters, products, logos and styles consistent across scenes.

Prompt-based video editing

Replace the protagonist, add or remove objects, repair parts of the frame or swap the soundtrack — with text alone.

Video extension

Continue any video forward or backward, with natural transitions and coherent light and motion.

Native synchronized audio

Precisely lip-synced voice, sound effects and music generated in parallel with the image — no post-production.

11 native languages

Prompts and spoken dialogue in English, Portuguese, Spanish, Japanese, Korean, Arabic, Thai and more — with lip sync.

Seedance 2.5 vs Seedance 2.0: what changed

CapabilitySeedance 2.5Seedance 2.0
Maximum duration per render30 seconds15 seconds
Multimodal references50 (30 img + 10 video + 10 audio)15 (9 img + 3 video + 3 audio)
Audio-only reference (no image)✓ Supported✗ Requires image/video
Output formatMP4 and MOV (4:4:4 chroma)MP4 only
Total duration of video/audio refs30 seconds15 seconds
Pre-generation planning (3D white-model)
Cinematic intent understandingFar superior (framing, camera language)Basic

The 6 creation modes (and the rules for each)

The engine classifies every request into one of six task types based on the uploaded references and the words in the prompt. The editing, extension and frame modes have mandatory aspect-ratio and duration constraints — respecting them avoids generation errors.

1 · Text to Video

How to trigger: All you need is a prompt — no files at all.

Rules: No restrictions: freely choose aspect ratio (16:9 … 9:16) and duration (4–30s).

"An FPV drone dives between skyscrapers in the rain, neon lights reflected on the asphalt. <rain and engines> (synthwave)"

2 · Image to Video (First Frame)

How to trigger: 1 image as the starting frame + a prompt describing the motion.

Rules: The aspect ratio automatically matches the image (adaptive). Free duration 4–30s.

"The person in the image comes to life: smiles as their hair moves in the breeze, camera orbiting 360°."

3 · Frame A → Frame B (First/Last)

How to trigger: 2 images: the first and last frames of the video. The engine interpolates the transition.

Rules: Aspect ratio inherited from the first frame (adaptive). Free duration 4–30s.

"Cinematic transition from the first to the last frame with light particles and realistic physics."

4 · Multimodal (References)

How to trigger: Combine up to 30 images (@Image1…), 10 videos (@Video1…) and 10 audio clips (@Audio1…).

Rules: Free aspect ratio and duration. ⚠ Avoid words like "edit" or "continue" in the prompt — the engine may reclassify the task and return an error.

"Reference @Image1 for the character. The character walks through a night market, camera movement from @Video1."

5 · Video Editing

How to trigger: Upload a video and use trigger words: edit, add, remove, replace, modify.

Rules: Aspect ratio: adaptive only (keeps the video's). Duration: automatic (≈ same as the original, 0.4s tolerance). The input video must be 4–30s. The uploaded video's duration counts toward billing.

"Video edit: remove everyone in @Video1 except the protagonist."

6 · Video Extension

How to trigger: Upload a video and use trigger words: continue, extend forward/backward, continue the story.

Rules: Aspect ratio: adaptive only (keeps the video's). Duration configurable 4–30s or automatic. The uploaded video's duration counts toward billing.

"Extend @Video1. After the window opens, move into the art gallery shown in @Video2."

Real example (official documentation): First frame → Last frame
FIRST FRAME (1:1)FIRST FRAME (1:1)
LAST FRAMELAST FRAME

Prompt used: "The girl in the frame says 'cheese' to the camera, with a 360-degree orbiting camera shot". The output video automatically inherits the 1:1 aspect ratio of the first frame.

Prompt syntax: the official formula

Subject + action/event + scene and environment + visual style + camera movement / cuts + sound

You may omit unnecessary parts, but keep this order — it's the structure the model was trained to follow.

( )Music

(epic orchestral music, building)

< >Sound effects (SFX)

<distant thunder> <waves crashing>

{ }Spoken dialogue

{Welcome to the future!} — state the language first if not obvious

【 】On-screen subtitles

【AVAILABLE NOW】

@References

@Image1, @Video2, @Audio1 — state what each asset provides (appearance, motion, timbre)

Real examples from the official documentation

Seedance 2.5 — 30s steampunk reference input

Continuous 30-second video with time blocks

A single 30s render in steampunk style, structured in three windows — [0-10s] a brass clock unfolding into gears, [10-20s] a flight through a zoetrope and copper cable car, [20-30s] a mechanical ship, a giant moon and a return to the clock, with the @Image1 logo appearing in the final second.

"A premium, highly cinematic 30-second 3D motion-graphics sequence in a refined
steampunk style... [0-10s]: A macro close-up of an antique brass clock face unfolds
into interlocking gear rings... [10-20s]: The camera glides into an ornate brass
zoetrope... [20-30s]: ...In the final second, a logo appears, referring to @Image1."
Seedance 2.5 — product ad reference image

Product ad with 1 image + 6 reference videos

A cookie ad where every reference has a role: @Image1 provides the product, @Video1 the opening composition, @Video2 the camera movement, @Video3 the slow-motion impact, @Video4-6 choreography, typography and the brand ending. This is how you keep brand consistency across an entire video.

"...The strawberry flavor refers to @Image1. The opening builds visual focus on the
fruit, referring to the composition of @Video1... one cookie snaps in half and enters
slow motion, referring to the impact of @Video3..."
Seedance 2.5 — multilingual music video reference

Music video with vocals in 8 languages and lip sync

A cinematic rap on a sunset beach where the vocalist sings "hello" in 8 languages — English, Chinese, Japanese, Korean, Portuguese, Thai, Spanish and Arabic — with precise lip sync, hard cuts on the beat and 8 shots described one by one with time windows. It demonstrates Seedance 2.5's native support for 11 languages.

"...Lyrics (precise lip sync): English: 'Hello' / Portuguese: 'Olá' / Japanese:
'こんにちは'... Shot 5 [0:10-0:13] - A musician by the shore; fast lateral dolly...
Lyric line 5 (Portuguese 'Olá'). Hard cut."

Technical output specifications

Aspect ratio480p720p
16:9854 × 4801280 × 720
4:3752 × 5601112 × 834
1:1640 × 640960 × 960
3:4560 × 752834 × 1112
9:16480 × 854720 × 1280
21:9992 × 4321470 × 630

Duration

From 4 to 30 seconds, or automatic — the engine picks the ideal duration for the prompt. In video editing, the output duration is always ≈ equal to the original video (0.4s maximum difference).

MP4 vs MOV

MP4 — maximum compatibility, ideal for publishing on social networks and the web. MOV Pro — H.264 + yuv444p chroma + PCM audio: superior color fidelity for color grading, keying and compositing. Recommended as input and output in editing/extension workflows.

Audio

Voice, sound effects and music generated natively and synchronized with the image, in 11 languages: English, Portuguese, Spanish, Chinese, Japanese, Korean, Arabic, Thai, Vietnamese, Indonesian and Malay.

Reference limits (what you can upload)

Reference images

  • Formats: JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC/HEIF
  • Dimensions: 300–6000 px per side; aspect ratio between 0.4 and 2.5
  • Size: up to 30 MB per image
  • Quantity: 1 (first frame) · 2 (first/last) · up to 30 (multimodal)

Reference videos

  • Formats: MP4 and MOV (H.264/H.265 + AAC/MP3)
  • Duration: 2–30s each; total of all videos ≤ 30s
  • Resolution 480p–720p · frame rate 24–60 fps · up to 200 MB
  • Quantity: up to 10 videos per request

Reference audio

  • Formats: WAV and MP3
  • Duration: 2–30s each; total of all audio ≤ 30s
  • Size: up to 15 MB per file
  • Quantity: up to 10 clips · works WITHOUT image/video (new in 2.5)

DXBuilder pricing & credit calculator

10s
0s
Estimated total cost450 credits
Rate: 45 cr/s · Billed time: 10s (10s generated)
How billing works:

Without an uploaded video: generated duration × base rate (45 cr/s at 720p · 21 cr/s at 480p)

With an uploaded video (editing/extension): (upload duration + generated duration) × reduced rate (27 cr/s at 720p · 13 cr/s at 480p)

The duration of the video you upload always adds to the billed time — trim the video before uploading if you only need part of it.

7 best practices for perfect renders

1

Follow the prompt formula

Subject + action/event + scene and environment + visual style + camera movement/cuts + sound. Omit what you don't need — but keep the order.

2

Split into time windows

For long videos, structure the prompt with [0-10s], [10-20s], [20-30s] markers. The engine respects each block's timing and the story flows.

3

Assign responsibilities to references

State explicitly what each asset provides: "@Image1 gives the character's appearance; @Video1 gives the camera movement; @Audio1 gives the voice timbre".

4

Beware of trigger words

"Edit", "remove", "continue" and "extend" change the TYPE of task. If you only want a style reference, avoid those verbs or the request may fail with a parameter error.

5

MOV for post-production

If you plan to do color grading, keying or compositing, generate in MOV (4:4:4 chroma + PCM). For publishing directly to social media, MP4 is enough and universal.

6

Adaptive is your friend

Adaptive aspect ratio + automatic duration let the engine pick the ideal format for the content. In editing/extension/frame modes it's actually mandatory.

7

No famous faces or brands

The engine rejects images with unauthorized real people's faces and trademarked content. Use generic characters or your own licensed assets.

Frequently asked questions about Seedance 2.5

What is ByteDance Seedance 2.5?

It is ByteDance's latest-generation AI video model, unveiled at Volcano Engine's FORCE 2026 conference. It generates continuous cinematic videos of up to 30 seconds with native synchronized audio, accepts up to 50 multimodal references (images, videos and audio) and performs prompt-based video editing and extension. On DXBuilder you get direct access with no API key required.

What is the difference between Seedance 2.5 and Seedance 2.0?

2.5 doubles the maximum duration (30s vs 15s), more than triples the references (50 vs 15), adds MOV output with 4:4:4 chroma, accepts audio-only references, and understands the cinematic language of references far better — framing, motion and creative intent.

How long can a Seedance 2.5 video be?

Between 4 and 30 seconds per render. You can also let the engine automatically pick the ideal duration for your prompt. In extension modes you can continue an existing video multiple times to tell longer stories.

Does Seedance 2.5 generate audio and voice?

Yes — natively. Use () for music, <> for sound effects, {} for lip-synced spoken dialogue and 【】 for subtitles. It supports voice in 11 languages, including English and Portuguese.

Which resolutions and formats does it support?

480p and 720p, in 7 aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 and adaptive. Output in MP4 (universal) or professional MOV with H.264 + yuv444p + PCM for color grading and compositing.

How does prompt-based video editing work?

You upload a video and write the change: "remove everyone except the protagonist", "replace the character in @Video1 with the one in @Image1", "remove the background music". The engine automatically keeps the original video's aspect ratio and duration.

How much does a Seedance 2.5 video cost on DXBuilder?

Billing is per second: 45 credits/s at 720p and 21 credits/s at 480p. If you upload a reference video, its duration adds to the billed time, but the per-second rate drops (27 cr/s at 720p, 13 cr/s at 480p). Use the calculator on this page for the exact amount.

Can I use photos of real people?

Only with authorization. The engine blocks faces of unauthorized real people and public figures. For human characters use AI-generated images, generic characters or your own licensed assets.

How many references can I combine in one video?

Up to 50: 30 images (@Image1–@Image30), 10 videos (@Video1–@Video10) and 10 audio clips (@Audio1–@Audio10). The total duration of reference videos cannot exceed 30 seconds, and the same applies to audio.

What can I create with Seedance 2.5?

30-second product launch ads, e-commerce videos, cinematic trailers and short films, virtual influencers with consistent characters, multi-scene brand storytelling, 9:16 content for TikTok/Reels/Shorts and multilingual music videos.

Ready to shoot with Seedance 2.5?

No API key, no setup — write the prompt, pick the format and the engine does the rest. If you're out of ideas, the AI invents one for you.

Open the Seedance 2.5 Studio