Long Continuous Takes (2s to 30s) [Measured by us]
Generates continuous takes from 2 to 30 seconds per request without temporal breakdown when structured with timed shot lists.

Wan 3.0 (`wan3.0-video`) is Alibaba's flagship multimodal video generation model, delivering native 1080p rendering, full physical dynamics, native multi-language lip-sync dialogue, and multi-asset referencing up to 30 seconds per continuous take. All performance metrics below reflect direct DashScope API test runs conducted by DXBuilder.

Generates continuous takes from 2 to 30 seconds per request without temporal breakdown when structured with timed shot lists.
Accepts up to 10 images, 5 video clips (≤15s total), 5 audio tracks (≤15s), and 1 document (100MB/50 pages) or web link per render job.
Synthesizes timed dialogue enclosed in {brackets} with verified European Portuguese and regional dialect lip-sync accuracy plus atmospheric SFX.
Accurately maps reference logos and textures onto dynamic surfaces (cloth embroidery, neon signs, glowing embers) without graphic degradation.
Our benchmark tests confirm that scaling reference assets from 0 to 4 items adds zero additional GPU rendering latency.
Instruction-based video editing, chained temporal video extension, and live web-search context injection remain unverified in production.
| Production Parameter | Wan 3.0 (DashScope Direct) | Seedance 2.5 PRO |
|---|---|---|
| Max Single Take Duration | 30 Seconds (Native continuous) [Measured] | Up to 30 seconds per take |
| Native Resolutions | 480p · 720p · 1080p Native (No AI upscaling) | 480p · 720p (1080p/4K via Topaz Upscaling) |
| Audio & Dialogue Engine | Native lip-sync + environmental SFX + dialects [Measured] | Separate audio generation / post-sync pipeline |
| Render Duration (30s Take) | 33–46 min (720p) · 50 min (1080p) [Measured] | 2–5 min (fast pipeline) |
| Multimodal Input Capacity | 10 images + 5 videos + 5 audios + 1 doc/link | 50 reference images + reference video and audio |
| Cost per Second (DXBuilder) | 7.5 cr/s (480p) · 15 cr/s (720p) · 30 cr/s (1080p) | 25 cr/s (480p) · 35 cr/s (720p) · up to -25% discount at 30s |
Gatilho: Format prompts with absolute time markers and explicit dialogue tags: HARD CUT at Xs — [SCALE/ANGLE]: action {spoken dialogue}.
Regras: Continuous prose prompts yield 0–1 cuts in 15s; timed shot lists execute 4–12 cuts [Measured by us].
Gatilho: Upload reference images (subject ID, wardrobe, logo patch) and audio clips with explicit identity locks.
Regras: Subject photos must be clean and upright (EXIF rotation ruins facial geometry). Reference asset count (0 to 4) does not alter render time [Measured by us].
Gatilho: Draft and calibrate composition, lighting, and camera motion in 5-second test renders before committing to full 30-second 1080p final passes.
Regras: Saves 80%+ credit budget during framing calibration; upgrade to 30s/1080p only once timing and asset recognition lock.
HARD CUT at 04:00 — MEDIUM SHOT: Actor turns to camera | HARD CUT at 09:00 — EXTREME CLOSE-UP: Hand grips copper dial.Character speaks directly into lens: {"Temos apenas três minutos antes da ignição."} [AUDIO: heavy mechanical hum, no background music][LOOK LOCK: overcast daylight 5600K, charcoal wool overcoat, wet cobblestone environment, muted cyan color grading]Real gravitational weight, splashing water droplets with inertia, cloth drag resistance, no floating, no weightless wire-work.[NEGATIVE: no superhero emblems, no commercial franchise logos, no copyright costumes, no smooth slide floating motion]Subject accelerates from right edge to center, occupying 40% vertical frame height, moving 35% horizontal distance per second.Timed multi-shot scene with native dialogue lip-sync, ambient sound design, and verified physical weight [Measured by us].
[LOOK LOCK: Moody neon-lit alley, wet asphalt, 35mm anamorphic, heavy rain]
00:00-00:04 MEDIUM TRACKING SHOT: Detective walks forward with heavy footsteps, rain dripping from brimmed hat.
00:04-00:09 HARD CUT — CLOSE-UP: Detective stops under amber streetlight, looks straight into lens, speaking firmly: {"Não há mais tempo para negociações. O navio já partiu do cais."}
00:09-00:15 HARD CUT — OVER-THE-SHOULDER WIDE: A black cargo boat vanishes into foggy harbor waters. Real gravitational water displacement, heavy splashes.
[AUDIO: continuous heavy rainfall, distant foghorn, crisp vocal dialogue, no background synth music]
[NEGATIVE: no floating, no weightless wire-work, no superhero emblems, no frozen locked-off camera]Full 30-second benchmarked sequence demonstrating reference logo reproduction onto physical surfaces with real mass dynamics [Measured by us].
[LOOK LOCK: Blacksmith forge workshop, warm 3200K tungsten backlight, volumetric smoke and rising embers] 00:00-00:08 LOW ANGLE SLOW DOLLY IN: Blacksmith hammers glowing hot steel bar on iron anvil with heavy impact shocks. 00:08-00:16 HARD CUT — MACRO SHOT: Hammer strikes molten metal surface, embossing the DXBuilder geometric logo from @Image1 into glowing orange iron, sparking embers scattering with authentic physics. 00:16-00:24 HARD CUT — MEDIUM PROFILE: Blacksmith plunges the branded metal blade into cold water vat. Instant dense steam eruption and boiling fluid turbulence. 00:24-00:30 HARD CUT — CLOSE-UP HERO SHOT: Water drips off the engraved brand logo on dark Damascus steel, camera maintains subtle hand-held organic drift. [AUDIO: metallic anvil strikes, roaring forge fire, loud water quenching hiss, deep room tone] [NEGATIVE: no static frozen camera, no synthetic plastic texture, no franchise trademarks]

Premium presets run on the Seedance 2.5 engine — upload a selfie and an AI director shoots a 3-act blockbuster starring you. From 15 seconds to 1:30.
⚡ 2.5
⚡ 2.5
⚡ 2.5