Thirty seconds, one take
Longer runtimes generated in a single pass rather than assembled from fragments. No seams where one clip ends and the next begins, and no drift in the subject halfway through.
Longer runtimes generated in a single pass rather than assembled from fragments. No seams where one clip ends and the next begins, and no drift in the subject halfway through.
Bring in up to ten images, five video clips, and five audio clips at once. Images hold the subject, video carries the movement, audio shapes the sound.
Define where the clip opens and where it closes, then let the model build the movement in between. The result lands where you planned it.
Audio is produced as part of the same generation rather than added later, so the clip arrives complete instead of silent.
Start creating with Alibaba’s WAN 3.0 by following a few simple steps directly inside the DaVinci AI Toolkit.
WAN 3.0 by Alibaba suits creators who need longer, sound-complete video with real control over subject and motion.

Fill a whole short-form slot with one generation instead of cutting three together. Keep the same character and the same setting from the first second to the last.

Produce full-length concept spots rather than fragments. Feed in product shots and a reference clip and get something close to a finished cut back.

Explore a scene at its real length. Set the opening and closing frames, hand over reference footage for the camera move, and watch the whole beat play out.

Fill a whole short-form slot with one generation instead of cutting three together. Keep the same character and the same setting from the first second to the last.

Produce full-length concept spots rather than fragments. Feed in product shots and a reference clip and get something close to a finished cut back.

Explore a scene at its real length. Set the opening and closing frames, hand over reference footage for the camera move, and watch the whole beat play out.
Enter a text prompt, upload an image, or add references, and generate a video with sound in seconds.
WAN 3.0 brings together long single-pass generation, multimodal references, and native audio to produce video from text or images.
Generate from a written prompt, animate a single still, or guide the result with images, video clips, and audio referenced directly in the prompt.
Longer clips generated in one go, so movement and identity stay consistent across the full runtime instead of resetting at every cut.
Sound is created alongside the visuals as part of the same pass, covering effects, ambience, and speech.
Set a starting image and an ending image and the model generates the transition between them, giving the clip a defined beginning and end.
Generate from a written prompt, animate a single still, or guide the result with images, video clips, and audio referenced directly in the prompt.
Longer clips generated in one go, so movement and identity stay consistent across the full runtime instead of resetting at every cut.
Sound is created alongside the visuals as part of the same pass, covering effects, ambience, and speech.
Set a starting image and an ending image and the model generates the transition between them, giving the clip a defined beginning and end.
