WAN 3.0 features for your next scene
Choose the input that gives your scene a clear starting point. Match your duration and resolution to the shot you want to make.
Turn a prompt or your own images into 4–30 second videos with sound. Choose up to 1080p and see the credit cost before generating.
WAN 3.0 is a video generation model in Alibaba’s Tongyi Wanxiang family. You can create a moving scene from text, animate a still image with an optional ending frame, or guide a new video with mixed reference material. Output options here include 480p, 720p and 1080p.
Use it for product teasers, social clips, storyboard motion tests and scenes that need more than a few seconds to unfold. Choose the input mode before uploading: first and last frames define endpoints, while reference images, video and audio guide the content. Those two workflows use separate sets of inputs.
Choose the input that gives your scene a clear starting point. Match your duration and resolution to the shot you want to make.
Watch text-to-video, first/last-frame animation and reference-guided creation in these official WAN 3.0 examples.
Create a moving scene from a written prompt, with sound and camera direction.
Guide a continuous transition between a planned opening and ending image.
Use a product image to guide a new presentation with a model.
Match the mode to the control you need. Use fixed frames for a planned transition and mixed references for broader creative guidance.
| Mode | Best starting material | Credits per second | 4-second output cost |
|---|---|---|---|
| Text to video | A written description of the scene Describe the subject, action, camera and sound. No upload is required. | 480p: 4 720p: 8 1080p: 16 | 480p: 16 credits 720p: 32 credits 1080p: 64 credits |
| Image to video | A first frame and optional last frame Upload the first frame before adding an ending frame. Do not combine these with reference assets. | 480p: 4 720p: 8 1080p: 16 | 480p: 16 credits 720p: 32 credits 1080p: 64 credits |
| Multi-reference | Images, video or audio with a prompt Explain each asset’s role. Video and audio each allow 15 seconds total; input video plus output must not exceed 30 seconds. | 480p: 4 720p: 8 1080p: 16 | 480p: 16 credits 720p: 32 credits 1080p: 64 credits |
Text, first-frame and image/audio-reference generations are charged by output duration. With a video reference, add its duration to the output length. The studio displays your total before you generate.
View credit plansStart with one coherent shot. Change one detail at a time so you can compare the results.
Turn product images into short reveals, launch clips and social advertisements.
Describe subjects, action, lighting and sound for a coherent shot lasting 4–30 seconds.
Combine image, video and audio references to guide subjects, motion and ambience.
A glass perfume bottle stands on dark stone. The camera slowly pushes closer while a soft reflection passes over the glass. Keep the bottle upright and its shape stable. Quiet studio ambience.
Use this promptA small sailboat crosses a calm lake at sunrise. A wide camera tracks gently from left to right. Mist drifts near the water, with soft wind and distant birds.
Use this promptUse Image 1 as the subject and Image 2 as the setting. Follow the gentle camera movement in Video 1. The subject turns toward the window while warm afternoon light fills the room.
Use this promptAnswers about WAN 3.0 inputs, duration, resolution and generation credits.
Yes. Upload a first-frame image and describe how the scene should move. You can also add a last frame to guide the ending. For multiple images used as creative references, choose multi-reference instead.
Choose an exact duration from 4 to 30 seconds. If you add reference video, the input video duration and requested output together must not exceed 30 seconds. For example, a 10-second reference leaves room for up to 20 seconds of output.
This generator offers 480p, 720p and 1080p. It does not offer a 4K output setting. Choose 480p for an initial motion test or 1080p when you need a larger output.
Yes. You can include dialogue, ambience or sound-effect directions in your prompt. Generated sound still needs review, particularly spoken wording and timing.
Multi-reference accepts up to 10 images, 5 video clips and 5 audio clips. Each video or audio clip must be 1–15 seconds, and videos and audio each have a 15-second total limit. First and last frames are used in the separate image-to-video mode.
Yes. Reference video duration is added to output duration. A 5-second output with a 3-second reference video is priced as 8 seconds at the selected resolution. Image and audio references do not add billed seconds.
Generation uses site credits. Check your current balance and the displayed generation cost before starting. Visit the pricing page if you need more credits; there is no unlimited free generation.
Credits reserved for a failed generation are returned automatically. If the task succeeds but the result is still being saved, give it time to finish and check My Works before submitting the same request again.