Alibaba Launches Wan3.0: Its 30-Second Video Model
Alibaba has launched Wan3.0, the latest version of its video generation model, days after raising roughly $10.2bn in a Hong Kong share placement, with all proceeds allocated to AI. This strategic move comes as the company emphasizes tangible outputs while Western attention primarily focuses on chatbots.
Key Features:
- Multimodal Input: Supports text, images, audio, video, web pages (PDFs, slides), and documents, allowing users to convert static data into dynamic videos up to 30 seconds long.
- High Quality and Resolution: Generates clips at resolutions up to 1080p, with detailed character renderings, props, spatial layout, and motion graphics consistent throughout.
- Synchronized Micro-Expressions: Captures and synchronizes facial expressions and provides multilingual voice output.
- Competitive Pricing: Offers $0.05/second at 480p, $0.10 at 720p, and $0.20 at 1080p, undercutting Google’s Veo 3.1 pricing.
Unique Selling Points:
- Versatile Applications: Targeted towards marketing departments, corporate communications teams, and various industries like film, robotics, and autonomous vehicles.
- Access and Availability: Currently accessible through Model Studio and Qwen Cloud on an application basis, with full API rollout planned. A consumer-facing site at wan.video is in the works as a members-only platform.
Notable Omissions:
- Open Weights: Unlike previous versions, Wan3.0’s weights are closed, marking a shift from open-sourcing most of Alibaba’s video models in the past. This decision also coincides with charging heavier users of its open Qwen models.
Remember that: While promising, independent benchmarks for Wan3.0 are currently unavailable, and its quality claims remain based on company-issued information.