阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频
TL;DR - Alibaba opened public beta of Wan 3.0, its video generation model, which now produces single 30-second clips and accepts structured documents (doc/xls/ppt/pdf/md) as input alongside text, image, audio, and video. It signals video generation shifting from single-shot content creation toward a productivity tool for courseware, product demos, and business reports.
- Single-pass generation extended to 30 seconds, enabling continuous camera moves and one-take shots; an "intelligent duration" feature auto-recommends clip length from the prompt.
- First support for document inputs (doc, xls, ppt, pdf, md), capped at 100MB and 50 pages per file or link, letting a slide deck plus a prompt drive a full video.
- Claimed gains in realistic human rendering (facial/skin detail, micro-expressions tied to body motion) and stable retention of character, props, audio, spatial relations, and style in reference-driven tasks; editing now covers visuals, plot, and dialogue.
- Available on Alibaba Cloud Bailian, Wanjing Yike, Wanxiang, Qwen PC creation, IF STUDIO, and Duiyou, with grayscale rollout in the Qwen app; API pricing is ¥0.3/0.6/1.2 per second for 480P/720P/1080P. Alibaba acknowledges audio quality and text accuracy remain weak points. Note: this piece is vendor-supplied content republished by QbitAI, so claims are unverified.