🛰️ Daily AI Frontier
‹ back to 2026-08-07

阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频

雷峰网 (AI科技评论) Multimodal & Generative 2026-08-07
Representative image for 阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频

TL;DR - Alibaba opened public beta of Wan 3.0, its video generation model, which extends single-shot generation to 30 seconds and adds document formats (doc/xls/ppt/pdf/md) as input modalities. It signals a shift from single-clip video generation toward document-to-video productivity tooling.

  • Generates up to 30s in one pass, enabling continuous camera moves and one-take shots; an "intelligent duration" feature auto-recommends length from the prompt.
  • Beyond text/image/audio/video, it ingests structured documents (≤100MB, ≤50 pages) — e.g., upload a product PPT plus a prompt to get a promo video, targeting courseware, product demos, and business reports.
  • Emphasizes real-world fidelity: distinct per-person faces, finer facial/skin detail, restrained emotion with linked micro-expressions and body motion; reference tasks hold character, props, sound, spatial relations, and style consistent.
  • Improved video editing (visuals, plot, dialogue); Alibaba admits audio quality and text accuracy still lag. Available on Bailian, 万镜一刻, 万相, Qwen PC creation, IF STUDIO, 堆友; API priced at 0.3/0.6/1.2 CNY per second for 480P/720P/1080P.

view merged work →