🛰️ Daily AI Frontier
‹ back to 2026-08-07

阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频

Industry & News Multimodal & Generative 🔗 2 sources

Ranking

Overall 54
Content 55
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频

Merged summary

TL;DR — Alibaba has opened public beta of Wan 3.0, a video generation model that now produces 30-second clips in a single pass and accepts structured documents (doc/xls/ppt/pdf/md) alongside text, image, audio, and video input. It marks a shift from single-shot video creation toward document-to-video productivity tooling for courseware, product demos, and business reports.

  • 30-second single-pass generation enables continuous camera moves and one-take shots; an "intelligent duration" feature auto-recommends clip length from the prompt.
  • First-ever document input support (doc, xls, ppt, pdf, md), capped at 100MB and 50 pages per file or link — e.g., upload a product PPT plus a prompt to get a promo video.
  • Improved human realism: distinct per-person faces, finer facial/skin detail, and restrained emotion with micro-expressions linked to body motion; reference-driven tasks hold character, props, audio, spatial relations, and style consistent.
  • Editing extended to visuals, plot, and dialogue, though Alibaba acknowledges audio quality and text accuracy remain weak points.
  • Availability and pricing: Alibaba Cloud Bailian, 万镜一刻, 万相, Qwen PC creation, IF STUDIO, and 堆友, with grayscale rollout in the Qwen app; API costs ¥0.3/0.6/1.2 per second for 480P/720P/1080P.

Sources differ slightly in emphasis: 量子位 flags the piece as vendor-supplied content with unverified claims and notes the Qwen app grayscale rollout, while 雷峰网 dwells more on real-world fidelity details.

Sources (2)

阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频

量子位 量子位的朋友们 2026-08-07
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-04 14:19:54.363533 UTC

TL;DR - Alibaba opened public beta of Wan 3.0, its video generation model, which now produces single 30-second clips and accepts structured documents (doc/xls/ppt/pdf/md) as input alongside text, image, audio, and video. It signals video generation shifting from single-shot content creation toward a productivity tool for courseware, product demos, and business reports.

  • Single-pass generation extended to 30 seconds, enabling continuous camera moves and one-take shots; an "intelligent duration" feature auto-recommends clip length from the prompt.
  • First support for document inputs (doc, xls, ppt, pdf, md), capped at 100MB and 50 pages per file or link, letting a slide deck plus a prompt drive a full video.
  • Claimed gains in realistic human rendering (facial/skin detail, micro-expressions tied to body motion) and stable retention of character, props, audio, spatial relations, and style in reference-driven tasks; editing now covers visuals, plot, and dialogue.
  • Available on Alibaba Cloud Bailian, Wanjing Yike, Wanxiang, Qwen PC creation, IF STUDIO, and Duiyou, with grayscale rollout in the Qwen app; API pricing is ¥0.3/0.6/1.2 per second for 480P/720P/1080P. Alibaba acknowledges audio quality and text accuracy remain weak points. Note: this piece is vendor-supplied content republished by QbitAI, so claims are unverified.
item →

阿里视频大模型Wan3.0开启公测:文档、ppt也能变视频

雷峰网 (AI科技评论) 2026-08-07
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-04 14:19:52.870238 UTC

TL;DR - Alibaba opened public beta of Wan 3.0, its video generation model, which extends single-shot generation to 30 seconds and adds document formats (doc/xls/ppt/pdf/md) as input modalities. It signals a shift from single-clip video generation toward document-to-video productivity tooling.

  • Generates up to 30s in one pass, enabling continuous camera moves and one-take shots; an "intelligent duration" feature auto-recommends length from the prompt.
  • Beyond text/image/audio/video, it ingests structured documents (≤100MB, ≤50 pages) — e.g., upload a product PPT plus a prompt to get a promo video, targeting courseware, product demos, and business reports.
  • Emphasizes real-world fidelity: distinct per-person faces, finer facial/skin detail, restrained emotion with linked micro-expressions and body motion; reference tasks hold character, props, sound, spatial relations, and style consistent.
  • Improved video editing (visuals, plot, dialogue); Alibaba admits audio quality and text accuracy still lag. Available on Bailian, 万镜一刻, 万相, Qwen PC creation, IF STUDIO, 堆友; API priced at 0.3/0.6/1.2 CNY per second for 480P/720P/1080P.
item →