🛰️ Daily AI Frontier
‹ back to 2026-08-03

阿里Qwen3.8正式发布,编程与办公再进化,推理更快更稳定

量子位 LLMs & Foundation Models 量子位的朋友们 2026-08-03
Representative image for 阿里Qwen3.8正式发布,编程与办公再进化,推理更快更稳定

TL;DR - Alibaba released Qwen3.8, a 2.4T-parameter sparse-MoE flagship (95B active) with 1M-token context and vision understanding, positioned second only to Anthropic's Claude on Arena and aimed squarely at agentic coding and professional "Cowork" tasks. It matters as a price-aggressive frontier alternative, with Qwen3.8-Max and Qwen3.8-27B slated for open-sourcing next week.

  • Architecture/serving: joint sparse-MoE + hybrid attention optimization scales totals to 2.4T with 95B activated; API priced at ¥12/M input and ¥36/M output (¥1.5 on implicit cache hits), stated as 40%/24% of Opus 5's international pricing; Alibaba's Zhenwu M890 supernode claims up to 1.5x agentic inference speedup.
  • Agentic coding: PaperBench 93.0 (+28.2 over prior gen), 4th on CodeArena; a demo run had the model autonomously build a self-evolving agent framework "oh-my-cli" over ~16 days from an empty folder, with the repo published on GitHub.
  • General agent + reasoning: WideSearch 81.9, Agent's Last Exam 52.4, IF Bench 82.8, GPQA Diamond 92.6, OSWorld-Verified 86.1 (claimed top among mainstream models).
  • Multimodal: 2nd on Vision Arena; BabyVision 82.0 without tools; handles 200-page PDFs and 100+ hour video, with a self-built RecreationBench for zero-source-code app replication via iterative visual coding.

Note: figures are vendor-supplied (the article is an authorized Alibaba-provided release), not independently verified.

view merged work →