华尔街实测8款全球主流Agent:千问办公综合排名第一
TL;DR - Jefferies tested eight mainstream AI agents on five real-world office tasks and ranked Alibaba’s Qianwen Office first overall. The results highlight agent harness engineering and per-task cost—not just underlying model quality—as key enterprise differentiators.
- Qianwen Office was reportedly the only product scoring above 90 across every evaluation dimension, including complex office work, browser control, and multimodal generation.
- Tasks covered annual-report summarization, web-based company comparisons, desktop browser operation, English presentation creation, and marketing-poster generation.
- Jefferies estimated that Qianwen had the strongest “implicit harness,” encompassing instructions, context management, tool use, safeguards, feedback, and error correction.
- Its Qwen 3.8 Max foundation model was reported to have lower API pricing than some leading overseas models, potentially improving cost per completed agent task.