千问开放平台上线!生态伙伴、开发者可自主接入AI智能体
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR — On Aug 10, Alibaba launched the Qwen (千问) Open Platform, letting ecosystem partners and third-party developers plug their own AI agents directly into the Qwen app; it marks Qwen's shift from a standalone chatbot to an agent marketplace and distribution channel for real-world service fulfillment.
- Multi-terminal access: Service access is opened across three endpoint types — mobile, PC, and AI glasses — rather than a single app surface.
- Agents as independent conversation spaces: Third-party agents run as their own conversation spaces inside the Qwen app, covering the full loop from consultation and recommendation through to order fulfillment.
- User-initiated invocation: Users trigger a service by @-mentioning it or tapping a "dot badge" in the page's upper-right corner.
- First-wave partners span 10+ verticals: logistics (SF Express, 闪送, 快递100), housing/rental (自如), local services (天鹅到家), finance (盈米且慢), mobility (哈啰租车, 嘟嘟巴士, 飞常准), IoT/home (美的美居), and weather (彩云天气).
- Undisclosed: No technical details on APIs, protocols, or revenue-sharing terms were provided.
Note: Only one supplied source (雷峰网/AI科技评论) actually covers this launch; the remaining four summaries describe unrelated work (MLS-Bench, AI-driven science automation, Tencent's AI game teammates, and enterprise AI ontology), so nothing from them was merged in.
Sources (5)
代码榜逐渐饱和,下一个前沿是AI科研,但大模型卡在了最关键一步
TL;DR - A multi-university team (Berkeley, Princeton, Tsinghua, CMU, et al.) introduced MLS-Bench, a benchmark testing whether frontier models can discover genuinely new, generalizable ML methods rather than just tune existing ones; none of the 5 evaluated models beat the strongest human methods.
- Design: 140 real research tasks across 12 ML domains (LM pre/post-training, vision generation, RL, robotics, systems, AI4Science, causal inference, time series, etc.). Each task pins a specific method component, requires ≥3 test conditions to check transfer, and re-implements ≥3 strong human methods (including SOTA) in the same code/training/scoring pipeline. Scoring anchors weakest reproduced baseline to 0 and strongest to 50.
- Attribution control: Editable vs. protected code is annotated line-by-line, and parameter counts are checked — with that check removed, models repeatedly won by simply scaling capacity, showing apparent "auto-discovery" is often just added compute.
- Findings: Models improve when prompted to optimize/debug but get worse when asked to invent new methods; expert review found submissions mostly recombine losses/modules/tricks from provided baselines. More sampling, evolutionary search, and test-time training saturate quickly and overfit to visible conditions.
- Deeper bottleneck: Given a flexible compute budget instead of three fixed 345M pretraining runs, models mostly performed worse — they failed to allocate compute to high-information experiments or revise plans from intermediate evidence. Web search and supplied papers/derivations helped little. Full run costs ~700 H100-hours; a 30-task subset (~100 H100-hours) is already used in some official frontier-model releases.
AI让科研变成“科学自动机”,副作用是什么?
TL;DR - HKU professor Zhang Zheng argues that the emerging "science automaton" — closed-loop AI hypothesis generation plus automated verification — will boost scientific throughput but carries two structural side effects: erosion of human scientific taste and a hard ceiling on genuine novelty.
- Closing the loop: generation is already industrialized (Shanghai Analemma's FARS produced 166 papers in 17 days, ~2h and ~$1k each, scoring 5.05 under Stanford's Agentic Reviewer vs. ICLR 2026 human submission mean 4.21); the hard part is verification — OpenAI's Lean 4 certificates for 10 open math problems, Google's Science One Framework evidence chains, and Shanghai AI Lab's Intern·Duanyan dry/wet closed loop with independent review agents all attack the same problem.
- Capability trap: taste and judgment only grow from years of hands-on struggle, especially writing ("writing is half of research" — cleaner designs surface only while drafting). Since institutions reward output, not capability formation, this step is being outsourced first and likely irreversibly; today's senior judgment is a non-renewable stock.
- Ceiling trap: LLMs are super pattern-completers bounded by generalization from existing corpora (author cites his own Comprehension Without Competence, TMLR 2025). Automata excel at combinatorial search in known spaces — materials screening, drug repurposing, weather (1000x faster than numerical methods), chip design (5–10x) — but "fill the floor" rather than raise the ceiling.
- Breakthroughs like negative numbers, complex numbers, non-Euclidean geometry, or AlphaFold's reframing of folding from physics computation to a prediction problem lie outside any existing search space; such outlier "spiky gradients" are exactly what training smooths away.
AI游戏搭子:和平精英领跑行业的新战场
TL;DR - Tencent's Game for Peace (和平精英) is scaling LLM-driven "AI teammates" into a mass-market battle royale, reporting 110M cumulative users and 17.7M weekend DAU, and showcased it at WAIC 2026 as a consumer-facing entry point to frontier AI.
- Moves beyond scripted NPCs/voice assistants: agents participate in combat (finishing kills, calling engagement lines, tracking enemies, locating vehicles/loot) and respond to natural-language voice commands like "hold position" or "search for enemies."
- 花傲天 is claimed as the first AI teammate with cross-match persistent memory; AI 小田 (launched May 2026, powered by Hunyuan 3.0) adds long-horizon preference modeling, personalized expression, and skill progression over time.
- Reported social effect: ~75% of players voluntarily use voice chat when squadding with AI teammates, framed as lowering the barrier to first interaction rather than replacing human social play.
- AI extends across the player lifecycle — pre-match advice via digital spokesperson 吉莉/"吉事通", post-match generated photos, and the 绿洲 (Oasis) AI editor / 绿洲启元 for natural-language UGC map and gameplay-logic creation.
AI 读懂了所有规则,公司却要重新理解自己
TL;DR - An essay arguing that when an AI agent correctly reads all of a company's conflicting rules (refund policy, accounting state, sales promise, risk-control threshold), it exposes organizational contradictions rather than knowledge gaps — forcing firms to re-specify their own authority, priorities, and accountability structures.
- AI shifts error structure, not error existence: it can reduce lookup/omission errors, surfaces pre-existing latent conflicts, adds new failure modes (fabrication, misread context, overconfidence), and amplifies bad definitions/metrics at software speed; shared models also correlate previously independent judgments.
- The right baseline is the post-AI human-machine system vs. today's real people/processes — measured on error severity, detectability, reversibility, correlated blast radius, who bears the cost, and business outcome, not just average accuracy.
- Data alone is insufficient: an "enterprise ontology" must encode object states, permitted actions, evidence thresholds, who executes/approves/grants exceptions, reversibility, and completion criteria, plus which value gets protected when commitments conflict.
- History is evidence, not authorization — logs record only the chosen branch (no counterfactuals); the proposed remedy is "decision memory" (evidence, rejected options, uncertainty, actual outcomes), with rollout staged: advisory-only first (tracking why humans reject suggestions), then low-risk reversible actions automated.
- Cites Brynjolfsson/Li/Raymond (5,179 agents, ~14% resolutions/hour, larger gains for novices), Dell'Acqua's "jagged frontier" (758 workers, worse accuracy past the capability boundary), Dietvorst's algorithm aversion, and a 776-professional experiment where AI-assisted individuals approached two-person team performance.
千问开放平台上线!生态伙伴、开发者可自主接入AI智能体
TL;DR - Alibaba's Qwen (千问) launched an open platform on Aug 10 that lets ecosystem partners and third-party developers plug their own AI agents into the Qwen app across phone, PC, and AI-glasses endpoints. It signals a shift from a standalone chatbot to an agent marketplace/distribution channel for real-world service fulfillment.
- Service access is opened across three terminal types — mobile, PC, and AI glasses — rather than a single app surface.
- Third parties build agents that run as independent conversation spaces inside the Qwen app, covering the full loop from consultation and recommendation through to order fulfillment.
- Invocation is user-initiated via @-mentioning a service or tapping a "dot badge" in the page's upper-right corner.
- First wave of partners spans 10+ verticals: logistics (SF Express, Shansong, 快递100), housing/rental (自如), local services (天鹅到家), finance (盈米且慢), mobility (哈啰租车, 嘟嘟巴士, 飞常准), plus IoT/home (美的美居) and weather (彩云天气); no technical details on APIs, protocols, or revenue terms are given.