🛰️ Daily AI Frontier
‹ back to 2026-07-27

人类研发时代结束?XYZ最强Search Agent横扫七大榜单,300个Agent跑通AI4AI全栈闭环

Industry & News LLM Agents

Ranking

Overall 71
Content 80
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 人类研发时代结束?XYZ最强Search Agent横扫七大榜单,300个Agent跑通AI4AI全栈闭环

Merged summary

TL;DR - XYZ AI Lab released two Deep Search agents—35B-parameter Aquila-mini and 397B-parameter Aquila-pro—developed through an AI4AI workflow that coordinated over 300 agents. The release demonstrates bounded, human-governed AI participation across data generation, training, runtime optimization, and evaluation.

  • Aquila-mini reportedly led same-scale systems across seven search benchmarks; Aquila-pro posted top public scores on BrowseComp-ZH, LiveBrowseComp, and WideSearch.
  • Human-defined goals, budgets, permissions, and acceptance rules constrain agents, while independent evaluators and gates approve changes and escalate risks.
  • The system retains append-only audit trails and shared memory so successful and failed experiments remain reproducible and reusable.
  • Agent-discovered improvements included proactive context organization to preserve evidence during long search trajectories and state-faithful supervised fine-tuning based on information actually visible at each step.

Sources (1)

人类研发时代结束?XYZ最强Search Agent横扫七大榜单,300个Agent跑通AI4AI全栈闭环

WeChat: 机器之心 2026-07-25
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-26 14:45:36.318238 UTC

TL;DR - XYZ AI Lab released two Deep Search agents—35B-parameter Aquila-mini and 397B-parameter Aquila-pro—developed through an AI4AI workflow that coordinated over 300 agents. The release demonstrates bounded, human-governed AI participation across data generation, training, runtime optimization, and evaluation.

  • Aquila-mini reportedly led same-scale systems across seven search benchmarks; Aquila-pro posted top public scores on BrowseComp-ZH, LiveBrowseComp, and WideSearch.
  • Human-defined goals, budgets, permissions, and acceptance rules constrain agents, while independent evaluators and gates approve changes and escalate risks.
  • The system retains append-only audit trails and shared memory so successful and failed experiments remain reproducible and reusable.
  • Agent-discovered improvements included proactive context organization to preserve evidence during long search trajectories and state-faithful supervised fine-tuning based on information actually visible at each step.
item →