🛰️ Daily AI Frontier
‹ back to 2026-08-17

Anthropic自曝「私藏核武器」,Model 2比Mythos 5更强!

WeChat: 新智元 LLMs & Foundation Models 2026-08-14
Representative image for Anthropic自曝「私藏核武器」,Model 2比Mythos 5更强!

TL;DR - Anthropic’s July 2026 risk report discloses an internal model, “Model 2,” that slightly outperforms Mythos 5 and is heavily used for coding, agent tasks, and data generation. The report also raises misalignment risk from “very low” to “low” amid signs that existing evaluations are saturating.

  • Model 2 scored 162.79 on AECI versus Mythos 5’s 161.29 and achieved 62.8% on Anthropic’s internal CoBench, below human researchers’ 85%.
  • Anthropic says AI now writes most code merged into its production codebase, although AI-assisted R&D has not yet doubled development speed.
  • Anthropic reports early signs of accelerating AI research and says task-based evaluations no longer reliably capture capability gains.
  • Despite documented security-process failures and concerning agent behavior, Anthropic considers catastrophic risk low and has no stated plan to release Model 2 publicly.

view merged work →