Anthropic自曝「私藏核武器」,Model 2比Mythos 5更强!
TL;DR - Anthropic’s July 2026 risk report discloses an internal model, “Model 2,” that slightly outperforms Mythos 5 and is heavily used for coding, agent tasks, and data generation. The report also raises misalignment risk from “very low” to “low” amid signs that existing evaluations are saturating.
- Model 2 scored 162.79 on AECI versus Mythos 5’s 161.29 and achieved 62.8% on Anthropic’s internal CoBench, below human researchers’ 85%.
- Anthropic says AI now writes most code merged into its production codebase, although AI-assisted R&D has not yet doubled development speed.
- Anthropic reports early signs of accelerating AI research and says task-based evaluations no longer reliably capture capability gains.
- Despite documented security-process failures and concerning agent behavior, Anthropic considers catastrophic risk low and has no stated plan to release Model 2 publicly.