🛰️ Daily AI Frontier
‹ back to 2026-07-27

半价干翻Fable 5?Opus 5实测炸场,网友:差点从椅子上摔下来

Industry & News LLMs & Foundation Models

Ranking

Overall 71
Content 80
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - Anthropic released Claude Opus 5, reporting major gains in coding, tool use, novel-problem solving, and computer operation at half the price of Fable 5. Early demonstrations also highlight stronger self-verification and enough model judgment for Claude Code to cut its system prompt by over 80% without benchmark regression.

  • Opus 5 costs $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8.
  • It scored 43.3% on Frontier-Bench, 30.2% on ARC-AGI-3, and 64.7% on tool-enabled Humanity’s Last Exam.
  • User tests showed strong generation of interactive games, simulations, CAD models, and single-file web experiences, often with accurate physics.
  • The model increasingly builds its own validation pipelines, measuring outputs against source data rather than relying only on visual inspection.

Sources (1)

半价干翻Fable 5?Opus 5实测炸场,网友:差点从椅子上摔下来

量子位 克雷西 2026-07-25
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-26 14:45:55.456752 UTC

TL;DR - Anthropic released Claude Opus 5, reporting major gains in coding, tool use, novel-problem solving, and computer operation at half the price of Fable 5. Early demonstrations also highlight stronger self-verification and enough model judgment for Claude Code to cut its system prompt by over 80% without benchmark regression.

  • Opus 5 costs $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8.
  • It scored 43.3% on Frontier-Bench, 30.2% on ARC-AGI-3, and 64.7% on tool-enabled Humanity’s Last Exam.
  • User tests showed strong generation of interactive games, simulations, CAD models, and single-file web experiences, often with accurate physics.
  • The model increasingly builds its own validation pipelines, measuring outputs against source data rather than relying only on visual inspection.
item →