打穿 AI 智商测试!GPT-6 Astra 的符号世界模型,是突破还是钞能力刷分?
TL;DR - OpenAI’s GPT-6 Astra reportedly approached 100% on ARC-AGI-3 by building symbolic world models, simulating plans, and creating task-specific tools. The result suggests stronger agentic reasoning, but ARC Prize cautions that costly harness infrastructure—not the model alone—helped produce the score and that it is not evidence of AGI.
- Astra reportedly represents environments with a custom symbolic DSL, tracks rules and state, and tests action sequences in simulated sandboxes before acting.
- In code-enabled evaluations, it created navigation, combat, and guard-patrol modeling tools tailored to individual games.
- Its system used persistent memory, programmatic analysis, hypothesis testing, state tracking, and context compression—capabilities traditionally supplied by external agent harnesses.
- The reported cost was about $360 per game and $18,000 for a full evaluation, raising major efficiency and accessibility concerns.