🛰️ Daily AI Frontier
‹ back to 2026-09-04

打穿 AI 智商测试!GPT-6 Astra 的符号世界模型,是突破还是钞能力刷分?

雷峰网 (AI科技评论) LLM Agents 2026-09-04
Representative image for 打穿 AI 智商测试!GPT-6 Astra 的符号世界模型,是突破还是钞能力刷分?

TL;DR - OpenAI’s GPT-6 Astra reportedly approached 100% on ARC-AGI-3 by building symbolic world models, simulating plans, and creating task-specific tools. The result suggests stronger agentic reasoning, but ARC Prize cautions that costly harness infrastructure—not the model alone—helped produce the score and that it is not evidence of AGI.

  • Astra reportedly represents environments with a custom symbolic DSL, tracks rules and state, and tests action sequences in simulated sandboxes before acting.
  • In code-enabled evaluations, it created navigation, combat, and guard-patrol modeling tools tailored to individual games.
  • Its system used persistent memory, programmatic analysis, hypothesis testing, state tracking, and context compression—capabilities traditionally supplied by external agent harnesses.
  • The reported cost was about $360 per game and $18,000 for a full evaluation, raising major efficiency and accessibility concerns.

view merged work →