🛰️ Daily AI Frontier
‹ back to 2026-07-31

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

Research LLM Agents

Ranking

Overall 86
Content 95
Popularity 63

Observed public metrics from 1 member.

Representative image for Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

Merged summary

TL;DR - Tycho is a coding-agent system that builds and selectively uses executable world models to solve interactive ARC-AGI-3 games efficiently. Its results suggest that deciding when to construct, repair, use, or bypass a model is as important as simulator accuracy.

  • Tycho models games as parameterized rendered deterministic Moore machines and separates actionable states from animation and terminal frames.
  • Actor-requested delegation to a model builder achieved the best tested orchestration result, averaging 88.49 Relative Human Action Efficiency across 25 public games.
  • With that policy, GPT-5.6 Sol and Opus 5 completed all 183 levels at 100.00 RHAE; Opus 5 used 61% fewer scored actions than aggregate official human baselines.
  • Automatic model repair improved transition reproduction but reached only 83.07 RHAE, showing that dynamics accuracy alone does not ensure effective planning or objective discovery.

Sources (1)

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

arXiv cs.AI Jens Lehmann, Andrei Aioanei, Sahar Vahdati 2026-07-30 arXiv:2607.28287
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-08-23 14:22:30.716462 UTC

TL;DR - Tycho is a coding-agent system that builds and selectively uses executable world models to solve interactive ARC-AGI-3 games efficiently. Its results suggest that deciding when to construct, repair, use, or bypass a model is as important as simulator accuracy.

  • Tycho models games as parameterized rendered deterministic Moore machines and separates actionable states from animation and terminal frames.
  • Actor-requested delegation to a model builder achieved the best tested orchestration result, averaging 88.49 Relative Human Action Efficiency across 25 public games.
  • With that policy, GPT-5.6 Sol and Opus 5 completed all 183 levels at 100.00 RHAE; Opus 5 used 61% fewer scored actions than aggregate official human baselines.
  • Automatic model repair improved transition reproduction but reached only 83.07 RHAE, showing that dynamics accuracy alone does not ensure effective planning or objective discovery.
item →