ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
Ranking
Overall
87
Content
100
Popularity
57
Observed public metrics from 1 member.
Merged summary
TL;DR - ExplorationBench evaluates whether AI systems can discover and apply unfamiliar rules through iterative experimentation rather than recall. Its executable, deliberately counterintuitive “Alien Worlds” make hypotheses exactly verifiable and reduce contamination from pretraining knowledge.
- The benchmark includes AlienCode and AlienLogic, totaling 55 discovery targets and 140 tasks.
- Each sandbox supplies a flawed manual, environment feedback, and a dedicated tool-call schema for exploration.
- Tests across 10 AI systems show that leading systems can learn unfamiliar rules, but results vary substantially between trajectories.
- Additional exploration does not reliably help: progress can stall or even reverse earlier gains.
Sources (1)
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
Public signals
Hugging Face upvotes 8
TL;DR - ExplorationBench evaluates whether AI systems can discover and apply unfamiliar rules through iterative experimentation rather than recall. Its executable, deliberately counterintuitive “Alien Worlds” make hypotheses exactly verifiable and reduce contamination from pretraining knowledge.
- The benchmark includes AlienCode and AlienLogic, totaling 55 discovery targets and 140 tasks.
- Each sandbox supplies a flawed manual, environment feedback, and a dedicated tool-call schema for exploration.
- Tests across 10 AI systems show that leading systems can learn unfamiliar rules, but results vary substantially between trajectories.
- Additional exploration does not reliably help: progress can stall or even reverse earlier gains.