🛰️ Daily AI Frontier
‹ back to 2026-09-25

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

arXiv cs.AI LLM Agents Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan 2026-09-24

TL;DR - ExplorationBench evaluates whether AI systems can discover and apply unfamiliar rules through iterative experimentation rather than recall. Its executable, deliberately counterintuitive “Alien Worlds” make hypotheses exactly verifiable and reduce contamination from pretraining knowledge.

  • The benchmark includes AlienCode and AlienLogic, totaling 55 discovery targets and 140 tasks.
  • Each sandbox supplies a flawed manual, environment feedback, and a dedicated tool-call schema for exploration.
  • Tests across 10 AI systems show that leading systems can learn unfamiliar rules, but results vary substantially between trajectories.
  • Additional exploration does not reliably help: progress can stall or even reverse earlier gains.

view merged work →