🛰️ Daily AI Frontier
‹ back to 2026-07-27

SceneActBench: Can Agents Act on the 3D Scenes They See?

arXiv cs.AI LLM Agents Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu 2026-07-24

TL;DR - SceneActBench evaluates VLM agents acting on complete, multi-object 3D scenes through a unified agent-environment loop. Results show substantial room for improvement, with no tested configuration performing consistently across tasks.

  • Covers five visually conditioned 3D tasks derived from 210 source instances and 520 task cases.
  • Accepts PNG images or sampled video frames, plus supplied 3D assets where applicable.
  • Scores final outputs against hidden ground truth using task-specific geometric metrics.
  • Eleven proprietary VLM configurations achieved overall scores of 38.6–50.2.

view merged work →