SceneActBench: Can Agents Act on the 3D Scenes They See?
TL;DR - SceneActBench evaluates VLM agents acting on complete, multi-object 3D scenes through a unified agent-environment loop. Results show substantial room for improvement, with no tested configuration performing consistently across tasks.
- Covers five visually conditioned 3D tasks derived from 210 source instances and 520 task cases.
- Accepts PNG images or sampled video frames, plus supplied 3D assets where applicable.
- Scores final outputs against hidden ground truth using task-specific geometric metrics.
- Eleven proprietary VLM configurations achieved overall scores of 38.6–50.2.