🛰️ Daily AI Frontier
‹ back to 2026-08-07

Visual Grounding in Zero-Shot Vision-Language Control

arXiv cs.RO Embodied AI & Robotics J. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà 2026-08-06
Representative image for Visual Grounding in Zero-Shot Vision-Language Control

TL;DR - An arXiv cs.RO study that stress-tests vision-language models used as zero-shot robot/driving controllers with input ablations, finding most "successful" trajectories are not actually grounded in visual input. It matters because it shows benchmark scores can be produced by simulator dynamics and conservative action priors rather than perception.

  • Evaluation spanned 32,874 scored calls across nine direct-action models, six structured local VLMs, and a VLM-MPC hierarchy, over two embodiments and three simulators, using blind-image controls, repeated inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks.
  • Direct-control results were largely negative: a constant-SLOW policy beat a scripted geometric controller, several models were image-invariant or near-constant, and models that detected longitudinal hazards still failed to swap LEFT/RIGHT under reflection; no local VLM met joint longitudinal and lateral grounding criteria.
  • Failures are modular, not inherent to the stimuli: an image-only deterministic positive control estimated lead gap at 0.090 m MAE with exact mirror equivariance, confirming sufficient visual information was present.
  • A leakage-controlled symmetry-consensus guardian (two models picked from 16 calibration frames, frozen 2-of-4 hazard vote across original and reflected views) hit 0.954 balanced accuracy on 272 held-out frames (95% CI [0.895, 0.990]), 0.973 when abstaining on ties at 0.824 coverage; offline modular replay reached 0.934 action agreement, supporting VLMs as bounded hazard assistants rather than monolithic controllers.

view merged work →