🛰️ Daily AI Frontier
‹ back to 2026-09-07

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Research Robotics AI

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - ROBORMBENCH reveals that vision-language reward models can assign contradictory rewards to identical robot trajectories when goal instructions are merely paraphrased. This fragility threatens the reliability of VLM-guided robot learning.

  • The benchmark contains 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases across lexical, syntactic, and action-goal rewrites.
  • Paraphrases can substantially change predicted progress scores and even flip the same behavior between failure and success.
  • Instability affects both proprietary and open-source VLMs, worsens with more divergent rewrites, and is not consistently mitigated by model scale or explicit reasoning.
  • Dedicated reward models using trajectory-grounded supervision are substantially more stable.

Sources (1)

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

arXiv cs.RO Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No 2026-09-04 arXiv:2609.05401
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:26:03.792556 UTC

TL;DR - ROBORMBENCH reveals that vision-language reward models can assign contradictory rewards to identical robot trajectories when goal instructions are merely paraphrased. This fragility threatens the reliability of VLM-guided robot learning.

  • The benchmark contains 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases across lexical, syntactic, and action-goal rewrites.
  • Paraphrases can substantially change predicted progress scores and even flip the same behavior between failure and success.
  • Instability affects both proprietary and open-source VLMs, worsens with more divergent rewrites, and is not consistently mitigated by model scale or explicit reasoning.
  • Dedicated reward models using trajectory-grounded supervision are substantially more stable.
item →