🛰️ Daily AI Frontier
‹ back to 2026-07-22

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

Research Multimodal & Generative

Ranking

Overall 62
Content 70
Popularity 44

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper introduces DiSG, an 86-category in-the-wild benchmark for open-vocabulary gaze object prediction, plus a framework for recognizing previously unseen gaze targets. It expands gaze understanding beyond fixed labels and scene-specific datasets.

  • Uses text-driven object discovery to identify potential gaze targets.
  • Applies gaze-guided selection to choose the attended object from candidates.
  • Introduces Gradient-Informed Selection Tuning (GIST) to selectively update vocabulary-relevant parameters.
  • The model performs effectively in open-vocabulary settings and reportedly surpasses prior methods in closed-vocabulary evaluation.

Sources (1)

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

arXiv cs.CV Binglu Wang, Sensen Niu, Ying Chen, Guangyu Guo 2026-07-21 arXiv:2607.18827
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-21 14:39:01.610470 UTC

TL;DR - This paper introduces DiSG, an 86-category in-the-wild benchmark for open-vocabulary gaze object prediction, plus a framework for recognizing previously unseen gaze targets. It expands gaze understanding beyond fixed labels and scene-specific datasets.

  • Uses text-driven object discovery to identify potential gaze targets.
  • Applies gaze-guided selection to choose the attended object from candidates.
  • Introduces Gradient-Informed Selection Tuning (GIST) to selectively update vocabulary-relevant parameters.
  • The model performs effectively in open-vocabulary settings and reportedly surpasses prior methods in closed-vocabulary evaluation.
item →