Open-Vocabulary Gaze Object Prediction: Benchmark and Method
Merged summary
TL;DR - This paper introduces DiSG, an 86-category in-the-wild benchmark for open-vocabulary gaze object prediction, plus a framework for recognizing previously unseen gaze targets. It expands gaze understanding beyond fixed labels and scene-specific datasets.
- Uses text-driven object discovery to identify potential gaze targets.
- Applies gaze-guided selection to choose the attended object from candidates.
- Introduces Gradient-Informed Selection Tuning (GIST) to selectively update vocabulary-relevant parameters.
- The model performs effectively in open-vocabulary settings and reportedly surpasses prior methods in closed-vocabulary evaluation.
Sources (1)
Open-Vocabulary Gaze Object Prediction: Benchmark and Method
TL;DR - This paper introduces DiSG, an 86-category in-the-wild benchmark for open-vocabulary gaze object prediction, plus a framework for recognizing previously unseen gaze targets. It expands gaze understanding beyond fixed labels and scene-specific datasets.
- Uses text-driven object discovery to identify potential gaze targets.
- Applies gaze-guided selection to choose the attended object from candidates.
- Introduces Gradient-Informed Selection Tuning (GIST) to selectively update vocabulary-relevant parameters.
- The model performs effectively in open-vocabulary settings and reportedly surpasses prior methods in closed-vocabulary evaluation.