Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
Merged summary
TL;DR - Mixture-of-Thought-Tokens (Motto) unifies spatial perception and reasoning for free-form multimodal grounding. It matters because existing approaches often trade precise localization for complex reasoning ability.
- Spatially-Grounded Thought Tokenization aligns special tokens with visual locations for interpretable spatial correspondence.
- A Context-Adaptive Chain-of-Tokens dynamically switches grounding modes within interleaved reasoning chains.
- PR-Bench evaluates the gap between perception and reasoning in referring-expression comprehension.
- The authors report state-of-the-art performance across diverse free-form grounding tasks.
Sources (1)
Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
TL;DR - Mixture-of-Thought-Tokens (Motto) unifies spatial perception and reasoning for free-form multimodal grounding. It matters because existing approaches often trade precise localization for complex reasoning ability.
- Spatially-Grounded Thought Tokenization aligns special tokens with visual locations for interpretable spatial correspondence.
- A Context-Adaptive Chain-of-Tokens dynamically switches grounding modes within interleaved reasoning chains.
- PR-Bench evaluates the gap between perception and reasoning in referring-expression comprehension.
- The authors report state-of-the-art performance across diverse free-form grounding tasks.