Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
Merged summary
TL;DR — AgentHOI is a training-free, agentic framework that orchestrates multiple vision foundation modules via multimodal LLM reasoning to detect human-object interactions (HOI) in open-world images, matching or beating supervised methods without any HOI training data.
- Reframes HOI detection away from supervised closed-set classifiers toward an agentic workflow that coordinates complementary vision foundation modules for open-ended semantic reasoning and spatial grounding.
- Introduces Context-aware Multi-round Reasoning to iteratively refine interaction hypotheses for more exhaustive, compositional interaction discovery.
- Introduces Multifaceted Interaction Localization, generating instance-specific descriptions combining semantic, spatial, and appearance cues to sharpen grounding.
- Reports superior real-world performance over state-of-the-art supervised and weakly-supervised baselines despite requiring no HOI-detection training data (specific benchmark numbers not provided in the abstract).
Sources (1)
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
TL;DR — AgentHOI is a training-free, agentic framework that orchestrates multiple vision foundation modules via multimodal LLM reasoning to detect human-object interactions (HOI) in open-world images, matching or beating supervised methods without any HOI training data.
- Reframes HOI detection away from supervised closed-set classifiers toward an agentic workflow that coordinates complementary vision foundation modules for open-ended semantic reasoning and spatial grounding.
- Introduces Context-aware Multi-round Reasoning to iteratively refine interaction hypotheses for more exhaustive, compositional interaction discovery.
- Introduces Multifaceted Interaction Localization, generating instance-specific descriptions combining semantic, spatial, and appearance cues to sharpen grounding.
- Reports superior real-world performance over state-of-the-art supervised and weakly-supervised baselines despite requiring no HOI-detection training data (specific benchmark numbers not provided in the abstract).