🛰️ Daily AI Frontier
‹ back to 2026-07-16

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

arXiv cs.CV Agents & Tool Use Ting Lei, Jialin Liu, Zhu Xu, Yuxin Peng, Yang Liu 2026-07-15

TL;DR — AgentHOI is a training-free, agentic framework that orchestrates multiple vision foundation modules via multimodal LLM reasoning to detect human-object interactions (HOI) in open-world images, matching or beating supervised methods without any HOI training data.

  • Reframes HOI detection away from supervised closed-set classifiers toward an agentic workflow that coordinates complementary vision foundation modules for open-ended semantic reasoning and spatial grounding.
  • Introduces Context-aware Multi-round Reasoning to iteratively refine interaction hypotheses for more exhaustive, compositional interaction discovery.
  • Introduces Multifaceted Interaction Localization, generating instance-specific descriptions combining semantic, spatial, and appearance cues to sharpen grounding.
  • Reports superior real-world performance over state-of-the-art supervised and weakly-supervised baselines despite requiring no HOI-detection training data (specific benchmark numbers not provided in the abstract).

view merged work →