🛰️ Daily AI Frontier
‹ back to 2026-07-22

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

Research Multimodal Graph Learning

Merged summary

TL;DR - OMG-VLM uses one vision-language backbone to learn across graphs with text, image, or mixed node attributes. It aims to replace modality-specific graph models while improving generalization to unseen graphs and schemas.

  • Structure-aware graph adapters integrate neighborhood information into the VLM’s embedding space.
  • A single model supports node classification and link prediction across heterogeneous attributed graphs.
  • Experiments report consistent gains over GNN- and LLM-based baselines.
  • The model demonstrates generalization to unseen graphs and varying modality configurations.

Sources (1)

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

arXiv cs.LG Jiayi Yang, Yifang Chen, Yuanfu Sun, Jiajin Liu, Qiaoyu Tan 2026-07-21 arXiv:2607.19128

TL;DR - OMG-VLM uses one vision-language backbone to learn across graphs with text, image, or mixed node attributes. It aims to replace modality-specific graph models while improving generalization to unseen graphs and schemas.

  • Structure-aware graph adapters integrate neighborhood information into the VLM’s embedding space.
  • A single model supports node classification and link prediction across heterogeneous attributed graphs.
  • Experiments report consistent gains over GNN- and LLM-based baselines.
  • The model demonstrates generalization to unseen graphs and varying modality configurations.
item →