One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models
Merged summary
TL;DR - OMG-VLM uses one vision-language backbone to learn across graphs with text, image, or mixed node attributes. It aims to replace modality-specific graph models while improving generalization to unseen graphs and schemas.
- Structure-aware graph adapters integrate neighborhood information into the VLM’s embedding space.
- A single model supports node classification and link prediction across heterogeneous attributed graphs.
- Experiments report consistent gains over GNN- and LLM-based baselines.
- The model demonstrates generalization to unseen graphs and varying modality configurations.
Sources (1)
One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models
TL;DR - OMG-VLM uses one vision-language backbone to learn across graphs with text, image, or mixed node attributes. It aims to replace modality-specific graph models while improving generalization to unseen graphs and schemas.
- Structure-aware graph adapters integrate neighborhood information into the VLM’s embedding space.
- A single model supports node classification and link prediction across heterogeneous attributed graphs.
- Experiments report consistent gains over GNN- and LLM-based baselines.
- The model demonstrates generalization to unseen graphs and varying modality configurations.