Pith. sign in

REVIEW 3 cited by

When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09132 v2 pith:XUI4K4R3 submitted 2024-10-11 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords multimodalattributesmagbtextitdatasetgraphgraphsattributed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multimodal Attributed Graphs (MAGs) are ubiquitous in real-world applications, encompassing extensive knowledge through multimodal attributes attached to nodes (e.g., texts and images) and topological structure representing node interactions. Despite its potential to advance diverse research fields like social networks and e-commerce, MAG representation learning (MAGRL) remains underexplored due to the lack of standardized datasets and evaluation frameworks. In this paper, we first propose MAGB, a comprehensive MAG benchmark dataset, featuring curated graphs from various domains with both textual and visual attributes. Based on MAGB dataset, we further systematically evaluate two mainstream MAGRL paradigms: $\textit{GNN-as-Predictor}$, which integrates multimodal attributes via Graph Neural Networks (GNNs), and $\textit{VLM-as-Predictor}$, which harnesses Vision Language Models (VLMs) for zero-shot reasoning. Extensive experiments on MAGB reveal following critical insights: $\textit{(i)}$ Modality significances fluctuate drastically with specific domain characteristics. $\textit{(ii)}$ Multimodal embeddings can elevate the performance ceiling of GNNs. However, intrinsic biases among modalities may impede effective training, particularly in low-data scenarios. $\textit{(iii)}$ VLMs are highly effective at generating multimodal embeddings that alleviate the imbalance between textual and visual attributes. These discoveries, which illuminate the synergy between multimodal attributes and graph topologies, contribute to reliable benchmarks, paving the way for future MAG research. The MAGB dataset and evaluation pipeline are publicly available at https://github.com/sktsherlock/MAGB.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RHEA: Reliability-Harmonized Reconstruction and Assignment for Robust Multimodal-Attributed Graph Clustering

    cs.LG 2026-08 conditional novelty 6.0 of 10

    RHEA estimates node-specific modality reliability from neighborhood agreement, reconstructs unreliable modalities, and uses reliability-aware optimal transport to cluster multimodal graphs, beating prior methods on fo...

  2. Disentangling Homophily and Heterophily in Multimodal Graph Clustering

    cs.AI 2025-07 conditional novelty 6.0 of 10

    DMGC clusters multimodal multi-relational graphs by disentangling homophilic and heterophilic edges and fusing low-pass and high-pass filtered representations in a self-supervised way, reporting SOTA accuracy on six b...

  3. Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A unified comparison across six multimodal graph datasets shows that fine-tuned multimodal LLMs used as direct predictors achieve the highest node classification accuracy, even without graph structure input.

Pith tools