Pith. sign in

REVIEW 2 cited by

DiffMM: Multi-Modal Diffusion Model for Recommendation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11781 v1 pith:BPGKBRKD submitted 2024-06-17 cs.IR

classification cs.IR
keywords multi-modaldiffmmdiffusionmodelgraphlearningmodelingsystems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rise of online multi-modal sharing platforms like TikTok and YouTube has enabled personalized recommender systems to incorporate multiple modalities (such as visual, textual, and acoustic) into user representations. However, addressing the challenge of data sparsity in these systems remains a key issue. To address this limitation, recent research has introduced self-supervised learning techniques to enhance recommender systems. However, these methods often rely on simplistic random augmentation or intuitive cross-view information, which can introduce irrelevant noise and fail to accurately align the multi-modal context with user-item interaction modeling. To fill this research gap, we propose a novel multi-modal graph diffusion model for recommendation called DiffMM. Our framework integrates a modality-aware graph diffusion model with a cross-modal contrastive learning paradigm to improve modality-aware user representation learning. This integration facilitates better alignment between multi-modal feature information and collaborative relation modeling. Our approach leverages diffusion models' generative capabilities to automatically generate a user-item graph that is aware of different modalities, facilitating the incorporation of useful multi-modal knowledge in modeling user-item interactions. We conduct extensive experiments on three public datasets, consistently demonstrating the superiority of our DiffMM over various competitive baselines. For open-sourced model implementation details, you can access the source codes of our proposed framework at: https://github.com/HKUDS/DiffMM .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Knowledge-aware Diffusion-Enhanced Multimedia Recommendation

    cs.MM 2025-07 conditional novelty 5.0 of 10

    KDiffE improves multimedia recommendation by weighting user-item edges with a random-walk Jaccard attention matrix and training a user-guided diffusion model to generate a denoised knowledge-graph contrastive view.

  2. Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal Recommendations

    cs.IR 2025-01 conditional novelty 5.0 of 10

    MoDiCF combines per-modality diffusion with modality-aware conditioning and a counterfactual re-scoring step, improving accuracy and item exposure fairness on incomplete multimodal recommendation datasets.

Pith tools