Pith. sign in

REVIEW 2 cited by

Triple Modality Fusion: Aligning Visual, Textual, and Graph Data with Large Language Models for Multi-Behavior Recommendations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12228 v2 pith:3IMGWJCQ submitted 2024-10-16 cs.IR cs.AIcs.CL

classification cs.IRcs.AIcs.CL
keywords datamodelsfusionitemuserbehaviorsfeaturesgraph
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrating diverse data modalities is crucial for enhancing the performance of personalized recommendation systems. Traditional models, which often rely on singular data sources, lack the depth needed to accurately capture the multifaceted nature of item features and user behaviors. This paper introduces a novel framework for multi-behavior recommendations, leveraging the fusion of triple-modality, which is visual, textual, and graph data through alignment with large language models (LLMs). By incorporating visual information, we capture contextual and aesthetic item characteristics; textual data provides insights into user interests and item features in detail; and graph data elucidates relationships within the item-behavior heterogeneous graphs. Our proposed model called Triple Modality Fusion (TMF) utilizes the power of LLMs to align and integrate these three modalities, achieving a comprehensive representation of user behaviors. The LLM models the user's interactions including behaviors and item features in natural languages. Initially, the LLM is warmed up using only natural language-based prompts. We then devise the modality fusion module based on cross-attention and self-attention mechanisms to integrate different modalities from other models into the same embedding space and incorporate them into an LLM. Extensive experiments demonstrate the effectiveness of our approach in improving recommendation accuracy. Further ablation studies validate the effectiveness of our model design and benefits of the TMF.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings

    cs.IR 2025-07 conditional novelty 6.0 of 10

    A production system that combines object-detection-based image cropping and LLM-based text rewriting with CLIP fine-tuning reports large gains in multimodal retrieval and online recommendation metrics at Walmart scale.

  2. Graph Foundation Models for Recommendation: A Comprehensive Survey

    cs.IR 2025-02 conditional novelty 4.0 of 10

    A comprehensive survey that categorizes graph foundation model approaches to recommendation into graph-augmented LLM, LLM-augmented graph, and LLM-graph harmonization.

Pith tools