REVIEW 9 cited by
A Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recommendation systems have become popular and effective tools to help users discover their interesting items by modeling the user preference and item property based on implicit interactions (e.g., purchasing and clicking). Humans perceive the world by processing the modality signals (e.g., audio, text and image), which inspired researchers to build a recommender system that can understand and interpret data from different modalities. Those models could capture the hidden relations between different modalities and possibly recover the complementary information which can not be captured by a uni-modal approach and implicit interactions. The goal of this survey is to provide a comprehensive review of the recent research efforts on the multimodal recommendation. Specifically, it shows a clear pipeline with commonly used techniques in each step and classifies the models by the methods used. Additionally, a code framework has been designed that helps researchers new in this area to understand the principles and techniques, and easily runs the SOTA models. Our framework is located at: https://github.com/enoche/MMRec
Forward citations
Cited by 9 Pith papers
-
Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders
Per-user modality weighting does not beat a single global modality weight once conditions are matched, and eval-time shuffle controls overstate personalization.
-
One Graph, Multiple Gains: Single High-Quality Item-Item Graph for Multimodal Recommendation
A single NCER-refined item-item graph, reused via adaptive gating, UI expansion, and discounted soft-positive BPR, improves multimodal recommendation accuracy and efficiency.
-
Binge Watch: Reproducible Multimodal Benchmarks Datasets for Large-Scale Movie Recommendation on MovieLens-10M and 20M
M3L-10M and M3L-20M add plot, poster, audio, and video embeddings to MovieLens and release them publicly as reproducible multimodal benchmarks.
-
Hi-SAM: A Hierarchical Structure-Aware Multi-modal Framework for Large-Scale Recommendation
Hi-SAM improves semantic-ID multimodal recommendation by disentangling shared versus modality-specific item codes and by letting transformers access history only through compressed anchor tokens.
-
The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal Recommendation
A multimodal recommender that trains without graph convolution and applies it only at test time outperforms graph-trained baselines while training much faster.
-
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
VRAgent-R1 uses an MLLM agent to summarize videos and a reinforcement-learned agent to simulate user choices, improving video recommendation and user-decision simulation on MicroLens-100K.
-
Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation
MambaRec improves multimodal recommendation accuracy on Baby, Sports, and Clothing datasets through local dilated-attention alignment and global MMD/contrastive alignment.
-
A Scenario-Oriented Survey of Federated Recommender Systems: Techniques, Challenges, and Future Directions
A scenario-oriented taxonomy of federated recommender systems that argues research should be organized around recommendation use cases rather than federated-learning abstractions.
-
FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation
FindRec combines Mamba temporal encoding, RBF-kernel cross-modal alignment, and expert routing to improve multimodal sequential recommendation, reporting 1.0 to 3.3 percent relative gains over baselines, with no proof...
Discussion (0). Continue with ORCID to comment.