A multimodal CNN on 87,547 Vogue images classifies fashion houses at 78.2% top-1 accuracy, decades at 88.6%, and years at 58.3% with 2.2-year mean error, and shows texture and luminance carry most of the house-identity signal.
A comprehensive survey on multimodal recommender systems: Taxonomy, evaluation, and future directions
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 5roles
background 1polarities
background 1representative citing papers
JBM-Diff applies conditional graph diffusion to remove preference-irrelevant multimodal noise and false-positive/negative behaviors, then augments training data via partial-order credibility scoring.
URecJPQ compresses user and item embeddings via joint product quantization for multimodal top-k recommendation, cutting checkpoint size 86-98% and parameters 98-99% with average 8.5% recall drop across three datasets.
MAIL constructs modality-aware ID-free identities via dynamic positional encoding modulation and applies counterfactual structure learning with popularity penalization, yielding 7.81% Recall@10 and 12.81% NDCG@10 gains on five Amazon datasets.
citing papers explorer
-
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing
A multimodal CNN on 87,547 Vogue images classifies fashion houses at 78.2% top-1 accuracy, decades at 88.6%, and years at 58.3% with 2.2-year mean error, and shows texture and luminance carry most of the house-identity signal.
-
Joint Behavior-guided and Modality-coherence Conditional Graph Diffusion Denoising for Multi Modal Recommendation
JBM-Diff applies conditional graph diffusion to remove preference-irrelevant multimodal noise and false-positive/negative behaviors, then augments training data via partial-order credibility scoring.
-
URecJPQ: Memory-efficient Multimodal Recommendation Models through RecJPQ in Large-Scale Scenarios
URecJPQ compresses user and item embeddings via joint product quantization for multimodal top-k recommendation, cutting checkpoint size 86-98% and parameters 98-99% with average 8.5% recall drop across three datasets.
-
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
MAIL constructs modality-aware ID-free identities via dynamic positional encoding modulation and applies counterfactual structure learning with popularity penalization, yielding 7.81% Recall@10 and 12.81% NDCG@10 gains on five Amazon datasets.
- TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning