Pith. sign in

REVIEW 4 major objections 4 minor 76 references

Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal Recommendations

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Incomplete multimodal recommendations can be made more accurate and fairer by generating missing modality features with a conditioned diffusion process and subtracting an estimated visibility-bias effect from ranking scores.

desk verdict A credible engineering contribution with an over-claimed causal core: the diffusion completion part holds up, but the counterfactual debiasing is a tuned product-form reranking. read the letter →

arxiv 2501.11916 v4 pith:LQMYP4XA submitted 2025-01-21 cs.IR

classification cs.IR
keywords multimodalrecommendationsmissingmodalitiesvisibilitybiasdiffusionmodelscounterfactualinferencefairnessdatacompletionrecommendation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tackles two problems that appear together when item modalities are missing in multimodal recommenders: generated replacements for missing modalities are often poor, and items with missing modalities receive less exposure than their preference alignment justifies. MoDiCF combines a modality-conditioned diffusion module that generates and iteratively refines missing modality features from per-modality latent distributions with a counterfactual recommendation module that subtracts an item-only predictor's contribution from ranking scores. On three real-world datasets, the framework reports consistent gains over existing methods in both accuracy metrics and a fairness metric measuring whether incomplete items appear in top-K lists at their population rate. If the claim holds, MoDiCF offers a practical recipe for treating incomplete items fairly without sacrificing recommendation quality.

What carries the argument

The framework rests on two mechanisms. The first is a modality-diffused data completion module: a denoising diffusion probabilistic model run separately for each modality in a latent space, conditioned on the other observed modalities through cross-modal attention fusion, and followed by iterative refinement in which generated values are reinserted into the training data each epoch. This module is what lets missing modalities be generated from modality-specific distributions rather than generic imputation. The second is the counterfactual recommendation module: a multimodal recommender plus a content-only item predictor whose output, scaled by a sigmoid and a coefficient, is subtracted from the recommender score to remove the estimated direct effect of visibility bias. The item predictor is the component that is supposed to isolate the bias.

What would settle it

A direct control experiment would settle the question: on the same incomplete datasets, replace Eq. 17 with a non-causal reranker that boosts incomplete items to match MoDiCF's exposure level, and compare the fairness-accuracy tradeoff. If the naive reranker matches MoDiCF, the counterfactual form is not doing the causal work claimed.

Watch

Extended reading notes

Core claim

MoDiCF's central claim is that visibility bias in incomplete multimodal recommendation is a natural direct effect from item features to ranking scores, and that it can be removed by a counterfactual adjustment. The framework trains a standard multimodal recommender on completed data and a separate item predictor that sees only item content, then combines their outputs as $\hat{y}_{u,i}\,\mathrm{sigm}(\hat{y}_i) - \gamma\,\mathrm{sigm}(\hat{y}_i)$, where $\gamma$ is an empirically chosen constant. In the paper's causal graph, the item predictor estimates the direct path $I^* \to Y$ that makes incomplete items lose exposure. The accompanying diffusion module supplies the missing modality features so that the recommender can exploit multimodal information instead of defaulting to complete items.

Load-bearing premise

The debiasing step assumes the item-only predictor measures the visibility-bias effect rather than item quality, popularity, or missingness patterns, and that the product-form subtraction removes the bias without any causal identification argument or ablation against a non-causal reranker.

Editorial extensions

If this is right

  • Items with missing modalities receive top-K exposure closer to their population share, as measured by F@K, while accuracy on Recall@K, Precision@K, and NDCG@K is maintained or improved.
  • The modality-diffused completion module outperforms mean, zero, random, and nearest-neighbor imputation, indicating that capturing per-modality distributions matters for downstream ranking quality.
  • Removing the counterfactual module causes a sharp drop in fairness scores but a smaller drop in accuracy, showing that debiasing and completion contribute in different proportions to the two goals.
  • The framework can be instantiated with different multimodal recommender backbones, and both tested variants improve over their original models, suggesting the two modules transfer across recommendation architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The counterfactual adjustment can be read as a re-ranking rule: it up-weights items whose content-only predictor score is low compared with the recommender score. If the item predictor captures content quality or popularity rather than missingness, the reported fairness gain could be a re-ranking artifact rather than true causal debiasing.
  • A natural testable extension is to condition fairness evaluation on the number or type of missing modalities, since the paper's F@K metric treats all incomplete items equally and may hide residual bias against items with more severe incompleteness.
  • The iterative refinement loop reuses generated features as conditions for other modalities during training, so errors in one generated modality can propagate; probing this error propagation would clarify when refinement helps versus hurts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes MoDiCF, a framework for incomplete multimodal recommendation that combines a modality-diffused data completion module (MDDC) with a counterfactual multimodal recommendation module (CFMR). MDDC generates missing modality features through a latent diffusion process with modality-aware conditioning and iterative refinement. CFMR pairs a graph-based multimodal recommender with an item predictor trained on multimodal content alone, and then combines the two scores using Eq. (17), which is presented as a counterfactual adjustment that removes the natural direct effect of item incompleteness. The paper evaluates MoDiCF on Baby, Tiktok, and Allrecipes under random missing-modality settings, reporting consistent gains in Recall, Precision, NDCG, a newly defined exposure fairness metric F@K, and a combined Ffuse@K metric, with 10 repetitions, standard deviations, and paired t-tests. The code and processed datasets are released. The reported accuracy gains are credible and the engineering is thorough, but the central causal claim that Eq. (17) implements the TE/NDE/TIE decomposition in Eqs. (9)-(10) is not derived, and the experiments do not include a non-causal re-ranking control that would distinguish counterfactual debiasing from a tuned exposure rescoring.

Significance. If the causal claim were established, MoDiCF would be a practically significant contribution: it provides a working diffusion-based generator for missing multimodal content, a reusable fairness metric, and a framework that can be instantiated with different backbone recommenders, all validated with unusually careful experimental protocol (10 runs, standard deviations, significance tests, and released code). The measured gains in accuracy and exposure fairness are real and worth reporting. However, the paper's main conceptual selling point is the counterfactual elimination of visibility bias, and that claim currently rests on an estimator, Eq. (17), whose connection to the NDE/TIE decomposition is asserted rather than shown. The study is therefore stronger as an empirical system paper with a heuristic exposure adjustment than as a causal debiasing method, and the framing needs to be corrected or the derivation supplied.

major comments (4)
  1. [§4.2, Eqs. (9)-(10); §4.2.2, Eq. (17)] Equation (17) does not follow from the TE/NDE/TIE decomposition. The text defines TIE = TE - NDE = Y_{i,M_i} - Y_{i,M_i*}, i.e., the difference between the ranking score under complete item features and the score under incomplete item features with the mediator fixed to its complete value. The implemented estimator is y_{u,i} = hat y_{u,i} * sigm(hat y_i) - gamma * sigm(hat y_i). Neither term is identified with Y_{i,M_i} or Y_{i,M_i*}; the product form and the gamma-scaled subtraction are new modeling choices, and gamma is stated to be 'usually chosen empirically.' Without a derivation or an identification argument, the paper's claim that NDE is eliminated and TIE is isolated is not supported. This is the load-bearing step for the fairness contribution, so it must either be derived from the causal model or explicitly reframed as a heuristic exposure adjustment.
  2. [§4.2.2, item predictor; Figure 2(c)] The item predictor is an MLP trained with BPR on the same user-item interaction matrix used for evaluation, using multimodal content as input. Under Figure 2(c), the path I* -> Y is claimed to represent visibility bias, but a model fit to observed interactions can equally encode item popularity, content quality, or user preference confounded with missingness. No identification condition is stated: the graph in Figure 2(c) is asserted, not checked, and the missingness mechanism is random by construction in the experiments, which is a favorable case that does not test the causal graph. If the item predictor captures content quality rather than missingness, Eq. (17) becomes a content-based re-ranking rule, and the fairness gain is not evidence of counterfactual debiasing.
  3. [§5.4, Table 5; §5.2.3, gamma tuning] The ablation MoDiCF-C only removes the entire counterfactual adjustment; it does not include a non-causal control that re-ranks by item-content scores without the causal framing. A control such as y_{u,i} = hat y_{u,i} - beta * sigm(hat y_i) with beta tuned on the validation set would show whether the fairness gain comes from the causal mechanism or from any content-based rescoring. The absence of this control is important because gamma itself is tuned per dataset (0.01, 20, 20 in Table 4), so the fairness result is partly a function of the hyperparameter choice. Please add such a control or explicitly limit the claim to 'exposure adjustment improves F@K.'
  4. [§5.3, Table 2] The significance reporting for the fairness metric is internally difficult to interpret. On Baby at K=10, MoDiCF's F@10 is 87.24, which is below LightGCN's 87.44 and AutoCF's 87.79, even though unimodal methods are described as the ideal fairness reference and MoDiCF is marked with an asterisk for significant improvement. Please clarify the comparison direction for F@K: is the significance test relative to all listed methods, or only to multimodal/incomplete-MMRec baselines? If the claim is 'fairness close to the unimodal reference,' the current asterisk is misleading; if the claim is 'better than all multimodal baselines,' the test should exclude LightGCN and AutoCF from the comparison.
minor comments (4)
  1. [§5.2.3] There is a typo in 'paramter' and, in Appendix C.1, 'dominently'; please proofread the final version.
  2. [Table 5] The first ablation row is labeled 'MMoDiCF-D+M' in the table header while the text refers to 'MoDiCF-D+M'; the inconsistent naming should be fixed.
  3. [§4.2.2 and Eq. (17)] The notation y_{u,i*} and the subscript i* are used without a precise definition of the counterfactual item state. Define Y_{i,M_i}, Y_{i*,M_i*}, and Y_{i,M_i*} explicitly in terms of the item predictor and recommender outputs, or replace them with notation that matches the implementation.
  4. [§5.2.2, Eq. (19)-(20)] The F@K metric is defined as a proportion ratio, but the harmonic-mean combination Ffuse@K treats F@K and Precision@K as if they were in the same scale. Please state explicitly that both are in [0,1] and whether P_r@K is computed over the same K for which Precision is reported.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the accuracy and fairness gains are measured on independent held-out test sets, and although Eq. 17 is under-identified as a counterfactual adjustment, it is not equivalent to its inputs by construction.

full rationale

The paper's central empirical claims are not circular: MoDiCF is compared against 13 baselines on fixed incomplete versions of Baby, Tiktok and Allrecipes with held-out test sets, and hyperparameters including the counterfactual coefficient gamma are tuned on the validation set, not on the test metrics reported in Table 2. The MDDC completion module is a conditional DDPM trained with a standard diffusion objective on observed modalities, and its contribution is tested by ablations (MoDiCF-D+{M,Z,R,N}, MoDiCF-con) against independent imputation baselines; no stage of the derivation defines the output in terms of the target metric. The weakest point is the counterfactual module: the paper defines TE/NDE/TIE in Eqs. 9-10 and then states that Eq. 17 follows from them, but the product form y = y_hat * sigm(y_hat_i) - gamma * sigm(y_hat_i) is not entailed by the NDE/TIE decomposition, and the paper itself admits gamma is 'usually chosen empirically [47]' (Section 4.2.2). That is an identification/validity gap, not a circular reduction: Eq. 17 is not equal to Eq. 10 by construction, and the fairness gains are not statistically forced because the test set is independent of the gamma tuning. The self-citations ([16], [52]) are contextual and are not load-bearing for the accuracy/fairness results. The absence of a non-causal re-ranking control is a legitimate experimental concern to be weighed under correctness risk, not under circularity.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

MoDiCF introduces no new physical or conceptual entities. The central claims rest instead on a set of tuned hyperparameters, especially gamma, and on unvalidated modeling assumptions about the causal graph, the MCAR missingness process, and the item predictor as a direct-effect estimator.

free parameters (6)
  • gamma = 0.01 (Baby), 20 (Tiktok), 20 (Allrecipes)
    Counterfactual adjustment strength in Eq. 17, tuned per dataset on validation Ffuse; directly controls the fairness-accuracy trade-off.
  • alpha1 = 1, 0.7, 0.6
    Balances diffusion loss and reconstruction loss in Eq. 5; tuned per dataset in {0.1, ..., 1.0}.
  • alpha2 = 0.7, 0.3, 0.5
    Balances item-predictor BPR loss against recommender BPR loss; tuned per dataset.
  • eta = 0.7, 0.6, 0.3
    Fusion weight in Eq. 14 controlling how much modality-aware embeddings enter the graph initialization.
  • delta = 0.4, 0.3, 0.4
    Fusion weight combining modality features with ID embeddings in the final user-item representation.
  • lambda1 = 0.09, 0.06, 0.15
    Weight of the cross-modal contrastive loss in Eq. 18, searched over {0.01, ..., 0.20}.
assumptions (5)
  • standard math DDPM forward and reverse processes with Gaussian noise and 1000 steps capture modality-specific latent distributions.
    Standard diffusion assumption invoked in Section 4.1.1, following Ho et al. [12].
  • domain assumption Missing modalities in evaluation are missing completely at random, generated by random dropping with at least one observed modality per item.
    Section 5.1 constructs incomplete data this way; real-world missingness may correlate with item quality, popularity, or collection cost, which would change the visibility bias mechanism.
  • ad hoc to paper The causal graph in Figure 2(c), with a direct path I* to Y, correctly represents incomplete MMRec and has no hidden confounders.
    Assumed without identification conditions; the direct effect of incompleteness is never intervened upon in the experiments.
  • ad hoc to paper The item predictor approximates the natural direct effect NDE = Y_{i,M_i*} - Y_{i*,M_i*}.
    Section 4.2.2 states the item predictor directly measures visibility bias, but no proof or identification argument is provided.
  • ad hoc to paper Equation 17 follows from the TE, NDE, and TIE decomposition in Eqs. 9 and 10.
    The product form with sigmoid(\\hat y_i) and the empirically chosen gamma is not derived from the decomposition.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal Recommendations." pith.science (2026). https://pith.science/paper/LQMYP4XA

@misc{pith2026250111916,
  author       = {Pith},
  title        = {Pith review of: Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal Recommendations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LQMYP4XA}},
  note         = {Machine review of arXiv:2501.11916}
}
read the original abstract

Incomplete scenario is a prevalent, practical, yet challenging setting in Multimodal Recommendations (MMRec), where some item modalities are missing due to various factors. Recently, a few efforts have sought to improve the recommendation accuracy by exploring generic structures from incomplete data. However, two significant gaps persist: 1) the difficulty in accurately generating missing data due to the limited ability to capture modality distributions; and 2) the critical but overlooked visibility bias, where items with missing modalities are more likely to be disregarded due to the prioritization of items' multimodal data over user preference alignment. This bias raises serious concerns about the fair treatment of items. To bridge these two gaps, we propose a novel Modality-Diffused Counterfactual (MoDiCF) framework for incomplete multimodal recommendations. MoDiCF features two key modules: a novel modality-diffused data completion module and a new counterfactual multimodal recommendation module. The former, equipped with a particularly designed multimodal generative framework, accurately generates and iteratively refines missing data from learned modality-specific distribution spaces. The latter, grounded in the causal perspective, effectively mitigates the negative causal effects of visibility bias and thus assures fairness in recommendations. Both modules work collaboratively to address the two aforementioned significant gaps for generating more accurate and fair results. Extensive experiments on three real-world datasets demonstrate the superior performance of MoDiCF in terms of both recommendation accuracy and fairness. The code and processed datasets are released at https://github.com/JinLi-i/MoDiCF.

Figures

Figures reproduced from arXiv: 2501.11916 by the authors.

Figure 1
Figure 1. Examples of visibility bias. (a): Exposure of incom [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Causal graph for (a) RS, (b) MMRec and (c) incom [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The MoDiCF framework comprises two key modules: a modality-diffused data completion module and a counterfactual [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Comparison with different MRs on three datasets. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 7
Figure 7. Figure 7: Comparison of running time of different methods [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Impacts of balance parameters (a) 𝛼1, (b) 𝛼2, fusion weights (c) 𝜂, (d) 𝛿, (e) the number of layers 𝐿, (f) the number of attention heads 𝐻 and (g) the dimensionality 𝑑𝑚. 8 (a) and (b). We can find that 𝛼1 and 𝛼2 are relatively more stable on the Baby and Allrecipes dat…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

76 extracted references · 63 canonical work pages

  1. [1]

    Haoyue Bai, Le Wu, Min Hou, Miaomiao Cai, Zhuangzhuang He, Yuyang Zhou, Richang Hong, and Meng Wang. 2024. Multimodality Invariant Learning for Multimedia-Based New Item Recommendation. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 677–686

  2. [2]

    Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng- Ann Heng, and Stan Z. Li. 2024. A Survey on Generative Diffusion Models. IEEE Transactions on Knowledge and Data Engineering 36, 7 (2024), 2814–2830

  3. [3]

    Jingyuan Chen, Hanwang Zhang, Xiangnan He, Liqiang Nie, Wei Liu, and Tat- Seng Chua. 2017. Attentive Collaborative Filtering: Multimedia Recommendation with Item- and Component-Level Attention. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 335–344

  4. [4]

    Giannis Daras, Kulin Shah, Yuval Dagan, Aravind Gollakota, Alex Dimakis, and Adam R. Klivans. 2023. Ambient Diffusion: Learning Clean Distributions from Corrupted Data. In Advances in Neural Information Processing Systems , Vol. 36. 288–313

  5. [5]

    Yashar Deldjoo, Markus Schedl, Paolo Cremonesi, and Gabriella Pasi. 2021. Rec- ommender Systems Leveraging Multimedia Content. Comput. Surveys 53, 5 (2021), 106:1–106:38

  6. [6]

    Chongming Gao, Shiqi Wang, Shijun Li, Jiawei Chen, Xiangnan He, Wenqiang Lei, Biao Li, Yuan Zhang, and Peng Jiang. 2024. CIRS: Bursting Filter Bubbles by Coun- terfactual Interactive Recommender System. ACM Transactions on Information Systems 42, 1 (2024), 14:1–14:27

  7. [7]

    Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Deep Sparse Rectifier Neural Networks. In Proceedings of the International Conference on Artificial Intelligence and Statistics, Vol. 15. 315–323

  8. [8]

    Tsang, and Yilong Yin

    Yongshun Gong, Zhibin Li, Wei Liu, Xiankai Lu, Xinwang Liu, Ivor W. Tsang, and Yilong Yin. 2023. Missingness-Pattern-Adaptive Learning With Incomplete Data. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 11053–11066

Show all 76 references
  1. [9]

    Zhiqiang Guo, Jianjun Li, Guohui Li, Chaoyang Wang, Si Shi, and Bin Ruan. 2024. LGMRec: Local and Global Graph Learning for Multimodal Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence . 8454–8462

  2. [10]

    Ruining He and Julian J. McAuley. 2016. VBPR: Visual Bayesian Personalized Ranking from Implicit Feedback. In Proceedings of the AAAI Conference on Artifi- cial Intelligence. 144–150

  3. [11]

    Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yong-Dong Zhang, and Meng Wang. 2020. LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 639–648

  4. [12]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems, Vol. 33. 6840–6851

  5. [13]

    Liting Huang, Zhihao Zhang, Yiran Zhang, Xiyue Zhou, and Shoujin Wang. 2024. RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection. CoRR abs/2406.04906 (2024)

  6. [14]

    Yangqin Jiang, Lianghao Xia, Wei Wei, Da Luo, Kangyi Lin, and Chao Huang

  7. [15]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimiza- tion. In Proceedings of the International Conference on Learning Representations

  8. [16]

    Aggarwal

    Jin Li, Shoujin Wang, Qi Zhang, Longbing Cao, Fang Chen, Xiuzhen Zhang, Dietmar Jannach, and Charu C. Aggarwal. 2024. Causal Learning for Trustworthy Recommender Systems: A Survey. CoRR abs/2402.08241 (2024)

  9. [17]

    Shuaiyang Li, Dan Guo, Kang Liu, Richang Hong, and Feng Xue. 2023. Multi- modal Counterfactual Learning Network for Multimedia-based Recommendation. In Proceedings of the International ACM SIGIR Conference on Research and Devel- opment in Information Retrieval . 1539–1548

  10. [18]

    Yunqi Li, Hanxiong Chen, Shuyuan Xu, Yingqiang Ge, and Yongfeng Zhang. 2021. Towards Personalized Fairness based on Causal Notion. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. 1054–1063

  11. [19]

    Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2024. Foundations & Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions. Comput. Surveys 56, 10 (2024), 264

  12. [20]

    Zhenghong Lin, Yanchao Tan, Yunfei Zhan, Weiming Liu, Fan Wang, Chaochao Chen, Shiping Wang, and Carl Yang. 2023. Contrastive Intra- and Inter-Modality Generation for Enhancing Incomplete Multimedia Recommendation. In Proceed- ings of the ACM International Conference on Multim...

  13. [21]

    Qidong Liu, Jiaxi Hu, Yutian Xiao, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Qing Li, and Jiliang Tang. 2024. Multimodal Recommender Systems: A Survey. Comput. Surveys (2024). https://doi.org/10.1145/3695461

  14. [22]

    Qidong Liu, Fan Yan, Xiangyu Zhao, Zhaocheng Du, Huifeng Guo, Ruiming Tang, and Feng Tian. 2023. Diffusion Augmentation for Sequential Recommendation. In Proceedings of the ACM International Conference on Information and Knowledge Management. 1576–1586

  15. [23]

    Qijiong Liu, Jieming Zhu, Yanting Yang, Quanyu Dai, Zhaocheng Du, Xiao-Ming Wu, Zhou Zhao, Rui Zhang, and Zhenhua Dong. 2024. Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and ...

  16. [24]

    Xinwang Liu, Miaomiao Li, Chang Tang, Jingyuan Xia, Jian Xiong, Li Liu, Marius Kloft, and En Zhu. 2021. Efficient and Effective Regularized Incomplete Multi- View Clustering. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 8 (2021), 2634–2646

  17. [25]

    Jing Long, Guanhua Ye, Tong Chen, Yang Wang, Meng Wang, and Hongzhi Yin

  18. [26]

    Haokai Ma, Ruobing Xie, Lei Meng, Xin Chen, Xu Zhang, Leyu Lin, and Zhan- hui Kang. 2024. Plug-In Diffusion Model for Sequential Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence . 8886–8894

  19. [27]

    In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining

    Diffusion-Based Cloud-Edge-Device Collaborative Learning for Next POI Recommendations. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 2026–2036

  20. [28]

    Haokai Ma, Yimeng Yang, Lei Meng, Ruobing Xie, and Xiangxu Meng. 2024. Multimodal Conditioned Diffusion Model for Recommendation. In Companion Proceedings of the ACM Web Conference . 1733–1740

  21. [29]

    Haokai Ma, Ruobing Xie, Lei Meng, Yimeng Yang, Xingwu Sun, and Zhanhui Kang. 2024. SeeDRec: Sememe-based Diffusion for Sequential Recommendation. In Proceedings of the International Joint Conference on Artificial Intelligence . 1–9

  22. [30]

    J Pearl. 2009. Causality. Cambridge university press

  23. [31]

    Malliaros, and Tommaso Di Noia

    Daniele Malitesta, Emanuele Rossi, Claudio Pomo, Fragkiskos D. Malliaros, and Tommaso Di Noia. 2024. Dealing with Missing Modalities in Multimodal Rec- ommendation: a Feature Propagation-based Approach. CoRR abs/2403.19841 (2024)

  24. [32]

    Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme

  25. [33]

    Yifang Qin, Hongjun Wu, Wei Ju, Xiao Luo, and Ming Zhang. 2024. A Diffusion Model for POI Recommendation. ACM Transactions on Information Systems 42, 2 (2024), 54:1–54:27

  26. [34]

    Yu Shang, Chen Gao, Jiansheng Chen, Depeng Jin, and Yong Li. 2024. Improving Item-side Fairness of Multimodal Recommendation via Modality Debiasing. In Proceedings of the ACM Web Conference . 4697–4705

  27. [35]

    Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising Diffusion Implicit Models. In Proceedings of the International Conference on Learning Repre- sentations

  28. [36]

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention, Vol. 9351. 234–241

  29. [37]

    Xuemeng Song, Chun Wang, Changchang Sun, Shanshan Feng, Min Zhou, and Liqiang Nie. 2023. MM-FRec: Multi-Modal Enhanced Fashion Item Recommen- dation. IEEE Transactions on Knowledge and Data Engineering 35, 10 (2023), 10072–10084

  30. [38]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All You Need. In Advances in Neural Information Processing Systems . 5998–6008

  31. [39]

    Wenzhuo Song, Shoujin Wang, Yan Wang, Kunpeng Liu, Xueyan Liu, and Ming- hao Yin. 2023. A Counterfactual Collaborative Session-based Recommender System. In Proceedings of the ACM Web Conference . 971–982

  32. [40]

    Meirui Wang, Pengjie Ren, Lei Mei, Zhumin Chen, Jun Ma, and Maarten de Rijke

  33. [41]

    Shoujin Wang, Ninghao Liu, Xiuzhen Zhang, Yan Wang, Francesco Ricci, and Bamshad Mobasher. 2022. Data Science and Artificial Intelligence for Responsible Recommendations. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 4904–4905

  34. [42]

    Cheng Wang, Mathias Niepert, and Hui Li. 2018. LRMM: Learning to Recommend with Missing Modalities. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . 3360–3370

  35. [43]

    Shoujin Wang, Yan Wang, Fikret Sivrikaya, Sahin Albayrak, and Vito Walter Anelli. 2023. Data science for next-generation recommender systems. Interna- tional Journal of Data Science and Analytics 16, 2 (2023), 135–145

  36. [44]

    Shoujin Wang, Xiuzhen Zhang, Yan Wang, and Francesco Ricci. 2024. Trustworthy Recommender Systems. ACM Transactions on Intelligent Systems and Technology 15, 4 (2024), 84:1–84:20

  37. [45]

    Wenjie Wang, Fuli Feng, Liqiang Nie, and Tat-Seng Chua. 2022. User-controllable Recommendation Against Filter Bubbles. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 1251– 1261

  38. [46]

    Shoujin Wang, Wentao Wang, Xiuzhen Zhang, Yan Wang, Huan Liu, and Fang Chen. 2024. A Hierarchical and Disentangling Interest Learning Framework for Unbiased and True News Recommendation. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3200–3211

  39. [47]

    Tianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu, Jinfeng Yi, and Xiangnan He

  40. [48]

    Wei Wei, Chao Huang, Lianghao Xia, and Chuxu Zhang. 2023. Multi-Modal Self-Supervised Learning for Recommendation. In Proceedings of the ACM Web Conference. 790–800

  41. [49]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2020. Graph-Refined Convolutional Network for Multimedia Recommendation with Im- plicit Feedback. In Proceedings of the ACM International Conference on Multimedia. 3541–3549

  42. [50]

    Yuanzhi Wang, Yong Li, and Zhen Cui. 2023. Incomplete Multimodality-Diffused Emotion Recognition. In Advances in Neural Information Processing Systems , Conference’17, July 2017, Washington, DC, USA Jin Li et al. Vol. 36. 17117–17128

  43. [51]

    Tong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng, Yeyun Gong, Yelong Shen, Jian Jiao, Juntao Li, Zhongyu Wei, Jian Guo, Nan Duan, and Weizhu Chen. 2023. AR- Diffusion: Auto-Regressive Diffusion Model for Text Generation. In Advances in Neural Information Processing Systems , Vol. ...

  44. [52]

    Zhangkai Wu, Xuhui Fan, Jin Li, Zhilin Zhao, Hui Chen, and Longbing Cao. 2024. ParamReL: Learning Parameter Space Representation via Progressively Encoding Bayesian Flow Networks. CoRR abs/2405.15268 (2024)

  45. [53]

    Zihao Wu, Xin Wang, Hong Chen, Kaidong Li, Yi Han, Lifeng Sun, and Wenwu Zhu. 2023. Diff4Rec: Sequential Recommendation with Curriculum-scheduled Diffusion Augmentation. In Proceedings of the ACM International Conference on Multimedia. 9329–9335

  46. [54]

    Lianghao Xia, Chao Huang, Chunzhen Huang, Kangyi Lin, Tao Yu, and Ben Kao

  47. [55]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal Graph Convolution Network for Personal- ized Recommendation of Micro-video. In Proceedings of the ACM International Conference on Multimedia. 1437–1445

  48. [56]

    Zhengyi Yang, Jiancan Wu, Zhicai Wang, Xiang Wang, Yancheng Yuan, and Xiangnan He. 2023. Generate What You Prefer: Reshaping Sequential Recom- mendation via Guided Diffusion. In Advances in Neural Information Processing Systems, Vol. 36. 24247–24261

  49. [57]

    Penghang Yu, Zhiyi Tan, Guanming Lu, and Bing-Kun Bao. 2023. LD4MRec: Simplifying and Powering Diffusion Model for Multimedia Recommendation. CoRR abs/2309.15363 (2023)

  50. [58]

    Yun-Hao Yuan, Jin Li, Yun Li, Jipeng Qiang, Yi Zhu, Xiaobo Shen, and Jianping Gou. 2022. Learning Canonical F-Correlation Projection for Compact Multiview Representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 19238–19247

  51. [59]

    Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang

  52. [60]

    Shanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen, Jinghui Qin, and Liang Lin. 2024. Mirror Gradient: Towards Robust Multimodal Recommender Sys- tems via Exploring Flat Local Minima. In Proceedings of the ACM Web Conference . 3700–3711

  53. [61]

    Yiyan Xu, Wenjie Wang, Fuli Feng, Yunshan Ma, Jizhi Zhang, and Xiangnan He. 2024. Diffusion Models for Generative Outfit Recommendation. In Proceed- ings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. 1350–1359

  54. [62]

    Xin Zhou. 2023. MMRec: Simplifying Multimodal Recommendation. In ACM Multimedia Asia Workshops. 6:1–6:2

  55. [63]

    Xin Zhou and Zhiqi Shen. 2023. A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal Recommendation. In Proceedings of the ACM International Conference on Multimedia . 935–943

  56. [64]

    Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, and Feijun Jiang. 2023. Bootstrap Latent Representations for Multi- modal Recommendation. In Proceedings of the ACM Web Conference . 845–854

  57. [65]

    Zhizhuo Zhou and Shubham Tulsiani. 2023. SparseFusion: Distilling View- Conditioned Diffusion for 3D Reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12588–12597. A ALGORITHM In this section, we summarize the proposed MoDiC...

  58. [66]

    In Proceedings of the ACM Multimedia Conference

    Mining Latent Structures for Multimedia Recommendation. In Proceedings of the ACM Multimedia Conference . 3872–3880

  59. [68]

    Hongyu Zhou, Xin Zhou, Zhiwei Zeng, Lingzi Zhang, and Zhiqi Shen. 2023. A Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions. CoRR abs/2302.04473 (2023)

  60. [73]

    We set the number of layers to 3 and the weight decay to 10−4

    Unimodal RS methods • LightGCN [11]: a classic unimodal RS method with simplified and efficient graph convolutional networks. We set the number of layers to 3 and the weight decay to 10−4. • AutoCF [54]: a self-supervised learning method incorporating adaptive data augmentatio...

  61. [74]

    We set the numbers of GCN layers for the user-item graph and the item-item graph, respectively, to 2 and 1

    MMRec methods: • FREEDOM [63]: an efficient design for MMRec by reducing the complexity of graph structure learning. We set the numbers of GCN layers for the user-item graph and the item-item graph, respectively, to 2 and 1. • MMSSL [48]: an advanced MMRec using adversarial au...

  62. [75]

    The learning rate is set to 0.0005 and the debiasing strength is set to 0.5

    as its backbone. The learning rate is set to 0.0005 and the debiasing strength is set to 0.5

  63. [76]

    The learning rate is set to 10−4

    Incomplete MMRec methods: • LRMM [39]: a classic method that uses autoencoders to recover incomplete multimodal data. The learning rate is set to 10−4. • CI2MG 1 [20]: a method that uses clustering-based imputation to recover missing features and performs cross-modal transport...

  64. [2009]

    InProceedings of the Conference on Uncertainty in Artificial Intelligence

    BPR: Bayesian Personalized Ranking from Implicit Feedback. InProceedings of the Conference on Uncertainty in Artificial Intelligence . 452–461

  65. [2019]

    In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval

    A Collaborative Session-based Recommendation Approach with Parallel Memory Modules. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 345–354

  66. [2021]

    In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining

    Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender System. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1791–1800

  67. [2023]

    In Proceedings of the ACM Web Conference

    Automated Self-Supervised Learning for Recommendation. In Proceedings of the ACM Web Conference. 992–1002

  68. [2024]

    CoRR abs/2406.11781 (2024)

    DiffMM: Multi-Modal Diffusion Model for Recommendation. CoRR abs/2406.11781 (2024)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.