Pith. sign in

REVIEW 4 major objections 4 minor 99 references

Generating Negative Samples for Multi-Modal Recommendation

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper proposes NegGen, which uses a multi-modal large language model to turn positive item descriptions into masked-and-completed counterfactual negatives, and argues that these generated negatives, combined with a causal contrast…

desk verdict A useful negative-sampling pipeline with an overstated causal claim; the false-negative issue needs an audit before the gains are fully trusted. read the letter →

arxiv 2501.15183 v3 pith:SDGFBPJO submitted 2025-01-25 cs.IR

classification cs.IR
keywords multi-modalrecommendationnegativesamplinglargelanguagemodelsattributemaskingcompletioncounterfactualnegativescausallearningBayesianpersonalizedranking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Multi-modal recommender systems often train on weak negatives: random uninteracted items or GAN-generated ones that do not exploit the image and text information attached to products. This paper argues that the right negative for a multi-modal recommender is a counterfactual item that resembles the positive in structure but differs in key attributes, and that such negatives can be produced by a multi-modal large language model with a generate-mask-complete prompt pipeline. The proposed NegGen framework first describes an item's image in text, masks its salient feature words, and asks the model to fill the masks with alternatives; it then encodes both original and generated attributes and trains a LightGCN base recommender with a BPR loss plus a causal contrast module. The paper reports that NegGen outperforms both state-of-the-art multi-modal recommenders and existing negative-sampling methods on four Amazon categories, and that the visual-plus-textual setting beats either modality alone. If correct, the payoff is a practical way to turn cheap MLLM generation into harder, more informative training signals for multimedia recommendation.

What carries the argument

The load-bearing mechanism is the three-prompt MLLM pipeline followed by a causal contrast module. Description Generation turns an item image into detailed text; Attribute Masking identifies and replaces core features, distinctive characteristics, and key specifications with [MASK] tokens; Attribute Completion fills each mask with plausible alternative words, producing a structurally parallel but semantically shifted description. The original and generated attribute strings are mapped by a frozen text encoder into embeddings $e_m$ and $e^*_m$, and a self-attention layer produces context-aware versions $\tilde{e}_m$ and $\tilde{e}^*_m$. The total causal effect of moving from positive to negative attributes is instantiated as their difference $e_t = \tilde{e}_m - \tilde{e}^*_m$, which is projected into the collaborative embedding space and used in the BPR loss $-\log\sigma(\lambda e_u^\top e_c + e_u^\top e_i - e_u^\top e^*_c)$; this difference is what the paper claims disentangles intervened key features from irrelevant attributes.

What would settle it

Ask a held-out set of users to choose between each positive item and its MLLM-generated negative description; if a large share of users prefer the generated variant for items they have interacted with, those training pairs are false negatives and the claimed supervision signal is corrupted. A cleaner quantitative check is to retrain NegGen with only generated negatives that pass a per-user preference filter and compare Recall and NDCG against the unfiltered run.

Watch

Extended reading notes

Core claim

The central discovery claimed is that a generate-mask-complete procedure over multi-modal item attributes yields negative samples that are simultaneously cohesive (semantically related to the positive) and hard (difficult for the recommender to distinguish), and that this improves top-K recommendation beyond what uniform, hard-negative, GAN-based, or diffusion-based samplers achieve. NegGen operationalizes this by using an MLLM to convert visual content into natural-language descriptions, masking key feature words, and completing the masks with alternative words; the resulting text attributes are encoded with a pretrained text encoder and fed through a self-attention module into a causal-effect embedding $e_t = \tilde{e}_m - \tilde{e}^*_m$ that is added to the collaborative score. The training objective combines a BPR ranking loss and a contrastive alignment loss, with the same generated negative embedding used for all users who interacted with the positive item. The paper supports the claim with experiments on Baby, Beauty, Clothing, and Sports, where NegGen reports consistent gains over second-best baselines in Recall and NDCG at K=10 and K=20.

Load-bearing premise

The load-bearing assumption is that, for every user who interacted with a positive item, the MLLM-generated altered description of that item is a genuine negative that the user does not want, even though the paper reports no per-user filtering or human evaluation of generated negatives.

Editorial extensions

If this is right

  • On all four datasets, NegGen's Recall and NDCG at K=10 and K=20 exceed every baseline, with relative gains over the best baseline ranging from roughly 2.3% to 9.0%.
  • The ablation that removes the negative generation module and falls back to uniform sampling drops performance, showing that the generated attributes carry a substantial part of the improvement.
  • The visual-plus-textual variant of NegGen beats both unimodal variants, indicating that the mask-complete pipeline reduces modality imbalance rather than simply adding text.
  • The paper reports that NegGen's negatives reach a given NDCG in fewer epochs than uniform sampling, IRGAN, DNS, DENS, AHNS, and MixGCF, which it reads as evidence that the negatives are both cohesive and hard.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: a per-user false-negative audit is the natural stress test: if a generated 'negative' (for example, the same shirt in a different color) is actually preferred by a user, that BPR pair is mislabeled, and adding a preference filter before training could either strengthen or bound NegGen's gains.
  • Editorial extension: the same generate-mask-complete recipe could be transplanted to other implicit-feedback problems with rich side information, such as sequential recommendation or product search, wherever the bottleneck is the informativeness of negatives rather than the model architecture.
  • Editorial extension: the paper claims the causal module disentangles key features from irrelevant attributes, but the implementation subtracts two learned embeddings; a direct test would intervene on a single attribute such as color, brand, or size and check whether the score changes only along that dimension.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes NegGen, a framework for negative sampling in multi-modal recommendation. NegGen uses a multi-modal large language model (MLLM) in three steps to generate counterfactual item descriptions (description generation, attribute masking, attribute completion), encodes them with a text encoder, and then combines a LightGCN base recommender with a 'causal learning' module that computes a difference between positive and negative attribute embeddings. The combined model is trained with a BPR-style loss and an alignment loss. Experiments on four Amazon subsets (Baby, Beauty, Clothing, Sports) report consistent improvements over several negative-sampling and multi-modal baselines, with paired t-tests.

Significance. The empirical results are potentially useful: the idea of using MLLMs to generate hard, modality-balanced negatives is timely, and the consistent gains across datasets (e.g., 9.02% R@10 on Baby) are non-trivial. The paper provides ablations and hyperparameter sensitivity, which support the reproducibility of the main pipeline. However, the claimed causal contribution is not supported by the evidence, and the validity of generated negatives as true negatives is not established. If these issues are addressed, the paper could be a solid contribution to multi-modal recommender systems.

major comments (4)
  1. [Section 3.2 and Eq. (19)] The generated negative e*_c is produced per item from the item's attributes, independent of the target user, and is then used in the BPR loss for every user u who interacted with i. This design assumes the counterfactual description is a valid negative for each such user, but the paper provides no per-user filtering, no human evaluation of generated negatives, and no analysis of the false-negative rate. Since the headline improvements in Table 2 rest on this negative-sampling module, the claim that NegGen generates high-quality negatives (RQ3) is not yet established; the ablation in Table 5 cannot separate generation quality from false-negative contamination.
  2. [Section 3.3, Eqs. (12)–(16)] The paper labels e_t = e_m_tilde - e*_m_tilde as an instantiation of the 'total causal effect' (TE). However, Eq. (12) is simply the definition of a difference between two potential outcome values under a deterministic SCM, and Eq. (16) is a difference of two learned embeddings. No identification step, no backdoor adjustment, and no do-calculus operation is performed; the SCM in Figure 5 is assumed without unobserved confounding. Therefore the claim that the module 'disentangle[s] the effect of intervened key features and irrelevant item attributes' is unsupported. The module may still be a useful contrastive feature, but it should be framed as such, not as causal inference.
  3. [Eq. (19)] The text states that e_c is 'the positive embedding of item i' and e*_c is 'the corresponding negative embedding', but Eqs. (17)–(18) define e_c as a projection of e_t = e_m_tilde - e*_m_tilde, i.e., a difference embedding, while e*_c is the projection of the negative attribute embedding. This discrepancy makes the training objective ambiguous: it is unclear whether the causal score is a difference term or an item-level positive score. The alignment loss in Eq. (20) suffers from the same ambiguity. The authors should correct the notation and clarify the exact forward pass.
  4. [Section 2.2, Eq. (1), Proposition 2] The lower bound on NDCG in Eq. (1) is stated as a sum of sigmoid terms over positive items and their corresponding negatives, but NDCG is a listwise ranking metric and the bound is not generally valid without additional assumptions about the remaining items and ranking positions. The citation to [32] does not supply the precise conditions. If the lemma is incorrect, the modality-imbalance analysis in Section 2.2 and the motivation for the generation pipeline are weakened. Please provide a proof or a precise citation, or soften the claims accordingly.
minor comments (4)
  1. [Section 2.1] The phrase 'precious study' should be 'previous study'.
  2. [Figure 6 caption] The caption contains 'Ours Ours' duplicated; only one 'Ours' is needed.
  3. [Table 4] In the Beauty row, V&T reports R@10 = 0.0832, but Table 2 lists NegGen Beauty R@10 = 0.0823; please reconcile this inconsistency.
  4. [Section 4.4] The text says 'we remove the the causal learning module' with a duplicated 'the'; also clarify what 'Ours' denotes in Figure 6 relative to the baselines.

Circularity Check

1 steps flagged · score 4.0 of 10

Main performance claim is evaluated on held-out interactions and is not circular, but the 'total causal effect' of the causal module is constructed as the difference between positive and negative embeddings, so the causal disentangling claim reduces to a relabeled contrastive subtraction.

  1. self definitional [Section 3.3, Eqs. (12) and (16)]
    "TE = E[Y|do(V =v,T =t)]− E[Y|do(V =v∗,T =t∗)] = fY(fM(v,t),u)− fY(fM(v∗,t∗),u). (12) ... Then we instantiate the total causal effect in Equation 12: e_t = ˜e_m− ˜e∗_m. (16)"

    The 'total causal effect' is defined in Eq. 12 as the difference between the predicted scores under positive attributes (v,t) and generated negative attributes (v*,t*). Eq. 16 'instantiates' this TE by subtracting the encoded negative attribute embedding from the encoded positive attribute embedding, followed only by a learned projection. No intervention, counterfactual model, or confounding control is introduced beyond this subtraction. Consequently, the claimed causal disentanglement of 'intervened key features and irrelevant item attributes' is not derived from any independent causal estimate; it is exactly the contrastive positive-minus-negative operation written in causal vocabulary, so the causal contribution is true by definition rather than by identification.

full rationale

The headline result—NegGen outperforming MMRS and negative sampling baselines—is a genuine empirical claim tested on held-out 80/10/10 interactions, so it is not circular. The ablation studies and hyperparameter analyses are also self-contained empirical evidence. No load-bearing self-citation was found; the cited prior work used for Lemma 1 and the causal survey [12] is external, not authored by the present authors. The main circular step is the causal learning module: the paper defines total causal effect as the score difference between positive and negative attribute settings (Eq. 12), then implements it as the embedding difference (Eq. 16) and treats this as causal disentanglement. This is a relabeling of a contrastive representation, so the causal contribution reduces by construction. The reviewer-flagged false-negative issue (per-item generated negatives applied to all interacting users without per-user filtering) is a validity and data-quality concern, not a circularity, and is therefore not scored here.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard LightGCN/BPR machinery, on unvalidated MLLM generation quality, on a per-item negative that is assumed valid for all users, and on a simplified causal graph. Three hyperparameters are tuned on validation data, but their final values are not reported.

free parameters (3)
  • lambda (causal score weight) = 0.4 to 0.6, exact value not reported
    Controls the contribution of the multi-modal causal score to the final prediction; tuned on validation in Section 4.6.
  • alpha (alignment loss weight) = range 1e-4 to 1, exact value not reported
    Balances the recommendation loss and the multi-modal alignment loss; tuned on validation in Section 4.6.
  • tau (alignment temperature) = 0.07 to 0.3, exact value not reported
    Temperature in the alignment objective; smaller values were found to work better in Section 4.6.
assumptions (4)
  • standard math LightGCN with averaged layer embeddings and BPR loss is an adequate base recommender
    Section 3.1 adopts LightGCN in Eq 6-7 and BPR in Eq 2 without justification; this is standard in prior work.
  • domain assumption The MLLM's generated descriptions faithfully capture visual content and coherent alternative attributes
    Section 3.2.1 to 3.2.3 relies on Llama 3.2-11B-Vision outputs as reliable item attributes and mask completions; no quality check is reported.
  • domain assumption A generated negative description is a valid negative for every user who interacted with the source item
    Equation 19 applies the same generated negative embedding e*_c in the BPR loss for all users; false negatives are not filtered per user.
  • ad hoc to paper The SCM in Figure 5 has no unobserved confounding and do-calculus applies to the V,T to M to Y structure
    Equation 12 claims a total causal effect without controlling for confounders such as user preference influencing exposure to item attributes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generating Negative Samples for Multi-Modal Recommendation." pith.science (2026). https://pith.science/paper/SDGFBPJO

@misc{pith2026250115183,
  author       = {Pith},
  title        = {Pith review of: Generating Negative Samples for Multi-Modal Recommendation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SDGFBPJO}},
  note         = {Machine review of arXiv:2501.15183}
}
read the original abstract

Multi-modal recommender systems (MMRS) have gained significant attention due to their ability to leverage information from various modalities to enhance recommendation quality. However, existing negative sampling techniques often struggle to effectively utilize the multi-modal data, leading to suboptimal performance. In this paper, we identify two key challenges in negative sampling for MMRS: (1) producing cohesive negative samples contrasting with positive samples and (2) maintaining a balanced influence across different modalities. To address these challenges, we propose NegGen, a novel framework that utilizes multi-modal large language models (MLLMs) to generate balanced and contrastive negative samples. We design three different prompt templates to enable NegGen to analyze and manipulate item attributes across multiple modalities, and then generate negative samples that introduce better supervision signals and ensure modality balance. Furthermore, NegGen employs a causal learning module to disentangle the effect of intervened key features and irrelevant item attributes, enabling fine-grained learning of user preferences. Extensive experiments on real-world datasets demonstrate the superior performance of NegGen compared to state-of-the-art methods in both negative sampling and multi-modal recommendation.

Figures

Figures reproduced from arXiv: 2501.15183 by the authors.

Figure 1
Figure 1. Illustration of effective negative sampling in MMRS. “Co [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 3
Figure 3. Gradient magnitudes of negative sample modalities over [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. The overall architecture of NegGen comprises three main components: (1) Base Recommender Training, which captures collaborative filtering signals from user-item interactions, (2) Negative Sample Generation, which produces contrasting item attributes via MLLM, and (3) Causal Learning, which models causal relationships between multi-modal item characteristics and recommendation outcomes. item 𝑖 are derived by averagin… view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Causal graph illustrating the relationships between multi [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Visualization of the convergence epoch and NDCG@20 [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Effect of different hyper-parameters on NegGen. 5 Related Work Hard Negative Sampling in RS. Negative sampling is widely used across fields like computer vision [28, 74, 75], natural language pro￾cessing [8, 59, 77], and information retrieval [44, 48, 49]. In RS, hard …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

99 extracted references · 66 canonical work pages

  1. [32]

    Riwei Lai, Rui Chen, Qilong Han, Chi Zhang, and Li Chen. 2025. Adaptive hardness negative sampling for collaborative filtering (AAAI’24). Article 961, 8 pages

  2. [1]

    Haoyue Bai, Le Wu, Min Hou, Miaomiao Cai, Zhuangzhuang He, Yuyang Zhou, Richang Hong, and Meng Wang. 2024. Multimodality Invariant Learning for Multimedia-Based New Item Recommendation (SIGIR ’24). 677–686

  3. [2]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Jun- yang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision- Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv:2308.12966 [cs.CV]

  4. [3]

    Jiawei Chen, Chengquan Jiang, Can Wang, Sheng Zhou, Yan Feng, Chun Chen, Martin Ester, and Xiangnan He. 2021. CoSam: An Efficient Collaborative Adaptive Sampler for Recommendation. ACM Trans. Inf. Syst. 39, 3 (2021), 34:1–34:24

  5. [4]

    Jiawei Chen, Can Wang, Sheng Zhou, Qihao Shi, Yan Feng, and Chun Chen. 2019. SamWalker: Social Recommendation with Informative Sampling Strategy(WWW ’19). 228–239

  6. [5]

    Jingyuan Chen, Hanwang Zhang, Xiangnan He, Liqiang Nie, Wei Liu, and Tat- Seng Chua. 2017. Attentive Collaborative Filtering: Multimedia Recommendation with Item- and Component-Level Attention (SIGIR ’17). 335–344

  7. [6]

    Ádám Tibor Czapp, Mátyás Jani, Bálint Domián, and Balázs Hidasi. 2024. Dynamic Product Image Generation and Recommendation at Scale for Personalized E- commerce (RecSys ’24). 768–770

  8. [7]

    Jianfeng Deng, Qingfeng Chen, Debo Cheng, Jiuyong Li, Lin Liu, and Xiaojing Du. 2024. Mitigating Dual Latent Confounding Biases in Recommender Systems. arXiv:2410.12451 [cs.IR]

Show all 99 references
  1. [8]

    Jacob Devasier, Yogesh Gurjar, and Chengkai Li. 2024. Robust Frame-Semantic Models with Lexical Unit Trees and Negative Samples (ACL ’24)

  2. [9]

    Jingtao Ding, Yuhan Quan, Xiangnan He, Yong Li, and Depeng Jin. 2019. Re- inforced negative sampling for recommendation with exposure data (IJCAI’19). 99–109

  3. [10]

    Xiaoyu Du, Zike Wu, Fuli Feng, Xiangnan He, and Jinhui Tang. 2022. Invariant Representation Learning for Multimedia Recommendation (MM ’22). 619–628

  4. [11]

    Lu Fan, Jiashu Pu, Rongsheng Zhang, and Xiao-Ming Wu. 2023. Neighborhood- based Hard Negative Mining for Sequential Recommendation (SIGIR’23). 2042–2046

  5. [12]

    Chen Gao, Yu Zheng, Wenjie Wang, Fuli Feng, Xiangnan He, and Yong Li. 2024. Causal Inference in Recommender Systems: A Survey and Future Directions. ACM Trans. Inf. Syst. 42, 4, Article 88 (Feb. 2024), 32 pages

  6. [13]

    Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research). PMLR

  7. [14]

    Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

    Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative Adversarial Networks. arXiv:1406.2661 [stat.ML]

  8. [15]

    Jie Gui, Zhenan Sun, Yonggang Wen, Dacheng Tao, and Jieping Ye. 2023. A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications.IEEE Trans. on Knowl. and Data Eng. 35, 4 (April 2023), 3313–3332

  9. [16]

    Zhiqiang Guo, Jianjun Li, Guohui Li, Chaoyang Wang, Si Shi, and Bin Ruan. 2024. LGMRec: Local and Global Graph Learning for Multimodal Recommendation (AAAI ’24). 8454–8462

  10. [17]

    Ruining He and Julian McAuley. 2016. VBPR: visual Bayesian Personalized Ranking from implicit feedback (AAAI’16). 144–150

  11. [18]

    Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, YongDong Zhang, and Meng Wang. 2020. LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation (SIGIR ’20). 639–648

  12. [19]

    Xiangnan He, Yang Zhang, Fuli Feng, Chonggang Song, Lingling Yi, Guohui Ling, and Yongdong Zhang. 2022. Addressing Confounding Feature Issue for Causal Recommendation. ACM Trans. Inf. Syst. (2022)

  13. [20]

    Tinglin Huang, Yuxiao Dong, Ming Ding, Zhen Yang, Wenzheng Feng, Xinyu Wang, and Jie Tang. 2021. MixGCF: An Improved Training Method for Graph Neural Network-based Recommender Systems (KDD ’21)

  14. [21]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. 20, 4 (Oct. 2002), 422–446

  15. [22]

    Deyi Ji, Feng Zhao, Hongtao Lu, Feng Wu, and Jieping Ye. 2025. Structural and Statistical Texture Knowledge Distillation and Learning for Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 47, 5 (Jan. 2025), 3639–3656. https://doi.org/10. 1109/TPAMI.2025.3536481

  16. [23]

    Deyi Ji, Feng Zhao, Lanyun Zhu, Wenwei Jin, Hongtao Lu, and Jieping Ye

  17. [24]

    Deyi Ji, Lanyun Zhu, Siqi Gao, Peng Xu, Hongtao Lu, Jieping Ye, and Feng Zhao

  18. [25]

    Yanbiao Ji, Yue Ding, Chang Liu, Yuxiang Lu, Xin Xin, and Hongtao Lu

  19. [26]

    arXiv:2411.08516 [cs.CL] https://arxiv.org/abs/2411.08516

    Tree-of-Table: Unleashing the Power of LLMs for Enhanced Large-Scale Table Understanding. arXiv:2411.08516 [cs.CL] https://arxiv.org/abs/2411.08516

  20. [27]

    Binbin Jin, Defu Lian, Zheng Liu, Qi Liu, Jianhui Ma, Xing Xie, and Enhong Chen

  21. [28]

    arXiv:2411.13892 [cs.IR]

    Topology-Aware Popularity Debiasing via Simplicial Complexes. arXiv:2411.13892 [cs.IR]

  22. [29]

    Yangqin Jiang, Lianghao Xia, Wei Wei, Da Luo, Kangyi Lin, and Chao Huang

  23. [30]

    7591–7599

    DiffMM: Multi-Modal Diffusion Model for Recommendation (MM ’24). 7591–7599

  24. [31]

    Riwei Lai, Li Chen, Yuhan Zhao, Rui Chen, and Qilong Han. 2023. Disentangled Negative Sampling for Collaborative Filtering (WSDM ’23). 96–104

  25. [33]

    Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus. 2020. Hard Negative Mixing for Contrastive Learning (NeurIPS ’20). Article 1829, 12 pages

  26. [34]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization (ICLR ’15)

  27. [35]

    Beinke, Gerrit Y

    Thorsten Krause, Alina Deriyeva, Jan H. Beinke, Gerrit Y. Bartels, and Oliver Thomas. 2024. Mitigating Exposure Bias in Recommender Systems—A Compara- tive Analysis of Discrete Choice Models. ACM Trans. Recomm. Syst. 3, 2, Article 19 (Nov. 2024), 37 pages

  28. [36]

    Yang Li, Qi’Ao Zhao, Chen Lin, Jinsong Su, and Zhilin Zhang. 2024. Who To Align With: Feedback-Oriented Multi-Modal Alignment in Recommendation Systems (SIGIR ’24). 667–676

  29. [37]

    Qiang Liu, Shu Wu, and Liang Wang. 2017. DeepStyle: Learning User Preferences for Visual Recommendation (SIGIR ’17). 841–844

  30. [38]

    Hongkang Li, Meng Wang, Tengfei Ma, Sijia Liu, ZAIXI ZHANG, and Pin-Yu Chen

  31. [39]

    What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding (ICML’24)

  32. [40]

    Shuaiyang Li, Dan Guo, Kang Liu, Richang Hong, and Feng Xue. 2023. Multimodal Counterfactual Learning Network for Multimedia-based Recommendation(SIGIR ’23). 1539–1548

  33. [41]

    Xingchen Li, Xiang Wang, Xiangnan He, Long Chen, Jun Xiao, and Tat-Seng Chua. 2020. Hierarchical Fashion Graph Network for Personalized Outfit Recom- mendation. arXiv:2005.12566 [cs.IR]

  34. [42]

    Masoud Mansoury, Bamshad Mobasher, and Herke van Hoof. 2024. Mitigating Exposure Bias in Online Learning to Rank Recommendation: A Novel Reward Model for Cascading Bandits (CIKM’24). 1638–1648

  35. [43]

    Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel

  36. [44]

    Yiyu Liu, Qian Liu, Yu Tian, Changping Wang, Yanan Niu, Yang Song, and Chenliang Li. 2021. Concept-Aware Denoising Graph Neural Network for Micro- Video Recommendation (CIKM ’21). 1099–1108

  37. [45]

    Haokai Ma, Ruobing Xie, Lei Meng, Xin Chen, Xu Zhang, Leyu Lin, and Jie Zhou

  38. [46]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...

  39. [47]

    Haokai Ma, Ruobing Xie, Lei Meng, Fuli Feng, Xiaoyu Du, Xingwu Sun, Zhanhui Kang, and Xiangxu Meng. 2024. Negative Sampling in Recommendation: A Survey and Future Directions. arXiv:2409.07237 [cs.IR]

  40. [48]

    Haokai Ma, Yimeng Yang, Lei Meng, Ruobing Xie, and Xiangxu Meng. 2024. Multimodal Conditioned Diffusion Model for Recommendation (WWW ’24) . 1733–1740

  41. [49]

    Thilina Chaturanga Rajapakse, Andrew Yates, and Maarten de Rijke. 2024. Nega- tive Sampling Techniques for Dense Passage Retrieval in a Multilingual Setting (SIGIR ’24). 575–584

  42. [50]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (EMNLP ’19). 3982–3992. MM ’25, October 27–31, 2025, Dublin, Ireland Yanbiao Ji et al

  43. [51]

    Steffen Rendle and Christoph Freudenthaler. 2014. Improving pairwise learning for item recommendation from implicit feedback (WSDM ’14). 273–282

  44. [52]

    Jack McKechnie, Graham McDonald, and Craig Macdonald. 2024. Bi-Objective Negative Sampling for Sensitivity-Aware Search (SIGIR ’24). 2296–2300

  45. [53]

    Trung-Kien Nguyen and Yuan Fang. 2024. Diffusion-based Negative Sampling on Graphs for Link Prediction (WWW ’24). 948–958

  46. [54]

    Zihua Si, Xueran Han, Xiao Zhang, Jun Xu, Yue Yin, Yang Song, and Ji-Rong Wen. 2022. A Model-Agnostic Causal Learning Framework for Recommendation using Search Data (WWW ’22). 224–233

  47. [55]

    Judea Pearl. 2009. Causal inference in statistics: An overview. Statistics Surveys 3, none (2009), 96 – 146

  48. [56]

    Gustavo Penha and Claudia Hauff. 2023. Do the Findings of Document and Passage Retrieval Generalize to the Retrieval of Responses for Dialogues? (ECIR ’23). 132–147

  49. [57]

    Jun Wang, Lantao Yu, Weinan Zhang, Yu Gong, Yinghui Xu, Benyou Wang, Peng Zhang, and Dell Zhang. 2017. IRGAN: A Minimax Game for Unifying Generative and Discriminative Information Retrieval Models (SIGIR ’17). 515–524

  50. [58]

    Qinyong Wang, Hongzhi Yin, Zhiting Hu, Defu Lian, Hao Wang, and Zi Huang

  51. [59]

    Tianqi Wang, Lei Chen, Xiaodan Zhu, Younghun Lee, and Jing Gao. 2023. Weighted Contrastive Learning With False Negative Control to Help Long-tailed Product Classification (ACL ’23). 6930–6941

  52. [60]

    Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme

  53. [61]

    Xiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao, Meng Wang, and Tat-Seng Chua. 2020. Reinforced Negative Sampling over Knowledge Graph for Recom- mendation (WWW ’20)

  54. [62]

    Wentao Shi, Jiawei Chen, Fuli Feng, Jizhi Zhang, Junkang Wu, Chongming Gao, and Xiangnan He. 2023. On the Theories Behind Hard Negative Sampling for Recommendation (WWW ’23). 812–822

  55. [63]

    Tianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu, Jinfeng Yi, and Xiangnan He

  56. [64]

    Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang, Lifang Yang, Xianglin Huang, and Tat-Seng Chua. 2023. Self-Supervised Learning for Multimedia Recommen- dation. IEEE Transactions on Multimedia 25 (2023), 5107–5116

  57. [65]

    Feng Wang and Huaping Liu. 2021. Understanding the Behaviour of Contrastive Loss (CVPR ’21). 2495–2504

  58. [66]

    Mike Wu, Milan Mosse, Chengxu Zhuang, Daniel Yamins, and Noah Goodman

  59. [67]

    Songli Wu, Liang Du, Jia-Qi Yang, Yuai Wang, De-Chuan Zhan, SHUANG ZHAO, and Zixun Sun. 2024. RE-SORT: Removing Spurious Correlation in Multilevel Interaction for CTR Prediction. In The 40th Conference on Uncertainty in Artificial Intelligence (UAI’24). Article 178, 13 pages

  60. [68]

    Shuyuan Xu, Da Xu, Evren Korpeoglu, Sushant Kumar, Stephen Guo, Kannan Achan, and Yongfeng Zhang. 2024. Causal Structure Learning for Recommender System. ACM Trans. Recomm. Syst. 3, 1, Article 8 (Oct. 2024), 23 pages

  61. [69]

    Jiahao Xun, Shengyu Zhang, Zhou Zhao, Jieming Zhu, Qi Zhang, Jingjie Li, Xiuqiang He, Xiaofei He, Tat-Seng Chua, and Fei Wu. 2021. Why Do We Click: Visual Impression-aware News Recommendation (MM ’21). 3881–3890

  62. [70]

    Wenjie Wang, Fuli Feng, Xiangnan He, Xiang Wang, and Tat-Seng Chua. 2021. Deconfounded Recommendation for Alleviating Bias Amplification (KDD ’21). 1717–1725

  63. [71]

    Yuan Yao, Tianyu Yu, Ao Zhang, and et al. 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone. arXiv:2408.01800 [cs.CV]

  64. [73]

    Junliang Yu, Hongzhi Yin, Xin Xia, Tong Chen, Jundong Li, and Zi Huang. 2024. Self-Supervised Learning for Recommender Systems: A Survey.IEEE Transactions on Knowledge and Data Engineering 36, 1 (2024), 335–355

  65. [74]

    Le Zhang, Rabiul Awal, and Aishwarya Agrawal. 2023. Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Fine- grained Understanding. arXiv preprint arXiv:2306.08832 (2023)

  66. [75]

    1791–1800

    Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender System (KDD’21). 1791–1800

  67. [76]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2020. Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit Feedback (MM ’20). 3541–3549

  68. [77]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal Graph Convolution Network for Personalized Recommendation of Micro-video (MM ’19). 1437–1445

  69. [78]

    Tong Zhao, Julian McAuley, and Irwin King. 2014. Leveraging Social Connections to Improve Personalized Ranking for Collaborative Filtering(CIKM ’14). 261–270

  70. [79]

    Conditional Negative Sampling for Contrastive Learning of Visual Repre- sentations (ICLR’21)

  71. [80]

    Hongyu Zhou, Xin Zhou, and Zhiqi Shen. 2023. Enhancing Dyadic Relations with Homogeneous Graphs for Multimodal Recommendation. ArXiv abs/2301.12097 (2023)

  72. [81]

    Xin Zhou and Zhiqi Shen. 2023. A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal Recommendation (MM ’23). 935–943

  73. [82]

    Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, and Feijun Jiang. 2023. Bootstrap Latent Representations for Multi- modal Recommendation (WWW ’23). 845–854

  74. [83]

    Zhen Yang, Ming Ding, Tinglin Huang, Yukuo Cen, Junshuai Song, Bin Xu, Yuxiao Dong, and Jie Tang. 2024. Does Negative Sampling Matter? a Review With Insights Into its Theory and Applications. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 8 (2024), 5692–5711

  75. [84]

    Lanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu, De Wen Soh, and Jun Liu. 2025. CPCF: A Cross-Prompt Contrastive Framework for Referring Multimodal Large Language Models. In Forty-second International Conference on Machine Learning . https://openreview.net/forum?id=0MpGi6IwZr

  76. [85]

    Hamilton, and Jure Leskovec

    Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L. Hamilton, and Jure Leskovec. 2018. Graph Convolutional Neural Networks for Web-Scale Recommender Systems (KDD ’18). 974–983

  77. [86]

    Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, and Ken Zheng. 2024. Balancing Multimodal Learning via Online Logit Modulation. 5753–5761

  78. [88]

    Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z. Li. 2020. Bridg- ing the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection (CVPR ’20). 9756–9765

  79. [89]

    Weinan Zhang, Tianqi Chen, Jun Wang, and Yong Yu. 2013. Optimizing top-n collaborative filtering via dynamic negative item sampling (SIGIR ’13). 785–788

  80. [90]

    Yanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu, and Yongyi Mao

  81. [93]

    Yu Zheng, Chen Gao, Jianxin Chang, Yanan Niu, Yang Song, Depeng Jin, and Yong Li. 2022. Disentangling Long and Short-Term Interests for Recommendation (WWW ’22). 2256–2267

  82. [97]

    Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, and Feijun Jiang. 2023. Bootstrap Latent Representations for Multi- Modal Recommendation (WWW ’23). 845–854

  83. [99]

    Ziwei Zhu, Yun He, Xing Zhao, and James Caverlee. 2021. Popularity Bias in Dynamic Recommendation (KDD ’21). 2439–2449

  84. [2009]

    BPR: Bayesian personalized ranking from implicit feedback (UAI ’09). 452–461

  85. [2015]

    Image-Based Recommendations on Styles and Substitutes (SIGIR ’15) . 43–52

  86. [2018]

    2467–2475

    Neural Memory Streaming Recommender Networks with Adversarial Training (KDD ’18). 2467–2475

  87. [2020]

    Article 1897, 11 pages

    Sampling-decomposable generative adversarial recommender (NIPS ’20). Article 1897, 11 pages

  88. [2021]

    1791–1800

    Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender System (KDD ’21). 1791–1800

  89. [2022]

    Proceedings of the AAAI Conference on Artificial Intelligence 36, 10 (Jun

    Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives. Proceedings of the AAAI Conference on Artificial Intelligence 36, 10 (Jun. 2022), 11730–11738

  90. [2023]

    Exploring False Hard Negative Sample in Cross-Domain Recommendation (RecSys’23). 502–514

  91. [2024]

    arXiv:2406.10475 [cs.CV] https://arxiv.org/abs/2406.10475

    Discrete Latent Perspective Learning for Segmentation and Detection. arXiv:2406.10475 [cs.CV] https://arxiv.org/abs/2406.10475

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.