REVIEW 4 major objections 6 minor 64 references
FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that a Stein-kernel Integrated Information Coordination Module provably aligns image and text features with ID interaction streams, and that the resulting FindRec framework beats state-of-the-art multimodal sequential…
desk verdict A plausible multimodal sequential recommender whose advertised theoretical guarantee is contradicted by its own equations, and whose implementation section is duplicated with conflicting settings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Integrated Information Coordination Module (IICM), built on an RBF-Stein kernel $K(z_{\text{img}}, z_{\text{text}}) = \exp\left(-\tfrac{\|z_{\text{img}} - z_{\text{text}}\|^2}{2\sigma^2}\right)$ with an adaptively estimated bandwidth $\sigma$, whose batch expectation becomes the alignment loss. The Stein gradient estimator $\nabla_z \log q(z)$ is the second half of the mechanism: it supplies an unbiased, score-based estimate of the entropy gradient used to maximize the differential entropy of the embeddings, which the paper argues prevents representation collapse and shortcut learning. Around this sits a cross-modal expert router: the aligned features are split into $H$ low-dimensional heads, cross-attention is computed per head, and a lightweight gating network routes each head's output through $N_e$ specialized experts before concatenation and final fusion with the Mamba-processed ID stream.
What would settle it
Run FindRec on Amazon Beauty with the IICM alignment loss replaced by a plain squared-error loss between the same two embedding vectors, keeping the router, Mamba backbone, and tuning budget identical; if NDCG@5 does not drop significantly, the Stein kernel and entropy regularization are not carrying the claimed alignment. A more direct check: on held-out features, measure a distribution discrepancy such as maximum mean discrepancy between image and text embeddings with and without IICM training, and compare the score estimate $\nabla_z \log q(z)$ against a Monte Carlo kernel density estimate; if the discrepancy does not shrink or the score estimates diverge, the claimed theoretical guarantee is not what the experiments show.
Extended reading notes
Core claim
The paper's central discovery, on its own terms, is that an RBF-Stein kernel similarity between the final image and text embeddings, averaged over the batch as an alignment loss, plus an unbiased Stein score-based entropy gradient, can serve as a theoretically grounded 'gatekeeper' for cross-modal information in sequential recommendation. FindRec's IICM simultaneously minimizes the distributional discrepancy between modality pairs and regularizes the latent distributions via differential entropy maximization, which the authors argue prevents over-alignment and preserves modality-specific information. Combined with cross-attention over H subspace heads and a router mixture-of-experts that filters noisy or irrelevant signals, the aligned features are fused with Mamba-processed ID sequences; the paper reports NDCG@5 of 0.0438 on MicroLens, 0.2325 on MovieLens-100K, and 0.0843 on Amazon Beauty, with relative gains of 1.04–3.34% over the best baseline. Ablations on Amazon Beauty show that removing the IICM costs 5.7% NDCG@5 and removing the expert routing costs 6.9% NDCG@5, so both components are load-bearing for the reported result.
Load-bearing premise
The paper assumes, citing earlier Stein-method work, that the Stein gradient estimator is unbiased for the deep-network features used here and that the adaptive bandwidth $\sigma$ is estimated reliably; if that score approximation is biased or $\sigma$ is poorly chosen, the claimed unbiased alignment and entropy maximization do not follow.
Editorial extensions
If this is right
- If the IICM guarantee holds, multimodal sequential recommenders can align heterogeneous features to ID streams without negative sampling, large batches, or contrastive heads, removing a common source of training instability.
- The adaptive expert router gives a concrete mechanism for noisy and unequal modalities: contextually irrelevant features are down-weighted, so fusion degrades gracefully on short or sparse interaction histories.
- Because Mamba layers provide linear-complexity temporal modeling, the full flow scales to much longer interaction sequences than transformer-based multimodal recommenders at comparable quality.
- The reported gains are largest where behavioral evidence is weakest (5–10 interactions), suggesting the alignment substitutes for missing sequential signal, not merely adds capacity.
Reading between the lines
- The IICM is not specific to image–text pairs; a natural extension is applying the same Stein-kernel alignment to user-side and item-side towers or to cross-domain embeddings, where a provable alignment could stabilize cold-start recommendations that lack ID streams.
- If entropy maximization is doing what the paper claims, FindRec's embeddings should resist modality collapse better than contrastively aligned representations, which is testable by measuring the rank of the embedding covariance matrix or the uniformity of nearest-neighbor distances on held-out features.
- The paper's own limitation section notes that temporal periodicity is ignored; feeding Fourier features of timestamps or holiday calendars into the Mamba layers would let the alignment and routing adapt to seasonal modality importance, for example images mattering more during fashion shopping seasons.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FindRec, a multimodal sequential recommender that combines Mamba-based temporal encoding, an Integrated Information Coordination Module (IICM) intended to align image and text features with ID-based behavior streams, and a cross-modal expert routing mechanism with multi-head cross-attention. The central claims are that IICM 'theoretically guarantees distribution consistency' and provides 'unbiased, score-based estimation of the entropy gradients', and that FindRec outperforms state-of-the-art baselines on MicroLens, MovieLens-100K, and Amazon Beauty. The experiments report consistent gains, with relative NDCG@5 improvements of 1.40-3.34% over the best baseline, plus ablations, sequence-length analysis, and hyperparameter studies.
Significance. If the theoretical claims were substantiated, the paper would offer a novel principled alignment mechanism for multimodal sequential recommendation, and the empirical results would be a useful system-level contribution. The manuscript has several strengths: a broad baseline comparison, multiple datasets, significance testing over ten seeds, ablations of all main components, and additional hyperparameter analyses in the appendix. However, the central theoretical guarantee is not established by any derivation or formal statement, and the implementation section contains two contradictory descriptions of the experimental configuration. These issues must be resolved before the contribution can be evaluated reliably.
major comments (4)
- [Abstract, Sec. 1, Sec. 2.2, Eqs. (1)-(2)] The central claim that IICM 'theoretically guarantees distribution consistency' and provides 'unbiased, score-based estimation of the entropy gradients' is not supported by the paper's equations. The only IICM objective, Eq. (2), is E[K(z_img,z_text)] with the RBF kernel in Eq. (1); its gradient is a deterministic pairwise kernel-alignment term and contains no score function, no Stein operator, and no entropy term. No theorem, lemma, or formal statement connects this objective to Stein gradient estimation or to differential entropy maximization, and the 'adaptive sigma updated via the Stein kernel mechanism' is never defined. The opening of Sec. 2.2 also mentions KL divergence regularization, but no KL term appears in Eq. (5) or elsewhere. The statement in Sec. 1 that Mamba layers 'maintain theoretical guarantees for distribution alignment' is likewise unexplained. Please either provide a derivation or proof substantiating the Stein and entropy claims, or remove these claims from the abstract, contributions, Sec. 2.2, and Sec. 5.
- [Sec. 3.3] There are two mutually contradictory copies of the implementation details. The first specifies learning rate 0.002, batch size 256, ID embedding dimensions 128 or 64, hidden dimension 512, and maximum sequence lengths 50 or 100; the second specifies learning rate 0.001, weight decay 0.01, batch size 128, 1024-dimensional BLIP/RoBERTa features projected to 256 dimensions, ID dimension 128, maximum length 50, two Mamba layers with state dimension 16 and expansion factor 2, 100 epochs, cosine scheduling, warmup, and gradient clipping. It is impossible to tell which configuration produced Table 2 and the ablations. Please merge these into one unambiguous implementation description.
- [Sec. 3.1, Tables 1 and 4] The Amazon Beauty statistics are inconsistent across tables: Table 1 reports 4,323 users, 2,424 items, 60,276 interactions, and 99.42% sparsity, while Table 4 reports 3,914 users, 2,344 items, 60,583 interactions, and 99.34% sparsity. If Table 4 is a filtered subset for the sequence-length analysis, the filtering rule must be stated; as written, the discrepancy calls into question the validity of the RQ2 results. In addition, Table 5 does not say which dataset the ablation study uses; the numbers indicate Amazon Beauty, but this should be stated explicitly.
- [Sec. 3.5-3.6] The ablation study in Table 5 appears to be run on a single dataset, yet Sec. 3.6 draws general conclusions such as 'The Mixture of Experts component shows the most crucial impact' and claims validation of 'theoretical insights' about entropy maximization. Without ablations on at least one more dataset, or an explicit statement that the conclusions are dataset-specific, the claims are overgeneralized. Also, the text in Sec. 3.5 says '>30 interactions: 0.1327 vs. 0.1077', but Table 3 reports 0.1328 for Ours; correct the number.
minor comments (6)
- [Sec. 3.1] The phrase 'The interaction density is 99.96%' for Micro-lens is inconsistent with Table 1, which lists sparsity as 99.96%; presumably this should read 'sparsity'.
- [Abstract, Sec. 5] The paradigm name is inconsistent: the abstract says 'information flow-control-output', while the conclusion says 'flow-control-align-output'; unify the terminology.
- [Sec. 3.3] The sentence 'The remaining implementation details align with the configurations from the original papers' appears twice in the first paragraph, and the relation between the two repeated paragraphs is unclear.
- [Sec. 2.3] The citation '[30? ]' is malformed; please replace it with the intended reference or remove the placeholder.
- [Appendix A, Table 8] The header 'Layers' should be 'Dimension', because the table reports alignment-dimension values, not layer counts.
- [Sec. 3.4] The RQ1 discussion calls the model 'MMMamba4Rec', whereas Table 2 labels it 'Mamba4Rec (MM)'; use a single name throughout.
Circularity Check
The IICM 'theoretical guarantee' is the name of its own RBF alignment loss; the entropy-gradient unbiasedness claim is imported from same-group prior work.
-
self definitional
[Sec. 2.2, Eqs. (1)-(2); Sec. 2.4 Eq. (5)]
"The average similarity across the batch is then defined as the alignment loss L_IICM = E[K(z_img,last,z_txt,last)], (2), which encourages the representations from both modalities to converge toward a coherent, unified space while preserving their unique discriminative features. ... A key design highlight of our IICM is its dynamic control over the information flow. The Stein kernel not only facilitates the computation of K(z_img,last,z_txt,last) but also provides an unbiased, score-based estimation of the entropy gradients with respect to the model parameters."
The only IICM term appearing in the optimized objective, L = L_rec + lambda*L_IICM in Eq. (5), is the RBF-kernel expectation in Eq. (2). No differential-entropy term, no score function, and no Stein operator is defined or optimized anywhere in the paper. Therefore the promised 'unbiased distribution alignment' and 'entropy maximization' are not consequences derived from the model; they are simply the properties attributed to the exact loss being maximized. If 'distribution consistency' means high RBF similarity, the guarantee is true by construction of the objective, and if it means anything stronger, the paper supplies no independent quantity, theorem, or experiment to establish it.
-
self citation load bearing
[Sec. 2.2 (IICM); Refs. [28,46]; Acknowledgments]
"In essence, the Stein gradient estimator [28, 46] taps into the local geometry of the feature distributions by approximating the score function ∇z log q(z). This approximation is crucial for maximally increasing the differential entropy of the embeddings, thus promoting global uniformity across latent space."
The load-bearing claim that FindRec obtains an unbiased, score-based entropy-gradient estimate is not derived in this paper; it is delegated to Refs. [28] and [46]. The relevant entropy-maximizing estimator is MVEB [46], prior work by the same research group (coauthor Zenglin Xu; the author of [46], Liangjian Wen, is also thanked in the acknowledgments), and Refs. [45] and [47] are the same group's related papers. No implementation-level equation connecting Eq. (2) to a Stein score estimator or entropy gradient is provided, so the central theoretical premise rests on a same-author citation chain rather than on a derivation or independently verifiable mechanism in this manuscript.
full rationale
The empirical core of FindRec—consistent gains over external baselines on MovieLens-100k, MicroLens, and Amazon Beauty, supported by ablations and hyperparameter analyses—is a self-contained systems result and is not circular. The benchmark comparisons, relative improvements, and code release stand independently of the theoretical narrative. However, the paper's headline theoretical claim that IICM 'theoretically guarantees distribution consistency' and provides 'unbiased distribution alignment' via entropy maximization is, by the paper's own equations, the name of the loss being optimized: Eq. (2) is an RBF-kernel similarity expectation, and it is the only IICM term in the total loss of Eq. (5). No separate entropy objective, Stein operator, or score-function term is defined, so the promised property reduces to the objective itself rather than following from it. The additional unbiasedness and entropy-gradient assertions are imported from same-group prior work [46] (and related papers [45,47]) without in-paper derivation or implementation details. Because the empirical results are externally validated, the overall circularity is partial rather than total; the theoretical contribution, however, is at least partly circular and partly self-citation-dependent, warranting a 6.
Assumptions & free parameters
free parameters (2)
- Kernel bandwidth sigma =
not specified, adaptively estimated
- IICM loss weight lambda =
1e-3
assumptions (3)
- domain assumption RBF kernel similarity maximization encourages distribution alignment between two representation spaces.
- standard math The Stein gradient estimator from references [28,46] provides an unbiased estimate of the entropy gradient with respect to model parameters in this deep learning setting.
- domain assumption Pretrained BLIP and RoBERTa features are informative for user preference and do not introduce harmful distribution shift.
Cite this review
Pith. "Pith review of FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation." pith.science (2026). https://pith.science/paper/VJYMCLGG
@misc{pith2026250704651,
author = {Pith},
title = {Pith review of: FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VJYMCLGG}},
note = {Machine review of arXiv:2507.04651}
}
read the original abstract
Modern recommendation systems face significant challenges in processing multimodal sequential data, particularly in temporal dynamics modeling and information flow coordination. Traditional approaches struggle with distribution discrepancies between heterogeneous features and noise interference in multimodal signals. We propose \textbf{FindRec}~ (\textbf{F}lexible unified \textbf{in}formation \textbf{d}isentanglement for multi-modal sequential \textbf{Rec}ommendation), introducing a novel "information flow-control-output" paradigm. The framework features two key innovations: (1) A Stein kernel-based Integrated Information Coordination Module (IICM) that theoretically guarantees distribution consistency between multimodal features and ID streams, and (2) A cross-modal expert routing mechanism that adaptively filters and combines multimodal features based on their contextual relevance. Our approach leverages multi-head subspace decomposition for routing stability and RBF-Stein gradient for unbiased distribution alignment, enhanced by linear-complexity Mamba layers for efficient temporal modeling. Extensive experiments on three real-world datasets demonstrate FindRec's superior performance over state-of-the-art baselines, particularly in handling long sequences and noisy multimodal inputs. Our framework achieves both improved recommendation accuracy and enhanced model interpretability through its modular design. The implementation code is available anonymously online for easy reproducibility~\footnote{https://github.com/Applied-Machine-Learning-Lab/FindRec}.
Figures
Reference graph
Works this paper leans on
-
[1]
Minahz Habib Afsar, Th Trong Duy Le, and Weiqing Wang. 2021. Deep Rein- forcement Learning for Search, Recommendation, and Online Advertising: A Survey. ACM Comput. Surv. (2021), 60:1–60:36
work page 2021
-
[2]
Chengfeng An, Defu Lian, Enhong Chen, Xu Liu, Lin Li, and Chunfeng Yuan
-
[3]
Shuqing Bian, Xingyu Pan, Wayne Xin Zhao, Jinpeng Wang, Chuyuan Wang, and Ji-Rong Wen. 2023. Multi-modal mixture of experts representation learning for sequential recommendation. In Proc. of CIKM
work page 2023
-
[4]
Christopher M Bishop and Nasser M Nasrabadi. 2006. Pattern recognition and machine learning. Springer
work page 2006
-
[5]
Xiang-Rong Cai, Weinan Zhang, Qing Tan, Zhaokun Wang, and Jun Wang
-
[6]
Ahmet Alper Cetin, Chuhan Wu, Ahtsham Ishaq, Srijan Kumar, and Ehtsham Elahi. 2023. Hamur: Hyper Adapter for Multi-Domain Recommendation. In Seventeenth ACM Conference on Recommender Systems, RecSys 2023, Singapore, Singapore, September 18-22, 2023 . 54–65
work page 2023
-
[7]
Peng Chen and Omar Ghattas. 2020. Projected Stein variational gradient descent. Proc. of NeurIPS (2020), 1947–1958
work page 2020
-
[8]
Qiang Cui, Shu Wu, Qiang Liu, Wen Zhong, and Liang Wang. 2018. MV-RNN: A multi-view recurrent neural network for sequential recommendation. IEEE Transactions on Knowledge and Data Engineering (2018), 317–331
work page 2018
Show all 64 references
-
[9]
Xinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu, Fuzhen Zhuang, Guan- feng Liu, Yanchi Liu, and Victor S Sheng. 2023. Frequency enhanced hybrid attention network for sequential recommendation. In Proc. of SIGIR. 78–88
2023
-
[10]
Ziwei Fan, Zhiwei Liu, Jiawei Zhang, Yun Xiong, Lei Zheng, and Philip S Yu. 2021. Continuous-time sequential recommendation with temporal graph collaborative transformer. In Proc. of CIKM. 433–442
2021
-
[11]
Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Jie Wang, and Joemon M Jose. 2024. IISAN: Efficiently Adapting Multimodal Rep- resentation for Sequential Recommendation with Decoupled PEFT. In Proc. of SIGIR. 687–697
2024
-
[12]
Chongming Gao, Shijun Li, Jiawei Chen, Bingsheng He, Xiangnan He, Jingsen Zhang, Zhenhui Li, and Philip S. Yu. 2022. KuaiSim: A Comprehensive Simulator for Recommender Systems. In Proc. of CIKM. 3003–3012
2022
-
[13]
Jingtong Gao, Xiangyu Zhao, Muyang Li, Minghao Zhao, Runze Wu, Ruocheng Guo, Yiding Liu, and Dawei Yin. 2024. SMLP4Rec: An Efficient all-MLP Architec- ture for Sequential Recommendations. ACM Transactions on Information Systems (2024), 1–23
2024
-
[14]
Albert Gu and Tri Dao. 2023. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. ArXiv (2023)
2023
-
[15]
Yejing Hao, Yujing Wang, Yunke Zhang, Zhaocheng Liu, Haifeng Zhang, YAMAN Kumar, Cen Chen, Yang Song, Cai-Zhi Weng, Tak chung Fu, Yuxuan Xiong, Wei-Wei Tu, Chen CHEN, Yiqi Wang, Wenjie Li, Wei Wu, Minghui Qiu, and Zhaowei Wang. 2023. PLATE: A Prompt-Enhanced Paradigm for Multi...
2023
-
[16]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk
-
[17]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. In Proc. of KDD. 585–593
2022
-
[18]
Hengchang Hu, Wei Guo, Yong Liu, and Min-Yen Kan. 2023. Adaptive multi- modalities fusion in sequential recommendation systems. In Proc. of CIKM. 843– 853
2023
-
[19]
Xilin Jiang, Cong Han, and Nima Mesgarani. 2024. Dual-path mamba: Short and long-term bidirectional selective structured state space models for speech separation. arXiv preprint arXiv:2403.18257 (2024)
2024 arXiv
-
[20]
Mengyuan Jing, Yanmin Zhu, Tianzi Zang, and Ke Wang. 2023. Contrastive self-supervised learning in recommender systems: A survey. ACM Transactions on Information Systems (2023), 1–39
2023
-
[21]
Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In Proc. of ICDM. 197–206
2018
-
[22]
Guangda Li, Manqing Dong, Changsheng Ma, Feng Duan, Ye Li, and Enhong Chen. 2022. Single-Shot Feature Selection for Multi-Task Recommendations. In Proc. of KDD. 898–907
2022
-
[23]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proc. of ICML. 12888–12900
2022
-
[24]
Jiacheng Li, Ming Wang, Jin Li, Jinmiao Fu, Xin Shen, Jingbo Shang, and Julian McAuley. 2023. Text is all you need: Learning language representations for sequential recommendation. In Proc. of KDD. 1258–1267
2023
-
[25]
Jiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang, Wanyu Wang, Haochen Liu, and Zitao Liu. 2023. Mmmlp: Multi-modal multilayer perceptron for sequen- tial recommendations. InProceedings of the ACM Web Conference 2023. 1109–1117
2023
-
[26]
Chang Liu, Xiaoguang Li, Guohao Cai, Zhenhua Dong, Hong Zhu, and Lifeng Shang. 2021. Noninvasive self-attention for side information fusion in sequential recommendation. In Proc. of AAAI. 4249–4256
2021
-
[27]
Chengkai Liu, Jianghao Lin, Jianling Wang, Hanzhou Liu, and James Caverlee
-
[28]
Qiang Liu and Dilin Wang. 2016. Stein variational gradient descent: A general purpose bayesian inference algorithm. Proc. of NeurIPS (2016)
2016
-
[29]
Weiming Liu, Jiajie Su, Chaochao Chen, and Xiaolin Zheng. 2021. Leveraging distribution alignment via stein path for cross-domain cold-start recommendation. Proc. of NeurIPS (2021), 19223–19234
2021
-
[30]
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang. 2016. Large-margin softmax loss for convolutional neural networks. arXiv preprint arXiv:1612.02295 (2016)
2016 arXiv
-
[31]
Xiaocong Liu, Zhengtao Wu, Chao Li, Lina Yao, Xiangmin Zhou, and Chengfei Liu. 2020. UserSim: User Simulation via Supervised Generative Adversarial Network. In Proc. of CIKM. 955–964
2020
-
[32]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020. RoBERTa: A Robustly Optimized BERT Pretraining Approach. In Proc. of ICLR
2020
-
[33]
I Loshchilov. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)
2017 arXiv
-
[34]
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jiyan Yang, Kr- ishnakumar Easwaran, Dmitry Kremez, Gema Parcerisas, Xiaodong Wang, Joe Isaacson, Roman Levenstein, Misha Smelyanskiy, Bill Jia, Xianjie Guan, Bor- Yiing Su, Narine Kokhlikyan, John Kim, Sam Naghshineh, and...
2020
-
[35]
Yongxin Ni, Yu Cheng, Xiangyan Liu, Junchen Fu, Youhua Li, Xiangnan He, Yongfeng Zhang, and Fajie Yuan. 2023. A Content-Driven Micro-Video Recom- mendation Dataset at Scale. arXiv preprint arXiv:2309.15379 (2023)
2023 arXiv
-
[36]
Haohao Qu, Liangbo Ning, Rui An, Wenqi Fan, Tyler Derr, Hui Liu, Xin Xu, and Qing Li. 2024. A survey of mamba. arXiv preprint arXiv:2408.01129 (2024)
2024 arXiv
-
[37]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[38]
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszko- reit, Mario Lucic, and Alexey Dosovitskiy. 2021. MLP-mixer: an all-MLP archi- tecture for vision. In Proc. of ICONIP
2021
-
[39]
Flavian Vasile, Elena Smirnova, and Alexis Conneau. 2016. Meta-prod2vec: Product embeddings using side-information for recommendation. In Proceedings of the 10th ACM conference on recommender systems . 225–232
2016
-
[40]
Jie Wang, Fajie Yuan, Mingyue Cheng, Joemon M Jose, Chenyun Yu, Beibei Kong, Zhijin Wang, Bo Hu, and Zang Li. 2024. Transrec: Learning transferable recom- mendation from mixture-of-modality feedback. In Asia-Pacific Web (APWeb) and Web-Age Information Management (W AIM) Joint ...
2024
-
[41]
Jinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang, Xingyu Lu, Tianxiang Li, Jun Yuan, Rui Zhang, Hai-Tao Zheng, and Shu-Tao Xia. 2023. MISSRec: Pre- training and transferring multi-modal interest-aware sequence representation for recommendation. In Proc. of ACM MM. 6548–6557
2023
-
[42]
Maolin Wang, Sheng Zhang, Ruocheng Guo, Wanyu Wang, Xuetao Wei, Zitao Liu, Hongzhi Yin, Yi Chang, and Xiangyu Zhao. 2025. STAR-Rec: Making Peace with Length Variance and Pattern Diversity in Sequential Recommendation. In Proceedings of the 48th International ACM SIGIR Conferen...
2025
-
[43]
Shoujin Wang, Liang Hu, Yan Wang, Longbing Cao, Quan Z Sheng, and Mehmet Orgun. 2019. Sequential recommender systems: challenges, progress and prospects. arXiv preprint arXiv:2001.04830 (2019)
2019 arXiv
-
[44]
Shuhan Wang, Bin Shen, Xu Min, Yong He, Xiaolu Zhang, Liang Zhang, Jun Zhou, and Linjian Mo. 2024. Aligned Side Information Fusion Method for Sequential Recommendation. In Companion Proceedings of the ACM on Web Conference 2024 . 112–120
2024
-
[45]
Liangjian Wen, Haoli Bai, Lirong He, Yiji Zhou, Mingyuan Zhou, and Zenglin Xu
-
[46]
Liangjian Wen, Xiasi Wang, Jianzhuang Liu, and Zenglin Xu. 2024. MVEB: Self- Supervised Learning With Multi-View Entropy Bottleneck. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[47]
Liangjian Wen, Yiji Zhou, Lirong He, Mingyuan Zhou, and Zenglin Xu. 2020. Mu- tual information gradient estimation for representation learning. arXiv preprint arXiv:2005.01123 (2020)
2020 arXiv
-
[48]
Ruobing Xin, Lixin Zou, Feng Zhang, Weidong Liu, Peng Wu, Le Wu, Chao- liang Zhong, Yang Cao, and Senzhang Wang. 2023. User Retention-oriented Recommendation with Decision Transformer. In Proc. of KDD. 2552–2562. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Maolin Wang, et al
2023
-
[49]
Jiyuan Yang, Yuanzi Li, Jingyu Zhao, Hanbing Wang, Muyang Ma, Jun Ma, Zhaochun Ren, Mengqi Zhang, Xin Xin, Zhumin Chen, et al. 2024. Uncovering Se- lective State Space Model’s Capabilities in Lifelong Sequential Recommendation. arXiv preprint arXiv:2403.16371 (2024)
2024 arXiv
-
[50]
Wenwen Ye, Shuaiqiang Wang, Xu Chen, Xuepeng Wang, Zheng Qin, and Dawei Yin. 2020. Time matters: Sequential recommendation with complex temporal information. In Proc. of SIGIR. 1459–1468
2020
-
[51]
Knowledge- Based Systems (2021), 107046
Gradient estimation of information measures in deep learning. Knowledge- Based Systems (2021), 107046
2021
-
[52]
Sheng Zhang, Maolin Wang, Wanyu Wang, Jingtong Gao, Xiangyu Zhao, Yu Yang, Xuetao Wei, Zitao Liu, and Tong Xu. 2025. GLINT-RU: Gated Lightweight Intelligent Recurrent Units for Sequential Recommender Systems. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discov...
2025
-
[53]
Sheng Zhang, Maolin Wang, Xiangyu Zhao, Ruocheng Guo, Yao Zhao, Chenyi Zhuang, Jinjie Gu, Zijian Zhang, and Hongzhi Yin. 2024. DNS-Rec: Data-aware Neural Architecture Search for Recommender Systems. In Proceedings of the 18th ACM Conference on Recommender Systems . 591–600
2024
-
[54]
Zijian Zhang, Shuchang Liu, Jiaao Yu, Qingpeng Cai, Xiangyu Zhao, Chunxu Zhang, Ziru Liu, Qidong Liu, Hongwei Zhao, Lantao Hu, et al . 2024. M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation Framework. In Proc. of SIGIR. 893–902
2024
-
[55]
Xiangyu Zhao, Chong Wang, Ming Chen, Xiaofeng Yi, Dawei Yin, and Jiliang Tang. 2019. Jointly Learning to Recommend and Advertise. In Proc. of WWW . 2379–2389
2019
-
[56]
Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Dawei Yin, Yihong Eric Zhao, and Jil- iang Tang. 2018. Deep Reinforcement Learning for Page-wise Recommendations. In Proceedings of the 12th ACM Conference on Recommender Systems, RecSys ’18, Vancouver, BC, Canada, October 2-7, 2018 . 95–103
2018
-
[57]
Shengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang, and Hui Xiong. 2025. Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recom- mendation. arXiv preprint arXiv:2501.14269 (2025)
2025 arXiv
-
[58]
Hongyu Zhou, Xin Zhou, Zhiwei Zeng, Lingzi Zhang, and Zhiqi Shen. 2023. A Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions. arXiv preprint arXiv:2302.04473 (2023). A More Hyperparameter Experimets Our extensive hyperparameter ...
2023 arXiv
-
[63]
Xiangyu Zhao, Lixin Zou, Mengying Tai, Haochen Liu, Dawei Yin, and Jiliang Tang. 2019. Recommendations with Negative Feedback via Pairwise Deep Rein- forcement Learning. In Proc. of CIKM. 329–338
2019
-
[2015]
arXiv preprint arXiv:1511.06939 (2015)
Session-based recommendations with recurrent neural networks. arXiv preprint arXiv:1511.06939 (2015)
2015 arXiv
-
[2019]
BERT4Rec: Sequential recommendation with bidirectional encoder repre- sentations from transformer. In Proc. of CIKM. 1441–1450
-
[2020]
In Proceedings of the 13th International Conference on Web Search and Data Mining, WSDM ’20, Houston, TX, USA, February 3-7, 2020
DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender Systems. In Proceedings of the 13th International Conference on Web Search and Data Mining, WSDM ’20, Houston, TX, USA, February 3-7, 2020 . 71–79
2020
-
[2021]
STRec: Sparse Transformer for Sequential Recommendations. In Proc. of CIKM. 42–51
-
[2024]
arXiv preprint arXiv:2403.03900 (2024)
Mamba4Rec: Towards Efficient Sequential Recommendation with Selective State Space Models. arXiv preprint arXiv:2403.03900 (2024)
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.