REVIEW 38 references
Model-agnostic post-hoc explainability for recommender systems
T0 review · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Deletion diagnostics reveal which users and items move a recommender's performance.
desk verdict Deletion diagnostics applied to recommenders, but the implementation deletes the target user/item before the train/test split, so the reported influence scores conflate training removal with evaluation-population shift. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the deletion diagnostic, formalized in Algorithms 1 and 2: iterate over users (or items), generate the dataset without that observation, retrain the model, evaluate with MAP, and store the difference from the original model's MAP. The difference is computed as $\mathrm{Influence}(-i)=\mathrm{eval}-\mathrm{eval}^{(-i)}$. This is a leave-one-out perturbation analysis adapted to recommendation metrics; its explanatory power comes from full retraining, which captures global effects that gradient-based or surrogate approximations can miss, at the cost of one retraining per observation.
What would settle it
Train the same recommender several times from different random seeds with and without a fixed user, and compare the spread of influence scores across seeds to the spread across users; if the seed-to-seed variation is comparable to or larger than the user-to-user variation, the deletion diagnostic is not isolating the effect of that user.
Extended reading notes
Core claim
The central claim is that influence in recommender systems is a measurable, model-agnostic quantity obtained by full retraining. For each user or item $i$, train the same architecture on the data with $i$ omitted, recompute a ranking metric, and record $\mathrm{eval}-\mathrm{eval}^{(-i)}$. Observations whose deletion lowers performance are positively influential; observations whose deletion raises performance are negatively influential. Because the procedure only requires a performance metric and the ability to retrain, it applies to any recommender, and the experiments demonstrate it on both a neural architecture and SVD. The reported results also show an asymmetry: removing the ten least influential users from the MovieLens NCF model improved MAP@K by 18.49% and Precision@K by 16.75%, suggesting that some training instances contribute noise rather than signal.
Load-bearing premise
The method assumes that the difference in evaluation metrics after one retraining is a stable property of the deleted observation, rather than a product of random initialization and training noise; the NCF runs use 10 epochs and no reported seeds, which is exactly the regime where that assumption can fail.
Editorial extensions
If this is right
- The recipe for influence is metric-agnostic and architecture-agnostic, so the same procedure can audit any recommender that can be retrained.
- Removing or down-weighting consistently negatively influential users and items is a data-curation strategy that can improve MAP, NDCG, and precision.
- The method gives developers a debugging signal: the most positively influential observations are the ones the model depends on most, and their characteristics can be inspected.
- Because it needs no gradients or internal parameters, the method works on proprietary or otherwise opaque recommenders.
- Retraining cost can be reduced by parallelization, subsampling, or focusing on top-K candidates, making the approach feasible at larger scale.
Reading between the lines
- The reported influence scores are single runs with no variance or seed information, so a natural next test is whether the ranking of influential users is stable across random seeds; without that stability, part of the signal may be training noise.
- The asymmetry between removing positive and negative influencers suggests that many recommenders are trained on a long tail of near-redundant interactions, and pruning that tail may offer larger gains than protecting the top.
- A testable extension is to compare the deletion-based ranking with influence estimates from cheaper approximations, asking whether exact retraining changes the practical decisions a data curator would make.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Circularity Check
Influence scores are constructed so that deleting a user changes the evaluation population; the ranking is algebraically tied to the user's own AP, and the validation is a consistency check of that construction.
-
self definitional
[Section 3.2, Eq. (9) and Algorithms 1-2; results in Section 4.1]
"The influence of a user or item i is computed as: Influence(−i) = eval − eval(−i) ... Algorithm 1: ... Data generation: Generate the data without participant u and split train/test; X(−u) = X\{u} ... Append to differences: MAP − MAP(−u)"
The split occurs after deleting u from X, so eval(−u) is MAP on a test set that excludes u, whereas eval includes u. Under no retraining effect, MAP − MAP(−u) = (AP_u − MAP)/(|U|−1). Hence the influence score is algebraically the centered AP of the deleted user (with an analogous composition effect for items), so the method's ranking of 'most/least influential' users is forced by the evaluation-population shift, not by the model's training dynamics. The quantity measured is not the promised influence of an observation on the recommender.
-
fitted input called prediction
[Section 4.1, Tables 5-6 and accompanying text]
"The results in Tables 5 and 6 demonstrate that the deletion diagnosis successfully identified both the most and, especially, the least influential users for this model. The improvement in performance following the removal of the least influential users suggests that these users contribution is minimal or even detrimental signal during training."
The 'validation' removes the users the algorithm selected as least influential and observes that MAP improves. But those users were selected because, by the Algorithm 1 construction, deleting them from the evaluation population raises MAP when their AP is below the original MAP. The check therefore restates the selection criterion rather than independently confirming that the method detects training-data influence. The same structure appears for items in Tables 8-9.
full rationale
The central definition Eq. (9) is a direct measurement and is not itself a fitted parameter, so the paper is not circular in the 'parameter fitted to a subset and then predicting that subset' sense. However, Algorithms 1 and 2 delete the user/item before the train/test split, so eval(−i) is computed on a different test population than eval. For MAP, Influence(−u) equals (AP_u − MAP)/(|U|−1) plus whatever genuine retraining effect exists; thus a user's label as positively or negatively influential is largely forced by their own average precision on the original test set. The subsequent 'validation' in Section 4.1 removes the 10 least influential users and reports improved MAP; since those users were selected via the same test-population-shifted difference, the improvement is partly a restatement of the construction, not independent evidence about training-data influence. No load-bearing self-citation, uniqueness theorem, or ansatz-smuggling was found. The paper's honest limitation statements about runtime and metric dependence do not affect this assessment.
Assumptions & free parameters
free parameters (1)
- NCF hyperparameters =
embedding_dim=4, layers=[16,8,4], lr=1e-3, batch=256, epochs=10
assumptions (3)
- domain assumption MAP is an adequate measure of recommendation quality for the purpose of computing influence.
- domain assumption Retraining without a single user or item yields a model sufficiently similar to the original for the metric difference to be attributable to that observation.
- domain assumption Removing an observation from training data does not change the train/test split in a way that biases the comparison.
Cite this review
Pith. "Pith review of Model-agnostic post-hoc explainability for recommender systems." pith.science (2026). https://pith.science/paper/RT2NMQEK
@misc{pith2026250910245,
author = {Pith},
title = {Pith review of: Model-agnostic post-hoc explainability for recommender systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/RT2NMQEK}},
note = {Machine review of arXiv:2509.10245}
}
read the original abstract
Recommender systems often benefit from complex feature embeddings and deep learning algorithms, which deliver sophisticated recommendations that enhance user experience, engagement, and revenue. However, these methods frequently reduce the interpretability and transparency of the system. In this research, we develop a systematic application, adaptation, and evaluation of deletion diagnostics in the recommender setting. The method compares the performance of a model to that of a similar model trained without a specific user or item, allowing us to quantify how that observation influences the recommender, either positively or negatively. To demonstrate its model-agnostic nature, the proposal is applied to both Neural Collaborative Filtering (NCF), a widely used deep learning-based recommender, and Singular Value Decomposition (SVD), a classical collaborative filtering technique. Experiments on the MovieLens and Amazon Reviews datasets provide insights into model behavior and highlight the generality of the approach across different recommendation paradigms.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
J. B. Schafer, J. Konstan, J. Riedl, Recom- mender systems in e-commerce, in: Proceedings of the 1st ACM Conference on Electronic Com- merce, EC ’99, Association for Computing Ma- chinery, New York, NY, USA, 1999, p. 158–166. doi:10.1145/336992.337035
arXiv 1999
-
[2]
A. Rivas, P. Chamoso, A. Gonz´ alez-Briones, J. Pav´ on, J. M. Corchado, Social network rec- ommender system, a neural network approach, in: Intelligent Data Engineering and Auto- mated Learning – IDEAL 2020: 21st Interna- tional Conference, Guimaraes, Portugal, Novem- ber 4–6, 2020, Proceedings, Part II, Springer- Verlag, Berlin, Heidelberg, 2020, p. 213–222
work page 2020
-
[3]
S. Chang, Y. Zhang, J. Tang, D. Yin, Y. Chang, M. A. Hasegawa-Johnson, T. S. Huang, Stream- ing recommender systems, in: Proceedings of the 26th International Conference on World Wide Web, WWW ’17, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 2017, p. 381–389. doi:10.1145/3038912.3052627
arXiv 2017
-
[5]
N. Tintarev, Explanations of recommendations, in: Proceedings of the 2007 ACM Conference on Recommender Systems, RecSys ’07, Association for Computing Machinery, New York, NY, USA, 2007, p. 203–206. doi:10.1145/1297231.1297275
arXiv 2007
-
[6]
G. Carenini, J. Moore, An empirical study of the influence of user tailoring on evaluative ar- gument effectiveness, in: Proceedings of the Sev- enteenth International Joint Conference on Ar- tificial Intelligence, IJCAI 2001, Seattle, Wash- ington, USA, August 4-10, 2001, 2001, pp. 1307– 1314
work page 2001
-
[7]
G. Zhou, C. Huang, X. Chen, X. Xu, C. Wang, L. Zhu, L. Yao, Contrastive counterfactual learning for causality-aware interpretable rec- ommender systems, in: Proceedings of the 32nd ACM International Conference on Infor- mation and Knowledge Management, CIKM ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 3564–3573. doi:10.1145/358378...
arXiv 2023
-
[8]
Y. Deldjoo, D. Jannach, A. Bellogin, A. Difonzo, D. Zanzonelli, Fairness in recommender systems: research landscape and future directions, User Modeling and User-Adapted Interaction 34 (1) (2024) 59–108
work page 2024
-
[9]
J. Chen, H. Dong, X. Wang, F. Feng, M. Wang, X. He, Bias and debias in recommender system: A survey and future directions, ACM Trans. Inf. Syst. 41 (3) (Feb. 2023). doi:10.1145/3564284
doi:10.1145/3564284 2023
Show all 38 references
-
[10]
Mansoury, H
M. Mansoury, H. Abdollahpouri, M. Pech- enizkiy, B. Mobasher, R. Burke, Feedback loop and bias amplification in recommender systems, in: Proceedings of the 29th ACM International Conference on Information & Knowledge Man- agement, CIKM ’20, Association for Comput- ing Machiner...
2020
-
[11]
Y. Li, H. Chen, S. Xu, Y. Ge, J. Tan, S. Liu, Y. Zhang, Fairness in recommendation: Foundations, methods, and applications, ACM Trans. Intell. Syst. Technol. 14 (5) (Oct. 2023). doi:10.1145/3610302
2023 doi
-
[12]
Zhang, X
Y. Zhang, X. Chen, Explainable recommenda- tion: A survey and new perspectives, Found. Trends Inf. Retr. 14 (1) (2020) 1–101
2020
-
[13]
J. Vig, S. Sen, J. Riedl, Tagsplanations: ex- plaining recommendations using tags, in: Pro- ceedings of the 14th International Conference on 15 Intelligent User Interfaces, IUI ’09, Association for Computing Machinery, New York, NY, USA, 2009, p. 47–56. doi:10.1145/1502650.1502661
2009
-
[14]
J. L. Herlocker, J. A. Konstan, J. Riedl, Explaining collaborative filtering recommenda- tions, in: Proceedings of the 2000 ACM Con- ference on Computer Supported Cooperative Work, CSCW ’00, Association for Comput- ing Machinery, New York, NY, USA, 2000, p. 241–250. doi:10.114...
-
[15]
B. M. Sarwar, G. Karypis, J. A. Konstan, J. Riedl, Item-based collaborative filtering rec- ommendation algorithms, in: V. Y. Shen, N. Saito, M. R. Lyu, M. E. Zurko (Eds.), Proceedings of the Tenth International World Wide Web Conference, WWW 10, Hong Kong, China, May 1-5, 2001...
2001
-
[16]
Bauman, B
K. Bauman, B. Liu, A. Tuzhilin, Aspect based recommendations: Recommending items with the most valuable aspects based on user reviews, in: Proceedings of the 23rd ACM SIGKDD In- ternational Conference on Knowledge Discov- ery and Data Mining, KDD ’17, Association for Computing...
2017
-
[17]
Y. Lu, R. Dong, B. Smyth, Coevolutionary recommendation model: Mutual learning be- tween ratings and reviews, in: Proceedings of the 2018 World Wide Web Conference, WWW ’18, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 2018, ...
2018
-
[18]
Zhang, G
Y. Zhang, G. Lai, M. Zhang, Y. Zhang, Y. Liu, S. Ma, Explicit factor models for explainable rec- ommendation based on phrase-level sentiment analysis, in: Proceedings of the 37th Inter- national ACM SIGIR Conference on Research & Development in Information Retrieval, SI- GIR ’...
2014
-
[19]
Abdollahi, O
B. Abdollahi, O. Nasraoui, Using explainability for constrained matrix factorization, in: Pro- ceedings of the Eleventh ACM Conference on Recommender Systems, RecSys ’17, Association for Computing Machinery, New York, NY, USA, 2017, p. 79–83. doi:10.1145/3109859.3109913
2017
-
[20]
M. T. Ribeiro, S. Singh, C. Guestrin, ”why should i trust you?”: Explaining the predic- tions of any classifier, in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 1135–1144
2016
-
[21]
Tohidi, M
N. Tohidi, M. Beheshti, Enhanced explanations in recommendation systems, in: 2024 IEEE In- ternational Symposium on Systems Engineering (ISSE), 2024, pp. 1–5
2024
-
[22]
P. W. Koh, P. Liang, Understanding black-box predictions via influence functions, in: Proceed- ings of the 34th International Conference on Ma- chine Learning, PMLR, 2017, pp. 1885–1894
2017
-
[23]
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, T.- S. Chua, Neural collaborative filtering, in: Pro- ceedings of the 26th International Conference on World Wide Web, WWW ’17, International World Wide Web Conferences Steering Commit- tee, Republic and Canton of Geneva, CHE, 2017, p. 173–182
2017
-
[24]
Ponnusamy, W.-K
C. Ponnusamy, W.-K. Wong, A. Raja, O. Kha- laf, A. Kiran, J. Babu, Health recommen- dation system using deep learning-based col- laborative filtering, Heliyon 9 (2023) e22844. doi:10.1016/j.heliyon.2023.e22844
2023 doi
-
[25]
H. L. Mulyana, F. Rumaisa, Course learn- ing recommendation system using neural col- laborative filtering, Brilliance: Research of Artificial Intelligence 4 (2) (2024) 517–524. doi:10.47709/brilliance.v4i2.4699
2024 doi
-
[26]
Marzuki, M
I. Marzuki, M. Hariadi, R. Rachmadi, Y. Arif, Neural collaborative filtering for improved tourism destination recommendation, in: 2024 8th International Conference on Information 16 Technology, Information Systems and Electri- cal Engineering (ICITISEE), 2024, pp. 481–486. doi...
2024
-
[27]
Wei, Tourist attraction image recognition and intelligent recommendation based on deep learning, Journal of computational methods in sciences and engineering (2025 FEB 7 2025)
C. Wei, Tourist attraction image recognition and intelligent recommendation based on deep learning, Journal of computational methods in sciences and engineering (2025 FEB 7 2025). doi:10.1177/14727978251318805
2025 doi
-
[28]
V. C, H. Oberoi, A. Goyal, N. Sikka, Re- recsys: An end-to-end system for recommend- ing properties in real-estate domain, in: Pro- ceedings of the 7th Joint International Confer- ence on Data Science; Management of Data (11th ACM IKDD CODS and 29th COMAD), CODS-COMAD 2024, AC...
2024
-
[29]
Koren, R
Y. Koren, R. Bell, C. Volinsky, Matrix factoriza- tion techniques for recommender systems, Com- puter 42 (8) (2009) 30–37
2009
-
[30]
Abdollahpouri, R
H. Abdollahpouri, R. Burke, B. Mobasher, Con- trolling popularity bias in learning-to-rank rec- ommendation, in: Proceedings of the 11th ACM Conference on Recommender Systems, ACM, 2017, pp. 42–46
2017
-
[31]
Steck, Calibrated recommendations, Pro- ceedings of the 12th ACM Conference on Rec- ommender Systems 2018 (2018) 154–162
H. Steck, Calibrated recommendations, Pro- ceedings of the 12th ACM Conference on Rec- ommender Systems 2018 (2018) 154–162
2018
-
[32]
Jadon, A
A. Jadon, A. Patil, A comprehensive survey of evaluation techniques for recommendation sys- tems (2024). arXiv:2312.16015. URLarxiv.org/html/2312.16015v2
2024 arXiv
-
[33]
Zangerle, C
E. Zangerle, C. Bauer, Evaluating recommender systems: Survey and framework, ACM Comput- ing Surveys 55 (8) (2022) 1–38
2022
-
[34]
F. M. Harper, J. A. Konstan, The movie- lens datasets: History and context, ACM Trans. Interact. Intell. Syst. 5 (4) (Dec. 2015). doi:10.1145/2827872
2015 doi
-
[35]
Rendle, W
S. Rendle, W. Krichene, L. Zhang, J. Anderson, Neural collaborative filtering vs. matrix factor- ization revisited (2020). arXiv:2005.09683. URLarxiv.org/abs/2005.09683
2020 arXiv
-
[36]
Kammoun, R
A. Kammoun, R. Slama, H. Tabia, T. Ouni, M. Abid, Generative adversar- ial networks for face generation: A sur- vey, ACM Computing Surveys (Mar. 2022). doi:10.1145/1122445.1122456
2022
-
[37]
H. Fan, M. Zhu, Y. Hu, H. Feng, Z. He, H. Liu, Q. Liu, Tim4rec: An efficient sequen- tial recommendation model based on time-aware structured state space duality model (2024). arXiv:2409.16182. URLarxiv.org/abs/2409.16182
2024 arXiv
-
[38]
Y. Chen, J. Tan, A. Zhang, Z. Yang, L. Sheng, E. Zhang, X. Wang, T.-S. Chua, On softmax direct preference optimization for recommenda- tion (2024). arXiv:2406.09215. URLarxiv.org/abs/2406.09215
2024 arXiv
-
[39]
McAuley, J
J. McAuley, J. Leskovec, Hidden factors and hidden topics: understanding rating dimensions with review text, in: Proceedings of the 7th ACM Conference on Recommender Systems, RecSys ’13, Association for Computing Machin- ery, New York, NY, USA, 2013, p. 165–172. doi:10.1145/25...
2013
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.