REVIEW 4 major objections 5 minor 195 references
A Comprehensive Study of Shapley Value in Data Analytics
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that Shapley value in data analytics reduces to four challenges and a small set of reusable technique blocks, and that hybrid combinations usually beat any single block.
desk verdict Useful survey and modular framework, but the headline efficiency and interpretability claims rest on comparisons that don't match accuracy or the actual definition of Shapley value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Shapley value formula itself, the weighted average over all $2^n$ coalitions of a player's marginal contribution to coalition utility (Equations 1 and 2), together with cooperative game modeling, in which the analyst decides who the players are (features, tuples, datasets, or derivatives such as trained models) and what the utility measures (goodness-of-fit scores or raw task outputs). The argument is carried by the decomposition of existing algorithms into reusable building blocks and by SVBench, a modular framework with configurable base algorithm, sampling strategy, optimization, and privacy modules that lets those blocks be recombined. What the machinery does is turn a sprawling literature into a testable design space: the paper's experiments implement dozens of combinations and adjudicate which blocks conflict and which cooperate.
What would settle it
Measure exact Shapley values and leave-one-out utility changes on the same task: for every player compute $\phi_i$, $U(N\setminus\{p_i\})-U(N)$, and $U(\{p_i\})-U(\emptyset)$, then compute the rank correlation between $\phi_i$ and each change across all players. The paper currently shows scatter-style plots that are read by eye; if a clean, high positive correlation appears across all four player types on, say, the DV-Wind and DSV-2Dplanes tasks the authors themselves flag as mismatched, the 'fluctuate and contradict' conclusion would be overturned.
Extended reading notes
Core claim
The core claim is that the Shapley value has become a general-purpose instrument for the whole data analytics lifecycle, and that the field is best understood not as a zoo of task-specific algorithms but as a small design space. The paper condenses four challenges—computation efficiency, approximation error, privacy preservation, and interpretability—and disentangles existing solutions into reusable techniques: iteration reduction (Monte Carlo, regression, multilinear extension, group testing, compressive permutation sampling, truncation, predictor learning), ML speedup (gradient approximation, test sample skip, model appraiser), variance reduction (stratified, antithetic, kernel-based), and privacy techniques (non-perturbation masking, homomorphic encryption, secure multiparty computation, quantization, dimension reduction, differential privacy). It then builds SVBench to assemble these blocks and evaluates them on result interpretation, data tuple valuation, dataset valuation, and federated learning. The quantitative study supports two findings the authors emphasize: hybrid algorithms that combine multiple efficiency techniques generally beat single-technique ones, and the mainstream utility-based interpretation of Shapley values—larger value means more impact on the task's overall utility—does not hold up when players are removed or added.
Load-bearing premise
The claim that utility-based Shapley interpretations are contradicted by the data rests on treating $U(N\setminus\{p_i\})-U(N)$ and $U(\{p_i\})-U(\emptyset)$ as measures of a player's impact on overall utility, even though the Shapley value averages marginal contributions over all $2^n$ coalitions and these two specific differences need not track that average.
Editorial extensions
If this is right
- If hybrids usually beat single techniques, then practitioners should default to combining truncation with a sampling-based base algorithm, and to adding gradient approximation and test sample skipping when utility computation involves costly model training.
- If utility-based interpretation is unreliable, decisions that rank data or models by Shapley value—pricing, selection, weighting, attribution—need to report variance of marginal contributions alongside the value, not just the value.
- If SVBench's modularity holds, new Shapley applications can be built by recombining existing blocks instead of designing bespoke algorithms from scratch, lowering the engineering cost of data marketplaces and federated learning incentives.
- If the approximation-error results generalize, convergence thresholds should be tuned dynamically during sampling, stopping early when the downstream task (e.g., top-k selection) stops changing, rather than fixed a priori.
- If privacy techniques distort Shapley rankings, then differentially private Shapley reporting needs a strength-setting guide that balances attack prevention against ranking fidelity.
Reading between the lines
- A direct extension the authors leave implicit: measuring impact by leave-one-out utility differences may itself be the wrong benchmark for the utility-based paradigm, since Shapley value averages over all coalitions; a fairer test would compare Shapley value against the expected marginal contribution over the actual coalition-size distribution of the task.
- The variance numbers reported (e.g., marginal-contribution variance 11.83–18.00 vs average 0.06–1.43 in DV-Wind) suggest a testable redesign: a variance-aware Shapley variant that weights players not only by mean contribution but by volatility would likely align better with removal/addition impacts.
- Because the evaluation covers four task types but relatively few datasets per type, the 'in most cases' conclusion could be stress-tested on large-scale, high-dimensional tasks and pre-trained models, where both utility costs and coalition structure differ.
- The modular decomposition also invites a transfer result: if the building-block taxonomy is right, then any efficiency technique proven on data valuation should port to dataset pricing or model weighting with only a change of utility function; SVBench is the natural vehicle to test that.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a survey and an empirical benchmark of Shapley value (SV) techniques across the data analytics (DA) workflow. It taxonomizes SV applications into pricing, selection, weighting, and attribution; identifies four challenges (computation efficiency, approximation error, privacy preservation, interpretability); and decomposes existing methods into reusable techniques such as iteration reduction, ML speedup, variance reduction, and privacy protection. The authors implement SVBench, an open-source modular framework, and use it to evaluate hybrid algorithms, sampling strategies, privacy-protection measures, and the utility-based interpretability paradigm across four DA task types (RI, DV, DSV, FL) on ten datasets. The paper concludes with findings and seven research directions.
Significance. The main value of the paper is organizational: it provides a comprehensive taxonomy of SV use in DA, disentangles monolithic algorithms into modular techniques, and makes an implementable benchmark (SVBench) available. The code and data release, the breadth of cited work (183 references), and the explicit enumeration of findings and research directions are useful for practitioners and researchers. If the experimental claims are made rigorous, the framework could serve as a common testbed for comparing SV algorithms. The four-challenge decomposition and the conflict analysis (e.g., efficiency vs. error, privacy vs. effectiveness) are conceptually helpful.
major comments (4)
- [§5.1, Research Direction 1] The headline efficiency claim in §5.1 compares algorithms at a fixed self-stability threshold Δφ < τ with τ = 0.05 rather than at matched approximation error. Because TC truncates once U(S) is close to U(N), and GA/TSS accelerate training/evaluation at the cost of increased variance, a hybrid algorithm's φ̂ can stabilize before the approximation error ε is small. The paper reports Nuc, time, and ε as separate panels in Figure 4 but does not compare cost at a common ε tolerance or plot accuracy-adjusted cost curves. The text itself concedes that TC 'may enlarge the approximation error' and GA/TSS 'may enlarge the variance.' Consequently, the observed reductions in Nuc and time for hybrid algorithms may reflect early stopping rather than genuine efficiency–accuracy improvements, and the conclusion that hybrid algorithms 'achieve better performance' is not established. Please add error-matched comparisons (e.g., minimum cost to reach a fixed ε, or accuracy-conditioned cost curves) or qualify the claim accordingly.
- [§5.4, Figure 7] The interpretability experiment in §5.4 uses the single-coalition differences U(N\{p_i}) - U(N) and U({p_i}) - U(∅) as the operationalization of a player's 'impact on the task's overall utility' (Figure 7). However, the Shapley value φ_i is the weighted average of marginal contributions over all 2^n coalitions, so there is no a priori reason that these two particular differences should track φ_i. The paper reports that the variance of marginal contributions is large (e.g., DV-Wind variance 11.83–18.00 vs. average 0.06–1.43) but does not explain why leave-one-out impact is the correct benchmark for the utility-based interpretation. The conclusion that 'the fluctuating results contradict the mainstream SV interpretations' therefore overstates the evidence; the results show only that SV does not correlate with these two specific marginal contributions. A valid test would require a more representative measure of 'impact' (e.g., average absolute marginal contribution over coalitions) or a direct correlation between SV and a task-specific notion of utility change.
- [§5 (Evaluation introduction) and §§5.2–5.4] The evaluation section states that MLE is used as the base algorithm in §§5.2–5.4 because it 'generally achieves better efficiency and accuracy performance' in §5.1, and that 'varying base algorithms would not influence conclusions in those subsections.' This choice is made after observing the §5.1 results, and the invariance claim is not tested. Since §§5.2–5.4 draw conclusions about approximation-error trade-offs, privacy effectiveness, and interpretability that are intended to generalize across SV algorithms, at least one robustness check (e.g., repeating the key experiments with MC, the most widely used base algorithm) or a clear scope limitation is needed before these conclusions can be considered general.
- [§5.1–§5.4] No random seeds, repetitions, or dispersion measures are reported for any of the stochastic experiments in §5.1–§5.4 (MC/MLE sampling, privacy noise, attack implementations). The comparisons rely on single runs, so observed differences—especially the 'in most cases' claim in §5.1 and the AUROC/MAE differences in §5.3—could be within run-to-run variation. Please report multiple seeds with means/variances or a statistical test, or at minimum justify why the reported numbers are stable.
minor comments (5)
- [§5.2] The 'score of impacts' definition is written as a signed sum Σ (φ̂_i/Σφ̂_i − φ_i/Σφ_i), which can be negative and cancel out; the figure shows nonnegative values, so the formula should be clarified as, presumably, the sum of absolute differences.
- [§3.2.3] Finding 11 uses the abbreviation 'SMCP' for secure multiparty computation, while the correct term 'SMPC' is used elsewhere; please fix the inconsistency.
- [Table 9] The ♠ superscript on RE, MLE, GT, CP is mentioned in the text ('the ♠-tagged base algorithm') but not defined in the table or caption; please explain what it marks.
- [Table 1] The cell entries '/reve' are not explained in the text or caption; the reader has to infer that they indicate coverage of a given purpose or solution. Define the notation.
- [§1] The paper claims to be 'the first comprehensive survey of SV applied throughout the DA workflow'; given the existing surveys cited in Table 1, the novelty claim would be more precise if stated as 'first to decompose techniques into building blocks and benchmark them in a unified framework.'
Circularity Check
No circularity found: the paper is a survey and empirical benchmark whose central claims are supported by external evidence and fresh experiments, not by fitted inputs or self-citation chains.
full rationale
The paper's contribution is a taxonomy of Shapley-value techniques and an open-source benchmark (SVBench). Its main quantitative claim (Research Direction 1, §5.1) that hybrid algorithms integrating multiple efficiency optimization techniques outperform single-technique algorithms is an empirical measurement of total cost (Nuc×Tuc) and complexity (Nuc) under a fixed convergence threshold; it is not derived by fitting a parameter to the predicted quantity, and no equation in the paper reduces the conclusion to its own input by construction. The findings in §3.2 are compiled from a broad set of prior works, and none of the load-bearing premises rely on a self-citation or on a uniqueness theorem imported from the authors' own previous papers. The §5.4 test of the utility-based interpretation compares SV values against leave-one-out utility differences (U(N\{p_i})−U(N) and U({p_i})−U(∅)); although one can question whether those two differences properly operationalize a coalition-averaged quantity, the paper does not define the utility-based paradigm as those differences, so the test is an empirical falsification attempt rather than a tautology. Since every checked claim has independent content and the experimental results are reported as measurements rather than as consequences of fitting, no significant circularity is present.
Assumptions & free parameters
free parameters (3)
- convergence threshold tau =
0.05
- privacy strength levels =
DP sigma 0.1/0.5/0.9, QT 0.9n/0.5n/0.1n, DR 0.1n/0.5n/0.9n
- DV subsample sizes =
n=18,15,18,22,18 for Iris, Wine, Bank, Ttt, Wind
assumptions (3)
- domain assumption Shapley value is the appropriate solution concept for fair contribution allocation in DA tasks
- domain assumption Utility functions can be evaluated for every coalition S subset of N, including singletons and the empty set
- ad hoc to paper The selected 10 datasets and 4 task types are representative of the DA domain
Cite this review
Pith. "Pith review of A Comprehensive Study of Shapley Value in Data Analytics." pith.science (2026). https://pith.science/paper/LQBTOFYU
@misc{pith2026241201460,
author = {Pith},
title = {Pith review of: A Comprehensive Study of Shapley Value in Data Analytics},
year = {2026},
howpublished = {\url{https://pith.science/paper/LQBTOFYU}},
note = {Machine review of arXiv:2412.01460}
}
read the original abstract
Over the recent years, Shapley value (SV), a solution concept from cooperative game theory, has found numerous applications in data analytics (DA). This paper presents the first comprehensive study of SV used throughout the DA workflow, clarifying the key variables in defining DA-applicable SV and the essential functionalities that SV can provide for data scientists. We condense four primary challenges of using SV in DA, namely computation efficiency, approximation error, privacy preservation, and interpretability, disentangle the resolution techniques from existing arts in this field, then analyze and discuss the techniques w.r.t. each challenge and the potential conflicts between challenges.We also implement SVBench, a modular and extensible open-source framework for developing SV applications in different DA tasks, and conduct extensive evaluations to validate our analyses and discussions. Based on the qualitative and quantitative results, we identify the limitations of current efforts for applying SV to DA and highlight the directions of future research and engineering.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Kjersti Aas, Martin Jullum, and Anders Løland. 2021. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artif. Intell. 298 (2021), 103502
2021
-
[2]
Kjersti Aas, Thomas Nagler, Martin Jullum, and Anders Løland. 2021. Explain- ing predictive models using Shapley values and non-parametric vine copulas. Dependence Modeling 9, 1 (2021), 62–81
2021
-
[3]
Anish Agarwal, Munther Dahleh, and Tuhin Sarkar. 2019. A marketplace for data: An algorithmic solution. In EC. 701–726
2019
-
[4]
Marco Ancona, Cengiz Öztireli, and Markus H. Gross. 2019. Explaining Deep Neural Networks with a Polynomial Time Algorithm for Shapley Value Ap- proximation. In ICML. PMLR, 272–281
2019
-
[5]
Meriem Arbaoui, Mohamed-El-Amine Brahmia, Abdellatif Rahmoun, and Mourad Zghal. 2024. Optimizing Shapley Value for Client Valuation in Federated Learning through Enhanced GTG-Shapley. In IWCMC. 1528–1533
2024
-
[6]
Santiago Andrés Azcoitia, Costas Iordanou, and Nikolaos Laoutaris. 2023. Un- derstanding the Price of Data in Commercial Data Marketplaces. InICDE. 3718– 3728
2023
-
[7]
Santiago Andrés Azcoitia and Nikolaos Laoutaris. 2022. A Survey of Data Marketplaces and Their Business Models. SIGMOD 51 (2022), 18 – 29
2022
-
[8]
Santiago Andrés Azcoitia, Marius Paraschiv, and Nikolaos Laoutaris. 2022. Computing the relative value of spatio-temporal data in data marketplaces. In SIGSPATIAL ’22. 1–11
2022
Show all 195 references
-
[9]
Zahra Batool, Kaiwen Zhang, and Matthew Toews. 2022. FL-MAB: client selection and monetization for blockchain-based federated learning. In SAC ’22. 299–307
2022
-
[10]
Daniel Beechey, Thomas MS Smith, and Özgür Şimşek. 2023. Explaining rein- forcement learning with shapley values. In ICML. 2003–2014
2023
-
[11]
Cruz, Mario Figueiredo, and Pedro Bizarro
Joao Bento, Pedro Saleiro, André F. Cruz, Mario Figueiredo, and Pedro Bizarro
-
[12]
Bertossi, Benny Kimelfeld, Ester Livshits, and Mikaël Monet
Leopoldo E. Bertossi, Benny Kimelfeld, Ester Livshits, and Mikaël Monet. 2023. The Shapley Value in Database Management. SIGMOD 52 (2023), 6 – 17
2023
-
[13]
Giovanna Bimonte, Maria Russolillo, Han Lin Shang, and Yang Yang. 2024. Mortality models ensemble via Shapley value. Decis. Econ. Finance (2024), 1–29
2024
-
[14]
Joost Bosker, Marc Gürtler, and Marvin Zöllner. 2024. Machine learning-based variable selection for clustered credit risk modeling. Journal of Business Eco- nomics (2024), 1–36
2024
-
[15]
Aso Bozorgpanah, Vicenç Torra, and Laya Aliahmadipour. 2022. Privacy and Explainability: The Effects of Data Protection on Shapley Values. Technologies 10, 6 (2022). https://doi.org/10.3390/technologies10060125
2022 doi
-
[16]
Andreas Brandsæter and Ingrid K Glad. 2024. Shapley values for cluster impor- tance: How clusters of the training data affect a prediction. Data Mining and Knowledge Discovery 38, 5 (2024), 2633–2664
2024
-
[17]
Covert, Scott M
Hugh Chen, Ian C. Covert, Scott M. Lundberg, and Su-In Lee. 2023. Algorithms to estimate Shapley value feature attributions. In IJCAI, Vol. 5. 590–601
2023
-
[18]
Hugh Chen, Scott M Lundberg, and Su-In Lee. 2022. Explaining a series of models by propagating Shapley values. Nat. Commun. 13, 1 (2022), 4512
2022
-
[19]
Lingjiao Chen, Paraschos Koutris, and Arun Kumar. 2019. Towards Model- based Pricing for Machine Learning in a Data Marketplace. In SIGMOD ’19. 1535–1552
2019
-
[20]
Cheng, Qiqi Lu, and Yichuan Deng
Jack C.P. Cheng, Qiqi Lu, and Yichuan Deng. 2016. Analytical review and evaluation of civil information modeling. Autom. Constr. 67 (2016), 31–47
2016
-
[21]
Ziwen Cheng, Yi Liu, Chao Wu, Yongqi Pan, Liushun Zhao, and Cheng Zhu
-
[22]
Carlin Chun Fai Chu and David Po Kin Chan. 2020. Feature Selection Using Approximated High-Order Interaction Components of the Shapley Value for Boosted Tree Classifier. IEEE Access 8 (2020), 112742–112750
2020
-
[23]
Shay Cohen, Gideon Dror, and Eytan Ruppin. 2007. Feature Selection via Coalitional Game Theory. Neural Comput. 19, 7 (2007), 1939–1961
2007
-
[24]
Shay B Cohen, Gideon Dror, and Eytan Ruppin. 2005. Feature selection based on the shapley value. In IJCAI. 1–6
2005
-
[25]
Ben Cottier, Robi Rahman, Loredana Fattorini, Nestor Maslej, and David Owen
-
[26]
Christie Courtnage. 2022. A Systematic Study of Semi-Supervised Learning Based on Shapley Value Data Valuation. https://www.diva-portal.org/smash/ get/diva2:1697410/FULLTEXT01.pdf
2022
-
[27]
Christie Courtnage and Evgueni Smirnov. 2021. Shapley-value data valuation for semi-supervised learning. In DS 2021. 94–108
2021
-
[28]
arXiv (2024)
The rising costs of training frontier AI models. arXiv (2024)
2024
-
[29]
Alahy Ratul, Islam Belmerabet, and Edoardo Serra
Alfredo Cuzzocrea, Qudrat E. Alahy Ratul, Islam Belmerabet, and Edoardo Serra
-
[30]
Konstantinos Demertzis, Lazaros Iliadis, Panagiotis Kikiras, and Elias Pimenidis
-
[31]
Ian Covert and Su-In Lee. 2020. Improving kernelSHAP: Practical shapley value estimation via linear regression. arXiv:2012.01536 (2020)
2020 arXiv
-
[32]
Hongbin Dong, Jing Sun, and Xiaohang Sun. 2021. A multi-objective multi-label feature selection algorithm based on shapley value. Entropy 23, 8 (2021), 1094
2021
-
[33]
Vaidotas Drungilas, Evaldas Vaičiukynas, Linas Ablonskis, and Lina Čeponieṅe
-
[34]
Alexandre Duval and Fragkiskos D Malliaros. 2021. Graphsvx: Shapley value explanations for graph neural networks. In ECML PKDD 2021. 302–318
2021
-
[35]
Zhenan Fan, Huang Fang, Zirui Zhou, Jian Pei, Michael P Friedlander, Changxin Liu, and Yong Zhang. 2022. Improving fairness for data valuation in horizontal federated learning. In ICDE. 2440–2453
2022
-
[36]
Daniel Deutch, Nave Frost, Benny Kimelfeld, and Mikaël Monet. 2022. Comput- ing the Shapley value of facts in query answering. In SIGMOD ’22. 1570–1583
2022
-
[37]
van Rijn, Arlind Kadra, Pieter Gijsbers, Neeratyoy Mallik, Sahithya Ravi, Andreas Mueller, Joaquin Vanschoren, and Frank Hutter
Matthias Feurer, Jan N. van Rijn, Arlind Kadra, Pieter Gijsbers, Neeratyoy Mallik, Sahithya Ravi, Andreas Mueller, Joaquin Vanschoren, and Frank Hutter. 2020. OpenML-Python: an extensible Python API for OpenML. arXiv 1911.02490 (2020). https://arxiv.org/pdf/1911.02490.pdf
2020 arXiv
-
[38]
Philip Hans Franses, Jiahui Zou, and Wendun Wang. 2024. Shapley-value-based forecast combination. Journal of Forecasting 43, 8 (2024), 3194–3202
2024
-
[39]
Applied Sciences 13, 12 (2023)
Shapley Values as a Strategy for Ensemble Weights Estimation. Applied Sciences 13, 12 (2023)
2023
-
[40]
Daniel Fryer, Inga Strümke, and Hien Nguyen. 2021. Shapley values for feature selection: The good, the bad, and the axioms. IEEE Access 9 (2021), 144352– 144360
2021
-
[41]
Felipe Garrido Lucero, Benjamin Heymann, Maxime Vono, Patrick Loiseau, and Vianney Perchet. 2024. Du-shapley: A shapley value proxy for efficient dataset valuation. Advances in Neural Information Processing Systems 37 (2024), 1973–2000
2024
-
[42]
Eitan Farchi, Ramasuri Narayanam, and Lokesh Nagalapatti. 2021. Ranking Data Slices for ML Model Validation: A Shapley Value Approach. In ICDE. 1937–1942
2021
-
[43]
Amirata Ghorbani and James Zou. 2019. Data shapley: Equitable valuation of data for machine learning. In ICML. 2242–2251
2019
-
[44]
Amirata Ghorbani, James Zou, and Andre Esteva. 2022. Data shapley valuation for efficient batch active learning. In ACSSC 2022. 1456–1462
2022
-
[45]
Christopher Frye, Damien de Mijolla, Tom Begley, Laurence Cowton, Megan Stanley, and Ilya Feige. 2021. Shapley explainability on the data manifold. In ICLR
2021
-
[46]
Eberhard Hechler, Maryela Weihrauch, and Yan (Catherine) Wu. 2023. Data Fabric Architecture Patterns. 231–255
2023
-
[47]
Alexandre Heuillet, Fabien Couthouis, and Natalia Díaz-Rodríguez. 2022. Collec- tive eXplainable AI: Explaining Cooperative Strategies and Agent Contribution in Multiagent Reinforcement Learning With Shapley Values. IEEE Computa- tional Intelligence Magazine 17, 1 (2022), 59–71
2022
-
[48]
Amirata Ghorbani, Michael Kim, and James Zou. 2020. A distributional frame- work for data valuation. In ICML. 3535–3544
2020
-
[49]
Jiyue Huang, Rania Talbi, Zilong Zhao, Sara Bouchenak, Lydia Yiyu Chen, and Stefanie Roos. 2020. An Exploratory Analysis on Users’ Contributions in Federated Learning. IEEE TPS 2020 (2020), 20–29
2020
-
[50]
Xuanxiang Huang and Joao Marques-Silva. 2024. On the failings of Shapley values for explainability. International Journal of Approximate Reasoning (2024), 109112
2024
-
[51]
Jiyang Guan, Zhuozhuo Tu, Ran He, and Dacheng Tao. 2022. Few-shot backdoor defense using shapley estimation. In CVPR. 13358–13367
2022
-
[52]
Al-Bakry
Wasnaa Kadhim Jawad and Abbas M. Al-Bakry. 2022. Big Data Analytics: A Survey. IJCI (2022)
2022
-
[53]
Neil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee, and Rajesh Ranganath. 2021. FastSHAP: Real-time shapley value estimation. In ICLR
2021
-
[54]
Jiyue Huang, Chi Hong, Lydia Y Chen, and Stefanie Roos. 2021. Is Shap- ley value fair? Improving client selection for mavericks in federated learning. arXiv:2106.10734 (2021)
2021 arXiv
-
[55]
Nailcan Kara, Yagiz Levent Gume, Umit Tigrak, Gokce Ezeroglu, Serdar Mola, Omer Burak Akgun, and Arzucan Özgür. 2022. A SHAP-based Active Learning Approach for Creating High-Quality Training Data. InIEEE BigData 2022. 4002– 4008
2022
-
[56]
Seo-Hee Kim, Sun Young Park, Hyungseok Seo, and Jiyoung Woo. 2024. Feature selection integrating Shapley values and mutual information in reinforcement learning: An application in the prediction of post-operative outcomes in patients with end-stage renal disease. Computer Meth...
2024
-
[57]
Lukas Huber, Marc Alexander Kühn, Edoardo Mosca, and Georg Groh. 2022. Detecting word-level adversarial text attacks via SHapley additive exPlanations. In RepL4NLP 2022. 156–166
2022
-
[58]
Patrick Kolpaczki, Georg Haselbeck, and Eyke Hüllermeier. 2024. How Much Can Stratification Improve the Approximation of Shapley Values?. InExplainable Artificial Intelligence. Cham, 489–512
2024
-
[59]
Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. Technical report (2009)
2009
-
[60]
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nez- ihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos. 2019. Towards efficient data valuation based on the shapley value. In AISTATS 2019. 1167–1176
2019
-
[61]
Indra Kumar, Carlos Scheidegger, Suresh Venkatasubramanian, and Sorelle Friedler. 2021. Shapley Residuals: Quantifying the limits of the Shapley value for explanations. NeurIPS 34 (2021), 26598–26608
2021
-
[62]
Friedler
Indra Elizabeth Kumar, Carlos Eduardo Scheidegger, Suresh Venkatasubrama- nian, and Sorelle A. Friedler. 2021. Shapley Residuals: Quantifying the limits of the Shapley value for explanations. In NeurIPS
2021
-
[63]
Patrick Kolpaczki, Viktor Bengs, Maximilian Muschalik, and Eyke Hüllermeier
-
[64]
Approximating the shapley value without marginal contributions. In Proc. AAAI Conf. Artif. Intell, Vol. 38. 13246–13255
-
[65]
Yongchan Kwon and James Y. Zou. 2022. WeightedSHAP: analyzing and im- proving Shapley based feature attributions. arXiv abs/2209.13429 (2022)
2022 arXiv
-
[66]
Yann LeCun and Corinna Cortes. 2010. MNIST handwritten digit database. (2010)
2010
-
[67]
H Kuhn and A Tucker. 1951. Nonlinear programming In Proceedings of 2nd Berkeley symposium (pp. 481–492). Berkeley: University of California Press. (1951)
1951
-
[68]
Meng Li, Hengyang Sun, Yanjun Huang, and Hong Chen. 2024. Shapley value: from cooperative game to explainable artificial intelligence. Auton. Intell. Syst. 4 (2024), 2
2024
-
[69]
Xiaoqiang Lin, Xinyi Xu, See-Kiong Ng, Chuan-Sheng Foo, and Bryan Kian Hsiang Low. 2023. Fair yet asymptotically equal collaborative learning. In ICML. 21223–21259
2023
-
[70]
I Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, and Sorelle Friedler. 2020. Problems with Shapley-value-based explanations as feature importance measures. In ICML. 5491–5500
2020
-
[71]
Yongchan Kwon, Manuel A Rivas, and James Zou. 2021. Efficient computation and analysis of distributional shapley values. In AISTATS. 793–801
2021
-
[72]
Zelei Liu, Yuanyuan Chen, Han Yu, Yang Liu, and Lizhen Cui. 2022. Gtg-shapley: Efficient and accurate participant contribution evaluation in federated learning. TIST 13, 4 (2022), 1–21
2022
-
[73]
Zelei Liu, Yuanyuan Chen, Yansong Zhao, Han Yu, Yang Liu, Renyi Bao, Jin- peng Jiang, Zaiqing Nie, Qian Xu, and Qiang Yang. 2022. Contribution-Aware Federated Learning for Smart Healthcare. Proc. AAAI Conf. Artif. Intell 36, 11 (2022), 12396–12404
2022
-
[74]
Jiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu, Long Chen, Fei Wu, and Jun Xiao. 2021. Shapley counterfactual credits for multi-agent reinforcement learning. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 934–942
2021
-
[75]
Ester Livshits, Leopoldo Bertossi, Benny Kimelfeld, and Moshe Sebag. 2021. Query games in databases. SIGMOD 50, 1 (2021), 78–85
2021
-
[76]
Ester Livshits, Leopoldo Bertossi, Benny Kimelfeld, and Moshe Sebag. 2021. The Shapley value of tuples in query answering. LMCS 17 (2021)
2021
-
[77]
Jinfei Liu, Jian Lou, Junxu Liu, Li Xiong, Jian Pei, and Jimeng Sun. 2021. Dealer: an end-to-end model marketplace with differential privacy. VLDB 14, 6 (2021)
2021
-
[78]
Yuan Liu, Zhengpeng Ai, Shuai Sun, Shuangfeng Zhang, Zelei Liu, and Han Yu. 2020. Fedcoin: A peer-to-peer payment system for federated learning. In Federated learning: privacy and incentive . 125–138
2020
-
[79]
Lundberg and Su-In Lee
Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In NeurIPS. 4768–4777
2017
-
[80]
Xinjian Luo, Yangfan Jiang, and X. Xiao. 2022. Feature Inference Attack on Shapley Values. ACM CCS 2022 (2022)
2022
-
[81]
Zhihong Liu, Hoang Anh Just, Xiangyu Chang, Xi Chen, and Ruoxi Jia. 2023. 2D-shapley: a framework for fragmented data valuation. InICML. 21730–21755
2023
-
[82]
Xuan Luo, Jian Pei, Cheng Xu, Wenjie Zhang, and Jianliang Xu. 2024. Fast Shapley Value Computation in Data Assemblage Tasks as Cooperative Simple Games. PACMMOD 2, 1 (2024), 1–28
2024
-
[83]
Shuaicheng Ma, Yang Cao, and Li Xiong. 2021. Transparent Contribution Evaluation for Secure Federated Learning on Blockchain. In ICDEW. 88–91
2021
-
[84]
Scott M Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee. 2020. From local explanations to global understanding with explainable AI for trees. Nat. Mach. Intell 2, 1 (2020), 56–67
2020
-
[85]
Lundberg and Su-In Lee
Scott M. Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions. In NeurIPS. 4768–4777
2017
-
[86]
Wilson E Marcílio and Danilo M Eler. 2020. From explanations to feature selection: assessing SHAP values as feature selection mechanism. In SIBGRAPI. 340–347
2020
-
[87]
Kolby Nottingham Markelle Kelly, Rachel Longjohn. 2025. The UCI Machine Learning Repository. https://archive.ics.uci.edu
2025
-
[88]
Xuan Luo and Jian Pei. 2024. Applications and Computation of the Shapley Value in Databases and Machine Learning. In SIGMOD/PODS ’24. 630–635
2024
-
[89]
Ayesh Meepaganithage, Suman Rath, Mircea Nicolescu, Monica Nicolescu, and Shamik Sengupta. 2024. Feature Selection Using the Advanced Shapley Value. In CCWC. 0207–0213
2024
-
[90]
Luke Merrick and Ankur Taly. 2020. The Explanation Game: Explaining Machine Learning Models Using Shapley Values. In MAKE. 17–38
2020
-
[91]
Sisi Ma and Roshan Tourani. 2020. Predictive and Causal Implications of using Shapley Value for Model Interpretation. In PMLR, Vol. 127. 23–38
2020
-
[92]
Srujana Maddula. 2024. An Introduction to Data Orchestration: Process and Benefits. https://www.datacamp.com/blog/introduction-to-data-orchestration- process-and-benefits
2024
-
[93]
Lokesh Nagalapatti and Ramasuri Narayanam. 2021. Game of gradients: Miti- gating irrelevant clients in federated learning. In Proc. AAAI Conf. Artif. Intell , Vol. 35. 9046–9054
2021
-
[94]
Quoc Phong Nguyen, Bryan Kian Hsiang Low, and Patrick Jaillet. 2022. Trade- off between Payoff and Model Rewards in Shapley-Fair Collaborative Machine Learning. In NeurIPS, Vol. 35. 30542–30553
2022
-
[95]
H. B. McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2016. Communication-Efficient Learning of Deep Networks from Decentralized Data. In AISTATS
2016
-
[96]
Lars H. B. Olsen, Ingrid K. Glad, Martin Jullum, and Kjersti Aas. 2022. Using Shapley Values and Variational Autoencoders to Explain Predictive Models with Dependent Mixed Features. JMLR 23, 213 (2022), 1–51
2022
-
[97]
Lars Henry Berge Olsen, Ingrid Kristine Glad, Martin Jullum, and Kjersti Aas
-
[98]
Rory Mitchell, Joshua Cooper, Eibe Frank, and Geoffrey Holmes. 2022. Sampling permutations for shapley value estimation. JMLR 23, 43 (2022), 1–46
2022
-
[99]
Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, and Eyke Hüller- meier. 2024. Beyond TreeSHAP: Efficient Computation of Any-Order Shapley Interactions for Tree Ensembles. Proc. AAAI Conf. Artif. Intell 38, 13 (2024), 14388–14396
2024
-
[100]
Khaoula Otmani, Rachid Elazouzi, and Vincent Labatut. 2024. FedSV: Byzantine- Robust Federated Learning via Shapley Value. In IEEE ICC
2024
-
[101]
Manisha Padala, Lokesh Nagalapatti, Atharv Tyagi, Ramasuri Narayanam, and Shiv Kumar Saini. 2025. Tab-Shapley: Identifying Top-k Tabular Data Quality Insights. arXiv preprint arXiv:2501.06685 (2025)
2025 arXiv
-
[102]
Ramin Okhrati and Aldo Lipani. 2020. A Multilinear Sampling Algorithm to Estimate Shapley Values. ICPR (2020), 7992–7999
2020
-
[103]
Jian Pei. 2020. A survey on data pricing: from economics to data science. IEEE TKDE 34, 10 (2020), 4586–4608
2020
-
[104]
Guilherme Dean Pelegrina, Miguel Couceiro, and Leonardo Tomazeli Duarte
-
[105]
Data Min
A comparative study of methods for estimating model-agnostic Shapley value explanations. Data Min. Knowl. Discov. 38, 4 (2024), 1782–1829
2024
-
[106]
Ontotext. 2023. What Is Data Fabric? https://www.ontotext.com/ knowledgehub/fundamentals/what-is-data-fabric/
2023
-
[107]
OpenAI. 2023. GPT-4 Technical Report. https://cdn.openai.com/papers/gpt- 4.pdf
2023
-
[108]
Market Research Report. 2024. Big Data Technology Market Size, Share & Industry Analysis. (2024). https://www.fortunebusinessinsights.com/data- analytics-market-108882
2024
-
[109]
Maria Rigaki and Sebastian Garcia. 2023. A survey of privacy attacks in machine learning. Comput. Surveys 56, 4 (2023), 1–34
2023
-
[110]
Konstantin D Pandl, Fabian Feiland, Scott Thiebes, and Ali Sunyaev. 2021. Trustworthy machine learning for health care: scalable data valuation with the shapley value. In ACM CHIL. 47–57
2021
-
[111]
Benedek Rozemberczki, Lauren Watson, Péter Bayer, Hao-Tsung Yang, Olivér Kiss, Sebastian Nilsson, and Rik Sarkar. 2022. The Shapley Value in Machine Learning. In IJCAI-22, Lud De Raedt (Ed.). 5572–5579. Survey Track
2022
-
[112]
Johansson
Franco Ruggeri, William Emanuelsson, Ahmad Terra, Rafia Inam, and Karl H. Johansson. 2024. Rollout-based Shapley Values for Explainable Cooperative Multi-Agent Reinforcement Learning. In 2024 IEEE International Conference on Machine Learning for Communication and Networking (I...
2024
-
[113]
ACM FACCT (2024)
A preprocessing Shapley value-based approach to detect relevant and disparity prone features in machine learning. ACM FACCT (2024)
2024
-
[114]
Guilherme Dean Pelegrina and Sajid Siraj. 2024. Shapley value-based approaches to explain the quality of predictions by classifiers.IEEE Transactions on Artificial Intelligence (2024)
2024
-
[115]
Guilherme Dean Pelegrina, Sajid Siraj, Leonardo Tomazeli Duarte, and Michel Grabisch. 2024. Explaining contributions of features towards unfairness in clas- sifiers: A novel threshold-dependent Shapley value-based approach.Engineering Applications of Artificial Intelligence 13...
2024
-
[116]
Annabelle Redelmeier, Martin Jullum, and Kjersti Aas. 2020. Explaining Pre- dictive Models with Mixed Features Using Shapley Values and Conditional Inference Trees. In MAKE. 117–137
2020
-
[117]
Yong Shi. 2022. Advances in Big Data Analytics: Theory, Algorithms and Practices. Advances in Big Data Analytics (2022)
2022
-
[118]
Yiwei Shi, Qi Zhang, Kevin McAreavey, and Weiru Liu. 2024. Counterfactual shapley values for explaining reinforcement learning. arXiv e-prints (2024), arXiv–2408
2024
-
[119]
Benedek Rozemberczki and Rik Sarkar. 2021. The shapley value of classifiers in ensemble games. In ACM CIKM. 1558–1567
2021
-
[120]
Zhuan Shi, Lan Zhang, Zhenyu Yao, Lingjuan Lyu, Cen Chen, Li Wang, Junhao Wang, and Xiang-Yang Li. 2022. FedFAIM: A Model Performance-based Fair Incentive Mechanism for Federated Learning. IEEE TBD (2022), 1–13
2022
-
[121]
Dongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner, Hyunwoo Kim, and Jongseong Jang. 2021. Online class-incremental continual learning with adver- sarial shapley value. In Proc. AAAI Conf. Artif. Intell , Vol. 35. 9630–9638
2021
-
[122]
Okegbile, and Jun Cai
Mohammadreza Salarbashishahri, Samuel D. Okegbile, and Jun Cai. 2022. A Shapley value-enhanced evaluation technique for effective aggregation in Fed- erated Learning. In FNWF. 88–93
2022
-
[123]
Stephanie Schoch, Haifeng Xu, and Yangfeng Ji. 2022. CS-Shapley: class-wise Shapley values for data valuation in classification. NeurIPS 35 (2022), 34574– 34585
2022
-
[124]
Carlos Sebastián and Carlos E González-Guillén. 2024. A feature selection method based on Shapley values robust for concept shift in regression. Neural. Comput. Appl. (2024), 1–23
2024
-
[125]
Lloyd S Shapley. 1953. A value for n-person games. Contribution to the Theory of Games 2 (1953)
1953
-
[126]
Ingredients
Rachael Hwee Ling Sim, Xinyi Xu, and Bryan Kian Hsiang Low. 2022. Data Valuation in Machine Learning:" Ingredients", Strategies, and Open Challenges. In IJCAI. 5607–5614
2022
-
[127]
Rachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, and Bryan Kian Hsiang Low. 2020. Collaborative Machine Learning with Incentive-Aware Model Rewards. In ICML PMLR, Vol. 119. 8927–8936
2020
-
[128]
Yiwei Shi, Qi Zhang, Kevin McAreavey, and Weiru Liu. 2024. Explaining Rein- forcement Learning: A Counterfactual Shapley Values Approach.arXiv preprint arXiv:2408.02529 (2024)
2024 arXiv
-
[129]
Tianshu Song, Yongxin Tong, and Shuyue Wei. 2019. Profit Allocation for Federated Learning. In IEEE BigData 2019. 2577–2586
2019
-
[130]
Qiheng Sun, Xiang Li, Jiayao Zhang, Li Xiong, Weiran Liu, Jinfei Liu, Zhan Qin, and Kui Ren. 2023. Shapleyfl: Robust federated learning based on shapley value. In ACM SIGKDD. 2096–2108
2023
-
[131]
Seyedamir Shobeiri and Mojtaba Aajami. 2021. Shapley value in convolutional neural networks (CNNs): A Comparative Study. AJMSE 2, 3 (2021), 9–14
2021
-
[132]
Seyedamir Shobeiri and Mojtaba Aajami. 2022. Shapley Value is an Equitable Metric for Data Valuation. IJEEE 18, 2 (2022)
2022
-
[133]
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017. Learn- ing important features through propagating activation differences. In ICML. 3145–3153
2017
-
[134]
Michelle Si and Jian Pei. 2024. Counterfactual Explanation of Shapley Value in Data Coalitions. pVLDB 17, 11 (2024), 3332–3345
2024
-
[135]
Nurbek Tastan, Samar Fares, Toluwani Aremu, Samuel Horvath, and Karthik Nandakumar. 2024. Redefining Contributions: Shapley-Driven Federated Learn- ing. arXiv:2406.00569 (2024)
2024 arXiv
-
[136]
Yingjie Tian, Yurong Ding, Saiji Fu, and Dalian Liu. 2022. Data boundary and data pricing based on the shapley value. IEEE Access 10 (2022), 14288–14300
2022
-
[137]
Pranava Singhal, Shashi Raj Pandey, and Petar Popovski. 2024. Greedy Shap- ley Client Selection for Communication-Efficient Federated Learning. IEEE Networking Letters 6, 2 (2024), 134–138
2024
-
[138]
Sandhya Tripathi, N Hemachandra, and Prashant Trivedi. 2020. Interpretable feature subset selection: A Shapley value based approach. In IEEE BigData 2020. 5463–5472
2020
-
[139]
van Rijn, Bernd Bischl, and Luis Torgo
Joaquin Vanschoren, Jan N. van Rijn, Bernd Bischl, and Luis Torgo. 2013. OpenML: networked science in machine learning. SIGKDD Explorations 15, 2 (2013), 49–60. https://doi.org/10.1145/2641190.2641198
2013
-
[140]
Qiheng Sun, Jiayao Zhang, Jinfei Liu, Li Xiong, Jian Pei, and Kui Ren. 2024. Shapley Value Approximation Based on Complementary Contribution. IEEE Transactions on Knowledge and Data Engineering 36, 12 (2024), 9263–9281
2024
-
[141]
Mukund Sundararajan and Amir Najmi. 2020. The many Shapley values for model explanation. In ICML. 9269–9278
2020
-
[142]
Zuoqi Tang, Zheqi Lv, and Chao Wu. 2020. A Brief SURVEY OF DATA PRICING FOR MACHINE LEARNING. In CS & IT Conference Proceedings , Vol. 10
2020
-
[143]
Zuoqi Tang, Feifei Shao, Long Chen, Yunan Ye, Chao Wu, and Jun Xiao. 2021. Optimizing Federated Learning on Non-IID Data Using Local Shapley Value. In Artif. Intell. 164–175
2021
-
[144]
Jiachen T Wang, Yuqing Zhu, Yu-Xiang Wang, Ruoxi Jia, and Prateek Mittal
-
[145]
Rui Wang, Xiaoqian Wang, and David I. Inouye. 2021. Shapley Explanation Networks. In ICLR
2021
-
[146]
Zhihua Tian, Jian Liu, Jingyu Li, Xinle Cao, Ruoxi Jia, Jun Kong, Mengdi Liu, and Kui Ren. 2022. Private data valuation and fair payment in data marketplaces. arXiv:2210.08723 (2022)
2022 arXiv
-
[147]
Yong Wang, Kaiyu Li, Yuyu Luo, Guoliang Li, Yunyan Guo, and Zhuo Wang
-
[148]
Ziming Wang, Changwu Huang, Yun Li, and Xin Yao. 2024. Multi-objective feature attribution explanation for explainable machine learning. ACM Trans- actions on Evolutionary Learning and Optimization 4, 1 (2024), 1–32
2024
-
[149]
Fangdi Wang, Jiaqi Jin, Jingtao Hu, Suyuan Liu, Xihong Yang, Siwei Wang, Xinwang Liu, and En Zhu. 2024. Evaluate then Cooperate: Shapley-based View Cooperation Enhancement for Multi-view Clustering. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems...
2024
-
[150]
Guan Wang, Charlie Xiaoqian Dang, and Ziye Zhou. 2019. Measure contribution of participants in federated learning. In IEEE BigData 2019. 2597–2604
2019
-
[151]
Junhao Wang, Lan Zhang, Anran Li, Xuanke You, and Haoran Cheng. 2022. Effi- cient Participant Contribution Evaluation for Horizontal and Vertical Federated Learning. In ICDE. 911–923
2022
-
[152]
Jiachen T Wang, Tianji Yang, James Zou, Yongchan Kwon, and Ruoxi Jia. 2024. Rethinking data shapley for data selection tasks: Misleads and merits. arXiv preprint arXiv:2405.03875 (2024)
2024 arXiv
-
[153]
Wikipedia. 2025. Law of Large Numbers. https://en.wikipedia.org/wiki/Law_ of_large_numbers
2025
-
[154]
arXiv:2308.15709 (2023)
Threshold knn-shapley: A linear-time and privacy-friendly approach to data valuation. arXiv:2308.15709 (2023)
2023 arXiv
-
[155]
Chengqian Wu, Xuemei Fu, Xiangli Yang, Ruonan Zhao, Qidong Wu, and Tinghua Zhang. 2023. CP-Decomposition Based Federated Learning with Shap- ley Value Aggregation. In ICPADS. 571–577
2023
-
[156]
Tianhao Wang, Johannes Rausch, Ce Zhang, Ruoxi Jia, and Dawn Song. 2020. A principled approach to data valuation for federated learning.Federated Learning: Privacy and Incentive (2020), 153–167
2020
-
[157]
Binhan Xi, Shaofeng Li, Jiachun Li, Hui Liu, Hong Liu, and Haojin Zhu. 2021. BatFL: Backdoor Detection on Federated Learning in e-Health. InIWQOS. 1–10
2021
-
[158]
Fast, Robust and Interpretable Participant Contribution Estimation for Federated Learning. In ICDE. 2298–2311
-
[159]
Haocheng Xia, Jinfei Liu, Jian Lou, Zhan Qin, Kui Ren, Yang Cao, and Li Xiong
-
[160]
Zexin Wang, Biwei Yan, and Anming Dong. 2022. Blockchain Empowered Federated Learning for Data Sharing Incentive Mechanism. Procedia Com- puter Science 202 (2022), 348–353. International Conference on Identification, Information and Knowledge in the internet of Things, 2021
2022
-
[161]
David Watson, Joshua O' Hara, Niek Tax, Richard Mudd, and Ido Guy. 2023. Explaining Predictive Uncertainty with Information Theoretic Shapley Values. In NeurIPS, Vol. 36. 7330–7350
2023
-
[162]
Lauren Watson, Rayna Andreeva, Hao Yang, and Rik Sarkar. 2022. Differentially Private Shapley Values for Data Evaluation. arXiv abs/2206.00511 (2022)
2022 arXiv
-
[163]
Shuyue Wei, Yongxin Tong, Zimu Zhou, and Tianshu Song. 2020. Efficient and fair data valuation for horizontal federated learning. Federated Learning: Privacy and Incentive (2020), 139–152
2020
-
[164]
Chengyi Yang, Zhaoxiang Hou, Sheng Guo, Hui Chen, and Zengxiang Li. 2023. SWATM: Contribution-Aware Adaptive Federated Learning Framework Based on Augmented Shapley Values. In ICME. 672–677
2023
-
[165]
Brian Williamson and Jean Feng. 2020. Efficient nonparametric statistical inference on population feature importance using Shapley values. In ICML PMLR, Vol. 119. 10282–10291
2020
-
[166]
Jilei Yang. 2021. Fast TreeSHAP: Accelerating SHAP Value Computation for Trees. arXiv abs/2109.09847 (2021)
2021 arXiv
-
[167]
Mengmeng Wu, Ruoxi Jia, Changle Lin, Wei Huang, and Xiangyu Chang. 2023. Variance reduced Shapley value estimation for trustworthy data valuation.COR 159 (2023), 106305
2023
-
[168]
Xun Yang, Shuwen Xiang, Changgen Peng, Weijie Tan, Zhuguo Li, Ningbo Wu, and Yan Zhou. 2023. Federated Learning Incentive Mechanism Design via Shapley Value and Pareto Optimality. Axioms 12 (2023), 636
2023
-
[169]
Haocheng Xia, Xiang Li, Junyuan Pang, Jinfei Liu, Kui Ren, and Li Xiong. 2024. P-Shapley: Shapley Values on Probabilistic Classifiers. VLDB 17, 7 (2024), 1737– 1750
2024
-
[170]
Dingze Yin, Dan Chen, Yunbo Tang, Heyou Dong, and Xiaoli Li. 2022. Adaptive feature selection with shapley and hypothetical testing: Case study of EEG feature engineering. INS 586 (2022), 374–390
2022
-
[171]
VLDB 16, 11 (2023), 3349–3362
Equitable data valuation meets the right to be forgotten in model markets. VLDB 16, 11 (2023), 3349–3362
2023
-
[172]
Haocheng Xia, Jiayao Zhang, Qiheng Sun, Jinfei Liu, Kui Ren, Li Xiong, and Jian Pei. 2025. Computing Shapley Values for Dynamic Data. IEEE Transactions on Knowledge and Data Engineering (2025)
2025
-
[173]
Lei Xu, Jiaqing Chen, Shan Chang, Cong Wang, and Bo Li. 2023. Toward Quality-aware Data Valuation in Learning Algorithms: Practices, Challenges, and Beyond. IEEE Network (2023), 1–1
2023
-
[174]
Xinyi Xu, Thanh Lam, Chuan Sheng Foo, and Bryan Kian Hsiang Low. 2024. Model Shapley: Equitable Model Valuation with Black-box Access. NeurIPS 36 (2024)
2024
-
[175]
Xinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao, Chuan Sheng Foo, and Bryan Kian Hsiang Low. 2021. Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine Learning. In NeurIPS, Vol. 34. 16104–16117
2021
-
[176]
Daniel Yue Zhang, Ziyi Kou, and Dong Wang. 2020. FairFL: A Fair Feder- ated Learning Approach to Reducing Demographic Bias in Privacy-Sensitive Classification Models. In IEEE BigData 2020. 1051–1060
2020
-
[177]
Chengyi Yang, Jia Liu, Hao Sun, Tongzhi Li, and Zengxiang Li. 2022. WTDP- Shapley: Efficient and effective incentive mechanism in federated learning for intelligent safety inspection. IEEE TBD (2022)
2022
-
[178]
Jiayao Zhang, Haocheng Xia, Qiheng Sun, Jinfei Liu, Li Xiong, Jian Pei, and Kui Ren. 2023. Dynamic shapley value computation. In ICDE. 639–652
2023
-
[179]
Xun Yang, Weijie Tan, Changgen Peng, Shuwen Xiang, Kun Niu, et al. 2022. Federated learning incentive mechanism design via enhanced shapley value method. Wireless Communications and Mobile Computing 2022 (2022)
2022
-
[180]
Quan Zheng, Ziwei Wang, Jie Zhou, and Jiwen Lu. 2022. Shap-CAM: Visual explanations for convolutional neural networks based on Shapley value. In ECCV. 459–474
2022
-
[181]
Xun Yang, Shuwen Xiang, Changgen Peng, Weijie Tan, Yue Wang, Hai Liu, and Hongfa Ding. 2024. Federated Learning Incentive Mechanism with Supervised Fuzzy Shapley Value. Axioms 13, 4 (2024), 254
2024
-
[182]
Haolin Zhu, Ziye Li, Dingzhi Zhong, Cheng Li, and Yong Yuan. 2023. Shapley- value-based Contribution Evaluation in Federated Learning: A Survey. 2023 IEEE 3rd International Conference on Digital Twins and Parallel Intelligence (DTPI) (2023), 1–5
2023
-
[183]
Zhaoyang You, Xinya Wu, Kexuan Chen, Xinyi Liu, and Chao Wu. 2021. Evalu- ate the Contribution of Multiple Participants in Federated Learning. In DEXA. 189–194
2021
-
[184]
Han Yu, Zelei Liu, Yang Liu, Tianjian Chen, Mingshu Cong, Xi Weng, Dusit Niyato, and Qiang Yang. 2020. A Fairness-aware Incentive Scheme for Federated Learning. In AIES ’20. 393–399
2020
-
[185]
Han Yu, Zelei Liu, Yang Liu, Tianjian Chen, Mingshu Cong, Xi Weng, Dusit Niyato, and Qiang Yang. 2020. A Sustainable Incentive Scheme for Federated Learning. IEEE Intelligent Systems 35, 4 (2020), 58–69
2020
-
[186]
Peng Yu, Albert Bifet, Jesse Read, and Chao Xu. 2022. Linear tree shap. In NeurIPS
2022
-
[187]
Pee, Shan L
Dan Zhang, L.G. Pee, Shan L. Pan, and Lili Cui. 2022. Big data analytics, resource orchestration, and digital sustainability: A case study of smart city development. Government Information Quarterly 39, 1 (2022), 101626
2022
-
[189]
Jiayao Zhang, Qiheng Sun, Jinfei Liu, Li Xiong, Jian Pei, and Kui Ren. 2023. Efficient sampling approaches to shapley value approximation. PACMMOD 1, 1 (2023), 1–24
2023
-
[191]
Ningsheng Zhao, Jia Yuan Yu, Krzysztof Dzieciolowski, and Trang Bui. 2024. Error Analysis of Shapley Value-Based Model Explanations: An Informative Perspective. In International Symposium on AI Verification . 29–48
2024
-
[193]
Shuyuan Zheng, Yang Cao, and Masatoshi Yoshikawa. 2022. Secure Shapley Value for Cross-Silo Federated Learning (Technical Report). arXiv:2209.04856 (2022)
2022 arXiv
-
[195]
Ye Zhu, Zhiqiang Liu, Peng Wang, and Chenglie Du. 2023. A dynamic incentive and reputation mechanism for energy-efficient federated learning in 6G.Digital Communications and Networks 9, 4 (2023), 817–826
2023
-
[2021]
In KDD ’21
TimeSHAP: Explaining Recurrent Models through Sequence Perturbations. In KDD ’21
-
[2022]
An explainable semi-personalized federated learning model. Integr. Comput.-Aided Eng. 29, 4 (2022), 335–350
2022
-
[2023]
Attribution Methods Assessment for Interpretable Machine Learning. In SEBD
-
[2024]
PoShapley-BCFL: A Fair and Robust Decentralized Federated Learning Based on Blockchain and the Proof of Shapley-Value. In NIPS. 531–549
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.