REVIEW 2 major objections 4 minor 68 references
Counterfactual Explanation of Shapley Value in Data Coalitions
T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proves that the Shapley-value gap between two data owners always admits a smallest record-transfer that flips the ranking, that finding it is NP-hard, and that the SV-Exp algorithm approximates it.
desk verdict New problem and a useful heuristic, but the NP-hardness proof uses utilities that violate the paper's own monotonicity assumption, so the theory as stated is not airtight. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing identity is the differential Shapley value, $\Psi_O(A,B) = \psi_O(A) - \psi_O(B)$, which Theorem 3 expresses as a weighted sum over coalitions $S$ containing neither $A$ nor $B$ of the utility differences $U(S \cup \{A\}) - U(S \cup \{B\})$. This identity turns a two-owner comparison into a single object, permits an unbiased Monte Carlo estimate from random permutations (Corollary 2), and defines the power of a data entry $x$: the expected change in the differential when $x$ is moved from $A$ to $B$. SV-Exp uses Thompson sampling to pick the entry currently estimated to have the largest power, moves that entry, and repeats until the estimated differential goes negative.
What would settle it
Enumerate all monotone utility functions over a small universe of data entries (say, all non-decreasing set functions on four entries) and check, by exact Shapley computation, every pair $A,B$ with $\psi(A) > \psi(B)$ to see whether a transfer $\Delta_A$ that flips the inequality always exists; finding one monotone instance with no flip would refute Proposition 2. The same search can test whether the NP-hardness reduction's survival of the monotone assumption matters, since that reduction's own utility is non-monotone.
Extended reading notes
Core claim
The central claim is that the difference in Shapley value between two data owners $A$ and $B$, with $\psi(A) > \psi(B)$, can be explained counterfactually: there is always at least one subset $\Delta_A \subseteq A$ whose transfer to $B$ makes $\psi(A \setminus \Delta_A) < \psi(B \cup \Delta_A)$, and the smallest such subset is a well-defined but NP-hard optimization problem. The existence proof rests on monotonicity, because moving all of $A$ empties it to Shapley value $0$ while $B$ gains $A$'s records and remains positive. The NP-hardness proof reduces the set cover problem to the search for a minimum flipping subset. To make the problem tractable, the paper proves that the differential Shapley value $\Psi(A,B) = \psi(A) - \psi(B)$ can be written as a weighted sum over coalitions containing neither owner, gives an unbiased permutation-based Monte Carlo estimator for it, and introduces the power of a data entry—the estimated drop in the differential when that record moves—so that a greedy iteration can assemble an approximate counterfactual. The resulting SV-Exp algorithm is demonstrated to beat a Monte Carlo subset-search baseline in runtime and flip success rate, and its outputs are shown to behave as feature selection and distribution-difference detectors.
Load-bearing premise
The load-bearing premise is that the utility function is monotone—adding more data never lowers a coalition's utility—and that is what guarantees a counterfactual explanation always exists; the paper's own NP-hardness reduction uses a utility that drops when a redundant element is added to a cover, and the experimental utilities are raw model errors, so outside the monotone regime feasibility and the interpretation of SV-Exp's outputs are not theoretically supported.
Editorial extensions
If this is right
- Under a monotone utility, every pair of unequal owners has at least one feasible counterfactual explanation, so the search is over how small the transfer can be, not whether one exists.
- Because exact minimization is NP-hard, practical deployments must approximate, and SV-Exp supplies a greedy sampling-based candidate that the experiments show flips rankings faster than a Monte Carlo subset-search baseline.
- Estimating the differential Shapley value directly removes the need to estimate two Shapley values separately and subtract them, reducing sampling cost and error.
- The size of a counterfactual carries semantic content: a small set of records signals that a few 'star' records dominate an owner's advantage, while a large set signals a distributed advantage.
- The same machinery can be applied as a feature-selection heuristic and as a detector of distributional differences between data owners, as the two case studies demonstrate.
Reading between the lines
- Beyond the paper: if monotonicity fails—as it does for raw model-error utilities—the existence guarantee collapses, so SV-Exp's outputs on such utilities should be read as empirical heuristics rather than theorem-backed explanations.
- Beyond the paper: records that repeatedly appear in small counterfactuals are candidate 'star' rows, so the power ranking could be reused as a data-pricing signal in markets that want to charge more for high-marginal-value data.
- Beyond the paper: the greedy transfer framework extends to deletion-only variants (the paper notes this) and to other cooperative valuation schemes such as beta Shapley, but each extension requires re-deriving the differential identity.
- Beyond the paper: the case study's finding that a less-correlated feature is chosen over a more-correlated one predicts that counterfactual explanations act as a multicollinearity-aware selection rule, a claim testable on synthetic data with controlled correlations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the counterfactual explanation of the Shapley value in data coalitions: given two data owners A and B with psi(A) > psi(B), it seeks a smallest subset of A's data whose transfer to B reverses the inequality. The paper claims that such an explanation always exists under monotone utilities, that finding the exact explanation is NP-hard, and it develops a differential-Shapley estimator and a greedy algorithm (SV-Exp) with Thompson sampling to approximate the explanation. Experiments on three datasets compare runtime and success rates against a Monte Carlo baseline and include two case studies on feature selection and distributional differences.
Significance. The problem is novel and practically motivated, and the differential-Shapley machinery in Theorems 3--5 and Corollaries 1--2 is a clean technical contribution: the direct estimators of the difference between two owners' Shapley values are unbiased and avoid estimating each value separately. The two case studies show that the proposed counterfactual can yield interpretable insights. However, the paper's central theoretical claim is not coherent as written: the NP-hardness reduction in Theorem 1 uses a non-monotone utility, contradicting the paper's own standing monotonicity assumption, and the experimental utilities are also non-monotone. These issues are load-bearing because feasibility (Proposition 2) and the intended problem setting both depend on monotonicity.
major comments (2)
- [Section 2, Theorem 1] The NP-hardness reduction in Theorem 1 constructs a utility U(S) = 0 for non-covers and U(S) = m - |S| + f(S) for covers. This utility is not monotone: for a cover S and a redundant element x, U(S union {x}) = m - |S| - 1 + f(S) + 2^{i_x}/2^{m+1} < m - |S| + f(S) = U(S), since 2^{i_x}/2^{m+1} <= 1/2. This directly contradicts the monotonicity assumption stated immediately before the theorem ('we from now on assume that the utility function is monotonic'). Consequently, the reduction establishes hardness only for a larger class of arbitrary utilities, not for the monotone problem whose feasibility is guaranteed by Proposition 2. The theorem should be reworked for the monotone setting (for example, by reasoning about the monotone closure U*(D) = max_{D' subset of D} U(D') mentioned in Section 2) or the paper must explicitly restrict the NP-hardness claim to non-monotone utilities and qualify the abstract accordingly.
- [Section 4.1 (with Section 2)] The experimental utilities are defined as eta minus test error, log loss, or MSE. Such utilities are not monotone in general, and the experiments do not replace them by the monotone closure U*(D) = max_{D' subset of D} U(D') that Section 2 invokes to justify the monotonicity assumption. Therefore the feasibility guarantee of Proposition 2 and the interpretation of the outputs as minimal transferring subsets are not theoretically supported in the evaluation. Moreover, the success-rate measurements in Table 4 and Figure 5 use a Monte Carlo estimate of the final differential Shapley value; under a non-monotone utility a true feasible counterfactual may not even exist, so a 'failure' can reflect nonexistence rather than algorithmic error. The experiments should either use utility functions that are monotone (or explicitly use the monotone closure) or be reframed as heuristic evidence for arbitrary utilities without invoking Proposition 2.
minor comments (4)
- [Section 2, Algorithm 1] Algorithm 1 loops over i = 1 to |A|-1 and has no fallback return. Since Proposition 2 guarantees that Delta_A = A is always feasible, the pseudocode should end with an explicit 'return A', as Algorithm 2 does.
- [Section 3.4, Algorithm 3] Algorithm 3 can exit its while loop without returning a counterfactual (for example, if the estimated differential d is converged and positive). The intended behavior should be clarified, and the convergence criterion for d should be defined precisely rather than left as 'converged'.
- [Section 2, Proposition 2] The proof of Proposition 2 asserts that because B union Delta_A is nonempty, its Shapley value is positive. Nonemptiness alone does not imply a positive Shapley value under monotonicity, since a nonempty owner can still be a dummy; the argument should instead use the fact that psi(A) > 0 ensures some positive marginal contribution from A's data, which then implies a positive marginal for the enlarged owner.
- [Throughout] There are several typographical errors: 'trail' should be 'trial' in Section 4.2, 'Wassertein' should be 'Wasserstein' in Section 4.7, and 'nueral networks' in reference [8] should be 'neural networks'.
Circularity Check
No derivation circularity; one minor self-referential success check in Section 4.5; Theorem 1 assumption mismatch is a correctness concern, not circularity.
-
other
[Section 4.5 (success rates) and Algorithm 3 Phase 2 (stopping rule)]
"after each trial, ψ(A\ΔA)−ψ(B∪ΔA) was estimated using Monte Carlo, where ΔA is the answer generated by MC or SV-Exp. If the difference was negative, the trial was marked a success (Sec. 4.5); Algorithm 3 returns only when its own Monte Carlo estimate d from Corollary 2 is negative."
For non-tiny datasets, the reported success rate is evaluated using the same Monte Carlo estimator of the differential Shapley value that the algorithms use as their termination test. SV-Exp returns ΔA only after its Monte Carlo estimate of d is negative, and Section 4.5 then marks a trial successful when a fresh Monte Carlo estimate of the same quantity is negative. The success numbers are therefore not an independent check that the true Shapley inequality is flipped; they are a re-run of the algorithm's stopping rule. This is a mild self-referential validation rather than a fitted parameter, and the paper mitigates it by explicitly acknowledging that the measure is approximate and by providing exact brute-force comparisons on tiny datasets (Tables 2–3).
full rationale
The central derivation chain is not circular. Theorem 3 is proven from the Shapley-value definition inside the paper; the reference to Jia et al. is contextual, not load-bearing. Corollaries 1–2 and Theorems 4–5 are algebraically derived unbiased estimators, and no parameter is fitted to force the outputs. Algorithm 1 is an exact enumerative solver, and Algorithms 2–3 are unbiased/heuristic approximations whose outputs are checked against brute-force exact Shapley computation on tiny datasets (Tables 2–3), which is an external ground truth. The only circular-adjacent element is the large-dataset success-rate protocol in Section 4.5: success is declared when a fresh Monte Carlo estimate of ψ(A\ΔA)−ψ(B∪ΔA) is negative, and the algorithms themselves stop exactly when their own Monte Carlo estimate of that same quantity is negative. This makes the success numbers partially self-referential, not an independent verification of the true inequality. The paper explicitly acknowledges that its success measure is also an approximation, and it does provide exact validation on small data, so the issue is minor and does not affect the correctness of the differential-Shapley derivation. Separately, and outside circularity: Theorem 1's NP-hardness reduction uses U(S)=m−|S|+f(S), which violates the Section 2 monotonicity assumption used by Proposition 2; that is a scope/correctness gap, not an input-output equivalence, and does not change the circularity score.
Assumptions & free parameters
free parameters (4)
- δ confidence threshold =
0.95 (example)
- ε confidence interval width threshold =
0.01 (example)
- number of Monte Carlo permutations per batch =
not reported
- η utility offset =
20 (case study), otherwise 'sufficiently large'
assumptions (4)
- domain assumption Utility function is monotonic
- standard math Shapley value is the valuation measure
- standard math Set cover is NP-hard
- domain assumption Data entries are atomic and transferable between owners
Cite this review
Pith. "Pith review of Counterfactual Explanation of Shapley Value in Data Coalitions." pith.science (2026). https://pith.science/paper/HYAIFM6J
@misc{pith2026250701267,
author = {Pith},
title = {Pith review of: Counterfactual Explanation of Shapley Value in Data Coalitions},
year = {2026},
howpublished = {\url{https://pith.science/paper/HYAIFM6J}},
note = {Machine review of arXiv:2507.01267}
}
abstract
The Shapley value is widely used for data valuation in data markets. However, explaining the Shapley value of an owner in a data coalition is an unexplored and challenging task. To tackle this, we formulate the problem of finding the counterfactual explanation of Shapley value in data coalitions. Essentially, given two data owners $A$ and $B$ such that $A$ has a higher Shapley value than $B$, a counterfactual explanation is a smallest subset of data entries in $A$ such that transferring the subset from $A$ to $B$ makes the Shapley value of $A$ less than that of $B$. We show that counterfactual explanations always exist, but finding an exact counterfactual explanation is NP-hard. Using Monte Carlo estimation to approximate counterfactual explanations directly according to the definition is still very costly, since we have to estimate the Shapley values of owners $A$ and $B$ after each possible subset shift. We develop a series of heuristic techniques to speed up computation by estimating differential Shapley values, computing the power of singular data entries, and shifting subsets greedily, culminating in the SV-Exp algorithm. Our experimental results on real datasets clearly demonstrate the efficiency of our method and the effectiveness of counterfactuals in interpreting the Shapley value of an owner.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Alessandro Acquisti, Curtis Taylor, and Liad Wagman. 2016. The Economics of Privacy. Journal of Economic Literature 54, 2 (June 2016), 442–92. https: //doi.org/10.1257/jel.54.2.442
-
[2]
Anish Agarwal, Munther Dahleh, and Tuhin Sarkar. 2019. A Marketplace for Data: An Algorithmic Solution. InProceedings of the 2019 ACM Conference on Economics and Computation (Phoenix, AZ, USA) (EC’19). Association for Computing Ma- chinery, New York, NY, USA, 701–726. https://doi.org/10.1145/3328526.3329589
arXiv 2019
- [3]
-
[4]
Charu C. Aggarwal and Philip S. Yu. 2008. Privacy-Preserving Data Mining: A Survey. In Handbook of Database Security: Applications and Trends, Michael Gertz and Sushil Jajodia (Eds.). Springer US, Boston, MA, 431–460. https://doi.org/10. 1007/978-0-387-48533-1_18
work page 2008
-
[5]
William Aiello, Yuval Ishai, and Omer Reingold. 2001. Priced Oblivious Transfer: How to Sell Digital Goods. In Advances in Cryptology - EUROCRYPT 2001, Inter- national Conference on the Theory and Application of Cryptographic Techniques, Innsbruck, Austria, May 6-10, 2001, Proceeding (Lecture Notes in Computer Science) , Vol. 2045. Springer, 119–135. http...
-
[6]
Arjun R Akula, Shuai Wang, and Song-Chun Zhu. 2020. CoCoX: Generating Conceptual and Counterfactual Explanations via Fault-Lines.. In AAAI. 2594– 2601
work page 2020
-
[7]
Nuno Antonio, Ana de Almeida, and Luis Nunes. 2019. Hotel booking demand datasets. Data in Brief 22 (2019), 41–49. https://doi.org/10.1016/j.dib.2018.11.126
-
[8]
Mohit Bajaj, Lingyang Chu, Zi Yu Xue, Jian Pei, Lanjun Wang, Peter Cho-Ho Lam, and Yong Zhang. 2021. Robust Counterfactual Explanations on Graph Neural Networks. In Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 5644–5655. https://proc...
work page 2021
Show all 68 references
-
[9]
Magdalena Balazinska, Bill Howe, and Dan Suciu. 2011. Data Markets in the Cloud: An Opportunity for the Database Community. PVLDB 4, 12 (2011), 1482–
2011
-
[10]
Elisa Bertino, Dan Lin, and Wei Jiang. 2008. A Survey of Quantification of Privacy Preserving Data Mining Algorithms. In Privacy-Preserving Data Mining: Models and Algorithms, Charu C. Aggarwal and Philip S. Yu (Eds.). Springer US, Boston, MA, 183–205. https://doi.org/10.1007/...
2008 doi
-
[11]
Lingjiao Chen, Paraschos Koutris, and Arun Kumar. 2019. Towards Model- based Pricing for Machine Learning in a Data Marketplace. In Proceedings of the 2019 International Conference on Management of Data, SIGMOD Conference 2019, Amsterdam, The Netherlands, June 30 - July 5, 201...
2019
-
[12]
Zicun Cong, Lingyang Chu, Yu Yang, and Jian Pei. 2021. Comprehensible Coun- terfactual Explanation on Kolmogorov-Smirnov Test. Proc. VLDB Endow. 14, 9 (oct 2021), 1583?1596. https://doi.org/10.14778/3461535.3461546
2021
-
[13]
David Dao, Dan Alistarh, Claudiu Musat, and Ce Zhang. 2018. DataBright: Towards a Global Exchange for Decentralized Data Ownership and Trusted Computation. CoRR abs/1802.04780 (2018). arXiv:1802.04780 http://arxiv.org/ abs/1802.04780
2018 arXiv
-
[14]
Papadimitriou
Xiaotie Deng and Christos H. Papadimitriou. 1994. On the Com- plexity of Cooperative Solution Concepts. Mathematics of Operations Research 19, 2 (1994), 257–266. https://doi.org/10.1287/moor.19.2.257 arXiv:https://doi.org/10.1287/moor.19.2.257
1994 doi
-
[15]
Daniel Deutch, Nave Frost, Benny Kimelfeld, and Mikaël Monet. 2022. Comput- ing the Shapley Value of Facts in Query Answering. In Proceedings of the 2022 International Conference on Management of Data (Philadelphia, PA, USA) (SIG- MOD ’22). Association for Computing Machinery,...
2022
-
[16]
Cynthia Dwork. 2008. Differential Privacy: A Survey of Results. In Theory and Applications of Models of Computation, Manindra Agrawal, Dingzhu Du, Zhenhua Duan, and Angsheng Li (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 1–19
2008
-
[17]
Franklin
Raul Castro Fernandez, Pranav Subramaniam, and Michael J. Franklin. 2020. Data Market Platforms: Trading Data Assets to Solve Data Problems. Proc. VLDB Endow. 13, 12 (sep 2020), 1933–1947. https://doi.org/10.14778/3407790.3407800
2020
-
[18]
M. A. Ferrag, L. Maglaras, and A. Ahmim. 2017. Privacy-Preserving Schemes for Ad Hoc Social Networks: A Survey. IEEE Communications Surveys Tutorials 19, 4 (2017), 3015–3045
2017
-
[19]
Fleischer and Yu-Han Lyu
Lisa K. Fleischer and Yu-Han Lyu. 2012. Approximately Optimal Auctions for Selling Privacy When Costs Are Correlated with Data. In Proceedings of the 13th ACM Conference on Electronic Commerce (Valencia, Spain) (EC’12). Association for Computing Machinery, New York, NY, USA, 5...
2012
-
[20]
Ruth C Fong and Andrea Vedaldi. 2017. Interpretable explanations of black boxes by meaningful perturbation. In Proceedings of the IEEE International Conference on Computer Vision. 3429–3437
2017
-
[21]
Benjamin C. M. Fung, Ke Wang, Rui Chen, and Philip S. Yu. 2010. Privacy- Preserving Data Publishing: A Survey of Recent Developments. ACM Comput. Surv. 42, 4, Article 14 (June 2010), 53 pages. https://doi.org/10.1145/1749603. 1749605
2010 doi
-
[22]
Arpita Ghosh, Katrina Ligett, Aaron Roth, and Grant Schoenebeck. 2014. Buying Private Data without Verification. In Proceedings of the Fifteenth ACM Conference on Economics and Computation (Palo Alto, California, USA) (EC’14). Association for Computing Machinery, New York, NY,...
2014
-
[23]
Goldberg, Jason D
Andrew V. Goldberg, Jason D. Hartline, and Andrew Wright. 2001. Competitive Auctions and Digital Goods. In Proceedings of the Twelfth Annual ACM-SIAM Symposium on Discrete Algorithms (Washington, D.C., USA) (SODA’01). Society for Industrial and Applied Mathematics, USA, 735–744
2001
-
[24]
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. MIT press
2016
-
[25]
David Harrison and Daniel L Rubinfeld. 1978. Hedonic housing prices and the demand for clean air. Journal of Environmental Economics and Management 5, 1 (1978), 81–102. https://doi.org/10.1016/0095-0696(78)90006-2
1978 doi
-
[26]
Miguel A Hernán and James M Robins. 2010. Causal inference
2010
-
[27]
Nick Hynes, David Dao, David Yan, Raymond Cheng, and Dawn Song. 2018. A Demonstration of Sterling: A Privacy-Preserving Data Marketplace. Proc. VLDB Endow. 11, 12 (Aug. 2018), 2086–2089. https://doi.org/10.14778/3229863.3236266
2018
-
[28]
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J. Spanos. 2019. Towards Efficient Data Valuation Based on the Shapley Value. In Proceedings of the Twenty-Second International Conference on Artificial ...
2019
-
[29]
Michael I Jordan and Tom M Mitchell. 2015. Machine learning: Trends, perspec- tives, and prospects. Science 349, 6245 (2015), 255–260
2015
-
[30]
Richard M. Karp. 1972. Reducibility among Combinatorial Problems . Springer US, Boston, MA, 85–103. https://doi.org/10.1007/978-1-4684-2001-2_9
1972 doi
-
[31]
Javen Kennedy, Pranav Subramaniam, Sainyam Galhotra, and Raul Castro Fernan- dez. 2022. Revisiting Online Data Markets in 2022: A Seller and Buyer Perspective. SIGMOD Rec. 51, 3 (nov 2022), 30–37. https://doi.org/10.1145/3572751.3572757
2022
-
[32]
Yongchan Kwon and James Zou. 2022. Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning. In Proceedings of The 25th International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research), Gustau Camps-Va...
2022
-
[33]
Yongchan Kwon and James Zou. 2023. Data-OOB: Out-of-Bag Estimate as a Simple and Efficient Data Value. InProceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). JMLR.org, Article 749, 18 pages
2023
-
[34]
Thai Le, Suhang Wang, and Dongwon Lee. 2020. GRACE: Generating Concise and Informative Contrastive Sample to Explain Neural Network Model’s Prediction. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) ...
2020
-
[35]
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436–444
2015
-
[36]
Chao Li, Daniel Yang Li, Gerome Miklau, and Dan Suciu. 2015. A Theory of Pricing Private Data. ACM Trans. Database Syst. 39, 4, Article 34 (Dec. 2015), 28 pages. https://doi.org/10.1145/2691190.2691191
2015
-
[37]
Xuan Luo, Jian Pei, Zicun Cong, and Cheng Xu. 2022. On shapley value in data assemblage under independent utility. Proc. VLDB Endow. 15, 11 (jul 2022), 2761–2773. https://doi.org/10.14778/3551793.3551829
2022
-
[38]
Xuan Luo, Jian Pei, Cheng Xu, Wenjie Zhang, and Jianliang Xu. 2024. Fast Shapley Value Computation in Data Assemblage Tasks as Cooperative Simple Games. In Proceedings of the 2024 ACM SIGMOD International Conference on Management of Data. Santiago, Chile. https://doi.org/10.11...
2024 doi
-
[39]
Sasan Maleki, Long Tran-Thanh, Greg Hines, Talal Rahwan, and Alex Rogers
-
[40]
Jonathan Moore, Nils Hammerla, and Chris Watkins. 2019. Explaining deep learn- ing models with constrained adversarial examples. In Pacific Rim International Conference on Artificial Intelligence. Springer, 43–56
2019
-
[41]
Raha Moraffah, Mansooreh Karami, Ruocheng Guo, Adrienne Raglin, and Huan Liu. 2020. Causal Interpretability for Machine Learning-Problems, Methods and Evaluation. ACM SIGKDD Explorations Newsletter 22, 1 (2020), 18–33
2020
-
[42]
Kobbi Nissim, Salil Vadhan, and David Xiao. 2014. Redrawing the Boundaries on Purchasing Data from Privacy-Sensitive Individuals. In Proceedings of the 5th Conference on Innovations in Theoretical Computer Science (Princeton, New Jersey, USA) (ITCS’14). Association for Computi...
2014
-
[43]
Chaoyue Niu, Zhenzhe Zheng, Fan Wu, Shaojie Tang, Xiaofeng Gao, and Guihai Chen. 2018. Unlocking the Value of Privacy: Trading Aggregate Statistics over Private Correlated Data. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining...
2018
-
[44]
Pantelis and L
K. Pantelis and L. Aija. 2013. Understanding the value of (big) data. In 2013 IEEE International Conference on Big Data . 38–42
2013
-
[45]
Judea Pearl. 2009. Causal inference in statistics: An overview. Statistics surveys 3 (2009), 96–146
2009
-
[46]
Judea Pearl. 2010. Causal inference. Causality: objectives and assessment (2010), 39–58
2010
-
[47]
J. Pei. 2022. A Survey on Data Pricing: From Economics to Data Science. IEEE Transactions on Knowledge and Data Engineering 34, 10 (oct 2022), 4586–4608. https://doi.org/10.1109/TKDE.2020.3045927
2022
-
[48]
Foster Provost and Tom Fawcett. 2013. Data science and its relationship to big data and data-driven decision making. Big data 1, 1 (2013), 51–59
2013
-
[49]
Paul Resnick and Hal R Varian. 1997. Recommender systems. Commun. ACM 40, 3 (1997), 56–58
1997
-
[50]
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen
-
[51]
Fabian Schomm, Florian Stahl, and Gottfried Vossen. 2013. Marketplaces for Data: An Initial Survey. SIGMOD Rec. 42, 1 (May 2013), 15–26. https://doi.org/ 10.1145/2481528.2481532
2013
-
[52]
Lloyd S. Shapley. 1952. A Value for n-Person Games . Technical Report P-295. RAND Corporation, Santa Monica, CA. https://www.rand.org/pubs/papers/ P0295.html
1952
-
[53]
Kacper Sokol and Peter A Flach. 2019. Counterfactual explanations of machine learning predictions: opportunities and challenges for AI safety. In SafeAI@ AAAI
2019
-
[54]
W. Starr. 2022. Counterfactuals. In The Stanford Encyclopedia of Philosophy (Win- ter 2022 ed.), Edward N. Zalta and Uri Nodelman (Eds.). Metaphysics Research Lab, Stanford University
2022
-
[55]
William R Thompson. 1933. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika 25, 3-4 (12 1933), 285–294. https://doi.org/10.1093/biomet/25.3- 4.285 arXiv:https://academic.oup.com/biomet/article-pdf/25/3-4/285/51372...
1933 doi
-
[56]
Thompson
William R. Thompson. 1935. On the Theory of Apportionment.American Journal of Mathematics 57, 2 (1935), 450–456. http://www.jstor.org/stable/2371219
1935
-
[57]
Reiter, and Thomas Ristenpart
Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart
-
[58]
Arnaud Van Looveren and Janis Klaise. 2019. Interpretable counterfactual expla- nations guided by prototypes. arXiv preprint arXiv:1907.02584 (2019)
2019 arXiv
-
[59]
Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2017. Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harv. JL & Tech. 31 (2017), 841
2017
-
[60]
Jiachen T Wang, Yuqing Zhu, Yu-Xiang Wang, Ruoxi Jia, and Prateek Mittal
-
[61]
Xintao Wu, Xiaowei Ying, Kun Liu, and Lei Chen. 2010. A Survey of Privacy- Preservation of Graphs and Social Networks. In Managing and Mining Graph Data, Charu C. Aggarwal and Haixun Wang (Eds.). Springer US, Boston, MA, 421–453. https://doi.org/10.1007/978-1-4419-6045-0_14
2010 doi
-
[62]
Bin Zhou, Jian Pei, and WoShun Luk. 2008. A Brief Survey on Anonymization Techniques for Privacy Preserving Publishing of Social Network Data. SIGKDD Explor. Newsl. 10, 2 (Dec. 2008), 12–22. https://doi.org/10.1145/1540276.1540279
2008
-
[63]
Matjaz Zwitter and Milan Soklic. 1988. Breast Cancer. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C51P4M
1988 doi
-
[1485]
http://dblp.uni-trier.de/db/journals/pvldb/pvldb4.html#BalazinskaHS11
-
[2013]
arXiv:1306.4265 [cs.GT]
Bounding the Estimation Error of Sampling-based Shapley Value Approxi- mation. arXiv:1306.4265 [cs.GT]
-
[2016]
In Proceedings of the 25th USENIX Conference on Security Symposium (Austin, TX, USA) (SEC’16)
Stealing Machine Learning Models via Prediction APIs. In Proceedings of the 25th USENIX Conference on Security Symposium (Austin, TX, USA) (SEC’16). USENIX Association, USA, 601–618
- [2020]
-
[2023]
arXiv preprint arXiv:2308.15709 (2023)
Threshold KNN-Shapley: A Linear-Time and Privacy-Friendly Approach to Data Valuation. arXiv preprint arXiv:2308.15709 (2023)
2023 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.