REVIEW 2 major objections 4 minor 58 references
Hyperparameters in Score-Based Membership Inference Attacks
T0 review · 2 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This paper shows that membership inference attacks can match the likelihood-ratio attack's power without knowing the target model's hyperparameters, by matching score distributions in shadow-model training, and finds no detectable extra…
desk verdict KL-LiRA is a genuinely useful attack that shows hiding hyperparameters does little in transfer learning, though the marginal-matching heuristic is underdetermined and the paper is appropriately honest about that. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is KL-LiRA's distribution-matching objective. For each candidate hyperparameter $\eta_j$, the attacker trains shadow models on shadow datasets, approximates the target model's and shadow model's logit-scaled score distributions by Gaussians $\mathcal{N}_T$ and $\mathcal{N}_S$, and minimizes $\varphi_{KL}(\mathcal{N}_T \| \mathcal{N}_S) = \frac{1}{2}[ (\mu_S-\mu_T)^2/\sigma_S^2 + \sigma_T^2/\sigma_S^2 - \ln(\sigma_T^2/\sigma_S^2) - 1]$. This selects the hyperparameter $\eta_{j^*}$ that is then used to train the shadow models feeding LiRA's IN/OUT likelihood-ratio test. The key assumption is that matching these marginal score distributions is enough to make the shadow models behave like the target on individual samples.
What would settle it
Construct two hyperparameter settings that produce near-identical marginal score distributions on a large probe set but different per-sample memorization, run KL-LiRA on both, and check whether the selected setting yields a true-positive rate at 0.1% false positives clearly below LiRA's; if it does, the distribution-matching premise is false.
Extended reading notes
Core claim
The paper's central claim is that target-model hyperparameters are not a prerequisite for a strong score-based membership inference attack in the transfer-learning setting. KL-LiRA, its proposed attack, recovers effective hyperparameters purely from the target model's output scores: it trains candidate shadow models, approximates per-dataset score distributions as Gaussians, and selects the hyperparameter that minimizes the KL divergence to the target's score distribution. With those hyperparameters, LiRA's likelihood-ratio test performs nearly as well as when the true target hyperparameters are known; at 0.1% false-positive rate the average true-positive rate is close to LiRA's, whereas accuracy-based ACC-LiRA drops to about a quarter of LiRA's TPR. Under differential privacy, the paper reports that performing hyperparameter optimization on the training data yields membership inference vulnerability statistically indistinguishable from optimization on an external dataset, after false-discovery-rate adjustment.
Load-bearing premise
The attack works only if matching the overall score distributions of the target and shadow models also matches the per-sample score differences that reveal membership; if two hyperparameter choices give the same overall scores but different memorization behavior, the attack fails.
Editorial extensions
If this is right
- Keeping the target model's hyperparameters secret does not by itself protect a non-private model: an attacker who sees only the final scores can recover equivalent attack power through KL-LiRA.
- Choosing shadow-model hyperparameters by validation accuracy is a poor strategy for LiRA; at 0.1% false-positive rate, ACC-LiRA's average true-positive rate falls to roughly a quarter of LiRA's, while KL-LiRA stays close to LiRA.
- A practical KL-LiRA attack needs about 16 candidate hyperparameter settings and at least one shadow model per candidate to match the full-grid attack in the tested regime.
- For differentially private models, the study finds no statistically significant extra membership leakage from doing hyperparameter optimization on the training data rather than on an external dataset, suggesting current privacy accounting for HPO may be conservative.
- With a mismatched architecture between target and shadow models, both ACC-LiRA and KL-LiRA lose about 93% of their low-false-positive-rate effectiveness, so architectural secrecy remains a meaningful obstacle.
Reading between the lines
- If distribution matching continues to work in other domains, then hiding training details beyond the model itself is unlikely to block score-based membership inference attacks; defenses should focus on mechanisms such as differential privacy or output perturbation.
- The marginal-matching heuristic is fragile in principle: two hyperparameter settings could yield nearly identical Gaussian score distributions while differing in how strongly individual samples are memorized, and a targeted experiment could test whether KL-LiRA ever selects a setting with poor IN/OUT separation.
- The null result for data-based HPO is measured through membership inference attacks, not through a direct privacy audit; stronger attacks or different training regimes could still reveal leakage, so the finding is best read as an empirical lower bound rather than proof of absence.
- KL-LiRA's cost could potentially be reduced by cheaper proxies such as training with early stopping or smaller shadow datasets, since the paper only explores cost in terms of the number of trained models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the role of hyperparameters in score-based membership inference attacks (MIAs) in the few-shot transfer learning setting. It proposes KL-LiRA, a variant of LiRA that selects shadow-model training hyperparameters without knowledge of the target model's hyperparameters, by minimizing the KL divergence between Gaussian approximations of the target and shadow models' score distributions on shadow data (Section 3.1.2, Algorithm 2, Eq. (4)). The authors report that KL-LiRA achieves attack performance nearly indistinguishable from LiRA with known target hyperparameters, and substantially better than ACC-LiRA (which selects hyperparameters by utility) and shadow-model-free baselines (Section 5). In a second contribution, the paper compares MIA vulnerability when hyperparameter optimization is performed on the training data (TD-HPO) versus on an external disjoint dataset (ED-HPO) under DP, finding no statistically significant increase in TPR after Benjamini-Yekutieli FDR control (Section 7, Tables A2 and A3). The paper includes extensive experiments across CIFAR-10/CIFAR-100, ViT-B and R-50 backbones, Head and FiLM parameterizations, non-DP and DP settings (epsilon = 8 and 1), with 10 repeats and Clopper-Pearson confidence intervals.
Significance. If the results hold, the paper makes two useful contributions. First, it challenges the common assumption that MIA attackers need to know the target model's hyperparameters, and offers a practical method (KL-LiRA) that recovers much of the attack performance in the tested few-shot transfer learning scenarios; this has implications for how practitioners should reason about hyperparameter secrecy. Second, it provides empirical evidence that performing HPO on training data under DP may not lead to a detectable increase in MIA vulnerability compared to using an external dataset, which informs the ongoing discussion about privacy accounting for hyperparameter optimization. The empirical methodology is a strength: the authors use multiple architectures, datasets, and parameterizations, report confidence intervals, run both paired t-tests and permutation tests, and apply FDR control. The code is publicly available. The null result is presented with appropriate caveats about absence of evidence, although one early sentence in Section 8.2 overstates its strength. The main weakness is that the headline claims are broader than the tested hyperparameter space, as discussed below.
major comments (2)
- [Abstract, Section 5.1, Appendix A.4] The abstract and Section 1 claim that 'knowledge of target hyperparameters is not a prerequisite for MIA in the transfer learning setting,' and Section 5.1 concludes that KL-LiRA performs on par with LiRA across all settings. However, in the main experiments the set of unknown hyperparameters is limited to learning rate, batch size, and (for DP) the clipping bound; Appendix A.4 shows that when the number of training epochs is added to this set, KL-LiRA's TPR at FPR=0.1% drops to roughly 0.62x of its value without epoch tuning. The broad claim is therefore not supported for the full hyperparameter space, and the abstract and conclusion should be explicitly scoped to the tested hyperparameter families, or supplemented with evidence that the method degrades gracefully as the search space grows.
- [Section 3.1.2, Eq. (4)] KL-LiRA's selection criterion matches the marginal score distribution of the target and shadow models, approximated by Gaussians. LiRA's attack statistic, however, is a per-sample likelihood ratio over the conditional IN and OUT score distributions. Matching the marginal mean and variance imposes only two constraints on the four conditional moments (mu_in, mu_out, sigma_in^2, sigma_out^2) that determine the likelihood ratio, so the objective is not guaranteed to reproduce the per-sample signal that LiRA exploits. The paper should either provide a theoretical condition under which marginal matching implies conditional matching, or explicitly frame KL-LiRA as a heuristic whose validity is demonstrated only for the tested settings. The current text in Section 3.1.2 presents the heuristic without discussing this gap, and Appendix A.4 is consistent with the concern that marginal matching becomes less sufficient as the search space grows.
minor comments (4)
- [Section 3.2 and Figure 9] The model-count formulas for KL-LiRA are inconsistent. Section 3.2 states a total of (C x T) + C x (N - 1) + M - N models, while Figure 9's caption uses (C x T) + C x (N - 1) - N for 'extra models needed.' Deriving from the described procedure (C HPO runs with T trials each, N shadow models per candidate, and M shadow models for the final attack) gives C x T + C x N + (M - N) total models, or C x T + N(C - 1) extra beyond the M attack models. Please correct the formulas and ensure the x-axis of Figure 9 is consistent with the main text.
- [Appendix A.4] The phrase 'tpr for KL-LiRA (when tuning for training epochs) is 0.62x lower than tpr for KL-LiRA (without tuning for training epochs)' is ambiguous; it should read '62% of' or '38% lower' to avoid confusion about the direction and magnitude of the drop.
- [Section 8.2] The sentence 'Our results suggest that the privacy bounds provided by current DP HPO methods are likely to be loose' overstates what a null result can establish. The immediately following paragraph correctly notes that absence of evidence is not evidence of absence; the earlier sentence should be softened to match this caveat.
- [Sections 2.2 and 3.1.2] The paper alternates between 'loss distributions' and 'score distributions' when describing the Gaussian approximations in LiRA and KL-LiRA. Since Eq. (1) defines logits of confidence scores, the terminology should be made consistent, particularly in Algorithm 2 and Eq. (4).
Circularity Check
No circularity: KL-LiRA's hyperparameter selection optimizes a distribution-matching objective, and the reported attack TPR is measured empirically rather than derived from that objective.
full rationale
No significant circularity. The paper's central claim — that KL-LiRA, which selects shadow-model hyperparameters by minimizing the KL divergence between Gaussian approximations of target and shadow score distributions (Eq. 4, Algorithm 2), achieves attack power comparable to LiRA with known target hyperparameters — is established by measuring TPR at fixed FPR (Figures 3 and 4), not by deriving TPR from the selection objective. The KL objective constrains only the marginal score means and variances on shadow datasets; it does not by construction reproduce the per-sample IN/OUT conditional separation that the LiRA likelihood ratio depends on, so near-parity in TPR is an empirical outcome. Indeed, Appendix A.4 shows the relationship is fragile: adding epochs to the unknown hyperparameter set drops KL-LiRA's TPR to 0.62x of its value without epoch tuning, which is inconsistent with a tautological fit. The paper's self-citations (e.g., Tobaben et al. [26] for the HPO protocol, and code adapted from prior work) are methodological and not load-bearing: neither the KL-selection heuristic nor the attack result reduces to those citations. The DP-HPO study (Sections 6-7) is likewise an empirical paired comparison with BY-adjusted p-values, not a derivation from an assumption. The stated limitations (disjoint shadow data, small hyperparameter search space) further confirm that the claims are contingent empirical findings rather than results forced by definition or by self-citation.
Assumptions & free parameters
assumptions (5)
- domain assumption Shadow datasets Di are drawn independently from the same data distribution D as the target model's training data.
- domain assumption The per-sample score distributions of target and shadow models are well approximated by Gaussians after logit scaling.
- domain assumption The attacker knows the set of hyperparameters and their plausible ranges (learning rate, batch size, gradient clipping bound), listed in Table A1.
- domain assumption For the HPO privacy comparison, the external dataset used for ED-HPO is drawn from the same population as the training data and is an appropriate baseline for the training-data HPO setting.
- standard math The PRV privacy accountant correctly computes the (epsilon, delta) budget for DP-SGD.
Cite this review
Pith. "Pith review of Hyperparameters in Score-Based Membership Inference Attacks." pith.science (2026). https://pith.science/paper/S3EMQ2ZX
@misc{pith2026250206374,
author = {Pith},
title = {Pith review of: Hyperparameters in Score-Based Membership Inference Attacks},
year = {2026},
howpublished = {\url{https://pith.science/paper/S3EMQ2ZX}},
note = {Machine review of arXiv:2502.06374}
}
read the original abstract
Membership Inference Attacks (MIAs) have emerged as a valuable framework for evaluating privacy leakage by machine learning models. Score-based MIAs are distinguished, in particular, by their ability to exploit the confidence scores that the model generates for particular inputs. Existing score-based MIAs implicitly assume that the adversary has access to the target model's hyperparameters, which can be used to train the shadow models for the attack. In this work, we demonstrate that the knowledge of target hyperparameters is not a prerequisite for MIA in the transfer learning setting. Based on this, we propose a novel approach to select the hyperparameters for training the shadow models for MIA when the attacker has no prior knowledge about them by matching the output distributions of target and shadow models. We demonstrate that using the new approach yields hyperparameters that lead to an attack near indistinguishable in performance from an attack that uses target hyperparameters to train the shadow models. Furthermore, we study the empirical privacy risk of unaccounted use of training data for hyperparameter optimization (HPO) in differentially private (DP) transfer learning. We find no statistically significant evidence that performing HPO using training data would increase vulnerability to MIA.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Membership Inference Attacks Against Machine Learning Models,
R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership Inference Attacks Against Machine Learning Models,” in 2017 IEEE Symposium on Secu- rity and Privacy, SP 2017 . IEEE Computer Society, 2017, pp. 3–18
work page 2017
-
[2]
Membership Inference Attacks From First Principles,
N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tram` er, “Membership Inference Attacks From First Principles,” in 43rd IEEE Symposium on Secu- rity and Privacy, SP 2022 . IEEE, 2022, pp. 1897– 1914
work page 2022
-
[3]
Enhanced Membership Inference At- tacks against Machine Learning Models,
J. Ye, A. Maddi, S. K. Murakonda, V. Bindschaedler, and R. Shokri, “Enhanced Membership Inference At- tacks against Machine Learning Models,” in Pro- ceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security, CCS 2022 . ACM, 2022, pp. 3093–3106
work page 2022
-
[4]
Low-cost High-power Membership Inference Attacks,
S. Zarifzadeh, P. Liu, and R. Shokri, “Low-cost High-power Membership Inference Attacks,” inForty- first International Conference on Machine Learning, ICML 2024 . OpenReview.net, 2024
work page 2024
-
[5]
Mem- bership Inference Attacks by Exploiting Loss Trajec- tory,
Y. Liu, Z. Zhao, M. Backes, and Y. Zhang, “Mem- bership Inference Attacks by Exploiting Loss Trajec- tory,” in Proceedings of the 2022 ACM SIGSAC Con- ference on Computer and Communications Security, CCS 2022 . ACM, 2022, pp. 2085–2098. 1https://github.com/cambridge-mlg/dp-few-shot (Tobaben et al. [26]) 2https://github.com/tensorflow/privacy/tree/master/ ...
work page 2022
-
[6]
A. Salem, Y. Zhang, M. Humbert, P. Berrang, M. Fritz, and M. Backes, “ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models,” in 26th An- nual Network and Distributed System Security Sym- posium, NDSS 2019 . The Internet Society, 2019
work page 2019
-
[7]
Calibrating Noise to Sensitivity in Private Data Analysis,
C. Dwork, F. McSherry, K. Nissim, and A. D. Smith, “Calibrating Noise to Sensitivity in Private Data Analysis,” in Theory of Cryptography, Third Theory of Cryptography Conference, TCC 2006, Proceedings, ser. Lecture Notes in Computer Science, vol. 3876. Springer, 2006, pp. 265–284
work page 2006
-
[8]
Private Selection From Private Candidates,
J. Liu and K. Talwar, “Private Selection From Private Candidates,” in Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing, STOC
Show all 58 references
-
[9]
Hyperparameter Tun- ing with Renyi Differential Privacy,
N. Papernot and T. Steinke, “Hyperparameter Tun- ing with Renyi Differential Privacy,” inThe Tenth In- ternational Conference on Learning Representations, ICLR 2022 . OpenReview.net, 2022
2022
-
[10]
Practical Differen- tially Private Hyperparameter Tuning with Subsam- pling,
A. Koskela and T. D. Kulkarni, “Practical Differen- tially Private Hyperparameter Tuning with Subsam- pling,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Informa- tion Processing Systems 2023, NeurIPS 2023 , 2023
2023
-
[11]
DP-HyPO: An Adaptive Private Framework for Hy- perparameter Optimization,
H. Wang, S. Gao, H. Zhang, W. J. Su, and M. Shen, “DP-HyPO: An Adaptive Private Framework for Hy- perparameter Optimization,” in Advances in Neural Information Processing Systems 36: Annual Confer- ence on Neural Information Processing Systems 2023, NeurIPS 2023 , 2023
2023
-
[12]
Tight Au- diting of Differentially Private Machine Learning,
M. Nasr, J. Hayes, T. Steinke, B. Balle, F. Tram` er, M. Jagielski, N. Carlini, and A. Terzis, “Tight Au- diting of Differentially Private Machine Learning,” in 32nd USENIX Security Symposium, USENIX Secu- rity 2023 . USENIX Association, 2023, pp. 1631– 1648
2023
-
[13]
Revisit- ing Differentially Private Hyper-parameter Tuning,
Z. Xiang, T. Wang, C. Wang, and D. Wang, “Revisit- ing Differentially Private Hyper-parameter Tuning,” CoRR, 2024
2024
-
[14]
Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations,
A. Vassilev, A. Oprea, A. Fordyce, and H. Andersen, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations,” in NIST Trustworthy and Responsible AI, National Institute of Standards and Technology, 2024
2024
-
[15]
Membership Inference Attacks on Ma- chine Learning: A Survey,
H. Hu, Z. Salcic, L. Sun, G. Dobbie, P. S. Yu, and X. Zhang, “Membership Inference Attacks on Ma- chine Learning: A Survey,” ACM Comput. Surv. , vol. 54, no. 11s, pp. 235:1–235:37, 2022. 15
2022
-
[16]
Compre- hensive Privacy Analysis of Deep Learning: Stand- alone and Federated Learning under Passive and Active White-box Inference Attacks,
M. Nasr, R. Shokri, and A. Houmansadr, “Compre- hensive Privacy Analysis of Deep Learning: Stand- alone and Federated Learning under Passive and Active White-box Inference Attacks,” CoRR, vol. abs/1812.00910, 2018
2018 arXiv
-
[17]
Label-only Membership Inference At- tacks,
C. A. Choquette-Choo, F. Tram` er, N. Carlini, and N. Papernot, “Label-only Membership Inference At- tacks,” in Proceedings of the 38th International Con- ference on Machine Learning, ICML 2021 , ser. Pro- ceedings of Machine Learning Research, vol. 139. PMLR, 2021, pp. 1964–1974
2021
-
[18]
A Differentially Pri- vate Stochastic Gradient Descent Algorithm for Mul- tiparty Classification,
A. Rajkumar and S. Agarwal, “A Differentially Pri- vate Stochastic Gradient Descent Algorithm for Mul- tiparty Classification,” in Proceedings of the Fif- teenth International Conference on Artificial Intelli- gence and Statistics, AISTATS 2012, ser. JMLR Pro- ceedings, vol. 2...
2012
-
[19]
Stochas- tic Gradient Descent With Differentially Private Up- dates,
S. Song, K. Chaudhuri, and A. D. Sarwate, “Stochas- tic Gradient Descent With Differentially Private Up- dates,” in IEEE Global Conference on Signal and In- formation Processing, GlobalSIP 2013 . IEEE, 2013, pp. 245–248
2013
-
[20]
Deep Learning with Differential Privacy,
M. Abadi, A. Chu, I. J. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep Learning with Differential Privacy,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Com- munications Security. ACM, 2016, pp. 308–318
2016
-
[21]
How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy,
N. Ponomareva, H. Hazimeh, A. Kurakin, Z. Xu, C. Denison, H. B. McMahan, S. Vassilvitskii, S. Chien, and A. G. Thakurta, “How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy,” J. Artif. Intell. Res., vol. 77, pp. 1113–1201, 2023
2023
-
[22]
Posi- tion: Considerations for Differentially Private Learn- ing with Large-scale Public Pretraining,
F. Tram` er, G. Kamath, and N. Carlini, “Posi- tion: Considerations for Differentially Private Learn- ing with Large-scale Public Pretraining,” in Forty- first International Conference on Machine Learning, ICML 2024 . OpenReview.net, 2024
2024
-
[23]
How Transferable Are Features in Deep Neural Networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How Transferable Are Features in Deep Neural Networks?” in Advances in Neural Information Processing Sys- tems 27: Annual Conference on Neural Information Processing Systems 2014, NeurIPS 2014 , 2014, pp. 3320–3328
2014
-
[24]
Fine-tuning with Differential Privacy Necessitates an Additional Hyperparameter Search,
Y. Cattan, C. A. Choquette-Choo, N. Papernot, and A. Thakurta, “Fine-tuning with Differential Privacy Necessitates an Additional Hyperparameter Search,” CoRR, vol. abs/2210.02156, 2022
2022 arXiv
-
[25]
Toward Training at Im- ageNet Scale with Differential Privacy,
A. Kurakin, S. Chien, S. Song, R. Geambasu, A. Terzis, and A. Thakurta, “Toward Training at Im- ageNet Scale with Differential Privacy,” CoRR, vol. abs/2201.12328, 2022
2022 arXiv
-
[26]
On the Efficacy of Differentially Private Few-shot Image Classification,
M. Tobaben, A. Shysheya, J. Bronskill, A. Paverd, S. Tople, S. Z. B´ eguelin, R. E. Turner, and A. Honkela, “On the Efficacy of Differentially Private Few-shot Image Classification,” Trans. Mach. Learn. Res., vol. 2023, 2023
2023
-
[27]
Privacy-aware Document Visual Question Answering,
R. Tito, K. Nguyen, M. Tobaben, R. Kerkouche, M. A. Souibgui, K. Jung, J. J¨ alk¨ o, V. P. D’Andecy, A. Joseph, L. Kang, E. Valveny, A. Honkela, M. Fritz, and D. Karatzas, “Privacy-aware Document Visual Question Answering,” in Document Analysis and Recognition - ICDAR 2024 - 1...
2024
-
[28]
Large Language Models Can Be Strong Differen- tially Private Learners,
X. Li, F. Tram` er, P. Liang, and T. Hashimoto, “Large Language Models Can Be Strong Differen- tially Private Learners,” in The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net, 2022
2022
-
[29]
Differ- entially Private Fine-tuning of Language Models,
D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, S. Yekhanin, and H. Zhang, “Differ- entially Private Fine-tuning of Language Models,” in The Tenth International Conference on Learning Representations, ICLR 2022. Open...
2022
-
[30]
LoRA: Low- rank Adaptation of Large Language Models,
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low- rank Adaptation of Large Language Models,” in The Tenth International Conference on Learning Repre- sentations, ICLR 2022 . OpenReview.net, 2022
2022
-
[31]
FiLM: Visual Reasoning with a General Conditioning Layer,
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville, “FiLM: Visual Reasoning with a General Conditioning Layer,” in Proceedings of the Thirty-Second AAAI Conference on Artificial Intel- ligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligen...
2018
-
[32]
Unlocking High-accuracy Differentially Pri- vate Image Classification through Scale,
S. De, L. Berrada, J. Hayes, S. L. Smith, and B. Balle, “Unlocking High-accuracy Differentially Pri- vate Image Classification through Scale,” CoRR, vol. abs/2204.13650, 2022
2022 arXiv
-
[33]
Adversary Instantiation: Lower Bounds for Differentially Private Machine Learning,
M. Nasr, S. Song, A. Thakurta, N. Papernot, and N. Carlini, “Adversary Instantiation: Lower Bounds for Differentially Private Machine Learning,” in 42nd IEEE Symposium on Security and Privacy, SP 2021 . IEEE, 2021, pp. 866–882. 16
2021
-
[34]
Auditing Differentially Private Machine Learning: How Private is Private SGD?
M. Jagielski, J. R. Ullman, and A. Oprea, “Auditing Differentially Private Machine Learning: How Private is Private SGD?” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, 2020
2020
-
[35]
The Compo- sition Theorem for Differential Privacy,
P. Kairouz, S. Oh, and P. Viswanath, “The Compo- sition Theorem for Differential Privacy,” in Proceed- ings of the 32nd International Conference on Machine Learning, ICML 2015, ser. JMLR Workshop and Con- ference Proceedings, vol. 37. JMLR.org, 2015, pp. 1376–1385
2015
-
[36]
Learning Multiple Layers of Features From Tiny Images,
A. Krizhevsky, “Learning Multiple Layers of Features From Tiny Images,” Master’s thesis, University of Toronto, 2009
2009
-
[37]
Adam: A Method for Stochastic Optimization,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in 3rd International Con- ference on Learning Representations, ICLR 2015, Conference Track Proceedings, 2015
2015
-
[38]
ImageNet Large Scale Visual Recognition Chal- lenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei- Fei, “ImageNet Large Scale Visual Recognition Chal- lenge,” Int. J. Comput. Vis. , vol. 115, no. 3, pp. 211– 252, 2015
2015
-
[39]
Big Transfer (BiT): General Visual Representation Learning,
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby, “Big Transfer (BiT): General Visual Representation Learning,” in Computer Vision - ECCV 2020 - 16th European Con- ference, ser. Lecture Notes in Computer Science, vol. 12350. Springer, 2020, pp...
2020
-
[40]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weis- senborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in 9th International Conference on Le...
2021
-
[41]
K for the Price of 1: Parameter- efficient Multi-task and Transfer Learning,
P. K. Mudrakarta, M. Sandler, A. Zhmoginov, and A. G. Howard, “K for the Price of 1: Parameter- efficient Multi-task and Transfer Learning,” in 7th In- ternational Conference on Learning Representations, ICLR 2019 . OpenReview.net, 2019
2019
-
[42]
Contextual Squeeze-and-excitation for Efficient Few-shot Image Classification,
M. Patacchiola, J. Bronskill, A. Shysheya, K. Hof- mann, S. Nowozin, and R. E. Turner, “Contextual Squeeze-and-excitation for Efficient Few-shot Image Classification,” in Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing S...
2022
-
[43]
FiT: Parameter Efficient Few-shot Transfer Learning for Personal- ized and Federated Image Classification,
A. Shysheya, J. Bronskill, M. Patacchiola, S. Nowozin, and R. E. Turner, “FiT: Parameter Efficient Few-shot Transfer Learning for Personal- ized and Federated Image Classification,” in The Eleventh International Conference on Learning Representations, ICLR 2023 . OpenReview.net, 2023
2023
-
[44]
Opacus: User-friendly Differential Privacy Library in PyTorch,
A. Yousefpour, I. Shilov, A. Sablayrolles, D. Testug- gine, K. Prasad, M. Malek, J. Nguyen, S. Ghosh, A. Bharadwaj, J. Zhao, G. Cormode, and I. Mironov, “Opacus: User-friendly Differential Privacy Library in PyTorch,” CoRR, vol. abs/2109.12298, 2021
2021 arXiv
-
[45]
Py- Torch: An Imperative Style, High-performance Deep Learning Library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Brad- bury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K¨ opf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Py- Torch: An Imperative Style...
2019
-
[46]
Numerical Composition of Differential Privacy,
S. Gopi, Y. T. Lee, and L. Wutschitz, “Numerical Composition of Differential Privacy,” in Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Sys- tems 2021, NeurIPS 2021 , 2021, pp. 11 631–11 642
2021
-
[47]
The Use of Confi- dence or Fiducial Limits Illustrated in the Case of the Binomial,
C. J. Clopper and E. S. Pearson, “The Use of Confi- dence or Fiducial Limits Illustrated in the Case of the Binomial,” Biometrika, vol. 26, no. 4, pp. 404–413, 12 1934
1934
-
[48]
Scalable Membership Inference At- tacks via Quantile Regression,
M. Bertr´ an, S. Tang, A. Roth, M. Kearns, J. Morgen- stern, and S. Wu, “Scalable Membership Inference At- tacks via Quantile Regression,” in Advances in Neural Information Processing Systems 36: Annual Confer- ence on Neural Information Processing Systems 2023, NeurIPS 2023 , 2023
2023
-
[49]
The Probable Error of a Mean,
Student, “The Probable Error of a Mean,” Biometrika, vol. 6, no. 1, pp. 1–25, 1908
1908
-
[50]
Permutation Methods: A Basis for Ex- act Inference,
M. D. Ernst, “Permutation Methods: A Basis for Ex- act Inference,” Statistical Science, vol. 19, no. 4, pp. 676–685, 2004
2004
-
[51]
The Control of the False Discovery Rate in Multiple Testing Under De- pendency,
Y. Benjamini and D. Yekutieli, “The Control of the False Discovery Rate in Multiple Testing Under De- pendency,” The Annals of Statistics , vol. 29, no. 4, pp. 1165 – 1188, 2001
2001
-
[52]
Differentially Private Learning Needs Hidden State (Or Much Faster Convergence),
J. Ye and R. Shokri, “Differentially Private Learning Needs Hidden State (Or Much Faster Convergence),” in Advances in Neural Information Processing Sys- tems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022 , 2022. 17
2022
-
[53]
Privacy of Noisy Stochastic Gradient Descent: More Iterations with- out More Privacy Loss,
J. M. Altschuler and K. Talwar, “Privacy of Noisy Stochastic Gradient Descent: More Iterations with- out More Privacy Loss,” in Advances in Neural In- formation Processing Systems 35: Annual Confer- ence on Neural Information Processing Systems 2022, NeurIPS 2022 , 2022
2022
-
[54]
BoTorch: A Framework for Efficient Monte-carlo Bayesian Op- timization,
M. Balandat, B. Karrer, D. R. Jiang, S. Daulton, B. Letham, A. G. Wilson, and E. Bakshy, “BoTorch: A Framework for Efficient Monte-carlo Bayesian Op- timization,” in Advances in Neural Information Pro- cessing Systems 33: Annual Conference on Neural In- formation Processing Sy...
2020
-
[55]
Optuna: A Next-generation Hyperpa- rameter Optimization Framework,
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A Next-generation Hyperpa- rameter Optimization Framework,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK. ACM, 2019, pp. 2623–2631
2019
-
[56]
Wide Residual Networks,
S. Zagoruyko and N. Komodakis, “Wide Residual Networks,” in Proceedings of the British Machine Vi- sion Conference 2016, BMVC 2016 . BMV A Press, 2016
2016
-
[57]
Gradient-based Learning Applied to Document Recognition,
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based Learning Applied to Document Recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278– 2324, 1998. A.1 Hyperparameter Optimization For a given input data set, D, we perform hyperparameter optimization (HPO) using 70% o...
1998
-
[2019]
ACM, 2019, pp. 298–309
2019
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.