Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

A Survey on Privacy Risks and Protection in Large Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This survey claims LLM privacy research can be organized by a unified classification of privacy leakage and privacy attacks.

desk verdict A broadly useful but sloppy survey: the taxonomy is clear, but Table 1 misassigns several papers to privacy categories, which undermines the central claim of a trusted map of the field. read the letter →

arxiv 2505.01976 v1 pith:FJZMYPDF submitted 2025-05-04 cs.CR

classification cs.CR
keywords largelanguagemodelsprivacyleakageattacksmembershipinferencemodelinversionbackdoorfederatedlearningdifferential
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the scattered privacy literature on large language models fits one organizing scheme: privacy leakage, which covers sensitive information, contextual information, and personal preferences, and privacy attacks, which target models, data, or users. It reviews existing defenses and maps them to these categories, arguing that the classification gives the field a shared vocabulary and a roadmap for future work. If the taxonomy is right, researchers and practitioners can file a reported vulnerability into a category and know which mitigation family has been studied for it.

What carries the argument

The central object is the taxonomy itself: a two-level classification of privacy risks according to how an attacker gains access to sensitive information. Privacy leakage names passive exposure through the model's ordinary behavior, while privacy attacks name active attempts to break into the model, the data, or the user. The taxonomy organizes the survey: each cell carries definitions, representative works, evaluation metrics, and candidate mitigations, so the entire review is structured by the classification.

What would settle it

Take the entries in Tables 1 through 4, read the abstract of each cited paper, and check whether the method and evaluation metric match the assigned category; several entries already listed under sensitive-information leakage are prompt-design and robustness studies, so if a substantial fraction of the other entries also fail this check, the taxonomy does not accurately map the field.

Watch

Extended reading notes

Core claim

The central claim is that LLM privacy work is not a set of isolated problems but a single two-branch landscape. Privacy leakage is the exploitation of LLM vulnerabilities to collect sensitive information, divided into sensitive information leakage, contextual leakage, and personal preferences leakage. Privacy attacks are active attempts to breach the model's defenses, divided into model-based attacks (backdoor, model inversion, model stealing), data-based attacks (data stealing, training data extraction), and user-based attacks (membership inference, attribute inference). The survey then pairs each category with representative papers and defense strategies, including data cleaning, inference detection, federated learning, differential privacy, backdoor removal, cryptography, and confidential computing, and states that these defenses correspond to the risks in the taxonomy.

Load-bearing premise

The survey's usefulness depends on each cited paper actually belonging to the category it is placed in, and since the survey gives no search or inclusion criteria, a reader cannot verify that the table assignments faithfully represent the literature.

Editorial extensions

If this is right

  • A practitioner who encounters a reported privacy incident can assign it to one leakage or attack category and immediately see which defense family has been tested against that category.
  • Privacy risk assessment frameworks for LLMs should measure the three leakage channels and the three attack targets named in the taxonomy, because the survey argues those cover the current literature.
  • Benchmarks for LLM privacy, which the survey calls for, would be comparable across papers if they are organized by the same categories.
  • The leakage-versus-attack split gives a natural mapping to governance: leakage concerns data-protection duties, while attacks concern adversarial threat modeling.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The boundary between leakage and attack is likely fuzzier in practice than the taxonomy suggests, since the same model behavior can be read as passive exposure or active exfiltration depending on the threat model and who initiates the interaction.
  • Because the survey does not state its search and inclusion criteria, the taxonomy's completeness is only as strong as the chosen sample of papers; a formal selection procedure and an audit of the table assignments would make the roadmap reproducible.
  • The contextual leakage category points toward a testable extension: privacy evaluation should compare information flows that have the same content but different recipients, senders, or transmission principles, rather than only measuring how much information is revealed.
  • If the taxonomy is adopted as a standard, it could be encoded as a machine-readable checklists for audit reports, letting regulators see which leakage and attack categories a model has actually been evaluated against.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a survey of privacy risks and protections for large language models. It proposes a taxonomy in which privacy issues are divided into privacy leakage (sensitive information leakage, contextual leakage, personal preferences leakage) and privacy attacks (model-based attacks: backdoor, model inversion, model stealing; data-based attacks: data stealing, training data extraction; user-based attacks: membership inference, attribute inference). It then reviews defense mechanisms including data cleaning, inference detection, federated learning, differential privacy, backdoor removal, cryptography, and confidential computing, and concludes with future research directions. The paper's stated contribution is a comprehensive and fine-grained classification that integrates privacy concerns and provides a roadmap for addressing them.

Significance. If the survey's taxonomy and literature assignments were accurate, the paper would be a useful organizing resource for researchers entering the LLM-privacy area: it covers a wide range of attacks and defenses, connects privacy leakage to contextual integrity, reproduces several key attack equations from the primary literature, and outlines sensible future directions such as privacy-preserving model compression, privacy risk assessment, and secure knowledge sharing. The main contribution is the taxonomy and the associated reference tables, so the correctness of those assignments is load-bearing. The miscategorizations identified below in Table 1 are not cosmetic; they directly affect whether a reader can trust the paper as a map of the field. The absence of a stated methodology for selecting and classifying the 92 references further weakens the comprehensiveness claim. The paper is salvageable with a systematic audit and a methodology statement, but in its current form the central empirical basis for the taxonomy is unreliable.

major comments (4)
  1. [Table 1, Section 3.1] Table 1 lists [28] Zamfirescu-Pereira et al. (CHI 2023) and [29] Wang et al. under 'Sensitive Information Leakage' as works establishing that category. [28] is a human-factors study of how non-AI experts design prompts, and [29] studies adversarial and out-of-distribution robustness of ChatGPT. Neither paper studies privacy leakage. The main text confirms this: [28] is cited only to support the general claim that chat-based interaction is useful for 'programming, academic writing, and medical diagnosis,' and [29] is cited only to support the claim that LLMs 'have achieved significant performance with various NLP tasks.' These are capability citations, not privacy-leakage citations, so the table entry misrepresents the evidence base for the category.
  2. [Table 1, Section 3.3] Table 1 places [33] Thomas et al. under 'Personal Preferences Leakage.' That paper is an information-retrieval study showing that LLMs can accurately predict searcher preferences; it does not demonstrate or measure privacy leakage. The text infers a privacy risk from the model's predictive ability, but presenting this work as an instance of a privacy-leakage category overstates what the cited paper establishes. At minimum the category needs a clearly labeled 'potential risk' entry, not a direct assignment.
  3. [Section 3, Tables 1-4] The survey does not state its search strategy, inclusion/exclusion criteria, databases, or time window for the 92 references, despite advertising a 'comprehensive overview.' Given that Table 1 demonstrably contains citations that do not support the categories to which they are assigned, the absence of a documented methodology makes it impossible for a reader to determine whether the remaining assignments were derived systematically or selectively. The comprehensiveness claim therefore needs either a methodology subsection or a substantial softening.
  4. [Table 4, Section 4.2] Table 4 appears to repeat the misassignment problem in the defense taxonomy: [67] is a fine-tuning-based backdoor defense, yet in the extracted table layout it is placed under 'Differential Privacy,' while [70] (THE-X) is a homomorphic-encryption method for transformer inference, yet it appears in the 'Backdoor Removal' row. If this is an artifact of the table's typesetting and column alignment, the table needs clearer category separators; if it is accurate, it provides further evidence that the category assignments are unreliable. Either way, the table as presented requires correction.
minor comments (5)
  1. [Equations (1)-(7)] Several equations are under-specified for a self-contained survey. In Eq. (2), Rl_c appears without definition; in Eq. (5), the symbols Xb, Sy, Tpre, and cPθ are not defined; and Eq. (3) contains a formatting artifact ('M AX(L)') that obscures the intended expression. Quoting equations from primary papers is acceptable, but each symbol should be defined or a precise reference given.
  2. [References [62] and [65]] References [62] and [65] are the same paper (Kim et al., ProPILE) and should be consolidated into a single entry to avoid confusing duplicate citations in Table 3 and the text.
  3. [Section 3.1, citation [24]] The sentence 'Some users believe that the information they provide is stored in the ChatGPT database...' is cited to [24], but [24] (Carlini et al.) is a training-data-extraction paper and does not report on user beliefs. A more appropriate citation would be [25] (Kshetri) or a user-survey source.
  4. [General presentation] The manuscript contains multiple LaTeX and typesetting artifacts, including 'T able' in table captions, missing spaces in phrases such as 'PersonalPreferencesLeakage,' and a stray '&' in Eq. (5). A careful proofreading pass is needed.
  5. [Data availability statement] For a survey paper, the statement that datasets 'are available from the corresponding author upon reasonable request' is misleading; the paper analyzes no datasets. This should be replaced with a statement that no new data were generated or analyzed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's taxonomy is assembled from cited prior work, with no derived quantities or fitted parameters that reduce to the paper's own inputs.

full rationale

This paper is a survey, not a derivation or prediction pipeline. Its central claim is that it provides a comprehensive, systematically organized overview of LLM privacy risks and defenses. That claim rests on the fidelity of its literature selection and category assignments in Tables 1-4, not on any mathematical or statistical derivation. The paper introduces no new quantity, model, or estimator that is defined in terms of its own output, and no result is predicted from a fitted parameter. The equations appearing in the text (e.g., the BadEdit objective in Eq. (1)-(2), the Text Revealer loss in Eq. (4), the data-stealing objective in Eq. (5), and the PPI definition in Eq. (7)) are quoted from cited prior work to illustrate known attack or defense mechanisms; they are not used to derive new conclusions. The taxonomy separating privacy leakage from privacy attacks is an organizational schema applied to externally published papers, and each substantive claim is attributed to a reference. The few self-citations present are not load-bearing, and no uniqueness theorem or prior result by the same authors is invoked to force a choice. The apparent miscategorizations flagged by the reader, such as listing [28] and [29] under Sensitive Information Leakage in Table 1, are accuracy and literature-mapping concerns rather than circularity: they concern whether cited papers support the stated category, not whether the category's content reduces to the paper's own assumptions. Therefore the survey is not circular, and a 0 score is appropriate.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities appear because the paper is a survey. The central assumptions are that LLM privacy risks are real and that the selected references are accurately represented by the taxonomy.

assumptions (2)
  • domain assumption LLMs can memorize and later emit sensitive training data or enable inference of personal attributes.
    The entire threat model rests on this premise, asserted in Sections 2 and 3 without independent verification in this paper.
  • ad hoc to paper The 92 references selected and the categories assigned in Tables 1-4 are representative of the wider literature.
    The survey provides no search strategy or inclusion criteria; the taxonomy's usefulness depends on this unstated selection assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Privacy Risks and Protection in Large Language Models." pith.science (2026). https://pith.science/paper/FJZMYPDF

@misc{pith2026250501976,
  author       = {Pith},
  title        = {Pith review of: A Survey on Privacy Risks and Protection in Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FJZMYPDF}},
  note         = {Machine review of arXiv:2505.01976}
}
read the original abstract

Although Large Language Models (LLMs) have become increasingly integral to diverse applications, their capabilities raise significant privacy concerns. This survey offers a comprehensive overview of privacy risks associated with LLMs and examines current solutions to mitigate these challenges. First, we analyze privacy leakage and attacks in LLMs, focusing on how these models unintentionally expose sensitive information through techniques such as model inversion, training data extraction, and membership inference. We investigate the mechanisms of privacy leakage, including the unauthorized extraction of training data and the potential exploitation of these vulnerabilities by malicious actors. Next, we review existing privacy protection against such risks, such as inference detection, federated learning, backdoor mitigation, and confidential computing, and assess their effectiveness in preventing privacy leakage. Furthermore, we highlight key practical challenges and propose future research directions to develop secure and privacy-preserving LLMs, emphasizing privacy risk assessment, secure knowledge transfer between models, and interdisciplinary frameworks for privacy governance. Ultimately, this survey aims to establish a roadmap for addressing escalating privacy challenges in the LLMs domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Reasoning Steering via Concept Chaining

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Reinforcement-learning-optimized concept-chain paragraphs covertly steer language-model multiple-choice preferences after continued pretraining, with far lower detectability than direct paraphrases.

  2. EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.

  3. An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    EASI-RAG is a structured agile method for deploying RAG tools in industrial SMEs, validated by one case study where a no-experience team built a working assistant in three weeks.

Reference graph

Works this paper leans on

92 extracted references · 41 canonical work pages · cited by 3 Pith papers

  1. [28]

    In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp

    Zamfirescu-Pereira, J., Wong, R.Y., Hartmann, B., Yang, Q.: Why johnny can’t prompt: how non-ai experts try (and fail) to design llm prompts. In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. 1–21 (2023)

  2. [29]

    arXiv preprint arXiv:2302.12095 (2023)

    Wang, J., Hu, X., Hou, W., Chen, H., Zheng, R., Wang, Y., Yang, L., Huang, H., Ye, W., Geng, X., et al.: On the robustness of chatgpt: An adversarial and out-of-distribution perspective. arXiv preprint arXiv:2302.12095 (2023)

  3. [33]

    In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp

    Thomas, P., Spielman, S., Craswell, N., Mitra, B.: Large language models can accurately predict searcher preferences. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1930–1940 (2024)

  4. [62]

    Advances in Neural Information Processing Systems 36 (2024)

    Kim, S., Yun, S., Lee, H., Gubri, M., Yoon, S., Oh, S.J.: Propile: Probing privacy leakage in large language models. Advances in Neural Information Processing Systems 36 (2024)

  5. [65]

    Advances in Neural Information Processing Systems 36, 20750–20762 (2023)

    Kim, S., Yun, S., Lee, H., Gubri, M., Yoon, S., Oh, S.J.: Propile: Probing privacy leakage in large language models. Advances in Neural Information Processing Systems 36, 20750–20762 (2023)

  6. [67]

    arXiv preprint arXiv:2212.09067 (2022)

    Sha, Z., He, X., Berrang, P., Humbert, M., Zhang, Y.: Fine-tuning is all you need to mitigate backdoor attacks. arXiv preprint arXiv:2212.09067 (2022)

  7. [70]

    arXiv preprint arXiv:2206.00216 (2022)

    Chen, T., Bao, H., Huang, S., Dong, L., Jiao, B., Jiang, D., Zhou, H., Li, J., Wei, F.: The-x: Privacy-preserving transformer inference with homomorphic encryption. arXiv preprint arXiv:2206.00216 (2022)

  8. [1]

    arXiv preprint arXiv:2303.08774 (2023) 22

    Achiam,J.,Adler,S.,Agarwal,S.,Ahmad,L.,Akkaya,I.,Aleman,F.L.,Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023) 22

Show all 92 references
  1. [2]

    ArXiv (2023)

    Bubeck, S., Chadrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y.T., Li, Y., Lundberg, S., et al.: Sparks of artificial general intelligence: Early experiments with gpt-4. ArXiv (2023)

  2. [3]

    arXiv preprint arXiv:2302.13971 (2023)

    Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al.: Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  3. [4]

    Learning and individual differences103, 102274 (2023)

    Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E.,et al.: Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences103, 1022...

  4. [5]

    IEEE Access10, 19621–19628 (2022)

    Chen, D., Hong, W., Zhou, X.: Transformer network for remaining useful life prediction of lithium-ion batteries. IEEE Access10, 19621–19628 (2022)

  5. [6]

    Advances in Neural Information Processing Systems36 (2024)

    Duan, H., Dziedzic, A., Papernot, N., Boenisch, F.: Flocks of stochastic parrots: Differentially private prompt learning for large language models. Advances in Neural Information Processing Systems36 (2024)

  6. [7]

    arXiv preprint arXiv:2204.09391 (2022)

    Plant, R., Giuffrida, V., Gkatzia, D.: You are what you write: Preserving privacy in the era of large language models. arXiv preprint arXiv:2204.09391 (2022)

  7. [8]

    arXiv preprint arXiv:2403.13355 (2024)

    Li, Y., Li, T., Chen, K., Zhang, J., Liu, S., Wang, W., Zhang, T., Liu, Y.: Badedit: Backdooring large language models by model editing. arXiv preprint arXiv:2403.13355 (2024)

  8. [9]

    Computers & Security135, 103476 (2023)

    Okey, O.D., Udo, E.U., Rosa, R.L., Rodríguez, D.Z., Kleinschmidt, J.H.: Investi- gating chatgpt and cybersecurity: A perspective on topic modeling and sentiment analysis. Computers & Security135, 103476 (2023)

  9. [10]

    arXiv preprint arXiv:2310.07298 (2023)

    Staab, R., Vero, M., Balunović, M., Vechev, M.: Beyond memorization: Violating privacy via inference with large language models. arXiv preprint arXiv:2310.07298 (2023)

  10. [11]

    IEEE Transactions on Big Data (2024)

    Xia, L., Fan, J., Parlikad, A., Huang, X., Zheng, P.: Unlocking large language model power in industry: Privacy-preserving collaborative creation of knowledge graph. IEEE Transactions on Big Data (2024)

  11. [12]

    Presented at Posters-at-the-Capitol, Northern Kentucky University, 2025 (2025)

    Dhungana, B., Ghimire, V., Shrestha Lama, J., Sadat, N., Caporusso, N., Doan, M.: Assessing Cybersecurity Awareness of ChatGPT’s New Memory Feature. Presented at Posters-at-the-Capitol, Northern Kentucky University, 2025 (2025)

  12. [13]

    arXiv preprint arXiv:2501.05965 (2025) 23

    Shu, Y., Li, S., Dong, T., Meng, Y., Zhu, H.: Model inversion in split learning for personalized llms: New insights from information bottleneck theory. arXiv preprint arXiv:2501.05965 (2025) 23

  13. [14]

    arXiv preprint arXiv:2403.05156 (2024)

    Yan, B., Li, K., Xu, M., Dong, Y., Zhang, Y., Ren, Z., Cheng, X.: On protect- ing the data privacy of large language models (llms): A survey. arXiv preprint arXiv:2403.05156 (2024)

  14. [15]

    ACM computing surveys56(11), 1–40 (2024)

    Mo, F., Tarkhani, Z., Haddadi, H.: Machine learning with confidential computing: A systematization of knowledge. ACM computing surveys56(11), 1–40 (2024)

  15. [16]

    Journal of Computer Science and Technology Studies7(1), 17–29 (2025)

    Nandagopal, S.: Securing retrieval-augmented generation pipelines: A compre- hensive framework. Journal of Computer Science and Technology Studies7(1), 17–29 (2025)

  16. [17]

    Information16(1), 49 (2025)

    Żarski, T.L., Janicki, A.: Enhancing privacy while preserving context in text transformations by large language models. Information16(1), 49 (2025)

  17. [18]

    ACM Computing Surveys57(6), 1–39 (2025)

    Das,B.C.,Amini,M.H.,Wu,Y.:Securityandprivacychallengesoflargelanguage models: A survey. ACM Computing Surveys57(6), 1–39 (2025)

  18. [19]

    High-Confidence Computing, 100211 (2024)

    Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, Z., Zhang, Y.: A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, 100211 (2024)

  19. [20]

    In: International Conference on Ubiquitous Security, pp

    Esmradi, A., Yip, D.W., Chan, C.F.: A comprehensive survey of attack tech- niques, implementation, and mitigation strategies in large language models. In: International Conference on Ubiquitous Security, pp. 76–95 (2023). Springer

  20. [21]

    Advances in Neural Information Processing Systems37, 131197–131223 (2025)

    Wang, J.T., Wu, T., Song, D., Mittal, P., Jia, R.: Greats: Online selection of high- quality data for llm training in every iteration. Advances in Neural Information Processing Systems37, 131197–131223 (2025)

  21. [22]

    Stanford Center for Research on Foundation Models

    Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., Hashimoto, T.B.: Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html3(6), 7 (2023)

  22. [23]

    Advances in neural information processing systems 35, 27730–27744 (2022)

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A.,et al.: Training language models to follow instructions with human feedback. Advances in neural information processing systems 35, 27730–27744 (2022)

  23. [24]

    In: 30th USENIX Security Symposium (USENIX Security 21), pp

    Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U.,et al.: Extracting training data from large language models. In: 30th USENIX Security Symposium (USENIX Security 21), pp. 2633–2650 (2021)

  24. [25]

    IT Professional 25(3), 9–13 (2023) 24

    Kshetri, N.: Cybercrime and privacy threats of large language models. IT Professional 25(3), 9–13 (2023) 24

  25. [26]

    Artificial Intelligence Review58(1), 1–47 (2025)

    Liu, Y., Huang, J., Li, Y., Wang, D., Xiao, B.: Generative ai model privacy: a survey. Artificial Intelligence Review58(1), 1–47 (2025)

  26. [27]

    arXiv preprint arXiv:2310.17884 (2023)

    Mireshghallah, N., Kim, H., Zhou, X., Tsvetkov, Y., Sap, M., Shokri, R., Choi, Y.: Can llms keep a secret? testing privacy implications of language models via contextual integrity theory. arXiv preprint arXiv:2310.17884 (2023)

  27. [30]

    arXiv preprint arXiv:2302.04023 (2023)

    Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., et al.: A multitask, multilingual, multimodal eval- uation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023 (2023)

  28. [31]

    Nature Medicine, 1–8 (2025)

    Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Amin, M., Hou, L., Clark,K.,Pfohl,S.R.,Cole-Lewis,H.,etal.:Towardexpert-levelmedicalquestion answering with large language models. Nature Medicine, 1–8 (2025)

  29. [32]

    arXiv preprint arXiv:2406.11149 (2024)

    Fan, W., Li, H., Deng, Z., Wang, W., Song, Y.: Goldcoin: Grounding large lan- guage models in privacy laws via contextual integrity theory. arXiv preprint arXiv:2406.11149 (2024)

  30. [34]

    arXiv preprint arXiv:2108.13888 (2021)

    Li, L., Song, D., Li, X., Zeng, J., Ma, R., Qiu, X.: Backdoor attacks on pre-trained models by layerwise weight poisoning. arXiv preprint arXiv:2108.13888 (2021)

  31. [35]

    arXiv preprint arXiv:2004.06660 (2020)

    Kurita, K., Michel, P., Neubig, G.: Weight poisoning attacks on pre-trained models. arXiv preprint arXiv:2004.06660 (2020)

  32. [36]

    In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp

    Zhang, Z., Ren, X., Su, Q., Sun, X., He, B.: Neural network surgery: Injecting data patterns into pre-trained models with minimal instance-wise side effects. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: ...

  33. [37]

    In: Proceedings of the 22nd ACMSIGSACConferenceonComputerandCommunicationsSecurity,pp.1322– 1333 (2015)

    Fredrikson, M., Jha, S., Ristenpart, T.: Model inversion attacks that exploit 25 confidence information and basic countermeasures. In: Proceedings of the 22nd ACMSIGSACConferenceonComputerandCommunicationsSecurity,pp.1322– 1333 (2015)

  34. [38]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Zhang, Y., Jia, R., Pei, H., Wang, W., Li, B., Song, D.: The secret revealer: Generative model-inversion attacks against deep neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 253–261 (2020)

  35. [39]

    arXiv preprint arXiv:2209.10505 (2022)

    Zhang, R., Hidano, S., Koushanfar, F.: Text revealer: Private text recon- struction via model inversion attacks against transformers. arXiv preprint arXiv:2209.10505 (2022)

  36. [40]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Truong, J.-B., Maini, P., Walls, R.J., Papernot, N.: Data-free model extraction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4771–4780 (2021)

  37. [41]

    arXiv preprint arXiv:2402.12959 (2024)

    Sha, Z., Zhang, Y.: Prompt stealing attacks against large language models. arXiv preprint arXiv:2402.12959 (2024)

  38. [42]

    Electronics13(14), 2858 (2024)

    He, J., Hou, G., Jia, X., Chen, Y., Liao, W., Zhou, Y., Zhou, R.: Data stealing attacks against large language models via backdooring. Electronics13(14), 2858 (2024)

  39. [43]

    In: 32nd USENIX Security Symposium (USENIX Security 23), pp

    Gao, X., Zhang, L.:{PCAT}: Functionality and data stealing from split learning by{Pseudo-Client} attack. In: 32nd USENIX Security Symposium (USENIX Security 23), pp. 5271–5288 (2023)

  40. [44]

    arXiv preprint arXiv:2405.05990 (2024)

    Bai, Y., Pei, G., Gu, J., Yang, Y., Ma, X.: Special characters attack: Toward scalable training data extraction from large language models. arXiv preprint arXiv:2405.05990 (2024)

  41. [45]

    arXiv preprint arXiv:2311.06062 (2023)

    Fu, W., Wang, H., Gao, C., Liu, G., Li, Y., Jiang, T.: Practical membership infer- ence attacks against fine-tuned large language models via self-prompt calibration. arXiv preprint arXiv:2311.06062 (2023)

  42. [46]

    Duan, M., Suri, A., Mireshghallah, N., Min, S., Shi, W., Zettlemoyer, L., Tsvetkov, Y., Choi, Y., Evans, D., Hajishirzi, H.: Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841 (2024)

  43. [47]

    In: 2021 IEEE European Symposium on Security and Privacy (EuroS&P), pp

    Zhao, B.Z.H., Agrawal, A., Coburn, C., Asghar, H.J., Bhaskar, R., Kaafar, M.A., Webb, D., Dickinson, P.: On the (in) feasibility of attribute inference attacks on machine learning models. In: 2021 IEEE European Symposium on Security and Privacy (EuroS&P), pp. 232–251 (2021). IEEE

  44. [48]

    ACM Transactions on Privacy and Security (TOPS)21(1), 1–30 (2018) 26

    Gong, N.Z., Liu, B.: Attribute inference attacks in online social networks. ACM Transactions on Privacy and Security (TOPS)21(1), 1–30 (2018) 26

  45. [49]

    arXiv preprint arXiv:2105.10909 (2021)

    Chen, C., He, X., Lyu, L., Wu, F.: Killing one bird with two stones: Model extraction and attribute inference attacks against bert-based apis. arXiv preprint arXiv:2105.10909 (2021)

  46. [50]

    In: 2017 IEEE Symposium on Security and Privacy (SP), pp

    Shokri, R., Stronati, M., Song, C., Shmatikov, V.: Membership inference attacks against machine learning models. In: 2017 IEEE Symposium on Security and Privacy (SP), pp. 3–18 (2017). IEEE

  47. [51]

    arXiv preprint arXiv:2108.07258 (2021)

    Bommasani, R., Hudson, D.A., Adeli, E., Altman, R., Arora, S., Arx, S., Bern- stein, M.S., Bohg, J., Bosselut, A., Brunskill, E., et al.: On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)

  48. [52]

    Journal of Energy Storage84, 110780 (2024)

    Chen, D., Zhou, X.: Attmoe: Attention with mixture of experts for remaining use- ful life prediction of lithium-ion batteries. Journal of Energy Storage84, 110780 (2024)

  49. [53]

    arXiv preprint arXiv:2304.14827 (2023)

    Chan, C., Cheng, J., Wang, W., Jiang, Y., Fang, T., Liu, X., Song, Y.: Chatgpt evaluation on sentence level relations: A focus on temporal, causal, and discourse relations. arXiv preprint arXiv:2304.14827 (2023)

  50. [54]

    arXiv preprint arXiv:2310.02469 (2023)

    Xiao, Y., Jin, Y., Bai, Y., Wu, Y., Yang, X., Luo, X., Yu, W., Zhao, X., Liu, Y., Gu, Q., et al.: Privacymind: large language models can be contextual privacy protection learners. arXiv preprint arXiv:2310.02469 (2023)

  51. [55]

    Apthorpe, N., Shvartzshnaider, Y., Mathur, A., Reisman, D., Feamster, N.: Discovering smart home internet of things privacy norms using contextual integrity.ProceedingsoftheACMoninteractive,mobile,wearableandubiquitous technologies 2(2), 1–23 (2018)

  52. [56]

    World Wide Web27(4), 42 (2024)

    Chen, J., Liu, Z., Huang, X., Wu, C., Liu, Q., Jiang, G., Pu, Y., Lei, Y., Chen, X., Wang, X.,et al.: When large language models meet personalization: Perspectives of challenges and opportunities. World Wide Web27(4), 42 (2024)

  53. [57]

    arXiv preprint arXiv:2212.10628 (2022)

    Yang, Y.: Holistic risk assessment of inference attacks in machine learning. arXiv preprint arXiv:2212.10628 (2022)

  54. [58]

    arXiv preprint arXiv:2406.18221 (2024)

    Venditti, D., Ruzzetti, E.S., Xompero, G.A., Giannone, C., Favalli, A., Romag- noli, R., Zanzotto, F.M.: Enhancing data privacy in large language models through private association editing. arXiv preprint arXiv:2406.18221 (2024)

  55. [59]

    IET Blockchain4, 706–724 (2024)

    Ullah, I., Hassan, N., Gill, S.S., Suleiman, B., Ahanger, T.A., Shah, Z., Qadir, J., Kanhere, S.S.: Privacy preserving large language models: Chatgpt case study based vision and framework. IET Blockchain4, 706–724 (2024)

  56. [60]

    arXiv preprint arXiv:2310.12214 (2023) 27

    Tong, M., Chen, K., Qi, Y., Zhang, J., Zhang, W., Yu, N.: Privinfer: Privacy-preserving inference for black-box large language model. arXiv preprint arXiv:2310.12214 (2023) 27

  57. [61]

    arXiv preprint arXiv:2402.08227 (2024)

    Yao, Y., Wang, F., Ravi, S., Chen, M.: Privacy-preserving language model inference with instance obfuscation. arXiv preprint arXiv:2402.08227 (2024)

  58. [63]

    arXiv preprint arXiv:2310.01467 (2023)

    Sun, J., Xu, Z., Yin, H., Yang, D., Xu, D., Chen, Y., Roth, H.R.: Fedbpt: Efficient federated black-box prompt tuning for large language models. arXiv preprint arXiv:2310.01467 (2023)

  59. [64]

    IEEE Internet of Things Journal 9(4), 2555–2565 (2021)

    Yuan, X., Ma, X., Zhang, L., Fang, Y., Wu, D.: Beyond class-level privacy leak- age: Breaking record-level privacy in federated learning. IEEE Internet of Things Journal 9(4), 2555–2565 (2021)

  60. [66]

    arXiv preprint arXiv:2307.08925 (2023)

    Chen, C., Feng, X., Zhou, J., Yin, J., Zheng, X.: Federated large language model: A position paper. arXiv preprint arXiv:2307.08925 (2023)

  61. [68]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

    Zhu,M.,Wei,S.,Shen,L.,Fan,Y.,Wu,B.:Enhancingfine-tuningbasedbackdoor defense with sharpness-aware minimization. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4466–4477 (2023)

  62. [69]

    In: International Symposium on Research in Attacks, Intrusions, and Defenses, pp

    Liu, K., Dolan-Gavitt, B., Garg, S.: Fine-pruning: Defending against backdooring attacks on deep neural networks. In: International Symposium on Research in Attacks, Intrusions, and Defenses, pp. 273–294 (2018). Springer

  63. [71]

    arXiv preprint arXiv:2205.03040 (2022)

    Dong, C., Weng, J., Liu, J.-N., Zhang, Y., Tong, Y., Yang, A., Cheng, Y., Hu, S.: Fusion: Efficient and secure inference resilient to malicious servers. arXiv preprint arXiv:2205.03040 (2022)

  64. [72]

    In: Findings of the Association for Computational Linguistics ACL 2024, pp

    Luo, J., Zhang, Y., Zhang, Z., Zhang, J., Mu, X., Wang, H., Yu, Y., Xu, Z.: Secformer: Fast and accurate privacy-preserving inference for transformer models via smpc. In: Findings of the Association for Computational Linguistics ACL 2024, pp. 13333–13348 (2024)

  65. [73]

    In: 32nd USENIX Security Symposium (USENIX Security 23), pp

    Chen, H., Chen, H.H., Sun, M., Li, K., Chen, Z., Wang, X.: A verified confidential 28 computing as a service framework for privacy preservation. In: 32nd USENIX Security Symposium (USENIX Security 23), pp. 4733–4750 (2023)

  66. [74]

    arXiv preprint arXiv:2401.09796 (2024)

    Huang, W., Wang, Y., Cheng, A., Zhou, A., Yu, C., Wang, L.: A fast, performant, secure distributed training framework for large language model. arXiv preprint arXiv:2401.09796 (2024)

  67. [75]

    1450–1465 (2020)

    Zhu, J., Hou, R., Wang, X., Wang, W., Cao, J., Zhao, B., Wang, Z., Zhang, Y., Ying, J., Zhang, L.,et al.: Enabling rack-scale confidential computing using het- erogeneoustrustedexecutionenvironment.In:2020IEEESymposiumonSecurity and Privacy (SP), pp. 1450–1465 (2020). IEEE

  68. [76]

    arXiv preprint arXiv:2110.06500 (2021)

    Yu, D., Naik, S., Backurs, A., Gopi, S., Inan, H.A., Kamath, G., Kulkarni, J., Lee, Y.T., Manoel, A., Wutschitz, L., et al.: Differentially private fine-tuning of language models. arXiv preprint arXiv:2110.06500 (2021)

  69. [77]

    ACM Computing Surveys (Csur) 51(4), 1–35 (2018)

    Acar, A., Aksu, H., Uluagac, A.S., Conti, M.: A survey on homomorphic encryp- tion schemes: Theory and implementation. ACM Computing Surveys (Csur) 51(4), 1–35 (2018)

  70. [78]

    In: Annual International Conference on the Theory and Applications of Cryptographic Techniques, pp

    Boyle, E., Gilboa, N., Ishai, Y.: Function secret sharing. In: Annual International Conference on the Theory and Applications of Cryptographic Techniques, pp. 337–367 (2015). Springer

  71. [79]

    In: 2015 IEEE Trustcom/BigDataSE/Ispa, vol

    Sabt, M., Achemlal, M., Bouabdallah, A.: Trusted execution environment: What it is, and what it is not. In: 2015 IEEE Trustcom/BigDataSE/Ispa, vol. 1, pp. 57–64 (2015). IEEE

  72. [80]

    arXiv preprint arXiv:2412.11640 (2024)

    Hu, G., Wu, Y., Chen, G., Dinh, T.T.A., Ooi, B.C.: Sesemi: Secure serverless model inference on sensitive data. arXiv preprint arXiv:2412.11640 (2024)

  73. [81]

    Transactions of the Association for Computational Linguistics 12, 1556–1577 (2024)

    Zhu, X., Li, J., Liu, Y., Ma, C., Wang, W.: A survey on model compression for large language models. Transactions of the Association for Computational Linguistics 12, 1556–1577 (2024)

  74. [82]

    arXiv preprint arXiv:2406.10861 (2024)

    Qin, L., Zhu, T., Zhou, W., Yu, P.S.: Knowledge distillation in federated learn- ing: A survey on long lasting challenges and new solutions. arXiv preprint arXiv:2406.10861 (2024)

  75. [83]

    In: Artificial Intelligence and Statistics, pp

    McMahan, B., Moore, E., Ramage, D., Hampson, S., Arcas, B.A.: Communication-efficient learning of deep networks from decentralized data. In: Artificial Intelligence and Statistics, pp. 1273–1282 (2017). PMLR

  76. [84]

    ACM Transactions on Modeling and Performance Evaluation of Computing Systems (2024) 29

    Chu, T., Yang, M., Laoutaris, N., Markopoulou, A.: Priprune: Quantifying and preserving privacy in pruned federated learning. ACM Transactions on Modeling and Performance Evaluation of Computing Systems (2024) 29

  77. [85]

    arXiv preprint arXiv:2305.10235 (2023)

    Ye, W., Ou, M., Li, T., Ma, X., Yanggong, Y., Wu, S., Fu, J., Chen, G., Wang, H., Zhao, J., et al.: Assessing hidden risks of llms: an empirical study on robustness, consistency, and credibility. arXiv preprint arXiv:2305.10235 (2023)

  78. [86]

    Expert Systems with Applications260, 125348 (2025)

    Su, J., Xu, B., Jiang, L., Liu, H., Chen, Y., Li, Y.,et al.: Cross-organizational knowledge sharing partner selection based on fogg behavioral model in probabilis- tic hesitant fuzzy environment. Expert Systems with Applications260, 125348 (2025)

  79. [87]

    Security and Safety1, 2021001 (2022)

    Feng,D.,Yang,K.:Concretelyefficientsecuremulti-partycomputationprotocols: survey and more. Security and Safety1, 2021001 (2022)

  80. [88]

    IEEE network35(4), 198–205 (2021)

    Sun, X., Yu, F.R., Zhang, P., Sun, Z., Xie, W., Peng, X.: A survey on zero- knowledge proof in blockchain. IEEE network35(4), 198–205 (2021)

  81. [89]

    In: Proceedings of the 2024 International Conference on Parallel Architectures and Compilation Techniques, pp

    Daftardar, A., Reagen, B., Garg, S.: Szkp: A scalable accelerator architecture for zero-knowledge proofs. In: Proceedings of the 2024 International Conference on Parallel Architectures and Compilation Techniques, pp. 271–283 (2024)

  82. [90]

    Computers and Electrical Engineering 120, 109698 (2024)

    Kibriya, H., Khan, W.Z., Siddiqa, A., Khan, M.K.: Privacy issues in large lan- guage models: a survey. Computers and Electrical Engineering 120, 109698 (2024)

  83. [91]

    A Practical Guide, 1st Ed., Cham: Springer International Publishing10(3152676), 10–5555 (2017)

    Voigt, P., Bussche, A.: The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing10(3152676), 10–5555 (2017)

  84. [92]

    arXiv preprint arXiv:2408.12787 (2024) 30

    Li, Q., Hong, J., Xie, C., Tan, J., Xin, R., Hou, J., Yin, X., Wang, Z., Hendrycks, D., Wang, Z., et al.: Llm-pbe: Assessing data privacy in large language models. arXiv preprint arXiv:2408.12787 (2024) 30

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.