{"id":"d2012bce-15fe-4d1f-be0d-56f747f213f1","arxiv_id":"2505.01976","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper surveys LLM privacy leaks and attacks, organizes them into a taxonomy, and reviews defenses without adding new empirical results.","lead":"This survey sorts known privacy risks of large language models into categories and pairs them with defenses. It is a reference map, not a new experiment or a new attack.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 miscategorizes non-privacy papers as Sensitive Information and Personal Preferences Leakage, so the survey's central taxonomy cannot currently be trusted as an accurate map of the literature.","rationale":"The reader's weakest assumption is that the papers selected and their assignments in Tables 1-4 accurately represent the literature. My review confirms that this assumption is violated in Table 1: at least three cited papers are not privacy-leakage studies, including refs [28], [29], and [33]. This is directly load-bearing because the survey's central claim is that its taxonomy provides a comprehensive and accurate organization of LLM privacy risks; if category assignments are unreliable, the taxonomy cannot fulfill that role. No independent re-derivation is needed because the cited papers are public and their abstracts are sufficient to test the assignments. I agree with the reader that the issues are fixable: replacing or relabeling misassigned entries and adding a methodology section with search and inclusion criteria would materially improve the survey. Since the reader's recommendation is CONDITIONAL and my analysis supports that same verdict, the appropriate output is UNCHANGED rather than a new verdict. I did not find evidence of deeper internal inconsistency in the defense taxonomy or in the attack categories beyond the Table 1 leakage misassignments, so no stronger rejection is warranted.","tokens_in":19478,"tokens_out":5055,"duration_ms":50694,"concrete_test":"For every entry in Table 1, retrieve the cited paper's abstract and determine whether the paper's primary contribution concerns privacy leakage, defined as the unauthorized exposure of private information through an LLM. Specifically, check refs [28], [29], [30], and [33] against this criterion; for example, [28] and [29] should be labeled non-privacy if their abstracts contain no privacy-leakage mechanism or privacy risk claim. Then recompute the fraction of Table 1 entries that are misassigned. If more than two of the seven Table 1 entries fail the privacy-leakage test, the category assignments should be revised and the survey should include an explicit inclusion criterion for cited works.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it offers a comprehensive overview of LLM privacy risks organized into a taxonomy of privacy leakage and privacy attacks. That claim requires the table entries to be faithful assignments of the cited literature to the stated categories. Table 1 fails this requirement in its first category: 'Sensitive Information Leakage' lists [28] Zamfirescu-Pereira et al. (CHI 2023), a human-factors study of prompt design by non-AI experts, and [29] Wang et al., a study of ChatGPT's adversarial and out-of-distribution robustness. Neither paper addresses privacy leakage in LLMs. The main text cites [28] and [29] only to support general capability statements ('tasks such as programming, academic writing, and medical diagnosis' and 'LLMs have achieved significant performance with various NLP tasks'), not to support the leakage category. Similarly, Table 1's 'Personal Preferences Leakage' includes [33] Thomas et al., an information-retrieval paper on predicting searcher preferences, which does not study privacy leakage. The absence of stated inclusion or selection criteria means these are not easily dismissed as isolated typos; they indicate that category assignments were not systematically derived from the cited content. Because the taxonomy is the survey's main contribution, a reader relying on it to organize the field will be misled about which papers establish each risk category. The taxonomy may be salvageable through corrections, but as written the central organizational claim is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of privacy risks and protections for large language models. It proposes a taxonomy in which privacy issues are divided into privacy leakage (sensitive information leakage, contextual leakage, personal preferences leakage) and privacy attacks (model-based attacks: backdoor, model inversion, model stealing; data-based attacks: data stealing, training data extraction; user-based attacks: membership inference, attribute inference). It then reviews defense mechanisms including data cleaning, inference detection, federated learning, differential privacy, backdoor removal, cryptography, and confidential computing, and concludes with future research directions. The paper's stated contribution is a comprehensive and fine-grained classification that integrates privacy concerns and provides a roadmap for addressing them.","tokens_in":19731,"tokens_out":6061,"duration_ms":55514,"significance":"If the survey's taxonomy and literature assignments were accurate, the paper would be a useful organizing resource for researchers entering the LLM-privacy area: it covers a wide range of attacks and defenses, connects privacy leakage to contextual integrity, reproduces several key attack equations from the primary literature, and outlines sensible future directions such as privacy-preserving model compression, privacy risk assessment, and secure knowledge sharing. The main contribution is the taxonomy and the associated reference tables, so the correctness of those assignments is load-bearing. The miscategorizations identified below in Table 1 are not cosmetic; they directly affect whether a reader can trust the paper as a map of the field. The absence of a stated methodology for selecting and classifying the 92 references further weakens the comprehensiveness claim. The paper is salvageable with a systematic audit and a methodology statement, but in its current form the central empirical basis for the taxonomy is unreliable.","major_comments":[{"comment":"Table 1 lists [28] Zamfirescu-Pereira et al. (CHI 2023) and [29] Wang et al. under 'Sensitive Information Leakage' as works establishing that category. [28] is a human-factors study of how non-AI experts design prompts, and [29] studies adversarial and out-of-distribution robustness of ChatGPT. Neither paper studies privacy leakage. The main text confirms this: [28] is cited only to support the general claim that chat-based interaction is useful for 'programming, academic writing, and medical diagnosis,' and [29] is cited only to support the claim that LLMs 'have achieved significant performance with various NLP tasks.' These are capability citations, not privacy-leakage citations, so the table entry misrepresents the evidence base for the category.","section":"Table 1, Section 3.1"},{"comment":"Table 1 places [33] Thomas et al. under 'Personal Preferences Leakage.' That paper is an information-retrieval study showing that LLMs can accurately predict searcher preferences; it does not demonstrate or measure privacy leakage. The text infers a privacy risk from the model's predictive ability, but presenting this work as an instance of a privacy-leakage category overstates what the cited paper establishes. At minimum the category needs a clearly labeled 'potential risk' entry, not a direct assignment.","section":"Table 1, Section 3.3"},{"comment":"The survey does not state its search strategy, inclusion/exclusion criteria, databases, or time window for the 92 references, despite advertising a 'comprehensive overview.' Given that Table 1 demonstrably contains citations that do not support the categories to which they are assigned, the absence of a documented methodology makes it impossible for a reader to determine whether the remaining assignments were derived systematically or selectively. The comprehensiveness claim therefore needs either a methodology subsection or a substantial softening.","section":"Section 3, Tables 1-4"},{"comment":"Table 4 appears to repeat the misassignment problem in the defense taxonomy: [67] is a fine-tuning-based backdoor defense, yet in the extracted table layout it is placed under 'Differential Privacy,' while [70] (THE-X) is a homomorphic-encryption method for transformer inference, yet it appears in the 'Backdoor Removal' row. If this is an artifact of the table's typesetting and column alignment, the table needs clearer category separators; if it is accurate, it provides further evidence that the category assignments are unreliable. Either way, the table as presented requires correction.","section":"Table 4, Section 4.2"}],"minor_comments":[{"comment":"Several equations are under-specified for a self-contained survey. In Eq. (2), Rl_c appears without definition; in Eq. (5), the symbols Xb, Sy, Tpre, and cPθ are not defined; and Eq. (3) contains a formatting artifact ('M AX(L)') that obscures the intended expression. Quoting equations from primary papers is acceptable, but each symbol should be defined or a precise reference given.","section":"Equations (1)-(7)"},{"comment":"References [62] and [65] are the same paper (Kim et al., ProPILE) and should be consolidated into a single entry to avoid confusing duplicate citations in Table 3 and the text.","section":"References [62] and [65]"},{"comment":"The sentence 'Some users believe that the information they provide is stored in the ChatGPT database...' is cited to [24], but [24] (Carlini et al.) is a training-data-extraction paper and does not report on user beliefs. A more appropriate citation would be [25] (Kshetri) or a user-survey source.","section":"Section 3.1, citation [24]"},{"comment":"The manuscript contains multiple LaTeX and typesetting artifacts, including 'T able' in table captions, missing spaces in phrases such as 'PersonalPreferencesLeakage,' and a stray '&' in Eq. (5). A careful proofreading pass is needed.","section":"General presentation"},{"comment":"For a survey paper, the statement that datasets 'are available from the corresponding author upon reasonable request' is misleading; the paper analyzes no datasets. This should be replaced with a statement that no new data were generated or analyzed.","section":"Data availability statement"}],"recommendation":"major_revision","confidential_remarks":"The number and pattern of citation-content mismatches in the tables is above the level I would treat as isolated typos. Before the paper can be considered for publication, the authors should perform a systematic audit of every table entry, verify that each cited work actually supports the category to which it is assigned, and add a methodology section describing how references were selected and classified. If the taxonomy is the main selling point, the tables are the evidence for it, and they currently contain demonstrable errors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This survey organizes LLM privacy into leakage (sensitive information, contextual, preferences) and attacks (model-, data-, user-based), then pairs each with defenses. The structure is genuinely readable, and the defense tables cover a wider range of techniques than some competing surveys. As a first stop for a practitioner wanting one overview, it has real value. That is the good part.\n\nThe soft spot is exactly what the stress-test flags. Table 1 places [28] (a prompt-design human-factors paper) and [29] (ChatGPT robustness) under Sensitive Information Leakage; in the text those citations only support general capability statements. [33] is an IR paper on predicting searcher preferences, not privacy leakage, and ends up under Personal Preferences Leakage. These are not isolated typos: they show the category assignments were not actually derived from the cited content. Because the taxonomy is the paper's main contribution, a reader using the table to navigate the literature will be misled. Also, [62] and [65] are the same paper (ProPILE) listed twice, and there is no stated inclusion or search methodology to back the comprehensiveness claim.\n\nThe equations quoted from primary sources are fine for a survey, though they add little. The defense sections are uneven in depth, but representative. I would not call the central idea wrong: a unified leakage/attack distinction is plausible and the paper holds together conceptually. The problems are in the execution, not the skeleton.\n\nBottom line: this deserves serious peer review, but as submitted it needs major revision—fix the tables, reconcile duplicate references, and add a short methodology paragraph saying how papers were selected. If those are addressed, it becomes a citable reference. As is, I would not put it in my own bibliography, but I would send it back for repair rather than reject it outright.","headline":"A broadly useful but sloppy survey: the taxonomy is clear, but Table 1 misassigns several papers to privacy categories, which undermines the central claim of a trusted map of the field.","tokens_in":20253,"tokens_out":1726,"would_cite":false,"duration_ms":19393,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims LLM privacy research can be organized by a unified classification of privacy leakage and privacy attacks.","keywords":["large language models","privacy leakage","privacy attacks","membership inference","model inversion","backdoor attacks","federated learning","differential privacy"],"falsifier":"Take the entries in Tables 1 through 4, read the abstract of each cited paper, and check whether the method and evaluation metric match the assigned category; several entries already listed under sensitive-information leakage are prompt-design and robustness studies, so if a substantial fraction of the other entries also fail this check, the taxonomy does not accurately map the field.","tokens_in":19267,"feed_emoji":"🔐","tokens_out":6780,"duration_ms":65004,"temperature":0.7,"pith_summary":"The paper tries to establish that the scattered privacy literature on large language models fits one organizing scheme: privacy leakage, which covers sensitive information, contextual information, and personal preferences, and privacy attacks, which target models, data, or users. It reviews existing defenses and maps them to these categories, arguing that the classification gives the field a shared vocabulary and a roadmap for future work. If the taxonomy is right, researchers and practitioners can file a reported vulnerability into a category and know which mitigation family has been studied for it.","feed_headline":"One taxonomy sorts LLM privacy risks into leaks and attacks","feed_subtitle":"The survey maps eleven risks to defense families, giving practitioners a shared vocabulary for LLM privacy.","key_machinery":"The central object is the taxonomy itself: a two-level classification of privacy risks according to how an attacker gains access to sensitive information. Privacy leakage names passive exposure through the model's ordinary behavior, while privacy attacks name active attempts to break into the model, the data, or the user. The taxonomy organizes the survey: each cell carries definitions, representative works, evaluation metrics, and candidate mitigations, so the entire review is structured by the classification.","core_discovery":"The central claim is that LLM privacy work is not a set of isolated problems but a single two-branch landscape. Privacy leakage is the exploitation of LLM vulnerabilities to collect sensitive information, divided into sensitive information leakage, contextual leakage, and personal preferences leakage. Privacy attacks are active attempts to breach the model's defenses, divided into model-based attacks (backdoor, model inversion, model stealing), data-based attacks (data stealing, training data extraction), and user-based attacks (membership inference, attribute inference). The survey then pairs each category with representative papers and defense strategies, including data cleaning, inference detection, federated learning, differential privacy, backdoor removal, cryptography, and confidential computing, and states that these defenses correspond to the risks in the taxonomy.","pith_inferences":["The boundary between leakage and attack is likely fuzzier in practice than the taxonomy suggests, since the same model behavior can be read as passive exposure or active exfiltration depending on the threat model and who initiates the interaction.","Because the survey does not state its search and inclusion criteria, the taxonomy's completeness is only as strong as the chosen sample of papers; a formal selection procedure and an audit of the table assignments would make the roadmap reproducible.","The contextual leakage category points toward a testable extension: privacy evaluation should compare information flows that have the same content but different recipients, senders, or transmission principles, rather than only measuring how much information is revealed.","If the taxonomy is adopted as a standard, it could be encoded as a machine-readable checklists for audit reports, letting regulators see which leakage and attack categories a model has actually been evaluated against."],"forward_implications":["A practitioner who encounters a reported privacy incident can assign it to one leakage or attack category and immediately see which defense family has been tested against that category.","Privacy risk assessment frameworks for LLMs should measure the three leakage channels and the three attack targets named in the taxonomy, because the survey argues those cover the current literature.","Benchmarks for LLM privacy, which the survey calls for, would be comparable across papers if they are organized by the same categories.","The leakage-versus-attack split gives a natural mapping to governance: leakage concerns data-protection duties, while attacks concern adversarial threat modeling."],"supporting_citations":[{"why":"Supplies the broad survey of LLM data-privacy protection that many attack and defense definitions in this paper are drawn from.","marker":"[14]"},{"why":"Provides a baseline security-and-privacy survey that the paper contrasts with its unified classification.","marker":"[18]"},{"why":"Offers the good-bad-ugly categorization of LLM security that this taxonomy extends by focusing specifically on privacy.","marker":"[19]"},{"why":"Demonstrates verbatim extraction of training data from language models, underpinning the training-data-extraction and sensitive-information-leakage categories.","marker":"[24]"},{"why":"Brings contextual integrity theory to LLMs and supplies the benchmark behind the contextual leakage category and the inference-detection defense.","marker":"[27]"},{"why":"Foundational model-inversion attack that establishes the model-based attack family.","marker":"[37]"},{"why":"Foundational membership-inference attack that defines the user-based attack family.","marker":"[50]"},{"why":"Shows that LLMs infer personal attributes from unstructured text, supporting the personal-preferences-leakage and attribute-inference categories.","marker":"[10]"}],"fun_headline_variants":["LLM privacy: one taxonomy, two attack fronts","Mapping LLM privacy: leakage vs. attacks","A single map for all LLM privacy risks and fixes","From leaks to attacks: a unified LLM privacy map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness depends on each cited paper actually belonging to the category it is placed in, and since the survey gives no search or inclusion criteria, a reader cannot verify that the table assignments faithfully represent the literature.","fun_headline_variants_meta":{"raw":{"variants":["LLM privacy: one taxonomy, two attack fronts","Mapping LLM privacy: leakage vs. attacks","A single map for all LLM privacy risks and fixes","From leaks to attacks: a unified LLM privacy map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000164,"raw_usage":{"total_tokens":1206,"prompt_tokens":862,"completion_tokens":344,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":280}},"tokens_in":478,"tokens_out":344,"duration_ms":3871,"temperature":1.0,"reasoning_tokens":280,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:04:04.759681+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the entries in Tables 1 through 4, read the abstract of each cited paper, and check whether the method and evaluation metric match the assigned category; several entries already listed under sensitive-information leakage are prompt-design and robustness studies, so if a substantial fraction of the other entries also fail this check, the taxonomy does not accurately map the field.","supporting_citations":[{"cited_title":"In: Proceedings of the 22nd ACMSIGSACConferenceonComputerandCommunicationsSecurity,pp.1322– 1333 (2015)","cited_arxiv_id":null,"evidence_quote":"Foundational model-inversion attack that establishes the model-based attack family."}],"review_version":1}