Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read All six major U.S. chatbot developers train on user chats by default, a 2025 policy review finds.

desk verdict Useful, carefully hedged policy survey with a real internal inconsistency: Table 1 codes Amazon 'Not Specified' for default chat training, yet the abstract and analysis say 'all six' train by default. read the letter →

arxiv 2509.05382 v1 pith:OHDVYJEX submitted 2025-09-05 cs.CY cs.AIcs.CR

classification cs.CYcs.AIcs.CR
keywords privacypolicieslargelanguagemodelschatbotstrainingdataconsentretentionCCPAchildren's
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Millions of people talk to chatbots, and those conversations can contain health details, biometric data, family information, and uploaded files. This paper asks whether the six largest U.S. frontier AI developers disclose in their privacy policies that they use those chats to train their models, and whether users can opt out. Working from the California Consumer Privacy Act's disclosure requirements, the authors coded 28 policy documents and linked notices as they stood in May 2025. They find that all six developers appear to use chat inputs for model training by default: Anthropic had just switched from opt-in to opt-out, Google and Meta offer no clear opt-out, and Amazon's written policies are silent, with training inferred from an interface notice. Three companies—Amazon, Meta, and OpenAI—appear to retain some chat data indefinitely, and four appear to include children's chat data in training. The paper's point is that privacy policies, the only comprehensive disclosure mechanism in U.S. law, are not telling consumers what happens to their chats.

What carries the argument

The analysis runs on a qualitative coding schema built from the CCPA's disclosure categories (categories of personal information collected, sources, business purposes, retention, sharing) plus LLM-specific categories such as default training, opt-out presence, human review, de-identification, and personalization. The schema was applied manually to 28 documents: each developer's main privacy policy, and every linked 'branch' sub-policy, FAQ, or interface notice needed to learn the chatbot's actual data practices. The key interpretive move is treating silence or 'not specified' as meaningful—so Amazon's default-training classification rests on a notice displayed in the Nova interface rather th

What would settle it

Create a fresh account at each developer with default settings, send a unique canary phrase in a chat, then submit a deletion request and probe publicly available versions of the models for memorization of that phrase; if the phrase appears, the company trained on default chats despite claiming not to, and if it never appears in any model, the default-training claim for that developer is weakened. For Amazon, checking whether the Nova interface notice appears without an account and whether the AWS privacy policy acknowledges training would directly test the paper's inference.

Watch

Extended reading notes

Core claim

The central claim is that, as of May 2025, every one of the six U.S. frontier chatbot developers—Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI—uses consumers' chat inputs to train and improve its AI models by default, with no affirmative consent required. Only some offer opt-out mechanisms; Anthropic's move from opt-in to opt-out closed the last opt-in default among the six. The paper also claims that Amazon, Meta, and OpenAI appear to retain chat data indefinitely, and that OpenAI, Meta, and Amazon likely include chats from users aged 13-17, while Google trains on teen chats only with opt-in and Microsoft excludes authenticated minors. The authors stress that these conclusions come

Load-bearing premise

The findings depend on privacy-policy text and linked notices being a complete and truthful map of what companies actually do with chat data, and on the Anthropic policy change being read as a switch from opt-in to opt-out; the manuscript's footnote states the opposite direction, so if that footnote is right, the 'all six train by default' claim loses one of its six data points.

Editorial extensions

If this is right

  • Opt-out, not opt-in, is the industry default: consumers who want their chats excluded must find and use a settings control, and Google and Meta offer no clear route at all.
  • Because chats can contain sensitive personal data and uploaded files, default training means highly revealing material enters training corpora without affirmative consent.
  • Indefinite retention at Amazon, Meta, and OpenAI creates a long-lived target for breaches and a permanent dossier of a user's conversations.
  • Children aged 13-17 using OpenAI, Meta, and Amazon are likely included in training data by default, raising consent questions that existing U.S. law does not answer.
  • The two-tier arrangement that excludes enterprise customers from default training while applying it to consumers is a direct consequence of current defaults.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: if privacy policies remain the only public record, the May 2025 snapshot likely understates current practice, because policies change faster than papers publish; the authors' own note that Anthropic flipped its default shows the clock running.
  • My inference: the paper's method implies a cheap test for regulators—request the CCPA-mandated disclosures from each developer and compare them to what the policies say; discrepancies would expose which companies are training without documenting it.
  • My inference: the 'not specified' categories for Amazon suggest that companies may be choosing silence strategically, so an enforcement action requiring explicit statements about training defaults could shift industry practice without new legislation.
  • Manuscript note: the paper's footnote about Anthropic states the direction of the policy change as opt-out-to-opt-in, while the body text states opt-in-to-opt-out; the 'all six train by default' conclusion depends on the body-text reading, so a reader should verify which direction is accurate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript analyzes the privacy policies, linked sub-policies, and interface notices of six U.S. frontier LLM developers (Amazon, Anthropic, Google, Meta, Microsoft, OpenAI) as of May 2025, using an inductive coding schema grounded in the CCPA and selected GDPR principles. It asks whether chatbot inputs and outputs are used for model training by default, what additional user data sources feed training, what opt-out mechanisms exist, how long chat data is retained, and how children's data are handled. The authors report that all six developers train on user chat data by default (after incorporating Anthropic's announced switch to opt-out), that some retain chat data indefinitely, and that four appear to train on children's chat data. They argue that developers' privacy policies are fragmented and incomplete, discuss normative implications of default training, indefinite retention, and children's data, and close with five policy recommendations.

Significance. If the findings are correct, the study is a timely and useful empirical contribution to AI privacy governance. It goes beyond general critiques of web scraping by focusing on the direct flow of consumer chat data into model training, and it compares six major firms with a consistent instrument. The paper's strengths include a transparent company-selection protocol, dual manual coding with cross-review, a clear external legal benchmark (CCPA), explicit limitation statements, and a decision not to use LLMs for coding. The policy recommendations are concrete and actionable. The main caveat is that the contribution's value depends on the coding being internally consistent and verifiable; in the current version, several table entries contradict the text, so the headline claims need strengthening before the empirical contribution can be fully credited.

major comments (4)
  1. [Abstract; §Analysis; Table 1] The headline claim that 'all six developers appear to employ their users' chat data to train and improve their models by default' (Abstract) is contradicted by Table 1, which codes Amazon's 'Chat input used for training by default' as 'Not Specified.' The §Analysis support is a Nova interface notice stating that interactions 'may be reviewed and retained' to 'provide, develop, and improve our services, including AI models.' This is permissive language about review and retention, not a disclosure of default training. Under the paper's own coding definition, ambiguous or undisclosed practices are 'Not Specified.' Since excluding Amazon reduces the finding to five of six, this classification is load-bearing and must be corrected or the claim narrowed.
  2. [Methodology; §Analysis; Discussion] The methodology fixes the policy snapshot at May 2025, and the §Analysis notes that Anthropic's privacy policy at that time required explicit opt-in for training. The Discussion then uses Anthropic's August/September 2025 announcement to assert that 'all six of the developers now train on their users' chat data by default.' The Abstract repeats 'all six' without a temporal qualifier. This mixes the analyzed snapshot with a later, announced state. Please state the date of the 'all six' claim explicitly and clearly separate the May 2025 snapshot from post-publication updates.
  3. [Children's data; Table 1] The children's data section says Amazon, Meta, and OpenAI allow users 13 and older and 'therefore likely train on it by default,' supporting the Abstract's 'four of six' count. Table 1, however, codes Amazon as 'No' for 'Allows accounts for children 13-18.' If Amazon does not allow minors, it should not be counted as training on children's data; if it does, Table 1 is wrong. This contradiction directly affects the reported number and must be resolved.
  4. [Policy selection protocol; Analysis] The empirical analysis is based on summaries of policy language, but the manuscript provides no direct quotations, URLs, or a shipped coding instrument/appendix. The crucial Amazon/Nova notice is paraphrased but not quoted, and references list 'Accessed: 2025-05-23' without links. For an empirical policy study, readers need to be able to check the coding. Please add a supplemental appendix with the codebook, relevant policy excerpts, URLs, and the coding decisions per company.
minor comments (4)
  1. [Footnote 1] Footnote 1 says Anthropic's policy update went 'from opt-out for model training to opt-in,' but the body text correctly says the change was from opt-in to opt-out. Reverse the wording.
  2. [Data retention] The text says 'Every chatbot developer that we analyzed appears to retain some chat data indefinitely,' which is inconsistent with Table 1 (only Amazon, Meta, and OpenAI are coded 'Yes' for indefinite retention) and with the later sentence 'Amazon, Meta, and OpenAI retain... indefinitely.' Clarify the definition of 'indefinitely' versus 'long-term' and align text with table.
  3. [Additional user-provided data sources; Table 2] The text says 'OpenAI, Google, and Microsoft also state that they may train on user-uploaded images,' but Table 2 codes Google as 'No' and Microsoft and OpenAI as 'Not Specified' for uploaded images. Align the text with Table 2 or correct the table.
  4. [References] Many references lack URLs despite 'Accessed' dates, and Google's two 2025 entries are both cited as (Google 2025), requiring 2025a/2025b disambiguation. Also fix the garbled 'Turow, Hennessy, and and 2018' reference and the malformed Staab et al. entry.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the paper's empirical findings are grounded in external policy documents and company notices, not in its own framework or prior self-citations.

full rationale

The paper's central claims—that frontier developers use chat data for training by default, retain data indefinitely, and lack clear opt-outs—are drawn from an inductive coding of 28 externally published privacy policies, sub-policies, FAQs, and interface notices as of May 2025. The CCPA is used as an external normative benchmark for evaluating comprehensiveness, not as a source from which the findings are derived. The authors' self-citations (e.g., Bommasani et al. 2023/2024/2025, King & Meinhardt 2024, Shen et al. 2025) appear only in background statements or in recommendations, and none is load-bearing for the empirical classification of any developer's practices. No equation, fitted parameter, or definitional identity is used to generate a predicted result from an input. The paper does contain an internal inconsistency—Table 1 codes Amazon's 'Chat input used for training by default' as 'Not Specified,' while the abstract and discussion count Amazon among 'all six' developers that train by default—and the Amazon classification rests on an interpretive reading of a Nova interface notice ('may be reviewed and retained' used to 'provide, develop, and improve our services, including AI models'). That is a validity and precision concern, not circularity: the claim is not equivalent to its input by construction, and the discrepancy does not involve the authors' framework substituting for external evidence. Similarly, the inference that Amazon, Meta, and OpenAI 'likely' train on children's chats because they allow 13+ accounts is an inference from policy facts, not a circular derivation. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This study does not introduce free parameters or invented entities. Its entire evidentiary base is qualitative: the text of privacy policies and sub-policies. The main epistemic commitments are that policy documents are truthful and comprehensive evidence of data practices, that absence of disclosure can be interpreted as absence of a user-visible opt-out, and that the CCPA provides the right normative lens for comparison.

assumptions (3)
  • domain assumption Privacy policies and linked sub-policies are comprehensive and accurate disclosures of companies' actual data collection and training practices.
    Introduced in the Introduction ('privacy policies remain the core resource for consumers seeking answers') and Methodology (treating policy text as evidence; coding 'Not Specified' as absence of a disclosed practice).
  • domain assumption Absence of a disclosed opt-out or exclusion implies that a practice, such as default training or training on children's chat data, is present.
    Used in the Analysis section when inferring 'all six train by default' from a mix of explicit statements and silence, and when inferring from account age policies that Amazon, Meta, and OpenAI 'likely' train on children's chats.
  • domain assumption The California Consumer Privacy Act is the appropriate normative framework for evaluating U.S. chatbot privacy policies.
    Methodology states the coding schema is 'grounded primarily' in CCPA because it is the most comprehensive U.S. law and all six companies serve California consumers; GDPR is mentioned but not the primary lens.

how reviews work

0 comments
Cite this review

Pith. "Pith review of User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies." pith.science (2026). https://pith.science/paper/OHDVYJEX

@misc{pith2026250905382,
  author       = {Pith},
  title        = {Pith review of: User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OHDVYJEX}},
  note         = {Machine review of arXiv:2509.05382}
}
read the original abstract

Hundreds of millions of people now regularly interact with large language models via chatbots. Model developers are eager to acquire new sources of high-quality training data as they race to improve model capabilities and win market share. This paper analyzes the privacy policies of six U.S. frontier AI developers to understand how they use their users' chats to train models. Drawing primarily on the California Consumer Privacy Act, we develop a novel qualitative coding schema that we apply to each developer's relevant privacy policies to compare data collection and use practices across the six companies. We find that all six developers appear to employ their users' chat data to train and improve their models by default, and that some retain this data indefinitely. Developers may collect and train on personal information disclosed in chats, including sensitive information such as biometric and health data, as well as files uploaded by users. Four of the six companies we examined appear to include children's chat data for model training, as well as customer data from other products. On the whole, developers' privacy policies often lack essential information about their practices, highlighting the need for greater transparency and accountability. We address the implications of users' lack of consent for the use of their chat data for model training, data security issues arising from indefinite chat data retention, and training on children's chat data. We conclude by providing recommendations to policymakers and developers to address the data privacy challenges posed by LLM-powered chatbots.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU

    cs.DC 2026-05 conditional novelty 7.0 of 10

    LlamaWeb is a WebGPU backend for llama.cpp that uses static memory planning, tunable kernels, and templated multi-precision support to cut memory use by 29-33% and raise decode throughput by 45-69% versus prior browse...

  2. Barriers to Evidence in AI-Related Cases and the Privatization of Proof

    cs.CY 2026-05 unverdicted novelty 5.0 of 10

    The paper identifies seven asymmetries in access to AI evidence and proposes a three-part test for courts to resolve disclosure disputes using proportionality and reasonable alternatives.

Reference graph

Works this paper leans on

94 extracted references · 61 canonical work pages · cited by 2 Pith papers

  1. [1]

    Anthropic . 2025 a . Anthropic’s Transparency Hub

  2. [2]

    Anthropic . 2025 b . Model Training Notice. Accessed: 2025-05-23

  3. [3]

    Anthropic . 2025 c . Privacy Policy. Accessed: 2025-05-23

  4. [4]

    Anthropic . 2025 d . Updates to our Privacy Policy. Accessed: 2025-05-23

  5. [5]

    Apple . 2024. Introducing Apple’s On-Device and Server Foundation Models. Machine Learning Research blog post. Updated July 29, 2024

  6. [6]

    Balebako, R.; Schaub, F.; Adjerid, I.; Acquisti, A.; and Cranor, L. 2015. The Impact of Timing on the Salience of Smartphone App Privacy Notices. In Proceedings of the 5th Annual ACM CCS Workshop on Security and Privacy in Smartphones and Mobile Devices, SPSM '15, 63–74. New York, NY, USA: Association for Computing Machinery. ISBN 9781450338196

  7. [7]

    Belanger, A. 2025. ChatGPT users shocked to learn their chats were in Google search results. Ars Technica

  8. [8]

    Bommasani, R.; Klyman, K.; Kapoor, S.; Longpre, S.; Xiong, B.; Maslej, N.; and Liang, P. 2025. The 2024 Foundation Model Transparency Index. arXiv:2407.12929

Show all 94 references
  1. [9]

    Bommasani, R.; Klyman, K.; Longpre, S.; Kapoor, S.; Maslej, N.; Xiong, B.; Zhang, D.; and Liang, P. 2023. The Foundation Model Transparency Index. arXiv:2310.12941

  2. [10]

    Bommasani, R.; Klyman, K.; Longpre, S.; Xiong, B.; Kapoor, S.; Maslej, N.; Narayanan, A.; and Liang, P. 2024. Foundation Model Transparency Reports. arXiv:2402.16268

  3. [11]

    Burgess, M. 2023. Generative AI's Biggest Security Flaw Is Not Easy to Fix. Wired

  4. [12]

    Burgess, M.; and Rogers, R. 2024. How to Stop Your Data From Being Used to Train AI. WIRED. Accessed: 2025-05-23

  5. [13]

    Cal. Bus. & Prof. Code § 22581 . 2024. Privacy Rights for California Minors in the Digital World

  6. [14]

    California Privacy Protection Agency . 2024. Title 11, Sec.7011(a)

  7. [15]

    California Privacy Protection Agency . 2025. Frequently Asked Questions (FAQs). Accessed: 2025-05-23

  8. [16]

    California State Legislature . 2023. California Civil Code §1798.130 (a)(5)(B)-(C). Accessed: 2025-05-23

  9. [17]

    Calo, M. R. 2011. The Boundaries of Privacy Harm. IND. L. J., 86(3): 1133

  10. [18]

    Chen, B. X. 2024. A.I. Devices Want More of Our Data. The New York Times

  11. [19]

    Chen, B. X. 2025. Google Is Going to Let Kids Use Its Gemini AI. The New York Times. Accessed: 2025-05-23

  12. [20]

    CIPL. 2019. Organizational Accountability in Light of FTC Consent Orders

  13. [21]

    F.; and Grimmelmann, J

    Cooper, A. F.; and Grimmelmann, J. 2025. The Files are in the Computer: Copyright, Memorization, and Generative AI. arXiv:2404.12590

  14. [22]

    Davies, P. 2024. Clearview AI fined by Dutch authorities for `illegal' facial recognition database

  15. [23]

    N.; Gerke, S.; and Kollnig, K

    Duffourc, M. N.; Gerke, S.; and Kollnig, K. 2024. Privacy of Personal Data in the Generative AI Data Lifecycle. NYU Journal of Intellectual Property & Entertainment Law, 13(2): 219--266. Accessed: 2025-05-23

  16. [24]

    A.; and Dodge, J

    Elazar, Y.; Bhagia, A.; Magnusson, I.; Ravichander, A.; Schwenk, D.; Suhr, A.; Walsh, P.; Groeneveld, D.; Soldaini, L.; Singh, S.; Hajishirzi, H.; Smith, N. A.; and Dodge, J. 2024. What's In My Big Data? arXiv:2310.20707

  17. [25]

    European Commission . 2025. What does `grounds of legitimate interest' mean?

  18. [26]

    Federal Trade Commission . 2022 a . Everalbum, Inc., In the Matter of

  19. [27]

    Federal Trade Commission . 2022 b . FTC Takes Action Against Company Formerly Known as Weight Watchers for Illegally Collecting Kids' Sensitive Health Data

  20. [28]

    Federal Trade Commission . 2023. Rite Aid Banned from Using AI Facial Recognition After FTC Says Retailer Deployed Technology without Reasonable Safeguards

  21. [29]

    A.; Rieser, V.; Iqbal, H.; Tomašev, N.; Ktena, I.; Kenton, Z.; Rodriguez, M.; El-Sayed, S.; Brown, S.; Akbulut, C.; Trask, A.; Hughes, E.; Bergman, A

    Gabriel, I.; Manzini, A.; Keeling, G.; Hendricks, L. A.; Rieser, V.; Iqbal, H.; Tomašev, N.; Ktena, I.; Kenton, Z.; Rodriguez, M.; El-Sayed, S.; Brown, S.; Akbulut, C.; Trask, A.; Hughes, E.; Bergman, A. S.; Shelby, R.; Marchal, N.; Griffin, C.; Mateos-Garcia, J.; Weidinger, L...

  22. [30]

    W.; Wallach, H.; III, H

    Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J. W.; Wallach, H.; III, H. D.; and Crawford, K. 2021. Datasheets for Datasets. arXiv:1803.09010

  23. [31]

    Goel, S.; and Webb, E. 2025. Contractors training Meta's AI say they read intimate talks with its chatbot --- and see data that identifies users Contractors training Meta's AI say they read intimate talks with its chatbot --- and see data that identifies users. Business Insider

  24. [32]

    Google . 2024. Google Privacy Policy. Accessed: 2025-05-23

  25. [33]

    Google. 2025. Gemini adds Temporary Chats and new personalization features

  26. [34]

    Google . 2025. Gemini Apps Privacy Hub. Accessed: 2025-05-23

  27. [35]

    M.; Kou, Y.; Battles, B.; Hoggatt, J.; and Toombs, A

    Gray, C. M.; Kou, Y.; Battles, B.; Hoggatt, J.; and Toombs, A. L. 2018. The Dark (Patterns) Side of UX Design. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI '18, 1–14. New York, NY, USA: Association for Computing Machinery. ISBN 9781450356206

  28. [36]

    Gumusel, E. 2025. A literature review of user privacy concerns in conversational chatbots: A social informatics approach: An Annual Review of Information Science and Technology (ARIST) paper. Journal of the Association for Information Science and Technology, 76(1): 121--154

  29. [37]

    Hern, A. 2020. This article is more than 5 years old Google says it will no longer save a complete record of every search. The Guardian

  30. [38]

    Higa, S. C. L., Haley; Bedikian. 2025. The Right to Be Forgotten Is Dead: Data Lives Forever in AI. Tech Policy Press

  31. [39]

    Hoofnagle, C. J. 2016. Federal Trade Commission Privacy Law and Policy. Cambridge University Press

  32. [40]

    Horwitz, J. 2025. Meta's AI rules have let bots hold `sensual' chats with kids, offer false medical info

  33. [41]

    of Privacy Professionals

    Int'l Assoc. of Privacy Professionals . 2024. State Comprehensive Privacy Law Comparison. Accessed: 2025-05-23

  34. [42]

    of Privacy Professionals

    Int'l Assoc. of Privacy Professionals . 2025. US State AI Governance Legislation Tracker 2025. Accessed: 2025-05-23

  35. [43]

    Javed, Y.; and Sajid, A. 2024. A Systematic Review of Privacy Policy Literature. ACM Comput. Surv., 57(2)

  36. [44]

    S.; Subramani, N.; Johnson, I.; Dupont, G.; Dodge, J.; Lo, K.; Talat, Z.; Radev, D.; Gokaslan, A.; Nikpoor, S.; Henderson, P.; Bommasani, R.; and Mitchell, M

    Jernite, Y.; Nguyen, H.; Biderman, S.; Rogers, A.; Masoud, M.; Danchev, V.; Tan, S.; Luccioni, A. S.; Subramani, N.; Johnson, I.; Dupont, G.; Dodge, J.; Lo, K.; Talat, Z.; Radev, D.; Gokaslan, A.; Nikpoor, S.; Henderson, P.; Bommasani, R.; and Mitchell, M. 2022. Data Governanc...

  37. [45]

    B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D

    Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020. Scaling Laws for Neural Language Models. arXiv:2001.08361

  38. [46]

    E.; Liang, P.; and Narayanan, A

    Kapoor, S.; Bommasani, R.; Klyman, K.; Longpre, S.; Ramaswami, A.; Cihon, P.; Hopkins, A.; Bankston, K.; Biderman, S.; Bogen, M.; Chowdhury, R.; Engler, A.; Henderson, P.; Jernite, Y.; Lazar, S.; Maffulli, S.; Nelson, A.; Pineau, J.; Skowron, A.; Song, D.; Storchan, V.; Zhang,...

  39. [47]

    G.; Cesca, L.; Bresee, J.; and Cranor, L

    Kelley, P. G.; Cesca, L.; Bresee, J.; and Cranor, L. F. 2010. Standardizing privacy notices: an online study of the nutrition label approach. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '10, 1573–1582. New York, NY, USA: Association for C...

  40. [48]

    Khan, M.; and Hanna, A. 2025. The Subjects and Stages of AI Dataset Development: A Framework for Dataset Accountability. The Ohio State Technology Law Journal. Accessed: 2025-05-23; see p. 220

  41. [49]

    King, J. 2018. Privacy, Disclosure, and Social Exchange Theory. Ph.D. thesis, University of California, Berkeley. ProQuest ID: King\_berkeley\_0028E\_17901; Merritt ID: ark:/13030/m5t77dzd

  42. [50]

    Becoming Part of Something Bigger

    King, J. 2019. "Becoming Part of Something Bigger": Direct to Consumer Genetic Testing, Privacy, and Personal Disclosure. Proc. ACM Hum.-Comput. Interact., 3(CSCW)

  43. [51]

    King, J.; and Meinhardt, C. 2024. Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World. Technical report, Stanford Institute for Human-Centered Artificial Intelligence. Accessed: 2025-05-23

  44. [52]

    Knibbs, K. 2024. Every AI Copyright Lawsuit in the US, Visualized. WIRED. Accessed: 2025-05-23

  45. [53]

    F.; and Grimmelmann, J

    Lee, K.; Cooper, A. F.; and Grimmelmann, J. 2024. Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain. arXiv:2309.08133

  46. [54]

    Leffer, L. 2023. Your Personal Information Is Probably Being Used to Train Generative AI Models. Scientific American

  47. [55]

    Longpre, S.; Mahari, R.; Lee, A.; Lund, C.; Oderinwale, H.; Brannon, W.; Saxena, N.; Obeng-Marnu, N.; South, T.; Hunter, C.; et al. 2024. Consent in crisis: The rapid decline of the ai data commons. Advances in Neural Information Processing Systems, 37: 108042--108087

  48. [56]

    M.; and Cranor, L

    McDonald, A. M.; and Cranor, L. F. 2008. The cost of reading privacy policies. Isjlp, 4: 543

  49. [57]

    Meta . 2024. Meta Privacy Policy. Accessed: 2025-05-23

  50. [58]

    Meta . 2025. How Meta uses information for generative AI models and features. Accessed: 2025-05-23

  51. [59]

    Microsoft . 2025 a . Microsoft Privacy Statement. Accessed: 2025-05-23

  52. [60]

    Microsoft . 2025 b . Privacy FAQ for Microsoft Copilot

  53. [61]

    Mireshghallah, N.; Antoniak, M.; More, Y.; Choi, Y.; and Farnadi, G. 2024. Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild. arXiv:2407.11438

  54. [62]

    D.; and Gebru, T

    Mitchell, M.; Wu, S.; Zaldivar, A.; Barnes, P.; Vasserman, L.; Hutchinson, B.; Spitzer, E.; Raji, I. D.; and Gebru, T. 2019. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, 220–229. ACM

  55. [63]

    F.; Ippolito, D.; Choquette-Choo, C

    Nasr, M.; Carlini, N.; Hayase, J.; Jagielski, M.; Cooper, A. F.; Ippolito, D.; Choquette-Choo, C. A.; Wallace, E.; Tramèr, F.; and Lee, K. 2023. Scalable Extraction of Training Data from (Production) Language Models. arXiv:2311.17035

  56. [64]

    Nolte, H.; Finck, M.; and Meding, K. 2025. Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data. arXiv:2503.01630

  57. [65]

    OpenAI . 2024. Privacy Policy. Accessed: 2025-05-23

  58. [66]

    OpenAI . 2025. Data Controls FAQ. Accessed: 2025-05-23

  59. [67]

    OpenAI. 2025. GPT-5 and the new era of work

  60. [68]

    OpenAI . 2025 a . How ChatGPT and our foundation models are developed. Accessed: 2025-05-23

  61. [69]

    OpenAI . 2025 b . What is memory. Accessed: 2025-05-23

  62. [70]

    D.; Bender, E

    Paullada, A.; Raji, I. D.; Bender, E. M.; Denton, E.; and Hanna, A. 2021. Data and its (dis)contents: A survey of dataset development and use in machine learning research. Patterns, 2(11): 100336

  63. [71]

    Proton . 2025. Lumo Privacy. Support page on Proton website

  64. [72]

    Reid, E. 2025. AI in Search: Going Beyond Information to Intelligence. Accessed: 2025-05-23

  65. [73]

    Sampson, P.; and Bogen, M. 2025. It's (Getting) Personal: How Advanced AI Systems Are Personalized. Technical report, Center for Democracy & Technology. Accessed: 2025-05-23

  66. [74]

    L.; and Cranor, L

    Schaub, F.; Balebako, R.; Durity, A. L.; and Cranor, L. F. 2015. A design space for effective privacy notices. In Proceedings of the Eleventh USENIX Conference on Usable Privacy and Security, SOUPS '15, 1--17. USA: USENIX Association. ISBN 9781931971249

  67. [75]

    Schomer, A. 2024. AI Content Licensing Deals With Publishers: Complete Updated Index

  68. [76]

    H.; Liu, K.; Wang, A.; Cen, S

    Shen, J. H.; Liu, K.; Wang, A.; Cen, S. H.; Zhang, A. K.; Meinhardt, C.; Zhang, D.; Klyman, K.; Bommasani, R.; and Ho, D. E. 2025. The Disclosure Delusion: Systemic Challenges in AI Data Transparency Policy. Copyright 2025 by the authors

  69. [77]

    A.; Zettlemoyer, L.; Koh, P

    Shi, W.; Bhagia, A.; Farhat, K.; Muennighoff, N.; Walsh, P.; Morrison, J.; Schwenk, D.; Longpre, S.; Poznanski, J.; Ettinger, A.; Liu, D.; Li, M.; Groeneveld, D.; Lewis, M.; tau Yih, W.; Soldaini, L.; Lo, K.; Smith, N. A.; Zettlemoyer, L.; Koh, P. W.; Hajishirzi, H.; Farhadi, ...

  70. [78]

    H.; Kumar, S.; Lucy, L.; Lyu, X.; Lambert, N.; Magnusson, I.; Morrison, J.; Muennighoff, N.; Naik, A.; Nam, C.; Peters, M

    Soldaini, L.; Kinney, R.; Bhagia, A.; Schwenk, D.; Atkinson, D.; Authur, R.; Bogin, B.; Chandu, K.; Dumas, J.; Elazar, Y.; Hofmann, V.; Jha, A. H.; Kumar, S.; Lucy, L.; Lyu, X.; Lambert, N.; Magnusson, I.; Morrison, J.; Muennighoff, N.; Naik, A.; Nam, C.; Peters, M. E.; Ravich...

  71. [79]

    Solove, D. J. 2025. Artificial Intelligence and Privacy. Florida Law Review, 77(1)

  72. [80]

    J.; and Hartzog, W

    Solove, D. J.; and Hartzog, W. 2025. The Great Scrape: The Clash Between Scraping and Privacy. California Law Review. Forthcoming

  73. [81]

    Staab, M. B. M. V. M., Robin; Vero. 2024. Beyond Memorization: Violating Privacy Via Inference with Large Language Models. arXiv:2310.07298

  74. [82]

    StatCounter. 2025. AI Chatbot Market Share Worldwide

  75. [83]

    Surfshark . 2025. AI Chatbots Ranked by Data They Collect. Accessed: 2025-05-23

  76. [84]

    H.; and Sunstein, C

    Thaler, R. H.; and Sunstein, C. R. 2008. Nudge: Improving Decisions About Health, Wealth, and Happiness. Yale University Press

  77. [85]

    Tinfoil . 2025. Tinfoil Enclaves: A Technical Overview. Blog post on Tinfoil website. Updated Sep. 3, 2025

  78. [86]

    Turow, J.; Hennessy, M.; and and, N. D. 2018. Persistent Misperceptions: Americans' Misplaced Confidence in Privacy Policies, 2003--2015. Journal of Broadcasting & Electronic Media, 62(3): 461--478

  79. [87]

    Tömekçe, B.; Vero, M.; Staab, R.; and Vechev, M. 2024. Private Attribute Inference from Images with Vision-Language Models. arXiv:2404.10618

  80. [88]

    UCLA ITLP, N. T. 2024. Letter to the CPPA Board

  81. [89]

    Villalobos, P.; Sevilla, J.; Heim, L.; Besiroglu, T.; Hobbhahn, M.; and Ho, A. 2024. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325, 1

  82. [90]

    Y.; and Ku, C

    Yang, J.; Chen, Y.-L.; Por, L. Y.; and Ku, C. S. 2023. A Systematic Literature Review of Information Security in Chatbots. Applied Sciences, 13(11)

  83. [91]

    Zanfir-Fortuna, G. 2023. How Data Protection Authorities are De Facto Regulating Generative AI. Future of Privacy Forum blog. Accessed June 13, 2025

  84. [92]

    Zhang, D.; Finckenberg-Broman, P.; Hoang, T.; Pan, S.; Xing, Z.; Staples, M.; and Xu, X. 2024. Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions. arXiv:2307.03941

  85. [93]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  86. [94]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.