{"id":"60e64355-1050-43e5-98bc-8573344959cc","arxiv_id":"2412.20231","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Training shared AI models on end-to-end encrypted messages is incompatible with E2EE confidentiality; AI inference on encrypted content is compatible only with endpoint-local processing or strict per-user, no-third-party conditions.","lead":"This paper argues that training shared AI models on end-to-end encrypted message content breaks the core confidentiality promise of encryption, and that cloud-based AI inference on such content only preserves that promise under narrow conditions. It offers a technical and legal framework that could shape how messaging apps design, label, and seek consent for AI features.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The categorical incompatibility claim depends on a derivative definition that the paper's own caveat-feature category already relaxes; without a stated threshold, Recommendation 1 is not entailed by the paper's framework.","rationale":"The reader's weakest assumption identifies the derivative definition in §3.1 as load-bearing; I agree, and push further: the paper's own Category 2 caveat-feature exception already permits third-party access to plaintext-dependent content in strictly limited cases. That means the derivative definition is not applied categorically, so the binary claim in Recommendation 1 cannot be derived solely from 'a model is a derivative'. One needs an account of why training-induced derivatives are qualitatively different from caveat-feature derivatives. The paper gestures at volume (§4.1.2, 'significant volumes') but does not define 'strictly limited' in a way that excludes a shared model trained on a tiny corpus while including link previews. This matters because Recommendation 1 anchors Recommendations 2, 3, and 4; if the incompatibility is actually a matter of degree or policy, the recommendations should be reframed. The paper's legal analysis, case studies, and disclosure/consent recommendations remain valuable and independently supported; the concern is specifically about the security claim's categorical framing. The requested threshold and formal predicate would settle whether the central claim is a theorem or a normative choice.","tokens_in":43962,"tokens_out":11176,"duration_ms":126425,"concrete_test":"Use the paper's own taxonomy to define a minimal shared model trained on exactly the same quantity of E2EE content that a link-preview feature exposes (e.g., one URL-length string from one message). Check whether the model is queryable by a third party and whether the resulting output is a derivative under §3.1. If the link-preview feature is permitted as a caveat while the identically-sized training set is forbidden, ask the authors to state the threshold that distinguishes them. A formal version: formalize the caveat-feature exception in §3.2 as a predicate C(D) on disclosed derivatives D; show that Recommendation 1 is equivalent to asserting C(D)=false for all training-induced models while C(url-preview)=true. If no such predicate is provided, the categorical claim is underdetermined.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Recommendation 1 asserts that training shared models on E2EE content is definitionally incompatible with E2EE because a model is a derivative of its training data (§7.1, §3.1). This conclusion rests on the §3.1 definition of E2EE content as including any derivative of protected data. However, §3.2 already permits Category 2 applications to expose strictly limited types or quantities of plaintext-dependent content to third parties through caveat features such as link previews, and still calls them E2EE in current practice. A URL sent to a link-preview service is part of the message content, and therefore at least as much a derivative as many model outputs. The paper gives no principled boundary that would make a shared model trained on a small amount of E2EE content categorically forbidden while a link preview is merely a caveat. If the derivative definition is taken literally, link previews should also be incompatible; if it is not taken literally, the incompatibility of shared-model training is not definitional. The central claim is thus not a theorem from standard E2EE confidentiality alone; it is a policy line requiring a quantitative or qualitative threshold that the paper does not supply.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether integrating AI assistants into end-to-end encrypted messaging systems, and training models on E2EE content, is compatible with the security guarantees of E2EE. It contributes a nine-term vocabulary and a five-category taxonomy (strict E2EE to no E2EE), a four-consideration technical evaluation framework (model location, privacy enhancement, guarantee type, shared vs. per-user), a survey of relevant US and EU legal frameworks and consent theory, and four numbered recommendations. Its headline normative claim is Recommendation 1: using E2EE content to train shared AI models is not compatible with E2EE. The paper applies its framework to Apple Intelligence, Samsung Galaxy AI, and Meta AI in WhatsApp, and is notably transparent about what vendor documentation does not disclose.","tokens_in":44159,"tokens_out":11451,"duration_ms":117977,"significance":"If the central claim is accepted, the paper provides a useful shared vocabulary and evaluation template for an urgent policy debate, and its legal synthesis is a valuable starting point for regulators. The paper's strengths include the explicit taxonomy, the structured evaluation framework with per-configuration verdicts, the candid \"What We Still Don't Know\" lists, and the willingness to identify open legal questions rather than overclaim certainty. It is a position/analysis paper rather than a formal security proof, and its central categorical recommendation needs refinement before the paper's claims can be fully endorsed; with that refinement, the paper would be a significant contribution to the cs.CR and security-policy literature.","major_comments":[{"comment":"Recommendation 1 is not entailed by the paper's own definitions. Section 3.1 defines E2EE content to include \"any derivatives\" of protected data (footnote 32), and Section 7.1 concludes that training a shared model on E2EE content \"definitionally undermines\" E2EE because the model is a derivative of its training data. However, Section 3.2 already grants Category 2 (\"E2EE in current practice\") to applications that expose plaintext-dependent content to third parties through caveat features such as link previews, and Section 2.1.5 concedes that link previewing \"technically violates E2EE confidentiality.\" A URL in a message is primary content, and the preview service's output is a derivative of that content; under the paper's own definition, a link preview is exactly the kind of disclosure that Recommendation 1 declares categorically incompatible. The only stated distinction is that caveat features are \"strictly limited\" while AI processing involves \"significant volumes\" (§4.1.2), and footnote 34 explicitly declines to define a threshold. Recommendation 1 therefore requires either a principled threshold separating caveat features from model training or an explicit restriction to \"strict E2EE\" (Category 1); the paper should supply one of these before the central claim can be evaluated.","section":"§3.1, §3.2, §7.1"},{"comment":"The opening of Section 4.2 states that \"existing privacy-enhancing techniques cannot address the security concerns we raise around integrating AI assistants with E2EE applications,\" but the subsequent analysis concludes that endpoint-local models are \"fully compatible with E2EE\" (Consideration #1) and that per-user models \"can be compatible with E2EE\" (Consideration #4), and Table 1 marks the on-device configuration with a green tag. The summary sentence is therefore overbroad and contradicts the paper's own framework. It should be qualified to refer to configurations that involve third-party or shared processing, or the compatible configurations should be explicitly acknowledged as exceptions. As written, the section's framing obscures the paper's main message and reduces the reader's ability to rely on the framework summary.","section":"§4.2, Consideration #1, Consideration #4, Table 1"},{"comment":"Recommendation 4 is in tension with Recommendation 1. Recommendation 1 states that training shared models on E2EE content is categorically incompatible with E2EE, but Recommendation 4 recommends opt-in consent as the \"standard mechanism for allowing messaging services to train AI on user data\" and explains how to design consent for such training. If the incompatibility is definitional, consent cannot change the security guarantee; if consent can make the processing acceptable in some circumstances, then the incompatibility in Recommendation 1 is not about E2EE confidentiality but about user agreement, and the paper should say so explicitly. The paper should clarify whether Recommendation 4 is a harm-reduction fallback for deployments that remain incompatible (which should be labeled as such) or a way to make training E2EE-compatible (which contradicts Recommendation 1).","section":"§7.2, Recommendation 4"}],"minor_comments":[{"comment":"The informal gloss of semantic security as saying that \"no computational adversary can guess any function of the message content\" should be stated more precisely, e.g., \"efficiently computable functions\" and \"with more than negligible advantage,\" to avoid overclaiming.","section":"§2.1.1, footnote 6"},{"comment":"The Zama estimate of roughly $5,000 per word is a vendor blog figure; it would be safer to cite peer-reviewed FHE performance measurements or to label the figure explicitly as a vendor-provided estimate, especially since the paper's argument does not depend on this particular number.","section":"§4.2, FHE paragraph"},{"comment":"The table row marks the on-device configuration as fully compatible, but it omits the qualification stated in §4.2 Consideration #1 that endpoint-local training is compatible only if E2EE messages are used solely for the sender's and recipients' endpoint-local models and not for fine-tuning models on other devices.","section":"Table 1, row \"On-device, fine-tuned model\""},{"comment":"The discussion of FTC \"precedent\" should be framed as agency enforcement actions and consent decrees rather than judicial precedent, since the paper is describing FTC practice and investigation patterns rather than binding court decisions.","section":"§6.2.1"},{"comment":"The Apple Intelligence analysis depends on version-specific facts (e.g., iOS 18.1 vs. iOS 18.3 default settings); the paper should include a clear \"status as of\" date or a general reminder to re-check these fast-moving product details, since the cited states will age quickly.","section":"§7.3"},{"comment":"There are several small presentation errors, including the duplicated \"aligned aligned\" in §8.2 and the stray spacing in the heading \"F ederated learning\" in §4.2.1; these should be corrected in a final pass.","section":"§8.2, §4.2.1"}],"recommendation":"major_revision","confidential_remarks":"This is a timely, well-sourced position paper whose main contribution is a vocabulary and evaluation framework for an important policy question. The central issue is that Recommendation 1 is stated categorically while the paper's own taxonomy admits caveat features, so the incompatibility claim needs reframing rather than rejection. I see no citation-pattern concerns, and the paper's transparency about unknowns is a genuine strength. The paper fits the journal's scope and should be revisable within a normal revision cycle."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth a look. It is the first systematic attempt I've seen to lay out what happens to E2EE guarantees when AI assistants get integrated into encrypted messaging, and it actually delivers a usable framework: the five-category taxonomy, the four-consideration evaluation flow, and the technical-legal pairing. The case studies on Apple, Samsung, and Meta AI are careful and flag what the vendors don't disclose. The legal reading is appropriately tentative, and the paper is honest about unknowns throughout.\n\nThe soft spot is exactly the one flagged in the stress-test note, and I think the note is right. Recommendation 1 asserts that training shared models on E2EE content is definitionally incompatible with E2EE. That conclusion rests on Section 3.1's definition of E2EE content as including any derivative of protected data. But Section 3.2 already permits Category 2 applications to expose strictly limited types or quantities of plaintext-dependent content to third parties through caveat features like link previews, and still calls them E2EE in current practice. A URL is part of the message content and is at least as much a derivative as many model outputs. If the derivative definition is literal, link previews are also incompatible; if it is not literal, the incompatibility of shared-model training is a policy line, not a theorem. The paper even admits the threshold problem in footnote 34, then punts. That's a real gap in the central argument.\n\nThat said, the gap is not fatal. The practical conclusions survive if you drop the 'definitional' language and instead ground them in the memorization and extraction literature, which the paper does cite. A model trained on other people's private messages is a channel by which those messages can leak to third parties, and that undermines the confidentiality promise in any meaningful sense. The paper would be stronger if it presented Recommendation 1 as a normative/definitional choice, with the boundary between caveat features and model training made explicit.\n\nMinor quibble: the FHE impracticality claim leans on a single vendor estimate ($5,000/word from Zama), which deserves a sanity check or a second source.\n\nWho this is for: regulators, engineers, and policy people thinking about AI features in encrypted apps. Deserves a serious referee; it will need revisions, but the core framework is worth publishing. I'd take it.","headline":"A genuinely useful framework for the E2EE/AI question, but Recommendation 1 overclaims by calling the incompatibility definitional when the paper's own caveat-feature category already relaxes the derivative definition.","tokens_in":44691,"tokens_out":2414,"would_cite":true,"duration_ms":24549,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that training a shared AI model on end-to-end encrypted content is definitionally incompatible with E2EE, and only local or strictly per-user inference can preserve the guarantee.","keywords":["end-to-end encryption","AI assistants","confidentiality","training data leakage","large language models","privacy-preserving machine learning","user consent","consumer protection"],"falsifier":"The claim would be refuted by a shared model trained on E2EE content with a formal, information-theoretic guarantee that no output and no parameter depends on any user's message content; it would also be dissolved if a court or standards body explicitly defined E2EE confidentiality to cover only direct access to plaintext rather than functions of it.","tokens_in":43775,"feed_emoji":"🔐","tokens_out":6681,"duration_ms":61742,"temperature":0.7,"pith_summary":"End-to-end encryption promises more than hiding message bytes: it promises that no one outside the sender and recipient can learn anything about a message's content, including any function of it. The paper argues that a shared AI model trained on end-to-end encrypted content is such a function, so letting other users query the model violates E2EE confidentiality by design. It distinguishes this from inference: processing encrypted content for AI is compatible with E2EE only when it happens on the user's own device, or when no third party can see or use the content and each user's content is used only for that user's request. The paper then argues that providers that ship AI features which use encrypted content must not market themselves as E2EE without qualification, and that AI features should be off by default and opt-in. The stakes are practical: Apple Intelligence, Samsung Galaxy AI, and Meta AI in WhatsApp are already deploying such integrations.","feed_headline":"Training AI on private messages breaks E2EE, paper argues","feed_subtitle":"A new analysis says shared models trained on encrypted chats violate the core promise of end-to-end encryption.","key_machinery":"The load-bearing object is the paper's definition of E2EE content, which includes any derivative of data the provider holds out as end-to-end encrypted; this definition is what makes the trained model itself E2EE content. The analytic device is a four-consideration framework: where the model runs, whether non-endpoint-local processing is privacy-enhanced, what type of confidentiality the enhancement provides, and whether the model is shared or per-user. A shared model fails the fourth consideration no matter what, an endpoint-local model passes all considerations, TEE-based processing supplies a different kind of security than E2EE, and per-user models can be compatible if the user's data is strictly isolated and statelessly processed.","core_discovery":"The paper's central claim is that using end-to-end encrypted content to train a shared AI model is definitionally incompatible with E2EE. E2EE confidentiality is defined to cover any derivative of message content, and a model trained on messages is a derivative of those messages; any user who can query the shared model can receive outputs that depend on other users' private messages. The paper argues that privacy-preserving training techniques, including differential privacy, data sanitization, federated learning, and multi-party computation, offer a spectrum of privacy rather than the binary guarantee of E2EE, so none of them can make shared-model training compatible. The same reasoning yields a four-part evaluation framework for AI assistants and a five-category taxonomy of applications ranging from strict E2EE to no E2EE.","pith_inferences":["The derivative argument extends beyond training: any server-side computation on E2EE plaintext, including client-side scanning or lawful-access designs, produces outputs that are derivatives of E2EE content; the paper gestures at this in its Crypto Wars discussion but does not fully develop it.","A testable extension would be to attempt to build a shared model whose outputs are information-theoretically independent of the E2EE content it was trained on; the paper's definitional claim predicts this is impossible, so a construction would force a redraw of the derivative definition.","The taxonomy could be used as a pre-launch checklist: classify any proposed AI feature against the five categories, and let the category determine what defaults and marketing language are permissible."],"forward_implications":["If the central claim is accepted, no shared model trained on E2EE content, no matter what privacy technique is used, can coexist with an honest claim that the service still provides E2EE.","Inference on E2EE content is compatible with E2EE only when it happens entirely on the user's device, or when the content is hidden from every third party and used exclusively to answer that user's request.","A messaging app that adds a cloud AI feature that is on by default and cannot be turned off is demoted, under the paper's taxonomy, to the no-E2EE category, even if its core messaging remains encrypted.","Providers that route E2EE content through AI features must qualify their E2EE marketing, or risk being deceptive under U.S. consumer-protection precedent.","AI features in E2EE systems should be off by default and activated only through explicit opt-in consent, with opt-out as easy as opt-in."],"supporting_citations":[{"why":"Supply the formal definitions of E2EE confidentiality that include derivatives of message content.","marker":"[82, 84, 100, 153]"},{"why":"Document that models reproduce training data, sometimes verbatim or nearly verbatim.","marker":"[58, 140, 114]"},{"why":"Shows adversarial queries can expose information in the training data.","marker":"[122]"},{"why":"Establishes training data extraction attacks against large language models.","marker":"[42]"},{"why":"Define membership inference attacks used to verify whether a data point is in the training set.","marker":"[156, 91]"},{"why":"Introduces differential privacy's epsilon-parameterized guarantee, which the paper contrasts with E2EE's binary guarantee.","marker":"[68]"},{"why":"Describes retrieval-augmented generation, the basis for per-user models that can maintain a base model independent of E2EE data.","marker":"[112]"},{"why":"Documents the current cost of fully homomorphic inference for LLMs, used to show that FHE is not yet practical.","marker":"[88]"}],"fun_headline_variants":["AI training on E2EE chats breaks encryption's core promise","Why E2EE and AI training are fundamentally incompatible","E2EE's confidentiality rules out AI training on user messages","Shared AI models on E2EE data: a definitional conflict","Paper: No privacy tool can reconcile AI training with E2EE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument rests on the definitional choice that E2EE confidentiality forbids any third party from learning any function of message content, and that a trained model is such a function; if a narrower definition of confidentiality is adopted, the claim that shared-model training is definitionally incompatible with E2EE does not follow.","fun_headline_variants_meta":{"raw":{"variants":["AI training on E2EE chats breaks encryption's core promise","Why E2EE and AI training are fundamentally incompatible","E2EE's confidentiality rules out AI training on user messages","Shared AI models on E2EE data: a definitional conflict","Paper: No privacy tool can reconcile AI training with E2EE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1653,"prompt_tokens":947,"completion_tokens":706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":620}},"tokens_in":563,"tokens_out":706,"duration_ms":7085,"temperature":1.0,"reasoning_tokens":620,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:25:09.608459+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The claim would be refuted by a shared model trained on E2EE content with a formal, information-theoretic guarantee that no output and no parameter depends on any user's message content; it would also be dissolved if a court or standards body explicitly defined E2EE confidentiality to cover only direct access to plaintext rather than functions of it.","supporting_citations":[{"cited_title":"What is retrieval-augmented generation? https : / / research","cited_arxiv_id":null,"evidence_quote":"Describes retrieval-augmented generation, the basis for per-user models that can maintain a base model independent of E2EE data."}],"review_version":1}