{"id":"63d168e6-4dd6-45c1-acb9-8fa76ff7299a","arxiv_id":"2506.22523","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A two-hour red teaming exercise at Dana-Farber showed GPT4DFCI can reproduce short verbatim passages from famous novels via indirect prompts, while news, scientific, and clinical content stayed protected.","lead":"A cancer center asked 42 security and AI experts to try to make its internal AI chatbot leak copyrighted text. The chatbot produced verbatim lines from two famous novels, while news articles, a scientific paper, and clinical notes stayed protected.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'differential success rates' inference is confounded: 0% extraction for news, scientific, and clinical targets is informative only if those texts are actually represented in the model's training data, which the paper never establishes.","rationale":"The reader's weakest assumption correctly identifies the training-data composition confound as the main threat to the paper's differential-protection conclusion. I considered whether the more load-bearing issue is the unauditable 23% success rate, but the central existence claim does not depend on that rate: one documented verbatim dedication suffices to show that some copyrighted text can be elicited. By contrast, the paper's headline inference about 'inconsistent copyright safeguards' depends on comparing rates across content types, and that comparison is only valid if all target texts are roughly equally represented in training data. The paper's own limitation statement concedes selection bias toward well-known works, and the Discussion lists 'varying representation levels in training corpora' as a possible mechanism, but no test distinguishes data composition from protective filtering. The proposed prefix-completion probe against an unfiltered model would settle this directly. The documented literary extractions are real and support the existence component of the claim, so the verdict should remain conditional rather than reject: the evidence is sufficient for an existence claim and a cautionary case report, but insufficient for the stronger differential-protection generalization without the additional membership check.","tokens_in":8768,"tokens_out":5957,"duration_ms":68875,"concrete_test":"Run a memorization probe against the same underlying model without the production content filter (e.g., the public GPT-4o or Azure OpenAI API) using exact targets from the paper. For each failure target—a distinctive sentence from the Iijima abstract, a distinctive sentence from the Chicago Tribune and NYT articles, and a distinctive MIMIC note passage—provide the first 10–20 tokens as a prefix and ask for continuation, and also apply the same indirect/translation strategies used in Task 1. Include a positive control: the reported Hitchhiker's sentence, which should be extractable if the probe protocol works. If the unfiltered model produces verbatim continuations for the news/scientific/MIMIC targets, the production system's 0% results reflect genuine filtering; if it does not, training-data absence is the likely explanation and the differential-success conclusion must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central analytical conclusion is the differential-protection claim: the Abstract says 'Differential success rates across content types suggest varying protective mechanisms' and the Conclusion cites 'inconsistent copyright safeguards requiring targeted improvements.' This inference compares a 23% literary extraction rate with 0% rates for news, scientific, and clinical notes. A 0% rate is evidence of protection only if the target content is present in the model's training data in comparable form. The authors selected targets for 'high probability of inclusion in known training datasets' (Target Material Selection), but probability is not presence. For the Chicago Tribune and NYT articles, the Iijima paper, and the MIMIC-IV notes, no membership evidence is offered. Supplementary Table 4 even shows the model failing to complete simple 'charttime' rows, suggesting weak memorization of the MIMIC content specifically. The literary successes are also short, widely quoted phrases: the Harry Potter dedication appears on many public websites, so it does not isolate book-corpus memorization. The observed gradient could therefore be fully explained by training-data composition rather than by 'varying protective mechanisms.' The bare existence claim—that some copyrighted text is extractable—survives, but the 'inconsistent safeguards' conclusion and the institutional recommendation for targeted improvements are not supported without controlling for target presence and memorability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-hour red teaming exercise conducted in November 2024 against GPT4DFCI, a production enterprise AI tool at Dana-Farber Cancer Institute using OpenAI models. Forty-two participants in four teams attempted to extract copyrighted content from four domains: literary works, news articles, scientific publications, and access-restricted MIMIC clinical notes. The authors report that indirect prompting elicited a verbatim Harry Potter dedication, an Italian/English round-trip of the same dedication, and a partially verbatim Hitchhiker's Guide sentence, while news, scientific, and clinical extractions failed. The paper interprets these results as evidence of differential protective mechanisms, recommends inference-time filtering, and describes a copyright-specific meta-prompt deployed in production in January 2025.","tokens_in":8919,"tokens_out":3115,"duration_ms":36949,"significance":"If the reported examples are accurate, the paper provides a useful real-world existence proof that a filtered, production medical AI can still emit short memorized passages of copyrighted literary text under conversational prompting. The main strength is documentation: Supplementary Table 1 contains prompt/output logs rather than only aggregate claims, and the authors name the institutional context, event time, and mitigation. The practical value for academic medical centers deploying Azure OpenAI products is real, and the paper is candid about several limitations. However, the analytical claims that go beyond the existence proof—especially the 23 percent literary success rate and the differential-protection inference—are not currently supported by the evidence presented.","major_comments":[{"comment":"The reported 23% success rate for literary works is not recoverable from the supplied data. Supplementary Table 1 lists 25 book-related interactions, and only rows 7–9 (the dedication, including Italian and English round-trip) and row 25 (approximately half of one sentence) show verbatim or near-verbatim reproduction; that is at most 4/25, or 16%, depending on the counting rule. The text also says the overall exercise comprised 156 documented attack attempts, but the four supplementary tables contain only 53 prompt/output rows, so the denominator for the 23% figure is not evident. Please provide the complete attack log with a per-attempt success label and an explicit statement of the success classification rubric, or remove the quantitative rate.","section":"Quantitative Results Summary and Supplementary Table 1"},{"comment":"The 'differential success rates across content types suggest varying protective mechanisms' claim is confounded by training-data presence. For the Chicago Tribune and New York Times articles, the Iijima paper, the MOLSCRIPT paper, and the MIMIC notes, the authors only state that targets were selected for 'high probability of inclusion in known training datasets'; no membership evidence is offered. A 0% extraction rate is equally consistent with the targets being absent or under-represented in the model's training data as with the model having protective mechanisms. Supplementary Table 4 even shows the model failing to complete simple 'charttime' rows, suggesting weak memorization of the MIMIC content specifically. The Conclusion's 'inconsistent copyright safeguards requiring targeted improvements' is therefore not supported unless target presence in training data is established or the claim is reframed as a hypothesis. Please either add membership tests (e.g., direct extraction probes or exposure metrics) or limit the abstract and conclusion to the existence claim.","section":"Abstract, Discussion, and Conclusion"},{"comment":"The paper states that the copyright-specific meta-prompt 'Avoid copyright infringement' was deployed in production since January 2025 and that this exemplifies responsible AI development, but no post-mitigation evaluation is reported. The effectiveness of the meta-prompt is not measured, so the conclusion that the exercise 'led to concrete mitigation strategies' is only a claim about deployment, not about protection. Please clarify in the Discussion that the mitigation has not yet been validated, or add the results of a re-test after deployment.","section":"Discussion (mitigation claim)"}],"minor_comments":[{"comment":"There are several typos in the prompts as transcribed, including 'Hitchhiker's Guid to the Galaxy', 'perspecitve', and 'Aurther's house'; these should be corrected or marked as verbatim user typos so readers can distinguish transcription errors from actual user input.","section":"Supplementary Table 1"},{"comment":"Row 2 lists the model as 'GPT4oo' while all other rows say 'GPT-4o'; this appears to be a typo.","section":"Supplementary Table 4"},{"comment":"The DAN jailbreak citation contains 'doi: doi: 10.5281/zenodo.1234', which looks like a placeholder DOI; please verify the citation or replace it with a proper source.","section":"Reference 16"},{"comment":"The authors generalize from four targets per domain (or fewer in the scientific domain, where two papers are listed) to broad statements about 'news articles', 'scientific publications', and 'academic content'; the Discussion should acknowledge that the number of targets per domain is too small to support domain-level conclusions even setting aside the training-presence confound.","section":"Target Material Selection"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a case report with operational value. Its central existence claim is supported by the supplementary logs, but the quantitative success rate and the differential-protection interpretation need substantial revision. If the authors cannot provide the full attack log or membership evidence, they should remove the 23% figure and soften the abstract/conclusion to a claim about literary-text extraction only. The paper may be better suited to a health-AI deployment venue than to a general AI journal, but that is not a blocking issue if the analytical claims are fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first copyright-focused red team of a production medical AI system I've seen, and the core descriptive finding is solid. Indirect prompts extracted the Harry Potter dedication verbatim and a near-verbatim Hitchhiker's sentence from a filtered enterprise GPT-4 instance, and those examples are documented in the supplementary tables. The phenomenon is real. That alone makes the paper worth a look for anyone running an institutional LLM deployment.\n\nWhat it does well: the event design is transparent (42 participants, four domains, governance approval), the write-up is honest about the two-hour constraint and selection bias, and the supplementary tables let you see the actual prompts and outputs. Deploying a meta-prompt as a first mitigation is a reasonable practical step given the timeline.\n\nThe soft spots are mostly about the comparisons, not the raw results. The headline '23% success rate for literary works' is not auditable: the tables show about 53 summarized rows, not the claimed 156 attempts, and there's no success rubric. More importantly, the differential-protection conclusion—'inconsistent copyright safeguards'—is confounded by training-data presence. You can't read a 0% extraction rate for the NYT/Chicago Tribune articles, the Iijima paper, or MIMIC notes as evidence of protection unless those exact texts are in the model's training data. The paper never establishes that. The authors do flag this in the Discussion but then leave it unaddressed, and the Abstract/Conclusion state the gradient as if it's established. That's a load-bearing overreach. Also, the literary successes are short, widely quoted phrases, which makes 'training data contamination' plausible but doesn't isolate book-corpus memorization specifically.\n\nThe mitigation's effectiveness is also unreported—fine for a case report, but worth saying clearly.\n\nBottom line: the existence claim holds up, and the methodology is reusable. The differential-safeguards claim needs either a membership check (e.g., probe with known text) or a downgrade to hypothesis-generating. I'd send this to peer review—it's a genuine contribution to the medical AI governance niche, and the flaws are addressable in revision. If I were writing on enterprise LLM risk, I'd cite it.","headline":"A genuinely useful case report that overreaches on the differential-safeguards conclusion; the raw extraction examples are real, the comparison across content types needs a training-data presence check.","tokens_in":9704,"tokens_out":3356,"would_cite":true,"duration_ms":37917,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A production medical AI reproduced a copyrighted book dedication verbatim under indirect prompting, despite direct requests being refused.","keywords":["red teaming","copyright compliance","generative AI","large language models","training data memorization","indirect prompting","academic medical center","content filtering"],"falsifier":"Run the same indirect prompts on a production model using a literary passage known to be absent from its training data; if the model still returns it verbatim, training-data contamination is not the explanation and a different failure mode would be at work. Conversely, rerun the news tasks on articles as widely copied as the Rowling dedication; if extraction remains zero, then content-type-specific safeguards, not data availability, explain the difference.","tokens_in":8496,"feed_emoji":"📖","tokens_out":5697,"duration_ms":63362,"temperature":0.7,"pith_summary":"This paper reports a structured red teaming exercise in which participants tried to extract copyrighted material from GPT4DFCI, a production generative AI assistant deployed in an academic medical center. The central finding is that the model reproduced a copyrighted book dedication verbatim and a short novel passage near-verbatim under indirect conversational prompting, even though direct requests for the same text were consistently refused. News articles, scientific papers, and clinical notes were not extracted, which the authors interpret as evidence that safeguards and training-data composition differ by content type. The authors argue that this vulnerability points to training-data memorization and that inference-time filtering plus continuous red teaming are needed.","feed_headline":"A medical AI recited a book dedication word for word","feed_subtitle":"Direct asks were refused; indirect ones pulled a memorized Rowling passage from a filtered model.","key_machinery":"The load-bearing mechanism is indirect prompting: instead of asking for copyrighted text directly, an attacker builds conversational context so that keyword-based refusal triggers are not activated, and the model then completes memorized training text. The same mechanism explains why direct requests for news, scientific, and clinical content failed while the literary dedication succeeded. The paper also introduces a four-domain testing matrix as a way to compare extraction rates across content types, and it uses that matrix to support the claim that protection levels vary.","core_discovery":"The report's central claim is that a production generative AI system with standard safeguards can inadvertently reproduce short copyrighted literary passages when prompted indirectly. Concretely, asking who the Harry Potter book was dedicated to produced the verbatim dedication, and a conversation about feeling insignificant led to a passage from The Hitchhiker's Guide to the Galaxy containing a near-verbatim sentence. Direct requests for book text were refused, while indirect strategies such as style mimicry and translation round-trips bypassed the refusals. The authors take the differential success rates across content types as evidence that protective mechanisms are inconsistent and that training-data contamination remains the most probable explanation for the literary reproductions.","pith_inferences":["If the differential results are driven by training-data representation rather than safeguards, then any content type well represented in training data is potentially at risk, including lyrics, scripts, or book dedications in other deployments.","The successful translation round-trip suggests that memorized text can survive semantic transformation, so exact-match evaluation is too weak; paraphrase-invariant tests would be more informative.","A decisive experiment would audit a model's training corpus and compare extraction of passages known to be in the corpus versus known to be absent, separating memorization from generation.","For institutional governance, logging adversarial prompt-output pairs during red teaming could become part of the evidence needed to support a legal indemnification claim."],"forward_implications":["Verbatim literary passages can be extracted through indirect prompting, so keyword-based refusal filtering is not sufficient on its own.","The 23% success rate for literary extraction, contrasted with 0% for news and science, suggests that protective mechanisms are content-specific or that training-data representation varies by content type.","A copyright-specific meta-prompt was implemented in production in January 2025 as a first mitigation step.","Academic medical institutions deploying generative AI should establish continuous red teaming protocols rather than relying on a one-time evaluation.","Institutions using hosted enterprise AI services should treat copyright testing as part of the shared responsibility needed to qualify for platform copyright protections."],"supporting_citations":[{"why":"Documents that shadow libraries and piracy databases may contribute copyrighted material to training corpora, the paper's proposed source of memorized passages.","marker":"[11]"},{"why":"Supplied the Harry Potter dedication that the model reproduced verbatim.","marker":"[12]"},{"why":"Supplied the Hitchhiker's Guide to the Galaxy passage that the model reproduced near-verbatim.","marker":"[13]"},{"why":"One of the paywalled news targets whose reproduction attempts all failed.","marker":"[14]"},{"why":"The other paywalled news target; the model refused direct requests and misattributed supplied quotes.","marker":"[15]"},{"why":"The DAN jailbreak template used in news-article attempts, which activated content safety rather than bypassing it.","marker":"[16]"},{"why":"The content-filtering configuration that the paper attributes the model's refusal behavior to.","marker":"[17]"},{"why":"The MIMIC-IV dataset used to test clinical-note and tabular-data reproduction.","marker":"[20]"},{"why":"The platform's customer copyright commitment, whose required user-side mitigations motivate the testing exercise.","marker":"[2]"}],"fun_headline_variants":["Hospital AI quoted Harry Potter dedication when nudged","Indirect prompts tricked medical AI into reciting books","Medical AI's literary recall exposed by red team","AI in cancer center reproduced book passage via indirect ask","Red team: indirect prompts made hospital AI recite novels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison of success rates across literary, news, science, and clinical content assumes those four kinds of material were similarly present in the model's training data; if the literary works were far better represented than the news and scientific articles, the observed differences would reflect data composition rather than protective mechanisms.","fun_headline_variants_meta":{"raw":{"variants":["Hospital AI quoted Harry Potter dedication when nudged","Indirect prompts tricked medical AI into reciting books","Medical AI's literary recall exposed by red team","AI in cancer center reproduced book passage via indirect ask","Red team: indirect prompts made hospital AI recite novels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000539,"raw_usage":{"total_tokens":2592,"prompt_tokens":961,"completion_tokens":1631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":1556}},"tokens_in":577,"tokens_out":1631,"duration_ms":11223,"temperature":1.0,"reasoning_tokens":1556,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:19:25.664839+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same indirect prompts on a production model using a literary passage known to be absent from its training data; if the model still returns it verbatim, training-data contamination is not the explanation and a different failure mode would be at work. Conversely, rerun the news tasks on articles as widely copied as the Rowling dedication; if extraction remains zero, then content-type-specific safeguards, not data availability, explain the difference.","supporting_citations":[],"review_version":1}