{"id":"9dc28b38-511e-4ddc-b310-d84b377b6f3a","arxiv_id":"2506.02063","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Privacy risk in clinical free text is context-dependent and cumulative, so safe reuse requires hybrid, continuously monitored de-identification embedded in auditable governance.","lead":"Using real Scottish NHS notes, this paper shows that identifying information in clinical free text varies sharply by hospital, note type, and documentation practice, and that de-identification tools lose accuracy when those practices change. It argues for hybrid, continuously monitored, publicly accountable de-identification instead of a one-time model.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The temporal-degradation claim rests on an underspecified proprietary-system comparison; without a controlled before/after evaluation, the 80%-versus-95% drop could reflect annotation or protocol mismatch.","rationale":"The reader's weakest_assumption correctly identifies annotator label noise as a threat to the quantitative evidence. I agree that low annotator F1 undermines the gold-standard assumption and that lack of inter-annotator agreement metrics is a serious gap. However, I see a more specific and more central problem: even if annotator agreement were perfect, the paper's single quantitative evidence for temporal degradation is an underspecified comparison against an unnamed proprietary system, with no historical baseline and no demonstration that the 80% F1 reflects documentation drift rather than annotation-schema mismatch or test-set selection. The reader's concern about the gold standard is a necessary condition for trusting the 80% figure, but it is not sufficient: a clean gold standard alone would not establish that the drop is caused by template changes unless the same system is evaluated on data from before and after those changes under identical protocols. Thus my concern partially overlaps with the reader's but is distinct in its emphasis on the missing controlled before/after comparison. The qualitative claims about context-dependence and cumulative risk are reasonable and consistent with prior work, so the paper does not warrant rejection. The conditional verdict stands, with the additional requirement that the proprietary-system comparison be fully documented and, ideally, re-run as a controlled temporal evaluation.","tokens_in":9727,"tokens_out":2563,"duration_ms":29845,"concrete_test":"Obtain the proprietary system's outputs on the 2,000 discharge summaries and 2,000 radiology reports used in the annotation study. Stratified random sample 200 reports with CHI and dates, and have two independent annotators label them with adjudication to measure inter-annotator agreement. Compute F1 for current data. Then retrieve a matched cohort of reports from before the documented template change (e.g., 2019–2021) from the same TREs, run the same proprietary system on that historical set, and apply the identical annotation protocol. If current F1 is ~80% with high annotator agreement and historical F1 is >95%, the degradation claim is confirmed; if the gap shrinks or disappears, the reported drop is likely an artifact of annotation inconsistency or protocol mismatch.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim that de-identification performance degrades over time as documentation changes is supported by one quantitative result: a proprietary de-identification system scored F1=80% on CHI and dates against human labels, versus an implementation target of >95% (Results). This single comparison carries the entire 'degradation over time' component of the argument, but the comparison is described in only two sentences: the system is unnamed, the sample size and selection are unspecified, no confidence intervals are given, and no baseline measurement of the same system on historical (pre-template-change) data is reported. The paper states that 'changes in how certain identifiers presented had occurred since the system's original training phase' but provides no direct evidence linking the 80% F1 to those changes rather than to differences in annotation criteria, entity boundary conventions, or the gold-standard labels themselves. This concern is compounded by Table 2, which shows annotator F1 scores as low as 0–15 for several entity types across TREs, indicating substantial label noise. If the 80% figure is an artifact of comparing a system tuned to one annotation schema against labels produced under a different schema—or of noisy reference labels—then the performance-degradation claim loses its only direct quantitative support. The surrounding qualitative evidence for context-dependence remains plausible, but the headline claim about temporal degradation needs a controlled comparison to be credible.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript synthesises the authors' previous studies on de-identification of clinical free text in Scottish Trusted Research Environments. It reports annotated counts of direct identifiers across three TREs and two record types (discharge summaries and radiology reports), a qualitative exploration of indirect identifiers via BERTopic clustering, findings from public engagement activities, and a prototype R Shiny dashboard for privacy-risk visualisation. The central claims are that privacy risk in clinical free text is context-dependent and cumulative, that identifier prevalence varies substantially across sites, record types, and documentation practices, and that de-identification model performance degrades over time as documentation templates change.","tokens_in":9915,"tokens_out":5392,"duration_ms":59798,"significance":"If the empirical claims are supported, the paper offers valuable real-world evidence on identifier variation across NHS sites and record types, and a useful argument that de-identification systems need continuous monitoring and adaptation. The use of real NHS data across multiple TREs, the development of a common annotation schema, and the integration of public engagement into governance tool design are notable strengths. However, the quantitative support is thin: Table 1 includes an extrapolation for one TRE from 20% of records, Table 2 reports annotator F1 scores as low as 0, 9, and 13 without confidence intervals or inter-annotator agreement measures, and the sole quantitative evidence for the temporal-degradation claim is an underspecified comparison of an unnamed proprietary system against human labels. The cumulative-risk claim about indirect identifiers is stated without supporting quantitative or coded evidence.","major_comments":[{"comment":"The paper reports annotator F1 scores that include extremely low values (0, 9, 13, and 15 across entity types and TREs) but does not specify how these F1 scores are computed, how many annotators participated, whether they measure pairwise inter-annotator agreement or agreement against a reference standard, or how disagreements were adjudicated. This matters directly for the later comparison between the proprietary system and 'the annotated data': if the human labels are noisy, the reported 80% F1 for CHI and dates may partly reflect annotation inconsistency rather than true system degradation. Please provide full details of the annotation procedure and report inter-annotator agreement measures (e.g., Cohen's kappa or pairwise F1 with confidence intervals).","section":"Results, Table 2"},{"comment":"The central claim that de-identification performance degrades over time as documentation changes rests on a single, underspecified quantitative result: a proprietary system achieved F1=80% for CHI and dates against human labels, compared to an implementation target of >95%. The system is unnamed, the sample size and selection are not reported, no confidence intervals are given, no baseline measurement on historical pre-change data is provided, and the asserted link to 'changes in how certain identifiers presented' is not directly evidenced. As written, the 80% figure could reflect differences in annotation criteria, entity boundary conventions, or a mismatched label schema. This is the only quantitative support for the temporal-degradation component of the abstract's claim. Please provide a controlled before/after evaluation or, at minimum, the system name, comparison dates, sample sizes, and a description of the specific template changes and how they were verified.","section":"Results, Direct Identifiers (proprietary system)"},{"comment":"The claim that indirect risks occur in 'cascading patterns' and that 'age range of patients and the frequency of attendance emerged as crucial factors influencing the cumulative risk of identifiability' is not supported by any quantitative or systematically coded evidence in the paper. The description of BERTopic sentence clustering and manual review is qualitative, and no counts, topic-label reliability measures, or inter-rater assessments are provided. Without such evidence, the 'cumulative risk' claim is an assertion rather than a finding. This is a load-bearing element of the paper's framing ('privacy risk is context-dependent and cumulative'), so please add the relevant evidence or temper the claim accordingly.","section":"Results, Implicit (Indirect) Identifiers"},{"comment":"The counts for TRE 3 are extrapolated from 20% of records to a full 2000, but the table presents these extrapolated values without any uncertainty or caveat, and the surrounding text discusses them on the same footing as the full counts from the other TREs. This is particularly problematic because TRE 3 has the highest counts for many entity types (Names, Contact, IDs), and these extremes drive the observed across-TRE variation that is a central empirical claim. Please mark all estimated values clearly, describe the extrapolation method and its assumptions, and discuss the sensitivity of the conclusions to this extrapolation.","section":"Table 1"},{"comment":"The public engagement activities are described only as 'a survey and workshops' with an independent facilitator, and the paper relies on external reports for details. Given that the paper uses public engagement as a justification for embedding 'public values' into the privacy-risk tool and as evidence that such tools are needed, the lack of basic methodological descriptors (number of participants, recruitment method, topic guide, analysis procedure) makes it difficult to assess the evidentiary weight of these claims. Please provide at least a brief summary of the methods and a direct citation to the specific report sections that contain the quoted findings.","section":"Methods and Results, Public Engagement"}],"minor_comments":[{"comment":"The entry '21,95' for Ward in TRE 2 Discharge Summaries appears to be a typo for '2,195'; please correct.","section":"Table 1"},{"comment":"Some cells in Table 2 contain values such as '09' in the Org Name column for TRE 1 Discharge Summaries, which appears to be a typo; also, the zero F1 values (e.g., Building Name for TRE 1) should be explained or annotated, as they may indicate zero occurrences or a total annotation mismatch.","section":"Table 2"},{"comment":"The paper uses the acronym CHI (Community Health Index) without defining it at first use; please expand the abbreviation in the introduction or methods.","section":"Abstract and Introduction"},{"comment":"The Methods section says 'for more details the original papers are referenced,' but only reference 21 is cited for the direct-identifier schema and data. Please add explicit citations for the indirect-identifier analysis (reference 22) and the privacy-risk dashboard (reference 28) in the corresponding methods subsections.","section":"Methods"},{"comment":"The comparison to Kraljevic et al. regarding passports and driving licenses is anecdotal ('This is something that we did not see in our data'); consider softening the statement or providing systematic evidence that such identifiers are absent in Scottish data.","section":"Discussion"},{"comment":"The paper lacks a dedicated limitations section. Given the extrapolation in Table 1, the low annotator F1 scores in Table 2, and the underspecified proprietary-system evaluation, a limitations paragraph explicitly addressing these uncertainties would strengthen the manuscript.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a synthesis of the authors' own prior work, and the reference list relies heavily on self-citations (e.g., refs. 4, 5, 21, 22, 23, 26, 27, 28). That is acceptable for a synthesis paper if the new framing adds value, but the novelty is limited to the integration of the results rather than new primary evidence. The paper's fit with a cryptography/security venue (cs.CR) is questionable; it is primarily a health-informatics and governance contribution. The most pressing concern is the unsupported temporal-degradation claim, which the authors may find difficult to substantiate without additional controlled experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a synthesis paper, not a new research contribution. The authors repackage their earlier annotation, public engagement, and dashboard work from Scottish TREs. That is fine if the contribution is framed as a governance and procurement-oriented summary, but the one genuinely new quantitative claim — that a proprietary de-identification system's F1 fell to 80% on CHI and dates versus a >95% target — is the weakest part of the paper.\n\nWhat the paper does well: it gives a concrete picture of how direct identifiers vary across record types, sites, and TRE configurations, and it is honest about annotator disagreement. The public engagement work is real and independently facilitated by Ipsos. The recommendation for hybrid, continuously monitored de-identification is sensible and consistent with the literature.\n\nWhere it gets soft: the degradation story is carried by a comparison described in two sentences. The system is unnamed; there is no sample size, no confidence interval, and no baseline measurement on pre-template-change data. The stress-test note is right: the 80% figure could reflect annotation-schema mismatch or noisy gold labels rather than true temporal drift. Table 2 makes that worry concrete, with annotator F1 scores as low as 0, 9, and 13. The paper does not report inter-annotator agreement, and the TRE 3 counts are extrapolated from 20% of records. These are addressable, but right now the quantitative support for the central claim is thin.\n\nThere is also a mild confirmatory loop: public engagement is used both to shape the tool and as evidence that the tool is needed. That should be acknowledged more explicitly.\n\nWho this is for: health-data governance leads, TRE operators, and NLP people who want a real-world cautionary tale about deployment drift. Methodologists will not find new techniques here.\n\nMy recommendation: send it to peer review, but the referee should insist on details of the proprietary-system comparison, inter-annotator agreement, and confidence intervals. With those additions, the paper would be a credible synthesis. Without them, the degradation claim should be downgraded to anecdotal.","headline":"Useful governance-focused synthesis of prior Scottish TRE work, but the headline temporal degradation claim rests on a single underspecified proprietary-system comparison.","tokens_in":10514,"tokens_out":2356,"would_cite":false,"duration_ms":24577,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Clinical note template changes degraded a live de-identification system's F1 score to 80%, far below its 95% target.","keywords":["clinical free text","de-identification","privacy risk","indirect identifiers","trusted research environments","natural language processing","documentation drift","public engagement"],"falsifier":"Re-measure the proprietary system's F1 on Community Health Index numbers and dates against a fresh sample drawn from the same documentation templates used at its original training time; a return to above 95% F1 on those older-format documents while current documents stay near 80% would confirm documentation drift, whereas similar low scores on both formats would point to annotation noise or a different failure cause.","tokens_in":9493,"feed_emoji":"🔒","tokens_out":7337,"duration_ms":71721,"temperature":0.7,"pith_summary":"This paper synthesises the authors' studies of clinical free text across Scottish NHS data providers to establish that privacy risk is context-dependent and cumulative, not a fixed property of a document. It shows that direct identifiers are two to three times more frequent in discharge summaries than in radiology reports, vary substantially across trusted research environments, and that a proprietary de-identification system scored 80% F1 (a combined measure of precision and recall) on Community Health Index numbers and dates against human labels, well below an implementation target of 95%, because documentation templates had changed since training. The paper identifies six recurring categories of indirect identifier risk and argues that implicit disclosures accumulate across a patient's records. It concludes that safe reuse of free text depends on hybrid de-identification pipelines, continuous monitoring, and governance tools that make risk decisions visible and auditable, in line with public expectations expressed in deliberative engagement.","feed_headline":"De-identification accuracy fell to 80% as NHS notes changed","feed_subtitle":"Real-world clinical notes drift over time, so de-identification needs re-validation and auditable governance.","key_machinery":"The load-bearing object is the multi-site annotated corpus built with a shared annotation schema for direct personal health identifiers, covering discharge summaries and radiology reports from three Scottish trusted research environments. The schema lets the authors compare identifier counts and annotator agreement across sites and document types, exposing how templates, system configuration, and processing workflows shape risk. Supporting machinery includes sentence-level neural topic modelling with class-based TF-IDF to surface indirect identifier categories, and a prototype dashboard that visualises cohort-level risk distributions, co-occurring risks, and patient-level profiles for governance review. The hybrid rule-plus-contextual-model pipeline is presented as the design response to the variation the corpus documents.","core_discovery":"The central claim is that effective de-identification of clinical free text cannot be achieved by any single static model: identifier presence and form are shaped by record type, hospital site, data-processing workflow, and changes in documentation practice over time, and privacy risk is cumulative across linked records. Evidence comes from annotated samples of 2000 discharge summaries and 2000 radiology reports across three Scottish trusted research environments, where entity counts varied by more than an order of magnitude for some categories, and from a manual review in which a proprietary system obtained an F1 of 80% on Community Health Index numbers and dates against human annotation, compared with a deployment target above 95%. The performance drop was traced to shifts in how identifiers are presented in updated clinical documentation templates. The paper also maps indirect identifiers into six categories—unique medical diseases or events, social circumstances, locations and organisations, mental health, police, and crime—and notes that age and attendance frequency raise cumulative identifiability. The intended consequence is that de-identification must be embedded in auditable, context-aware governance workflows rather than treated as a one-off model output.","pith_inferences":["A testable extension is to contractually require rolling re-benchmarking of commercial de-identification tools against newly annotated local samples, since the paper demonstrates drift but does not propose a procurement mechanism.","The six indirect-risk categories could seed a structured annotation task for estimating cumulative re-identification risk across linked records, giving governance teams a quantitative prior rather than qualitative categories.","Comparing an LLM-based redaction system trained on the same corpus would directly test whether sub-word tokenisation makes model degradation worse or better when template headers change; the paper raises degradation but does not compare architectures."],"forward_implications":["De-identification systems deployed in clinical settings need scheduled re-validation against fresh annotated samples, because identifier formatting changes with documentation templates and processing workflows.","Rule-based methods remain appropriate for well-structured identifiers such as postcodes, CHI numbers, and dates, while contextual identifiers such as patient names, hospital names, and occupations require models that use surrounding narrative.","Privacy-risk assessment should be cohort-level and cumulative, not per-document, because indirect risks cascade across reports and are compounded by patient age and frequency of attendance.","Governance workflows in trusted research environments should include visual, auditable risk dashboards that log de-risking decisions and support proportionate access decisions.","Public acceptance of free-text reuse is conditional on visible safeguards and explainable processes, so tooling and governance should be designed together with public engagement."],"supporting_citations":[{"why":"Defines the shared annotation schema for direct personal health identifiers across Scottish trusted research environments used to create the labelled corpus.","marker":"[21]"},{"why":"Sets out the method and earlier findings on indirect identifiers in clinical free text that this paper synthesises.","marker":"[22]"},{"why":"Provides the sentence-level neural topic modelling and class-based TF-IDF technique used to surface indirect-risk categories.","marker":"[24]"},{"why":"Documents that hybrid models outperform single approaches and that pre-trained embeddings do not always transfer, framing the paper's hybrid recommendation.","marker":"[12]"},{"why":"Supplies the MIMIC-IV corpus as a benchmark that is single-institution and sanitised, against which the paper contrasts real-world multi-site variation.","marker":"[14]"},{"why":"Shows through a citizens' jury that public acceptance of free-text research is conditional on transparency and safeguards, grounding the public-engagement claims.","marker":"[19]"},{"why":"Gives a UK de-identification example whose PHI categories differ from Scottish data, illustrating why local annotation evidence is needed.","marker":"[29]"}],"fun_headline_variants":["De-id accuracy drops to 80% as clinical notes change","Public values and risk detection for clinical text de-id","Cumulative identifier risk across records drives governance","Hybrid de-identification needed for evolving clinical notes","Auditable governance beats static models for clinical text privacy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Human annotators' labels are treated as the gold standard for measuring both identifier prevalence and model performance, even though some annotator agreement F1 scores in the paper are as low as 0, 9, 13, and 15 and no overall inter-annotator agreement statistic is reported.","fun_headline_variants_meta":{"raw":{"variants":["De-id accuracy drops to 80% as clinical notes change","Public values and risk detection for clinical text de-id","Cumulative identifier risk across records drives governance","Hybrid de-identification needed for evolving clinical notes","Auditable governance beats static models for clinical text privacy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000601,"raw_usage":{"total_tokens":2834,"prompt_tokens":1002,"completion_tokens":1832,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1755}},"tokens_in":618,"tokens_out":1832,"duration_ms":14382,"temperature":1.0,"reasoning_tokens":1755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:49:58.577082+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-measure the proprietary system's F1 on Community Health Index numbers and dates against a fresh sample drawn from the same documentation templates used at its original training time; a return to above 95% F1 on those older-format documents while current documents stay near 80% would confirm documentation drift, whereas similar low scores on both formats would point to annotation noise or a different failure cause.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the shared annotation schema for direct personal health identifiers across Scottish trusted research environments used to create the labelled corpus."},{"cited_title":"(2024, June 12)","cited_arxiv_id":null,"evidence_quote":"Sets out the method and earlier findings on indirect identifiers in clinical free text that this paper synthesises."},{"cited_title":"Deidentification of free-text medical records using pre- trained bidirectional transformers","cited_arxiv_id":null,"evidence_quote":"Documents that hybrid models outperform single approaches and that pre-trained embeddings do not always transfer, framing the paper's hybrid recommendation."},{"cited_title":"and Hua, W., 2015","cited_arxiv_id":null,"evidence_quote":"Shows through a citizens' jury that public acceptance of free-text research is conditional on transparency and safeguards, grounding the public-engagement claims."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives a UK de-identification example whose PHI categories differ from Scottish data, illustrating why local annotation evidence is needed."}],"review_version":1}