Pith. sign in

REVIEW 5 major objections 6 minor 29 references

Privacy-Aware, Public-Aligned: Embedding Risk Detection and Public Values into Scalable Clinical Text De-Identification for Trusted Research Environments

T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Clinical note template changes degraded a live de-identification system's F1 score to 80%, far below its 95% target.

desk verdict Useful governance-focused synthesis of prior Scottish TRE work, but the headline temporal degradation claim rests on a single underspecified proprietary-system comparison. read the letter →

arxiv 2506.02063 v1 pith:JHS6HDX4 submitted 2025-06-01 cs.CR

classification cs.CR
keywords clinicalfreetextde-identificationprivacyriskindirectidentifierstrustedresearchenvironmentsnaturallanguageprocessingdocumentationdriftpublicengagement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper synthesises the authors' studies of clinical free text across Scottish NHS data providers to establish that privacy risk is context-dependent and cumulative, not a fixed property of a document. It shows that direct identifiers are two to three times more frequent in discharge summaries than in radiology reports, vary substantially across trusted research environments, and that a proprietary de-identification system scored 80% F1 (a combined measure of precision and recall) on Community Health Index numbers and dates against human labels, well below an implementation target of 95%, because documentation templates had changed since training. The paper identifies six recurring categories of indirect identifier risk and argues that implicit disclosures accumulate across a patient's records. It concludes that safe reuse of free text depends on hybrid de-identification pipelines, continuous monitoring, and governance tools that make risk decisions visible and auditable, in line with public expectations expressed in deliberative engagement.

What carries the argument

The load-bearing object is the multi-site annotated corpus built with a shared annotation schema for direct personal health identifiers, covering discharge summaries and radiology reports from three Scottish trusted research environments. The schema lets the authors compare identifier counts and annotator agreement across sites and document types, exposing how templates, system configuration, and processing workflows shape risk. Supporting machinery includes sentence-level neural topic modelling with class-based TF-IDF to surface indirect identifier categories, and a prototype dashboard that visualises cohort-level risk distributions, co-occurring risks, and patient-level profiles for governance review. The hybrid rule-plus-contextual-model pipeline is presented as the design response to the variation the corpus documents.

What would settle it

Re-measure the proprietary system's F1 on Community Health Index numbers and dates against a fresh sample drawn from the same documentation templates used at its original training time; a return to above 95% F1 on those older-format documents while current documents stay near 80% would confirm documentation drift, whereas similar low scores on both formats would point to annotation noise or a different failure cause.

Watch

Extended reading notes

Core claim

The central claim is that effective de-identification of clinical free text cannot be achieved by any single static model: identifier presence and form are shaped by record type, hospital site, data-processing workflow, and changes in documentation practice over time, and privacy risk is cumulative across linked records. Evidence comes from annotated samples of 2000 discharge summaries and 2000 radiology reports across three Scottish trusted research environments, where entity counts varied by more than an order of magnitude for some categories, and from a manual review in which a proprietary system obtained an F1 of 80% on Community Health Index numbers and dates against human annotation, compared with a deployment target above 95%. The performance drop was traced to shifts in how identifiers are presented in updated clinical documentation templates. The paper also maps indirect identifiers into six categories—unique medical diseases or events, social circumstances, locations and organisations, mental health, police, and crime—and notes that age and attendance frequency raise cumulative identifiability. The intended consequence is that de-identification must be embedded in auditable, context-aware governance workflows rather than treated as a one-off model output.

Load-bearing premise

Human annotators' labels are treated as the gold standard for measuring both identifier prevalence and model performance, even though some annotator agreement F1 scores in the paper are as low as 0, 9, 13, and 15 and no overall inter-annotator agreement statistic is reported.

Editorial extensions

If this is right

  • De-identification systems deployed in clinical settings need scheduled re-validation against fresh annotated samples, because identifier formatting changes with documentation templates and processing workflows.
  • Rule-based methods remain appropriate for well-structured identifiers such as postcodes, CHI numbers, and dates, while contextual identifiers such as patient names, hospital names, and occupations require models that use surrounding narrative.
  • Privacy-risk assessment should be cohort-level and cumulative, not per-document, because indirect risks cascade across reports and are compounded by patient age and frequency of attendance.
  • Governance workflows in trusted research environments should include visual, auditable risk dashboards that log de-risking decisions and support proportionate access decisions.
  • Public acceptance of free-text reuse is conditional on visible safeguards and explainable processes, so tooling and governance should be designed together with public engagement.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to contractually require rolling re-benchmarking of commercial de-identification tools against newly annotated local samples, since the paper demonstrates drift but does not propose a procurement mechanism.
  • The six indirect-risk categories could seed a structured annotation task for estimating cumulative re-identification risk across linked records, giving governance teams a quantitative prior rather than qualitative categories.
  • Comparing an LLM-based redaction system trained on the same corpus would directly test whether sub-word tokenisation makes model degradation worse or better when template headers change; the paper raises degradation but does not compare architectures.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript synthesises the authors' previous studies on de-identification of clinical free text in Scottish Trusted Research Environments. It reports annotated counts of direct identifiers across three TREs and two record types (discharge summaries and radiology reports), a qualitative exploration of indirect identifiers via BERTopic clustering, findings from public engagement activities, and a prototype R Shiny dashboard for privacy-risk visualisation. The central claims are that privacy risk in clinical free text is context-dependent and cumulative, that identifier prevalence varies substantially across sites, record types, and documentation practices, and that de-identification model performance degrades over time as documentation templates change.

Significance. If the empirical claims are supported, the paper offers valuable real-world evidence on identifier variation across NHS sites and record types, and a useful argument that de-identification systems need continuous monitoring and adaptation. The use of real NHS data across multiple TREs, the development of a common annotation schema, and the integration of public engagement into governance tool design are notable strengths. However, the quantitative support is thin: Table 1 includes an extrapolation for one TRE from 20% of records, Table 2 reports annotator F1 scores as low as 0, 9, and 13 without confidence intervals or inter-annotator agreement measures, and the sole quantitative evidence for the temporal-degradation claim is an underspecified comparison of an unnamed proprietary system against human labels. The cumulative-risk claim about indirect identifiers is stated without supporting quantitative or coded evidence.

major comments (5)
  1. [Results, Table 2] The paper reports annotator F1 scores that include extremely low values (0, 9, 13, and 15 across entity types and TREs) but does not specify how these F1 scores are computed, how many annotators participated, whether they measure pairwise inter-annotator agreement or agreement against a reference standard, or how disagreements were adjudicated. This matters directly for the later comparison between the proprietary system and 'the annotated data': if the human labels are noisy, the reported 80% F1 for CHI and dates may partly reflect annotation inconsistency rather than true system degradation. Please provide full details of the annotation procedure and report inter-annotator agreement measures (e.g., Cohen's kappa or pairwise F1 with confidence intervals).
  2. [Results, Direct Identifiers (proprietary system)] The central claim that de-identification performance degrades over time as documentation changes rests on a single, underspecified quantitative result: a proprietary system achieved F1=80% for CHI and dates against human labels, compared to an implementation target of >95%. The system is unnamed, the sample size and selection are not reported, no confidence intervals are given, no baseline measurement on historical pre-change data is provided, and the asserted link to 'changes in how certain identifiers presented' is not directly evidenced. As written, the 80% figure could reflect differences in annotation criteria, entity boundary conventions, or a mismatched label schema. This is the only quantitative support for the temporal-degradation component of the abstract's claim. Please provide a controlled before/after evaluation or, at minimum, the system name, comparison dates, sample sizes, and a description of the specific template changes and how they were verified.
  3. [Results, Implicit (Indirect) Identifiers] The claim that indirect risks occur in 'cascading patterns' and that 'age range of patients and the frequency of attendance emerged as crucial factors influencing the cumulative risk of identifiability' is not supported by any quantitative or systematically coded evidence in the paper. The description of BERTopic sentence clustering and manual review is qualitative, and no counts, topic-label reliability measures, or inter-rater assessments are provided. Without such evidence, the 'cumulative risk' claim is an assertion rather than a finding. This is a load-bearing element of the paper's framing ('privacy risk is context-dependent and cumulative'), so please add the relevant evidence or temper the claim accordingly.
  4. [Table 1] The counts for TRE 3 are extrapolated from 20% of records to a full 2000, but the table presents these extrapolated values without any uncertainty or caveat, and the surrounding text discusses them on the same footing as the full counts from the other TREs. This is particularly problematic because TRE 3 has the highest counts for many entity types (Names, Contact, IDs), and these extremes drive the observed across-TRE variation that is a central empirical claim. Please mark all estimated values clearly, describe the extrapolation method and its assumptions, and discuss the sensitivity of the conclusions to this extrapolation.
  5. [Methods and Results, Public Engagement] The public engagement activities are described only as 'a survey and workshops' with an independent facilitator, and the paper relies on external reports for details. Given that the paper uses public engagement as a justification for embedding 'public values' into the privacy-risk tool and as evidence that such tools are needed, the lack of basic methodological descriptors (number of participants, recruitment method, topic guide, analysis procedure) makes it difficult to assess the evidentiary weight of these claims. Please provide at least a brief summary of the methods and a direct citation to the specific report sections that contain the quoted findings.
minor comments (6)
  1. [Table 1] The entry '21,95' for Ward in TRE 2 Discharge Summaries appears to be a typo for '2,195'; please correct.
  2. [Table 2] Some cells in Table 2 contain values such as '09' in the Org Name column for TRE 1 Discharge Summaries, which appears to be a typo; also, the zero F1 values (e.g., Building Name for TRE 1) should be explained or annotated, as they may indicate zero occurrences or a total annotation mismatch.
  3. [Abstract and Introduction] The paper uses the acronym CHI (Community Health Index) without defining it at first use; please expand the abbreviation in the introduction or methods.
  4. [Methods] The Methods section says 'for more details the original papers are referenced,' but only reference 21 is cited for the direct-identifier schema and data. Please add explicit citations for the indirect-identifier analysis (reference 22) and the privacy-risk dashboard (reference 28) in the corresponding methods subsections.
  5. [Discussion] The comparison to Kraljevic et al. regarding passports and driving licenses is anecdotal ('This is something that we did not see in our data'); consider softening the statement or providing systematic evidence that such identifiers are absent in Scottish data.
  6. [General] The paper lacks a dedicated limitations section. Given the extrapolation in Table 1, the low annotator F1 scores in Table 2, and the underspecified proprietary-system evaluation, a limitations paragraph explicitly addressing these uncertainties would strengthen the manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's empirical claims rest on annotated entity counts, annotator agreement scores, and an external comparison with a proprietary de-identification system; self-citations are methodological pointers rather than load-bearing reductions.

full rationale

The paper is a synthesis of prior empirical studies and does not present a formal derivation whose output is equivalent to its input. The direct-identifier variation claims rest on annotated counts across three TREs (Table 1) and annotator F1 scores (Table 2), which are external evidence rather than consequences of the paper's own assumptions. The temporal-degradation claim rests on a comparison between a third-party de-identification system and human labels (F1=80% versus an implementation target above 95%); although the comparison is underspecified, it is an external benchmark rather than a fitted parameter renamed as a prediction. The indirect-identifier categories are derived from BERTopic clustering followed by manual privacy-risk review, so the categories are empirical groupings, not definitions that presuppose the conclusion. Public engagement was independently facilitated and analysed by IPSOS Scotland and is used as evidence of societal expectations, not as proof of a technical derivation. Self-citations appear as pointers to the annotation schema, indirect-risk methods, and prior reports, but none of these citations is load-bearing in the sense that the central claim reduces to the citation alone; the central claims about identifier variation and performance degradation are supported by the presented data and the external system comparison. No circular step can be exhibited by quoting the paper and showing an equation or fitted parameter that is equivalent to the claimed result, so the correct finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rest on annotation quality, the indirect-risk discovery pipeline, public engagement representativeness, and the TRE 3 extrapolation. None of these are independently evidenced in this paper beyond citations to prior work, and the paper does not introduce new physical or conceptual entities.

assumptions (4)
  • domain assumption Human annotator labels are an acceptable gold standard for measuring de-identification performance.
    Used throughout Results for F1 scores and for the claimed 80% versus >95% drop in the proprietary system; Table 2 itself shows high annotator variability (F1 as low as 0, 9, 13, 15).
  • domain assumption BERTopic sentence clustering plus manual review of topic labels can surface indirect identifier categories.
    The six indirect risk categories in Figure 2 rest on this pipeline with no quantitative validation of sensitivity or recall.
  • domain assumption Public engagement participants' views represent broader societal expectations.
    Used to justify embedding public values in the tool and recommendations; sample and recruitment details are only in cited IPSOS and SARA reports.
  • domain assumption The 20% sample from TRE 3 is representative of the remaining records.
    Table 1 extrapolates TRE 3 counts from 20% of records to 2000 records, assuming the sampled subset reflects the full set.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Privacy-Aware, Public-Aligned: Embedding Risk Detection and Public Values into Scalable Clinical Text De-Identification for Trusted Research Environments." pith.science (2026). https://pith.science/paper/JHS6HDX4

@misc{pith2026250602063,
  author       = {Pith},
  title        = {Pith review of: Privacy-Aware, Public-Aligned: Embedding Risk Detection and Public Values into Scalable Clinical Text De-Identification for Trusted Research Environments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JHS6HDX4}},
  note         = {Machine review of arXiv:2506.02063}
}
read the original abstract

Clinical free-text data offers immense potential to improve population health research such as richer phenotyping, symptom tracking, and contextual understanding of patient care. However, these data present significant privacy risks due to the presence of directly or indirectly identifying information embedded in unstructured narratives. While numerous de-identification tools have been developed, few have been tested on real-world, heterogeneous datasets at scale or assessed for governance readiness. In this paper, we synthesise our findings from previous studies examining the privacy-risk landscape across multiple document types and NHS data providers in Scotland. We characterise how direct and indirect identifiers vary by record type, clinical setting, and data flow, and show how changes in documentation practice can degrade model performance over time. Through public engagement, we explore societal expectations around the safe use of clinical free text and reflect these in the design of a prototype privacy-risk management tool to support transparent, auditable decision-making. Our findings highlight that privacy risk is context-dependent and cumulative, underscoring the need for adaptable, hybrid de-identification approaches that combine rule-based precision with contextual understanding. We offer a comprehensive view of the challenges and opportunities for safe, scalable reuse of clinical free-text within Trusted Research Environments and beyond, grounded in both technical evidence and public perspectives on responsible data use.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

29 extracted references · 25 canonical work pages

  1. [1]

    Kovačević, A., Bašaragin, B., Milošević, N., & Nenadić, G. (2024). De-identification of clinical free text using natural language processing: A systematic review of current approaches. Artificial intelligence in medicine, 151, 102845. https://doi.org/10.1016/j.artmed.2024.102845

  2. [2]

    Automatic de-identification of textual documents in the electronic health record: a review of recent research

    Meystre SM, Friedlin FJ, South BR, Shen S, Samore MH. Automatic de-identification of textual documents in the electronic health record: a review of recent research. BMC medical research methodology. 2010 Dec;10:1-6

  3. [3]

    (2024),https://doi.org/10.5281/zenodo.13353747

    Sudlow, C. (2024),https://doi.org/10.5281/zenodo.13353747

  4. [6]

    A., Sheldon, E

    Kraljevic, Z., Shek, A., Yeung, J. A., Sheldon, E. J., Shuaib, H., Al-Agil, M., Bai, X., Noor, K., Shah, A. D., Dobson, R., & Teo, J. (2023). Validating Transformers for Redaction of Text from Electronic Health Records in Real-World Healthcare. Proceedings - 2023 IEEE 11th International Conference on Healthcare Informatics, ICHI 2023, 544-

  5. [7]

    Deidentifying Medical Documents with Local, Privacy-Preserving Large Language Models: The LLM-Anonymizer

    Wiest IC, Leßmann ME, Wolf F, Ferber D, Treeck MV , Zhu J, Ebert MP, Westphalen CB, Wermke M, Kather JN. Deidentifying Medical Documents with Local, Privacy-Preserving Large Language Models: The LLM-Anonymizer. NEJM AI. 2025 Mar 27;2(4):AIdbp2400537

  6. [8]

    Kraljevic, Z., Bean, D., Shel, A., Bendayan, R., Hemingway, H., Yeung, J.A., Balsten, A., Ross, J., Idowu, E., Teo, J.T., Dobson, R.J.B. Foresight- a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study, The Lancet Digital Health, V olume 6, Issue 4, e281-e290

  7. [9]

    OCR Privacy Summary Rule, (Dated 05/03) URL:www.hhs.gov/sites/default/files/privacysummary.pdf [Accessed: 23 May 2025]

  8. [10]

    Data Protection Act 2018, c.12., URL: www.legislation.gov.uk/ukpga/2018/12/contents [Accessed: 20 May 2025]

Show all 29 references
  1. [11]

    Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus

    Stubbs A, Uzuner Ö. Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus. Journal of biomedical informatics. 2015 Dec 1;58:S20-9

  2. [12]

    Deidentification of free-text medical records using pre- trained bidirectional transformers

    Johnson AE, Bulgarelli L, Pollard TJ. Deidentification of free-text medical records using pre- trained bidirectional transformers. In Proceedings of the ACM conference on health, inference, and learning 2020 Apr 2 (pp. 214-221)

  3. [13]

    Stubbs, A., Kotfila, C., & Uzuner, Ö. (2015). Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth shared task Track

  4. [14]

    and Lehman, L.W.H., 2023

    Johnson, A.E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T.J., Hao, S., Moody, B., Gow, B. and Lehman, L.W.H., 2023. MIMIC-IV , a freely accessible electronic health record dataset. Scientific data, 10(1), p.1

  5. [15]

    https://doi.org/10.1016/j.jbi.2015.06.007

    Journal of biomedical informatics, 58 Suppl(Suppl), S11–S19. https://doi.org/10.1016/j.jbi.2015.06.007

  6. [16]

    Word embeddings trained on published case reports are lightweight, effective for clinical tasks, and free of protected health information

    Flamholz ZN, Crane-Droesch A, Ungar LH, Weissman GE. Word embeddings trained on published case reports are lightweight, effective for clinical tasks, and free of protected health information. Journal of biomedical informatics. 2022 Jan 1;125:103971

  7. [17]

    An empirical test of GRUs and deep contextualized word representations on de-identification

    Lee K, Filannino M, Uzuner Ö. An empirical test of GRUs and deep contextualized word representations on de-identification. In MEDINFO 2019: Health and Wellbeing e-Networks for All 2019 (pp. 218-222). IOS Press

  8. [18]

    and Ardhanari, S., 2021

    Murugadoss, K., Rajasekharan, A., Malin, B., Agarwal, V ., Bade, S., Anderson, J.R., Ross, J.L., Faubion, W.A., Halamka, J.D., Soundararajan, V . and Ardhanari, S., 2021. Building a best-in-class automated de-identification tool for electronic health records through ensemble l...

  9. [19]

    and Hua, W., 2015

    He, B., Guan, Y ., Cheng, J., Cen, K. and Hua, W., 2015. CRFs based de-identification of medical records. Journal of biomedical informatics, 58, pp.S39-S46

  10. [20]

    K., Dobson, R., Roberts, A., Jones, K., Shah, A

    Fitzpatrick, N. K., Dobson, R., Roberts, A., Jones, K., Shah, A. D., Nenadic, G., & Ford, E. (2023). Understanding Views Around the Creation of a Consented, Donated Databank of Clinical Free Text to Develop and Train Natural Language Processing Models for Research: Focus Group...

  11. [21]

    Ford, E., Oswald, M., Hassan, L., Bozentko, K., Nenadic, G., & Cassell, J. (2020). Should free-text data in electronic medical records be shared for research? A citizens' jury study in the UK. Journal of medical ethics, 46(6), 367–377. https://doi.org/10.1136/medethics-2019- 105472

  12. [22]

    (2024, June 12)

    Casey, A., Falis, M., Franz, G., Tilbrook, A., Ford, E., & Harrison, K. (2024, June 12). Beyond the Surface – Exploring and Defining Indirect Risks in Clinical Free-text. HealTAC 2024, Lancaster. https://doi.org/10.5281/zenodo.15118875

  13. [23]

    S., Murrell, M., Nita, S., Tilbrook, A., Mayor, C., O’Sullivan, K., & Harrison, K

    Casey, A., Falis, M., Gruber, F. S., Murrell, M., Nita, S., Tilbrook, A., Mayor, C., O’Sullivan, K., & Harrison, K. (2024). Developing a Common Schema for De-identification of Personal Health Identifiers in EHRs across Scotland. HealTAC 2024, Lancaster, United Kingdom. https:/...

  14. [24]

    BERTopic: Neural topic modeling with a class-based TF-IDF procedure

    Grootendorst, M., 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794

  15. [25]

    Casey, A., Tilbrook, A., Dunbar, S., Linksted, P., Harrison, K., Mills, N., O'Sullivan, K., Markovic, M., Ciocarlan, A., Wilde, K., Mayor, C., Caldwell, J., Ford E. (2023). SARA: Semi-Automated Risk Assessment of Data Provenance and Clinical Free-text in trusted research envir...

  16. [26]

    Mulholland, C., Simpson, E., Abernethy, S. (2023). Risk assessment and mitigation in health data research. Findings from deliberative workshops and an online survey. Zenodo. https://doi.org/10.5281/zenodo.10229070

  17. [27]

    IPSOS Scotland, URL:https://www.ipsos.com/en-uk/scotland [accessed 01 June 2025]

  18. [28]

    (2024, June 12)

    Gruber, F., Falis, M., & Casey, A. (2024, June 12). A Privacy Risk Dashboard for Clinical Free-text. HealTAC2024, Lancaster. https://doi.org/10.5281/zenodo.15556569

  19. [29]

    Dunbar, S., O'Sullivan, K., Tilbrook, A., Ford, E., Linksted, P., Mayor, C., Caldwell, J., Markovic, M., Ciocarlan, A., Harrison, K., Mills, N., Wilde, K., Casey, A. (2023). SARA Public Involvement and Engagement Final Report. Zenodo. https://doi.org/10.5281/zenodo.10084410

  20. [31]

    and Teo, J., 2023, June

    Kraljevic, Z., Shek, A., Yeung, J.A., Sheldon, E.J., Shuaib, H., Al-Agil, M., Bai, X., Noor, K., Shah, A.D., Dobson, R. and Teo, J., 2023, June. Validating transformers for redaction of text from electronic health records in real-world healthcare. In 2023 IEEE 11th Internation...

  21. [549]

    https://doi.org/10.1109/ICHI57859.2023.00098

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.