REVIEW 5 major objections 5 minor 57 references
Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper validates PDSQI-9, a nine-item instrument for scoring LLM-generated clinical summaries, using 779 real-world EHR summaries.
desk verdict A useful PDQI-9 adaptation with real data, but the reliability evidence as reported doesn't support 'robust construct validity'—the average-measure ICC hides a single-measure value around 0.57. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the PDSQI-9 rubric, a nine-attribute adaptation of the Physician Documentation Quality Instrument that targets LLM-specific failure modes: Cited, Accurate, Thorough, Useful, Organized, Comprehensible, Succinct, Synthesized, and Stigmatizing. The instrument carries the argument by converting qualitative judgments about summary quality into quantifiable scores, which are then tested under Messick's validity framework through factor analysis, reliability coefficients, and group comparisons.
What would settle it
Re-run the factor analysis and reliability statistics on all 779 evaluations, or on a documented random sample, and check whether the four-factor structure and the alpha/ICC values hold; also compare the 118 factor-analysis summaries with the remaining 661 on model type, specialty, and note length to test for selection bias.
Extended reading notes
Core claim
The central claim is that the PDSQI-9 is a valid and reliable instrument for evaluating LLM-generated summaries of clinical documentation. The paper presents evidence from content validity established through a semi-Delphi process, substantive validity shown by expected correlations with note length, structural validity supported by Cronbach's $\alpha = 0.879$ and a four-factor model explaining 58% of variance, generalizability supported by ICC $= 0.867$, and discriminant validity distinguishing high-quality from low-quality summaries at $p < 0.001$. The authors conclude that the instrument demonstrates sound construct validity and is ready for use in clinical practice to evaluate LLM-generated summaries before integration into healthcare workflows.
Load-bearing premise
The validity evidence depends on the 118 summaries used in the factor analysis, while the study reports 779 evaluations; the paper never states how that subset was selected, so if it is not representative of the full corpus the psychometric claims do not generalize.
Editorial extensions
If this is right
- Hospitals and health systems can use PDSQI-9 as a standardized human-evaluation step before deploying an LLM summarization tool on real patient notes.
- Benchmark comparisons of different LLMs or prompt strategies can be reported as nine attribute scores plus a four-factor profile rather than a single overall number.
- Attribute-level reliability estimates identify where scoring is easiest (Cited, Succinct) and where it is least stable (Comprehensible, Synthesized), guiding targeted improvements to prompts or retrieval.
- The finding that longer inputs are associated with lower Organized, Succinct, and Thorough scores implies that deployment decisions should account for note length, not just model choice.
Reading between the lines
- The authors do not test whether higher PDSQI-9 scores predict better downstream clinical decisions or reduced physician workload; that link is the logical next validation step beyond the paper's claims.
- The moderate Krippendorff's alpha (0.575) alongside high ICC suggests overall reliability is driven partly by attributes with low score variance, so users may need attribute-specific thresholds before treating a summary as safe.
- Synthesized's weak factor loading and the small number of summaries judged to present abstraction opportunities suggest abstractive quality is the least well-measured construct and may require a dedicated subscale.
- Because the validity evidence depends on the 118 summaries used in the factor analysis, an external replication should report the full-sample analysis and the sampling rule for that subset.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the PDSQI-9, a nine-item instrument for rating the quality of LLM-generated clinical summaries, and reports a validation study using summaries generated by GPT-4o, Mixtral 8x7b, and Llama 3-8b from real UW Health EHR notes, scored by seven physician raters. Validation analyses include Cronbach's alpha, ICC, Krippendorff's alpha, factor analysis, substantive correlations, and discriminant comparisons. The paper concludes that the PDSQI-9 demonstrates robust construct validity and is ready for use in clinical practice.
Significance. If the validity evidence were sound, this would be a useful contribution: the instrument addresses a real gap in evaluation of LLM-generated clinical summaries, uses real-world multi-document EHR data, employs multiple LLMs, incorporates a semi-Delphi content-validity process, and makes the instrument and prompts publicly available. However, as presented, the psychometric evidence is substantially overstated. The reported average-measure ICC does not support single-rater clinical use, the Krippendorff's alpha is moderate and below common thresholds, and the factor analysis rests on an unexplained 118-summary subset. These issues are load-bearing for the central claim, so the manuscript requires major revision before it can support the stated conclusions.
major comments (5)
- [Section 4 and Table 3] The abstract and conclusion describe 'high inter-rater reliability (ICC = 0.867)' based on ICC(3,k), an average-measure coefficient across the five evaluators specified in the Table 3 caption. For the stated clinical use, in which a single provider applies the instrument, the relevant coefficient is the single-measure ICC. Using the Spearman-Brown formula, the reported average ICC of 0.867 implies ICC(3,1) of approximately 0.867 / [5 - 4(0.867)] = 0.57, below the 0.75 threshold commonly expected for clinical instruments. Additionally, the reported 95% CI (0.867-0.868) is implausibly narrow for an ICC with 779 observations and five raters, suggesting a calculation or reporting error. This point is load-bearing for the generalizability claim and must be corrected.
- [Section 4 and Abstract] Krippendorff's alpha is reported as 0.575 (95% CI: 0.539-0.609), a value the Discussion itself calls 'moderate' and which falls below the commonly accepted 0.667 threshold. The Abstract omits this value entirely, and the conclusion nonetheless calls the evidence 'robust.' Because inter-rater reliability is one of the two pillars of the generalizability claim, the manuscript must report this coefficient prominently and temper its conclusions accordingly.
- [Appendix B and Section 4] The factor analysis output in Appendix B reports a harmonic n.obs of 118 and a total n.obs of 118, whereas the main validation corpus consists of 779 summaries and 8,329 item responses. The Results also mention 117 summaries for the abstraction question. The manuscript does not explain how the 118-summary subset was selected or whether it is representative of the full corpus. Without this information, the structural-validity conclusions cannot be generalized to the entire evaluation set.
- [Section 3.5 and Appendix B] The Methods section describes 'Confirmatory factor analysis,' but Appendix B is clearly an exploratory factor analysis (minres extraction, varimax rotation, factor selection based on eigenvalues and scree plot). Moreover, the reported fit is not unambiguously strong: RMSEA = 0.05 with a 90% CI of 0 to 0.198, and the loading matrix in Table 6 shows that Synthesized has no positive loading on any factor (its largest loading is -0.347 on MR4). The claim that the four-factor model provides strong support for construct validity is therefore overstated.
- [Section 3.5 and Section 4] Discriminant validity is established by comparing summaries generated with GPT-4o and error-free prompts (defined as high quality) against summaries from Llama 3-8b and Mixtral 8x7b with error-prone prompts (defined as low quality). Because the high/low classification is built into the prompt design and the instrument was developed by the same team with the same conceptual framework, this comparison shows that the instrument can distinguish two groups the authors designed to differ, but it does not establish that the instrument detects externally defined quality. An external gold standard or an independent criterion-based validation is needed to support the discriminant-validity claim.
minor comments (5)
- [Sections 3.3 and 4] The corpus size is reported inconsistently: Section 3.3 states that the final corpus had 200 summaries (100 GPT-4o, 50 Mixtral, 50 Llama), while Section 4 reports 779 summaries and 8,329 questions. Please clarify whether 779 refers to unique summaries, summary-rater pairs, or some other unit.
- [Table 3] The table reports a 'Cronbach's α' value for each individual attribute, and these values are numerically identical to the attribute's ICC. Cronbach's alpha is a scale-level statistic and is not defined for a single item; this column should be removed or replaced with an appropriate item-level statistic (e.g., alpha-if-item-deleted).
- [Section 3.5 and Appendix B] The Methods section says 'Confirmatory factor analysis,' but the analysis is exploratory. Please use the correct terminology throughout.
- [Appendix B] The table titles are duplicated ('PDSQI-9 Scores Eigenvalues' appears twice), and the second table is actually a factor-loading matrix; the labels should be corrected.
- [Section 3.5 and Table 3] The agreement coefficient is spelled both 'Krippendorf' and 'Krippendorff' in different places; please standardize to 'Krippendorff.'
Circularity Check
Discriminant-validity test is a manipulation check: 'low-quality' summaries were built by injecting the exact errors measured by PDSQI-9 items, so one pillar of 'robust construct validity' reduces to the study design; the remaining psychometric evidence is independent but not circular.
-
self definitional
[Section 3.3 (LLM Summarizations) and Section 3.5 (Analysis Plan and Validation)]
"To generate lower-quality summaries, additional variations of the prompt removed instructions or encouraged the inclusion of false information. ... Anti-Rules introduced intentional errors (e.g., hallucinations, omissions) to include mistakes. ... The summaries generated by GPT-4o with error-free prompts were considered the highest quality, while those generated by Llama 3-8b and Mixtral 8x7b with error-prone prompts were considered the lowest quality."
The PDSQI-9 defines poor quality by the same defects the Anti-Rules prompts were engineered to insert: 'Accurate' penalizes falsification/fabrication and 'Thorough' penalizes omissions. The 'high-quality' versus 'low-quality' groups are therefore constructed from the instrument's own target dimensions. The significant Mann-Whitney result (p < 0.001) shows that raters noticed the deliberately injected errors; it is a manipulation check, not independent evidence that the instrument measures a latent construct. Presenting this as discriminant validity supporting 'robust construct validity' makes that particular validity claim self-fulfilling rather than externally grounded.
full rationale
The majority of the psychometric work is not circular: Cronbach's alpha, ICC, Krippendorff's alpha, the four-factor model, and the note-length correlations are internal statistics computed on the ratings themselves, and they would stand or fall regardless of how the summaries were produced. The circular element is isolated to discriminant validity, where the known-groups contrast is defined by prompt engineering that injected the very error types the instrument items measure; this does not invalidate the internal-consistency evidence but weakens the conclusion that construct validity is 'robust.' Two non-circular concerns also affect interpretation: the factor analysis is reported on an unexplained subset (harmonic n.obs = 118 in Appendix B versus 779 summaries in the Abstract and Results), and the abstract's 'high inter-rater reliability' relies on average-measure ICC(3,k) = 0.867 across five raters while single-measure reliability and Krippendorff's alpha = 0.575 are moderate; these are correctness and reporting issues rather than circularity. Overall score 4 reflects one central validity sub-claim reducing by construction while the remaining validity evidence retains independent content.
Assumptions & free parameters
assumptions (3)
- domain assumption Likert-scale item scores can be treated as continuous for Pearson correlation and factor analysis.
- domain assumption The semi-Delphi expert consensus establishes content validity.
- domain assumption The sample size calculation assumption of an even score distribution is valid despite observed skew.
Cite this review
Pith. "Pith review of Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models." pith.science (2026). https://pith.science/paper/2MVE6T7D
@misc{pith2026250108977,
author = {Pith},
title = {Pith review of: Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/2MVE6T7D}},
note = {Machine review of arXiv:2501.08977}
}
abstract
As Large Language Models (LLMs) are integrated into electronic health record (EHR) workflows, validated instruments are essential to evaluate their performance before implementation. Existing instruments for provider documentation quality are often unsuitable for the complexities of LLM-generated text and lack validation on real-world data. The Provider Documentation Summarization Quality Instrument (PDSQI-9) was developed to evaluate LLM-generated clinical summaries. Multi-document summaries were generated from real-world EHR data across multiple specialties using several LLMs (GPT-4o, Mixtral 8x7b, and Llama 3-8b). Validation included Pearson correlation for substantive validity, factor analysis and Cronbach's alpha for structural validity, inter-rater reliability (ICC and Krippendorff's alpha) for generalizability, a semi-Delphi process for content validity, and comparisons of high-versus low-quality summaries for discriminant validity. Seven physician raters evaluated 779 summaries and answered 8,329 questions, achieving over 80% power for inter-rater reliability. The PDSQI-9 demonstrated strong internal consistency (Cronbach's alpha = 0.879; 95% CI: 0.867-0.891) and high inter-rater reliability (ICC = 0.867; 95% CI: 0.867-0.868), supporting structural validity and generalizability. Factor analysis identified a 4-factor model explaining 58% of the variance, representing organization, clarity, accuracy, and utility. Substantive validity was supported by correlations between note length and scores for Succinct (rho = -0.200, p = 0.029) and Organized ($\rho = -0.190$, $p = 0.037$). Discriminant validity distinguished high- from low-quality summaries ($p < 0.001$). The PDSQI-9 demonstrates robust construct validity, supporting its use in clinical practice to evaluate LLM-generated summaries and facilitate safer integration of LLMs into healthcare workflows.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Patterson BW, Hekman DJ, Liao FJ, Hamedani AG, Shah MN, Afshar M. Call me Dr Ishmael: trends in electronic health record notes available at emergency department visits and admissions. JAMIA Open. 2024 Apr;7(2):ooae039
work page 2024
-
[2]
To Err is Human: Building a Safer Health System
Institute of Medicine (US) Committee on Quality of Health Care in America. To Err is Human: Building a Safer Health System. Kohn LT, Corrigan JM, Donaldson MS, editors. Washington (DC): National Academies Press (US); 2000. Available from: http://www.ncbi.nlm.nih.gov/books/ NBK225182/
work page 2000
-
[3]
Embi PJ, Weir C, Efthimiadis EN, Thielke SM, Hedeen AN, Hammond KW. Computerized provider documentation: findings and implications of a multisite study of clinicians and administrators. Journal of the American Medical Informatics Association: JAMIA. 2013 Jan;20(4):718
work page 2013
-
[4]
Lost in the Middle: How Language Models Use Long Contexts
Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024;12:157–173
work page 2024
-
[5]
Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review
Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2024 Oct:e2421700
work page 2024
-
[6]
Tam TYC, Sivarajkumar S, Kapoor S, Stolyar A V, Polanska K, McCarthy KR, et al. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digital Medicine. 2024 Sep;7(1):258
work page 2024
-
[7]
Kernberg A, Gold JA, Mohan V. Using ChatGPT-4 to Create Structured Medical Notes From Audio Recordings of Physician-Patient Encounters: Comparative Study. Journal of Medical Internet Research. 2024 Apr;26:e54419. 11
work page 2024
-
[8]
Owens LM, Wilda JJ, Grifka R, Westendorp J, Fletcher JJ. Effect of Ambient Voice Technology, Natural Language Processing, and Artificial Intelligence on the Patient-Physician Relationship. Applied Clinical Informatics. 2024 Aug;15(4):660–667
work page 2024
Show all 57 references
-
[9]
Ambient Artifi- cial Intelligence Scribes to Alleviate the Burden of Clinical Documentation
Tierney AA, Gayre G, Hoberman B, Mattern B, Ballesca M, Kipnis P, et al. Ambient Artifi- cial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst. 2024 Feb;5(3):CAT.23.0404
2024
-
[10]
Assessing Electronic Note Quality Using the Physician Documentation Quality Instrument (PDQI-9)
Stetson PD, Bakken S, Wrenn JO, Siegler EL. Assessing Electronic Note Quality Using the Physician Documentation Quality Instrument (PDQI-9). Applied Clinical Informatics. 2012 Apr;3(2):164
2012
-
[11]
A Survey of Large Language Models
Zhao WX, Zhou K, Li J, Tang T, Wang X, Hou Y, et al. A Survey of Large Language Models. 2023 Jun;(arXiv:2303.18223). ArXiv:2303.18223 [cs]. Available from: http://arxiv.org/abs/2303. 18223
2023 arXiv
-
[12]
The Delphi Method: Techniques and Applications
Turoff M, Linstone HA. The Delphi Method: Techniques and Applications
-
[13]
A Survey of Evaluation Metrics Used for NLG Systems
Sai AB, Mohankumar AK, Khapra MM. A Survey of Evaluation Metrics Used for NLG Systems. ACM Computing Surveys. 2023;55(2)
2023
-
[14]
Generation of Patient After-Visit Summaries to Support Physicians
Cai P, Liu F, Bajracharya A, Sills J, Kapoor A, Liu W, et al. Generation of Patient After-Visit Summaries to Support Physicians. In: Proceedings of the 29th International Conference on Computational Lin- guistics. Gyeongju, Republic of Korea: International Committee on Computa...
2022
-
[15]
A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital- Course Summarization
Adams G, Zucker J, Elhadad N. A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital- Course Summarization. 2023 Mar;(arXiv:2303.03948). ArXiv:2303.03948 [cs]. Available from:http: //arxiv.org/abs/2303.03948
2023 arXiv
-
[16]
Large language models encode clinical knowledge
Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023 Jul:1–9
2023
-
[17]
Med-HALT: Medical Domain Hallucination Test for Large Language Models
Umapathi LK, Pal A, Sankarasubbu M. Med-HALT: Medical Domain Hallucination Test for Large Language Models. 2023 Jul;(arXiv:2307.15343). ArXiv:2307.15343 [cs, stat]. Available from: http: //arxiv.org/abs/2307.15343
2023 arXiv
-
[18]
Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization; 2020
Wallace BC, Saha S, Soboczenski F, Marshall IJ. Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization; 2020. Available from: https: //arxiv.org/abs/2008.11293v2
2020 arXiv
-
[19]
The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature
Otmakhova Y, Verspoor K, Baldwin T, Lau JH. The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1:...
2022
-
[20]
Revisiting Summarization Evaluation for Scientific Articles
Cohan A, Goharian N. Revisiting Summarization Evaluation for Scientific Articles
-
[21]
Development of a Human Evaluation Framework and Correlation with Automated Metrics for Natural Language Generation of Medical Di- agnoses
Croxford E, Gao Y, Patterson B, To D, Tesch S, Dligach D, et al. Development of a Human Evaluation Framework and Correlation with Automated Metrics for Natural Language Generation of Medical Di- agnoses. 2024 Apr:2024.03.20.24304620. Available from: https://www.medrxiv.org/con...
2024 doi
-
[22]
Reinforcement Learning for Abstractive Question Summarization with Question-aware Semantic Rewards
Yadav S, Gupta D, Abacha AB, Demner-Fushman D. Reinforcement Learning for Abstractive Question Summarization with Question-aware Semantic Rewards. 2021 Jun;(arXiv:2107.00176). ArXiv:2107.00176 [cs]. Available from: http://arxiv.org/abs/2107.00176
2021 arXiv
-
[23]
Automated Lay Language Summarization of Biomedical Scientific Reviews
Guo Y, Qiu W, Wang Y, Cohen T. Automated Lay Language Summarization of Biomedical Scientific Reviews. 2022 Jan;(arXiv:2012.12573). ArXiv:2012.12573 [cs]. Available from: http://arxiv. org/abs/2012.12573
2022 arXiv
-
[24]
An Investigation of Evaluation Metrics for Automated Medical Note Generation
Abacha AB, Yim Ww, Michalopoulos G, Lin T. An Investigation of Evaluation Metrics for Automated Medical Note Generation. 2023 May;(arXiv:2305.17364). ArXiv:2305.17364 [cs]. Available from: http://arxiv.org/abs/2305.17364
2023 arXiv
-
[25]
Research Electronic Data Capture (REDCap) - A metadata-driven methodology and workflow process for providing translational research informatics support
Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research Electronic Data Capture (REDCap) - A metadata-driven methodology and workflow process for providing translational research informatics support. Journal of biomedical informatics. 2008 Sep;42(2):377
2008
-
[26]
The REDCap consortium: Building an international community of software platform partners
Harris PA, Taylor R, Minor BL, Elliott V, Fernandez M, O’Neal L, et al. The REDCap consortium: Building an international community of software platform partners. Journal of Biomedical Informatics. 2019 Jul;95:103208
2019
-
[27]
GPT-4 Technical Report
OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, et al. GPT-4 Technical Report. 2024 Mar;(arXiv:2303.08774). ArXiv:2303.08774. Available from: http://arxiv.org/abs/2303. 08774
2024 arXiv
-
[28]
Mixtral of Experts
Jiang AQ, Sablayrolles A, Roux A, Mensch A, Savary B, Bamford C, et al. Mixtral of Experts. 2024 Jan;(arXiv:2401.04088). ArXiv:2401.04088. Available from: http://arxiv.org/abs/2401. 04088
2024 arXiv
-
[29]
The Llama 3 Herd of Models
Grattafiori A, Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, et al. The Llama 3 Herd of Models. 2024 Nov;(arXiv:2407.21783). ArXiv:2407.21783. Available from:http://arxiv.org/abs/2407. 21783
2024 arXiv
-
[30]
Available from: https://huggingface.co/
2024. Available from: https://huggingface.co/
2024
-
[31]
kappaSize: Sample Size Estimation Functions for Studies of Interobserver Agreement
Rotondi MA. kappaSize: Sample Size Estimation Functions for Studies of Interobserver Agreement
-
[32]
Words Matter: Strategies to Reduce Bias in Electronic Health Records
Canonico M. Words Matter: Strategies to Reduce Bias in Electronic Health Records. Words Matter
-
[33]
Standards of Validity and the Validity of Standards in Performance Asessment
Messick S. Standards of Validity and the Validity of Standards in Performance Asessment. Educational Measurement: Issues and Practice. 1995. Available from: https://onlinelibrary.wiley.com/ doi/10.1111/j.1745-3992.1995.tb00881.x
1995
-
[34]
Content Analysis: An Introduction to Its Methodology
Krippendorff K. Content Analysis: An Introduction to Its Methodology. SAGE Publications; 2018. Google-Books-ID: nE1aDwAAQBAJ
2018
-
[35]
In: Kotz S, Johnson NL, editors
Fisher RA. In: Kotz S, Johnson NL, editors. Statistical Methods for Research Workers. New York, NY: Springer; 1992. p. 66–70. Available from: https://doi.org/10.1007/978-1-4612-4380-9_6
1992 doi
-
[36]
A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research
Koo TK, Li MY. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine. 2016 Mar;15(2):155. 13
2016
-
[37]
Coefficient alpha and the internal structure of tests
Cronbach LJ. Coefficient alpha and the internal structure of tests. Psychometrika. 1951 Sep;16(3):297–334
1951
-
[38]
Statistical Inference for Coefficient Alpha
Feldt LS, Woodruff DJ, Salih FA. Statistical Inference for Coefficient Alpha. Applied Psychological Measurement. 1987 Mar;11(1):93–103
1987
-
[39]
Intraclass Correlations: Uses in Assessing Rater Reliability
Shrout PE, Fleiss JL. Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological Bulletin. 1979
1979
-
[40]
Transformers: State-of-the-Art Natural Language Processing
Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. Transformers: State-of-the-Art Natural Language Processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Online: Association for Computational L...
2020
-
[41]
Natural Language Processing with Python
Bird EL Steven, Klein E. Natural Language Processing with Python. O’Reilly Media Inc. 2009
2009
-
[42]
smplot: an R package for easy and elegant data visualization
Min SH, Zhou J. smplot: an R package for easy and elegant data visualization. Frontiers in Genetics. 2021;12:802894
2021
-
[43]
ggplot2: Elegant Graphics for Data Analysis
Wickham H. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York; 2016. Available from: https://ggplot2.tidyverse.org
2016
-
[44]
psych: Procedures for Psychological, Psychometric, and Personality Research
William Revelle. psych: Procedures for Psychological, Psychometric, and Personality Research. Evanston, Illinois; 2024. R package version 2.4.12. Available from: https://CRAN.R-project. org/package=psych
2024
-
[45]
krippendorffsalpha: An R Package for Measuring Agreement Using Krippendorff’s Alpha Coefficient
Hughes J. krippendorffsalpha: An R Package for Measuring Agreement Using Krippendorff’s Alpha Coefficient. 2021 Mar;(arXiv:2103.12170). ArXiv:2103.12170. Available from:http://arxiv.org/ abs/2103.12170
2021 arXiv
-
[46]
DescTools: Tools for Descriptive Statistics; 2017
et mult al AS. DescTools: Tools for Descriptive Statistics; 2017. R package version 0.99.23. Available from: https://cran.r-project.org/package=DescTools
2017
-
[47]
Toward Clinical Generative AI: Conceptual Framework
Bragazzi NL, Garbarino S. Toward Clinical Generative AI: Conceptual Framework. JMIR AI. 2024 Jun;3:e55957
2024
-
[48]
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
Croxford E, Gao Y, Pellegrino N, Wong KK, Wills G, First E, et al. Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review. 2024 Sep;(arXiv:2409.18170). ArXiv:2409.18170. Available from: http://arxiv.org/abs/2409.18170
2024 arXiv
-
[49]
Physician Use of Stigmatizing Language in Patient Medical Records
Park J, Saha S, Chee B, Taylor J, Beach MC. Physician Use of Stigmatizing Language in Patient Medical Records. JAMA network open. 2021 Jul;4(7):e2117052
2021
-
[50]
the pa6ent experienced shortness of breath, had an elevated white blood cell count, and showed a right lower lobe infiltrate on a chest X-ray,
Klang E, Apakama D, Abbott EE, Vaid A, Lampert J, Sakhuja A, et al. A strategy for cost-effective large language model use at health system-scale. NPJ digital medicine. 2024 nov;7(1):320. 14 6 Appendices 6.1 Appendix A: Time Spent per Evaluation by Evaluator Experience Figure ...
2024
-
[53]
Is the summary thorough without any omissions? Iden6fy any per6nent or poten6ally per6nent omissions: a. Per/nent omissions refer to essen6al informa6on required for the specific use case or intended provider, where missing details could directly impact pa6ent care decisions (i...
-
[54]
All the informa6on is in there that is useful to the target provider/intended audience b
Is the summary useful? a. All the informa6on is in there that is useful to the target provider/intended audience b. The summary is extremely relevant, providing valuable informa6on and/or analysis 1: Not at All 2 3 4 5: Extremely No asser$ons are per$nent to the target user So...
-
[55]
The summary is well-formed and structured in a way that helps the reader understand the pa6ent’s clinical course
Is the summary organized? a. The summary is well-formed and structured in a way that helps the reader understand the pa6ent’s clinical course. 1: Not at All 2 3 4 5: Extremely All Asser$ons presented out of order and groupings incoherent (completely disorganized) Some asser$on...
-
[56]
the pa6ent had an HbA1c of 9.2%, was prescribed mekormin, and reported not following dietary recommenda6ons
Is the summary succinct with economy of language? a. The summary is brief, to the point, and without redundancy 1: Not at All 2 3 4 5: Extremely Too wordy across all asser$ons with redundancy in syntax and seman$c More than one asser$on has contextual seman$c redundancy At lea...
-
[57]
The pa6ent's mother said that the ‘pills don't work’ for her child
Is there presence of s+gma+zing language? a. Please follow the tool in this link: h^ps://www.chcs.org/media/Words-Ma^er-Strategies-to-Reduce-Bias-in-Electronic-Health-Records_102022.pdf b. Refrain from using discredi6ng or exaggerated words, such as claims, insists, or reporte...
-
[2018]
Available from: https://cran.r-project.org/web/packages/kappaSize/index.html
-
[2020]
p. 38-45. Available from: https://www.aclweb.org/anthology/2020.emnlp-demos.6
2020
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.