Pith. sign in

REVIEW 5 major objections 5 minor 57 references

Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper validates PDSQI-9, a nine-item instrument for scoring LLM-generated clinical summaries, using 779 real-world EHR summaries.

desk verdict A useful PDQI-9 adaptation with real data, but the reliability evidence as reported doesn't support 'robust construct validity'—the average-measure ICC hides a single-measure value around 0.57. read the letter →

arxiv 2501.08977 v2 pith:2MVE6T7D submitted 2025-01-15 cs.AI

classification cs.AI
keywords PDSQI-9largelanguagemodelsclinicalsummarizationelectronichealthrecordspsychometricvalidationconstructvalidityinter-raterreliabilityhumanevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces and validates the Provider Documentation Summarization Quality Instrument (PDSQI-9), a nine-item rubric for scoring how well large language models summarize electronic health record notes. Seven physician raters applied the instrument to 779 summaries generated from real-world EHR data by GPT-4o, Mixtral 8x7b, and Llama 3-8b, producing 8,329 item scores. The paper reports high internal consistency (Cronbach's $\alpha = 0.879$), high inter-rater reliability (ICC $= 0.867$), and a four-factor structure corresponding to organization, clarity, accuracy, and utility. The authors argue this constitutes strong construct validity and that the instrument can support safer integration of LLM summarization into clinical workflows.

What carries the argument

The central object is the PDSQI-9 rubric, a nine-attribute adaptation of the Physician Documentation Quality Instrument that targets LLM-specific failure modes: Cited, Accurate, Thorough, Useful, Organized, Comprehensible, Succinct, Synthesized, and Stigmatizing. The instrument carries the argument by converting qualitative judgments about summary quality into quantifiable scores, which are then tested under Messick's validity framework through factor analysis, reliability coefficients, and group comparisons.

What would settle it

Re-run the factor analysis and reliability statistics on all 779 evaluations, or on a documented random sample, and check whether the four-factor structure and the alpha/ICC values hold; also compare the 118 factor-analysis summaries with the remaining 661 on model type, specialty, and note length to test for selection bias.

Watch

Extended reading notes

Core claim

The central claim is that the PDSQI-9 is a valid and reliable instrument for evaluating LLM-generated summaries of clinical documentation. The paper presents evidence from content validity established through a semi-Delphi process, substantive validity shown by expected correlations with note length, structural validity supported by Cronbach's $\alpha = 0.879$ and a four-factor model explaining 58% of variance, generalizability supported by ICC $= 0.867$, and discriminant validity distinguishing high-quality from low-quality summaries at $p < 0.001$. The authors conclude that the instrument demonstrates sound construct validity and is ready for use in clinical practice to evaluate LLM-generated summaries before integration into healthcare workflows.

Load-bearing premise

The validity evidence depends on the 118 summaries used in the factor analysis, while the study reports 779 evaluations; the paper never states how that subset was selected, so if it is not representative of the full corpus the psychometric claims do not generalize.

Editorial extensions

If this is right

  • Hospitals and health systems can use PDSQI-9 as a standardized human-evaluation step before deploying an LLM summarization tool on real patient notes.
  • Benchmark comparisons of different LLMs or prompt strategies can be reported as nine attribute scores plus a four-factor profile rather than a single overall number.
  • Attribute-level reliability estimates identify where scoring is easiest (Cited, Succinct) and where it is least stable (Comprehensible, Synthesized), guiding targeted improvements to prompts or retrieval.
  • The finding that longer inputs are associated with lower Organized, Succinct, and Thorough scores implies that deployment decisions should account for note length, not just model choice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not test whether higher PDSQI-9 scores predict better downstream clinical decisions or reduced physician workload; that link is the logical next validation step beyond the paper's claims.
  • The moderate Krippendorff's alpha (0.575) alongside high ICC suggests overall reliability is driven partly by attributes with low score variance, so users may need attribute-specific thresholds before treating a summary as safe.
  • Synthesized's weak factor loading and the small number of summaries judged to present abstraction opportunities suggest abstractive quality is the least well-measured construct and may require a dedicated subscale.
  • Because the validity evidence depends on the 118 summaries used in the factor analysis, an external replication should report the full-sample analysis and the sampling rule for that subset.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The manuscript introduces the PDSQI-9, a nine-item instrument for rating the quality of LLM-generated clinical summaries, and reports a validation study using summaries generated by GPT-4o, Mixtral 8x7b, and Llama 3-8b from real UW Health EHR notes, scored by seven physician raters. Validation analyses include Cronbach's alpha, ICC, Krippendorff's alpha, factor analysis, substantive correlations, and discriminant comparisons. The paper concludes that the PDSQI-9 demonstrates robust construct validity and is ready for use in clinical practice.

Significance. If the validity evidence were sound, this would be a useful contribution: the instrument addresses a real gap in evaluation of LLM-generated clinical summaries, uses real-world multi-document EHR data, employs multiple LLMs, incorporates a semi-Delphi content-validity process, and makes the instrument and prompts publicly available. However, as presented, the psychometric evidence is substantially overstated. The reported average-measure ICC does not support single-rater clinical use, the Krippendorff's alpha is moderate and below common thresholds, and the factor analysis rests on an unexplained 118-summary subset. These issues are load-bearing for the central claim, so the manuscript requires major revision before it can support the stated conclusions.

major comments (5)
  1. [Section 4 and Table 3] The abstract and conclusion describe 'high inter-rater reliability (ICC = 0.867)' based on ICC(3,k), an average-measure coefficient across the five evaluators specified in the Table 3 caption. For the stated clinical use, in which a single provider applies the instrument, the relevant coefficient is the single-measure ICC. Using the Spearman-Brown formula, the reported average ICC of 0.867 implies ICC(3,1) of approximately 0.867 / [5 - 4(0.867)] = 0.57, below the 0.75 threshold commonly expected for clinical instruments. Additionally, the reported 95% CI (0.867-0.868) is implausibly narrow for an ICC with 779 observations and five raters, suggesting a calculation or reporting error. This point is load-bearing for the generalizability claim and must be corrected.
  2. [Section 4 and Abstract] Krippendorff's alpha is reported as 0.575 (95% CI: 0.539-0.609), a value the Discussion itself calls 'moderate' and which falls below the commonly accepted 0.667 threshold. The Abstract omits this value entirely, and the conclusion nonetheless calls the evidence 'robust.' Because inter-rater reliability is one of the two pillars of the generalizability claim, the manuscript must report this coefficient prominently and temper its conclusions accordingly.
  3. [Appendix B and Section 4] The factor analysis output in Appendix B reports a harmonic n.obs of 118 and a total n.obs of 118, whereas the main validation corpus consists of 779 summaries and 8,329 item responses. The Results also mention 117 summaries for the abstraction question. The manuscript does not explain how the 118-summary subset was selected or whether it is representative of the full corpus. Without this information, the structural-validity conclusions cannot be generalized to the entire evaluation set.
  4. [Section 3.5 and Appendix B] The Methods section describes 'Confirmatory factor analysis,' but Appendix B is clearly an exploratory factor analysis (minres extraction, varimax rotation, factor selection based on eigenvalues and scree plot). Moreover, the reported fit is not unambiguously strong: RMSEA = 0.05 with a 90% CI of 0 to 0.198, and the loading matrix in Table 6 shows that Synthesized has no positive loading on any factor (its largest loading is -0.347 on MR4). The claim that the four-factor model provides strong support for construct validity is therefore overstated.
  5. [Section 3.5 and Section 4] Discriminant validity is established by comparing summaries generated with GPT-4o and error-free prompts (defined as high quality) against summaries from Llama 3-8b and Mixtral 8x7b with error-prone prompts (defined as low quality). Because the high/low classification is built into the prompt design and the instrument was developed by the same team with the same conceptual framework, this comparison shows that the instrument can distinguish two groups the authors designed to differ, but it does not establish that the instrument detects externally defined quality. An external gold standard or an independent criterion-based validation is needed to support the discriminant-validity claim.
minor comments (5)
  1. [Sections 3.3 and 4] The corpus size is reported inconsistently: Section 3.3 states that the final corpus had 200 summaries (100 GPT-4o, 50 Mixtral, 50 Llama), while Section 4 reports 779 summaries and 8,329 questions. Please clarify whether 779 refers to unique summaries, summary-rater pairs, or some other unit.
  2. [Table 3] The table reports a 'Cronbach's α' value for each individual attribute, and these values are numerically identical to the attribute's ICC. Cronbach's alpha is a scale-level statistic and is not defined for a single item; this column should be removed or replaced with an appropriate item-level statistic (e.g., alpha-if-item-deleted).
  3. [Section 3.5 and Appendix B] The Methods section says 'Confirmatory factor analysis,' but the analysis is exploratory. Please use the correct terminology throughout.
  4. [Appendix B] The table titles are duplicated ('PDSQI-9 Scores Eigenvalues' appears twice), and the second table is actually a factor-loading matrix; the labels should be corrected.
  5. [Section 3.5 and Table 3] The agreement coefficient is spelled both 'Krippendorf' and 'Krippendorff' in different places; please standardize to 'Krippendorff.'

Circularity Check

1 steps flagged · score 4.0 of 10

Discriminant-validity test is a manipulation check: 'low-quality' summaries were built by injecting the exact errors measured by PDSQI-9 items, so one pillar of 'robust construct validity' reduces to the study design; the remaining psychometric evidence is independent but not circular.

  1. self definitional [Section 3.3 (LLM Summarizations) and Section 3.5 (Analysis Plan and Validation)]
    "To generate lower-quality summaries, additional variations of the prompt removed instructions or encouraged the inclusion of false information. ... Anti-Rules introduced intentional errors (e.g., hallucinations, omissions) to include mistakes. ... The summaries generated by GPT-4o with error-free prompts were considered the highest quality, while those generated by Llama 3-8b and Mixtral 8x7b with error-prone prompts were considered the lowest quality."

    The PDSQI-9 defines poor quality by the same defects the Anti-Rules prompts were engineered to insert: 'Accurate' penalizes falsification/fabrication and 'Thorough' penalizes omissions. The 'high-quality' versus 'low-quality' groups are therefore constructed from the instrument's own target dimensions. The significant Mann-Whitney result (p < 0.001) shows that raters noticed the deliberately injected errors; it is a manipulation check, not independent evidence that the instrument measures a latent construct. Presenting this as discriminant validity supporting 'robust construct validity' makes that particular validity claim self-fulfilling rather than externally grounded.

full rationale

The majority of the psychometric work is not circular: Cronbach's alpha, ICC, Krippendorff's alpha, the four-factor model, and the note-length correlations are internal statistics computed on the ratings themselves, and they would stand or fall regardless of how the summaries were produced. The circular element is isolated to discriminant validity, where the known-groups contrast is defined by prompt engineering that injected the very error types the instrument items measure; this does not invalidate the internal-consistency evidence but weakens the conclusion that construct validity is 'robust.' Two non-circular concerns also affect interpretation: the factor analysis is reported on an unexplained subset (harmonic n.obs = 118 in Appendix B versus 779 summaries in the Abstract and Results), and the abstract's 'high inter-rater reliability' relies on average-measure ICC(3,k) = 0.867 across five raters while single-measure reliability and Krippendorff's alpha = 0.575 are moderate; these are correctness and reporting issues rather than circularity. Overall score 4 reflects one central validity sub-claim reducing by construction while the remaining validity evidence retains independent content.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim of robust construct validity rests on standard psychometric assumptions (treating Likert scores as continuous, trusting semi-Delphi consensus for content validity, and assuming the planned sample size remains powered despite skewed distributions). No physical constants or fitted curves are involved. The validation is internal to the study: the instrument is developed and tested on the same authors' data, with no external gold standard.

assumptions (3)
  • domain assumption Likert-scale item scores can be treated as continuous for Pearson correlation and factor analysis.
    The paper computes Pearson correlations and performs factor analysis on 5-point Likert scores. Ordinal data can distort these statistics, but this is standard psychometric practice.
  • domain assumption The semi-Delphi expert consensus establishes content validity.
    Content validity is asserted from the agreement of nine stakeholders, some of whom are authors or employed by an EHR vendor (Epic Systems). No external clinical outcome or independent expert panel is used.
  • domain assumption The sample size calculation assumption of an even score distribution is valid despite observed skew.
    The sample size target assumed an even score distribution across the 5-point scale, but results show skewed distributions (e.g., Comprehensible had no score of 1, Accurate median 5.0). This assumption violation may affect power.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models." pith.science (2026). https://pith.science/paper/2MVE6T7D

@misc{pith2026250108977,
  author       = {Pith},
  title        = {Pith review of: Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2MVE6T7D}},
  note         = {Machine review of arXiv:2501.08977}
}
abstract

As Large Language Models (LLMs) are integrated into electronic health record (EHR) workflows, validated instruments are essential to evaluate their performance before implementation. Existing instruments for provider documentation quality are often unsuitable for the complexities of LLM-generated text and lack validation on real-world data. The Provider Documentation Summarization Quality Instrument (PDSQI-9) was developed to evaluate LLM-generated clinical summaries. Multi-document summaries were generated from real-world EHR data across multiple specialties using several LLMs (GPT-4o, Mixtral 8x7b, and Llama 3-8b). Validation included Pearson correlation for substantive validity, factor analysis and Cronbach's alpha for structural validity, inter-rater reliability (ICC and Krippendorff's alpha) for generalizability, a semi-Delphi process for content validity, and comparisons of high-versus low-quality summaries for discriminant validity. Seven physician raters evaluated 779 summaries and answered 8,329 questions, achieving over 80% power for inter-rater reliability. The PDSQI-9 demonstrated strong internal consistency (Cronbach's alpha = 0.879; 95% CI: 0.867-0.891) and high inter-rater reliability (ICC = 0.867; 95% CI: 0.867-0.868), supporting structural validity and generalizability. Factor analysis identified a 4-factor model explaining 58% of the variance, representing organization, clarity, accuracy, and utility. Substantive validity was supported by correlations between note length and scores for Succinct (rho = -0.200, p = 0.029) and Organized ($\rho = -0.190$, $p = 0.037$). Discriminant validity distinguished high- from low-quality summaries ($p < 0.001$). The PDSQI-9 demonstrates robust construct validity, supporting its use in clinical practice to evaluate LLM-generated summaries and facilitate safer integration of LLMs into healthcare workflows.

Figures

Figures reproduced from arXiv: 2501.08977 by the authors.

Figure 1
Figure 1. illustrates the average scores for each attribute of a summary, as evaluated by our raters, in relation to the length of the notes being summarized. As the length of the notes increased, the quality of the generated summaries was rated lower in the attributes of Organized (𝜌 = -0.190 with p-value = 0.037), Succinct (𝜌 = -0.200 with p-value = 0.029), and Thorough (𝜌 = -0.31 with p-value < 0.001). Additionally, the va… view at source ↗
Figure 2
Figure 2. Length of Patient Notes vs. Standard Deviation of Evaluator Scores Scatter plots of the score standard deviation among evaluators compared to the token length of the patient notes provided for summarization. Trend lines are superimposed on the plot along with the Spearman 𝜌 (denoted as R) coefficient and its p-value (denoted as p). Each plot corresponds to a single attribute from the PDSQI-9 instrument [PITH_FULL_I… view at source ↗
Figure 3
Figure 3. Likert Score Distributions by Attribute A density ridge across the 5 Likert scale points that evaluators used to score each attribute from the PDSQI-9 instrument. Each attribute is provided separately and identified on the y-axis. Distributions include the scores from every evaluator on every unique patient summary reported in this study. Attribute ICC Krippendorf’s 𝛼 Cronbach’s 𝛼 Accurate 0.791 0.394 0.791 (95% CI)… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Time Spent per Evaluation by Evaluator Experience A boxplot describing the amount of time it took an evaluator to complete each evaluation in full. We have provided the information for the junior and senior physicians separately for comparison. The time is represented …
Figure 5
Figure 5. Figure 5: PDSQI-9 Attribute Correlation Plot A visual representation of the correlations between the scores for each attribute in the PDSQI-9. The strength of the correlation is represented by the color and size of each circle [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Scree Plot for PDSQI-9 Evaluative Scores A scree plot to visualize the eigenvalues related with the number of components in a principal component analysis (PC) or factors in a factor analysis (FA). 16 [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 47 canonical work pages

  1. [1]

    Call me Dr Ishmael: trends in electronic health record notes available at emergency department visits and admissions

    Patterson BW, Hekman DJ, Liao FJ, Hamedani AG, Shah MN, Afshar M. Call me Dr Ishmael: trends in electronic health record notes available at emergency department visits and admissions. JAMIA Open. 2024 Apr;7(2):ooae039

  2. [2]

    To Err is Human: Building a Safer Health System

    Institute of Medicine (US) Committee on Quality of Health Care in America. To Err is Human: Building a Safer Health System. Kohn LT, Corrigan JM, Donaldson MS, editors. Washington (DC): National Academies Press (US); 2000. Available from: http://www.ncbi.nlm.nih.gov/books/ NBK225182/

  3. [3]

    Computerized provider documentation: findings and implications of a multisite study of clinicians and administrators

    Embi PJ, Weir C, Efthimiadis EN, Thielke SM, Hedeen AN, Hammond KW. Computerized provider documentation: findings and implications of a multisite study of clinicians and administrators. Journal of the American Medical Informatics Association: JAMIA. 2013 Jan;20(4):718

  4. [4]

    Lost in the Middle: How Language Models Use Long Contexts

    Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024;12:157–173

  5. [5]

    Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review

    Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2024 Oct:e2421700

  6. [6]

    A framework for human evaluation of large language models in healthcare derived from literature review

    Tam TYC, Sivarajkumar S, Kapoor S, Stolyar A V, Polanska K, McCarthy KR, et al. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digital Medicine. 2024 Sep;7(1):258

  7. [7]

    Using ChatGPT-4 to Create Structured Medical Notes From Audio Recordings of Physician-Patient Encounters: Comparative Study

    Kernberg A, Gold JA, Mohan V. Using ChatGPT-4 to Create Structured Medical Notes From Audio Recordings of Physician-Patient Encounters: Comparative Study. Journal of Medical Internet Research. 2024 Apr;26:e54419. 11

  8. [8]

    Effect of Ambient Voice Technology, Natural Language Processing, and Artificial Intelligence on the Patient-Physician Relationship

    Owens LM, Wilda JJ, Grifka R, Westendorp J, Fletcher JJ. Effect of Ambient Voice Technology, Natural Language Processing, and Artificial Intelligence on the Patient-Physician Relationship. Applied Clinical Informatics. 2024 Aug;15(4):660–667

Show all 57 references
  1. [9]

    Ambient Artifi- cial Intelligence Scribes to Alleviate the Burden of Clinical Documentation

    Tierney AA, Gayre G, Hoberman B, Mattern B, Ballesca M, Kipnis P, et al. Ambient Artifi- cial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst. 2024 Feb;5(3):CAT.23.0404

  2. [10]

    Assessing Electronic Note Quality Using the Physician Documentation Quality Instrument (PDQI-9)

    Stetson PD, Bakken S, Wrenn JO, Siegler EL. Assessing Electronic Note Quality Using the Physician Documentation Quality Instrument (PDQI-9). Applied Clinical Informatics. 2012 Apr;3(2):164

  3. [11]

    A Survey of Large Language Models

    Zhao WX, Zhou K, Li J, Tang T, Wang X, Hou Y, et al. A Survey of Large Language Models. 2023 Jun;(arXiv:2303.18223). ArXiv:2303.18223 [cs]. Available from: http://arxiv.org/abs/2303. 18223

  4. [12]

    The Delphi Method: Techniques and Applications

    Turoff M, Linstone HA. The Delphi Method: Techniques and Applications

  5. [13]

    A Survey of Evaluation Metrics Used for NLG Systems

    Sai AB, Mohankumar AK, Khapra MM. A Survey of Evaluation Metrics Used for NLG Systems. ACM Computing Surveys. 2023;55(2)

  6. [14]

    Generation of Patient After-Visit Summaries to Support Physicians

    Cai P, Liu F, Bajracharya A, Sills J, Kapoor A, Liu W, et al. Generation of Patient After-Visit Summaries to Support Physicians. In: Proceedings of the 29th International Conference on Computational Lin- guistics. Gyeongju, Republic of Korea: International Committee on Computa...

  7. [15]

    A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital- Course Summarization

    Adams G, Zucker J, Elhadad N. A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital- Course Summarization. 2023 Mar;(arXiv:2303.03948). ArXiv:2303.03948 [cs]. Available from:http: //arxiv.org/abs/2303.03948

  8. [16]

    Large language models encode clinical knowledge

    Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023 Jul:1–9

  9. [17]

    Med-HALT: Medical Domain Hallucination Test for Large Language Models

    Umapathi LK, Pal A, Sankarasubbu M. Med-HALT: Medical Domain Hallucination Test for Large Language Models. 2023 Jul;(arXiv:2307.15343). ArXiv:2307.15343 [cs, stat]. Available from: http: //arxiv.org/abs/2307.15343

  10. [18]

    Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization; 2020

    Wallace BC, Saha S, Soboczenski F, Marshall IJ. Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization; 2020. Available from: https: //arxiv.org/abs/2008.11293v2

  11. [19]

    The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature

    Otmakhova Y, Verspoor K, Baldwin T, Lau JH. The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1:...

  12. [20]

    Revisiting Summarization Evaluation for Scientific Articles

    Cohan A, Goharian N. Revisiting Summarization Evaluation for Scientific Articles

  13. [21]

    Development of a Human Evaluation Framework and Correlation with Automated Metrics for Natural Language Generation of Medical Di- agnoses

    Croxford E, Gao Y, Patterson B, To D, Tesch S, Dligach D, et al. Development of a Human Evaluation Framework and Correlation with Automated Metrics for Natural Language Generation of Medical Di- agnoses. 2024 Apr:2024.03.20.24304620. Available from: https://www.medrxiv.org/con...

  14. [22]

    Reinforcement Learning for Abstractive Question Summarization with Question-aware Semantic Rewards

    Yadav S, Gupta D, Abacha AB, Demner-Fushman D. Reinforcement Learning for Abstractive Question Summarization with Question-aware Semantic Rewards. 2021 Jun;(arXiv:2107.00176). ArXiv:2107.00176 [cs]. Available from: http://arxiv.org/abs/2107.00176

  15. [23]

    Automated Lay Language Summarization of Biomedical Scientific Reviews

    Guo Y, Qiu W, Wang Y, Cohen T. Automated Lay Language Summarization of Biomedical Scientific Reviews. 2022 Jan;(arXiv:2012.12573). ArXiv:2012.12573 [cs]. Available from: http://arxiv. org/abs/2012.12573

  16. [24]

    An Investigation of Evaluation Metrics for Automated Medical Note Generation

    Abacha AB, Yim Ww, Michalopoulos G, Lin T. An Investigation of Evaluation Metrics for Automated Medical Note Generation. 2023 May;(arXiv:2305.17364). ArXiv:2305.17364 [cs]. Available from: http://arxiv.org/abs/2305.17364

  17. [25]

    Research Electronic Data Capture (REDCap) - A metadata-driven methodology and workflow process for providing translational research informatics support

    Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research Electronic Data Capture (REDCap) - A metadata-driven methodology and workflow process for providing translational research informatics support. Journal of biomedical informatics. 2008 Sep;42(2):377

  18. [26]

    The REDCap consortium: Building an international community of software platform partners

    Harris PA, Taylor R, Minor BL, Elliott V, Fernandez M, O’Neal L, et al. The REDCap consortium: Building an international community of software platform partners. Journal of Biomedical Informatics. 2019 Jul;95:103208

  19. [27]

    GPT-4 Technical Report

    OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, et al. GPT-4 Technical Report. 2024 Mar;(arXiv:2303.08774). ArXiv:2303.08774. Available from: http://arxiv.org/abs/2303. 08774

  20. [28]

    Mixtral of Experts

    Jiang AQ, Sablayrolles A, Roux A, Mensch A, Savary B, Bamford C, et al. Mixtral of Experts. 2024 Jan;(arXiv:2401.04088). ArXiv:2401.04088. Available from: http://arxiv.org/abs/2401. 04088

  21. [29]

    The Llama 3 Herd of Models

    Grattafiori A, Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, et al. The Llama 3 Herd of Models. 2024 Nov;(arXiv:2407.21783). ArXiv:2407.21783. Available from:http://arxiv.org/abs/2407. 21783

  22. [30]

    Available from: https://huggingface.co/

    2024. Available from: https://huggingface.co/

  23. [31]

    kappaSize: Sample Size Estimation Functions for Studies of Interobserver Agreement

    Rotondi MA. kappaSize: Sample Size Estimation Functions for Studies of Interobserver Agreement

  24. [32]

    Words Matter: Strategies to Reduce Bias in Electronic Health Records

    Canonico M. Words Matter: Strategies to Reduce Bias in Electronic Health Records. Words Matter

  25. [33]

    Standards of Validity and the Validity of Standards in Performance Asessment

    Messick S. Standards of Validity and the Validity of Standards in Performance Asessment. Educational Measurement: Issues and Practice. 1995. Available from: https://onlinelibrary.wiley.com/ doi/10.1111/j.1745-3992.1995.tb00881.x

  26. [34]

    Content Analysis: An Introduction to Its Methodology

    Krippendorff K. Content Analysis: An Introduction to Its Methodology. SAGE Publications; 2018. Google-Books-ID: nE1aDwAAQBAJ

  27. [35]

    In: Kotz S, Johnson NL, editors

    Fisher RA. In: Kotz S, Johnson NL, editors. Statistical Methods for Research Workers. New York, NY: Springer; 1992. p. 66–70. Available from: https://doi.org/10.1007/978-1-4612-4380-9_6

  28. [36]

    A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research

    Koo TK, Li MY. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine. 2016 Mar;15(2):155. 13

  29. [37]

    Coefficient alpha and the internal structure of tests

    Cronbach LJ. Coefficient alpha and the internal structure of tests. Psychometrika. 1951 Sep;16(3):297–334

  30. [38]

    Statistical Inference for Coefficient Alpha

    Feldt LS, Woodruff DJ, Salih FA. Statistical Inference for Coefficient Alpha. Applied Psychological Measurement. 1987 Mar;11(1):93–103

  31. [39]

    Intraclass Correlations: Uses in Assessing Rater Reliability

    Shrout PE, Fleiss JL. Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological Bulletin. 1979

  32. [40]

    Transformers: State-of-the-Art Natural Language Processing

    Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. Transformers: State-of-the-Art Natural Language Processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Online: Association for Computational L...

  33. [41]

    Natural Language Processing with Python

    Bird EL Steven, Klein E. Natural Language Processing with Python. O’Reilly Media Inc. 2009

  34. [42]

    smplot: an R package for easy and elegant data visualization

    Min SH, Zhou J. smplot: an R package for easy and elegant data visualization. Frontiers in Genetics. 2021;12:802894

  35. [43]

    ggplot2: Elegant Graphics for Data Analysis

    Wickham H. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York; 2016. Available from: https://ggplot2.tidyverse.org

  36. [44]

    psych: Procedures for Psychological, Psychometric, and Personality Research

    William Revelle. psych: Procedures for Psychological, Psychometric, and Personality Research. Evanston, Illinois; 2024. R package version 2.4.12. Available from: https://CRAN.R-project. org/package=psych

  37. [45]

    krippendorffsalpha: An R Package for Measuring Agreement Using Krippendorff’s Alpha Coefficient

    Hughes J. krippendorffsalpha: An R Package for Measuring Agreement Using Krippendorff’s Alpha Coefficient. 2021 Mar;(arXiv:2103.12170). ArXiv:2103.12170. Available from:http://arxiv.org/ abs/2103.12170

  38. [46]

    DescTools: Tools for Descriptive Statistics; 2017

    et mult al AS. DescTools: Tools for Descriptive Statistics; 2017. R package version 0.99.23. Available from: https://cran.r-project.org/package=DescTools

  39. [47]

    Toward Clinical Generative AI: Conceptual Framework

    Bragazzi NL, Garbarino S. Toward Clinical Generative AI: Conceptual Framework. JMIR AI. 2024 Jun;3:e55957

  40. [48]

    Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review

    Croxford E, Gao Y, Pellegrino N, Wong KK, Wills G, First E, et al. Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review. 2024 Sep;(arXiv:2409.18170). ArXiv:2409.18170. Available from: http://arxiv.org/abs/2409.18170

  41. [49]

    Physician Use of Stigmatizing Language in Patient Medical Records

    Park J, Saha S, Chee B, Taylor J, Beach MC. Physician Use of Stigmatizing Language in Patient Medical Records. JAMA network open. 2021 Jul;4(7):e2117052

  42. [50]

    the pa6ent experienced shortness of breath, had an elevated white blood cell count, and showed a right lower lobe infiltrate on a chest X-ray,

    Klang E, Apakama D, Abbott EE, Vaid A, Lampert J, Sakhuja A, et al. A strategy for cost-effective large language model use at health system-scale. NPJ digital medicine. 2024 nov;7(1):320. 14 6 Appendices 6.1 Appendix A: Time Spent per Evaluation by Evaluator Experience Figure ...

  43. [53]

    Is the summary thorough without any omissions? Iden6fy any per6nent or poten6ally per6nent omissions: a. Per/nent omissions refer to essen6al informa6on required for the specific use case or intended provider, where missing details could directly impact pa6ent care decisions (i...

  44. [54]

    All the informa6on is in there that is useful to the target provider/intended audience b

    Is the summary useful? a. All the informa6on is in there that is useful to the target provider/intended audience b. The summary is extremely relevant, providing valuable informa6on and/or analysis 1: Not at All 2 3 4 5: Extremely No asser$ons are per$nent to the target user So...

  45. [55]

    The summary is well-formed and structured in a way that helps the reader understand the pa6ent’s clinical course

    Is the summary organized? a. The summary is well-formed and structured in a way that helps the reader understand the pa6ent’s clinical course. 1: Not at All 2 3 4 5: Extremely All Asser$ons presented out of order and groupings incoherent (completely disorganized) Some asser$on...

  46. [56]

    the pa6ent had an HbA1c of 9.2%, was prescribed mekormin, and reported not following dietary recommenda6ons

    Is the summary succinct with economy of language? a. The summary is brief, to the point, and without redundancy 1: Not at All 2 3 4 5: Extremely Too wordy across all asser$ons with redundancy in syntax and seman$c More than one asser$on has contextual seman$c redundancy At lea...

  47. [57]

    The pa6ent's mother said that the ‘pills don't work’ for her child

    Is there presence of s+gma+zing language? a. Please follow the tool in this link: h^ps://www.chcs.org/media/Words-Ma^er-Strategies-to-Reduce-Bias-in-Electronic-Health-Records_102022.pdf b. Refrain from using discredi6ng or exaggerated words, such as claims, insists, or reporte...

  48. [2018]

    Available from: https://cran.r-project.org/web/packages/kappaSize/index.html

  49. [2020]

    p. 38-45. Available from: https://www.aclweb.org/anthology/2020.emnlp-demos.6

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.