{"id":"d3c91e8a-8723-4188-88b3-9e037e22dddf","arxiv_id":"2501.08977","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PDSQI-9 is a 9-item instrument for rating LLM-generated clinical summaries; in a validation study with 7 physician raters and 779 summary evaluations, it showed Cronbach's alpha 0.879 and ICC 0.867, but Krippendorff's alpha was only 0.575.","lead":"Researchers developed and tested a nine-item scoring tool (PDSQI-9) for rating how well AI-generated summaries of patient charts capture accurate, organized, and useful information. In a study of 779 physician ratings of AI summaries from real electronic health records, the tool showed good internal consistency and inter-rater agreement, though the validation evidence is weaker than the conclusions claim.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Average-measure ICC (0.867) implies single-measure ICC ≈0.57, and Krippendorff's α=0.575 is moderate; 'high inter-rater reliability' and 'robust construct validity' are not supported for real-world single-rater use.","rationale":"I focused on the reliability evidence rather than the factor-analysis subset because the reported numbers are internally sufficient to test the claim, and the issue directly affects the intended use case. The abstract's 'robust construct validity' sentence hinges on 'high inter-rater reliability (ICC = 0.867).' The methods explicitly state ICC(3,k), and Table 3 says the coefficients are 'across our five evaluators.' Converting the average ICC to a single-rater value via Spearman-Brown gives approximately 0.57, which is below conventional thresholds for clinical measures. Additionally, Krippendorff's α=0.575 is reported in the results and is moderate at best; the discussion simultaneously calls it 'moderate' and 'robust.' These two observations are concrete, quantitative, and directly contradict the strength of the central claim. The study has genuine strengths: a priori power calculations, real-world EHR summaries, multiple LLMs, seven physician raters, semi-Delphi content development, and a publicly available instrument. The factor-analysis sample-size discrepancy (n=118 vs 779) and the engineered discriminant-validity manipulation are additional concerns that deserve transparent reporting, but they are secondary to the reliability gap. The verdict remains CONDITIONAL because the concern is addressable by reporting single-measure reliability and tempering the clinical-deployment conclusion.","tokens_in":15541,"tokens_out":10978,"duration_ms":109607,"concrete_test":"Request the raw rating matrix and recompute ICC(3,1) with 95% confidence intervals. Independently apply the Spearman-Brown formula to the reported average ICC(3,5)=0.867: ICC(3,1)=0.867/(5−4×0.867)=0.57. If the lower bound of the single-measure ICC 95% CI is below 0.75, or if Krippendorff's α remains below 0.667, the abstract's 'high inter-rater reliability' and 'supporting its use in clinical practice' claims are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the instrument being reliable in practical clinical use, where a single provider would apply it. The paper reports ICC(3,k), an average-measure coefficient across multiple raters; Table 3's caption specifies five evaluators. Using the Spearman-Brown formula, the reported average ICC of 0.867 corresponds to a single-measure ICC of 0.867/(5 − 4×0.867) ≈ 0.57, below the 0.75 threshold commonly expected for clinical instruments. The paper also reports Krippendorff's α = 0.575 (95% CI 0.539–0.609), which is below the 0.667 acceptable threshold; the discussion itself calls this 'moderate' yet the abstract omits it and the conclusion calls the evidence 'robust.' Even if the factor-analysis subset issue (n=118 vs 779) were resolved, the reliability evidence as reported does not justify the conclusion that the PDSQI-9 supports use in clinical practice.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces the PDSQI-9, a nine-item instrument for rating the quality of LLM-generated clinical summaries, and reports a validation study using summaries generated by GPT-4o, Mixtral 8x7b, and Llama 3-8b from real UW Health EHR notes, scored by seven physician raters. Validation analyses include Cronbach's alpha, ICC, Krippendorff's alpha, factor analysis, substantive correlations, and discriminant comparisons. The paper concludes that the PDSQI-9 demonstrates robust construct validity and is ready for use in clinical practice.","tokens_in":15659,"tokens_out":7254,"duration_ms":75226,"significance":"If the validity evidence were sound, this would be a useful contribution: the instrument addresses a real gap in evaluation of LLM-generated clinical summaries, uses real-world multi-document EHR data, employs multiple LLMs, incorporates a semi-Delphi content-validity process, and makes the instrument and prompts publicly available. However, as presented, the psychometric evidence is substantially overstated. The reported average-measure ICC does not support single-rater clinical use, the Krippendorff's alpha is moderate and below common thresholds, and the factor analysis rests on an unexplained 118-summary subset. These issues are load-bearing for the central claim, so the manuscript requires major revision before it can support the stated conclusions.","major_comments":[{"comment":"The abstract and conclusion describe 'high inter-rater reliability (ICC = 0.867)' based on ICC(3,k), an average-measure coefficient across the five evaluators specified in the Table 3 caption. For the stated clinical use, in which a single provider applies the instrument, the relevant coefficient is the single-measure ICC. Using the Spearman-Brown formula, the reported average ICC of 0.867 implies ICC(3,1) of approximately 0.867 / [5 - 4(0.867)] = 0.57, below the 0.75 threshold commonly expected for clinical instruments. Additionally, the reported 95% CI (0.867-0.868) is implausibly narrow for an ICC with 779 observations and five raters, suggesting a calculation or reporting error. This point is load-bearing for the generalizability claim and must be corrected.","section":"Section 4 and Table 3"},{"comment":"Krippendorff's alpha is reported as 0.575 (95% CI: 0.539-0.609), a value the Discussion itself calls 'moderate' and which falls below the commonly accepted 0.667 threshold. The Abstract omits this value entirely, and the conclusion nonetheless calls the evidence 'robust.' Because inter-rater reliability is one of the two pillars of the generalizability claim, the manuscript must report this coefficient prominently and temper its conclusions accordingly.","section":"Section 4 and Abstract"},{"comment":"The factor analysis output in Appendix B reports a harmonic n.obs of 118 and a total n.obs of 118, whereas the main validation corpus consists of 779 summaries and 8,329 item responses. The Results also mention 117 summaries for the abstraction question. The manuscript does not explain how the 118-summary subset was selected or whether it is representative of the full corpus. Without this information, the structural-validity conclusions cannot be generalized to the entire evaluation set.","section":"Appendix B and Section 4"},{"comment":"The Methods section describes 'Confirmatory factor analysis,' but Appendix B is clearly an exploratory factor analysis (minres extraction, varimax rotation, factor selection based on eigenvalues and scree plot). Moreover, the reported fit is not unambiguously strong: RMSEA = 0.05 with a 90% CI of 0 to 0.198, and the loading matrix in Table 6 shows that Synthesized has no positive loading on any factor (its largest loading is -0.347 on MR4). The claim that the four-factor model provides strong support for construct validity is therefore overstated.","section":"Section 3.5 and Appendix B"},{"comment":"Discriminant validity is established by comparing summaries generated with GPT-4o and error-free prompts (defined as high quality) against summaries from Llama 3-8b and Mixtral 8x7b with error-prone prompts (defined as low quality). Because the high/low classification is built into the prompt design and the instrument was developed by the same team with the same conceptual framework, this comparison shows that the instrument can distinguish two groups the authors designed to differ, but it does not establish that the instrument detects externally defined quality. An external gold standard or an independent criterion-based validation is needed to support the discriminant-validity claim.","section":"Section 3.5 and Section 4"}],"minor_comments":[{"comment":"The corpus size is reported inconsistently: Section 3.3 states that the final corpus had 200 summaries (100 GPT-4o, 50 Mixtral, 50 Llama), while Section 4 reports 779 summaries and 8,329 questions. Please clarify whether 779 refers to unique summaries, summary-rater pairs, or some other unit.","section":"Sections 3.3 and 4"},{"comment":"The table reports a 'Cronbach's α' value for each individual attribute, and these values are numerically identical to the attribute's ICC. Cronbach's alpha is a scale-level statistic and is not defined for a single item; this column should be removed or replaced with an appropriate item-level statistic (e.g., alpha-if-item-deleted).","section":"Table 3"},{"comment":"The Methods section says 'Confirmatory factor analysis,' but the analysis is exploratory. Please use the correct terminology throughout.","section":"Section 3.5 and Appendix B"},{"comment":"The table titles are duplicated ('PDSQI-9 Scores Eigenvalues' appears twice), and the second table is actually a factor-loading matrix; the labels should be corrected.","section":"Appendix B"},{"comment":"The agreement coefficient is spelled both 'Krippendorf' and 'Krippendorff' in different places; please standardize to 'Krippendorff.'","section":"Section 3.5 and Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a potentially useful instrument and a substantial real-world evaluation dataset, but the headline reliability and validity claims are not supported by the statistics as currently reported. I would ask the authors to re-run or provide the underlying analyses to (1) correct the ICC confidence interval, (2) report single-measure ICC and interpret it in light of intended use, (3) explain the 118-summary factor-analysis subset, and (4) revise the conclusions to match the actual magnitude of the reliability estimates. If, after reanalysis, the single-measure ICC and Krippendorff's alpha remain at the levels currently reported, the 'robust' conclusion should be substantially softened or the instrument's intended use should be restricted to settings with multiple raters."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the PDSQI-9 is a sensible adaptation of PDQI-9, and the dataset is real and fairly large. But the paper's headline claim—'robust construct validity'—is not supported by the numbers as reported. The ICC CI of 0.867–0.868 is too narrow to be real, and more importantly the paper uses ICC(3,k), an average-measure coefficient over five raters. Back out the single-measure value and you get roughly 0.57, below the usual 0.75 threshold for clinical instruments. Krippendorff's alpha is 0.575, which the discussion itself calls moderate; the abstract omits it and the conclusion calls it robust. That matters because in practice a single provider would apply the tool. The stress-test note lands: the reliability evidence doesn't support single-rater clinical use.\n\nWhat is genuinely new: two attributes added to PDQI-9 (citations, stigmatizing language) that target known LLM failure modes, validation on real multi-document EHR summaries across multiple specialties and three LLMs, an a priori power calculation, and a semi-Delphi item development process. The appendix gives the full rubric, and the instrument and prompts are posted. That is reproducible groundwork.\n\nSoft spots beyond the ICC issue: the factor analysis output reports a harmonic n.obs of 118, while the study says 779 summaries were evaluated; the paper never explains the subset. The abstraction analysis uses 117 summaries. The discriminant validity test compares summaries the authors intentionally engineered to be bad against ones engineered to be good; that is a manipulation check, not evidence the instrument can distinguish unaided quality differences. There is no external gold standard, so criterion validity is unaddressed. None of these are fatal—they are fixable—but together they mean the conclusion should be softened from 'supports use in clinical practice' to 'shows promise pending external validation and corrected reliability estimates.'\n\nWho gets value: clinical NLP and health informatics readers designing evaluation instruments, and anyone who needs a concrete example of how ICC average vs single-measure reporting can change a conclusion. I'd send it to peer review, but with a major-revision request.","headline":"A useful PDQI-9 adaptation with real data, but the reliability evidence as reported doesn't support 'robust construct validity'—the average-measure ICC hides a single-measure value around 0.57.","tokens_in":16344,"tokens_out":2553,"would_cite":false,"duration_ms":26271,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper validates PDSQI-9, a nine-item instrument for scoring LLM-generated clinical summaries, using 779 real-world EHR summaries.","keywords":["PDSQI-9","large language models","clinical summarization","electronic health records","psychometric validation","construct validity","inter-rater reliability","human evaluation"],"falsifier":"Re-run the factor analysis and reliability statistics on all 779 evaluations, or on a documented random sample, and check whether the four-factor structure and the alpha/ICC values hold; also compare the 118 factor-analysis summaries with the remaining 661 on model type, specialty, and note length to test for selection bias.","tokens_in":15272,"feed_emoji":"🩺","tokens_out":5867,"duration_ms":55896,"temperature":0.7,"pith_summary":"This paper introduces and validates the Provider Documentation Summarization Quality Instrument (PDSQI-9), a nine-item rubric for scoring how well large language models summarize electronic health record notes. Seven physician raters applied the instrument to 779 summaries generated from real-world EHR data by GPT-4o, Mixtral 8x7b, and Llama 3-8b, producing 8,329 item scores. The paper reports high internal consistency (Cronbach's $\\alpha = 0.879$), high inter-rater reliability (ICC $= 0.867$), and a four-factor structure corresponding to organization, clarity, accuracy, and utility. The authors argue this constitutes strong construct validity and that the instrument can support safer integration of LLM summarization into clinical workflows.","feed_headline":"Nine-item scale for AI clinical summaries passes validity testing","feed_subtitle":"Seven physicians scored 779 AI summaries; high reliability and four quality factors support clinical use.","key_machinery":"The central object is the PDSQI-9 rubric, a nine-attribute adaptation of the Physician Documentation Quality Instrument that targets LLM-specific failure modes: Cited, Accurate, Thorough, Useful, Organized, Comprehensible, Succinct, Synthesized, and Stigmatizing. The instrument carries the argument by converting qualitative judgments about summary quality into quantifiable scores, which are then tested under Messick's validity framework through factor analysis, reliability coefficients, and group comparisons.","core_discovery":"The central claim is that the PDSQI-9 is a valid and reliable instrument for evaluating LLM-generated summaries of clinical documentation. The paper presents evidence from content validity established through a semi-Delphi process, substantive validity shown by expected correlations with note length, structural validity supported by Cronbach's $\\alpha = 0.879$ and a four-factor model explaining 58% of variance, generalizability supported by ICC $= 0.867$, and discriminant validity distinguishing high-quality from low-quality summaries at $p < 0.001$. The authors conclude that the instrument demonstrates sound construct validity and is ready for use in clinical practice to evaluate LLM-generated summaries before integration into healthcare workflows.","pith_inferences":["The authors do not test whether higher PDSQI-9 scores predict better downstream clinical decisions or reduced physician workload; that link is the logical next validation step beyond the paper's claims.","The moderate Krippendorff's alpha (0.575) alongside high ICC suggests overall reliability is driven partly by attributes with low score variance, so users may need attribute-specific thresholds before treating a summary as safe.","Synthesized's weak factor loading and the small number of summaries judged to present abstraction opportunities suggest abstractive quality is the least well-measured construct and may require a dedicated subscale.","Because the validity evidence depends on the 118 summaries used in the factor analysis, an external replication should report the full-sample analysis and the sampling rule for that subset."],"forward_implications":["Hospitals and health systems can use PDSQI-9 as a standardized human-evaluation step before deploying an LLM summarization tool on real patient notes.","Benchmark comparisons of different LLMs or prompt strategies can be reported as nine attribute scores plus a four-factor profile rather than a single overall number.","Attribute-level reliability estimates identify where scoring is easiest (Cited, Succinct) and where it is least stable (Comprehensible, Synthesized), guiding targeted improvements to prompts or retrieval.","The finding that longer inputs are associated with lower Organized, Succinct, and Thorough scores implies that deployment decisions should account for note length, not just model choice."],"supporting_citations":[{"why":"The Physician Documentation Quality Instrument (PDQI-9) that PDSQI-9 adapts; supplies the baseline attribute set and the reliability comparison point.","marker":"[10]"},{"why":"Messick's framework of validity organizes the substantive, structural, generalizability, content, and discriminant validity analyses.","marker":"[33]"},{"why":"Sample-size estimation for inter-observer agreement used to set the number of evaluations needed for 80% power.","marker":"[31]"},{"why":"Defines the ICC(3,k) two-way mixed-effects model used to compute inter-rater reliability.","marker":"[36]"},{"why":"Defines coefficient alpha, the internal-consistency statistic reported for the instrument and each attribute.","marker":"[37]"},{"why":"Defines Krippendorff's alpha, the chance-corrected agreement measure reported alongside ICC.","marker":"[34]"},{"why":"Shrout and Fleiss procedure used to construct confidence intervals for the ICC estimates.","marker":"[39]"},{"why":"Systematic review of human evaluation of LLMs in healthcare that motivates the sample-size, rater-number, and training standards the study follows.","marker":"[6]"}],"fun_headline_variants":["New 9-item scale validates AI clinical summaries","PDSQI-9: reliable tool for scoring AI summaries","Physician-tested scale evaluates AI clinical notes","Validated scale for LLM-made clinical summaries","Nine-item tool assesses AI summary quality with high reliability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The validity evidence depends on the 118 summaries used in the factor analysis, while the study reports 779 evaluations; the paper never states how that subset was selected, so if it is not representative of the full corpus the psychometric claims do not generalize.","fun_headline_variants_meta":{"raw":{"variants":["New 9-item scale validates AI clinical summaries","PDSQI-9: reliable tool for scoring AI summaries","Physician-tested scale evaluates AI clinical notes","Validated scale for LLM-made clinical summaries","Nine-item tool assesses AI summary quality with high reliability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000702,"raw_usage":{"total_tokens":3224,"prompt_tokens":1060,"completion_tokens":2164,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":2089}},"tokens_in":676,"tokens_out":2164,"duration_ms":14424,"temperature":1.0,"reasoning_tokens":2089,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:13:11.218512+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the factor analysis and reliability statistics on all 779 evaluations, or on a documented random sample, and check whether the four-factor structure and the alpha/ICC values hold; also compare the 118 factor-analysis summaries with the remaining 661 on model type, specialty, and note length to test for selection bias.","supporting_citations":[{"cited_title":"Assessing Electronic Note Quality Using the Physician Documentation Quality Instrument (PDQI-9)","cited_arxiv_id":null,"evidence_quote":"The Physician Documentation Quality Instrument (PDQI-9) that PDSQI-9 adapts; supplies the baseline attribute set and the reliability comparison point."},{"cited_title":"Standards of Validity and the Validity of Standards in Performance Asessment","cited_arxiv_id":null,"evidence_quote":"Messick's framework of validity organizes the substantive, structural, generalizability, content, and discriminant validity analyses."},{"cited_title":"kappaSize: Sample Size Estimation Functions for Studies of Interobserver Agreement","cited_arxiv_id":null,"evidence_quote":"Sample-size estimation for inter-observer agreement used to set the number of evaluations needed for 80% power."},{"cited_title":"A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research","cited_arxiv_id":null,"evidence_quote":"Defines the ICC(3,k) two-way mixed-effects model used to compute inter-rater reliability."},{"cited_title":"Coefficient alpha and the internal structure of tests","cited_arxiv_id":null,"evidence_quote":"Defines coefficient alpha, the internal-consistency statistic reported for the instrument and each attribute."},{"cited_title":"Content Analysis: An Introduction to Its Methodology","cited_arxiv_id":null,"evidence_quote":"Defines Krippendorff's alpha, the chance-corrected agreement measure reported alongside ICC."},{"cited_title":"Intraclass Correlations: Uses in Assessing Rater Reliability","cited_arxiv_id":null,"evidence_quote":"Shrout and Fleiss procedure used to construct confidence intervals for the ICC estimates."},{"cited_title":"A framework for human evaluation of large language models in healthcare derived from literature review","cited_arxiv_id":null,"evidence_quote":"Systematic review of human evaluation of LLMs in healthcare that motivates the sample-size, rater-number, and training standards the study follows."}],"review_version":1}