{"id":"66bb2bbe-d48c-4c0a-aa01-29badf655a1a","arxiv_id":"2505.10409","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Human-written plain language summaries led to significantly better reader comprehension than LLM-generated ones, despite similar subjective ratings, and most automated metrics did not predict comprehension.","lead":"This study tested whether plain language summaries written by large language models are actually understood by readers, using 150 crowd workers who rated and answered questions about the summaries. It found that readers scored higher on comprehension questions after reading human-written summaries, even though they rated LLM-written summaries just as highly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The main human-vs-LLM MCQ comparison depends on an undisclosed input source for the comprehension questions; if LLaMA-3 was prompted with the human-written PLS, the headline result is a measurement artifact.","rationale":"The reader's weakest assumption identifies the same load-bearing point: Section 4.2 never states the source text for MCQ generation. I read the paper in good faith and the central claim is plausible, but this omission is decisive because the MCQ instrument is the sole objective support for the claim. If the questions came from the human-written PLS, the comparison is circular in a practical sense: the test key would be derived from the very text that is claimed to be easier to understand. The paper does provide some independent support for robustness—the rigorous subset shows consistent patterns, and the human-advantage result is at least not contradicted by the recall measure—but none of that addresses instrument bias. The completion-time filter and the automated-metric overstatement are secondary; they affect peripheral conclusions, not the central comprehension comparison. I therefore see no reason to move beyond the reader's CONDITIONAL verdict: the claim should be accepted only with the source-text condition resolved. If the authors disclose that the input was the abstract, the concern largely evaporates; if not, the main finding needs re-testing with abstract-derived questions.","tokens_in":13433,"tokens_out":7784,"duration_ms":79102,"concrete_test":"Request from the authors the exact input passed to the LLaMA-3 comprehension-question prompt for the 50 abstracts. If the input is the scientific abstract, the concern is resolved as a documentation gap. If it is the human-written PLS, or if the input is not recoverable, run a small replication: for 10-15 abstracts, regenerate the three MCQs from the scientific abstract using the same LLaMA-3 prompt and the same crowdsourced procedure, then compare human-vs-LLM accuracy. The concern is confirmed if the human advantage disappears or shrinks substantially; it is refuted if the advantage persists with abstract-derived questions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.4's central result—human-authored PLSs yield significantly higher MCQ accuracy than every LLM version—requires that the three comprehension questions measure understanding of the underlying study content equally across versions. Section 4.2 gives the exact LLaMA-3 prompt ('Create three multiple-choice questions... assess (1) the motivation of the study, (2) the methods used, and (3) the main results') but omits what text was put into that prompt. The attention-check prompt explicitly refers to 'the text', while the comprehension prompt refers to 'the study' with no input field shown. Because the same three questions were held fixed across all PLS versions, any source-specific bias propagates into every human-vs-LLM comparison. If the input was the human-authored PLS, the correct answers are keyed to the structure, vocabulary, and possibly added background of the human text, making the human condition artificially easier. The authors used LLaMA-3 to avoid GPT-4 self-preference, but source-neutrality is a separate requirement that is not documented. This is not a minor reporting omission: recall differences were not significant, so the MCQ is the only objective measure supporting the central claim, and the claim stands or falls on this instrument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a crowdsourced evaluation of plain language summaries (PLSs) generated by GPT-4 against human-authored PLSs from the CELLS corpus. Using 50 abstract–PLS pairs, six LLM-generated PLS variants (criteria-agnostic, simplification, informativeness, coherence, faithfulness, and all combined), and 150 Amazon Mechanical Turk workers, the authors collect subjective Likert ratings, multiple-choice comprehension questions, and recall responses. They find that LLM-generated PLSs receive subjective ratings comparable to human-written PLSs, but that readers answer multiple-choice questions significantly more accurately after reading human-written PLSs. They also report that most automated evaluation metrics are not significantly associated with comprehension, with QAEval being the only significant predictor in their rigorous subset. The paper argues for comprehension-centered evaluation of PLSs rather than reliance on surface-level metrics or subjective ratings.","tokens_in":13608,"tokens_out":4120,"duration_ms":42969,"significance":"If the central finding is valid, this is a valuable contribution to health communication and NLP evaluation. The study is unusually large for PLS evaluation, uses both subjective and objective measures, and makes a concrete attempt to avoid model self-preference by generating comprehension questions with LLaMA-3 rather than GPT-4. The mixed-effects modeling framework and the comparison of ten automated metrics against human comprehension are also useful. The main claim—that perceived quality does not imply actual comprehension—is important and testable. However, the validity of the headline result rests entirely on the multiple-choice instrument, and a key detail about how that instrument was constructed is missing from the manuscript.","major_comments":[{"comment":"The comprehension-question generation procedure does not disclose the input text. The attention-check prompt explicitly refers to \"the text\", but the comprehension prompt says only \"Create three multiple-choice questions in plain language that assess (1) the motivation of the study, (2) the methods used, and (3) the main results\" and shows no input field. Because the same three questions are used across all PLS versions, the source of these questions is load-bearing for the central claim in Section 2.4. If LLaMA-3 was prompted with the human-written PLS, the correct answers would be keyed to the structure, vocabulary, and possibly added background of the human text, making the human condition artificially easier. Since recall differences were not significant, the MCQ is the only objective measure supporting the claim that human-written PLSs lead to significantly better comprehension. The authors must state what text was supplied to LLaMA-3 when generating the questions, or otherwise demonstrate source-neutrality; without this, the main finding may be a measurement artifact.","section":"Section 4.2"},{"comment":"The statistical comparison in Section 2.4 uses paired t-tests, but the pairing unit is unspecified. Participants were told that each person evaluated only one version, and the study design assigns each participant to one batch; it is unclear whether the pairing is across abstracts (e.g., averaged ratings per abstract per version), across participants, or across annotation pairs. The validity of the reported p-values depends on the correct pairing structure, and the description in Section 4.4 does not define it. Please specify the exact pairing and explain how the design supports it.","section":"Section 4.4"},{"comment":"The \"rigorous evaluation subset\" is constructed by removing responses whose completion time falls outside the 25th to 75th percentile, but no justification is given for this percentile filter, and the filter is applied after other exclusions. This arbitrary choice affects the secondary analyses in Tables 2 and 3, even though the full-dataset results in Appendix Tables 4 and 5 are directionally consistent. The authors should justify the filter, state whether it was applied per participant or per batch, and report whether the conclusions of Section 2.5 and 2.6 are sensitive to the specific percentile threshold.","section":"Section 2 and Tables 2–3"}],"minor_comments":[{"comment":"There is a typo: \"the our Institutional Review Board\" should be \"our Institutional Review Board\".","section":"Section 2.1"},{"comment":"The word \"comperehensively\" in the last sentence of Section 4.2 is misspelled; it should be \"comprehensively\".","section":"Section 4.2"},{"comment":"In the description of mixed-effects models, \"differences in abstract difficult\" should be \"differences in abstract difficulty\".","section":"Section 4.4"},{"comment":"The text says subjective ratings \"predict\" comprehension, but the models are contemporaneous associations within the same reading episode; consider using \"are associated with\" or \"are predictive of\" only if a temporal or out-of-sample analysis is provided.","section":"Section 2.5"},{"comment":"The caption mentions \"factuality\" as a rated dimension, while the text consistently uses \"faithfulness\"; please align the terminology.","section":"Figure 2 caption"},{"comment":"The discussion of readability metrics says human-written scientific abstracts received higher scores than PLSs \"but the difference was not statistically significant\", while Figure 1 asterisks indicate only comparisons of LLM-generated PLSs to human PLSs; please clarify whether the abstract comparison is included in the multiple-comparison adjustment.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"This is a well-motivated study with a potentially important finding, but the undisclosed source of the LLaMA-3-generated comprehension questions is a load-bearing methodological gap. If the authors can show that the questions were generated from the scientific abstract or from a neutral common text, the paper could become publishable after relatively minor revision. If they cannot, the headline claim may need to be substantially softened or the experiment repeated. I would also encourage the editors to ask for the pairing details of the paired t-test, since the current description is not sufficient to verify the reported significance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, large-scale study that probably gets the direction right, but the headline comparison rests on an undocumented choice — what text went into the LLaMA-3 prompt that generated the comprehension questions. That has to be fixed or verified before I'd trust the main claim.\n\nWhat's new: first crowdsourced evaluation to combine Likert ratings, MCQ, and recall for both human and six LLM-generated PLSs, plus the first alignment analysis of ten automated metrics against comprehension. The main empirical pattern—LLM PLSs rated as fluent as human ones but yield lower MCQ accuracy—is a real contribution, assuming the instrument is neutral. They were also careful to use LLaMA-3 for question generation to avoid GPT-4 self-preference, and they used a decent sample (150 MTurk participants, 1,346 annotations, 50 abstract-PLS pairs).\n\nSoft spots: (1) The MCQ generation prompt in §4.2 does not disclose its input. If LLaMA-3 was given the human-written PLS, the correct answers are keyed to that text's structure and wording, and the human-vs-LLM comparison is biased. This is not a minor omission because recall differences were not significant—the MCQ is the only objective measure carrying the central claim. The authors need to state the input source and ideally release the instrument. (2) The abstract overstates the automated-metric result: in the full dataset (Table 5), QAEval, Coherence-chn, BERTScore, Coherence-bag, and ROUGE all predict comprehension; the 'fail to reflect human judgment' narrative depends on the rigorous subset created by a completion-time filter (25th–75th percentile) that is arbitrary and excludes two-thirds of the data. The paper does acknowledge this in §2.6, but the abstract doesn't.\n\nOther notes: the recall measure is token-overlap with the PLS, which is a weak proxy for comprehension, but since it shows no differences it's not load-bearing. The sample diversity and quality controls are fine. The paper should also state data/code availability; the instrument hasn't been released, which is exactly what's needed to resolve the first point.\n\nVerdict: worth serious peer review, conditional on the authors disclosing the question-generation input and either releasing the questions or running a robustness check (e.g., regenerate questions from the abstract). The paper is a good contribution to health NLP evaluation; it just needs to close this gap.","headline":"A large-scale, useful evaluation of LLM plain language summaries whose central human-vs-LLM comparison needs one missing detail disclosed (the input text for the MCQ generator) before the result can be trusted.","tokens_in":14170,"tokens_out":2378,"would_cite":true,"duration_ms":21492,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human-written plain language summaries lead to higher reader comprehension than LLM-generated summaries.","keywords":["plain language summaries","LLM generation","comprehension evaluation","crowdsourced evaluation","health communication","automated evaluation metrics","multiple-choice comprehension","readability measures"],"falsifier":"Generate the comprehension questions from a neutral common source, such as the scientific abstract alone or alternating randomly between versions, keep raters blind to which summary they read, and measure multiple-choice accuracy again. If the human-written advantage disappears, the reported comprehension gap is an artifact of question provenance rather than of summary quality.","tokens_in":13175,"feed_emoji":"🩺","tokens_out":8173,"duration_ms":77104,"temperature":0.7,"pith_summary":"This paper asks whether large language models can write plain language summaries of medical research that lay readers actually understand, not merely summaries that look understandable. Across 50 biomedical abstracts, crowd workers rated six LLM-generated summaries per abstract as similar to the human-written versions in simplicity, coherence, informativeness, and faithfulness, yet the same readers scored significantly lower on identical multiple-choice comprehension questions after reading the LLM versions. Ten automated metrics were checked against comprehension, and most failed; only a question-answering-based metric tracked reader understanding. The paper's point is that perceived quality and actual comprehension diverge, so plain language summaries should be evaluated by what readers can do with the text, not by how fluent it appears.","feed_headline":"Human-written summaries beat LLM summaries on comprehension","feed_subtitle":"A crowdsourced test finds readers rate LLM summaries highly but understand them less.","key_machinery":"The controlled within-abstract comparison is the load-bearing mechanism: each abstract has one human-written PLS and six LLM-generated variants, and the same three multiple-choice questions are attached to every version, so any difference in accuracy can be attributed to the summary text itself. The comprehension questions were generated by a separate LLM and kept identical across versions; recall was scored by token overlap with the original summary. The statistical core is a set of paired t-tests between human and LLM versions plus linear mixed-effects models with random intercepts for participants and abstracts, used to test which human ratings and automated metrics predict multiple-choice accuracy.","core_discovery":"The central claim is that human-written plain language summaries convey biomedical content more effectively than LLM-generated summaries, even though the two are subjectively indistinguishable. The study compares human-written PLSs to six LLM-generated variants—one unoptimized and five optimized for simplification, informativeness, coherence, faithfulness, or all combined—using the same 50 scientific abstracts. All versions were paired with the same three multiple-choice questions about the study's motivation, methods, and results, and participants answered significantly more accurately after reading the human versions. The paper also reports that lexical-overlap metrics (ROUGE, BLEU, METEOR, SARI), perplexity-based fluency, and most model-based metrics do not significantly predict comprehension, while the QA-based metric QAEval does.","pith_inferences":["One testable extension is to use answerability as a training signal: generate summaries, ask a QA model questions about the source abstract, and reward summaries from which those questions can be answered correctly. If the comprehension gap closes, answerability is a useful objective.","The non-significant recall difference suggests the human advantage may be specific to retrieving targeted facts about motivation, methods, and results, not to general memorability. A direct test would score recall at the level of individual claims rather than token overlap.","All six LLM variants behaved similarly regardless of optimization prompt, which leaves open that current prompt-level optimization does not steer comprehension-relevant content. A stricter comparison would vary decoding or add retrieval-augmented background information.","Since only 50 abstracts were sampled, a stratified replication across topics and reader familiarity levels could show whether the human advantage is driven mainly by the background information human authors add."],"forward_implications":["Comprehension questions should become a standard part of plain language summary evaluation, since Likert ratings in this study did not reveal the comprehension gap.","Lexical and perplexity-based automated metrics should not be used as proxies for PLS quality; QA-based metrics such as QAEval align better with reader understanding.","LLM generation strategies should be optimized for comprehension, for example by rewarding summaries that support correct answers to questions, rather than for readability or n-gram similarity.","Including necessary background information seems to help lay readers, because background ratings were among the strongest predictors of multiple-choice accuracy.","Diverse lay audiences matter in evaluation: moderately familiar participants outperformed self-identified experts, so expert panels may misjudge what is understandable."],"supporting_citations":[{"why":"Supplies the paired scientific abstracts and human-written plain language summaries used as the baseline.","marker":"[3]"},{"why":"Documents the specific LLM used to generate the six LLM variants from each abstract.","marker":"[37]"},{"why":"Motivates using a different LLM for question generation, to avoid self-preference bias toward the generation model's own text.","marker":"[38]"},{"why":"Defines QAEval, the question-answering-based metric that was the only automated score significantly associated with comprehension.","marker":"[44]"},{"why":"Provides the LLM-based coherence metrics tested for association with comprehension outcomes.","marker":"[13]"},{"why":"Grounds the subjective evaluation dimensions in an earlier perturbation-based testbed for PLS metrics.","marker":"[18]"}],"fun_headline_variants":["LLM summaries pass subjective test but fail comprehension","Crowdsourced study: Human summaries aid understanding more than LLM","Readers can't tell LLM from human summaries but understand humans better","Automated metrics fail to predict comprehension of LLM summaries","Large-scale test finds LLM summaries rate well but teach less"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main result assumes the three fixed multiple-choice questions do not favor the human-written summaries, but the paper never states what text was used to generate those questions.","fun_headline_variants_meta":{"raw":{"variants":["LLM summaries pass subjective test but fail comprehension","Crowdsourced study: Human summaries aid understanding more than LLM","Readers can't tell LLM from human summaries but understand humans better","Automated metrics fail to predict comprehension of LLM summaries","Large-scale test finds LLM summaries rate well but teach less"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000975,"raw_usage":{"total_tokens":4137,"prompt_tokens":933,"completion_tokens":3204,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":3119}},"tokens_in":549,"tokens_out":3204,"duration_ms":24322,"temperature":1.0,"reasoning_tokens":3119,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:08:55.517931+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate the comprehension questions from a neutral common source, such as the scientific abstract alone or alternating randomly between versions, keep raters blind to which summary they read, and measure multiple-choice accuracy again. If the human-written advantage disappears, the reported comprehension gap is an artifact of question provenance rather than of summary quality.","supporting_citations":[{"cited_title":"& Cohen, T","cited_arxiv_id":null,"evidence_quote":"Supplies the paired scientific abstracts and human-written plain language summaries used as the baseline."},{"cited_title":"Gpt-4 technical report","cited_arxiv_id":null,"evidence_quote":"Documents the specific LLM used to generate the six LLM variants from each abstract."},{"cited_title":"& Feng, S","cited_arxiv_id":null,"evidence_quote":"Motivates using a different LLM for question generation, to avoid self-preference bias toward the generation model's own text."},{"cited_title":"& Roth, D","cited_arxiv_id":null,"evidence_quote":"Defines QAEval, the question-answering-based metric that was the only automated score significantly associated with comprehension."},{"cited_title":"& Leroy, G","cited_arxiv_id":null,"evidence_quote":"Provides the LLM-based coherence metrics tested for association with comprehension outcomes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the subjective evaluation dimensions in an earlier perturbation-based testbed for PLS metrics."}],"review_version":1}