{"id":"396b60df-2a1b-4ce2-b3e5-8429a11cfeb2","arxiv_id":"2506.17163","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MedPerturb finds that LLMs are more sensitive to gender and style changes in clinical text, while medical students are more sensitive to LLM-generated summaries and dialogues, in triage decisions.","lead":"This paper introduces MedPerturb, a dataset of 800 clinical vignettes with controlled gender, style, and format perturbations, plus triage decisions from four LLMs and medical student annotators. It reports that LLMs shift treatment choices more than humans when gender or style changes, while humans shift more than LLMs when input is summarized or turned into a dialogue.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unpaired human annotator design invalidates the paired human statistics underlying both headline sensitivity claims.","rationale":"The reader's rejection is well-founded: the paper's central comparative claims rest on paired human statistics that the design cannot support. Section 3.3's no-overlap rule means that no human annotator judged both a baseline and a perturbed version of the same clinical case. Appendix K.3 nonetheless justifies the paired t-test by saying 'each treatment decision under the baseline is naturally paired with a corresponding decision under the perturbation for the same annotator or model instance.' That statement is true only for LLMs. For humans, any baseline–perturbed pairing must be imposed by the analyst; if pairing is by context ID, the difference includes inter-annotator variability, so the significant human shifts in Case Study II may not reflect sensitivity to format at all. The same flaw contaminates Case Study I: the mutual information formula in Section 4.1 is defined for joint decisions of the same decision-maker, but humans provide only unpaired labels. A human MI computed from independent annotators is not a measure of stability and cannot be compared to LLM MI to conclude 'clinicians produce more stable decisions' under gender/style perturbations. Thus both headline claims—LLM sensitivity to gender/style and human sensitivity to format—are unsupported by the reported statistics. This is a correctness risk, not a disagreement with the field; the released dataset and code may allow the authors to repair the analysis. If a re-analysis with unpaired human comparisons preserves the effects, the paper could be reconsidered. For now the reader's REJECT verdict should stand.","tokens_in":36454,"tokens_out":7311,"duration_ms":72866,"concrete_test":"Rerun Section 4.2's human ATR/PC comparisons with an unpaired two-sample permutation test over annotators (reflecting Section 3.3's no-overlap design), rather than the paired t-test in Appendix K.3; if the summary-vs-vignette shift is no longer significant at p<0.01, the claim that humans are more sensitive to format perturbations is unsupported. A parallel rerun of Case Study I human MI with permuted baseline–perturbed pairing would settle whether the gender/style sensitivity gap also depends on the invalid pairing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 states that 'no labeler was provided with two clinical contexts corresponding to the same baseline clinical context.' Consequently, for humans there is no baseline–perturbed pair from the same annotator, so the paired t-tests and percent-change computations in Case Study II (Section 4.2, Appendix K.3) cannot be computed as written: Appendix K.3 explicitly relies on 'same annotator or model instance' pairing. Pairing by context ID instead mixes inter-annotator variability with perturbation effects, so the claimed human shift toward self-management under summaries (Section 4.2) may be an artifact of annotator heterogeneity. The same design flaw undermines Case Study I: the mutual information in Section 4.1 requires a joint distribution P(t_base=a, t_pert=b) of decisions by the same individual; with disjoint annotators, human MI measures agreement across different people, not decision stability, so the claim that humans are less sensitive to gender/style perturbations than LLMs is unsupported. Because both headline claims compare LLM (paired) with human (unpaired) statistics, the central comparative message is not supported by the reported tests.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces MedPerturb, a dataset of 800 clinical vignettes perturbed along three axes—gender, style, and format—with 7,200 human clinician reads and 28,800 LLM reads across three binary triage questions (MANAGE, VISIT, RESOURCE). Two case studies compare human and LLM sensitivity to these perturbations using average treatment rates (ATR), mutual information (MI), and percent change of recommendations (PC). The paper claims that LLMs are more sensitive to gender and style perturbations while human annotators are more sensitive to LLM-generated format perturbations such as clinical summaries, and that LLMs underrecommend self-management and overallocate resources relative to clinicians. The dataset and code are publicly released with detailed documentation of prompts, sampling procedures, annotator demographics, and IRB/consent information.","tokens_in":36682,"tokens_out":10789,"duration_ms":119485,"significance":"The public release of a systematically perturbed clinical vignette dataset with paired LLM and human clinician judgments is a potentially valuable resource for studying non-content robustness in clinical LLMs and human-AI alignment. The paper documents its data creation pipeline, model sampling, and annotator demographics in unusual detail, and the reproducibility materials (code, Hugging Face dataset, website) are concrete strengths. If the comparative sensitivity claims were valid, the resource would support a meaningful, falsifiable empirical message about how surface-level changes affect LLM and human treatment decisions. However, as detailed below, the key statistical comparisons for human annotators rest on a pairing assumption that the data collection design explicitly violates, so the headline findings are not currently supported. The paper's own limitations section candidly acknowledges some related concerns, but it does not address the specific unpaired-annotator problem that undermines both case studies.","major_comments":[{"comment":"Section 3.3 states that 'no labeler was provided with two clinical contexts corresponding to the same baseline clinical context.' The Percent Change metric PC_q in Section 4.2 is defined as (1/N)Σ|t_pert_i,q - t_base_i,q|, which requires a baseline–perturbed pair from the same decision-maker i. For humans, no such pairs exist, so the computation must either pair different annotators arbitrarily or use context-level aggregates; neither choice supports the paired t-test whose validity Appendix K.3 justifies by 'the same annotator or model instance.' Consequently, the reported ~30% increase in human self-management recommendations under summaries and the ~20% decrease in resource allocation are not established by the tests as written, and the second headline claim—that human annotators are more sensitive to format perturbations—is unsupported.","section":"§3.3 and §4.2, Appendix K.3"},{"comment":"The mutual information MI_q is defined through the joint probability P(t_base_q=a, t_pert_q=b) of decisions under baseline and perturbed conditions. For LLMs, this can be computed for the same model instance across sampling seeds; for humans, the disjoint annotator design means the joint distribution is actually a cross-annotator agreement table, not a measure of decision stability. A high or low human MI in this setting reflects inter-annotator (dis)agreement, so the comparison between human and LLM MI does not support the conclusion that 'clinicians tend to produce more stable and internally consistent treatment decisions.' In addition, the Mann–Whitney U test described in Appendix K.4 is underspecified: the main text reports one MI value per treatment question, but a U test requires a sample of MI values per group, and the constitution of these samples is never described. The first headline claim—that LLMs are more sensitive to gender and style perturbations than humans—therefore lacks a valid statistical basis.","section":"§4.1 and Appendix K.4"},{"comment":"The unit of analysis is ambiguous throughout the case studies. Section 4.1 defines t_i,q as the treatment selected by 'annotator or LLM instance i' and N as the number of prompts, but the dataset provides three clinician reads and twelve LLM reads (four models × three seeds) per prompt. It is never stated whether the analyses use individual reads, model runs, majority votes, or some other aggregation. This ambiguity is load-bearing because both headline claims depend on the paired tests, and without a clear definition of what is being paired, the reported p-values and error bars cannot be verified. The authors should specify the exact observations entering each test, or re-run the analyses with an explicitly defined and appropriate aggregation.","section":"§4.1–4.2"}],"minor_comments":[{"comment":"The introduction closes with 'We provide the MedPeturb dataset,' which should be corrected to 'MedPerturb.'","section":"Section 1"},{"comment":"The format perturbation paragraph contains the typo 'multiturn converesations'; it should read 'multiturn conversations.'","section":"Section 3.1"},{"comment":"The heading 'Text Analysis' introduces 'how the texts are changed by perturabtion'; this should be 'perturbation.'","section":"Appendix E.3"},{"comment":"The notation 'Fleiss’κp' is not defined; the subscript p should be explained or removed, as it is not standard for Fleiss' kappa.","section":"Section 3.4"},{"comment":"The sentence 'We do not that for models, longer summaries tend to receive less visitation recommendations' appears to be missing the word 'note' and should be revised.","section":"Appendix N.1.2"}],"recommendation":"reject","confidential_remarks":"The central issue is the mismatch between the unpaired human annotation design and the paired statistics used for both headline claims. The dataset itself may be reusable with a corrected analysis that uses explicitly defined context-level aggregates and appropriate unpaired or paired-by-context tests, but as submitted the main comparative findings are unsupported. A revision would need either new paired human data or a substantial reframing of the claims, which goes beyond local fixes. This is a fit issue as well: the resource is potentially valuable, but the current case studies do not yet support the paper's advertised conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my take on MedPerturb. The dataset is a real contribution: 800 clinical contexts, perturbed along gender, style, and format axes, with parallel labels from four LLMs and three medical-student annotators per context. That resource, plus the released code and data, is the strongest part of the paper. The aggregate finding that LLMs under-recommend self-management and over-allocate resources is plausibly robust and worth reporting.\n\nThe trouble is that both headline comparative claims are statistically unsupported as written. Section 3.3 states that no labeler saw two contexts from the same baseline. That means the human baseline and perturbed responses are unpaired. Yet Appendix K.3 defends the paired t-test on the grounds that each decision under baseline is paired with a decision under perturbation for the same annotator or model instance. That assumption is false for humans. The paired-t comparisons in Case Study II (Section 4.2) therefore mix inter-annotator variability with perturbation effects; the claimed human shift toward self-management under summaries could be annotator heterogeneity. The mutual information analysis in Case Study I has the same problem: MI requires a joint distribution of decisions by the same individual, and with disjoint annotators the human MI measures cross-person agreement, not decision stability. So the claim that humans are less sensitive to gender and style than LLMs is not established by the reported tests. The Mann-Whitney U comparison of MI values is also underspecified, since the paper reports a single MI per question, making that test very low powered.\n\nThere are smaller issues. The annotators are medical students, not practicing clinicians; calling them 'human expert reads' and 'clinicians' throughout oversells the population. The perturbations themselves are LLM-generated, which the authors acknowledge in the limitations, but that confound is worth taking seriously when interpreting the style and format results.\n\nNone of this kills the dataset. The resource is reusable, the documentation is thorough, and a careful revision that either uses unpaired tests, reports per-annotator shifts correctly, or softens the comparative claims could make the paper publishable. As it stands, the central message is not supported.\n\nI would send this to peer review, not desk-reject it, because the dataset deserves scrutiny and the methodological problems are fixable. But I would expect major revision before acceptance.\n\nBest.","headline":"Valuable dataset, but the headline human-LLM sensitivity claims are undermined by paired statistics applied to unpaired human annotations.","tokens_in":37164,"tokens_out":3107,"would_cite":true,"duration_ms":29811,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Same clinical content, rephrased, moves LLM and human triage decisions in opposite directions.","keywords":["MedPerturb","clinical robustness","LLM evaluation","perturbation sensitivity","gender bias","style variation","format perturbation","human versus LLM decision-making"],"falsifier":"Have each clinician read the original vignette and the AI summary of the same case in randomized order, then compute within-clinician flip rates for self-management and resource allocation. If the flip rate drops toward the rate seen for repeated reads of unchanged vignettes, the reported human sensitivity to format is an artifact of using different annotators for the two conditions.","tokens_in":36274,"feed_emoji":"🩺","tokens_out":5290,"duration_ms":55896,"temperature":0.7,"pith_summary":"This paper introduces MedPerturb, a dataset of 800 clinical vignettes in which the same medical content is re-expressed along three non-clinical axes: gender markers, language style (uncertain or colorful), and format (multiturn conversation or LLM summary). Using three triage questions—self-manage at home, seek a visit, and use extra resources—the authors compare four LLMs to 36 medical students across 7,200 human reads and 28,800 LLM reads. They find that LLMs change treatment recommendations more than clinicians when gender or tone changes, while clinicians change recommendations more than LLMs when the same case is presented as an AI summary or agentic conversation. The point of the paper is that surface-level rephrasing is not neutral: it can move model and human treatment decisions, and it moves them in opposite directions.","feed_headline":"Gender and style sway LLMs; AI summaries sway clinicians","feed_subtitle":"Surface rephrasing of the same case flips model and human triage choices in opposite directions.","key_machinery":"The load-bearing object is the MedPerturb dataset and its paired baseline-perturbation structure: 200 base vignettes expanded to 800 contexts by three controlled transformations, each read by three clinicians and four LLMs on three binary triage questions, yielding 7,200 human and 28,800 model decisions. The argument runs on three metrics computed pairwise: Average Treatment Ratio (the fraction of yes answers per question), mutual information between baseline and perturbed decisions (stability), and percent change (flip rate). The dataset carries the claim because every comparison is within the same clinical content, so any measured shift is attributed to the perturbation rather than to new medical information.","core_discovery":"The central discovery is that identical clinical content, rewritten along non-clinical axes, changes treatment decisions in opposite directions for machines and people. Across gender-swapped, gender-removed, uncertain-style, and colorful-style versions of cancer and patient-forum vignettes, the four LLMs altered their yes/no recommendations on self-management, visits, and resource allocation substantially more than clinicians did; clinician decisions had higher mutual information with the baseline, meaning they stayed more stable. The pattern reverses for format: when a vignette is turned into an AI-fabricated doctor-patient dialogue or a third-person LLM summary, clinicians shifted—recommending roughly 20–30% more self-management and fewer resource allocations—while LLM aggregate recommendations barely moved. The paper interprets this as evidence that static benchmarks miss the real failure mode: not whether the model knows the content, but whether surface cues hijack either the model's or the clinician's judgment.","pith_inferences":["The opposite sensitivity directions imply that LLM-as-judge evaluations of clinical text could systematically misestimate human impact: models are insensitive to the exact format changes that move human clinicians.","A within-subject replication, in which each clinician reads both the original vignette and its AI summary, would test whether the reported human format-sensitivity is a genuine content effect or partly an artifact of comparing different annotator pools.","The correlation between conversation turn count and clinician self-management recommendations suggests that redundancy may signal lower acuity to humans; controlled rewrites that preserve information while varying redundancy could separate that signal from actual information loss.","Deployment studies for clinical LLMs should measure clinician behavior changes, not just model self-ratings, because the present results indicate that model and human sensitivities are negatively correlated."],"forward_implications":["Static accuracy benchmarks can hide clinically meaningful brittleness: two systems with identical aggregate treatment rates may still differ sharply in how much their decisions move under surface changes.","LLM-based triage tools may systematically under-recommend self-management and over-order labs and referrals, which could strain health systems if deployed without human oversight.","Gender and style cues that human clinicians ignore can change LLM recommendations, making fairness auditing of clinical LLMs a requirement rather than an afterthought.","Because AI summaries and agentic conversations move human clinicians toward more self-management and fewer resources, evaluating clinical summarizers by faithfulness scores alone will not capture their downstream effect on care decisions."],"supporting_citations":[{"why":"Supplies the OncQA cancer patient message baseline used for gender and style perturbations.","marker":"[47]"},{"why":"Supplies the r/AskaDocs informal patient post baseline used for gender and style perturbations.","marker":"[48]"},{"why":"Supplies the USMLE and Derm vignettes used as the baseline for format perturbations.","marker":"[60]"},{"why":"Provides the CRAFT-MD framework adapted to generate multiturn conversations and summaries.","marker":"[78]"},{"why":"Defines the three triage questions (MANAGE, VISIT, RESOURCE) that structure all human and LLM reads.","marker":"[50]"},{"why":"Llama-3 models generate the gender and style perturbations and supply two of the four LLM judge models.","marker":"[79]"},{"why":"GPT-4 generates the multiturn and summarized variants and provides one of the four LLM judge models.","marker":"[84]"}],"fun_headline_variants":["LLMs sway on style, humans on summaries","Same case, flipped choices: LLMs vs clinicians","Gender and style trip LLMs; summaries trip clinicians","Non-clinical rewrites flip AI and human decisions","Doctors and models diverge on surface cues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that clinicians are more sensitive to format perturbations assumes baseline and perturbed clinician reads are paired, but the study assigned different annotators to each condition; if those annotator pools differ, the paired tests can attribute annotator differences to the perturbation instead of to the text format.","fun_headline_variants_meta":{"raw":{"variants":["LLMs sway on style, humans on summaries","Same case, flipped choices: LLMs vs clinicians","Gender and style trip LLMs; summaries trip clinicians","Non-clinical rewrites flip AI and human decisions","Doctors and models diverge on surface cues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1194,"prompt_tokens":985,"completion_tokens":209,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":135}},"tokens_in":601,"tokens_out":209,"duration_ms":2408,"temperature":1.0,"reasoning_tokens":135,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:10:22.517633+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have each clinician read the original vignette and the AI summary of the same case in randomized order, then compute within-clinician flip rates for self-management and resource allocation. If the flip rate drops toward the rate seen for repeated reads of unchanged vignettes, the reported human sensitivity to format is an artifact of using different annotators for the two conditions.","supporting_citations":[],"review_version":2}