{"id":"eb91e454-083b-49e2-9745-c8e07a6debb4","arxiv_id":"2501.06269","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On a four-image chest X-ray test, GPT-4o correctly identified only the COVID-19 image in a composite view and only the healthy image when shown individually, a 25% success rate in each condition.","lead":"Researchers asked OpenAI's ChatGPT-4o to read chest X-rays from a public COVID-19 dataset and compared its answers to the known labels. The model was right on only a minority of images, so the authors say it is not reliable enough to diagnose alone but might help doctors.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 25% success rate depends on an exact-etiology rubric; by a clinically standard sick/healthy rubric ChatGPT-4o scored 3/4, so the central negative claim is not established without a human-reader baseline.","rationale":"The reader identified unverified ground-truth labels and composite layout as the weakest assumption, which I agree matters. My additional load-bearing concern is that the paper's scoring rubric is not clinically calibrated: it counts as failures cases where ChatGPT correctly detected pneumonia but named a different etiology, and it provides no human-reader baseline to show that exact etiology can be determined from these chest X-rays at all. The 25% figure is therefore partly an artifact of the rubric rather than a pure measure of model capability. The 'assist' claim is also untested, consistent with the reader's overclaim flag. I would not change the CONDITIONAL verdict because the conditions needed—independent label verification, repeated trials, a radiologist baseline, and clinician assessment—are exactly the conditions the reader already requires. The concern strengthens the justification for those conditions without turning a small case report into a fully rejected finding. I set agreement to partial because the reader's weakest assumption is narrower than the rubric/baseline issue I find most load-bearing.","tokens_in":990,"tokens_out":868,"duration_ms":115119,"concrete_test":"Recruit three board-certified radiologists, blinded to the dataset labels, to classify the four Figure 1 images into the same four categories used in the paper and to record whether a clavicle fracture is present in image (a). Then re-run the individual-image protocol 10 times per image at a fixed temperature. If the radiologists do not confirm the dataset labels, or if their exact-etiology accuracy is no better than ChatGPT's, or if ChatGPT's modal diagnoses match the radiologist-confirmed labels more often than 25%, the paper's conclusion must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ChatGPT-4o is 'not sufficient and accurate' is derived from exact-category scoring of four individual chest X-rays (bacteria vs healthy vs virus vs COVID), yielding a 25% success rate. However, the verbatim responses show that the model correctly identified disease in 3 of 4 images: the viral image was read as pneumonia (wrong cause), the COVID image was read as pneumonia (not specifically COVID), and only the bacterial case was missed as an infectious process (read as clavicle trauma). Since chest X-ray has limited ability to differentiate bacterial, viral, and COVID-19 pneumonia on a single image, requiring exact etiology may be an inappropriate accuracy criterion. No radiologist baseline is provided to show what accuracy is achievable on these same four images with the same rubric. The ground-truth labels are taken from covid-chestxray-dataset metadata without independent verification, and the prompts were run once on a stochastic model. The composite trial's 25% is explicitly conditional on the authors' own assumption that ChatGPT did not misread image numbering: 'If ChatGPT did not examine the wrong images due to a misunderstanding of the locations.' Moreover, the positive claim that ChatGPT 'can provide interpretations that can assist medical doctors or clinicians' is untested: no clinician used the outputs, so decision-support value is not measured. The paper's own caveats and missing baselines mean the evidence does not bear the abstract's conclusion as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a small case study in which ChatGPT-4o was asked to interpret four chest X-ray images from the covid-chestxray-dataset. The authors first present a composite test with all four images and then individual tests, comparing the model's free-text diagnoses with dataset labels. They report a 25% success rate in the composite test, note that the model correctly identified only the COVID-19 case, and conclude that ChatGPT is not sufficiently accurate for standalone diagnosis but could assist clinicians. The paper includes verbatim prompts and responses, a discussion of each case, and recommendations for future work.","tokens_in":11408,"tokens_out":3790,"duration_ms":33081,"significance":"If the central claim were established, the paper would provide a useful cautionary data point about general-purpose vision-language models in radiology. The authors are transparent about their prompts and record the raw outputs, and they use a public dataset, which supports reproducibility. However, the study is only a four-image, single-trial case report without a human-reader baseline or clinician evaluation, and the exact-etiology rubric conflates triage ability with etiologic specificity. The assistive claim is not measured at all. These gaps make the current evidence too weak to support the abstract's conclusions, although the descriptive material is a reasonable starting point.","major_comments":[{"comment":"The central negative claim that ChatGPT is 'not sufficient and accurate' is derived from an exact-etiology scoring of four images (25% in the composite test). However, the verbatim responses in the Implementation section show that the model correctly distinguished sick from healthy in three of the four individual cases: the healthy image was read as normal, the viral pneumonia image was read as pneumonia (wrong cause), and the COVID-19 image was read as pneumonia (not specifically COVID-19); only the bacterial case was missed entirely as a clavicle fracture. Because chest radiography has limited ability to differentiate bacterial, viral, and COVID-19 pneumonia from a single image, the exact-etiology rubric may not be an appropriate accuracy criterion. Without a radiologist baseline on the same four images under the same rubric, the reported 25% cannot support the conclusion that the model is not sufficiently accurate for triage or for identifying the presence of disease.","section":"Abstract and Implementation (composite test)"},{"comment":"The paper's positive claim that ChatGPT 'can provide interpretations that can assist medical doctors or clinicians' is untested. No clinician or radiologist reviewed the model's outputs for usefulness, correctness of the reasoning, or safety in a decision-support workflow. The Conclusion's statement that the verbatim responses provide 'strong evidence' for an active role in medical image processing is therefore unsupported by the reported methods, which only compare outputs with dataset labels.","section":"Conclusion"},{"comment":"The accuracy result is conditional on two unverified assumptions. First, the ground-truth labels are taken from the covid-chestxray-dataset metadata without independent radiologist confirmation that each image is correctly labeled (Data and Method, Table 1). Second, the composite test requires that ChatGPT interpreted the numbered positions in Figure 1 in the intended order; the paper itself states that the 25% success rate applies only 'If ChatGPT did not examine the wrong images due to a misunderstanding of the locations.' If either assumption fails, the reported success rate and the per-case conclusions do not hold.","section":"Implementation"},{"comment":"The study consists of a single trial per image and per composite presentation with no repeated sampling of the stochastic model. There is no statistical analysis, confidence interval, or human-reader comparison, so the reported accuracy estimates cannot be separated from sampling variation. Adding multiple independent runs, a small set of experienced readers, and a pre-specified scoring rubric would be needed to make the central claim robust.","section":"Data and Method"}],"minor_comments":[{"comment":"The in-text citation 'Chon et al. (2020)' should be 'Cohen et al. (2020)'; the reference list correctly cites Cohen et al., so the in-text name is inconsistent.","section":"Data and Method (Table 1)"},{"comment":"There is a typo in the reference for Sechopoulos et al.: 'digital übreast tomosynthesis' should be 'digital breast tomosynthesis.'","section":"References"},{"comment":"The paper uses inconsistent terminology for the virus ('virus' vs 'COVID-19' vs 'Covid'); standardizing to 'COVID-19' throughout would improve readability.","section":"General"},{"comment":"Several listed references (e.g., Lecler et al., 2023) are not cited in the body, and the related-work section would benefit from an explicit discussion of existing vision-language model evaluations in chest X-ray interpretation.","section":"References and Related Work"},{"comment":"Some figures (e.g., Figure 2 and Figures 3–6) are not explicitly called out in the text at the point where their analysis is discussed; adding in-text references would help the reader map the responses to the images.","section":"Figures"},{"comment":"The claim 'The success rate is very low' should be quantified (e.g., 1/4 in the composite test) so that the reader can assess the magnitude without re-reading the Implementation section.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"This is a short descriptive case study rather than a full experimental paper. The main risk is that the conclusions in the abstract and conclusion overstate what a four-image, single-trial analysis can support. I would encourage the editor to ask for the claims to be scaled back or for the additional experiments described in the major comments. The citation list contains several references that are not cited in the body, which may require attention during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a four-image case study of GPT-4o reading chest X-rays from the covid-chestxray-dataset. What the paper does well is show its work: every prompt and full response is quoted. That transparency is real. It is also honest about the tiny sample and says more experiments are needed.\n\nWhat is new is almost nothing. Prior evaluations of GPT-4 on chest X-rays already exist, and the paper cites Lecler et al. The only addition is a composite-image prompt (four X-rays in one figure) plus individual prompts, with n=4 and no repeated sampling.\n\nThe soft spots are bigger than the positive content. The 25% success rate is computed against an exact-etiology rubric (bacterial vs viral vs COVID vs healthy). Under a clinically more standard sick/healthy rubric, the model read three of four images as showing pneumonia or disease; only the bacterial case was read as trauma. Chest X-ray has limited ability to separate bacterial, viral, and COVID-19 pneumonia on a single image, so requiring exact etiology may not be a fair accuracy criterion. Without a radiologist baseline on these same four images under the same rubric, the central negative claim is not established. The ground-truth labels come from dataset metadata with no independent verification. The composite trial also depends on the assumption that ChatGPT mapped the numbered labels to the correct positions, which the authors themselves flag. And the \"can assist doctors\" claim is untested: no clinician ever looked at the outputs.\n\nI'd still say the paper is a serious attempt; it doesn't hide its failures. But the evidence is too thin and the design too cursory to support the abstract's conclusions. A peer reviewer would need major revisions: repeated trials, a human-reader baseline, label verification, and a clinician assessment of the assist value.\n\nMy recommendation: I wouldn't send this to peer review as a full research paper. It is more of a preprint-level case note. It might be worth a reading group discussion about evaluation rubrics, but it shouldn't be cited as evidence for or against GPT-4 performance.","headline":"A transparent but tiny four-image case study; the 25% figure depends on an exact-etiology rubric and no human baseline, so the conclusions outrun the evidence.","tokens_in":11937,"tokens_out":2796,"would_cite":false,"duration_ms":26987,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o correctly identified only one of four chest X-rays, so the paper concludes it cannot yet make final diagnoses on its own.","keywords":["OpenAI","ChatGPT","GPT-4","Artificial Intelligence","Decision Support System","Medical Doctor","chest X-ray interpretation","covid-chestxray-dataset"],"falsifier":"Take the same four chest X-rays, have two radiologists independently confirm the labels, then prompt GPT-4o again with the four panels in a randomized order and at full resolution; if the model gets more than one of four correct, the paper's 25% success rate does not generalize.","tokens_in":10959,"feed_emoji":"🩻","tokens_out":7635,"duration_ms":64942,"temperature":0.7,"pith_summary":"This paper asks whether OpenAI's GPT-4o can interpret chest X-rays well enough to replace a medical doctor or to serve as a decision-support tool. The authors test four labeled lung images from a public COVID-era dataset, presenting them both as a four-panel composite and individually. GPT-4o correctly identified only the COVID-19 film in the composite (a 25% success rate) and, when shown individual images, correctly recognized the healthy case while misreading bacterial pneumonia as a clavicle fracture, viral pneumonia as bacterial, and COVID-19 as generalized pneumonia. The authors conclude that ChatGPT is not sufficiently accurate to make a final diagnosis in its current form, but that its structured, criterion-based responses could assist doctors and clinicians.","feed_headline":"AI chatbot gets 1 of 4 chest X-rays right in doctor test","feed_subtitle":"Four chest films, one correct call: ChatGPT can assist doctors but cannot diagnose on its own.","key_machinery":"The carrying mechanism is a prompted image-interpretation protocol: GPT-4o receives chest X-rays as image inputs, is instructed to act as a radiologist, and is asked to classify each film as healthy or sick and, if sick, by cause (bacteria, virus, COVID-19, or other). The model's outputs are compared against the ground-truth labels of the covid-chestxray-dataset, with the same images tested in composite and individually to probe consistency.","core_discovery":"On the paper's own terms, the discovery is that a state-of-the-art general-purpose multimodal chatbot, prompted to act as a radiologist, produces fluent and anatomically structured readouts yet fails on three of four chest X-rays when judged against the dataset's labels. The model's only unambiguous success was identifying a healthy film as normal; it missed bacterial pneumonia, confused viral pneumonia with bacterial infection, and did not name COVID-19 as the likely cause of the fourth film even though it flagged pneumonia. The paper therefore establishes a negative result about standalone use and a weakly evidenced positive result about assistance: the model can enumerate radiographic criteria and recommend next steps, but its diagnostic accuracy is too low for autonomous deployment.","pith_inferences":["The measured 25% success rate could be an artifact of the composite layout: if GPT-4o misread which panel was labeled 1, 2, 3, or 4, its diagnoses might have been correct for the wrong images. A repeat with clearly labeled, full-resolution panels would separate spatial reasoning from image interpretation.","The ground-truth labels are taken from the dataset's filenames without independent radiologist re-reads; a label error in any of the four films would change the success count. Re-verifying the four images is a cheap way to test the paper's core number.","A larger follow-up with dozens of films, multiple prompt phrasings, and statistical confidence intervals would be needed before using these results to decide real-world deployment.","The paper's optimism about assistance rests on process quality (fluent, structured reports), not on measured accuracy; unless assistance is evaluated separately, the claim that the model 'can assist' remains unproven."],"forward_implications":["General-purpose multimodal LLMs should not be deployed as standalone chest X-ray readers in their current form.","Their structured radiology-style reports could still serve as a triage or educational aid, but every output needs clinician verification.","Diagnostic performance is sensitive to how images are presented, so evaluations must standardize layout and resolution.","Specialized training on infectious etiologies and integration of clinical data are prerequisite steps before such tools can be relied upon.","The model's correct read of the healthy film hints that ruling out obvious disease may be a more attainable near-term use than pinpointing etiology."],"supporting_citations":[{"why":"Supplies the covid-chestxray-dataset repository from which the four test films were drawn.","marker":"Github, 2020"},{"why":"Provides the ground-truth labels (bacterial, healthy, viral, COVID) used as the accuracy benchmark in Table 1.","marker":"Chon et. al, 2020"},{"why":"Marshals the prior result that dedicated AI matches radiologists in breast cancer screening, the motivation for testing whether GPT-4o can do similar work.","marker":"Rodriguez-Ruiz et al., 2019a"}],"fun_headline_variants":["GPT-4 aces only one in four chest X-rays","ChatGPT flunks three of four radiology reads","AI doctor test: GPT-4 misses pneumonia, COVID","GPT-4 sees 1 of 4 X-rays correctly","Study: ChatGPT not reliable for chest X-ray reads"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole accuracy count depends on the four images' labels being correct and the composite figure being laid out in the assumed order, since neither was independently verified by a radiologist.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 aces only one in four chest X-rays","ChatGPT flunks three of four radiology reads","AI doctor test: GPT-4 misses pneumonia, COVID","GPT-4 sees 1 of 4 X-rays correctly","Study: ChatGPT not reliable for chest X-ray reads"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000329,"raw_usage":{"total_tokens":1790,"prompt_tokens":853,"completion_tokens":937,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":856}},"tokens_in":469,"tokens_out":937,"duration_ms":8890,"temperature":1.0,"reasoning_tokens":856,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:12:16.254750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same four chest X-rays, have two radiologists independently confirm the labels, then prompt GPT-4o again with the four panels in a randomized order and at full resolution; if the model gets more than one of four correct, the paper's 25% success rate does not generalize.","supporting_citations":[{"cited_title":"https://github.com/ieee8023/covid-chestxray-dataset Harrer, S., Shah, P., Antony, B., & Hu, J","cited_arxiv_id":null,"evidence_quote":"Supplies the covid-chestxray-dataset repository from which the four test films were drawn."}],"review_version":1}