{"id":"874b6750-fe5b-4066-8d44-db64d2d90dda","arxiv_id":"2508.21777","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-5 outperformed older GPT models on a radiation oncology exam, but expert ratings of its treatment plans were variable with low inter-rater agreement.","lead":"The study benchmarked GPT-5 on a 300-question radiation oncology exam and on 60 real patient case vignettes. GPT-5 scored 92.8% on the exam and received mostly positive but variable expert ratings on treatment plans, with low inter-rater agreement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Vignette pillar of the central claim is unsupported: expert ratings have Fleiss' kappa=0.083 for correctness and negative kappa for comprehensiveness/hallucinations (Table 2), so the reported means and 10% hallucination rate are not demonstrably reproducible. TXIT gains may stand, but 'highly rated,","rationale":"The strongest part of the paper is the TXIT benchmark: five repeated runs, a tight accuracy range, the same item pool and constrained output format as the reproduced GPT-3.5/4 baselines, and consistency with prior work. That portion of the central claim—measurable gains on the exam—is credible and would survive the concern raised here. The real-world vignette claim, however, has a decisive weakness located inside the paper's own results: the four expert ratings that constitute the outcome measure have Fleiss' kappa = 0.083 for correctness, -0.016 for comprehensiveness, and -0.111 for hallucinations (Section 3.4, Table 2). The authors do not define 'hallucination' in Methods 2.3, do not provide an itemized scoring rubric, and do not compare against human reference plans for the same cases. The Discussion and Limitations acknowledge inter-rater variability but then proceed to use the aggregated Likert means and the 10% hallucination rate as stable measurements. This is internally inconsistent, and it is exactly the condition that must hold for the claim that GPT-5's recommendations were rated highly with rare hallucinations. The reader's weakest assumption identifies the same issue. I would not reject the paper: the TXIT result and the qualitative error examples still support a conditional conclusion. But the vignette pillar should be treated as unverified until a reproducibility check with a clearer rubric or a second panel confirms the aggregate ratings. Since the reader already returned CONDITIONAL, my stress-test does not change that verdict; it reinforces it.","tokens_in":13771,"tokens_out":9948,"duration_ms":119927,"concrete_test":"Conduct a second independent rating study: recruit four different board-certified radiation oncologists, give them the same 60 GPT-5 vignette outputs along with a pre-specified itemized scoring rubric (each required plan element—stage, intent, modality, dose/fractionation, volumes/OAR, toxicity, follow-up—scored as correct/incorrect/uncertain) and a predefined hallucination definition. Compute mean correctness, hallucination rate, and Fleiss' kappa/ICC. If the reproduced mean correctness falls outside the reported 95% CI (3.11–3.38), or the hallucination rate exceeds the reported 10% by a meaningful margin, or agreement remains near zero (kappa < 0.2), the original vignette conclusion is rater-dependent and cannot support the central claim. If the mean replicates and kappa > 0.4, the concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The real-world claim that GPT-5's treatment recommendations are 'highly rated' with 'rare' hallucinations rests entirely on the 60-vignette ratings (Section 3.4, Table 2). The paper itself reports Fleiss' kappa = 0.083 for correctness, -0.016 for comprehensiveness, and -0.111 for hallucinations. This is not a minor caveat: the four board-certified raters did not agree on which outputs were correct beyond chance, and their binary hallucination flags were negatively correlated. The mean correctness (3.24/4) and comprehensiveness (3.59/4) are averages of ratings that the manuscript shows are not reproducible, and the '10% hallucination rate' is derived from the same unreliable flags. The statement that 'no case reached majority consensus' is used to downplay hallucinations, but with kappa = -0.111 this is exactly what would be expected from idiosyncratic raters, not evidence of rarity. The Discussion (Section 4) and Limitations (Section 5) acknowledge inter-rater variability but still treat the aggregate ratings as objective measurements. No human baseline, no consensus/adjudicated reference, and no explicit definition of 'hallucination' are given in Methods Section 2.3. Therefore, unless the ratings can be shown reproducible under a clearer rubric, the vignette portion of the central claim is not established. The TXIT benchmark, by contrast, is comparatively robust: five runs, a tight accuracy range (92.3–93.0%), and the same item pool and constrained response format as the reproduced GPT-3.5/4 baselines.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a benchmark of GPT-5 in radiation oncology using two evaluation settings: (i) the ACR Radiation Oncology In-Training Examination (TXIT) 2021, where GPT-5 achieved a mean accuracy of 92.8% over five runs, compared with published GPT-4 (78.8%) and GPT-3.5 (62.1%) baselines; and (ii) 60 real-world clinical vignettes rated by four board-certified radiation oncologists for correctness (mean 3.24/4), comprehensiveness (mean 3.59/4), and hallucinations (10% of ratings). The authors report low inter-rater reliability (Fleiss' kappa = 0.083 for correctness, -0.016 for comprehensiveness, -0.111 for hallucinations) but nevertheless conclude that GPT-5 can generate coherent, comprehensive drafts for supervised use in education, pre-board preparation, and tumor-board draft generation, while requiring expert oversight.","tokens_in":14197,"tokens_out":5059,"duration_ms":56835,"significance":"If the claims are substantiated, this would be one of the first comprehensive evaluations of GPT-5 in radiation oncology and would provide a useful comparison point for clinical LLM benchmarks. The TXIT part is valuable because it uses a public exam and a reproducible API pipeline, and the five-run range is narrow. The vignette dataset is also a worthwhile resource, covering a broad set of real-world disease sites and treatment intents. The paper ships prompts, seeds, and raw outputs in the supplementary material, which is a strength for reproducibility. However, the vignette-based central claim is currently undermined by the reported lack of inter-rater agreement, and the TXIT comparison needs clarification regarding the item denominator and statistical reporting.","major_comments":[{"comment":"The GPT-5 comparison is asymmetric: GPT-5 was run on the full 300-item TXIT including seven image-based questions, whereas GPT-4/GPT-3.5 baselines were computed on 293 text-only items. The reported mean of 92.8% appears to be computed over the 293 text items, because adding the 7 image items (2/7 correct) would lower the overall accuracy to ~91.3%. Please state explicitly which denominator was used. If GPT-5 is compared on 293 text items, the comparison is fair but must be stated; if it is compared on 300 items, the baselines need to be recomputed on the same item pool. Also report the SD/CI for the five runs and a test (e.g., bootstrap or McNemar) that the difference vs. GPT-4 is significant.","section":"§2.2/§3.1"},{"comment":"The vignette results rest on ratings with essentially no inter-rater reliability: Fleiss' kappa = 0.083 for correctness, -0.016 for comprehensiveness, -0.111 for hallucinations. The reported means (3.24, 3.59) and the 10% hallucination rate are averages of ratings whose reproducibility is not demonstrated. No adjudicated reference standard, consensus procedure, or per-rater agreement analysis is provided. The statement that no case reached majority consensus for hallucinations is expected when kappa is near zero and cannot be used as evidence of rarity. Please either add a consensus/adjudication step and report agreement on final labels, or weaken the vignette conclusions to descriptive exploratory findings.","section":"§3.4/Table 2"},{"comment":"The conclusion 'Hallucinations were rare and not a substantive concern' is not supported: 24/240 ratings (10%) flagged hallucinations, with 40% of cases receiving at least one flag and no case reaching majority. Given the negative kappa for hallucination flags, this could reflect idiosyncratic rater behavior rather than a low true hallucination rate. At minimum, define 'hallucination' operationally, provide examples of flagged outputs, and report the rate under a stricter/consensus definition.","section":"§4/§6"},{"comment":"No human baseline is reported for the vignette task. Without clinician performance on the same 60 cases, the practical meaning of a mean correctness of 3.24/4 cannot be interpreted. A small panel of additional radiation oncologists rating the same vignettes (or a subset) would also help establish whether the low reliability is in the rubric or in the outputs.","section":"§3.4"}],"minor_comments":[{"comment":"For GPT-5 only the range is reported without SD, while Figure 1's caption says error bars show SD. Add the SD for the five runs.","section":"§3.1"},{"comment":"The item-count description is confusing: 'Fourteen questions included medical images, of which seven required visual interpretation.' Clarify whether the other seven image questions were text-scorable and whether they were included in the 293-item baseline.","section":"§2.2"},{"comment":"The phrase 'no case reaching majority consensus for their presence' is misleading in the abstract, since low kappa makes exactly this pattern likely by chance.","section":"Abstract"},{"comment":"The text says exploratory subgroup analyses were prespecified, but no analysis plan or protocol is given. Specify whether subgroups were defined before seeing the data and report the number of comparisons.","section":"§2.4"},{"comment":"The prompt listing includes a JSON schema and German-language instructions. Please confirm this is the exact prompt used for all cases and report decoding parameters (temperature, top-p, etc.), which are not given in §2.1.","section":"Supplementary Material"},{"comment":"Typo: 'artifical intelligence' should be 'artificial intelligence'.","section":"§2.1"}],"recommendation":"major_revision","confidential_remarks":"The TXIT comparison is potentially robust and useful, but the vignette pillar needs substantial reanalysis before the central 'highly rated' and 'hallucinations rare' conclusions can be accepted. The inter-rater agreement problem is acknowledged in the limitations but not allowed to constrain the interpretation. I would be willing to see a revised version that either provides a consensus/adjudication procedure or substantially reduces the strength of the vignette claims. The paper's subject area is clinical NLP, so 'cs.CV' appears to be an unusual classification, though that is not a reason to reject the work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,  \nRead this with the reader's take in front of me, and I largely agree with the stress-test note. The TXIT comparison is the paper's real contribution. GPT-5 at 92.8% (range 92.3–93.0) versus GPT-4 at 78.8% and GPT-3.5 at 62.1% is a striking, plausible gain, and the authors reproduced the baselines under their own API pipeline with a constrained response format. The five-run reproducibility and the item-level scoring make that part credible. Minor quibble: GPT-5 also took the seven image questions, which the baselines couldn't, so the comparison is slightly conservative in favour of the baselines. That's the kind of problem you want to have.  \n\nThe 60-case vignette set is new and could be a useful benchmark asset, but the ratings as analysed don't support the conclusions drawn. Fleiss' κ = 0.083 for correctness, −0.016 for comprehensiveness, and −0.111 for hallucinations means the four board-certified raters were essentially working from different rubrics. Reporting means of those ratings as if they were measurements is questionable. The hallucination rate of 10% is an average of flags across cases, and with negative agreement, \"no case reached majority consensus\" tells you nothing about the actual rate. The authors acknowledge the low κ but then proceed to interpret the aggregate numbers and subgroup patterns without correcting for it. There's no human baseline on the same cases, no consensus/adjudication, and no explicit definition of \"hallucination\" in Methods. That's a load-bearing flaw, not a cosmetic one.  \n\nWhat the paper does well: the TXIT benchmark is reproducible, the prompting is transparent, and the limitations section is mostly honest. The vignette set itself has good coverage and the examples give a concrete feel for the errors. It would be a stronger paper if the real-world claim were scaled back and the vignette ratings were re-analysed with a consensus approach or at least reported as the range of individual raters rather than as a single mean.  \n\nWho should read this: anyone tracking LLM performance in clinical decision support, and radiation oncology researchers thinking about how to evaluate models for tumor-board drafting. It deserves a serious referee, but the referee should insist on either a proper inter-rater analysis (e.g., rating a subset by consensus, then measuring agreement) or a substantially softened conclusion about the vignette results. I'd happily cite the TXIT numbers; I'd be careful about citing the hallucination rates.","headline":"The TXIT comparison is solid and timely; the vignette ratings, as analysed, are not reliable enough to support the real-world conclusions.","tokens_in":14696,"tokens_out":3270,"would_cite":true,"duration_ms":35940,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-5 reaches 92.8% on a radiation oncology in-training exam and earns high expert ratings on real treatment plans, while errors in complex cases still demand expert oversight.","keywords":["GPT-5","radiation oncology","large language models","treatment recommendation","hallucination","in-training exam","clinical decision support","expert evaluation"],"falsifier":"Re-run the same 60 vignettes through an independent panel of four board-certified radiation oncologists using the same rubric. If the new mean correctness falls outside the reported 3.11–3.38 confidence interval, or if a majority of the new panel flags hallucinations in any case, the vignette conclusion fails. On the exam side, run GPT-5 on a different year of the in-training exam under the same fixed prompt; an accuracy much below 92.8% would show the result is exam-specific rather than a general capability.","tokens_in":13732,"feed_emoji":"🩻","tokens_out":9693,"duration_ms":106249,"temperature":0.7,"pith_summary":"GPT-5, the latest reasoning-oriented large language model, is being marketed for oncology use, and this paper asks whether it can actually support radiation oncology decisions. On 293 multiple-choice questions from the 2021 in-training examination (TXIT), it reports a mean accuracy of 92.8%, versus 78.8% for GPT-4 and 62.1% for GPT-3.5 under the same protocol. On 60 real-patient vignettes, four board-certified radiation oncologists rated GPT-5's treatment plans 3.24/4 for correctness and 3.59/4 for comprehensiveness, with hallucinations flagged in 10% of ratings and no case attracting a majority hallucination flag. The paper's conclusion is that GPT-5 is a credible supervised assistant for education, pre-board preparation, and tumor-board draft generation, but not for autonomous clinical decision-making.","feed_headline":"GPT-5 hits 92.8% on radiation oncology exam","feed_subtitle":"Four experts rated GPT-5's treatment plans highly, but complex cases still need clinician sign-off.","key_machinery":"Two complementary instruments carry the argument. The first is TXIT, a 300-item multiple-choice in-training exam used as a knowledge benchmark; items are mapped to knowledge domains and clinical care-path categories and scored with a fixed API prompt over five runs, with the same prompts and item pool applied to GPT-3.5, GPT-4, and GPT-5 so that score gaps can be attributed to model capability. The second is a new set of 60 anonymized real-patient vignettes balanced across six tumor sites; GPT-5 is prompted to produce a structured therapeutic plan, and four board-certified radiation oncologists independently rate correctness and comprehensiveness on 4-point scales and flag hallucinations. In","core_discovery":"The central claim is that GPT-5 has crossed a practical threshold in radiation oncology knowledge while still falling short of independent clinical reliability. Under an item pool and scoring identical to prior model tests, GPT-5 reached 92.8% mean accuracy on the TXIT exam, with the largest domain gains in dose specification and diagnosis and persistent weaknesses in gynecology, brachytherapy, dosimetry, and trial-specific details; on seven image-based questions it answered only two correctly. On the 60-vignette benchmark, its structured treatment plans were rated highly for comprehensiveness (3.59/4) and moderately high for correctness (3.24/4), with hallucinations rare and never endorsed","pith_inferences":["Editorial inference: because the four raters agreed on correctness only at a near-chance level (kappa 0.083), the 3.24/4 mean is a panel-specific estimate; a different expert panel might move it meaningfully, and the 10% hallucination figure should be read as fragile.","Editorial inference: the 92.8% exam score likely overstates bedside competence; the 60-case ratings, especially the low-scoring rectal/anal and lung re-irradiation subgroups, are the more honest measure of real-world utility.","Editorial inference: a natural testable extension is to let the model abstain or state confidence on ambiguous cases and see whether expert correctness ratings rise; another is to feed it the same 60 vignettes with live guideline/trial lookup and measure how many of the flagged errors disappear.","Editorial inference: the low agreement may partly reflect that the Likert rubric cannot separate 'reasonable alternative plan' from 'wrong plan'; a consensus-based outcome such as 'would this plan pass tumor-board review' might yield higher reliability and clearer results."],"forward_implications":["GPT-5 could be used to generate first-draft tumor-board plans and study materials, provided a clinician reviews and corrects them before use.","The remaining error clusters define a concrete checklist for supervision: verify trial-specific evidence, dose/fractionation details, multimodal sequencing, and biomarker-driven decisions.","Because hallucinations were rare but nonzero, the practical risk shifts from fabricated content to subtle guideline mismatches that a non-expert might miss.","On exam-style knowledge, GPT-5 has reduced but not eliminated domain gaps; gynecologic oncology, brachytherapy, and dosimetry remain the places to look for wrong answers.","Without tool use or external retrieval, the model's performance on evolving trial knowledge is capped; adding live access to guidelines would be the natural next test."],"supporting_citations":[{"why":"Provides the GPT-3.5/GPT-4 TXIT baselines and the exact 293-item protocol that GPT-5's 92.8% is compared against.","marker":"39"},{"why":"Defines the model family under test as a reasoning-oriented, mixture-of-experts system, supporting the attribution of exam gains to training and architecture.","marker":"37"},{"why":"Supplies the TXIT content-analysis framework used to split results by knowledge domain and care-path category.","marker":"40"},{"why":"Offers the specialist-level oncology assistant comparison that motivates evaluating GPT-5 on real, multi-disease vignettes rather than synthetic cases.","marker":"33"},{"why":"Shows LLM performance on radiation oncology physics depends on item structure, supporting the need for identical item formats in the comparison.","marker":"29"},{"why":"Shows answer-option shuffling changes LLM accuracy, reinforcing the claim that GPT-5's improvement is not a format artifact.","marker":"30"},{"why":"Review evidence that oncology LLM outputs require human oversight and vary by topic, used to frame GPT-5's supervised role.","marker":"17"},{"why":"Meta-analysis on LLM integrations in cancer decision-making, used to position the expected safe use as supervised draft generation with retrieval.","marker":"18"}],"fun_headline_variants":["GPT-5 scores 92.8% on radiation oncology exam, but needs oversight","GPT-5 tops radiation oncology exam, yet complex cases still stump it","GPT-5 improves on radiotherapy exam, but expert sign-off still required","GPT-5's radiation oncology gains are real, but so are its errors"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The vignette conclusions assume that averaging four doctors' scores yields a stable measure of treatment-plan quality; since the doctors barely agreed with each other (kappa 0.083), those averages could look different with another panel.","fun_headline_variants_meta":{"raw":{"variants":["GPT-5 scores 92.8% on radiation oncology exam, but needs oversight","GPT-5 tops radiation oncology exam, yet complex cases still stump it","GPT-5 improves on radiotherapy exam, but expert sign-off still required","GPT-5's radiation oncology gains are real, but so are its errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000601,"raw_usage":{"total_tokens":2722,"prompt_tokens":898,"completion_tokens":1824,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":1741}},"tokens_in":642,"tokens_out":1824,"duration_ms":15014,"temperature":1.0,"reasoning_tokens":1741,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:56:00.591171+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same 60 vignettes through an independent panel of four board-certified radiation oncologists using the same rubric. If the new mean correctness falls outside the reported 3.11–3.38 confidence interval, or if a majority of the new panel flags hallucinations in any case, the vignette conclusion fails. On the exam side, run GPT-5 on a different year of the in-training exam under the same fixed prompt; an accuracy much below 92.8% would show the result is exam-specific rather than a general capability.","supporting_citations":[],"review_version":1}