Pith. sign in

REVIEW 4 major objections 6 minor

Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read GPT-5 reaches 92.8% on a radiation oncology in-training exam and earns high expert ratings on real treatment plans, while errors in complex cases still demand expert oversight.

desk verdict The TXIT comparison is solid and timely; the vignette ratings, as analysed, are not reliable enough to support the real-world conclusions. read the letter →

arxiv 2508.21777 v1 pith:2FPD7ZZX submitted 2025-08-29 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords GPT-5radiationoncologylargelanguagemodelstreatmentrecommendationhallucinationin-trainingexamclinicaldecisionsupportexpertevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

GPT-5, the latest reasoning-oriented large language model, is being marketed for oncology use, and this paper asks whether it can actually support radiation oncology decisions. On 293 multiple-choice questions from the 2021 in-training examination (TXIT), it reports a mean accuracy of 92.8%, versus 78.8% for GPT-4 and 62.1% for GPT-3.5 under the same protocol. On 60 real-patient vignettes, four board-certified radiation oncologists rated GPT-5's treatment plans 3.24/4 for correctness and 3.59/4 for comprehensiveness, with hallucinations flagged in 10% of ratings and no case attracting a majority hallucination flag. The paper's conclusion is that GPT-5 is a credible supervised assistant for education, pre-board preparation, and tumor-board draft generation, but not for autonomous clinical decision-making.

What carries the argument

Two complementary instruments carry the argument. The first is TXIT, a 300-item multiple-choice in-training exam used as a knowledge benchmark; items are mapped to knowledge domains and clinical care-path categories and scored with a fixed API prompt over five runs, with the same prompts and item pool applied to GPT-3.5, GPT-4, and GPT-5 so that score gaps can be attributed to model capability. The second is a new set of 60 anonymized real-patient vignettes balanced across six tumor sites; GPT-5 is prompted to produce a structured therapeutic plan, and four board-certified radiation oncologists independently rate correctness and comprehensiveness on 4-point scales and flag hallucinations. In

What would settle it

Re-run the same 60 vignettes through an independent panel of four board-certified radiation oncologists using the same rubric. If the new mean correctness falls outside the reported 3.11–3.38 confidence interval, or if a majority of the new panel flags hallucinations in any case, the vignette conclusion fails. On the exam side, run GPT-5 on a different year of the in-training exam under the same fixed prompt; an accuracy much below 92.8% would show the result is exam-specific rather than a general capability.

Watch

Extended reading notes

Core claim

The central claim is that GPT-5 has crossed a practical threshold in radiation oncology knowledge while still falling short of independent clinical reliability. Under an item pool and scoring identical to prior model tests, GPT-5 reached 92.8% mean accuracy on the TXIT exam, with the largest domain gains in dose specification and diagnosis and persistent weaknesses in gynecology, brachytherapy, dosimetry, and trial-specific details; on seven image-based questions it answered only two correctly. On the 60-vignette benchmark, its structured treatment plans were rated highly for comprehensiveness (3.59/4) and moderately high for correctness (3.24/4), with hallucinations rare and never endorsed

Load-bearing premise

The vignette conclusions assume that averaging four doctors' scores yields a stable measure of treatment-plan quality; since the doctors barely agreed with each other (kappa 0.083), those averages could look different with another panel.

Editorial extensions

If this is right

  • GPT-5 could be used to generate first-draft tumor-board plans and study materials, provided a clinician reviews and corrects them before use.
  • The remaining error clusters define a concrete checklist for supervision: verify trial-specific evidence, dose/fractionation details, multimodal sequencing, and biomarker-driven decisions.
  • Because hallucinations were rare but nonzero, the practical risk shifts from fabricated content to subtle guideline mismatches that a non-expert might miss.
  • On exam-style knowledge, GPT-5 has reduced but not eliminated domain gaps; gynecologic oncology, brachytherapy, and dosimetry remain the places to look for wrong answers.
  • Without tool use or external retrieval, the model's performance on evolving trial knowledge is capped; adding live access to guidelines would be the natural next test.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because the four raters agreed on correctness only at a near-chance level (kappa 0.083), the 3.24/4 mean is a panel-specific estimate; a different expert panel might move it meaningfully, and the 10% hallucination figure should be read as fragile.
  • Editorial inference: the 92.8% exam score likely overstates bedside competence; the 60-case ratings, especially the low-scoring rectal/anal and lung re-irradiation subgroups, are the more honest measure of real-world utility.
  • Editorial inference: a natural testable extension is to let the model abstain or state confidence on ambiguous cases and see whether expert correctness ratings rise; another is to feed it the same 60 vignettes with live guideline/trial lookup and measure how many of the flagged errors disappear.
  • Editorial inference: the low agreement may partly reflect that the Likert rubric cannot separate 'reasonable alternative plan' from 'wrong plan'; a consensus-based outcome such as 'would this plan pass tumor-board review' might yield higher reliability and clearer results.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports a benchmark of GPT-5 in radiation oncology using two evaluation settings: (i) the ACR Radiation Oncology In-Training Examination (TXIT) 2021, where GPT-5 achieved a mean accuracy of 92.8% over five runs, compared with published GPT-4 (78.8%) and GPT-3.5 (62.1%) baselines; and (ii) 60 real-world clinical vignettes rated by four board-certified radiation oncologists for correctness (mean 3.24/4), comprehensiveness (mean 3.59/4), and hallucinations (10% of ratings). The authors report low inter-rater reliability (Fleiss' kappa = 0.083 for correctness, -0.016 for comprehensiveness, -0.111 for hallucinations) but nevertheless conclude that GPT-5 can generate coherent, comprehensive drafts for supervised use in education, pre-board preparation, and tumor-board draft generation, while requiring expert oversight.

Significance. If the claims are substantiated, this would be one of the first comprehensive evaluations of GPT-5 in radiation oncology and would provide a useful comparison point for clinical LLM benchmarks. The TXIT part is valuable because it uses a public exam and a reproducible API pipeline, and the five-run range is narrow. The vignette dataset is also a worthwhile resource, covering a broad set of real-world disease sites and treatment intents. The paper ships prompts, seeds, and raw outputs in the supplementary material, which is a strength for reproducibility. However, the vignette-based central claim is currently undermined by the reported lack of inter-rater agreement, and the TXIT comparison needs clarification regarding the item denominator and statistical reporting.

major comments (4)
  1. [§2.2/§3.1] The GPT-5 comparison is asymmetric: GPT-5 was run on the full 300-item TXIT including seven image-based questions, whereas GPT-4/GPT-3.5 baselines were computed on 293 text-only items. The reported mean of 92.8% appears to be computed over the 293 text items, because adding the 7 image items (2/7 correct) would lower the overall accuracy to ~91.3%. Please state explicitly which denominator was used. If GPT-5 is compared on 293 text items, the comparison is fair but must be stated; if it is compared on 300 items, the baselines need to be recomputed on the same item pool. Also report the SD/CI for the five runs and a test (e.g., bootstrap or McNemar) that the difference vs. GPT-4 is significant.
  2. [§3.4/Table 2] The vignette results rest on ratings with essentially no inter-rater reliability: Fleiss' kappa = 0.083 for correctness, -0.016 for comprehensiveness, -0.111 for hallucinations. The reported means (3.24, 3.59) and the 10% hallucination rate are averages of ratings whose reproducibility is not demonstrated. No adjudicated reference standard, consensus procedure, or per-rater agreement analysis is provided. The statement that no case reached majority consensus for hallucinations is expected when kappa is near zero and cannot be used as evidence of rarity. Please either add a consensus/adjudication step and report agreement on final labels, or weaken the vignette conclusions to descriptive exploratory findings.
  3. [§4/§6] The conclusion 'Hallucinations were rare and not a substantive concern' is not supported: 24/240 ratings (10%) flagged hallucinations, with 40% of cases receiving at least one flag and no case reaching majority. Given the negative kappa for hallucination flags, this could reflect idiosyncratic rater behavior rather than a low true hallucination rate. At minimum, define 'hallucination' operationally, provide examples of flagged outputs, and report the rate under a stricter/consensus definition.
  4. [§3.4] No human baseline is reported for the vignette task. Without clinician performance on the same 60 cases, the practical meaning of a mean correctness of 3.24/4 cannot be interpreted. A small panel of additional radiation oncologists rating the same vignettes (or a subset) would also help establish whether the low reliability is in the rubric or in the outputs.
minor comments (6)
  1. [§3.1] For GPT-5 only the range is reported without SD, while Figure 1's caption says error bars show SD. Add the SD for the five runs.
  2. [§2.2] The item-count description is confusing: 'Fourteen questions included medical images, of which seven required visual interpretation.' Clarify whether the other seven image questions were text-scorable and whether they were included in the 293-item baseline.
  3. [Abstract] The phrase 'no case reaching majority consensus for their presence' is misleading in the abstract, since low kappa makes exactly this pattern likely by chance.
  4. [§2.4] The text says exploratory subgroup analyses were prespecified, but no analysis plan or protocol is given. Specify whether subgroups were defined before seeing the data and report the number of comparisons.
  5. [Supplementary Material] The prompt listing includes a JSON schema and German-language instructions. Please confirm this is the exact prompt used for all cases and report decoding parameters (temperature, top-p, etc.), which are not given in §2.1.
  6. [§2.1] Typo: 'artifical intelligence' should be 'artificial intelligence'.

Circularity Check

0 steps flagged · score 1.0 of 10

Empirical benchmark with no fitted parameters or derivation chain; no circular reduction found.

full rationale

The paper is a direct empirical benchmark. The TXIT results are generated by running GPT-5, GPT-4, and GPT-3.5 through the same fixed API prompt and scoring rule (Sections 2.2 and 3.1); the prior-model baselines were independently reproduced in this study: 'Using the application programming interface (API) with a fixed prompt over five repeated runs, GPT-3.5 achieved 62.1% ± 1.1% and GPT-4 78.8% ± 0.9%, consistent with earlier reports (39)'. Thus the comparison does not reduce to a self-citation or to fitted inputs. The vignette evaluation is also observational: Likert ratings and hallucination flags are raw expert judgments summarized in Table 2. The low Fleiss' kappa values (0.083, -0.016, -0.111) are reported as limitations and concern the reliability of the measurement, but they do not make the summary statistics definitionally equivalent to the model's outputs. The self-citation to the group's prior GPT-4 TXIT work (ref 39) is contextual and non-load-bearing because the relevant baseline numbers are regenerated in this paper. No equation or quantity is defined in terms of the target claim, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is an empirical benchmark with no fitted parameters and no new theoretical entities. Its central claims rest on the validity of the TXIT exam, the reliability of expert ratings, the representativeness of the vignette sample, and the stability of the GPT-5 API model.

assumptions (4)
  • domain assumption The ACR TXIT 2021 item pool is a valid measure of radiation oncology knowledge.
    Used as the primary endpoint in Section 2.2 and interpreted as 'genuine advances in model capability' in Section 3.1.
  • domain assumption Aggregated Likert ratings by four board-certified radiation oncologists are a valid gold standard for treatment-plan correctness and comprehensiveness.
    Sections 2.3 and 3.4 use these ratings as ground truth despite low Fleiss' kappa.
  • domain assumption API interactions using 'GPT-5 family' configuration represent the GPT-5 model as marketed for oncology.
    Section 2.1 defines the model only by the system card, without a pinned version or release hash.
  • domain assumption The 60 vignettes are representative of real-world radiation oncology decision-making.
    Section 2.3 claims a broad spectrum and balance, but cases come from a single institution and may carry survivorship/documentation biases (acknowledged in Limitations).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight." pith.science (2026). https://pith.science/paper/2FPD7ZZX

@misc{pith2026250821777,
  author       = {Pith},
  title        = {Pith review of: Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2FPD7ZZX}},
  note         = {Machine review of arXiv:2508.21777}
}
read the original abstract

Introduction: Large language models (LLM) have shown great potential in clinical decision support. GPT-5 is a novel LLM system that has been specifically marketed towards oncology use. Methods: Performance was assessed using two complementary benchmarks: (i) the ACR Radiation Oncology In-Training Examination (TXIT, 2021), comprising 300 multiple-choice items, and (ii) a curated set of 60 authentic radiation oncologic vignettes representing diverse disease sites and treatment indications. For the vignette evaluation, GPT-5 was instructed to generate concise therapeutic plans. Four board-certified radiation oncologists rated correctness, comprehensiveness, and hallucinations. Inter-rater reliability was quantified using Fleiss' \k{appa}. Results: On the TXIT benchmark, GPT-5 achieved a mean accuracy of 92.8%, outperforming GPT-4 (78.8%) and GPT-3.5 (62.1%). Domain-specific gains were most pronounced in Dose and Diagnosis. In the vignette evaluation, GPT-5's treatment recommendations were rated highly for correctness (mean 3.24/4, 95% CI: 3.11-3.38) and comprehensiveness (3.59/4, 95% CI: 3.49-3.69). Hallucinations were rare with no case reaching majority consensus for their presence. Inter-rater agreement was low (Fleiss' \k{appa} 0.083 for correctness), reflecting inherent variability in clinical judgment. Errors clustered in complex scenarios requiring precise trial knowledge or detailed clinical adaptation. Discussion: GPT-5 clearly outperformed prior model variants on the radiation oncology multiple-choice benchmark. Although GPT-5 exhibited favorable performance in generating real-world radiation oncology treatment recommendations, correctness ratings indicate room for further improvement. While hallucinations were infrequent, the presence of substantive errors underscores that GPT-5-generated recommendations require rigorous expert oversight before clinical implementation.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.