{"id":"21efeeec-d4da-45fa-a5dd-1b2b7b6748fe","arxiv_id":"2608.09142","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A domain-adapted 8B LLM with agentic guideline retrieval matched UF Health oncologists' ratings on correctness, currency, and safety for colorectal cancer treatment plans in a 79-case blinded evaluation.","lead":"GatorOnco is an 8-billion-parameter language model tailored to colorectal cancer treatment planning using 282 billion tokens of biomedical text and an agentic retrieval system that pulls current NCCN guidelines into its reasoning. In a blinded test with five UF Health oncologists, it was rated on par with expert-written plans for correctness, currency, and safety, and higher for readability and completeness.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The expert-level claim rests on a same-institution gold standard annotated and rated by overlapping UF oncologists; parity with that reference is not independent evidence of expert-level care.","rationale":"The reader's weakest-assumption identification is accurate, and my read reinforces it rather than introducing a new objection. The paper's engineering contributions are real: a large de-identified corpus, exclusion of CRC-cohort notes from pre-training to prevent leakage, a held-out test set, a mixed-effects model, and explicit reasoning traces. Those elements support the internal claim that GatorOnco matches the UF annotators' own judgments on this cohort. What they do not support is the general claim of expert-level treatment planning. The test set is from the same institution, the gold-standard plans were produced or edited by the same oncologist group, and the author contributions show direct overlap between annotators and raters (TJ and TL). Because the model's SFT and RL objective was to reproduce these experts' annotations, high agreement with their edited notes could reflect alignment with a specific institutional template and a specific set of preferences. Blinding mitigates but cannot eliminate recognition: raters may identify UF notes or their own phrasing, and no post-hoc recognition probe is reported. The statistical framing adds a second weakness: equivalence is asserted from non-significant P-values without a non-inferiority margin or confidence intervals, so even on the internal data the conclusion 'expert-level' is stronger than the analysis supports. I would therefore keep the reader's CONDITIONAL verdict, with the condition being an external, independent evaluation and a proper non-inferiority analysis. This is not a finding of misconduct or a claim that the model is unsafe; it is a statement that the central evidentiary link, the reference standard, is not yet independent enough to carry the headline claim.","tokens_in":13684,"tokens_out":7259,"duration_ms":66296,"concrete_test":"Conduct a preregistered external validation: recruit 3-5 oncologists from two or more outside institutions with no connection to UF Health, blind them to model identity, and have them rate GatorOnco and de-identified UF gold-standard plans for the same 79 cases on the same five dimensions. If GatorOnco no longer shows non-inferiority within a predefined margin (e.g., 0.5 on the 5-point scale) on correctness and safety, the expert-level claim is not supported. As a supplementary sensitivity check on the existing data, re-run the linear mixed-effects model excluding the two raters who also annotated cases and report whether the parity result survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GatorOnco is expert-level for CRC treatment planning depends on the validity of the comparison reference. The reference plans are not independent clinical ground truth: they were reviewed and, when incomplete or erroneous, edited by UF Health oncologists (Methods, 'Expert annotation and quality control'), and the author contributions show that two of the five human raters (TJ, TL) also participated in case annotation. GatorOnco was trained through SFT and RL to reproduce the same experts' annotations, using a hierarchical reward co-designed with UF oncologists. The blind protocol prevents raters from knowing which plans are AI-generated, but it does not prevent recognition of familiar UF note templates or one's own or colleagues' notes, and no recognition check is reported. Parity with such a reference is therefore partly a measure of fit to the training signal and institutional style rather than an independent demonstration of expert-level care. In addition, 'statistically comparable' is inferred from non-significant P-values (correctness 4.09 vs 4.11, P=0.921; safety 4.22 vs 4.22, P=0.999) without a pre-specified non-inferiority margin or confidence intervals, so equivalence is not established. A single-institution, 79-case design with overlapping annotator and rater roles cannot support the general claim as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"GatorOnco is an 8B-parameter Llama-3.1-based LLM adapted to colorectal cancer treatment planning through continued pretraining on a 282B-token biomedical corpus (166B tokens from UF Health EHRs), model merging, supervised fine-tuning, and agentic reinforcement learning with interleaved retrieval of NCCN guidelines. The model is evaluated on 79 held-out UF Health CRC cases for treatment-category determination (F1 up to 0.924) and narrative plan generation using automatic lexical, semantic, and faithfulness metrics, and in a blinded human evaluation by five UF Health oncologists that compared GatorOnco, Llama-3.1-8B-Instruct, and expert gold-standard plans across five Likert dimensions. The authors report that GatorOnco matches expert oncologists on correctness, currency, and safety and exceeds them on readability and completeness.","tokens_in":13930,"tokens_out":7759,"duration_ms":69417,"significance":"If the results held as stated, GatorOnco would be an important demonstration that a compact domain-adapted model with agentic retrieval can approach expert-level treatment planning. The work combines an unusually large clinical corpus, a structured reward hierarchy that enforces modality-to-regimen reasoning, and a stratified held-out test design with appropriate mixed-effects and aligned-rank-transform analyses. The authors also commit to releasing code and a de-identified test set. However, the headline 'expert-level' claim is currently supported by a same-institution reference standard rated in part by oncologists who also contributed to the training annotations, and the parity conclusion is drawn from non-significant p-values without an equivalence margin. These issues do not undermine the engineering contributions or the automatic-evaluation results, but they do constrain how the central claim can be stated.","major_comments":[{"comment":"The central expert-level claim rests on a reference standard that is not independent of the model's training signal. The gold-standard plans were written or edited by UF Health oncologists (Methods, 'Expert annotation and quality control'), the model was trained with SFT and RL to reproduce those same annotations, and two of the five blind raters (TJ, TL) are listed both as case annotators and as human evaluators in the Author Contributions. Consequently, parity with the gold standard could reflect raters' familiarity with their own or colleagues' notes and with UF-specific documentation conventions rather than independent evidence of expert-level care. The blind protocol prevents raters from knowing which system produced a plan, but no check is reported for whether raters recognized the source of ground-truth notes. Please report such a recognition check, re-analyze the primary comparisons excluding raters who participated in annotation, or obtain independent external oncologist ratings; at minimum, the 'expert-level' claim should be rephrased to reflect that the comparison is to institutional expert annotations.","section":"Methods, 'Expert annotation and quality control' and 'Blind evaluation by UF Health Oncologists'; Author contributions"},{"comment":"The conclusion that GatorOnco is 'statistically comparable' or 'on par' with experts for correctness (4.09 vs. 4.11, P = 0.921), currency (4.04 vs. 3.98, P = 0.478), and safety (4.22 vs. 4.22, P = 0.999) is inferred entirely from non-significant p-values. This does not establish equivalence: a large confidence interval or a low-powered test can also yield P > 0.05. The manuscript provides neither a pre-specified non-inferiority margin nor two one-sided tests (TOST) or confidence intervals for the differences. In addition, the composite mixed-model contrast shows GatorOnco significantly higher than experts (β = 0.136, P < 0.01), which is inconsistent with a pure parity claim. Please add equivalence testing with an explicit margin and report confidence intervals for the dimension-specific differences.","section":"Results, 'GatorOnco achieved expert-level treatment planning for CRC...' and Methods, 'Blind evaluation by UF Health…"},{"comment":"The human evaluation compares only GatorOnco, Llama-3.1-8B-Instruct, and the expert ground truth; MedGemma-27B and Llama-3.3-70B-Instruct were evaluated only with automatic metrics. The abstract's claim that GatorOnco 'significantly outperformed open-source LLMs' is therefore supported for only a single human-evaluated baseline, not for the class of open-source LLMs. Either include the additional baselines in the blind human evaluation or restrict the claim to the parameter-matched Llama-3.1-8B model.","section":"Methods, 'Blind evaluation by UF Health Oncologists'; Abstract"},{"comment":"The abstract and title claim expert-level treatment planning without institutional or temporal qualification, yet the evaluation is a single-institution, single-disease, 79-case study with a reference standard that was edited by the same group to conform to contemporaneous NCCN guidelines. The Limitations section acknowledges the single-institution scope, but the abstract should carry the same qualification, for example by stating 'expert-level on this institutional test set' or 'comparable to UF Health oncologists' rather than the unqualified 'expert-level treatment planning'.","section":"Discussion (Limitations) and Abstract"}],"minor_comments":[{"comment":"The phrase 'batter ratings' should be 'better ratings'.","section":"Results, 'GatorOnco achieved expert-level treatment planning for CRC...'"},{"comment":"The GRPO loss equation does not render in the text, and the reward formula is garbled (e.g., the definition of R and the Dice coefficient term are not readable); please provide a clean mathematical formulation, as the reward design is central to the method's reproducibility.","section":"Methods, 'Optimization Objective and Hierarchical Reward'"},{"comment":"The 'Overall' score is described as the arithmetic mean of the metrics, but it is not stated whether all seven metrics are weighted equally; please specify the aggregation rule explicitly.","section":"Table 2 and Methods, 'Generating narrative treatment plan sections'"},{"comment":"The overlap between annotators (TJ, TL, TG, LE) and raters (TJ, TL, LG, CS, OM) is not disclosed in the Methods or Limitations; please add an explicit disclosure and discuss its implications for the blind evaluation.","section":"Author Contributions and Methods"},{"comment":"The code availability section refers to 'GatorTronGPT' training code, while the model is called GatorTronLlama; please align the model names to avoid confusion.","section":"Code availability"},{"comment":"The phrase 'statistically comparable' should be replaced with 'not significantly different' unless equivalence tests with confidence intervals are added, to avoid a common statistical misinterpretation.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering contribution and likely of interest to the journal, but the framing of the result as 'expert-level' is ahead of the evidence. The overlap between annotators and raters should be transparent in the main text, not only inferable from the author contributions. I would encourage the editor to require either external or non-overlapping rater validation, or a substantial softening of the central claim. If external validation is not feasible, a revised version that frames the contribution as 'matching institutional expert annotations' could be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core contribution here is genuinely useful: an 8B model that, after domain adaptation on 282B tokens (including UF Health's clinical text), model merging, SFT, GRPO with a hierarchical clinical reward, and agentic RAG over NCCN guidelines, beats much larger general-purpose and medical LLMs on automatic metrics for CRC treatment planning. The human evaluation is also a real effort: five oncologists, 79 cases, blinded, mixed-effects modeling. The completeness and readability gains over expert-written notes are credible and consistent with prior work on LLM clarity.\n\nThe soft spots are exactly where the reader's stress test lands. The strongest claim—expert-level parity on correctness, currency, and safety—is built on a comparison against a gold standard written or edited by UF oncologists, and two of the five raters were among the annotators. The blind protocol hides AI vs. human authorship, but it cannot hide institutional note style or one's own prose, and no recognition check is reported. So the headline conclusion overreaches: this shows GatorOnco matches its own training signal and institutional conventions, not that it matches an independent standard of care. The \"statistically comparable\" conclusion is also just non-significance, with no pre-specified non-inferiority margin or confidence intervals. And only one baseline (Llama-3.1-8B) went into the human eval; the 70B and MedGemma comparisons are automatic only.\n\nNone of that makes the paper weak. The engineering is careful, the automatic gains are large and consistent across multiple metrics, and the limitations section is honest about single-institution, single-disease scope and prototype status. The missing ablation isolating the agentic RAG or the RL is a minor complaint; the integrated system is the contribution. The major gap is external validation and a cleaner human benchmark, which is exactly what peer review should push for.\n\nThis paper deserves a serious referee. It's a meaningful advance for clinical NLP and a good example of what domain adaptation can do with a modest model. I'd cite it for the training recipe and the evaluation framework, though not for the parity claim as stated.","headline":"A serious, well-engineered clinical LLM paper whose expert-level claim is real but rests on a same-institution gold standard and overlapping evaluators; worth refereeing, not desk-rejecting.","tokens_in":14551,"tokens_out":1539,"would_cite":true,"duration_ms":16780,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A domain-adapted 8-billion-parameter language model generates colorectal cancer treatment plans that blind oncologists rate as statistically indistinguishable from expert-written plans.","keywords":["colorectal cancer","treatment planning","agentic LLM","retrieval-augmented generation","reinforcement learning","domain adaptation","NCCN guidelines","clinical evaluation"],"falsifier":"A cross-institution study in which treatment plans from GatorOnco, from oncologists at a different health system, and from a nationally reviewed gold standard are rated by an independent panel unaware of the source would settle the claim: if GatorOnco's correctness or safety ratings fall significantly below the expert plans, the expert-level claim is refuted. A simpler check is to remove UF Health clinical notes from the pre-training corpus and see whether the parity advantage disappears, which would indicate dependence on the specific training data.","tokens_in":13472,"feed_emoji":"🩺","tokens_out":5701,"duration_ms":46936,"temperature":0.7,"pith_summary":"GatorOnco is an 8-billion-parameter language model adapted to oncology that generates treatment plans for colorectal cancer. The paper's core claim is that, in a blind evaluation by five oncologists, GatorOnco's plans are rated statistically on par with expert-written plans for correctness, currency, and safety, and higher for readability and completeness, while outperforming much larger open-source models. The significance would be that specialized, locally deployable models, rather than massive general-purpose ones, can meet expert-level standards in a high-stakes clinical task, provided they are grounded in current guidelines through agentic retrieval. The evidence comes from 79 real patient cases at a single health system.","feed_headline":"8B AI model's cancer treatment plans rated on par with oncologists","feed_subtitle":"A domain-adapted LLM with agentic guideline retrieval matches expert ratings on correctness, currency, and safety.","key_machinery":"The central mechanism is an agentic retrieval-augmented generation loop in which the model, after continued pre-training and instruction-following fusion via model merging, is optimized with group relative policy optimization under a hierarchical clinical reward. During generation, the model emits a structured search action when it has an information gap, the environment appends top-ranked chunks from an indexed NCCN guideline database, and the model resumes its chain-of-thought before writing the plan. The hierarchical reward first requires the correct high-level therapy modality and only then rewards correct drug-entity alignment, so that a plan is never rewarded if the wrong modality is chosen. This design grounds reasoning in time-sensitive guidelines and enforces a safety-first ordering of clinical decisions.","core_discovery":"On the paper's own terms, GatorOnco achieves expert-level treatment planning for colorectal cancer as judged by the same institution's oncologists. In the blinded comparison, GatorOnco scored 4.09 versus 4.11 for correctness (P = 0.921), 4.04 versus 3.98 for currency (P = 0.478), and 4.22 versus 4.22 for safety (P = 0.999); it exceeded oncologists on readability (4.46 vs 4.19) and completeness (3.91 vs 3.52). The model also outperformed Llama-3.3-70B and MedGemma-27B on treatment-option determination and narrative plan quality. The paper interprets this as evidence that domain adaptation on real clinical text, combined with safety-oriented reinforcement learning and agentic retrieval of NCCN guidelines, can close the gap between generative AI and expert performance in cancer treatment planning.","pith_inferences":["The same pipeline applied to other cancer types and other health systems would test whether the expert-level parity is due to the method or to the single-institution training and rating context.","Because raters and ground-truth writers belong to the same institution, the parity scores may partly reflect shared documentation style; a cross-institutional panel could separate house style from clinical quality.","The authors' call for machine-readable guidelines (e.g., FHIR) points to a concrete engineering bottleneck: manual parsing of PDF guidelines is substantial, so any shift to structured guideline releases could accelerate deployment of such agents.","The hierarchical reward design could transfer to other safety-critical generation tasks, where fine-grained rewards are gated behind coarse categorical correctness."],"forward_implications":["If the central claim is correct, an 8-billion-parameter model with agentic retrieval can match expert oncologists on treatment planning, implying that scale is not the main bottleneck for high-stakes clinical text generation.","The full pipeline (continued pre-training, model merging, two-stage post-training, and agentic RL with hierarchical rewards) becomes a reusable recipe for other disease domains.","Because the agent retrieves from current NCCN guideline versions, treatment plans can be audited by linking each recommendation to specific guideline chunks and patient-specific evidence.","The reported readability and completeness advantages over expert notes suggest AI-generated plans may serve as drafting aids that reduce documentation burden.","The authors state that GatorOnco remains a prototype; practical deployment would require fail-safes, uncertainty quantification, and deferral policies."],"supporting_citations":[{"why":"Supplies the base Meta-Llama-3.1-8B model that is continually pre-trained and instruction-tuned.","marker":"[18]"},{"why":"Provides the evolutionary model-merging method that combines clinical knowledge with instruction-following ability.","marker":"[20]"},{"why":"Defines GRPO, the reinforcement learning algorithm used to optimize the agentic policy.","marker":"[21]"},{"why":"Search-R1 framework implements the interleaved retrieval-generation agent used for guideline grounding.","marker":"[44]"},{"why":"FineWeb general-domain corpus that, with the UF Health clinical corpus, forms the 282-billion-token pre-training set.","marker":"[35]"},{"why":"E5-large-v2 text embedder used to index NCCN guideline chunks for retrieval.","marker":"[41]"},{"why":"FAISS vector database that stores and retrieves guideline chunks during agentic search.","marker":"[42]"},{"why":"GPT-oss-120B distills stepwise reasoning trajectories used in post-training data construction.","marker":"[43]"}],"fun_headline_variants":["Agentic LLM matches oncologists on colorectal cancer plans","Cancer AI plans judged as safe and correct as oncologists","GatorOnco matches expert oncologists in CRC treatment plans","AI colorectal cancer planner ties expert oncologists"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the oncologist-written or oncologist-edited treatment plans from UF Health are a valid and unbiased gold standard, and that the same institution's oncologists, several of whom helped annotate the training and test cases, can rate plans blindly; if the gold-standard notes are not the true standard of care or raters recognize their own or colleagues' phrasing, then statistical parity with those notes does not prove expert-level performance.","fun_headline_variants_meta":{"raw":{"variants":["Agentic LLM matches oncologists on colorectal cancer plans","Cancer AI plans judged as safe and correct as oncologists","GatorOnco matches expert oncologists in CRC treatment plans","AI colorectal cancer planner ties expert oncologists"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000811,"raw_usage":{"total_tokens":3624,"prompt_tokens":1081,"completion_tokens":2543,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":697,"completion_tokens_details":{"reasoning_tokens":2478}},"tokens_in":697,"tokens_out":2543,"duration_ms":15326,"temperature":1.0,"reasoning_tokens":2478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:39:57.776923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A cross-institution study in which treatment plans from GatorOnco, from oncologists at a different health system, and from a nationally reviewed gold standard are rated by an independent panel unaware of the source would settle the claim: if GatorOnco's correctness or safety ratings fall significantly below the expert plans, the expert-level claim is refuted. A simpler check is to remove UF Health clinical notes from the pre-training corpus and see whether the parity advantage disappears, which would indicate dependence on the specific training data.","supporting_citations":[],"review_version":1}