{"id":"e26b81c1-35de-43cb-a97b-02e34c6e7d6a","arxiv_id":"2607.03931","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An open multi-cancer FDG-PET/CT lesion segmentation model trained on 1,563 scans outperforms public benchmarks on 185 external scans with fewer false positives and robust TTB/TLG.","lead":"GLOW-FDG is an open-source deep-learning model that segments cancer lesions on whole-body FDG-PET/CT across multiple cancer types. It beats public benchmarks on external scans by cutting false positives while keeping tumor-burden measures close to expert agreement.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged external-cohort and inter-observer limits.","rationale":"The reader’s strongest claim matches the abstract and Results §2.1–2.5. The weakest assumption correctly isolates the partial independence of the validation set (detection-only BAMF-rechecked cohorts, institutional melanoma split, n=10 dual-reader subset). After re-reading Materials & Methods §4.3–4.8, Tables 2–18, and the Discussion limitations, that remains the single load-bearing soft spot; no stronger technical flaw (e.g., metric definition error, unacknowledged train–test leakage beyond the disclosed SINERGIA split, or contradiction between detection and quantification results) is present. Open weights/code and public baselines further support soundness. Therefore the ACCEPT / MODERATE verdict needs no adjustment; the concrete check above simply stress-tests the claim under the strictest subset of the already-reported data.","tokens_in":22638,"tokens_out":612,"duration_ms":5112,"concrete_test":"On the three fully human-segmented external cohorts only (CHESS n=31, HECKTOR-USZ n=63, SINERGIA val n=20), recompute patient- and lesion-wise F1 and Dice for GLOW-FDG vs the three public baselines using the same one-to-one mask-matching rule (§4.7). If GLOW-FDG no longer ranks first on F1 (or the F1 gap shrinks below the reported 95% CI separation), the multi-cohort superiority claim weakens; otherwise the claim holds under the stricter independence filter.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that GLOW-FDG consistently leads public baselines on patient- and lesion-wise detection F1 by raising precision while keeping high recall, with robust TTB/TLG ICC and performance near dual-reader variability—is supported by the reported tables (Tables 2, 4–6, 13–18) and the disclosed evaluation design. The softest supporting condition is exactly the one the reader named: two of five external cohorts (QIN-Breast, ACRIN-NSCLC) are detection-only after BAMF mask re-check rather than full independent human re-segmentation (§4.3–4.4), SINERGIA melanoma is a same-institution train/val split, and the human-reference analysis uses only 10 melanoma cases with deliberately different PET- vs CT-centric styles (§4.8). Those constraints limit how far “generalizable clinical-grade multi-cancer” can be pushed, but they are already stated in the Discussion and do not invert the head-to-head gains on the fully human-segmented external cohorts (CHESS, HECKTOR-USZ, melanoma). No additional internal inconsistency or hidden circularity appears in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"GLOW-FDG is an open-source dual-head ResEncL nnU-Net for whole-body FDG-PET/CT cancer lesion segmentation, pretrained MultiTalent-style then fine-tuned on 1,563 multi-cancer scans with organ supervision and PET–CT misalignment augmentation. On 185 external scans across breast, nonmetastatic and oligometastatic lung, head and neck, and metastatic melanoma, it reports the highest patient- and lesion-wise detection F1 among three public baselines (AutoPET DKFZ, AutoPET IKIM, onlyPET), mainly by higher precision (fewer FPs) at high recall, with strong TTB/TLG ICC and Dice, and performance near dual-reader variability on a 10-case melanoma subset.","tokens_in":22958,"tokens_out":832,"duration_ms":6578,"significance":"If the external gains hold, this is a practically useful contribution: a released multi-cancer FDG-PET/CT lesion model with public weights/code, head-to-head comparison against recent AutoPET-class baselines, and clinically oriented metrics (lesion F1, TTB/TLG RPD/ARPD/ICC) rather than Dice alone. Strengths include open release, multi-institution external testing, bootstrap CIs, TP/FN/FP characterization, and an explicit inter-observer reference. The work addresses a real gap between challenge-trained models and broader multi-cancer whole-body use.","major_comments":[{"comment":"§4.3–4.4 and Results §2.1: Two of five external cohorts (QIN-Breast, ACRIN-NSCLC) support only detection after BAMF mask re-check, not full independent human re-segmentation; SINERGIA melanoma is a same-institution train/val split. The abstract’s “185 external scans from independent institutions” and “generalizable” framing should be tightened so claims rest primarily on fully human-segmented external cohorts (CHESS, HECKTOR-USZ, melanoma) and detection-only cohorts are labeled as such.","section":null},{"comment":"§2.5 and §4.8: Inter-observer context is limited to 10 metastatic melanoma cases with deliberately different PET- vs CT-centric styles. The claim that performance “approached the variability observed between expert radiation oncologists” is directionally useful but over-extended for multi-cancer clinical-grade generalization; restrict or qualify this comparison and avoid treating it as a multi-reader multi-cancer standard.","section":null}],"minor_comments":[{"comment":"Figure 3 caption: clarify that breast and nonmetastatic lung Dice are omitted because only detection labels were used after BAMF re-check.","section":null},{"comment":"Tables 7–9 vs 10–12: numbering and cross-references to TP/FN/FP distributions are slightly inconsistent in the text; align table numbers and in-text citations.","section":null},{"comment":"Discussion: ultra-low-dose and non-FDG limitations are appropriately noted; a short explicit statement on scanner/protocol diversity in the external set would help readers bound expected domain shift.","section":null},{"comment":"Minor typos and formatting: “publically avaliable,” “adress,” mixed SU V BW notation, and a few broken hyphenations (e.g., “H¨ ullner”) should be cleaned in production.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Central empirical claim is supported on the fully human-segmented external cohorts; the main risk is over-claiming generality from mixed detection-only and same-institution data. I would not require new multi-center multi-reader studies for acceptance if the abstract/discussion are tightened. Fit for a methods/imaging AI venue is good given open weights and clinically oriented metrics."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a careful, usable open model rather than a new architecture paper. GLOW-FDG is nnU-Net ResEncL with MultiTalent-style pretraining, organ multitask heads, and the misalignment augmentation already used in AutoPET III work. What is actually new is the curated multi-cancer training set (1,563 scans), the external multi-cohort evaluation against three named public baselines, the emphasis on precision/false positives and TTB/TLG (RPD/ARPD/ICC), and the public weights/code.\n\nThe head-to-head numbers hold up on the manuscript’s own terms. Across breast, non- and oligometastatic lung, head and neck, and melanoma, GLOW-FDG leads patient- and lesion-wise F1 mainly by raising precision while keeping high recall. Dice and especially TLG/TTB ICCs look strong on the fully human-segmented cohorts (CHESS, HECKTOR-USZ, melanoma). The dual-reader melanoma subset (n=10) is small and deliberately PET- vs CT-centric, but it is presented as context, not as a definitive human ceiling, and the model sits near that inter-observer range. Tables and bootstrap CIs are coherent; circularity is low.\n\nSoft spots are real but already disclosed and do not invert the main claim. QIN-Breast and ACRIN-NSCLC are detection-only after BAMF re-check, not full independent re-segmentation. SINERGIA melanoma is a same-institution train/val split. Ultra-low-dose and non-FDG tracers are out of scope. Private cohorts limit full re-run. Those constraints mean “generalizable clinical-grade multi-cancer” should be read as “strong on these external FDG phenotypes,” not as universal scanner/protocol proof. That is proportionate, not fatal.\n\nWho it is for: people who need an open whole-body FDG lesion tool for staging support, oligometastatic work, RT contouring assist, or large-scale MTV/TLG studies, and who care about false-positive rate. Citation pattern is appropriate (AutoPET, HECKTOR, nnU-Net, MultiTalent, related clinical PET work). Math is standard supervised segmentation; free parameters are the usual training knobs, not hidden load-bearing tricks.\n\nI would send this to peer review. It is important enough within nuclear medicine / radiation oncology AI, evidentially sharp enough on the stated cohorts, and open enough to be useful. Expect referees to press on detection-only cohorts and the small inter-observer set; the paper can absorb that. Engage with it if you work on PET lesion AI or quantitative biomarkers.","headline":"Solid open multi-cancer FDG-PET/CT lesion segmenter with real external head-to-heads and fewer false positives; novelty is mostly data scale and release, not architecture.","tokens_in":23677,"tokens_out":648,"would_cite":true,"duration_ms":5890,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An open-source AI model segments whole-body FDG-PET/CT cancer lesions across multiple cancer types with fewer false positives than public benchmarks and approaches expert agreement.","keywords":["FDG-PET/CT","whole-body lesion segmentation","deep learning","total tumor burden","total lesion glycolysis","false-positive reduction","multi-cancer external validation","open-source model"],"falsifier":"Run the released model on a fully independent multi-center multi-cancer FDG-PET/CT set with complete expert lesion masks (including ultra-low-dose or scanners/protocols never seen in training) and check whether patient- and lesion-wise F1 and TTB/TLG agreement still beat the same public benchmarks and stay inside the dual-reader range.","tokens_in":23486,"feed_emoji":"🔬","tokens_out":697,"duration_ms":5412,"temperature":0.7,"pith_summary":"Manual outlining of cancer on whole-body FDG-PET/CT is slow and variable, so large-scale use of tumor burden biomarkers has stayed limited. This paper introduces GLOW-FDG, an open model trained on 1,563 multi-cancer scans and tested on 185 scans from other institutions covering breast, lung, head and neck, and melanoma. Across those cohorts it beats three public benchmark models on patient- and lesion-level detection, mainly by cutting false-positive lesions while keeping high recall. Total tumor burden and total lesion glycolysis stay close to human references with high agreement scores, and on a melanoma subset the model’s detection and overlap sit near the range of two radiation oncologists reading the same cases. The authors release weights and code so others can use automated pre-segmentation and biomarker extraction without fixed SUV thresholds.","feed_headline":"Open AI cuts false-positive cancer lesions on whole-body PET/CT","feed_subtitle":"Trained on 1,563 multi-cancer scans, GLOW-FDG beats public models and nears expert agreement on tumor burden.","key_machinery":"GLOW-FDG: a dual-headed residual U-Net (lesion head plus organ-supervision head) trained with PET–CT misalignment augmentation and multi-dataset pretraining, then fine-tuned so physiologic high-uptake organs are learned separately from pathology, reducing false positives without fixed SUV thresholds.","core_discovery":"GLOW-FDG, trained on a diverse multi-cancer FDG-PET/CT corpus of 1,563 scans and evaluated on 185 external scans, consistently delivers the highest patient- and lesion-wise detection F1 among compared public models by raising precision (fewer false positives) while holding high recall, with robust total tumor burden and total lesion glycolysis quantification and performance that approaches inter-observer variability between expert readers on a metastatic melanoma subset.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["GLOW-FDG cuts false-positive lesions on multi-cancer FDG-PET/CT","Open-source model leads public benchmarks in whole-body lesion detection","GLOW-FDG nears expert agreement on tumor burden and TLG","Fewer false positives while holding recall across cancer types","External validation shows top F1 for GLOW-FDG cancer segmentation"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the five external cohorts—two used only for detection after partial re-check, one an institutional train/validation split, plus a ten-case dual-reader melanoma set—are enough to claim generalizable multi-cancer whole-body performance across scanners and disease types not fully represented in training.","fun_headline_variants_meta":{"raw":{"variants":["GLOW-FDG cuts false-positive lesions on multi-cancer FDG-PET/CT","Open-source model leads public benchmarks in whole-body lesion detection","GLOW-FDG nears expert agreement on tumor burden and TLG","Fewer false positives while holding recall across cancer types","External validation shows top F1 for GLOW-FDG cancer segmentation"]},"model":"grok-4.5","effort":"low","cost_usd":0.007048,"raw_usage":{"total_tokens":1748,"prompt_tokens":764,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":70480000,"prompt_tokens_details":{"text_tokens":764,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":904,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":764,"tokens_out":80,"duration_ms":7183,"temperature":1.0,"reasoning_tokens":904,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T22:55:25.684651+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the released model on a fully independent multi-center multi-cancer FDG-PET/CT set with complete expert lesion masks (including ultra-low-dose or scanners/protocols never seen in training) and check whether patient- and lesion-wise F1 and TTB/TLG agreement still beat the same public benchmarks and stay inside the dual-reader range.","supporting_citations":[],"review_version":1}