{"id":"1eb2f19e-0627-45db-8567-94acbf30809f","arxiv_id":"2607.22861","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A uniform fine-tuning step improves acquisition robustness and downstream performance across ten pathology foundation models, but the paper never discloses the fine-tuning recipe.","lead":"This paper reports that fine-tuning ten pathology AI models with a single (undisclosed) recipe makes their tile features more robust to differences in scanners and staining, while also improving or preserving downstream task scores. It is a candidate recipe for making digital pathology tools work across different laboratories, but the method itself is not described in the manuscript.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim cannot be checked: the fine-tuning recipe and training data are never described, so independence of the evaluation benchmarks is unverifiable.","rationale":"The reader identified the same weakest assumption: fine-tuning data could overlap with evaluation benchmarks. I agree and add that the omission is not merely a data-hygiene detail; it is the epistemic core of the paper. The entire contribution is an empirical generalization claim, and that claim is only meaningful if the recipe is disclosed. Section 3 never specifies the training data, objective, or hyperparameters; the Methods section is absent; no code or config is released. The released checkpoints (Phaet, Mascaret) are evidence that the fine-tuning happened, but they do not let a reviewer check whether the evaluation sets were contaminated. The reported numbers are internally consistent, and the Figure 2 CAMELYON analysis does show a genuine shift toward biological clustering for Phikon-v2, which is real independent evidence for the mechanism on that one model. However, the cross-benchmark claim over ten models rests on the unstated disjointness assumption. This is severe, but it is not a demonstrated fraud; it is an unverifiable claim. Therefore REJECT is too strong as a scientific verdict: the right verdict is UNVERDICTED until the recipe and data provenance are disclosed. A REJECT verdict presumes the claim is false or the work is unusable; an UNVERDICTED verdict accurately states that the central claim cannot be assessed with the provided information. If the authors supply the omitted details and the disjointness check passes, the paper should move to CONDITIONAL or ACCEPT depending on the results of the concrete test.","tokens_in":26368,"tokens_out":1998,"duration_ms":17413,"concrete_test":"Obtain from the authors a precise list of the training slides/cohorts used for fine-tuning, including TCGA and Tolkach membership, and the loss/objective and hyperparameters. Then re-run the PathoROB evaluation on a held-out TCGA and Tolkach split that is provably disjoint from the fine-tuning set by slide identifier; if the RI gains shrink below the reported 23% average or fail on either held-out cohort, the generalization claim is unsupported.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper's central claim is that a single fine-tuning step improves robustness and downstream performance on unseen acquisition sources. The single load-bearing condition is that the fine-tuning data are disjoint from the evaluation benchmarks. The paper only states that the five CAMELYON centers in Figure 2 were unseen during fine-tuning (Figure 2 caption). It never lists the fine-tuning data, the loss function, or the training protocol, so the TCGA and Tolkach components of PathoROB, the HEST/THUNDER/Patho-Bench cohorts, and the PLISM/SCORPION analysis sets could all have contributed training tiles. If they did, the robustness gains are partly in-distribution and the claimed generalization collapses. Section 3.2 defines the PathoROB benchmark (TCGA, CAMELYON, Tolkach), Section 4.1 only excludes the five CAMELYON centers for Figure 2, and Section 5 uses PLISM and SCORPION only as analysis probes with no statement that they were held out. Because the recipe is absent, even the reviewer cannot distinguish a genuine invariance mechanism from evaluation leakage or from a subtle form of self-distillation on benchmark-related data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript claims a novel fine-tuning recipe that, applied to ten pathology foundation models, jointly improves robustness to acquisition factors and downstream task performance with no observed trade-off. The evaluation uses PathoROB for robustness, and HEST, THUNDER, and Patho-Bench for performance, reporting consistent gains across all ten pairs, with Wilcoxon signed-rank tests, and releases two fine-tuned models (Phaet and Mascaret). The paper also presents analyses on PLISM and SCORPION to argue that acquisition shifts are near-linear in feature space and that fine-tuning instills invariance across transformer depth.","tokens_in":26547,"tokens_out":3154,"duration_ms":28784,"significance":"If the empirical claims hold, the paper would be practically valuable: it offers a model-agnostic way to improve robustness across very different pathology FMs, with gains also on downstream tasks, and it releases models that practitioners could adopt directly. The benchmark coverage is broad and external (PathoROB, HEST, THUNDER, Patho-Bench), the set of base FMs includes models from several independent groups, and the internal numbers are reproducible in the sense that the tables are detailed and the paired Wilcoxon tests support the direction of the robustness effect. However, the central contribution, the fine-tuning recipe itself, is never described, and the training data are not disclosed, so the generalization claim cannot currently be evaluated or reproduced.","major_comments":[{"comment":"The manuscript never specifies the fine-tuning recipe that the abstract and introduction present as the central contribution. There is no loss function, no fine-tuning dataset, no hyperparameters, no optimization details, and no compute budget anywhere in Sections 3–5. As a result, a reader cannot reproduce the method, cannot determine what makes it 'novel', and cannot assess whether the recipe is distinct from existing robustness methods discussed in Section 2. This is load-bearing: the entire paper is an evaluation of a method that is never stated. A complete method section, including training data provenance, objective, and hyperparameters, is required.","section":"Section 3 (Experimental setup)"},{"comment":"The only statement that evaluation data were unseen during fine-tuning concerns the five CAMELYON centers shown in Figure 2. No statement is made for the TCGA and Tolkach components of PathoROB (Section 3.2), for HEST, THUNDER, or Patho-Bench (Section 3.3), or for PLISM and SCORPION (Section 5). If any of these cohorts or slides were used for fine-tuning, then the claimed robustness gains on 'unseen acquisition sources' are partly in-distribution and the central generalization claim collapses. The paper must provide an explicit list of the fine-tuning data and a per-benchmark statement of disjointness.","section":"Section 4.1 and Figure 2 caption"},{"comment":"The claim 'no observed trade-off' is contradicted by the paper's own tables. Table 2 shows that H0-mini's THUNDER rank sum worsens from 61 to 65 and AquaViT's is unchanged at 60. Table 3 shows that GenBio-PathFM's average HEST Pearson correlation decreases from 0.4197 to 0.4178. These are regressions or ties on individual benchmarks, even if aggregate cross-benchmark ranks improve. The text should be qualified to say that overall performance improves on aggregate, while individual benchmark regressions do occur, rather than claiming no trade-off at all.","section":"Section 4.1 and Conclusion"},{"comment":"The extended leaderboards mix official published values for base models with in-house computed values for fine-tuned models and for some mixed-precision base models. If the evaluation environments differ (e.g., full precision vs. mixed precision, different preprocessing or hardware), the ranks in Table 1 are not strictly comparable across rows. Please state explicitly which rows were computed in-house, which were taken from official leaderboards, and whether official leaderboard values were produced under the same precision and preprocessing protocol.","section":"Appendix A (Tables 5–7)"}],"minor_comments":[{"comment":"Section 3.3 lists nine cancer types including 'hepatocellular carcinoma (liver cancer, HCC)', but Table 3's columns are labeled IDC, PRAD, PAAD, SKCM, COAD, READ, CCRCC, LUNG, LYMPH IDC. The LUNG column appears in place of HCC. Please align the list and the table headers.","section":"Section 3.3 vs. Table 3"},{"comment":"Figure 3 shows 225 points per scanner but the text does not explain how this subset of the 16,278 PLISM tiles was sampled. A brief sampling description would help interpret the PCA projection.","section":"Section 5, Figure 3"},{"comment":"The THUNDER rank sums in Table 2 (e.g., UNI2-h base 26, fine-tuned 20) differ from the extended leaderboard in Table 5 (UNI2-h base 35, fine-tuned 31) because Table 5 includes more models and uses a different ranking pool. The relationship between the two tables should be explained in the text.","section":"Section 4.2, Table 2"},{"comment":"The paper refers to 'wearewaiv.github.io/histoboard/models' for detailed model information. This is not a stable scientific citation; please provide a persistent repository, versioned release, or arXiv reference for the model details.","section":"Section 3.1 and Appendix A"},{"comment":"Several hyperlinks are written as raw URLs (e.g., huggingface.co/wearewaiv/models). Please format them properly and ensure they are accessible at the time of publication.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The empirical study is extensive and the reported numbers appear internally consistent, but the missing method section and incomplete data-disjointness statement are fundamental. I would be willing to consider a revised version that adds a full fine-tuning specification and an explicit training-data audit. If the authors cannot disclose the training data or the loss, the paper would need to be reframed as an empirical robustness study of a proprietary recipe, which would still require enough detail for the claims to be testable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, this is a paper with a large claim and a missing heart. What's new: a single fine-tuning step applied uniformly to ten pathology FMs, including externally trained ones like UNI2-h and Virchow2, improves PathoROB robustness across the board and mostly improves or preserves performance on three downstream benchmarks. That's a strong, non-circular empirical result, and releasing Phaet and Mascaret is concrete. The feature-space analysis with PLISM and SCORPION is a nice touch, and the depth-wise retrieval curves give the robustness gain some mechanistic plausibility.\n\nThe problem is that the paper never says what the fine-tuning recipe actually is. There is no method section, no loss function, no training data, no hyperparameters, no code. The central contribution is a black box. You cannot reproduce it, compare it with existing robustness methods, or judge whether it is 'novel' beyond a routine invariance loss. That is load-bearing absence, not a stylistic choice.\n\nThe leakage concern is real but not yet proven: the fine-tuning data are never listed, so the TCGA and Tolkach components of PathoROB, and possibly the HEST/THUNDER/Patho-Bench cohorts, could have contributed training tiles. The paper only guarantees that the five CAMELYON centers in Figure 2 were unseen. If TCGA slides were used, the PathoROB robustness gains are partly in-distribution. The authors need to either disclose the fine-tuning data or demonstrate disjointness.\n\nAlso, 'no observed trade-off' is overstated. The paper's own limitations admit H0-mini regresses on THUNDER and GenBio-PathFM on HEST. That is not fatal—the aggregate trend is still positive—but the abstract should say 'small, task-specific regressions observed.'\n\nMinor: base model results appear to be quoted from official leaderboards while fine-tuned results are computed in-house; that is a potential protocol mismatch in the comparison, though the paper claims to use official implementations.\n\nOverall, this is a plausible and potentially useful empirical result, but as written it cannot be validated. I would send it to peer review, not desk reject, because the finding is significant if the recipe is disclosed and the leakage question resolved. Expect a major revision or rejection if the authors cannot supply the missing method.","headline":"Big empirical claim, but the fine-tuning recipe is missing and the fine-tuning data are undisclosed, so the central result cannot be checked.","tokens_in":27125,"tokens_out":3595,"would_cite":false,"duration_ms":31280,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single fine-tuning step applied to ten pathology foundation models raises average robustness by 23% and cross-benchmark performance by 43%, with no observed trade-off.","keywords":["pathology foundation models","robustness","fine-tuning","scanner variability","stain variability","domain shift","feature space invariance","histopathology"],"falsifier":"Check the fine-tuning data manifest against PathoROB (including its TCGA and Tolkach cohorts), HEST, THUNDER, and Patho-Bench: if any of those slides or tiles appear in fine-tuning, the robustness and performance gains are partly in-distribution rather than evidence of generalization to unseen acquisition sources. If the data are clean, rerunning the recipe with those benchmark cohorts explicitly held out should reproduce the average 23% PathoROB gain and the 43% cross-benchmark improvement.","tokens_in":26116,"feed_emoji":"🔬","tokens_out":7419,"duration_ms":58619,"temperature":0.7,"pith_summary":"This paper tries to establish that a single fine-tuning step can make pathology foundation models robust to scanner and staining variability without sacrificing their generic quality. Applied uniformly to ten existing encoders, the step raised the average PathoROB robustness index from 0.72 to 0.87 and improved combined HEST, THUNDER, and Patho-Bench performance by 43%, with every model moving up on robustness and overall rank. If true, it means a cheap, label-free procedure can decouple pathology representations from acquisition confounders, so a single encoder can be deployed across laboratories with different scanners and stain protocols. The paper also gives a mechanistic picture: fine-tuning reorients the dominant directions of feature space from 'where the slide was digitized' to 'what the morphology is.'","feed_headline":"Fine-tuning lifts pathology AI robustness by 23%","feed_subtitle":"A single label-free fine-tune also raised downstream performance on all ten models, with no trade-off.","key_machinery":"The load-bearing object is the fine-tuned encoder's reorganized feature geometry, produced by a fine-tuning step applied uniformly to all ten models. The paper supports its mechanism with two observations: on the PLISM dataset, scanner shift appears as a near-linear offset in feature space, so a feature-level correction seems possible in principle; but cross-scanner retrieval on SCORPION shows that invariance builds up across transformer depth, with fine-tuning reaching a given retrieval quality roughly eight blocks earlier and a higher asymptote (mAP about 0.99 versus 0.91 at the final block). This locates acquisition invariance deep in the transformer and identifies fine-tuning as re-purposing the dominant feature-space directions from acquisition site to biological class.","core_discovery":"The central claim is that acquisition robustness is not a property that must be bought with pretraining scale or traded off against downstream utility: fine-tuning the encoder itself is sufficient. Across ten pathology foundation models spanning different architectures and pretraining recipes, every model improved its PathoROB robustness index after fine-tuning (one-sided Wilcoxon signed-rank $p<10^{-4}$), and every model improved its overall rank on the combined HEST, THUNDER, and Patho-Bench leaderboards ($p<10^{-4}$); the best fine-tuned encoder, UNI2-h, moved from total rank 21 to 5. On the CAMELYON subset of PathoROB, the leading axis of Phikon-v2's feature space switched from clustering by medical center (Adjusted Rand Index 0.46 against center, 0.00 against metastasis) to clustering by metastasis status (ARI 0.67 against metastasis, 0.01 against center), for five centers that were not seen during fine-tuning. The paper interprets the joint up-and-to-the-right shift as evidence that scanner- and stain-related directions are nuisance dimensions: removing them frees capacity for biologically relevant structure.","pith_inferences":["If scanner shift really is a near-linear offset, a testable extension is to combine the fine-tuning recipe with an explicit linear correction estimated on paired multi-scanner slides; the paper's depth analysis suggests the linear correction alone would be partial, because the invariance is assembled deep in the transformer.","Because the fine-tuning data and recipe hyperparameters are not disclosed in the main text, external replication on completely unseen scanners will determine whether the no-trade-off claim is a property of the method or of the specific evaluation setup.","The PathoROB index measures feature-space dominance of biology over confounders, not end-task robustness; a natural next check is whether fine-tuned encoders reduce site-specific errors in biomarker or survival tasks under external validation, which the paper explicitly leaves for future work.","If the gains hold under strict data separation, the recipe becomes a general post-processing step that could be applied to any new pathology encoder, including vision-language models, rather than a one-off training scheme."],"forward_implications":["A laboratory can adopt a robustified encoder as a drop-in feature extractor and expect it to keep working when the scanner or stain protocol changes, without retraining the encoder.","Fine-tuning acts as an equalizer: the largest robustness gains land on the least robust base models, so older or smaller encoders can be brought closer to top-tier performance without larger pretraining corpora.","Because every fine-tuned model improves or matches its base model on aggregate benchmarks, robustification can be applied before downstream heads are trained, with no observed performance tax.","The released robust versions of Phikon-v2 and Midnight-12k (Phaet and Mascaret) make the effect immediately available for other pipelines; Mascaret ranks first on PathoROB and second on average downstream performance among publicly available models.","Since scanner invariance emerges roughly eight transformer blocks earlier after fine-tuning, even models that read intermediate features inherit part of the robustness gain."],"supporting_citations":[{"why":"Defines PathoROB and the robustness index used to measure acquisition robustness.","marker":"[34]"},{"why":"Provides HEST, the spatial gene-expression prediction benchmark used for downstream performance.","marker":"[29]"},{"why":"Provides THUNDER, the tile-level benchmark used for downstream rank sums.","marker":"[39]"},{"why":"Provides Patho-Bench, the 63-task slide-level benchmark used for downstream grand averages.","marker":"[63]"},{"why":"Supplies PLISM, the registered multi-scanner multi-stain dataset used to observe that scanner shift is a near-linear offset in feature space.","marker":"[42]"},{"why":"Supplies SCORPION, the aligned scanner-pair patches used to show that fine-tuning builds scanner invariance across transformer depth.","marker":"[45]"}],"fun_headline_variants":["Fine-tune once, boost all 10 pathology models","23% robust gain, 43% perf gain: fine-tune pathology FMs","No trade-off: fine-tuning lifts pathology AI robustness 23%","Single fine-tune makes every pathology model more robust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that robustness generalizes to unseen acquisition sources depends on the fine-tuning data being disjoint from every evaluation benchmark; the paper demonstrates this for five CAMELYON centers but never states whether the TCGA or Tolkach parts of PathoROB, or the data behind HEST, THUNDER, and Patho-Bench, were excluded.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tune once, boost all 10 pathology models","23% robust gain, 43% perf gain: fine-tune pathology FMs","No trade-off: fine-tuning lifts pathology AI robustness 23%","Single fine-tune makes every pathology model more robust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001073,"raw_usage":{"total_tokens":4496,"prompt_tokens":949,"completion_tokens":3547,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":3473}},"tokens_in":565,"tokens_out":3547,"duration_ms":25217,"temperature":1.0,"reasoning_tokens":3473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:28:26.775767+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the fine-tuning data manifest against PathoROB (including its TCGA and Tolkach cohorts), HEST, THUNDER, and Patho-Bench: if any of those slides or tiles appear in fine-tuning, the robustness and performance gains are partly in-distribution rather than evidence of generalization to unseen acquisition sources. If the data are clean, rerunning the recipe with those benchmark cohorts explicitly held out should reproduce the average 23% PathoROB gain and the 43% cross-benchmark improvement.","supporting_citations":[{"cited_title":"Scorpion: Addressing scanner-induced variability in histopathology","cited_arxiv_id":null,"evidence_quote":"Supplies SCORPION, the aligned scanner-pair patches used to show that fine-tuning builds scanner invariance across transformer depth."}],"review_version":1}