{"id":"08939f57-b5ab-42f1-99b8-9086e8064229","arxiv_id":"2607.17227","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"In a harmonised benchmark of six transcriptomics foundation models, performance is strongly context-dependent, with no model dominating across modalities, tasks, and metrics.","lead":"A large benchmark tested six AI models trained on single-cell and spatial gene-expression data across five biological tasks, including cell-type identification, tissue-domain recovery, and perturbation prediction. The result: no single model wins everywhere; the best choice depends on the data type, the task, and the metric used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ranking evidence for context-dependence is not protected against protocol confounds: Leiden resolution is tuned per model and preprocessing/tokenisation are never varied independently.","rationale":"The reader's weakest assumption identified essentially the same load-bearing concern: cross-model comparability is undermined by model-specific preprocessing, tokenisation, and clustering protocols, with Leiden resolution tuned to match label counts. My stress-test agrees and sharpens it. The paper's strongest claim is that 'foundation' is a conditional property and that rankings shift with factors including preprocessing and tokenisation. The tables do show that different models perform best on different tasks, which supports a weaker version of the conditional conclusion. However, the stronger causal version—that preprocessing, tokenisation, and biological prior drive the shifts—is not tested: those factors are fixed within each model's recommended pipeline and never varied while holding the model constant. The lack of simple baselines and uncertainty intervals compounds this, because a ranking based on single-seed point estimates may not be stable. I therefore cannot call the central argument fully established, but the qualitative observation is plausible and the paper already frames itself as a benchmark rather than a mechanistic decomposition. CONDITIONAL remains the appropriate verdict: the concern is real but addressable, not fatal. I recommend no change to the reader's verdict, with the concrete test above clarifying what would settle the issue.","tokens_in":25114,"tokens_out":4138,"duration_ms":46153,"concrete_test":"Re-run the zero-shot and continual-pretraining clustering tables on BMMC and HBCA1 with a single fixed Leiden resolution (e.g., resolution=1.0) across all models instead of sweeping to match the reference label count, and repeat with 20 random seeds to obtain bootstrap 95% confidence intervals for ARI and NMI. If the rank order of models changes materially, or if confidence intervals are wider than the observed between-model gaps, the current evidence is protocol-dependent and the central conditional-ranking claim needs to be softened or re-evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion—that model performance is strongly conditional and rankings shift with modality, preprocessing, tokenisation, biological prior, domain shift and metric choice—rests on cross-model ranking comparisons in Tables 2–5. Those comparisons are not shielded from evaluation-protocol artifacts. First, zero-shot clustering uses model-specific preprocessing and graph construction (Methods, 'Zero-shot cell type clustering'), then 'Leiden resolution was swept until the number of clusters matched the number of reference labels'; for Novae, the number of assigned domains is set to match labels. Matching cluster count does not make resolution comparable: models with different intrinsic granularities are evaluated at different operating points of the ARI/NMI curve, and a small resolution change can reorder models. Second, the abstract and Discussion attribute ranking shifts to 'preprocessing, tokenisation, biological prior', but the design never varies these factors for a fixed model; they are fully confounded with model identity, so those causal attributions are not supported by the experiments. Third, all metrics are single-seed point estimates without bootstrap confidence intervals or simple baselines (e.g., PCA+Leiden) for the clustering tasks, even though the Discussion itself recommends reporting 'simple baselines' and 'uncertainty intervals.' Consequently, observed gaps such as scELMo ARI=0.50 vs CellPLM ARI=0.31 on BMMC (Table 2) may reflect resolution tuning or seed sensitivity rather than representational quality. This does not necessarily overturn the qualitative 'no model dominates' observation, but it does weaken the strong version of the claim—that ranking shifts are driven by the listed biological and design factors.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript benchmarks six foundation models (Nicheformer, CellPLM, scGPT-spatial, GenePT-w, scELMo, Novae) across scRNA-seq, spatial transcriptomics and Perturb-seq, using zero-shot clustering, continual pretraining, supervised annotation, marker-gene concordance and perturbation prediction. The central claim is that no model dominates and that ‘foundation’ is a conditional property: rankings shift with modality, dataset, biological granularity, metric and preprocessing/tokenisation. The empirical basis is mainly Tables 2–5, with interpretation framed around the conditional-generalisation thesis in the Discussion.","tokens_in":25509,"tokens_out":4714,"duration_ms":45770,"significance":"If the cross-model comparisons were clean, this would be a valuable and timely resource: the paper covers a broad and well-chosen model panel, uses external datasets, fixes random seeds, applies shared test sets for supervised annotation, uses shared evaluated gene sets for perturbation prediction, and provides a public code repository. It also explicitly argues for biological, perturbation-grounded evaluation rather than leaderboard ranking, which is an important message for the field. However, the central ranking evidence is not yet shielded from evaluation-protocol confounds. The paper’s own Discussion recommends simple baselines and uncertainty intervals, but the analyses do not include them, and several methodological choices—especially resolution sweeping and model-specific preprocessing—directly affect the reported ranking shifts. The significance of the claims is therefore conditional on additional sensitivity analyses and reframing.","major_comments":[{"comment":"The central ranking evidence in Table 2 is not protected against clustering-protocol artifacts. Leiden resolution is ‘swept until the number of clusters matched the number of reference labels’, and for Novae the number of assigned domains is set to the number of labels. Matching the cluster count does not make the operating point comparable across models: each model is evaluated at a different point on the ARI/NMI-versus-resolution curve, and small resolution changes can reorder models. The abstract’s attribution of ranking shifts to ‘preprocessing, tokenisation’ is also not supported by the design, since these factors are not varied independently of model identity. The same concern applies to the continual-pretraining clustering results in Table 3. I request sensitivity analyses with a fixed resolution and/or multiple resolutions, simple baselines (e.g., PCA + Leiden), and bootstrap con","section":"Methods — Zero-shot cell type clustering; Table 2"},{"comment":"The perturbation-prediction comparison is architecture- and decoder-confounded. CellPLM and scGPT-spatial are evaluated with their native end-to-end models, whereas GenePT-w and scELMo are evaluated by feeding their gene embeddings into a shared GEARS decoder. The shared gene set is defined as an intersection of model vocabularies/embedding spaces, but the downstream prediction machinery is not shared across all four models. Consequently, the claim that ‘language-derived gene embeddings were competitive for selected perturbation-response metrics’ conflates gene-representation quality with decoder compatibility and training protocol. In addition, the Discussion cites Ahlmann-Eltze et al. and Systema as showing that deep perturbation models do not yet outperform simple linear baselines, yet Table 5 includes no such baselines. Global Pearson correlations are all ≈0.98–0.99, so the response-","section":"Methods — Perturbation Prediction; Table 5"},{"comment":"Supervised annotation results in Table 4 are also difficult to interpret as representation comparisons because each model uses a different supervised head and training protocol: full fine-tuning for Nicheformer, CellPLM and scGPT-spatial; kNN classifiers for GenePT-w and scELMo; and a frozen encoder with a one-epoch MLP for Novae. The Novae row on SEAAD (accuracy 0.30, macro-F1 0.03, ROC-AUC 0.50) is presented as a representation-scale mismatch, but it is equally explained by the very weak one-epoch MLP protocol. The claim that ‘supervised performance depends … on whether the model’s pretraining objective and representation scale match the biological resolution’ would be better supported by a common classifier (e.g., logistic regression or kNN on embeddings) applied to all models, with the native protocols reported as secondary information.","section":"Methods — Cell type annotation; Table 4"},{"comment":"The paper’s own Discussion recommends that future benchmarks report ‘simple baselines, leave-domain-out splits, uncertainty intervals’. None of these are provided for the central tables. All results appear to be single-seed point estimates, and many ranking differences are small (e.g., Table 3, BMMC NMI: scGPT-spatial 0.66 vs CellPLM 0.66; Table 5, Replogle Pearson Delta DE: CellPLM 0.4604 vs scELMo 0.4579). With no bootstrap or repeated-seed variance, claims that ‘rankings shifted’ and ‘no model dominated’ are over-precise and may be affected by noise. The authors should add confidence intervals, repeated seeds, or an explicit noise-aware analysis, and at minimum soften the abstract and Discussion claims until such evidence is available.","section":"Discussion; Tables 2–5"}],"minor_comments":[{"comment":"The phrase ‘rankings shifted with modality, preprocessing, tokenisation, biological prior, domain shift and metric choice’ overstates what the design can show: preprocessing and tokenisation are model-specific and never varied independently. Consider wording such as ‘rankings differed across models, whose preprocessing and tokenisation differ’.","section":"Abstract"},{"comment":"Minor grammatical error: ‘using harmonised evaluation framework’ should be ‘using a harmonised evaluation framework’.","section":"Results — first paragraph"},{"comment":"The statement that fixing random seeds at 42 makes results reproducible is not a substitute for variance estimation; please clarify that reproducibility and statistical certainty are distinct.","section":"Methods — Zero-shot cell type clustering"},{"comment":"The cell-state labels mix formatting: ‘AS DC’ appears as ‘AS-DC’ elsewhere. Please standardise across figure, text and tables.","section":"Figure 4 caption"},{"comment":"A versioned release of the code repository (e.g., a DOI or release tag) would improve reproducibility, since the current link points to an unversioned GitHub repository.","section":"Data and Tool Availability"},{"comment":"The k-means sensitivity check is reported only for GenePT-w. It would be informative to report k-means results for the other models as well, given the resolution-sweeping concern in the main clustering analysis.","section":"Supplementary Table S2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about many of its limitations in the Discussion, but those limitations are load-bearing for the central claim rather than cosmetic. The title and abstract promise a ‘harmonised’ benchmark, yet preprocessing, tokenisation, adaptation protocols and supervised classifiers differ across models, and the paper’s own recommended controls (simple baselines, uncertainty intervals, leave-domain-out splits) are absent. I see this as a correctable problem: the benchmarking effort is broad and potentially valuable, and adding sensitivity analyses, common classifiers/decoders, fixed-resolution clustering, and variance estimates would substantially strengthen it. I would not reject the paper, but I would require these additions before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is the paper I owe you: a benchmark of six single-cell/spatial foundation models across scRNA-seq, spatial transcriptomics, and Perturb-seq, with zero-shot clustering, continual pretraining, supervised annotation, marker-gene concordance, and perturbation prediction. The headline—no model dominates, and rankings shift with dataset, task, and metric—is supported by the tables and is worth taking seriously. The breadth is genuinely new: no prior benchmark in the cited literature covers this exact model panel across these modalities with a shared evaluation pipeline. They also ship code, which puts them ahead of many benchmarks in the field.\n\nWhat the paper does well: it treats \"foundation\" as a hypothesis rather than a given, it checks biological interpretability via marker-gene overlap, and it separates global expression reconstruction from perturbation-response prediction. The practical message—evaluate task-specific, don't trust leaderboards—is solid and actionable.\n\nThe soft spots are mostly in the methodology, not the conclusion. \"Harmonised\" is doing a lot of work: each model uses its own preprocessing and tokenisation, and the paper never varies those factors for a fixed model. The attributions in the abstract and discussion—that ranking shifts are driven by \"preprocessing, tokenisation, biological prior\"—are not tested by ablations; those factors are fully confounded with model identity. Second, zero-shot clustering sweeps Leiden resolution until cluster count matches the reference label count. That makes models comparable in name but not in operating point: a model with different intrinsic granularity is being evaluated at a different point on the ARI curve, and a small change in resolution can reorder models. Third, there are no simple baselines (e.g., PCA+Leiden), no confidence intervals or bootstrap, and no multi-seed variability, even though the Discussion itself recommends exactly these. The perturbation prediction comparison also mixes native workflows with a shared GEARS decoder, which helps but doesn't fully standardise the decoders.\n\nMy read: these issues weaken the strong version of the claim—that ranking shifts are caused by the listed design factors—but the qualitative observation that no single model dominates across tasks and modalities is likely robust. It would take more than a resolution sweep to explain away scGPT-spatial being best on ATAA and scELMo best on BMMC, or GenePT winning on Norman perturbation response. Still, the paper as written oversells the causal attributions.\n\nThis deserves a serious referee. A revision with simple baselines, error bars, and a preprocessing ablation would turn a useful benchmark into a really solid one. I'd take it to a reading group—there's a lot to argue about, which is exactly what makes it useful.\n\nRecommendation: send it to peer review, with a request for those additions.","headline":"Broad benchmark with a real headline observation—no foundation model dominates—but the strongest version of the claim is undercut by protocol tuning, missing baselines, and single-seed metrics.","tokens_in":26030,"tokens_out":2320,"would_cite":true,"duration_ms":21905,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No single foundation model dominates across single-cell and spatial tasks, a harmonised benchmark finds.","keywords":["foundation models","single-cell transcriptomics","spatial transcriptomics","Perturb-seq","benchmarking","zero-shot clustering","cell type annotation","perturbation prediction"],"falsifier":"A concrete check: re-run zero-shot clustering on the three scRNA-seq datasets using a different principled resolution rule for every model, such as a fixed resolution across all models or each model's native clustering method with default parameters. If the rank order of ARI or NMI changes substantially—for example, a model that was weakest becomes strongest—the harmonised pipeline, not the biology, is driving the reported conditional generalisation.","tokens_in":24992,"feed_emoji":"🧬","tokens_out":3311,"duration_ms":30416,"temperature":0.7,"pith_summary":"The paper tests whether single-cell and spatial 'foundation models' truly generalise across modalities and analytical tasks by running six models through one harmonised benchmark spanning scRNA-seq, spatial transcriptomics and Perturb-seq. It claims that 'foundation' is a conditional property, not a consequence of model scale or pretraining: rankings shift with task, modality, preprocessing, tokenisation and metric. Expression-trained cell-level transformers resolve many cell-identity tasks best, spatial and graph-aware models better preserve tissue architecture, and text-derived gene embeddings are competitive for selected perturbation-response metrics. If the paper is right, model selection must be task- and modality-specific, and single-number leaderboard evaluations are misleading.","feed_headline":"No single foundation model wins across cell and tissue tasks","feed_subtitle":"Harmonised benchmark of six models shows rankings shift with task, modality and metric","key_machinery":"A harmonised benchmarking workflow that combines each model's own recommended preprocessing and tokenisation with shared downstream tasks, shared metrics, fixed random seeds, and—for perturbation prediction—shared evaluated gene sets. The load-bearing design choice is Leiden clustering with the resolution swept until the cluster count matches the number of reference labels, used for both zero-shot and continually pretrained clustering. This workflow is what lets the paper attribute ranking differences to model design rather than inconsistent evaluation, and it is also the step on which cross-model comparability rests.","core_discovery":"On the paper's own terms, the central discovery is that generalisation in single-cell and spatial foundation models is context-dependent rather than universal. Across five downstream tasks, no architecture consistently beats the others: the strongest cell-identity clustering in scRNA-seq came from expression-trained cell-level transformers such as scGPT-spatial and CellPLM, while tissue-domain recovery favoured spatial and graph-based models such as Novae, with scELMo leading selected settings, and language-derived gene embeddings (GenePT-w, scELMo) led response-focused perturbation metrics. The paper argues this pattern shows that current models capture useful but partial biological represe","pith_inferences":["A direct extension would be systematic leave-one-domain-out evaluation (leave-one-tissue-out, leave-one-platform-out, leave-one-species-out), since the paper argues domain shift is the central biological problem yet its datasets are varied rather than fully crossed.","The result predicts that task-specialised models encoding a specific biological hypothesis—such as immune-context, developmental programme or tissue niche—will outperform generic foundation models on those tasks, pointing toward problem-specific rather than universal models.","Because the conclusion depends on the harmonised pipeline, a testable extension is to vary resolution rules and preprocessing choices per model and check whether the observed ranking inversions persist; stability across those choices would strengthen the claim that the rankings reflect biology rather than evaluation artefacts."],"forward_implications":["Model selection for cell-identity tasks should favour expression-trained cell-level transformers over text-derived or spatial-domain models in many scRNA-seq settings.","Spatial-domain discovery benefits from graph-based or spatially aware models; using a cell-identity model for tissue segmentation can produce fragmented or over-segmented domains.","Global post-perturbation expression reconstruction is not sufficient: models must be judged on delta-based, DE-gene metrics to test whether they capture actual perturbation response.","Continual pretraining is not a default improvement; it can sharpen fine-grained immune states and tumour domains but degrade other settings, so adaptation protocols must be evaluated per architecture and dataset.","Any single leaderboard score misleads; rankings must be reported per modality, biological granularity and evaluation metric."],"fun_headline_variants":["No universal winner in cell foundation models","Cell model rankings shift with task and data","Context rules which cell model performs best","Benchmark finds no all-around cell AI model","Cell AI models: best pick depends on job"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The benchmark assumes that a single harmonised pipeline—model-specific preprocessing plus Leiden resolution swept to match label counts—compares models fairly; if these choices are not interchangeable across models, the observed ranking shifts could be evaluation artefacts rather than genuine differences in biological generalisation.","fun_headline_variants_meta":{"raw":{"variants":["No universal winner in cell foundation models","Cell model rankings shift with task and data","Context rules which cell model performs best","Benchmark finds no all-around cell AI model","Cell AI models: best pick depends on job"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000135,"raw_usage":{"total_tokens":951,"prompt_tokens":688,"completion_tokens":263,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":197}},"tokens_in":432,"tokens_out":263,"duration_ms":3428,"temperature":1.0,"reasoning_tokens":197,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T18:38:09.071454+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: re-run zero-shot clustering on the three scRNA-seq datasets using a different principled resolution rule for every model, such as a fixed resolution across all models or each model's native clustering method with default parameters. If the rank order of ARI or NMI changes substantially—for example, a model that was weakest becomes strongest—the harmonised pipeline, not the biology, is driving the reported conditional generalisation.","supporting_citations":[],"review_version":1}