Pith. sign in

REVIEW 2 major objections 4 minor 40 references

Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop

T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper argues that a lack of standardized, cross-domain benchmarks is the core barrier to trustworthy AI models of biological systems, and that the recommendations from the CZI Virtual Cells workshop provide a concrete path to building…

desk verdict A solid, well-referenced workshop consensus document that accurately lists benchmarking challenges but leaves the hard collective-action problem unresolved; worth peer review only as a position piece. read the letter →

arxiv 2507.10502 v2 pith:QKILSGSV submitted 2025-07-14 cs.LG cs.AI

classification cs.LGcs.AI
keywords benchmarkingAIvirtualcellsreproducibilityevaluationmetricsdataleakagecross-domainbenchmarksmulti-omicscommunitystandards
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the absence of standardized, cross-domain benchmarks is the main obstacle to building AI models of cells that are robust and trustworthy, and that the recommendations from a recent CZI Virtual Cells workshop can remove that obstacle. Synthesizing input from experts across imaging, transcriptomics, proteomics, and genomics, it identifies recurring bottlenecks: small and noisy datasets, batch effects, data leakage, weak reproducibility incentives, single-metric evaluation, fragmented resources, and biases in what gets studied. It then proposes concrete interventions: invest in curated benchmarking datasets, standardize tools and documentation, use multiple metrics reviewed by domain experts, create a centralized or federated platform, and sustain an interdisciplinary community with recurring CASP-style assessments. If adopted, these recommendations would make model performance comparable across biological tasks and modalities, which the paper sees as a prerequisite for AI-driven Virtual Cells.

What carries the argument

The carrying mechanism is a proposed benchmarking ecosystem rather than a mathematical object. Its core pieces are curated benchmark datasets designed to be withheld from training; the three-part reproducibility framework of technical replicability (sharing versioned, containerized code), statistical replicability (proper data splitting and resampling), and conceptual replicability (documented workflows and metadata); multi-faceted evaluation metrics chosen jointly by machine learning researchers and biological domain experts; a centralized or federated platform with common formats; and recurring community assessments modeled on CASP. Each recommendation is meant to feed a loop in which data informs models, models are scored against benchmarks, and benchmark plateaus signal where new data generation is needed.

What would settle it

Track whether the proposed platform and standards are actually adopted: if, within a few years of release, the majority of new AI-for-biology papers still evaluate models on private, purpose-built datasets or incompatible public ones, the claim that these recommendations will accelerate robust benchmarking is contradicted.

Watch

Extended reading notes

Core claim

The paper's central claim is that evaluating AI models of biological systems has not kept pace with model development, and that this gap, rather than any single modeling deficiency, is what currently blocks the field from achieving reliable, general-purpose Virtual Cells. Across imaging, transcriptomics, proteomics, and genomics, the authors find a common pattern: datasets are small, heterogeneous, and biased; code and workflows are hard to reproduce; metrics are narrow and often detached from biological questions; and benchmarks, leaderboards, and data are scattered across incompatible platforms. They conclude that model innovation alone will not produce trustworthy biology AI without a coordinated benchmarking ecosystem, and they propose a set of community-coordinated standards, tools, and incentives to build one.

Load-bearing premise

The entire recommendation set rests on the assumption that funders, companies, and academic labs will be motivated to contribute data, maintain tools, and follow shared standards even though the report itself documents that their incentives currently pull in opposite directions.

Editorial extensions

If this is right

  • If benchmarks are standardized across modalities, models trained on imaging, transcriptomics, proteomics, and genomics can be compared directly on tasks requiring cross-domain biological knowledge.
  • A centralized or federated platform with common data formats would let researchers discover relevant benchmarks and reproduce results without rebuilding pipelines from scratch.
  • Multi-faceted, expert-reviewed metrics would reduce the risk that optimizing a single number produces models that look good on a leaderboard but fail in biological context.
  • Recurring community assessments modeled on CASP would give the field a shared signal: performance plateaus would pinpoint exactly where new data generation is needed.
  • Incentivizing reproducible workflows would make published model results verifiable, rather than accepted on the strength of a paper's reported metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the incentive alignment problem is not solved first, the proposed centralized platform may become a low-traffic archive rather than the community hub the recommendations envision, because the report itself documents that academic, pharma, and nonprofit actors want different things from benchmarks.
  • Beyond the paper: a testable extension is to compare adoption rates of benchmarks hosted on the proposed platform against decentralized efforts; if domain-specific, decentralized benchmarks grow faster, federated cross-referencing may prove more viable than centralization.
  • Beyond the paper: the CASP analogy implies that the main bottleneck for virtual cells is data design rather than model architecture, which would argue for shifting funding from model-centric projects toward benchmark-data generation and curation.
  • Beyond the paper: the same recommendations could serve as a checklist for auditing existing biological AI benchmarks, not just for creating future ones.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper reports on a CZI-hosted workshop on benchmarking AI models in biology, with a focus on AI-driven virtual cells. It identifies technical and systemic challenges (data heterogeneity, reproducibility, metric relevance, fragmented ecosystem, biases, benchmark maintenance, community development) and proposes eight recommendations, including investing in curated benchmark data, standardized tooling, multi-faceted metrics, a centralized platform, community engagement, updating/deprecating benchmarks, and cross-sector collaboration. The central claim is that the lack of standardized, cross-domain benchmarks is a key barrier to robust virtual cell models, and that the proposed recommendations will accelerate the development of such benchmarks.

Significance. If its recommendations are adopted, the paper could help coordinate community investment and resource allocation in a rapidly growing area. Its strengths are the breadth of the author list spanning academia, industry, and non-profits, and the extensive referencing of existing benchmark efforts (CASP, DREAM, OpenProblems, Polaris, CACHE, CAGI, D3R). The paper does not claim new quantitative results; it is a position/workshop report. Its value lies in synthesizing known issues and proposing an actionable agenda. However, the causal link between addressing the identified barriers and accelerating virtual cell development is asserted rather than demonstrated, and the recommendations presuppose a collective-action capacity that the paper itself shows is lacking.

major comments (2)
  1. [Community Development and Recommendations (Create a centralized platform)] The paper identifies 'differing incentive structures' and a 'fragmented ecosystem' as central barriers, yet the first concrete recommendation (create a centralized platform) assigns no actor, no funding source, no governance model, and no mechanism for sustaining the platform over time. The same gap applies to the calls to reward reproducible tooling, share data, and deprecate stale benchmarks: these list desired outcomes without specifying who changes the incentives that the paper itself describes as misaligned. Because the stated value of the paper is to guide community investment, this missing mechanism is load-bearing: if the incentive misalignment is real, a platform without an organizational home or enforcement mechanism is unlikely to be adopted or sustained. The cited precedents (DREAM, OpenProblems, Polaris, CACHE, CAGI, D3R) are largely domain-specific or project-funded, and none is shown to be a sustained cross-domain hub. The paper should either specify a possible governance/funding model or explicitly acknowledge that the collective-action problem remains open.
  2. [Introduction] The paper motivates the entire recommendation set with the claim that CASP drove the advances of AlphaFold. This is a historically plausible but post-hoc narrative; the paper provides no evidence that benchmarking was the causal driver rather than an enabler, and it does not substantiate that the same dynamic will transfer to the much harder, cross-domain, multi-modal virtual cell setting, where the paper itself notes that 'new challenges emerge while existing ones are compounded.' The recommendations may be reasonable, but the abstract's implied causal claim—that these recommendations 'will ultimately advance the field toward integrated models'—goes beyond what the paper supports. The paper should temper this claim or supply additional evidence (e.g., documented cases where benchmarking infrastructure directly accelerated progress in other cross-disciplinary fields).
minor comments (4)
  1. [Summary/Abstract] The full-text Summary lists the domains covered as 'imaging, proteomics, and genomics,' omitting transcriptomics, which appears in the abstract; the two lists should be consistent.
  2. [Biases] The citation string '31,32; 32' contains a stray semicolon and duplicated reference number; it should read '31,32.'
  3. [Author affiliations] Affiliation 7 is spelled 'Bringham and Women’s Hospital'; the correct spelling is 'Brigham and Women’s Hospital.'
  4. [Methodology] The paper does not describe the workshop's process (number of participants, selection criteria, breakout structure, or how consensus was reached), which would help readers assess the representativeness of the recommendations.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a workshop-report and recommendations document making no quantitative predictions.

full rationale

The paper contains no derivation chain with fitted parameters or predicted quantities. Its central claim is that a lack of standardized, cross-domain benchmarks impedes robust AI in biology and that community recommendations can address this. This is an argumentative position rather than a result derived from data. The cited benchmarks (CASP, DREAM, OpenProblems, Polaris, CACHE, CAGI, D3R) are offered as motivating examples of prior benchmarking ecosystems, not as inputs that entail the recommendations. No uniqueness theorem is invoked, no ansatz is smuggled in via citation, and no known result is renamed as a new finding. The only self-citation (Ash et al., ref. 12, by a co-author) is used as an example of statistical-replicability tooling and is not load-bearing. The paper explicitly acknowledges the incentive-alignment difficulty it recommends solving, so the practical concern that the recommendations lack a governance or funding mechanism is an unsupported-practical-premise critique, not circularity. Thus no step in the paper reduces to its own inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters, invented entities, or new quantities appear; the paper is a qualitative position statement relying on background assumptions about the scientific community and benchmark design.

assumptions (3)
  • domain assumption CASP-style community challenges are an effective engine for scientific progress in structure prediction and can be emulated for virtual cells.
    Invoked in the Introduction and Recommendations as inspiration; if this assumption is false, the recommendation to build CASP-like frameworks weakens.
  • domain assumption Biological data heterogeneity, noise, and small sample sizes are inherent and require special benchmarking practices.
    Treated as background consensus supported by cited single-cell and genomics literature; not proven in this paper.
  • ad hoc to paper The workshop participants' consensus is representative of the broader community's needs and priorities.
    The paper generalizes from a single workshop without describing participant selection, the agenda, or dissenting views, so this representativeness is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop." pith.science (2026). https://pith.science/paper/QKILSGSV

@misc{pith2026250710502,
  author       = {Pith},
  title        = {Pith review of: Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QKILSGSV}},
  note         = {Machine review of arXiv:2507.10502}
}
read the original abstract

Artificial intelligence holds immense promise for transforming biology, yet a lack of standardized, cross domain, benchmarks undermines our ability to build robust, trustworthy models. Here, we present insights from a recent workshop that convened machine learning and computational biology experts across imaging, transcriptomics, proteomics, and genomics to tackle this gap. We identify major technical and systemic bottlenecks such as data heterogeneity and noise, reproducibility challenges, biases, and the fragmented ecosystem of publicly available resources and propose a set of recommendations for building benchmarking frameworks that can efficiently compare ML models of biological systems across tasks and data modalities. By promoting high quality data curation, standardized tooling, comprehensive evaluation metrics, and open, collaborative platforms, we aim to accelerate the development of robust benchmarks for AI driven Virtual Cells. These benchmarks are crucial for ensuring rigor, reproducibility, and biological relevance, and will ultimately advance the field toward integrated models that drive new discoveries, therapeutic insights, and a deeper understanding of cellular systems.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 39 canonical work pages

  1. [1]

    Bunne, C. et al. How to build the virtual cell with artificial intelligence: Priorities and opportunities. Cell 187 , 7045–7063 (2024)

  2. [3]

    Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596 , 583–589 (2021)

  3. [4]

    & Saez-Rodriguez, J

    Meyer, P. & Saez-Rodriguez, J. Advances in systems biology modeling: 10 years of crowdsourcing DREAM challenges. Cell Syst. 12 , 636–653 (2021)

  4. [5]

    Luecken, M. D. et al. Defining and benchmarking open problems in single-cell analysis. Res Sq (2024) doi:10.21203/rs.3.rs-4181617/v1

  5. [6]

    Jain, Y. et al. Segmentation of human functional tissue units in support of a Human Reference Atlas. Commun. Biol. 6 , 717 (2023)

  6. [7]

    Le, T. et al. Analysis of the Human Protein Atlas weakly supervised single-Cell Classification competition. Nat. Methods 19 , 1221–1229 (2022)

  7. [8]

    Ouyang, W. et al. Analysis of the Human Protein Atlas image classification competition. Nat. Methods 16 , 1254–1261 (2019)

  8. [9]

    Ackloo, S. et al. CACHE (Critical Assessment of Computational Hit-finding Experiments): A public-private partnership benchmarking initiative to enable the development of computational methods for hit-finding. Nat. Rev. Chem. 6 , 287–295 (2022)

Show all 40 references
  1. [10]

    CAGI, the Critical Assessment of Genome Interpretation, establishes progress and prospects for computational genetic variant interpretation methods

    Critical Assessment of Genome Interpretation Consortium. CAGI, the Critical Assessment of Genome Interpretation, establishes progress and prospects for computational genetic variant interpretation methods. Genome Biol. 25 , 53 (2024)

  2. [11]

    Gathiaka, S. et al. D3R grand challenge 2015: Evaluation of protein-ligand pose and affinity predictions. J. Comput. Aided Mol. Des. 30 , 651–668 (2016)

  3. [12]

    Ash, J. R. et al. Practically significant method comparison protocols for machine learning in small molecule drug discovery. ChemRxiv (2024) doi:10.26434/chemrxiv-2024-6dbwv-v2

  4. [13]

    & Xing, E

    Song, L., Segal, E. & Xing, E. Toward AI-Driven Digital Organism: Multiscale foundation models for predicting, simulating and programming biology at all levels. arXiv [cs.AI] (2024)

  5. [14]

    Clark, T. et al. Cell Maps for Artificial Intelligence: AI-ready maps of human cell architecture from disease-relevant cell lines. bioRxivorg (2024) doi:10.1101/2024.05.21.589311

  6. [15]

    Johnson, G. T. et al. Building the next generation of virtual cells to understand cellular biology. Biophys. J. 122 , 3560–3569 (2023)

  7. [16]

    Carr, A. et al. AI: A transformative opportunity in cell biology. Mol. Biol. Cell 35 , e4 (2024)

  8. [17]

    Sun, R. et al. A perturbation proteomics-based foundation model for virtual cell construction. bioRxiv (2025) doi:10.1101/2025.02.07.637070

  9. [18]

    & Narayanan, A

    Kapoor, S. & Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns (N. Y.) 4 , 100804 (2023)

  10. [19]

    McDermott, M. B. A. et al. Reproducibility in machine learning for health research: Still a ways to go. Science Translational Medicine (2021) doi:10.1126/scitranslmed.abb1655

  11. [20]

    Luecken, M. D. et al. Benchmarking atlas-level data integration in single-cell genomics. Nature Methods 19 , 41–50 (2021)

  12. [21]

    Lähnemann, D. et al. Eleven grand challenges in single-cell data science. Genome Biology 21 , 1–35 (2020)

  13. [22]

    Jiang, R., Sun, T., Song, D. & Li, J. J. Statistics or biology: the zero-inflation controversy about scRNA-seq data. Genome Biology 23 , 1–24 (2022)

  14. [23]

    & Shi, L

    Yu, Y., Mai, Y., Zheng, Y. & Shi, L. Assessing and mitigating batch effects in large-scale omics studies. Genome Biology 25 , 1–27 (2024)

  15. [24]

    Thomas, R. L. & Uminsky, D. Reliance on metrics is a fundamental challenge for AI. Patterns (N Y) 3 , 100476 (2022)

  16. [25]

    Whalen, S., Schreiber, J., Noble, W. S. & Pollard, K. S. Navigating the pitfalls of applying machine learning in genomics. Nature reviews. Genetics 23 , (2022)

  17. [26]

    Walters, W. P. & Murcko, M. Assessing the impact of generative AI on medicinal chemistry. Nature Biotechnology 38 , 143–145 (2020)

  18. [27]

    & Prainsack, B

    Pot, M., Spahl, W. & Prainsack, B. The gender of biomedical data: Challenges for Personalised and Precision Medicine. Somatechnics 9 , 170–187 (2019)

  19. [28]

    Corpas, M. et al. Bridging genomics’ greatest challenge: The diversity gap. Cell Genom 5 , 100724 (2025)

  20. [29]

    Stoeger, T., Gerlach, M., Morimoto, R. I. & Nunes Amaral, L. A. Large-scale investigation of the reasons why potentially important genes are ignored. PLOS Biology 16 , e2006643 (2018)

  21. [30]

    Munafò, M. R. et al. A manifesto for reproducible science. Nature Human Behaviour 1 , 1–9 (2017)

  22. [31]

    Mendez, D. et al. ChEMBL: towards direct deposition of bioassay data. Nucleic Acids Res 47 , D930–D940 (2019)

  23. [32]

    Positive directions from negative results

    Gan, L. Positive directions from negative results. Journal of Cell Science (2023) doi:10.1242/jcs.261594

  24. [33]

    Vendetti, J. et al. BioPortal: an open community resource for sharing, searching, and utilizing biomedical ontologies. Nucleic Acids Res gkaf402 (2025)

  25. [34]

    Bernett, J. et al. Guiding questions to avoid data leakage in biological machine learning applications. Nat. Methods 21 , 1444–1453 (2024)

  26. [35]

    Heil, B. J. et al. Reproducibility standards for machine learning in the life sciences. Nat. Methods 18 , 1132–1135 (2021)

  27. [36]

    Maier-Hein, L. et al. Metrics reloaded: recommendations for image analysis validation. Nat Methods 21 , 195–212 (2024)

  28. [37]

    Reinke, A. et al. Understanding metric-related pitfalls in image analysis validation. Nat Methods 21 , 182–194 (2024)

  29. [38]

    Wiesenfarth, M. et al. Methods and open-source toolkit for analyzing and visualizing challenge results. Sci Rep 11 , 2369 (2021)

  30. [39]

    T., Judson, R

    Moult, J., Pedersen, J. T., Judson, R. & Fidelis, K. A large-scale experiment to assess protein structure prediction methods. Proteins 23 , ii–v (1995)

  31. [40]

    & Moult, J

    Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K. & Moult, J. Critical assessment of methods of protein structure prediction (CASP)-Round XIV. Proteins 89 , 1607–1617 (2021)

  32. [41]

    & Moult, J

    Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K. & Moult, J. Critical assessment of methods of protein structure prediction (CASP)-Round XV. Proteins 91 , 1539–1549 (2023)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.