REVIEW 3 major objections 3 minor 26 references
Virtual Cells: Predict, Explain, Discover
T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Virtual cells must predict, explain, discover to aid drug discovery.
desk verdict A solid perspective worth reading—the P-E-D framework is a real contribution, but the causal identifiability issue is not confronted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Predict-Explain-Discover (P-E-D) capability triad, supported by three design principles. Principle 1: predict relative changes, meaning the model's output is a change conditioned on the measured state of the cell before perturbation. Principle 2: explain perturbations as dynamic changes to a minimal set of key biomolecular interactions, using ML-based structural predictions and targeted atomistic simulation as anchors. Principle 3: lab-in-the-loop falsification, in which the virtual cell is treated as a falsifiable theory and agentic systems design experiments to falsify its hypotheses, with results updating the model. A fourth piece of machinery is the proposed benchmarking framework, which maps capabilities (specific, core, and unobserved biology; intrinsic and extrinsic context; time; causality) onto three performance levels for virtual cells.
What would settle it
A concrete test: take a virtual cell's mechanistic explanation for a drug's effect, say the claim that the response is caused by inhibition of a specific kinase that then rewires transcription, and in the lab knock out or chemically silence that kinase before applying the drug; if the observed response in the silenced cell is the same as with the drug alone, the claimed key interaction is not load-bearing for that response, and the explanation is falsified, with a benchmark built from many such prospective tests settling whether the P-E-D bar is reachable.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that prediction alone is not enough: a therapeutically relevant virtual cell must explain its predictions as modifications to key biomolecular interactions, and must use that understanding to discover novel, actionable biology. The authors position the virtual cell as a falsifiable theory of human cellular physiology, built not by simulating all roughly one hundred trillion atoms of a eukaryotic cell from first principles, but by training AI/ML models on interventional omics data, conditioning predictions on the pre-perturbation cell state, anchoring explanations with ML-based structural biology and targeted molecular dynamics, and refining the model through a lab-in-the-loop cycle in which agents design experiments to falsify the model's claims. They further claim that this framework, with explicit performance levels for observational, contextual, and explanatory capabilities, is the right scaffold for benchmarking progress and for building virtual tissues, organs, and patients.
Load-bearing premise
The load-bearing premise is that training AI models on large collections of before-and-after measurements from perturbed cells can yield both accurate predictions of cellular responses and reliable explanations couched in key biomolecular interactions, even though the models never simulate the cell from physical first principles; if the data alone cannot reveal which molecular interactions actually cause a response, the explanation and discovery capabilities lose their foundation even when prediction succeeds.
Editorial extensions
If this is right
- If the P-E-D triad is the right bar, current perturbation-prediction models that only forecast gene expression or morphology are incomplete: they must add falsifiable mechanistic explanations before they count as virtual cells.
- Benchmarks that measure predictive accuracy on transcriptomic readouts alone would be demoted in favor of suites spanning multiple modalities, cellular contexts, perturbation types, and the explanation and discovery capabilities.
- The lab-in-the-loop design turns virtual cell development into an active-learning problem: experiments are chosen to falsify model claims, so progress is measured by how quickly prediction and explanation improve per experiment.
- Because the framework conditions on the pre-perturbation state and predicts relative changes, models become trainable across heterogeneous datasets without aggressive batch correction.
- The same Predict-Explain-Discover logic extends to virtual tissues, organs, and patients, meaning patient-level models should predict clinical outcomes and explain them via lower-level mechanisms.
Reading between the lines
- An implication the authors leave implicit: the 'explain' requirement pushes evaluation away from post-hoc interpretability tools like saliency maps and toward counterfactual consistency, so a model that gives the right answer but cannot say which key interaction change caused it would fail their bar.
- A testable extension would be to score explanations prospectively: take a virtual cell's predicted key interaction for a drug, disrupt that interaction in the lab, and ask whether the observed response matches the prediction, making explanation quality a measurable quantity.
- The vision implies a data-strategy shift, since the largest transcriptomic perturbation atlases alone may be insufficient for explanations that require protein-level and post-translational information, so benchmarks may need to combine phenomics with targeted proteomic or structural measurements.
- Applied to virtual patients, the framework implies a need for interventional patient-level data and a much harder causal inference problem: explaining clinical outcomes must bridge patient-level observations to cell-level mechanisms.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This perspective paper argues that virtual cells for drug discovery should be developed and evaluated according to three core capabilities: predicting the functional response of cells to perturbations, explaining those responses in terms of key biomolecular interactions and mechanisms, and discovering novel biological insights through lab-in-the-loop experimentation. The authors propose design principles (predicting relative changes, explaining via dynamic interaction changes, and falsification-driven iteration), introduce a capability taxonomy with performance levels for benchmarking, and suggest that the same Predict-Explain-Discover framework could extend to virtual tissues, organs, and patients. The paper contains no new empirical results and is explicitly framed as a vision statement.
Significance. If the proposed framework is adopted, it could provide a useful common vocabulary and benchmarking structure for a rapidly growing field. The paper's notable strengths include its explicit emphasis on falsifiability, its caution about simple baselines and data leakage, its call for prospective evaluation, and its recognition that transcriptomic benchmarks alone are insufficient. The authors are also candid about the current gap between vision and reality, repeatedly labeling parts of the proposal as aspirational. The framework is not circular: its logical validity does not depend on the authors' own prior results. However, the central claim that explanations must be mechanistically grounded is defended only loosely, and the paper does not confront the identifiability problems that undermine causal explanation from sparse interventional omics data. Given the paper's role as a perspective, this is a load-bearing gap that should be addressed before the framework can serve as a reliable guide for the field.
major comments (3)
- [Section 2.2 / Box 2] The abstract and Box 1 define 'explain' as identifying 'key biomolecular interactions, causal pathways, and context-dependent regulatory mechanisms', which implies a causal, mechanistic account. Yet Box 2 explicitly allows 'structural-statistical explanations' that are valid 'even in the absence of complete mechanistic detail', and Section 2.2 states that 'we do not require virtual cells to replicate canonical biological pathways' and that they 'may uncover alternative structures that better fit the data'. As written, these statements are compatible with a model that generates post-hoc, non-causal narratives, which collapses the distinction between the proposed Explain capability and ordinary model interpretability. The paper should either commit to a falsifiable, mechanistically grounded notion of explanation and state the assumptions under which such explanations are identifiable (e.g., correct intervention targets, faithfulness, no unobserved confounders), or explicitly reposition explanations as model-based hypotheses that are useful for guiding experiments rather than as recovered ground-truth mechanisms.
- [Section 3.3 / Table 2] The benchmark framework lists 'Predicts response causally' as a capability and assigns it to VC level 3, but it provides no evaluation protocol or measurable criterion for this capability. The text only acknowledges that 'Causal learning, however, is a notoriously difficult problem in general, and particularly so in the high-dimensional omics datasets available within biology.' Because the paper argues that benchmarks should shape the field, a capability that cannot currently be evaluated is not actionable. The authors should propose concrete assessment strategies, such as prospective interventional validation, consistency across contexts, counterfactual agreement with known perturbation outcomes, or explicit criteria for what would count as a causal explanation.
- [Section 3 / Section 2.3] There is a tension between Section 2.3, which presents Discover as a distinct lab-in-the-loop process involving active learning and falsification, and Section 3, which states that 'we view Discover not as a separate axis to benchmark in isolation, but as the natural consequence of predictive and explanatory models applied to therapeutic contexts.' If discovery is a consequence of prediction and explanation, then the design principle of lab-in-the-loop falsification is not really a third capability but a workflow built on the first two. The paper should clarify whether Discover is an independent capability with its own evaluation criteria or a downstream process, because the title and central P-E-D framing depend on this distinction.
minor comments (3)
- [Section 1.1] The first example contains a typo: 'V orinostat' should be 'Vorinostat'.
- [References] The reference to Wenkel et al. (2025) is incomplete: the arXiv URL is given as 'https://arxiv.org/abs/XXXX.XXXXX' with a placeholder ID and the paper is marked as 'Preprint, not yet peer-reviewed'. If the paper is not yet publicly available, this reference should be marked as 'in preparation' or omitted.
- [Section 2.2 / Appendix A.1] The statement in Section 2.2 that virtual cells 'may uncover alternative structures that better fit the data' sits uneasily with Appendix A.1's principle to 'Be biologically consistent', which says models 'should recapitulate widely-accepted biology'. The authors should explain how these two standpoints are reconciled, for example by distinguishing between well-established pathway knowledge and flexible mechanistic hypotheses.
Circularity Check
No significant circularity: the paper's normative P-E-D claim is argued from drug-discovery goals, and self-citations are illustrative rather than load-bearing.
full rationale
This is a perspective/position paper, not a derivation. Its central claim—that virtual cells should predict functional responses, explain them in terms of key biomolecular interactions, and enable lab-in-the-loop discovery—is argued from the stated drug-discovery goal and presented as a design principle; it is not derived from any fitted quantity, benchmark result, or self-cited theorem. The paper's self-citations (TxPert, RxRx3, GFlowNets, Recursion's automated lab) appear as supporting examples of capabilities that could be achieved, not as premises that force the P-E-D conclusion. For instance, the TxPert reference is used only to illustrate relative-change prediction in Figure 3 and Design principle 1, and the GFlowNets citation supports active-learning methodology in Section 2.3; neither carries the argument's normative weight. The TxPert reference has a placeholder arXiv identifier and is labeled 'Preprint, not yet peer-reviewed'; this is a provenance/completeness defect, but it does not make the argument circular because the paper does not rest on TxPert's results to establish the P-E-D framework. No equation or definition in the paper reduces a claimed output to an input: 'explain' is stipulated in Box 2 and Section 2.2 as a pluralistic, falsifiable account, and the design principles are recommendations rather than derivations. Appendix A's cautions about simple baselines, data leakage, and reusable transcriptomic benchmarks further indicate that the evaluation framework is meant to be independently tested rather than assumed. Verdict: no significant circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption Functional responses of cells are adequately captured by changes in omics readouts such as transcriptomics and phenomics.
- domain assumption AI/ML models trained on massive interventional datasets can approximate cellular responses without full mechanistic simulation.
- ad hoc to paper Explanations in terms of modifications to key biomolecular interactions are necessary and sufficient for therapeutically useful virtual cells.
- domain assumption Lab-in-the-loop falsification improves virtual cell accuracy and leads to new discoveries.
- ad hoc to paper Benchmarks defined by the proposed capability taxonomy can guide progress toward useful virtual cells.
Cite this review
Pith. "Pith review of Virtual Cells: Predict, Explain, Discover." pith.science (2026). https://pith.science/paper/GKZGPWL6
@misc{pith2026250514613,
author = {Pith},
title = {Pith review of: Virtual Cells: Predict, Explain, Discover},
year = {2026},
howpublished = {\url{https://pith.science/paper/GKZGPWL6}},
note = {Machine review of arXiv:2505.14613}
}
read the original abstract
Drug discovery is fundamentally a process of inferring the effects of treatments on patients, and would therefore benefit immensely from computational models that can reliably simulate patient responses, enabling researchers to generate and test large numbers of therapeutic hypotheses safely and economically before initiating costly clinical trials. Even a more specific model that predicts the functional response of cells to a wide range of perturbations would be tremendously valuable for discovering safe and effective treatments that successfully translate to the clinic. Creating such virtual cells has long been a goal of the computational research community that unfortunately remains unachieved given the daunting complexity and scale of cellular biology. Nevertheless, recent advances in AI, computing power, lab automation, and high-throughput cellular profiling provide new opportunities for reaching this goal. In this perspective, we present a vision for developing and evaluating virtual cells that builds on our experience at Recursion. We argue that in order to be a useful tool to discover novel biology, virtual cells must accurately predict the functional response of a cell to perturbations and explain how the predicted response is a consequence of modifications to key biomolecular interactions. We then introduce key principles for designing therapeutically-relevant virtual cells, describe a lab-in-the-loop approach for generating novel insights with them, and advocate for biologically-grounded benchmarks to guide virtual cell development. Finally, we make the case that our approach to virtual cells provides a useful framework for building other models at higher levels of organization, including virtual patients. We hope that these directions prove useful to the research community in developing virtual models optimized for positive impact on drug discovery outcomes.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[3]
doi: 10.1038/s41592-024-02241-6. URL https://doi.org/10. 1038/s41592-024-02241-6 . Chen, Y. and Zou, J. Genept: a simple but effective foundation model for genes and cells built from chatgpt. bioRxiv, pp. 2023–10,
-
[8]
doi: https://doi.org/10.1016/j.shpsc.2016.02.002
ISSN 1369-8486. doi: https://doi.org/10.1016/j.shpsc.2016.02.002. URL https://www.sciencedirect. com/science/article/pii/S1369848616000200. Grimme, S., Hansen, A., Brandenburg, J. G., and Bannwarth, C. Dispersion-corrected mean-field electronic structure methods. Chemical reviews, 116(9):5105–5154,
-
[9]
doi: 10.1101/2024.07. 01.600583. URL https://www.biorxiv.org/content/early/2024/12/31/2024.07.01.600583. Heiser, K., McLean, P. F., Davis, C. T., Fogelson, B., Gordon, H. B., Jacobson, P., Hurst, B., Miller, B., Alfa, R. W., Earnshaw, B. A., et al. Identification of potential treatments for covid-19 through artificial intelligence-enabled phenomic analysi...
doi:10.1101/2024.07 2024
-
[10]
Generative modeling of molecular dynamics trajectories
Jing, B., St¨ ark, H., Jaakkola, T., and Berger, B. Generative modeling of molecular dynamics trajectories. arXiv preprint arXiv:2409.17808 ,
-
[12]
ISSN 1744-4292. doi: 10.15252/msb.202211517. URL http: //dx.doi.org/10.15252/msb.202211517. Machamer, P., Darden, L., and Craver, C. F. Thinking about mechanisms. Philosophy of science , 67 (1):1–25,
-
[13]
Maritan, M., Autin, L., Karr, J., Covert, M
URL https://arxiv.org/abs/2504.20955. Maritan, M., Autin, L., Karr, J., Covert, M. W., Olson, A. J., and Goodsell, D. S. Building structural models of a whole mycoplasma cell. Journal of molecular biology , 434(2):167351,
-
[14]
Omenn, G. S., Lane, L., Overall, C. M., Lindskog, C., Pineau, C., Packer, N. H., Cristea, I. M., Weintraub, S. T., Orchard, S., Roehrl, M. H. A., Nice, E., Guo, T., Van Eyk, J. E., Liu, S., Bandeira, N., Aebersold, R., Moritz, R. L., and Deutsch, E. W. The 2023 report on the proteome from the hupo human proteome project. Journal of Proteome Research, 23(2...
work page 2023
-
[15]
doi: 10.1021/acs.jproteome.3c00591
ISSN 1535-3907. doi: 10.1021/acs.jproteome.3c00591. URL http://dx.doi.org/10.1021/acs.jproteome.3c00591. Pang, Y. T., Kuo, K. M., Yang, L., and Gumbart, J. C. Deeppath: Overcoming data scarcity for protein transition pathway prediction using physics-based deep learning. bioRxiv, pp. 2025–02,
Show all 26 references
-
[16]
doi: 10.1093/nar/gkae1142
ISSN 1362-4962. doi: 10.1093/nar/gkae1142. URL https://doi.org/10.1093/nar/gkae1142. Rafelski, S. M. and Theriot, J. A. Establishing a conceptual framework for holistic cell states and state transitions. Cell, 187(11):2633–2651,
-
[17]
The human cell atlas white paper
Regev, A., Teichmann, S., Rozenblatt-Rosen, O., Stubbington, M., Ardlie, K., Amit, I., Arlotta, P., Bader, G., Benoist, C., Biton, M., et al. The human cell atlas white paper. arXiv preprint arXiv:1810.05192,
-
[18]
doi: 10.1038/s41587-023-01905-6
ISSN 1546-1696. doi: 10.1038/s41587-023-01905-6. URL http://dx.doi.org/10.1038/s41587-023-01905-6 . Ross, L. N. Causal concepts in biology: How pathways differ from mechanisms and why it matters. The British Journal for the Philosophy of Science ,
-
[19]
Goal-conditioned gflownets for controllable multi-objective molecular design
Roy, J., Bacon, P.-L., Pal, C., and Bengio, E. Goal-conditioned gflownets for controllable multi-objective molecular design. arXiv preprint arXiv:2306.04620 ,
-
[21]
Modeling and predicting single-cell multi-gene perturbation responses with sclambda
Wang, G., Liu, T., Zhao, J., Cheng, Y., and Zhao, H. Modeling and predicting single-cell multi-gene perturbation responses with sclambda. bioRxiv, pp. 2024–12,
2024
-
[22]
Preprint, not yet peer-reviewed
URL https://arxiv.org/abs/XXXX.XXXXX. Preprint, not yet peer-reviewed. Winnifrith, A., Outeiral, C., and Hie, B. Generative artificial intelligence for de novo protein design. arXiv preprint arXiv:2310.09685 ,
-
[23]
Boltz-1: Democratizing biomolecular interaction modeling
Wohlwend, J., Corso, G., Passaro, S., Reveiz, M., Leidal, K., Swiderski, W., Portnoi, T., Chinn, I., Silterra, J., Jaakkola, T., et al. Boltz-1: Democratizing biomolecular interaction modeling. bioRxiv, pp. 2024–11,
2024
-
[24]
A., Borja, R
Zhang, J., Ubas, A. A., Borja, R. de, Svensson, V., Thomas, N., Thakar, N., Lai, I., Winters, A., Khan, U., Jones, M. G., Tran, V., Pangallo, J., Papalexi, E., Sapre, A., Nguyen, H., Sanderson, O., Nigos, M., Kaplan, O., Schroeder, S., Hariadi, B., Marrujo, S., Salvino, C. C. ...
-
[25]
org/abs/2409.07594
URL https://arxiv. org/abs/2409.07594. 24 Virtual Cells: Predict, Explain, Discover A Considerations for benchmarking virtual cells Here we present both scientific and practical considerations for benchmarking virtual cells. We begin by acknowledging that creating useful bench...
2012 arXiv
-
[2009]
doi: 10.1007/s10838-009-9091-3
ISSN 1572-8587. doi: 10.1007/s10838-009-9091-3. URL http://dx.doi.org/10. 1007/s10838-009-9091-3 . Cretu, M., Harris, C., Igashov, I., Schneuing, A., Segler, M., Correia, B., Roy, J., Bengio, E., and Lio, P. Synflownet: Design of diverse and novel molecules with synthesis cons...
-
[2011]
On the scalability of gnns for molecular graphs
Sypetkowski, M., Wenkel, F., Poursafaei, F., Dickson, N., Suri, K., Fradkin, P., and Beaini, D. On the scalability of gnns for molecular graphs. arXiv preprint arXiv:2404.11568 ,
-
[2012]
S., Bobrovskiy, D., Zimmermann, L., Becker, S., Palma, A., Dony, L., Tejada- Lapuerta, A., Huguet, G., Lin, H.-C., et al
Klein, D., Fleck, J. S., Bobrovskiy, D., Zimmermann, L., Becker, S., Palma, A., Dony, L., Tejada- Lapuerta, A., Huguet, G., Lin, H.-C., et al. Cellflow enables generative single-cell phenotype modeling with flow matching. bioRxiv, pp. 2025–04,
2025
-
[2016]
doi: 10.1088/0953-8984/ 28/39/393001
ISSN 1361-648X. doi: 10.1088/0953-8984/ 28/39/393001. URL http://dx.doi.org/10.1088/0953-8984/28/39/393001. 18 Virtual Cells: Predict, Explain, Discover Corfield, D., Sch¨ olkopf, B., and Vapnik, V. Falsificationism and statistical learning theory: Comparing the popper and vap...
-
[2018]
already offer strong baselines when making predictions about transcriptomic data (Bendidi et al., 2024). We therefore strongly recommend that all virtual cell models begin their evaluations by comparing to simple baselines in order to assess whether they can sufficiently repre...
2024
-
[2021]
B., Mesbahi, Y
Bendidi, I., Whitfield, S., Kenyon-Dean, K., Yedder, H. B., Mesbahi, Y. E., Noutahi, E., and Denton, A. K. Benchmarking transcriptomics foundation models for perturbation analysis: one pca still rules them all. arXiv preprint arXiv:2410.13956 ,
-
[2023]
URL https://www.biorxiv.org/content/early/2023/02/08/2023.02.07.527350
doi: 10.1101/2023.02.07.527350. URL https://www.biorxiv.org/content/early/2023/02/08/2023.02.07.527350. Georgouli, K., Yeom, J.-S., Blake, R. C., and Navid, A. Multi-scale models of whole cells: progress and challenges. Frontiers in Cell and Developmental Biology , 11:1260507,
2023 doi
-
[2024]
Chai-1: Decoding the molecular interactions of life
Boitreaud, J., Dent, J., McPartlon, M., Meier, J., Reis, V., Rogozhonikov, A., and Wu, K. Chai-1: Decoding the molecular interactions of life. BioRxiv, pp. 2024–10,
2024
-
[2025]
net/forum?id=uvHmnahyp1
URL https://openreview. net/forum?id=uvHmnahyp1. Cuccarese, M. F., Earnshaw, B. A., Heiser, K., Fogelson, B., Davis, C. T., McLean, P. F., Gordon, H. B., Skelly, K.-R., Weathersby, F. L., Rodic, V., et al. Functional immune mapping with deep-learning enabled phenomics applied ...
2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.