Pith. sign in

REVIEW 3 major objections 4 minor 78 references

Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper argues that current AI systems can only automate isolated stages of scientific discovery, not the full cycle, and proposes a four-part research agenda to bridge that gap.

desk verdict A readable, broad roadmap for AI-for-science whose modularity assumption is asserted rather than argued, and that carries one clear misattribution; still worth refereeing as a position paper. read the letter →

arxiv 2412.11427 v2 pith:FYTV63DL submitted 2024-12-16 cs.LG cs.AI

classification cs.LGcs.AI
keywords scientificdiscoverygenerativeAIlargelanguagemodelsagentsbenchmarksmultimodalrepresentationlearningtheoremprovingneuro-symbolicreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a position statement about generative AI for scientific discovery. Its central claim is that current AI systems can automate pieces of the scientific process—reading literature, generating hypotheses, designing experiments, discovering equations—but no existing system can sustain an entire research loop from problem specification to experimental validation and back. The authors argue that this integration gap, not any single capability, is the bottleneck, and they lay out four research directions to close it: discovery-oriented benchmarks, science-focused AI agents, multimodal scientific representations, and unified frameworks that combine theorem proving with data-driven modeling. A sympathetic reader would care because the paper gives a concrete agenda for turning today's fragmented AI tools into collaborative partners for scientists.

What carries the argument

The central organizing object is the iterative discovery cycle shown in Figure 1, which decomposes scientific inquiry into a closed loop of specification, retrieval, generation, experimentation, and evaluation. Around that cycle the paper builds a second framework, Figure 2, for science-focused AI agents that integrates multimodal inputs, tool use, and expert feedback. The four proposed challenge areas—benchmarks, agents, multimodal representations, and theory-data unification—are the paper's proposed machinery for moving from isolated tools to a unified system.

What would settle it

A concrete test would be a single AI system that, without using the modular agent architecture proposed here, carries out a complete scientific discovery—formulating a new hypothesis, designing and running an experiment, and validating a reproducible result—in a setting where the answer is not in its training data. If such systems appear and perform sustained research despite ignoring the proposed separation of modules, the paper's bottleneck claim would be undermined. Until then, a simpler observation: if the field's fragmented tools, when simply chained together, already reproduce a known scientific result end-to-end on a standard task like rediscovering a physics law with experimental feedback, that would partially contradict the 'integration gap' assertion.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is an assessment of the field's state: more than a decade of advances in language models, symbolic reasoning, and data-driven modeling has produced strong point solutions, but no integrated system that can carry out autonomous long-term research. The authors frame scientific discovery as an iterative cycle—problem specification, context retrieval, hypothesis generation, experiment design, evaluation, and refinement—and argue that all current work attacks isolated stages of that cycle. They then propose that the field's priority should be to close the integration gap by building science-focused agents, designing benchmarks that test novel discovery rather than rediscovery, developing representations that span text, images, graphs, and numeric data, and unifying deductive reasoning with empirical modeling. The paper's forward-looking claim is that these four directions, pursued together, will yield AI systems that accelerate discovery across disciplines.

Load-bearing premise

The argument assumes that scientific discovery decomposes into the modular iterative cycle of Figure 1, and that improving each module in isolation will eventually combine into an integrated discovery system; the paper asserts this cycle without defending it.

Editorial extensions

If this is right

  • Evaluation will shift from rediscovering known laws to testing novel discovery in configurable simulated domains, reducing the risk that models simply recite memorized training data.
  • AI agents that interface with specialized scientific tools, such as reaction predictors or simulation codes, will become a standard way to combine LLM reasoning with external validation.
  • Multimodal representation learning will let discovery systems search hypotheses in low-dimensional latent spaces, making exploration of combinatorial scientific spaces more efficient.
  • Unified frameworks that combine logical reasoning with data-driven modeling will produce hypotheses that are not just predictive but derivable, improving generalization to out-of-distribution settings.
  • In the near term, the payoff is likely to be AI assistants that augment human scientists rather than fully autonomous systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable implication of the paper's modular-cycle assumption is that end-to-end benchmarks will reward systems that iterate across stages more than systems that excel at one stage.
  • If the modular decomposition holds, the field should expect diminishing returns from scaling LLMs alone; the next performance leaps would come from architectural innovations in tool integration and multimodal reasoning rather than from larger models.
  • A natural extension of the evaluation metrics (novelty, generalizability, alignment) would be a composite 'discovery score' that weights all three; the paper does not propose one, but its framework makes such a metric the obvious next step.
  • The paper's critique of memorization suggests that any benchmark with static, published answers will eventually be leaked into training sets, pointing toward dynamically generated benchmarks as a necessary, though unstated, consequence.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper is a position/vision essay on the use of generative AI for scientific discovery. It argues that despite progress in applying large language models and related techniques to literature analysis, theorem proving, experimental design, and data-driven discovery, no current AI system integrates the full cycle of scientific inquiry. The paper structures this cycle as problem specification, context retrieval, hypothesis generation, experiment design, and evaluation (Figure 1), and it proposes four research directions: improved benchmarks and evaluation, science-focused AI agents, multimodal scientific representations, and unified frameworks that combine reasoning, theorem proving, and data-driven modeling. The intended contribution is to synthesize recent advances and to identify critical gaps that should guide future research.

Significance. If the agenda is accepted, the paper could help focus investment and research effort in AI for science. Its strength is a broad and readable consolidation of recent work across several subfields, and it raises several concrete and plausible concerns, including the risk that current benchmarks reward memorization rather than discovery, the need for domain-expert involvement in evaluation, and the value of latent-space search and derivable hypotheses. The paper does not offer machine-checked proofs, new empirical results, or reproducible code; its value is as an agenda-setting essay. The significance therefore depends on the accuracy of its survey summaries and on whether its central framing of scientific discovery as a modular cycle is credible. Those two points are where the manuscript currently needs the most work.

major comments (3)
  1. [Introduction and Figure 1] The paper's central claim—that "we still lack AI systems capable of integrating the diverse cognitive processes involved in sustained scientific research and discovery"—rests on treating the Figure 1 cycle as the correct decomposition of scientific inquiry. The cycle (problem specification, context retrieval, hypothesis generation, experiment design, evaluation) is asserted rather than defended, and the claim that "most work has focused on narrow aspects of scientific reasoning in isolation" is not reconciled with the paper's own citation of The AI Scientist (Lu et al. 2024), an end-to-end system, in the same Introduction. If scientific discovery does not decompose into these independently improvable modules, the proposed directions (modular agents, stage-specific benchmarks) lose their rationale. Please add an explicit argument for the modularity assumption, or temper the claim by presenting the cycle as a heuristic and discussing why existing end-to-end attempts fall short.
  2. [Data-driven Discovery] The drug-discovery paragraph states that "recent works employed generative (Mak, Wong, and Pichika 2023; Callaway 2024) and multimodal representation learning (Gao et al. 2024) models to discover a novel antibiotic, effective against a wide range of bacteria, by searching and screening millions of molecules in the representation space (Gao et al. 2024)." This is a factual misattribution: Gao et al. (2024), DrugCLIP, is a contrastive protein-molecule representation model for virtual screening and does not report discovering an antibiotic; the other two references likewise do not report such a discovery. The sentence conflates virtual screening with generative antibiotic discovery and should be rewritten with the correct citation(s) and a clearer separation of screening from generative discovery. This matters because the survey's credibility depends on accurate summaries of the cited systems.
  3. [Theory and Data Unification] The opening of this section asserts that "most existing AI approaches to scientific tasks focus on just one of these aspects" (theoretical reasoning, empirical observation, mathematical modeling), but the paper's own survey describes systems such as ChemCrow (M. Bran et al. 2024), AtomAgents (Ghafarollahi and Buehler 2024a), and AI-Descartes (Cornelio et al. 2023) that already combine data-driven modeling with tool use, reasoning, or logical derivation. The reader needs a precise definition of the proposed "unification" and an explicit gap analysis showing why these integrated systems are insufficient; otherwise the fourth research direction is difficult to evaluate.
minor comments (4)
  1. [Benchmarks for Scientific Discovery] The sentence "they may be vulnerable to reciting or memorization by large language models... (Carlini et al. 2021; Shojaee et al. 2024b)" cites Shojaee et al. 2024b (LLM-SR), an equation-discovery method, as evidence about benchmark memorization; this citation does not support the claim and should be replaced with a relevant memorization study.
  2. [Theory and Data Unification] The heading "Reasoning discovery uncertainty in formal frameworks" appears to be missing a word; it should read "Reasoning about discovery uncertainty in formal frameworks" or similar.
  3. [References] Several reference entries are incomplete or inconsistently formatted: the entry for "Brown 2020" should be "Brown et al. 2020," the entry for "Biggio et al. 2021" ends with "Pmlr" as a publisher, and the entry for "Draft, Sketch, and Prove" lacks a year. Please normalize the bibliography.
  4. [Figure 2] The caption of Figure 2 uses circled labels (a⃝–d⃝) that are not explained in the surrounding text; please add a legend or replace them with conventional labels.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the paper is a survey/position piece whose self-citations are non-load-bearing examples of published progress.

full rationale

This paper is a survey and research-agenda essay, not a derivation with fitted parameters or a formal chain from premises to predictions. The central claim—that the field lacks integrated AI systems for sustained scientific discovery—is a gap analysis, and the proposed challenge areas are recommendations rather than consequences derived from the cited systems. The self-citations (Shojaee et al. 2024a, 2024b; Meidani et al. 2024) appear as illustrative examples of equation-discovery progress and as one supporting citation alongside Carlini et al. for the concern about memorization in benchmarks. These are published, externally evaluated results and are not load-bearing for the paper's agenda: removing them would not change the central argument that integration is missing. The modular cycle in Figure 1 is an organizing assumption, not a result claimed to follow from the cited work; an unargued assumption is a scope/correctness concern, not circularity. The DrugCLIP antibiotic misattribution in the Data-driven Discovery section is a factual issue, not a circular-reasoning issue. No step in the paper reduces by construction to its own inputs or to a self-citation chain, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is a position paper, so there are no fitted parameters or introduced entities. The axioms listed are the unstated or lightly stated presuppositions of the proposed research agenda.

assumptions (4)
  • ad hoc to paper Scientific discovery can be decomposed into a repeatable cycle of context retrieval, hypothesis generation, experiment design, and evaluation (Figure 1).
    The paper's framework assumes this modularity without argument. The entire research agenda depends on improving stages independently and then integrating them.
  • domain assumption Large language models are the appropriate substrate for building science-focused discovery agents.
    The survey and proposed agenda are centered on LLMs and their extension. Whether LLM-based agents can maintain long-horizon rigor is an open empirical question the paper does not address.
  • domain assumption Simulated domains (synthetic physics, synthetic chemistry) are valid proxies for real scientific discovery in benchmarks.
    The paper recommends configurable simulated environments (e.g., M. Bran et al. 2024, Shojaee et al. 2024b) as discovery benchmarks. Transfer from simulation to real discovery is assumed.
  • ad hoc to paper The four enumerated challenge areas are the critical bottlenecks for AI-driven scientific discovery.
    These are the authors' organizational priorities; other surveys emphasize different obstacles (e.g., data quality, reproducibility, ethics). The paper does not justify why these four are primary.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges." pith.science (2026). https://pith.science/paper/FYTV63DL

@misc{pith2026241211427,
  author       = {Pith},
  title        = {Pith review of: Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FYTV63DL}},
  note         = {Machine review of arXiv:2412.11427}
}
read the original abstract

Scientific discovery is a complex cognitive process that has driven human knowledge and technological progress for centuries. While artificial intelligence (AI) has made significant advances in automating aspects of scientific reasoning, simulation, and experimentation, we still lack integrated AI systems capable of performing autonomous long-term scientific research and discovery. This paper examines the current state of AI for scientific discovery, highlighting recent progress in large language models and other AI techniques applied to scientific tasks. We then outline key challenges and promising research directions toward developing more comprehensive AI systems for scientific discovery, including the need for science-focused AI agents, improved benchmarks and evaluation metrics, multimodal scientific representations, and unified frameworks combining reasoning, theorem proving, and data-driven modeling. Addressing these challenges could lead to transformative AI tools to accelerate progress across disciplines towards scientific discovery.

Figures

Figures reproduced from arXiv: 2412.11427 by the authors.

Figure 1
Figure 1. Overview of the AI-driven scientific discovery framework. The cycle illustrates the iterative process of scientific inquiry. The framework begins with user-defined problem specifications, retrieves relevant scientific context from literature and databases, and utilizes generative AI sys￾tems to produce new hypotheses and experimental designs. These AI-generated concepts are then evaluated and refined through experim… view at source ↗
Figure 2
Figure 2. A comprehensive framework for science-focused AI agents. The diagram illustrates ⃝a the multi-modal nature of scientific data, ⃝b the inputs for scientific tasks, ⃝c the key actions performed by AI agents in scientific discovery, and ⃝d the evaluation metrics for assessing scientific outcomes. This framework highlights the integration of diverse data sources, AI￾driven tools, and human experts in advancing scientifi… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 41 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    N.; Urban, N

    Abeer, A. N.; Urban, N. M.; Weil, M. R.; Alexander, F. J.; and Yoon, B.-J. 2024. Multi-objective latent space optimization of generative molecular design models. Patterns

  4. [4]

    Ahmed, K.; Teso, S.; Chang, K.-W.; Van den Broeck, G.; and Vergari, A. 2022. Semantic probabilistic layers for neuro-symbolic learning. Advances in Neural Information Processing Systems, 35: 29944--29959

  5. [5]

    M.; Wu, Y.; and Krenn, M

    Arlt, S.; Duan, H.; Li, F.; Xie, S. M.; Wu, Y.; and Krenn, M. 2024. Meta-Designing Quantum Experiments with Language Models. arXiv preprint arXiv:2406.02470

  6. [6]

    Baldi, P.; Sadowski, P.; and Whiteson, D. 2014. Searching for exotic particles in high-energy physics with deep learning. Nature communications, 5(1): 4308

  7. [7]

    Beltagy, I.; Lo, K.; and Cohan, A. 2019. SciBERT: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676

  8. [8]

    Biggio, L.; Bendinelli, T.; Neitz, A.; Lucchi, A.; and Parascandolo, G. 2021. Neural symbolic regression that scales. In International Conference on Machine Learning, 936--945. Pmlr

Show all 78 references
  1. [9]

    Birhane, A.; Kasirzadeh, A.; Leslie, D.; and Wachter, S. 2023. Science in the age of large language models. Nature Reviews Physics, 5(5): 277--280

  2. [10]

    B \"o hme, S.; and Nipkow, T. 2010. Sledgehammer: judgement day. In Automated Reasoning: 5th International Joint Conference, IJCAR 2010, Edinburgh, UK, July 16-19, 2010. Proceedings 5, 107--121. Springer

  3. [11]

    A.; MacKnight, R.; Kline, B.; and Gomes, G

    Boiko, D. A.; MacKnight, R.; Kline, B.; and Gomes, G. 2023. Autonomous chemical research with large language models. Nature, 624(7992): 570--578

  4. [12]

    A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M

    Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258

  5. [13]

    Brown, T. B. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165

  6. [14]

    T.; Li, Y.; Lundberg, S.; et al

    Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712

  7. [15]

    W.; Charton, F.; Nolte, N.; Wilhelm, M.; Cranmer, K.; and Dixon, L

    Cai, T.; Merz, G. W.; Charton, F.; Nolte, N.; Wilhelm, M.; Cranmer, K.; and Dixon, L. J. 2024. Transforming the bootstrap: Using transformers to compute scattering amplitudes in planar n= 4 super yang-mills theory. Machine Learning: Science and Technology

  8. [16]

    Callaway, E. 2024. Major AlphaFold upgrade offers boost for drug discovery. Nature, 629(8012): 509--510

  9. [17]

    Carlini, N.; Ippolito, D.; Jagielski, M.; Lee, K.; Tramer, F.; and Zhang, C. 2022. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646

  10. [18]

    Carlini, N.; Tramer, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, U.; et al. 2021. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), 2633--2650

  11. [19]

    Castro, E.; Godavarthi, A.; Rubinfien, J.; Givechian, K.; Bhaskar, D.; and Krishnaswamy, S. 2022. Transformer-based protein generation with regularized latent space optimization. Nature Machine Intelligence, 4(10): 840--851

  12. [20]

    Chen, A.; Wang, Z.; Vidaurre, K. L. L.; Han, Y.; Ye, S.; Tao, K.; Wang, S.; Gao, J.; and Li, J. 2024. Knowledge-Reuse Transfer Learning Methods in Molecular and Material Science. arXiv preprint arXiv:2403.12982

  13. [21]

    R.; Goncalves, J.; Clarkson, K

    Cornelio, C.; Dash, S.; Austel, V.; Josephson, T. R.; Goncalves, J.; Clarkson, K. L.; Megiddo, N.; El Khadir, B.; and Horesh, L. 2023. Combining data and theory for derivable scientific discovery with AI-Descartes. Nature Communications, 14(1): 1777

  14. [22]

    Cranmer, M.; Greydanus, S.; Hoyer, S.; Battaglia, P.; Spergel, D.; and Ho, S. 2020 a . Lagrangian neural networks. arXiv preprint arXiv:2003.04630

  15. [23]

    Cranmer, M.; Sanchez Gonzalez, A.; Battaglia, P.; Xu, R.; Cranmer, K.; Spergel, D.; and Ho, S. 2020 b . Discovering symbolic models from deep learning with inductive biases. Advances in neural information processing systems, 33: 17429--17442

  16. [24]

    De Raedt, L.; and Kimmig, A. 2015. Probabilistic (logic) programming concepts. Machine Learning, 100: 5--47

  17. [25]

    De Raedt, L.; Kimmig, A.; and Toivonen, H. 2007. ProbLog: a probabilistic prolog and its application in link discovery. In Proceedings of the 20th International Joint Conference on Artifical Intelligence, IJCAI'07, 2468–2473. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc

  18. [26]

    Edwards, C.; Zhai, C.; and Ji, H. 2021. Text2mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 595--607

  19. [27]

    Gao, B.; Qiang, B.; Tan, H.; Jia, Y.; Ren, M.; Lu, M.; Liu, J.; Ma, W.-Y.; and Lan, Y. 2024. Drugclip: Contrasive protein-molecule representation learning for virtual screening. Advances in Neural Information Processing Systems, 36

  20. [28]

    d.; and Lamb, L

    Garcez, A. d.; and Lamb, L. C. 2023. Neurosymbolic AI: The 3 rd wave. Artificial Intelligence Review, 56(11): 12387--12406

  21. [29]

    Ghafarollahi, A.; and Buehler, M. J. 2024 a . AtomAgents: Alloy design and discovery through physics-aware multi-modal multi-agent artificial intelligence. arXiv preprint arXiv:2407.10022

  22. [30]

    Ghafarollahi, A.; and Buehler, M. J. 2024 b . SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning. arXiv preprint arXiv:2409.05556

  23. [31]

    Golkar, S.; Pettee, M.; Eickenberg, M.; Bietti, A.; Cranmer, M.; Krawezik, G.; Lanusse, F.; McCabe, M.; Ohana, R.; Parker, L.; et al. 2023. xval: A continuous number encoding for large language models. arXiv preprint arXiv:2310.02989

  24. [32]

    Gu, Y.; Tinn, R.; Cheng, H.; Lucas, M.; Usuyama, N.; Liu, X.; Naumann, T.; Gao, J.; and Poon, H. 2021. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1): 1--23

  25. [33]

    M.; Rute, J.; Wu, Y.; Ayers, E

    Han, J. M.; Rute, J.; Wu, Y.; Ayers, E. W.; and Polu, S. 2021. Proof artifact co-training for theorem proving with language models. arXiv preprint arXiv:2102.06203

  26. [34]

    Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874

  27. [35]

    Huang, J.; and Chang, K. C.-c. 2023. Towards Reasoning in Large Language Models: Survey, Implication, and Reflection. In The 61st Annual Meeting Of The Association For Computational Linguistics

  28. [36]

    A.; Yin, D.; Shah, M.; Zhou, D.; Altman, R.; Wang, M.; and Cong, L

    Huang, K.; Qu, Y.; Cousins, H.; Johnson, W. A.; Yin, D.; Shah, M.; Zhou, D.; Altman, R.; Wang, M.; and Cong, L. 2024. Crispr-GPT: An LLM agent for automated design of gene-editing experiments. arXiv preprint arXiv:2404.18021

  29. [37]

    Ji, H.; Wang, Q.; Downey, D.; and Hope, T. 2024. SCIMON: Scientific Inspiration Machines Optimized for Novelty. In ACL Anthology: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 279--299. University of Illi...

  30. [38]

    Q.; Welleck, S.; Zhou, J

    Jiang, A. Q.; Welleck, S.; Zhou, J. P.; Lacroix, T.; Liu, J.; Li, W.; Jamnik, M.; Lample, G.; and Wu, Y. 2023. Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs. In The Eleventh International Conference on Learning Representations

  31. [39]

    Jumper, J.; Evans, R.; Pritzel, A.; Green, T.; Figurnov, M.; Ronneberger, O.; Tunyasuvunakool, K.; Bates, R.; Z \' dek, A.; Potapenko, A.; et al. 2021. Highly accurate protein structure prediction with AlphaFold. nature, 596(7873): 583--589

  32. [40]

    Kamienny, P.-A.; d'Ascoli, S.; Lample, G.; and Charton, F. 2022. End-to-end symbolic regression with transformers. Advances in Neural Information Processing Systems, 35: 10269--10281

  33. [41]

    Lample, G.; Lacroix, T.; Lachaux, M.-A.; Rodriguez, A.; Hayat, A.; Lavril, T.; Ebner, G.; and Martinet, X. 2022. Hypertree proof search for neural theorem proving. Advances in neural information processing systems, 35: 26337--26349

  34. [42]

    S.; Yang, J.; Glatt, R.; Santiago, C

    Landajuela, M.; Lee, C. S.; Yang, J.; Glatt, R.; Santiago, C. P.; Aravena, I.; Mundhenk, T.; Mulcahy, G.; and Petersen, B. K. 2022. A unified framework for deep symbolic regression. Advances in Neural Information Processing Systems, 35: 33985--33998

  35. [43]

    H.; and Kang, J

    Lee, J.; Yoon, W.; Kim, S.; Kim, D.; Kim, S.; So, C. H.; and Kang, J. 2020. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4): 1234--1240

  36. [44]

    Liu, P.; Ren, Y.; Tao, J.; and Ren, Z. 2024 a . Git-mol: A multi-modal large language model for molecular science with graph, image, and text. Computers in biology and medicine, 171: 108073

  37. [45]

    Y.; and Tegmark, M

    Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Solja c i \'c , M.; Hou, T. Y.; and Tegmark, M. 2024 b . Kan: Kolmogorov-arnold networks. arXiv preprint arXiv:2404.19756

  38. [46]

    T.; Foerster, J.; Clune, J.; and Ha, D

    Lu, C.; Lu, C.; Lange, R. T.; Foerster, J.; Clune, J.; and Ha, D. 2024. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292

  39. [47]

    Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35: 2507--2521

  40. [48]

    Luo, R.; Sun, L.; Xia, Y.; Qin, T.; Zhang, S.; Poon, H.; and Liu, T.-Y. 2022. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23(6): bbac409

  41. [49]

    Bran, A.; Cox, S.; Schilter, O.; Baldassari, C.; White, A

    M. Bran, A.; Cox, S.; Schilter, O.; Baldassari, C.; White, A. D.; and Schwaller, P. 2024. Augmenting large language models with chemistry tools. Nature Machine Intelligence, 1--11

  42. [50]

    B.; Rus, D.; Gan, C.; and Matusik, W

    Ma, P.; Wang, T.-H.; Guo, M.; Sun, Z.; Tenenbaum, J. B.; Rus, D.; Gan, C.; and Matusik, W. 2024. LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery. In Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J...

  43. [51]

    MacColl, H. 1897. Symbolic reasoning. Mind, 6(24): 493--510

  44. [52]

    Mak, K.-K.; Wong, Y.-H.; and Pichika, M. R. 2023. Artificial intelligence in drug discovery and development. Drug Discovery and Evaluation: Safety and Pharmacokinetic Assays, 1--38

  45. [53]

    K.; and Farimani, A

    Meidani, K.; Shojaee, P.; Reddy, C. K.; and Farimani, A. B. 2024. SNIP : Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training. In The Twelfth International Conference on Learning Representations

  46. [54]

    S.; Aykol, M.; Cheon, G.; and Cubuk, E

    Merchant, A.; Batzner, S.; Schoenholz, S. S.; Aykol, M.; Cheon, G.; and Cubuk, E. D. 2023. Scaling deep learning for materials discovery. Nature, 624(7990): 80--85

  47. [55]

    Miret, S.; and Krishnan, N. 2024. Are LLMs Ready for Real-World Materials Discovery? arXiv preprint arXiv:2402.05200

  48. [56]

    H.; Callahan, T

    Park, N. H.; Callahan, T. J.; Hedrick, J. L.; Erdmann, T.; and Capponi, S. 2024. Leveraging Chemistry Foundation Models to Facilitate Structure Focused Retrieval Augmented Generation in Multi-Agent Workflows for Catalyst and Materials Design. arXiv preprint arXiv:2408.11793

  49. [57]

    Polu, S.; and Sutskever, I. 2020. Generative language modeling for automated theorem proving. arXiv preprint arXiv:2009.03393

  50. [58]

    O.; Pitera, J

    Pyzer-Knapp, E. O.; Pitera, J. W.; Staar, P. W.; Takeda, S.; Laino, T.; Sanders, D. P.; Sexton, J.; Smith, J. R.; and Curioni, A. 2022. Accelerating materials discovery using artificial intelligence, high performance computing and robotics. npj Computational Materials, 8(1): 84

  51. [59]

    V.; and Katritch, V

    Sadybekov, A. V.; and Katritch, V. 2023. Computational approaches streamlining drug discovery. Nature, 616(7958): 673--685

  52. [60]

    Schmidt, M.; and Lipson, H. 2009. Symbolic regression of implicit equations. In Genetic programming theory and practice VII, 73--85. Springer

  53. [61]

    H.; Preuss, M.; and Waller, M

    Segler, M. H.; Preuss, M.; and Waller, M. P. 2018. Planning chemical syntheses with deep neural networks and symbolic AI. Nature, 555(7698): 604--610

  54. [62]

    Sheth, A.; Roy, K.; and Gaur, M. 2023. Neurosymbolic artificial intelligence (why, what, and how). IEEE Intelligent Systems, 38(3): 56--62

  55. [63]

    Shojaee, P.; Meidani, K.; Barati Farimani, A.; and Reddy, C. 2024 a . Transformer-based planning for symbolic regression. Advances in Neural Information Processing Systems, 36

  56. [64]

    B.; and Reddy, C

    Shojaee, P.; Meidani, K.; Gupta, S.; Farimani, A. B.; and Reddy, C. K. 2024 b . Llm-sr: Scientific equation discovery via programming with large language models. arXiv preprint arXiv:2404.18400

  57. [65]

    Si, C.; Yang, D.; and Hashimoto, T. 2024. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. arXiv preprint arXiv:2409.04109

  58. [66]

    S.; Wei, J.; Chung, H

    Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S. S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. 2023. Large language models encode clinical knowledge. Nature, 620(7972): 172--180

  59. [67]

    Topol, E. J. 2023. As artificial intelligence goes multimodal, medical applications multiply

  60. [68]

    Udrescu, S.-M.; Tan, A.; Feng, J.; Neto, O.; Wu, T.; and Tegmark, M. 2020. AI Feynman 2.0: Pareto-optimal symbolic regression exploiting graph modularity. Advances in Neural Information Processing Systems, 33: 4860--4871

  61. [69]

    Udrescu, S.-M.; and Tegmark, M. 2020. AI Feynman: A physics-inspired method for symbolic regression. Science Advances, 6(16): eaay2631

  62. [70]

    Wang, H.; Fu, T.; Du, Y.; Gao, W.; Huang, K.; Liu, Z.; Chandak, P.; Liu, S.; Van Katwyk, P.; Deac, A.; et al. 2023 a . Scientific discovery in the age of artificial intelligence. Nature, 620(7972): 47--60

  63. [71]

    Wang, H.; Yuan, Y.; Liu, Z.; Shen, J.; Yin, Y.; Xiong, J.; Xie, E.; Shi, H.; Li, Y.; Li, L.; et al. 2023 b . Dt-solver: Automated theorem proving with dynamic-tree sampling guided by proof-level value function. In Proceedings of the 61st Annual Meeting of the Association for C...

  64. [72]

    Wang, R.; Zelikman, E.; Poesia, G.; Pu, Y.; Haber, N.; and Goodman, N. 2024. Hypothesis Search: Inductive Reasoning with Language Models. In The Twelfth International Conference on Learning Representations

  65. [73]

    R.; Zhang, S.; Sun, Y.; and Wang, W

    Wang, X.; Hu, Z.; Lu, P.; Zhu, Y.; Zhang, J.; Subramaniam, S.; Loomba, A. R.; Zhang, S.; Sun, Y.; and Wang, W. 2023 c . Scibench: Evaluating college-level scientific problem-solving abilities of large language models. arXiv preprint arXiv:2307.10635

  66. [74]

    Wu, Z.; Qiu, L.; Ross, A.; Aky \"u rek, E.; Chen, B.; Wang, B.; Kim, N.; Andreas, J.; and Kim, Y. 2023. Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks. arXiv preprint arXiv:2307.02477

  67. [75]

    Xu, M.; Yuan, X.; Miret, S.; and Tang, J. 2023. Protst: Multi-modality learning of protein sequences and biomedical texts. In International Conference on Machine Learning, 38749--38767. PMLR

  68. [76]

    J.; and Anandkumar, A

    Yang, K.; Swope, A.; Gu, A.; Chalamala, R.; Song, P.; Yu, S.; Godil, S.; Prenger, R. J.; and Anandkumar, A. 2024. Leandojo: Theorem proving with retrieval-augmented language models. Advances in Neural Information Processing Systems, 36

  69. [77]

    Zhang, D.; Hu, Z.; Zhoubian, S.; Du, Z.; Yang, K.; Wang, Z.; Yue, Y.; Dong, Y.; and Tang, J. 2024. SciGLM: Training Scientific Language Models with Self-Reflective Instruction Annotation and Tuning. arXiv:2401.07950

  70. [78]

    Zheng, W.; Li, J.; and Zhang, Y. 2023. Desirable molecule discovery via generative latent space exploration. Visual Informatics, 7(4): 13--21

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.