Pith. sign in

REVIEW 4 major objections 5 minor 91 references

Large language models for partial differential equation workflows

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read LLMs earn their place in PDE research by wiring together the workflow, not by solving equations themselves.

desk verdict A useful organizational review of LLM-for-PDE work with a sensible taxonomy; the abstract overpromises somewhat but the body is careful and the gaps are fixable. read the letter →

arxiv 2608.03600 v1 pith:W4HXHL7N submitted 2026-08-04 cs.AI

classification cs.AI
keywords largelanguagemodelspartialdifferentialequationsscientificworkflowautomationequationdiscoveryPDEsolvergenerationPDE-constrainedoptimizationAItaxonomy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review argues that large language models are most credible in partial differential equation research as workflow-level coordinators rather than as standalone numerical solvers. It organizes the emerging field into three stages—Discovery, Solving, and Optimization—and shows that LLMs are currently used to connect natural-language intent, mathematical representations, scientific knowledge, code, computational tools, and feedback across those stages. The review's central claim is that the value of LLMs lies in making expert-designed procedures more executable, adaptive, and inspectable, while numerical solvers and physical validation remain essential. It also identifies the field's main limitations: scarce high-quality datasets and benchmarks, especially for knowledge discovery and real-world applications, and a persistent gap between simulation results and practical scientific and engineering systems. If this picture is right, progress should be measured not by isolated task accuracy but by whether LLM assistance reduces expert burden and preserves numerical and physical credibility across complete workflows.

What carries the argument

The organizing device is the Discovery–Solving–Optimization workflow taxonomy, built on the abstract PDE problem F(x,t;u,∂_t u,∇u,…;λ)=0 with initial and boundary conditions. Each stage corresponds to a different unresolved component of this same mathematical object: Discovery identifies or formulates F, Solving computes the state u, and Optimization selects λ to minimize an objective L under the PDE constraint. This taxonomy does the argument's work by turning a heterogeneous set of LLM-for-PDE systems into comparable categories defined by where the LLM intervenes and what purpose it serves, and it grounds the review's central claim that LLMs contribute at the workflow level rather than as

What would settle it

A direct test would be to take a representative sample of the surveyed systems, run them on held-out, unfamiliar PDE problems with new geometries, boundary conditions, software interactions, and physical regimes, and measure whether expert intervention is actually reduced and whether failures can be repaired without human framing. If such systems succeeded end-to-end while embedding the LLM as a standalone solver component rather than as a workflow orchestrator, the claim that LLMs contribute credibly only at the workflow level would be weakened.

Watch

Extended reading notes

Core claim

The paper's central discovery is a field-level pattern: LLMs are entering PDE scientific computing at the level of the workflow, not the solver. Across equation discovery, solver construction, and optimization, LLMs act as interfaces that translate informal scientific intent into executable computational pipelines, generate and revise code or configurations, and interpret solver feedback. The review codifies this pattern in a three-stage taxonomy—Discovery (identifying or formulating the PDE operator), Solving (computing the state given the PDE), and Optimization (choosing parameters, controls, or designs subject to the PDE constraint)—and maps representative systems onto it. The authors con

Load-bearing premise

The review's central claim rests on assuming that its three-stage Discovery–Solving–Optimization taxonomy and the representative papers chosen in Table 1 faithfully capture the emerging field; if substantial LLM-for-PDE work falls outside this workflow pipeline, or if the surveyed systems mostly reflect curated demonstrations rather than robust end-to-end practice, the characterization of what LLMs currently do would overgeneralize.

Editorial extensions

If this is right

  • Evaluation of LLM-based PDE systems should emphasize workflow-level outcomes—executability, physical validity, inspectability, robustness, and reduction of expert effort—rather than solution accuracy alone.
  • Benchmarks and datasets for the Discovery and Optimization stages need substantial investment, since they lag behind Solving-stage benchmarks and currently force reliance on curated demonstrations.
  • LLM assistance will redistribute expert effort toward modelling choices, challenge of assumptions, and design of informative tests, rather than eliminating the need for numerical expertise.
  • Reliable deployment requires preserving assumptions, code changes, solver states, and validation evidence over long workflows, and moving from reacting to predefined errors toward identifying scientifically meaningful failure modes.
  • When surrogate models serve as evaluators, LLM-guided search may exploit their weaknesses; high-fidelity numerical and physical validation must remain coupled to the optimization loop.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable prediction follows: systems that combine provenance-aware memory with executable tool use should show larger gains on unfamiliar, long-horizon PDE workflows than on curated benchmark cases, where current systems already do well.
  • The taxonomy implies a difficulty ordering the field has not yet formalized: Discovery-stage scientific reasoning is the least mature and may benefit most from benchmarks that measure downstream usability of a proposed formulation, not just expression recovery.
  • If the workflow-level claim is right, a useful practical metric would be reduction in human interventions per completed simulation or design task; existing executability and accuracy metrics only partially capture this.
  • The review's framing suggests that LLM-assisted optimization should be treated as a proposal mechanism operating inside a solve–evaluate–revise loop, which makes surrogate-model exploitation a central risk worth explicit diagnostic study.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This review organizes the emerging literature on LLM use in PDE research into three workflow stages—Discovery, Solving, and Optimization—and maps representative systems onto this taxonomy in Table 1. Its central claim, stated in the Abstract and §7, is that current LLM-based PDE systems act primarily as workflow-level interfaces: they connect natural-language intent, mathematical representations, scientific knowledge, code, solver outputs, and feedback, rather than replacing numerical solvers. The paper describes representative systems in each stage, discusses benchmarks and evaluation metrics (Table 2), and identifies open challenges such as dependence on expert-designed structures, long-horizon coherence, and the gap between simulation and real-world validity. The overall tone is appropriately cautious, and the authors explicitly acknowledge that current systems depend strongly on expert framing and validation.

Significance. If the review's framing is accepted, it provides a useful conceptual organization for a fast-growing and scattered field. The workflow-level perspective is a genuine contribution: it shifts attention from isolated accuracy comparisons to the coordination problem in scientific computing, and Table 2's separation of evaluation metrics by workflow direction is practically valuable. The paper gives explicit credit to the field's own limitations, including the persistent need for expert-defined operators, interfaces, and validation pipelines. However, the review's evidentiary basis is not systematic: the selection of papers in Table 1 lacks a stated search protocol, inclusion criteria, or annotation methodology, and the body itself contains systems that do not fit the workflow-level characterization. The central generalization therefore currently rests more on the authors' organizing scheme than on a robust empirical survey. The contribution is timely and likely to be useful as a position piece, but the strength of the claims must be aligned with the strength of the evidence.

major comments (4)
  1. [§4.3 and Table 1] The central claim that 'current systems act primarily as workflow-level interfaces' (Abstract; §7) depends on Table 1 being representative of the field, but the review provides no search protocol, database sources, inclusion/exclusion criteria, screening procedure, or annotation scheme for the 40+ systems mapped in the table. The taxonomy is defined in §2 and the table is then assembled to fit it, making the taxonomy not independently testable from the evidence presented. As a review, this is load-bearing: without a reproducible selection method, the reader cannot tell whether the workflow-level conclusion reflects the literature or the authors' sampling. Please add a Methods subsection describing how Table 1 was constructed, or explicitly reframe the paper as proposing a taxonomy with illustrative examples rather than as an empirical characterization of 'current systems'.
  2. [§7 and reference list] There is an internal inconsistency between the text and the table that directly affects the central claim. In §4.3, UPS, Unisolver, FLUID-LLM, Text2PDE, and FLUID-GPT are described as 'less workflow-oriented,' with the LLM contribution 'mainly in representation transfer or cross-modal conditioning rather than in reasoning over the solver workflow itself.' Yet these systems are categorized in Table 1 under 'Neural solvers' in the Solving Stage and are then folded into the general statement that current systems act as workflow-level interfaces. Either these systems should be placed in a separate category (e.g., 'representation/conditioning models') so that the workflow-level conclusion applies only to the genuinely workflow-oriented subset, or the conclusion should be explicitly qualified to acknowledge that a substantial fraction of the Solving-stage literature is not workflow-level. As w
  3. [§6 and §8] Many of the load-bearing examples cited in support of the current-state claim are unreviewed preprints or papers with incomplete bibliographic information (including [37], [40], [48], [50], [58], [59], [63], [67], [71], [72], [82], [83], and [86]). For a review that concludes 'the most credible contribution of LLMs to PDE research lies in supporting and coordinating the workflow as a whole,' the reliability of the primary studies matters. Unreviewed preprints may be reasonable to cite in an emerging area, but the review should either tag their status in Table 1, restrict the current-state claim to peer-reviewed/archival work, or perform a sensitivity check showing the conclusion is robust when preprints are excluded. Without this, the word 'credible' in the central claim is not fully earned.
  4. [Multiple] The paper acknowledges in §6 that current evaluation is uneven and that benchmarks often emphasize proxy outcomes rather than expert burden reduction, but it treats these as future benchmark problems rather than as constraints on the review's own current-state generalization. The Abstract states that 'current systems act primarily as workflow-level interfaces' without the qualifications that appear later (e.g., §7: 'current systems still operate within structures largely designed by domain experts'; §4.3: several systems are not workflow-oriented). Please calibrate the Abstract and Conclusion so that the central claim is presented as a proposed interpretation of a selected, heterogeneous body of work, not as a settled empirical finding. A formulation such as 'workflow-level assistance is the most distinctive and promising role demonstrated so far, but the evidence base is limited and het
minor comments (5)
  1. [§2, Fig. 1] The terminology is inconsistent: §2 and the text use 'Discovery–Solving–Application' in the opening paragraphs, while the section headings and Table 1 use 'Optimization Stage' and 'LLMs for the Optimization Stage.' Figure 1 labels the third stage 'Application' in panel (a) and 'Application Stage' in the caption, but the text refers to 'Optimization.' Please unify the terms.
  2. [Table 2] The columns mix different kinds of resources (benchmarks, datasets, fine-tuning corpora, workflow systems). Consider adding a 'Status' column indicating whether each resource is peer-reviewed, a preprint, or a project artifact, and whether it has released code/data. This would help readers judge the maturity of each direction.
  3. [§4.3] The phrase 'less workflow-oriented' is used to describe UPS, Unisolver, FLUID-LLM, Text2PDE, and FLUID-GPT, but these systems are diverse: some use LLM-style architectures but are not actually LLM-assisted workflows, while others use text conditioning. A brief sentence defining what counts as 'workflow-oriented' (e.g., presence of an active LLM that makes decisions across multiple components) would clarify the boundary.
  4. [References] Several references have incomplete venue information or placeholder journal names (e.g., [66] and [81] list 'Journal Name'; [86] is an arXiv-style entry without an identifier). Please complete or annotate these entries. Also, some citations (e.g., [45] with 'Nature, 655:497–505, 2026') appear unusual for the stated publication date; verifying bibliographic accuracy would strengthen the review.
  5. [Eq. (1)-(3)] The mathematical notation is generally clear, but the role of λ is overloaded: in Eq. (1) it denotes system parameters, in Eq. (2) it denotes optimization/control variables, and in Eq. (3) it is a decision variable with constraints C(u,λ) ≤ 0. A short notational remark distinguishing physical parameters from control/design variables would avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the workflow taxonomy is an organizing frame, and the central claim is an externally grounded qualitative synthesis with explicit internal qualifications.

full rationale

The paper is a qualitative review, not a derivation. Its central claim—that LLMs currently assist PDE research by coordinating workflows across Discovery, Solving, and Optimization—is a synthesis of external primary studies (e.g., [37,39,40,42,62,71,81]) and is not derived from the taxonomy by construction. Section 2 introduces the three-stage taxonomy as an organizing scheme ('This Review therefore organizes LLMs for PDE scientific computing into three categories'), and Table 1 maps independent papers into it; the conclusion in Section 8 restates the survey's qualitative finding, not a formal consequence of the taxonomy. No fitted parameter is renamed as a prediction, and no uniqueness theorem is imported. The only self-citations ([29] and [46]) are illustrative examples and are not load-bearing premises. The paper also explicitly qualifies its own claim: Section 4.3 lists 'a related but less workflow-oriented direction' of representation models (UPS, Unisolver, FLUID-LLM, Text2PDE, FLUID-GPT), and Section 7 concedes that 'current demonstrations therefore show primarily that LLMs can make expert-designed procedures more executable and accessible' rather than establishing autonomous workflow-level capability. These internal qualifications show the taxonomy is not being used to force a conclusion. Concerns about sample representativeness and reliance on unreviewed preprints are evidence-quality concerns, not circularity under the hard rules. Hence score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities. The central claim rests on structural assumptions about the workflow decomposition and on the representativeness of the cited literature.

assumptions (3)
  • domain assumption PDE workflows naturally decompose into Discovery, Solving, and Optimization stages.
    Section 2 defines the taxonomy and the entire review is organized around it. If workflows do not decompose this way, the central claim about where LLMs contribute loses its organizing power.
  • domain assumption The papers listed in Table 1 are representative of current LLM-assisted PDE research.
    The abstract and conclusion generalize from these cited systems to 'current systems'. No inclusion criteria are given, so representativeness is assumed rather than shown.
  • standard math Standard PDE notation and problem specification (equation 1) apply to the surveyed systems.
    Equation (1) and the initial/boundary condition formalism are conventional and serve as a neutral frame for the taxonomy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large language models for partial differential equation workflows." pith.science (2026). https://pith.science/paper/W4HXHL7N

@misc{pith2026260803600,
  author       = {Pith},
  title        = {Pith review of: Large language models for partial differential equation workflows},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W4HXHL7N}},
  note         = {Machine review of arXiv:2608.03600}
}
read the original abstract

Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect modelling assumptions, governing equations, numerical solvers, diagnostics, and decisions. Large language models (LLMs) are beginning to support such workflows by linking natural language, symbolic mathematics, code, solver outputs, and feedback. Here we examine recent advances in LLM-assisted PDE research across three stages: the discovery and formulation of governing models, the generation and revision of executable numerical solvers, and the use of simulation feedback to support control, design, and optimization. Across these stages, current systems act primarily as workflow-level interfaces. Despite this progress, the field remains limited by the scarcity of high-quality datasets and benchmarks, especially for knowledge discovery and real-world applications, where expert annotation, executable problem construction, and task-level feedback require substantial domain effort. A further challenge is the persistent gap between simulation-based results and real-world scientific and engineering systems, which limits the direct transfer of numerical simulations, control policies, and optimized designs to practical settings. These challenges make LLM-assisted PDE workflows a critical testbed for developing scientific AI systems that can connect language, computation, physical constraints, and real-world decision-making.

Figures

Figures reproduced from arXiv: 2608.03600 by the authors.

Figure 1
Figure 1. Comparison of traditional, deep-learning, and LLM-assisted workflows for PDE scientific [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Representative LLM-assisted workflows in the Discovery Stage. (a) [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. LLMs for PDE solving-stage workflows. a, In traditional solver workflows, LLM agents translate PDE specifications into solver configurations, execute external tools, and revise failed cases using logs and diagnostics. b, In solver-code generation, LLMs generate, test, and debug numerical programmes from mathematical or natural￾language PDE descriptions. c, In deep-learning PDE solvers, LLMs assist architecture desig… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: LLMs for PDE optimization-stage workflows. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

91 extracted references · 63 canonical work pages

  1. [37]

    Llm4ed: Large language models for automatic equation discovery.arXiv preprint arXiv:2405.07761, 2024

    Mengge Du, Yuntian Chen, Zhongzheng Wang, Longfeng Nie, and Dongxiao Zhang. Llm4ed: Large language models for automatic equation discovery.arXiv preprint arXiv:2405.07761, 2024

  2. [40]

    Foam-agent: Towards automated intelligent cfd workflows

    Ling Yue, Nithin Somasekharan, Yadi Cao, and Shaowu Pan. Foam-agent: Towards automated intelligent cfd workflows. 2025

  3. [48]

    Agentic symbolic search: Characterizing pdes beyond hand-crafted expressions, meshes, and neural networks, 2026

    Zongmin Yu and Liu Yang. Agentic symbolic search: Characterizing pdes beyond hand-crafted expressions, meshes, and neural networks, 2026

  4. [50]

    Drsr: Llm based scientific equation discovery with dual reasoning from data and experience.arXiv preprint arXiv:2506.04282, 2025

    Runxiang Wang, Boxiao Wang, Kai Li, Yifan Zhang, and Jian Cheng. Drsr: Llm based scientific equation discovery with dual reasoning from data and experience.arXiv preprint arXiv:2506.04282, 2025

  5. [58]

    Automated code development for pde solvers using large language models

    Haoyang Wu, Xinxin Zhang, and Lailai Zhu. Automated code development for pde solvers using large language models. 2025

  6. [59]

    Autonumerics: An autonomous, pde-agnostic multi-agent pipeline for scientific computing, 2026

    Jianda Du, Youran Sun, and Haizhao Yang. Autonumerics: An autonomous, pde-agnostic multi-agent pipeline for scientific computing, 2026

  7. [63]

    Ai cfd scientist: Toward open-ended computational fluid dynamics discovery with physics-aware ai agents, 2026

    Nithin Somasekharan, Rabi Pathak, Manushri Dhanakoti, Tingwen Zhang, Ling Yue, Andy Zhu, and Shaowu Pan. Ai cfd scientist: Toward open-ended computational fluid dynamics discovery with physics-aware ai agents, 2026

  8. [67]

    Chatcfd: An llm-driven agent for end-to-end cfd automation with domain-specific structured reasoning

    E Fan, Kang Hu, Zhuowen Wu, Jiangyang Ge, Jiawei Miao, Yuzhi Zhang, He Sun, Weizong Wang, and Tianhan Zhang. Chatcfd: An llm-driven agent for end-to-end cfd automation with domain-specific structured reasoning. 2025

  9. [71]

    Pinnsagent: Automated pde surrogation with large language models.arXiv preprint arXiv:2501.12053, 2025

    Qingpo Wuwu, Chonghan Gao, Tianyu Chen, Yihang Huang, Yuekai Zhang, Jianing Wang, Jianxin Li, Haoyi Zhou, and Shanghang Zhang. Pinnsagent: Automated pde surrogation with large language models.arXiv preprint arXiv:2501.12053, 2025. 25

  10. [72]

    Lang-pinn: From language to physics-informed neural networks via a multi-agent framework

    Xin He, Liangliang You, Hongduan Tian, Bo Han, Ivor Tsang, and Yew-Soon Ong. Lang-pinn: From language to physics-informed neural networks via a multi-agent framework. 2025

  11. [82]

    Self-evolving scientific agent discovers generalizable physically-reasoned fluid control, 2026

    Boai Sun, Wenjin Guo, Zongmin Yu, and Liu Yang. Self-evolving scientific agent discovers generalizable physically-reasoned fluid control, 2026

  12. [83]

    Zhongxin Yang, Yuanwei Bin, Yipeng Shi, and Xiang I. A. Yang. Large language model driven development of turbulence models. 2025

  13. [86]

    Think like a scientist: Physics-guided llm agent for equation discovery, 2026

    Jianke Yang, Ohm Venkatachalam, Mohammad Kianezhad, Sharvaree Vadgama, and Rose Yu. Think like a scientist: Physics-guided llm agent for equation discovery, 2026

Show all 91 references
  1. [1]

    Evans.Partial Differential Equations, volume 19 ofGraduate Studies in Mathe- matics

    Lawrence C. Evans.Partial Differential Equations, volume 19 ofGraduate Studies in Mathe- matics. American Mathematical Society, Providence, RI, 2 edition, 2010

  2. [2]

    K. W. Morton and D. F. Mayers.Numerical Solution of Partial Differential Equations: An Introduction. Cambridge University Press, 2 edition, 2005

  3. [3]

    Oberkampf and Timothy G

    William L. Oberkampf and Timothy G. Trucano. Verification and validation in computational fluid dynamics.Progress in Aerospace Sciences, 38(3):209–272, 2002. 20

  4. [4]

    Prentice Hall, Upper Saddle River, NJ, 2 edition, 1999

    Lennart Ljung.System Identification: Theory for the User. Prentice Hall, Upper Saddle River, NJ, 2 edition, 1999

  5. [5]

    LeVeque.Finite Difference Methods for Ordinary and Partial Differential Equations: Steady-State and Time-Dependent Problems

    Randall J. LeVeque.Finite Difference Methods for Ordinary and Partial Differential Equations: Steady-State and Time-Dependent Problems. Society for Industrial and Applied Mathematics, Philadelphia, PA, 2007

  6. [6]

    H. K. Versteeg and W. Malalasekera.An Introduction to Computational Fluid Dynamics: The Finite Volume Method. Pearson Education Limited, Harlow, England, 2 edition, 2007

  7. [7]

    Brenner and L

    Susanne C. Brenner and L. Ridgway Scott.The Mathematical Theory of Finite Element Meth- ods, volume 15 ofTexts in Applied Mathematics. Springer, New York, 3 edition, 2008

  8. [8]

    Trefethen.Spectral Methods in MATLAB

    Lloyd N. Trefethen.Spectral Methods in MATLAB. Society for Industrial and Applied Math- ematics, Philadelphia, PA, 2000

  9. [9]

    Review of discontinuous galerkin finite element methods for partial differential equations on complicated domains

    PaolaF.Antonietti, AndreaCangiani, JenniferCollis, ZhaonanDong, EmmanuilH.Georgoulis, Stefano Giani, and Paul Houston. Review of discontinuous galerkin finite element methods for partial differential equations on complicated domains. InBuilding Bridges: Connections and Challen...

  10. [10]

    A review of mesh adaptation technology applied to computational fluid dynamics.Fluids, 10(5):129, 2025

    Guglielmo Vivarelli, Ning Qin, and Shahrokh Shahpar. A review of mesh adaptation technology applied to computational fluid dynamics.Fluids, 10(5):129, 2025

  11. [11]

    Springer, Berlin, Heidelberg, 1971

    Jacques Louis Lions.Optimal Control of Systems Governed by Partial Differential Equations, volume 170 ofGrundlehren der mathematischen Wissenschaften. Springer, Berlin, Heidelberg, 1971

  12. [12]

    Springer, Dor- drecht, 2009

    Michael Hinze, Rene Pinnau, Michael Ulbrich, and Stefan Ulbrich.Optimization with PDE Constraints, volume 23 ofMathematical Modelling: Theory and Applications. Springer, Dor- drecht, 2009

  13. [13]

    Gunzburger.Perspectives in Flow Control and Optimization

    Max D. Gunzburger.Perspectives in Flow Control and Optimization. Advances in Design and Control. Society for Industrial and Applied Mathematics, Philadelphia, PA, 2002

  14. [14]

    Bendsøe and Ole Sigmund.Topology Optimization: Theory, Methods, and Applica- tions

    Martin P. Bendsøe and Ole Sigmund.Topology Optimization: Theory, Methods, and Applica- tions. Springer, Berlin, Heidelberg, 2 edition, 2004

  15. [15]

    Brunton, Joshua L

    Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Discovering governing equations from data by sparse identification of nonlinear dynamical systems.Proceedings of the National Academy of Sciences, 113(15):3932–3937, 2016

  16. [16]

    Rudy, Steven L

    Samuel H. Rudy, Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Data-driven discovery of partial differential equations.Science Advances, 3(4):e1602614, 2017

  17. [17]

    Data-driven equation discovery of ocean mesoscale closures

    Laure Zanna and Thomas Bolton. Data-driven equation discovery of ocean mesoscale closures. Geophysical Research Letters, 47(17):e2020GL088376, 2020

  18. [18]

    Formulating turbulence closures using sparse regression with embedded form invariance.Physical Review Fluids, 5(8):084611, 2020

    Sarah Beetham and Jesse Capecelatro. Formulating turbulence closures using sparse regression with embedded form invariance.Physical Review Fluids, 5(8):084611, 2020. 21

  19. [19]

    Data-driven discovery of coarse-grained equations

    Joseph Bakarji and Daniel M Tartakovsky. Data-driven discovery of coarse-grained equations. Journal of Computational Physics, 434:110219, 2021

  20. [20]

    Maziar Raissi, Paris Perdikaris, and George E Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.Journal of Computational Physics, 378:686–707, 2019

  21. [21]

    Yohai Bar-Sinai, Stephan Hoyer, Jason Hickey, and Michael P. Brenner. Learning data-driven discretizations for partial differential equations.Proceedings of the National Academy of Sci- ences, 116(31):15344–15349, 2019

  22. [22]

    Smith, Ayya Alieva, Qing Wang, Michael P

    Dmitrii Kochkov, Jamie A. Smith, Ayya Alieva, Qing Wang, Michael P. Brenner, and Stephan Hoyer. Machine learning-accelerated computational fluid dynamics.Proceedings of the National Academy of Sciences, 118(21):e2101784118, 2021

  23. [23]

    Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators

    Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nat. Mach. Intell., 3(3):218–229, 2021

  24. [24]

    Fourier neural operator for parametric partial differential equations

    Zongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede liu, Kaushik Bhat- tacharya, Andrew Stuart, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations. InInternational Conference on Learning Representations, 2021

  25. [25]

    Factorized Fourier neural operators

    Alasdair Tran, Alexander Mathews, Lexing Xie, and Cheng Soon Ong. Factorized Fourier neural operators. InInternational Conference on Learning Representations, 2023

  26. [26]

    Physics-informed neural networks with hard constraints for inverse design.SIAM Journal on Scientific Computing, 43(6):B1105–B1132, 2021

    Lu Lu, Raphael Pestourie, Wenjie Yao, Zhicheng Wang, Francesc Verdugo, and Steven G Johnson. Physics-informed neural networks with hard constraints for inverse design.SIAM Journal on Scientific Computing, 43(6):B1105–B1132, 2021

  27. [27]

    Gradient-enhanced physics- informed neural networks for forward and inverse pde problems.Computer Methods in Applied Mechanics and Engineering, 393:114823, 2022

    Jeremy Yu, Lu Lu, Xuhui Meng, and George Em Karniadakis. Gradient-enhanced physics- informed neural networks for forward and inverse pde problems.Computer Methods in Applied Mechanics and Engineering, 393:114823, 2022

  28. [28]

    Encoding physics to learn reaction–diffusion processes.Nature Machine Intelligence, 5(7):765–779, 2023

    Chengping Rao, Pu Ren, Qi Wang, Oral Buyukozturk, Hao Sun, and Yang Liu. Encoding physics to learn reaction–diffusion processes.Nature Machine Intelligence, 5(7):765–779, 2023

  29. [29]

    Pesanet: Physics-encoded spec- tral attention network for simulating pde-governed complex systems

    Han Wan, Rui Zhang, Qi Wang, Yang Liu, and Hao Sun. Pesanet: Physics-encoded spec- tral attention network for simulating pde-governed complex systems. InInternational Joint Conference on Artificial Intelligence, pages 1–7, 2025

  30. [30]

    Artificial neural networks trained through deep reinforcement learning discover control strategies for active flow control.Journal of Fluid Mechanics, 865:281–302, 2019

    Jean Rabault, Miroslav Kuchta, Atle Jensen, Ulysse Réglade, and Nicolas Cerardi. Artificial neural networks trained through deep reinforcement learning discover control strategies for active flow control.Journal of Fluid Mechanics, 865:281–302, 2019

  31. [31]

    Learning to control pdes with differentiable physics, 2020

    Philipp Holl, Nils Thuerey, and Vladlen Koltun. Learning to control pdes with differentiable physics, 2020

  32. [32]

    Direct shape optimization through deep reinforcement learning.Journal of Computational Physics, 428:110080, 2021

    Jonathan Viquerat, Jean Rabault, Alexander Kuhnle, Hassan Ghraieb, Aurélien Larcher, and Elie Hachem. Direct shape optimization through deep reinforcement learning.Journal of Computational Physics, 428:110080, 2021. 22

  33. [33]

    Stachenfeld, Alvaro Sanchez-Gonzalez, Pe- ter Battaglia, Jessica B

    Kelsey Allen, Tatiana Lopez-Guevara, Kimberly L. Stachenfeld, Alvaro Sanchez-Gonzalez, Pe- ter Battaglia, Jessica B. Hamrick, and Tobias Pfaff. Inverse design for fluid-structure interac- tions using graph network simulators. InAdvances in Neural Information Processing Systems...

  34. [34]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...

  35. [35]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. InProceedings of the 36th International Conference on Neural Information Processing Syst...

  36. [36]

    Mielke, Yonatan Belinkov, Barak Lenz, Omer Lieber, et al

    Ekin Karpas, Yoav Levine, Sabrina J. Mielke, Yonatan Belinkov, Barak Lenz, Omer Lieber, et al. Mrkl systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.arXiv preprint arXiv:2205.00445, 2022

  37. [38]

    Physpde: Rethinking pde discovery and a physical hypothesis selection benchmark

    MingquanFeng, YixinHuang, YizhouLiu, BofangJiang, andJunchiYan. Physpde: Rethinking pde discovery and a physical hypothesis selection benchmark. InThe Thirteenth International Conference on Learning Representations, 2025

  38. [39]

    Codepde: An inference framework for llm-driven pde solver generation

    Shanda Li, Tanya Marwah, Junhong Shen, Weiwei Sun, Andrej vectors Risteski, Yiming Yang, and Ameet Talwalkar. Codepde: An inference framework for llm-driven pde solver generation. Transactions on Machine Learning Research, 2026. Accepted by TMLR

  39. [41]

    Pde-sharp: Pde solver hybrids through analysis and refinement passes.arXiv preprint arXiv:2511.00183, 2025

    Shaghayegh Fazliani and Madeleine Udell. Pde-sharp: Pde solver hybrids through analysis and refinement passes.arXiv preprint arXiv:2511.00183, 2025

  40. [42]

    Pde-controller: Llms for autoformalization and reasoning of pdes

    Mauricio Soroco, Jialin Song, Mengzhou Xia, Kye Emond, Weiran Sun, and Wuyang Chen. Pde-controller: Llms for autoformalization and reasoning of pdes. InInternational Conference on Machine Learning. PMLR, 2025

  41. [43]

    Using large language models for parametric shape op- timization.Physics of Fluids, 37(8):083601, 2025

    Xinxin Zhang, Zhuoqun Xu, Guangpu Zhu, Chien Ming Jonathan Tay, Yongdong Cui, Boo Cheong Khoo, and Lailai Zhu. Using large language models for parametric shape op- timization.Physics of Fluids, 37(8):083601, 2025. 23

  42. [44]

    Accelerating scientific discovery with co-scientist.Nature, 655:487–496, 2026

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, et al. Accelerating scientific discovery with co-scientist.Nature, 655:487–496, 2026

  43. [45]

    Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J

    Ali E. Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J. Szostkiewicz, Dmytro Shved, Gavin J. Gyimesi, Jon M. Laurent, Samantha M. Wright, Muhammed T. Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. A multi-agent syst...

  44. [46]

    Evaluating llms’ divergent thinking capabilities for scientific idea generation with minimal context.Nature Communications, 17(1):3625, 2026

    Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang, Yang Liu, and Hao Sun. Evaluating llms’ divergent thinking capabilities for scientific idea generation with minimal context.Nature Communications, 17(1):3625, 2026

  45. [47]

    Llm assisted mathematical modeling: Homogeneous laplace equation in cylinder with the complete electrode model

    Agah Drajat Garnadi. Llm assisted mathematical modeling: Homogeneous laplace equation in cylinder with the complete electrode model. 2025

  46. [49]

    The impact of large language models on scientific discovery: a preliminary study using gpt-4

    Microsoft Research AI4Science and Microsoft Azure Quantum. The impact of large language models on scientific discovery: a preliminary study using gpt-4. 2023

  47. [51]

    Llm-sr: Scientific equation discovery via programming with large language models

    Parshin Shojaee, Kazem Meidani, Shashank Gupta, Amir Barati Farimani, and Chandan K Reddy. Llm-sr: Scientific equation discovery via programming with large language models. arXiv preprint arXiv:2404.18400, 2024

  48. [52]

    From equations to insights: Unraveling symbolic structures in pdes with llms.arXiv preprint arXiv:2503.09986, 2025

    Rohan Bhatnagar, Ling Liang, Krish Patel, and Haizhao Yang. From equations to insights: Unraveling symbolic structures in pdes with llms.arXiv preprint arXiv:2503.09986, 2025

  49. [53]

    Llm and simulation as bilevel optimizers: A new paradigm to advance physical scientific discovery.arXiv preprint arXiv:2405.09783, 2024

    Pingchuan Ma, Tsun-Hsuan Wang, Minghao Guo, Zhiqing Sun, Joshua B Tenenbaum, Daniela Rus, Chuang Gan, and Wojciech Matusik. Llm and simulation as bilevel optimizers: A new paradigm to advance physical scientific discovery.arXiv preprint arXiv:2405.09783, 2024

  50. [54]

    Jiachen Guo, Chanwook Park, Dong Qian, Thomas J. R. Hughes, and Wing Kam Liu. Large language model-empowered next-generation computer-aided engineering. 2025

  51. [55]

    Pdeagent-bench: A multi-metric, multi-library benchmark for pde solver generation, 2026

    Zhen Hang, Yushan Yashengjiang, Junhui Li, Huanshuo Dong, Yang Wei, Zhezheng Hao, Jiangtao Ma, Songlin Bai, Haozhong Kai, Xihang Yue, Gangzong Si, Dongming Jiang, Chao Yao, Zhanhua Hu, Jiangqing Zhang, Pengwei Liu, Yaomin Shen, Xingyu Ren, Lei Liu, Zikang Xu, Han Li, Qingsong ...

  52. [56]

    Deepseek vs

    Qile Jiang, Zhiwei Gao, and George Em Karniadakis. Deepseek vs. chatgpt vs. claude: A comparative study for scientific computing and scientific machine learning tasks.Theoretical and Applied Mechanics Letters, 15(3):100583, 2025

  53. [57]

    All-fem: Agentic large language models fine-tuned for 24 finite element methods.Computer Methods in Applied Mechanics and Engineering, 457:118985, 2026

    Rushikesh Deotale, Adithya Srinivasan, Mahmoud Golestanian, Yuan Tian, Tianyi Zhang, Pavlos Vlachos, and Hector Gomez. All-fem: Agentic large language models fine-tuned for 24 finite element methods.Computer Methods in Applied Mechanics and Engineering, 457:118985, 2026

  54. [60]

    Evaluations of large language models in computa- tional fluid dynamics: Leveraging, learning and creating knowledge.Theoretical and Applied Mechanics Letters, 15(3):100597, 2025

    Long Wang, Lei Zhang, and Guowei He. Evaluations of large language models in computa- tional fluid dynamics: Leveraging, learning and creating knowledge.Theoretical and Applied Mechanics Letters, 15(3):100597, 2025

  55. [61]

    Cfdllmbench: A benchmark suite for evaluating large language models in computational fluid dynamics

    Nithin Somasekharan, Ling Yue, Yadi Cao, Weichao Li, Patrick Emami, Pochinapeddi Sai Bhargav, Anurag Acharya, Xingyu Xie, and Shaowu Pan. Cfdllmbench: A benchmark suite for evaluating large language models in computational fluid dynamics. 2025

  56. [62]

    Openfoamgpt: A retrieval-augmented large language model (llm) agent for openfoam-based computational fluid dynamics.Physics of Fluids, 37(3), 2025

    Sandeep Pandey, Ran Xu, Wenkang Wang, and Xu Chu. Openfoamgpt: A retrieval-augmented large language model (llm) agent for openfoam-based computational fluid dynamics.Physics of Fluids, 37(3), 2025

  57. [64]

    Openfoamgpt 2.0: End-to-end, trustworthy automation for computational fluid dynamics.International Journal of Heat and Fluid Flow, 120:110399, 2026

    Jingsen Feng, Ran Xu, and Xu Chu. Openfoamgpt 2.0: End-to-end, trustworthy automation for computational fluid dynamics.International Journal of Heat and Fluid Flow, 120:110399, 2026

  58. [65]

    Metaopenfoam: an llm-based multi-agent framework for cfd

    Yuxuan Chen, Xu Zhu, Hua Zhou, and Zhuyin Ren. Metaopenfoam: an llm-based multi-agent framework for cfd. 2024

  59. [66]

    Metaopenfoam 2.0: Large language model driven chain of thought for automating cfd simulation and post-processing.Journal Name, 2025

    Yuxuan Chen, Xu Zhu, Hua Zhou, and Zhuyin Ren. Metaopenfoam 2.0: Large language model driven chain of thought for automating cfd simulation and post-processing.Journal Name, 2025

  60. [68]

    Fine-tuning a large language model for automating com- putational fluid dynamics simulations.Theoretical and Applied Mechanics Letters, 15:100594, 2025

    Zhehao Dong, Zhen Lu, and Yue Yang. Fine-tuning a large language model for automating com- putational fluid dynamics simulations.Theoretical and Applied Mechanics Letters, 15:100594, 2025

  61. [69]

    Physics simulation capabilities of llms.Physica Scripta, 99(11):116003, oct 2024

    Mohamad Ali-Dib and Kristen Menou. Physics simulation capabilities of llms.Physica Scripta, 99(11):116003, oct 2024

  62. [70]

    Mycrunchgpt: Achatgptassistedframeworkforscientificmachinelearning.Journal of Machine Learning for Modeling and Computing, 4(4):41–72, January 2023

    Varun Kumar, Leonard Gleyzer, Adar Kahana, Khemraj Shukla, and George Karniadakis. Mycrunchgpt: Achatgptassistedframeworkforscientificmachinelearning.Journal of Machine Learning for Modeling and Computing, 4(4):41–72, January 2023

  63. [73]

    Text-trained llms can zero-shot extrapolate pde dynamics.arXiv preprint arXiv:2509.06322, 2025

    JiajunBao, NicolasBoullé, ToniJBLiu, RaphaëlSarfati, andChristopherJEarls. Text-trained llms can zero-shot extrapolate pde dynamics.arXiv preprint arXiv:2509.06322, 2025

  64. [74]

    Unisolver: Pde- conditional transformers towards universal neural pde solvers

    Hang Zhou, Yuezhou Ma, Haixu Wu, Haowen Wang, and Mingsheng Long. Unisolver: Pde- conditional transformers towards universal neural pde solvers. InForty-second International Conference on Machine Learning, 2025

  65. [75]

    UPS: Efficiently building foundation models for PDE solving via cross-modal adaptation.Transactions on Machine Learning Re- search, 2024

    Junhong Shen, Tanya Marwah, and Ameet Talwalkar. UPS: Efficiently building foundation models for PDE solving via cross-modal adaptation.Transactions on Machine Learning Re- search, 2024

  66. [76]

    Fluid-llm: Learning computational fluid dynamics with spatiotemporal-aware large language models

    Max Zhu, Adrián Bazaga, and Pietro Liò. Fluid-llm: Learning computational fluid dynamics with spatiotemporal-aware large language models. 2024

  67. [77]

    Buchanan, and Amir Barati Farimani

    Anthony Zhou, Zijie Li, Michael Schneier, John R. Buchanan, and Amir Barati Farimani. Text2PDE: Latent diffusion models for accessible physics simulation. InThe Thirteenth Inter- national Conference on Learning Representations (ICLR), 2025

  68. [78]

    Yang, Zulfikhar A

    Steve D. Yang, Zulfikhar A. Ali, and Bryan M. Wong. Fluid-gpt (fast learning to understand and investigate dynamics with a generative pre-trained transformer): Efficient predictions of particle trajectories and erosion.Industrial & Engineering Chemistry Research, 62(37):15278–...

  69. [79]

    Aeroagent: A vision- physics-decision framework for aerodynamic vehicle design

    Ye Liu, Shouyi Liu, Huiyu Yang, Jianghang Gu, Wenhao Fan, Zhongxin Yang, Ding Wang, Simeng Chen, Zirun Jiang, Yuanwei Bin, Shiyi Chen, and Yuntian Chen. Aeroagent: A vision- physics-decision framework for aerodynamic vehicle design. InProceedings of the IEEE/CVF Conference on ...

  70. [80]

    Shapebench: A scalable benchmark and diagnostic suite for standardized evaluation in aerodynamic shape optimization, 2026

    Shaghayegh Fazliani, Krissh Chawla, Jack Guo, Yiren Shen, Matthias Ihme, and Madeleine Udell. Shapebench: A scalable benchmark and diagnostic suite for standardized evaluation in aerodynamic shape optimization, 2026

  71. [81]

    Optmetaopenfoam: Large language model driven chain of thought for sensitivity analysis and parameter optimization based on cfd.Journal Name, 2025

    Yuxuan Chen, Long Zhang, Xu Zhu, Hua Zhou, and Zhuyin Ren. Optmetaopenfoam: Large language model driven chain of thought for sensitivity analysis and parameter optimization based on cfd.Journal Name, 2025

  72. [84]

    Toward knowledge-guided ai for inverse design in manufacturing: A perspective on domain, physics, and human–ai synergy

    Hugon Lee, Hyeonbin Moon, Junhyeong Lee, and Seunghwa Ryu. Toward knowledge-guided ai for inverse design in manufacturing: A perspective on domain, physics, and human–ai synergy. Advanced Intelligent Discovery, 2025

  73. [85]

    Toward autonomous engineering design: A knowledge-guided multi-agent framework, 2025

    Varun Kumar and George Em Karniadakis. Toward autonomous engineering design: A knowledge-guided multi-agent framework, 2025. 26

  74. [87]

    Callaghan, and Dongxiao Zhang

    Hao Xu, Yuntian Chen, Rui Cao, Tianning Tang, Mengge Du, Jian Li, Adrian H. Callaghan, and Dongxiao Zhang. Generative discovery of partial differential equations by learning from math handbooks, 2025

  75. [88]

    Osher, and Hayden Schaeffer

    Elisa Negrini, Yuxuan Liu, Liu Yang, Stanley J. Osher, and Hayden Schaeffer. A multimodal pde foundation model for prediction and scientific text descriptions. 2025

  76. [89]

    Jasak, A

    H. Jasak, A. Jemcov, Z. Tukovic, et al. Openfoam: A C++ library for complex physics simulations. InInternational Workshop on Coupled Methods in Numerical Dynamics, volume 1000, pages 1–20, Dubrovnik, Croatia, 2007

  77. [90]

    PDEBench: AnExtensiveBenchmarkforScientificMachine Learning

    Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Dan MacKinlay, Francesco Alesiani, DirkPflüger, andMathiasNiepert. PDEBench: AnExtensiveBenchmarkforScientificMachine Learning. In36th Conference on Neural Information Processing Systems (NeurIPS 2022) Track on Datasets and...

  78. [91]

    Pinnacle: a comprehensive benchmark of physics-informed neural networks for solving pdes

    Zhongkai Hao, Jiachen Yao, Chang Su, Hang Su, Ziao Wang, Fanzhi Lu, Zeyu Xia, Yichi Zhang, Songming Liu, Lu Lu, and Jun Zhu. Pinnacle: a comprehensive benchmark of physics-informed neural networks for solving pdes. InProceedings of the 38th International Conference on Neural I...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.