Pith. sign in

REVIEW 7 cited by

Poseidon: Efficient Foundation Models for PDEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19101 v2 pith:HD7TYSJB submitted 2024-05-29 cs.LG

classification cs.LG
keywords poseidonpdesdownstreammodelpretrainingfoundationwellcamlab-ethz
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Poseidon, a foundation model for learning the solution operators of PDEs. It is based on a multiscale operator transformer, with time-conditioned layer norms that enable continuous-in-time evaluations. A novel training strategy leveraging the semi-group property of time-dependent PDEs to allow for significant scaling-up of the training data is also proposed. Poseidon is pretrained on a diverse, large scale dataset for the governing equations of fluid dynamics. It is then evaluated on a suite of 15 challenging downstream tasks that include a wide variety of PDE types and operators. We show that Poseidon exhibits excellent performance across the board by outperforming baselines significantly, both in terms of sample efficiency and accuracy. Poseidon also generalizes very well to new physics that is not seen during pretraining. Moreover, Poseidon scales with respect to model and data size, both for pretraining and for downstream tasks. Taken together, our results showcase the surprising ability of Poseidon to learn effective representations from a very small set of PDEs during pretraining in order to generalize well to unseen and unrelated PDEs downstream, demonstrating its potential as an effective, general purpose PDE foundation model. Finally, the Poseidon model as well as underlying pretraining and downstream datasets are open sourced, with code being available at https://github.com/camlab-ethz/poseidon and pretrained models and datasets at https://huggingface.co/camlab-ethz.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. Probabilistic operator learning: generative modeling and uncertainty quantification for foundation models of differential equations

    stat.ML 2025-09 conditional novelty 6.0 of 10

    ICON is shown to compute the posterior predictive mean of differential equation solutions, and a generative extension, GenICON, provides samples from this distribution for uncertainty quantification.

  2. Physics-informed, boundary-constrained Gaussian process regression for the reconstruction of fluid flow fields

    physics.flu-dyn 2025-07 conditional novelty 6.0 of 10

    A spectral boundary-constraining method for Gaussian processes is extended to incompressible flow reconstruction, giving divergence-free, slip-condition-satisfying priors that need no profile-boundary observations.

  3. Autoregressive regularized score-based diffusion models for multi-scenarios fluid flow prediction

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A regularized autoregressive score-based diffusion model predicts turbulent flows across multiple scenarios, with the variance-preserving SDE formulation performing best.

  4. A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A multimodal transformer predicts ODE/PDE solutions and generates correct scientific text descriptions from numerical and symbolic inputs, with low error on in-distribution and out-of-distribution tests.

  5. Towards Foundational Models for Dynamical System Reconstruction: Hierarchical Meta-Learning via Mixture of Experts

    cs.LG 2025-02 conditional novelty 6.0 of 10

    MixER uses K-means clustering and least squares to route each dynamical system to a specialist expert, improving reconstruction across loosely related ODE families in low-data regimes while underperforming on closely ...

  6. Mondrian: Transformer Operators via Domain Decomposition

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Mondrian applies transformer attention to subdomain-restricted functions, decoupling the model from the grid resolution.

  7. Latent Mamba Operator for Partial Differential Equations

    cs.LG 2025-05 conditional novelty 5.0 of 10

    LaMO replaces attention in latent-token neural operators with bidirectional state-space models and reports consistent accuracy gains on six PDE benchmarks.

Pith tools