Pith. sign in

REVIEW 3 cited by

LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06512 v2 pith:HLEN25SL submitted 2025-03-09 cond-mat.mtrl-sci

classification cond-mat.mtrl-sci
keywords domainknowledgellm-feynmandatadiscoveryformulaformulasmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Distilling underlying principles from data has historically driven scientific breakthroughs. However, conventional data-driven machine learning often produces complex models that lack interpretability and generalization due to insufficient domain expertise. Here, we present LLM-Feynman, a novel framework that leverages large language models (LLMs) alongside systematic optimization to derive concise, interpretable formulas from data and domain knowledge. Our method integrates automated feature engineering, LLM-guided symbolic regression with self-evaluation, and Monte Carlo tree search to enhance formula discovery and clarity. The embedding of domain knowledge simplifies the formula, while self-evaluation based on this knowledge further minimizes prediction errors, surpassing conventional symbolic regression in accuracy and interpretability. Our LLM-Feynman successfully rediscovered over 90% of fundamental physical formulas and demonstrated its efficacy in key materials science applications, including classification of two-dimensional material and perovskite synthesizability and determination of the Green's function and screened Coulomb interaction bandgaps, and prediction of ionic conductivity in lithium solid-state electrolytes. By transcending mere data fitting through the integration of deep domain knowledge, this LLM-Feynman offers a transformative paradigm for the automated discovery of generalizable scientific formulas and theories across disciplines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

    cs.LG 2025-05 reject novelty 6.0 of 10

    A new benchmark of 380 principle-based physics problems shows that state-of-the-art LLMs struggle to apply symmetry, conservation, and dimensional-analysis shortcuts, achieving under 50 percent average accuracy with h...

  2. Data-driven Discovery of Digital Twins in Biomedical Research

    q-bio.QM 2025-08 conditional novelty 2.0 of 10

    Sparse regression, especially Bayesian approaches, generally outperform symbolic regression for ODE-based digital twin discovery in biology, though the evidence is qualitative and non-systematic.

  3. Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research

    cs.RO 2025-06 accept novelty 1.0 of 10

    A perspective article reviews the state of using foundation models for laboratory automation and proposes a roadmap for fully autonomous experiments.

Pith tools