Pith. sign in

REVIEW 9 cited by

Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.04871 v4 pith:U6MSTIQI submitted 2019-12-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords symbolicexpressionsregressiondeepmathematicalperformancepolicyrisk-seeking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Discovering the underlying mathematical expressions describing a dataset is a core challenge for artificial intelligence. This is the problem of $\textit{symbolic regression}$. Despite recent advances in training neural networks to solve complex tasks, deep learning approaches to symbolic regression are underexplored. We propose a framework that leverages deep learning for symbolic regression via a simple idea: use a large model to search the space of small models. Specifically, we use a recurrent neural network to emit a distribution over tractable mathematical expressions and employ a novel risk-seeking policy gradient to train the network to generate better-fitting expressions. Our algorithm outperforms several baseline methods (including Eureqa, the gold standard for symbolic regression) in its ability to exactly recover symbolic expressions on a series of benchmark problems, both with and without added noise. More broadly, our contributions include a framework that can be applied to optimize hierarchical, variable-length objects under a black-box performance metric, with the ability to incorporate constraints in situ, and a risk-seeking policy gradient formulation that optimizes for best-case performance instead of expected performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neuro-Symbolic ODE Discovery with Latent Grammar Flow

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Latent Grammar Flow embeds grammar-based ODE representations into a discrete latent space with a behavioural loss and samples candidate equations via discrete flow to fit observed data.

  2. LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization

    cs.LG 2026-02 conditional novelty 7.0 of 10

    PiT-PO adaptively fine-tunes an LLM during symbolic regression search using physics-validity and token-level redundancy constraints, reporting state-of-the-art benchmark results and a periodic-hill turbulence closure.

  3. Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective

    cs.LG 2026-07 accept novelty 6.5 of 10

    Rank-conditioned Horvitz–Thompson reuses all C(n,K) subsets of one Gumbel-Top-n pool for unbiased Plackett–Luce best-of-K value and score-function gradient, with an exact Max-specific DP collapse to a 1-D integral.

  4. MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    MOT-SR combines tool-augmented data analysis with multi-objective Pareto selection to discover symbolic equations, outperforming LLM-based and classical SR baselines on benchmarks and an EMRI orbital-correction task.

  5. LLM-PDESR: Robust PDE Discovery via Subdomain Weighted Residuals and LLM-Guided Symbolic Hypothesis Generation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    LLM-PDESR pairs LLM structural hypotheses with C4 quintic splines and subdomain weighted residuals to recover PDEs from noisy data more robustly than prior symbolic methods.

  6. Symbolic Regression for Shared Expressions: Introducing Partial Parameter Sharing

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Introduces partially-shared parameters for symbolic regression with multiple categorical variables, matching prior fit quality on a supernovae dataset with fewer parameters.

  7. SymMatika: Structure-Aware Symbolic Discovery

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A structure-aware symbolic regression framework combining multi-island genetic programming with reusable motif libraries reports state-of-the-art recovery rates on Nguyen and Feynman benchmarks, including 61% on Nguyen-12.

  8. Sparse Interpretable Deep Learning with LIES Networks for Symbolic Regression

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LIES, a fixed network with logarithm, identity, exponential, and sine activations, recovers exact symbolic formulas from AI Feynman benchmark data more often than comparison methods on a selected 61-equation subset.

  9. Drag modelling for flows through assemblies of spherical particles with machine learning: A comparison of approaches

    physics.comp-ph 2025-07 conditional novelty 4.0 of 10

    Applying genetic programming to outputs of a graph neural network yields compact symbolic drag-variation formulas at Reynolds numbers up to 280, though with lower accuracy than the network.

Pith tools