Pith. sign in

REVIEW 1 cited by

Model-based Policy Optimization using Symbolic World Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13518 v1 pith:R65LBD6T submitted 2024-07-18 cs.LG

classification cs.LG
keywords symbolicmodelmodel-baseddynamicsefficiencylearningsampleapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The application of learning-based control methods in robotics presents significant challenges. One is that model-free reinforcement learning algorithms use observation data with low sample efficiency. To address this challenge, a prevalent approach is model-based reinforcement learning, which involves employing an environment dynamics model. We suggest approximating transition dynamics with symbolic expressions, which are generated via symbolic regression. Approximation of a mechanical system with a symbolic model has fewer parameters than approximation with neural networks, which can potentially lead to higher accuracy and quality of extrapolation. We use a symbolic dynamics model to generate trajectories in model-based policy optimization to improve the sample efficiency of the learning algorithm. We evaluate our approach across various tasks within simulated environments. Our method demonstrates superior sample efficiency in these tasks compared to model-free and model-based baseline methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. M3PO: Massively Multi-Task Model-Based Policy Optimization

    cs.LG 2025-06 reject novelty 4.0 of 10

    A hybrid model-based RL algorithm combining TDMPC2-style world models, PPO clipped updates, and POME-style exploration bonuses reports strong benchmark results but omits reproducible evidence.

Pith tools