Pith. sign in

REVIEW 1 cited by

Programmatic Reinforcement Learning: Navigating Gridworlds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11650 v2 pith:YXA3Q5VN submitted 2024-02-18 cs.LG cs.LOcs.PL

classification cs.LGcs.LOcs.PL
keywords programmaticpolicieslearningoptimaltheoreticalalgorithmclassenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of reinforcement learning (RL) is concerned with algorithms for learning optimal policies in unknown stochastic environments. Programmatic RL studies representations of policies as programs, meaning involving higher order constructs such as control loops. Despite attracting a lot of attention at the intersection of the machine learning and formal methods communities, very little is known on the theoretical front about programmatic RL: what are good classes of programmatic policies? How large are optimal programmatic policies? How can we learn them? The goal of this paper is to give first answers to these questions, initiating a theoretical study of programmatic RL. Considering a class of gridworld environments, we define a class of programmatic policies. Our main contributions are to place upper bounds on the size of optimal programmatic policies, and to construct an algorithm for synthesizing them. These theoretical findings are complemented by a prototype implementation of the algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simplicity Lies in the Eye of the Beholder: A Strategic Perspective on Controllers in Reactive Synthesis

    cs.LO 2025-09 accept novelty 3.0 of 10

    A survey of memory and randomness complexity for strategies in reactive synthesis, arguing that Mealy-machine-based measures of simplicity are representation-dependent.

Pith tools