Pith. sign in

REVIEW

Learning in POMDPs is Sample-Efficient with Hindsight Observability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.13857 v2 pith:4MXN7UML submitted 2023-01-31 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords hindsightlearningobservabilitypomdpsdecisionduringevenintractable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

POMDPs capture a broad class of decision making problems, but hardness results suggest that learning is intractable even in simple settings due to the inherent partial observability. However, in many realistic problems, more information is either revealed or can be computed during some point of the learning process. Motivated by diverse applications ranging from robotics to data center scheduling, we formulate a Hindsight Observable Markov Decision Process (HOMDP) as a POMDP where the latent states are revealed to the learner in hindsight and only during training. We introduce new algorithms for the tabular and function approximation settings that are provably sample-efficient with hindsight observability, even in POMDPs that would otherwise be statistically intractable. We give a lower bound showing that the tabular algorithm is optimal in its dependence on latent state and observation cardinalities.

Discussion (0). Sign in to comment.

Pith tools