Pith. sign in

REVIEW 1 cited by

Online algorithms for POMDPs with continuous state, action, and observation spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1709.06196 v6 pith:4BCRSOBO submitted 2017-09-18 cs.AI cs.ROcs.SYeess.SY

classification cs.AIcs.ROcs.SYeess.SY
keywords algorithmsspacesstateactionchallengecontinuousobservationonline
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Online solvers for partially observable Markov decision processes have been applied to problems with large discrete state spaces, but continuous state, action, and observation spaces remain a challenge. This paper begins by investigating double progressive widening (DPW) as a solution to this challenge. However, we prove that this modification alone is not sufficient because the belief representations in the search tree collapse to a single particle causing the algorithm to converge to a policy that is suboptimal regardless of the computation time. This paper proposes and evaluates two new algorithms, POMCPOW and PFT-DPW, that overcome this deficiency by using weighted particle filtering. Simulation results show that these modifications allow the algorithms to be successful where previous approaches fail.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Guided Probabilistic Program Induction for POMDP Model Estimation

    cs.AI 2025-05 conditional novelty 6.0 of 10

    LLM-guided probabilistic program induction can learn low-complexity POMDP models from ten demonstrations and outperform tabular learning, behavior cloning, and direct LLM planning in simulated and real robot domains.

Pith tools