Pith. sign in

REVIEW 1 cited by

Testing Stationarity and Change Point Detection in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.01707 v4 pith:PLTL7NHW submitted 2022-03-03 stat.ML cs.LG

classification stat.MLcs.LG
keywords datastationarityassumptionchangedetectiondevelopenvironmentsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider offline reinforcement learning (RL) methods in possibly nonstationary environments. Many existing RL algorithms in the literature rely on the stationarity assumption that requires the system transition and the reward function to be constant over time. However, the stationarity assumption is restrictive in practice and is likely to be violated in a number of applications, including traffic signal control, robotics and mobile health. In this paper, we develop a consistent procedure to test the nonstationarity of the optimal Q-function based on pre-collected historical data, without additional online data collection. Based on the proposed test, we further develop a sequential change point detection method that can be naturally coupled with existing state-of-the-art RL methods for policy optimization in nonstationary environments. The usefulness of our method is illustrated by theoretical results, simulation studies, and a real data example from the 2018 Intern Health Study. A Python implementation of the proposed procedure is available at https://github.com/limengbinggz/CUSUM-RL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Off-Policy Evaluation and Learning for the Future under Non-Stationarity

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A new importance-weighted estimator, OPFV, estimates and optimizes future policy value in non-stationary bandit environments by leveraging recurring time features in historical logs.

Pith tools