Pith. sign in

REVIEW 4 cited by

Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03578 v3 pith:VN2IPSBT submitted 2024-04-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords robusttrainingalgorithmcollectiondataenvironmentenvironmentslearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distributionally robust RL, often framed as a robust Markov decision process (RMDP). In this framework, the objective is to find a robust policy that achieves good performance under the worst-case scenario among all environments within a pre-specified uncertainty set centered around the training environment. Unlike previous work, which relies on a generative model or a pre-collected offline dataset enjoying good coverage of the deployment environment, we tackle robust RL via interactive data collection, where the learner interacts with the training environment only and refines the policy through trial and error. In this robust RL paradigm, two main challenges emerge: managing distributional robustness while striking a balance between exploration and exploitation during data collection. Initially, we establish that sample-efficient learning without additional assumptions is unattainable owing to the curse of support shift; i.e., the potential disjointedness of the distributional supports between the training and testing environments. To circumvent such a hardness result, we introduce the vanishing minimal value assumption to RMDPs with a total-variation (TV) distance robust set, postulating that the minimal value of the optimal robust value function is zero. We prove that such an assumption effectively eliminates support shift pathologies for RMDPs with a TV distance robust set, and present an algorithm with near-optimal sample complexity. To demonstrate the breadth of our framework, we extend our algorithm and theory to new robust set formulations and robust Markov games. To illustrate the operational relevance, we apply our algorithm to data-driven robust inventory control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

    cs.LG 2026-08 accept novelty 8.0 of 10

    For average-reward MDPs with total-variation uncertainty, the minimax sample complexity is SA/epsilon^2 times min{H0,Hsigma}, with an extra SA sigma Hsigma^2/epsilon^2 term in the low-tolerance regime, and the paper p...

  2. Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

    cs.LG 2025-05 reject novelty 7.0 of 10

    RHI is claimed to find an epsilon-optimal robust policy under the average-reward criterion with about SAH^2/epsilon^2 samples under the communicating assumption, with a parameter-free variant that avoids knowing H.

  3. Causality-Inspired Robustness for Nonlinear Models via Representation Learning

    stat.ML 2025-05 reject novelty 6.0 of 10

    The proposed two-step method, CIRRL, learns an affine-equivalent latent representation of causal structure and applies a distributionally robust linear regression on it, claiming minimax robustness against bounded dis...

  4. Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A pessimism-based framework for zero-shot transfer RL builds conservative proxies from robust MDPs, yielding lower-bound performance guarantees and distributed algorithms that mitigate negative transfer.

Pith tools