Pith. sign in

REVIEW 3 cited by

Sample-Efficient Reinforcement Learning via Counterfactual-Based Data Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.09092 v1 pith:MWMXDSJR submitted 2020-12-16 cs.LG stat.ML

classification cs.LGstat.ML
keywords dataalgorithmslearningpoliciescounterfactualcounterfactual-basedlearnonly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) algorithms usually require a substantial amount of interaction data and perform well only for specific tasks in a fixed environment. In some scenarios such as healthcare, however, usually only few records are available for each patient, and patients may show different responses to the same treatment, impeding the application of current RL algorithms to learn optimal policies. To address the issues of mechanism heterogeneity and related data scarcity, we propose a data-efficient RL algorithm that exploits structural causal models (SCMs) to model the state dynamics, which are estimated by leveraging both commonalities and differences across subjects. The learned SCM enables us to counterfactually reason what would have happened had another treatment been taken. It helps avoid real (possibly risky) exploration and mitigates the issue that limited experiences lead to biased policies. We propose counterfactual RL algorithms to learn both population-level and individual-level policies. We show that counterfactual outcomes are identifiable under mild conditions and that Q- learning on the counterfactual-based augmented data set converges to the optimal value function. Experimental results on synthetic and real-world data demonstrate the efficacy of the proposed approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A New Perspective On AI Safety Through Control Theory Methodologies

    cs.AI 2025-06 conditional novelty 6.0 of 10

    This paper outlines a new conceptual paradigm, data control, which transfers control-theoretic system analysis and properties to AI systems to support generic AI safety assurance.

  2. RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations

    cs.RO 2024-11 conditional novelty 6.0 of 10

    A combined causal, SE(3)-equivariant, and visual data augmentation method improves behavior-cloning policy performance and generalization across five simulated manipulation tasks.

  3. MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A model-based goal augmentation method, MGDA, improves the stitching ability of offline goal-conditioned weighted supervised learning on maze benchmarks by filtering augmented goals through a locally Lipschitz dynamics model.

Pith tools