Pith. sign in

REVIEW 3 cited by

Causal Reinforcement Learning using Observational and Interventional Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.14421 v1 pith:L7LHOMQK submitted 2021-06-28 cs.LG

Causal Reinforcement Learning using Observational and Interventional Data

classification cs.LG
keywords learningcausalagentdataofflineenvironmentexperiencesmodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Learning efficiently a causal model of the environment is a key challenge of model-based RL agents operating in POMDPs. We consider here a scenario where the learning agent has the ability to collect online experiences through direct interactions with the environment (interventional data), but has also access to a large collection of offline experiences, obtained by observing another agent interacting with the environment (observational data). A key ingredient, that makes this situation non-trivial, is that we allow the observed agent to interact with the environment based on hidden information, which is not observed by the learning agent. We then ask the following questions: can the online and offline experiences be safely combined for learning a causal model ? And can we expect the offline experiences to improve the agent's performances ? To answer these questions, we import ideas from the well-established causal framework of do-calculus, and we express model-based reinforcement learning as a causal inference problem. Then, we propose a general yet simple methodology for leveraging offline data during learning. In a nutshell, the method relies on learning a latent-based causal transition model that explains both the interventional and observational regimes, and then using the recovered latent variable to infer the standard POMDP transition model via deconfounding. We prove our method is correct and efficient in the sense that it attains better generalization guarantees due to the offline data (in the asymptotic case), and we illustrate its effectiveness empirically on synthetic toy problems. Our contribution aims at bridging the gap between the fields of reinforcement learning and causality.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety

    stat.ML 2026-05 unverdicted novelty 6.0

    Introduces a regularized estimator achieving optimal MSE rates under a new relative balancedness condition while providing safety guarantees that match independent learning when tasks are unrelated.

  2. Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety

    stat.ML 2026-05 unverdicted novelty 6.0

    Matrix-weighted regularization for robust multi-task regression achieves optimal MSE under weaker spectral assumptions and performs no worse than independent learning when balancedness is poor.

  3. CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery

    cs.SE 2026-07 conditional novelty 5.0

    Causal-guided PPO training produces policies that cooperate with rule-based recovery, yielding significant gains in reward, distance, and velocity over non-causal baselines in CARLA driving scenarios.