Pith. sign in

REVIEW 3 cited by

Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.09907 v3 pith:KKV4A5JH submitted 2021-02-19 stat.ML cs.LG

classification stat.MLcs.LG
keywords dataobservationalvariablesconfoundeddynamicsofflinetransitionalgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In offline reinforcement learning (RL) an optimal policy is learned solely from a priori collected observational data. However, in observational data, actions are often confounded by unobserved variables. Instrumental variables (IVs), in the context of RL, are the variables whose influence on the state variables is all mediated by the action. When a valid instrument is present, we can recover the confounded transition dynamics through observational data. We study a confounded Markov decision process where the transition dynamics admit an additive nonlinear functional form. Using IVs, we derive a conditional moment restriction through which we can identify transition dynamics based on observational data. We propose a provably efficient IV-aided Value Iteration (IVVI) algorithm based on a primal-dual reformulation of the conditional moment restriction. To our knowledge, this is the first provably efficient algorithm for instrument-aided offline RL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression

    stat.ML 2025-01 conditional novelty 8.0 of 10

    DFIV achieves minimax-optimal rates in Besov spaces for nonparametric IV regression, and beats fixed-feature estimators for spatially inhomogeneous targets.

  2. Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

    cs.LG 2024-12 conditional novelty 7.0 of 10

    A two-way deconfounder algorithm that models unmeasured confounders as per-trajectory and per-timestep latent factors and uses a neural tensor network for off-policy evaluation.

  3. The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An optimism-based algorithm with nonparametric instrumental variables learns an epsilon-optimal policy under information asymmetry and knowledge transfer with O~(1/epsilon^2) sample complexity.

Pith tools