Pith. sign in

REVIEW 1 cited by

Structured Difference-of-Q via Orthogonal Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08697 v3 pith:35HMAGKJ submitted 2024-06-12 stat.ML cs.LGmath.OCstat.ME

classification stat.MLcs.LGmath.OCstat.ME
keywords estimationlearningcausalcontrastfunctionfunctionsleveragemany
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Offline reinforcement learning is important in many settings with available observational data but the inability to deploy new policies online due to safety, cost, and other concerns. Many recent advances in causal inference and machine learning target estimation of causal contrast functions such as CATE, which is sufficient for optimizing decisions and can adapt to potentially smoother structure. We develop a dynamic generalization of the R-learner (Nie and Wager 2021, Lewis and Syrgkanis 2021) for estimating and optimizing the difference of $Q^\pi$-functions, $Q^\pi(s,1)-Q^\pi(s,0)$ (which can be used to optimize multiple-valued actions). We leverage orthogonal estimation to improve convergence rates in the presence of slower nuisance estimation rates and prove consistency of policy optimization under a margin condition. The method can leverage black-box nuisance estimators of the $Q$-function and behavior policy to target estimation of a more structured $Q$-function contrast.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Estimating the behavior policy from longer histories provably reduces the asymptotic variance of importance-sampling based off-policy evaluation estimators at the cost of increased finite-sample bias, with different e...

Pith tools