{"id":"1c8ebb86-cc00-4766-9939-3fac8cc1c1ed","arxiv_id":"2505.00304","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper derives a bridge-function identification result that supports root-n-rate policy value estimation and policy-gradient policy learning for continuous actions under unmeasured confounding.","lead":"This paper proposes methods to evaluate and learn policies from offline data when actions are continuous and some confounding variables are never observed. It gives an identification formula, estimators with error guarantees, and tests them on simulations and German relationship panel data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed n^{-1/2} evaluation and regret rates are not established because Theorem 4.1 has no sample-size term and Theorem 4.2 is not a genuine finite-sample bound.","rationale":"I read the paper as making three connected claims: identification of Q/V-bridges under Assumptions 1–4, an estimator with O(n^{-1/2}) error for a fixed policy, and a learned policy with O(p^{1/2} n^{-1/2}) regret. The identification step is not my main objection: the conditional-moment equation (3.3) is consistent with the stated definitions and the completeness assumptions, and even though completeness is strong and unverifiable, it is a standard identifiability condition in proximal causal inference. The more concrete and immediately visible weakness is in Section 4. Theorem 4.1's displayed Bellman-error bound is missing any dependence on the sample size, so it cannot be the full finite-sample guarantee needed for Theorem 4.2. The use of o_P(n^{-1/2}) inside a claimed high-probability finite-sample bound is not a valid probabilistic statement, and the tuning condition λ_n=o(n^{-1/(1+α)}) does not imply that nuisance estimation error is negligible at the n^{-1/2} scale. Since Theorem 4.3's regret bound is built on a uniform extension of Theorem 4.2, the headline rates are unsupported as written. This is a serious but potentially fixable gap: if the full proofs supply the missing stochastic terms and show the Riesz-representer cancellation that yields n^{-1/2} despite slower Bellman-error convergence, the paper's conclusions may hold. The reader's conditional verdict is therefore appropriate, but the binding reason is the unproven finite-sample rates rather than completeness alone.","tokens_in":21267,"tokens_out":24037,"duration_ms":259772,"concrete_test":"Independently re-derive the finite-sample Bellman error for the estimator (3.5) using standard empirical-process bounds on the sample analogue of the minimax objective. If the resulting high-probability bound contains a term of order (n λ_n)^{-1/2} or n^{-1/(2+α)} rather than only λ_n, then Theorem 4.2's O(n^{-1/2}) conclusion cannot follow from Theorem 4.1; this would settle that the rate claims currently lack support. A complementary check is to simulate the RKHS setting of Section 5 with n increased by a factor of 4 and verify whether the empirical RMSE of TJ(π) actually halves, as the claimed n^{-1/2} rate requires.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claims are the O(n^{-1/2}) policy-evaluation error and the O(p^{1/2} n^{-1/2}) policy-regret bound. These rest on Theorem 4.1, whose displayed Bellman-error bound is “∥Tπ(·,·,·;Q̂π)-Q̂π∥² ≲ (1/κ²)[λ_n(1+h_1²(Q_π)){1+ log(1/δ)}]”. The right-hand side contains no n-dependent stochastic term. Taken literally, sending λ_n→0 arbitrarily fast would drive the error to zero at fixed n, which is impossible for an estimator built from n finite trajectories. The missing term is the empirical-process/optimization error of the minimax estimator (3.5).\n\nTheorem 4.2 then claims |TJ(π)-J(π)| ≤ C√(log(2/δ)/(2n))+o_P(n^{-1/2}) with probability at least 1-δ. This is internally inconsistent: an o_P remainder cannot appear in a high-probability finite-sample bound, and the condition λ_n=o(n^{-1/(1+α)}) does not, by itself, control the Q-bridge estimation error at o_P(n^{-1/2}) because Theorem 4.1 supplies no such control. Theorem 4.3 inherits the problem, since the regret bound requires a uniform version of the missing n^{-1/2} evaluation bound over the policy class Π. The identification argument in Theorem 3.1 is plausible under Assumptions 1–4, but the paper's headline rates are not supported by the displayed theorems.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies offline policy evaluation and learning in an infinite-horizon confounded POMDP with continuous actions. It proposes Q- and V-bridge functions using time-dependent proxy variables, derives the conditional moment equation (3.3), and constructs a minimax estimator with an RKHS-based critic together with a policy-gradient algorithm. The headline theoretical claims are an O(n^{-1/2}) policy-evaluation error (Theorem 4.2), a regret bound O(p^{1/2}n^{-1/2}) (Theorem 4.3), and a supporting Bellman-error bound (Theorem 4.1). The methodology is illustrated with simulations and an application to the Pairfam data.","tokens_in":21584,"tokens_out":8558,"duration_ms":76839,"significance":"If the stated rates were rigorously established, this would be a valuable extension of proximal causal inference to continuous-action infinite-horizon settings, with a computationally attractive kernel decoupling of the minimax estimator and a practically relevant policy-gradient scheme. The identification strategy is plausible and follows the proximal template, and the empirical section, including the Pairfam application, is a strength. However, the displayed theorems in Section 4 do not establish the advertised rates, so the theoretical contribution is currently incomplete.","major_comments":[{"comment":"The Bellman-error bound in Theorem 4.1 has no sample-size term: the displayed right-hand side depends only on λ_n, κ, h_1(Qπ) and δ. Taken literally, sending λ_n→0 would drive the bound to zero at fixed n, which is impossible for an estimator constructed from n finite trajectories. The missing empirical-process/optimization error must appear in the bound; without it, the claimed rate O(n^{-1/(1+α)}) and all subsequent results that rely on Theorem 4.1 are not supported.","section":"Section 4, Theorem 4.1"},{"comment":"Theorem 4.2 is not a genuine finite-sample high-probability bound because it contains an o_P(n^{-1/2}) remainder inside a statement that holds with probability at least 1−δ. In addition, the condition λ_n = o(n^{-1/(1+α)}) does not imply a root-n error for the Q-bridge estimation, since Theorem 4.1 provides no n-dependent control of ∥Q̂π−Qπ∥. The O(n^{-1/2}) policy-evaluation claim is therefore not established.","section":"Section 4, Theorem 4.2"},{"comment":"Theorem 4.3's regret bound requires a uniform version of the root-n evaluation error over the policy class Π. Proposition 4.1 extends Theorem 4.1 uniformly over Π but inherits the same defect: its bound is λ_n times a policy-complexity factor with no n-dependent term. Consequently the advertised O(p^{1/2}n^{-1/2}) regret rate is unsupported.","section":"Section 4, Theorem 4.3 and Proposition 4.1"},{"comment":"The identification theorem is the foundation of the paper, but its proof is not given in the main text and the 'regularity conditions' are left unspecified. The authors should present a complete proof with explicit conditions, either in the main text or in a supplement that is actually included and verifiable; as written, the existence of the bridge functions and the validity of Eq. (3.3) cannot be checked from the manuscript.","section":"Section 3.1, Theorem 3.1"},{"comment":"The policy-gradient formula treats θ as a function of ζ but the displayed derivative does not account for the dependence of Q̂π on ζ through the estimation procedure and through the Bellman equation defining the bridge. The authors should clarify whether Algorithm 2 is an exact gradient method for the objective in (3.10) or a heuristic alternating scheme, and if the latter, what guarantees apply.","section":"Section 3.3, Eq. (3.11)"}],"minor_comments":[{"comment":"There is a typo: 'the the curse of dimensionality' should read 'the curse of dimensionality'.","section":"Introduction"},{"comment":"There is a typo: 'demnstrated' should read 'demonstrated'.","section":"Section 4"},{"comment":"The completeness assumption is unverifiable from data; the authors should more explicitly acknowledge that the identification and all subsequent results depend on it, and discuss its plausibility in the Pairfam application.","section":"Assumption 4"},{"comment":"The y-axis label 'l g MSE' should be 'log MSE'.","section":"Figure 2"},{"comment":"The sentence 'As shown in Figure 2, our proposed OPE method exhibits the smallest bias...' appears to reference the simulation figure rather than a figure in the application section; please check the cross-reference.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper has a plausible identification framework and a useful empirical component, but the theoretical section as written does not support the advertised rates. The missing stochastic term in Theorem 4.1 and the o_P remainder in Theorem 4.2 are substantive proof gaps rather than presentation issues. I recommend major revision with a full rewrite of Section 4 and a clarification of the algorithm-to-theory mapping. I do not see grounds for rejection, because the identification result and the proposed estimator are defensible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline claim of root-n policy evaluation and regret rates is not actually proven in this draft. Theorem 4.1's displayed bound has no sample-size term—the right-hand side is just λ_n times a log(1/δ) factor, so taken literally it lets you drive the Bellman error to zero at fixed n by shrinking λ_n. That cannot hold for an estimator built from n trajectories, and it means the missing empirical-process term is load-bearing. Theorem 4.2 then mixes a high-probability finite-sample bound with an o_P(n^{-1/2}) remainder, which is internally inconsistent, and Theorem 4.3 inherits the problem. So the headline rates should not be accepted as they stand.\n\nWhat is genuinely new is the combination of continuous actions, infinite horizon, non-Markovian confounding, and bridge-function identification using the previous state-action pair as action proxy and W_t as reward proxy. Theorem 3.1 is plausible and follows the proximal inference template; the kernel-based minimization in Theorems 3.3–3.4 gives a tractable estimator; and the simulations plus Pairfam application offer a reasonable sanity check where the method beats MDP baselines. The citation pattern looks fine, with proper credit to Shi et al., Tchetgen Tchetgen et al., and the proximal RL literature.\n\nThe soft spots beyond the rates: the policy-gradient update in (3.11) leaves ∇_ζ Qπ unspecified, so the algorithm is not fully implementable as written. The text calls the estimator \"unbiased\" in Section 3.3 despite the regularization in (3.5). And while Assumption 4 completeness is standard for this literature, Assumption 2—W_t independent of the action given the full state—is a substantive domain assumption that is quite strong in the Pairfam example.\n\nIf the supplementary proofs supply the missing empirical-process term and Theorem 4.2 is corrected to a genuine finite-sample bound, this would be a solid contribution. As it stands, the identification and algorithm are worth reading, but the theoretical claims are unsubstantiated. I'd send it to a serious referee: the fixes seem addressable, and the identification part deserves scrutiny. I wouldn't cite the rate claims until they are fixed.","headline":"The identification and algorithm are a genuine extension to continuous-action infinite-horizon POMDPs, but the root-n rate claims are not supported by the displayed theorems.","tokens_in":22148,"tokens_out":3703,"would_cite":false,"duration_ms":34449,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62G05","68T05","90C40"],"pacs":[],"model":"deepseek-v4-flash","headline":"Bridge functions recover policy values and near-optimal policies in continuous-action RL with hidden confounders.","keywords":["off-policy evaluation","off-policy learning","confounded POMDP","unmeasured confounding","continuous action space","bridge functions","minimax estimation","policy gradient"],"falsifier":"Run the estimator on a synthetic confounded POMDP that satisfies Assumptions 1-3 but generates $S_t$ independently of $(O_{t-1},A_{t-1})$ given $(O_t,A_t)$, so completeness fails; if the estimated $J(\\pi)$ does not converge to the true policy value as $n$ grows, identification is genuinely doing the work. A data-only probe is to fit the conditional moment equation from multiple random initializations and with different kernel bandwidths: if the resulting policy-value estimates differ by more than the claimed $n^{-1/2}$ scale, the equation is admitting multiple bridge functions and the central identification is not operational.","tokens_in":21063,"feed_emoji":"🎯","tokens_out":7316,"duration_ms":69939,"temperature":0.7,"pith_summary":"This paper tackles offline reinforcement learning when the behavior policy depends on an unobserved state variable, actions are continuous, and the horizon is infinite. It claims that even though the true state is hidden, proxy variables observed alongside the rewards allow the value of any target policy to be identified from batch data, without discretizing the action space. On top of that identification, it constructs a penalized minimax estimator for the value and a policy-gradient procedure that searches the best policy within a parametric class. The central promise is practical: consistent policy evaluation at $O(n^{-1/2})$ error and learned policies with $O(p^{1/2}n^{-1/2})$ regret, in a setting where existing methods either assume away confounding or restrict to discrete actions.","feed_headline":"Hidden confounders no longer block continuous-action offline RL","feed_subtitle":"Proxy variables recover policy values at the optimal nonparametric rate and drive near-optimal learned policies.","key_machinery":"The load-bearing object is the pair of bridge functions: $Q^\\pi(O_t,W_t,A_t)$ and $V^\\pi(O_t,W_t)$, defined so that their conditional expectations given the full state equal the state-action and state value of the target policy. The identification equation uses the previous observed state-action pair $(O_{t-1},A_{t-1})$ as an action-inducing proxy and the reward proxy $W_t$ to condition away the unobserved state, yielding a conditional moment restriction with the Bellman residual. The estimation machinery is a penalized minimax problem whose inner maximization, over critic functions in a bounded RKHS, collapses to a closed form; this reduces the procedure to a single-stage regularized minimization that supports stochastic-gradient optimization and a policy-gradient loop over the policy parameters.","core_discovery":"The paper's central claim is Theorem 3.1: in an infinite-horizon confounded POMDP satisfying Assumptions 1-4, there exist Q-bridge and V-bridge functions, and a particular pair solves the conditional moment equation $E[Q^\\pi(O_t,W_t,A_t)-R_t-\\gamma V^\\pi(O_{t+1},W_{t+1})\\mid O_{t-1},A_{t-1},O_t,A_t]=0$. Consequently the policy value $J(\\pi)=E[V^\\pi(O_0,W_0)]$ is nonparametrically identified from observed trajectories even though the behavior policy depends on the unobserved state. The paper then shows that the associated minimax estimator is consistent, attains $O(n^{-1/2})$ finite-sample error for a fixed policy (Theorem 4.2), and that the policy-gradient learner has regret $O(p^{1/2}n^{-1/2})$ (Theorem 4.3). The same identification also gives a tractable algorithm because the inner maximization over critic functions decouples into a closed-form kernel expression when critics are modelled in a reproducing kernel Hilbert space.","pith_inferences":["One implication left implicit is that the completeness assumption is untestable from data, so a practical safeguard would be to estimate the bridge functions under several candidate proxy sets and report the spread of the resulting policy values; the paper does not develop such a sensitivity analysis.","The conditional-moment identification is not tied to the discounted infinite-horizon objective; the same Q/V-bridge argument should extend to average-reward and finite-horizon objectives with only the Bellman residual changed.","A direct testable extension would be to evaluate the learned Pairfam policy by comparing its estimated long-term satisfaction against couples whose actual intimacy frequency matches the policy's recommended distribution in a hold-out wave.","The policy-gradient formulation avoids the per-iteration argmax over a continuous action space; the same trick could be imported into proximal value-based methods for POMDPs that currently rely on discrete actions."],"forward_implications":["Policy evaluation in confounded POMDPs with continuous actions no longer requires discretizing the action space; the bridge-function estimator is consistent and achieves $O(n^{-1/2})$ error for fixed target policies.","The learned in-class policy attains regret $O(p^{1/2}n^{-1/2})$, so batch data with unmeasured confounders can support near-optimal treatment or intervention recommendations.","The RKHS decoupling of the minimax problem yields a closed-form inner maximization, making the method computationally feasible for continuous states and actions rather than a purely theoretical identification result.","Because the identification uses reward proxies $W_t$ that need not cause $R_t$, the method applies to settings such as survey panels where auxiliary variables are correlated with outcomes but not driven by the action."],"supporting_citations":[{"why":"Supplies the proximal causal inference setting and the completeness conditions under which bridge functions exist.","marker":"Tchetgen Tchetgen et al. (2020)"},{"why":"Establishes proxy-variable identification with an unmeasured confounder, the template for the reward-proxy argument.","marker":"Miao et al. (2018)"},{"why":"The confounded-POMDP minimax OPE method whose rates and direction-function analysis are extended to continuous actions and a wider function class.","marker":"Shi et al. (2022a)"},{"why":"Provides the minimax Q-function and direction-function framework used for the finite-sample error bound.","marker":"Uehara et al. (2020)"},{"why":"Supplies the well-posedness (kappa) condition that controls the Bellman-error-to-Q-error conversion.","marker":"Chen and Qi (2022)"},{"why":"Shows proximal identification and estimation in POMDPs, the discrete-action baseline that this paper generalizes.","marker":"Bennett and Kallus (2023)"},{"why":"Gives minimax-optimal policy learning under unobserved confounding, the benchmark for the regret analysis.","marker":"Kallus and Zhou (2021)"},{"why":"Introduces policy gradient methods for confounded POMDPs, the starting point for the continuous-action policy search.","marker":"Hong et al. (2023)"}],"fun_headline_variants":["Continuous-action offline RL works despite unmeasured confounders","Offline RL with continuous actions: hidden confounders no longer block","Identification result enables optimal policy under confounding","Minimax estimator for confounded continuous-action offline RL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on an unverifiable completeness condition: the observed proxies must be rich enough that any two different hidden-state configurations produce different predictions from the observed variables. If that fails, the bridge functions are not identifiable and the error and regret guarantees collapse.","fun_headline_variants_meta":{"raw":{"variants":["Continuous-action offline RL works despite unmeasured confounders","Offline RL with continuous actions: hidden confounders no longer block","Identification result enables optimal policy under confounding","Minimax estimator for confounded continuous-action offline RL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1424,"prompt_tokens":914,"completion_tokens":510,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":444}},"tokens_in":530,"tokens_out":510,"duration_ms":5814,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:46:06.289524+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the estimator on a synthetic confounded POMDP that satisfies Assumptions 1-3 but generates $S_t$ independently of $(O_{t-1},A_{t-1})$ given $(O_t,A_t)$, so completeness fails; if the estimated $J(\\pi)$ does not converge to the true policy value as $n$ grows, identification is genuinely doing the work. A data-only probe is to fit the conditional moment equation from multiple random initializations and with different kernel bandwidths: if the resulting policy-value estimates differ by more than the claimed $n^{-1/2}$ scale, the equation is admitting multiple bridge functions and the central identification is not operational.","supporting_citations":[{"cited_title":"(2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp","cited_arxiv_id":null,"evidence_quote":"Provides the minimax Q-function and direction-function framework used for the finite-sample error bound."},{"cited_title":"and Qi, Z","cited_arxiv_id":null,"evidence_quote":"Supplies the well-posedness (kappa) condition that controls the Bellman-error-to-Q-error conversion."},{"cited_title":"and Kallus, N","cited_arxiv_id":null,"evidence_quote":"Shows proximal identification and estimation in POMDPs, the discrete-action baseline that this paper generalizes."}],"review_version":1}