{"id":"a7c761ef-1c16-4b63-9722-1e330bedc945","arxiv_id":"2501.00854","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"The paper proves graphical identification criteria under which the value of an adaptive policy can be estimated from off-policy data, and shows how to choose state variables to satisfy them.","lead":"This paper gives graphical rules for choosing which historical variables to include as the state when learning a decision policy from data generated by a different policy. It unifies dynamic treatment regimes and offline reinforcement learning, and shows that a poorly chosen state can produce policies worse than doing nothing.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mismatch between two m-separation definitions in §2.2 and §1 may invalidate Lemma 5/6 and hence Theorem 1.","rationale":"The reader's central concern was the reliance on the unreviewed Theorem 3 of Zhao (2024b) for Proposition 2. That is a valid dependency, but my stress-test found a more basic and potentially fatal issue: the paper appears to use two different m-separation notions interchangeably. If the §2.2 walk-blocking definition is indeed not equivalent to the standard ancestral-blocking definition of §1, then the proofs of the key lemmas (which rely on the stronger, paper-specific separations) do not imply the conditional independences needed for Theorem 7. This is not an external-consensus disagreement but an internal inconsistency. The proposed test is concrete: a small graph where the two definitions provably diverge. If the divergence is confirmed, the proof of the central claim is incomplete as written, and no amount of verifying Zhao (2024b) would fix it. I therefore recommend UNVERDICTED rather than CONDITIONAL: the paper's simulation and unification are valuable, but the correctness of the main theorem is not established until the equivalence of the m-separation notions is settled and the lemmas re-verified under the correct definition. This concern is raised in good faith; if the authors can show the equivalence (e.g., by restricting to the relevant d-SWIGs), the verdict should return to CONDITIONAL.","tokens_in":26411,"tokens_out":17138,"duration_ms":154462,"concrete_test":"Construct the small ADMG with vertices A, B, C, D, edges A <-> C, C <-> B, C -> D, and take L = {D}. Under the paper's §2.2 walk-blocking rule, verify that A and B are m-separated given L; under the standard ancestral-blocking definition of §1 (and Richardson 2003), verify they are m-connected. If the two definitions disagree, the claimed equivalence in §2.2 fails. Then, as a second check, enumerate all walks in the d-SWIG of Figure 4b that are m-separated under the paper's definition but not under ancestral blocking, and test whether any such walk would invalidate Lemma 5 or Lemma 6 by providing a counterexample to the claimed conditional independence.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper defines m-separation in two ways that are not equivalent. Section 1 (and the ADMG literature) uses ancestral blocking: a path is blocked if it contains a collider that is not an ancestor of the conditioning set L. Section 2.2 instead blocks a walk whenever it contains a collider not in L, ignoring whether that collider has a descendant in L. These differ. Example: in the ADMG with A <-> C <-> B and C -> D, conditioning on L = {D}, standard m-separation says A and B are m-connected because C is an ancestor of L, but the paper's walk-definition declares them m-separated because C is not in L. Since the paper's proofs of Lemmas 3 and 4 (and hence Lemmas 5 and 6, Theorem 7, and Theorem 1) use the walk-blocking notion, while Proposition 2's global Markov property is imported from Zhao (2024b) in the standard sense, the chain from graphical separation to conditional independence is not established. If the equivalence claim in §2.2 is false, the central identification result rests on an invalid step, independent of whether Zhao's theorem itself is correct.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a graphical framework for off-policy identification in sequential decision processes. With state sets S_t, decision variables A_t, and reward variables R_t, the authors propose three assumptions—nested states, memorylessness, and a dynamic back-door criterion—and argue that, under dynamic consistency and positivity, the joint distribution of states, actions, and rewards under an adaptive policy g is identified by a g-computation-type formula from the observational distribution. The paper then relates dynamic unconfoundedness to sequential ignorability in dynamic treatment regimes, discusses the implicit randomized-decision assumption in MDP/RL practice and the general non-identifiability of POMDPs, and presents a dynamic pricing simulation in which policy iteration with a state set violating the criteria produces worse policies than the null policy.","tokens_in":26684,"tokens_out":15173,"duration_ms":150505,"significance":"The main result, if correct, is a useful unification: it gives explicit, checkable graphical conditions that span both DTR and MDP settings and clarifies why POMDPs are generally not identified without additional structure. The recursive proof via d-SWIGs is natural, and the simulation study with a realistic container-pricing environment and detailed parameters is a practical asset. The paper also makes the important point that state choice is part of the identification problem, not merely a modeling convenience. The central limitation is that a load-bearing proof step is imported from an unreviewed preprint, which makes independent verification difficult.","major_comments":[{"comment":"The proof of Proposition 2 is a citation to Theorem 3 of Zhao (2024b), an unreviewed arXiv preprint. This proposition is the only bridge from m-separation in G(g) to the conditional independence statements in Lemmas 5 and 6, which in turn drive Theorem 7 and Theorem 1. The manuscript therefore has a load-bearing proof step that is not self-contained. I recommend either proving the needed global-Markov statement in the supplement or restating the cited theorem in full and verifying that its hypotheses apply to the d-SWIG with natural counterfactuals A^-(g) and policy vertices A(g).","section":"Appendix B, proof of Proposition 2"},{"comment":"The displayed formula in Theorem 1 has a notational mismatch: the left-hand side conditions on A_T(g)=a_T, a single decision variable, while the right-hand side is a product over t=1,...,T involving all actions a_t. The left-hand side should be \\bar A_T(g)=\\bar a_T, the full action history. As printed, the central identification formula is not dimensionally consistent, and the subsequent remark claiming that the left side contains all decision variables is only true with the overline correction.","section":"Section 1.1, Theorem 1"}],"minor_comments":[{"comment":"The walk-based blocking definition is easy to misread as conflicting with the ancestral-blocking definition in Section 1; the equivalence relies on allowing non-simple walks, which can pass through descendants of colliders. For example, with A<->C<->B and C->D, conditioning on L={D}, the walk A<->C->D<-C<->B is unblocked, so A and B are m-connected under both definitions. A sentence making this role of non-simple walks explicit would prevent confusion.","section":"Section 2.2"},{"comment":"The scenario labels in Table C.1 all appear as \\rightarrow G with no distinction between G_1, G_2, and G_3, even though Section 6.2 defines different arrow types for the three scenarios. Please render the arrows consistently so that the parameter columns can be matched to the graphs.","section":"Table C.1"},{"comment":"The column 'Regret (%)' appears to report improvement relative to the null policy, with negative values indicating worse-than-null performance; please define this quantity explicitly, since standard regret is nonnegative.","section":"Table 1"},{"comment":"There are numerous typographical errors that should be corrected in revision, including 'resemblence', 'identifiying', 'denscendants', 'sophiscated', 'dicussion', and 'oer' (Section 6.3).","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central identification claim is plausible and the paper is a good fit for the journal, but the proof depends on an unreviewed self-cited preprint. I would ask the authors to make the proof self-contained or at least explicitly state and verify the imported theorem. The remaining issues are local and should be addressable in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere is my take on arXiv:2501.00854. The paper offers a graphical framework for selecting state variables in sequential off-policy learning, spanning DTRs and offline RL. The dynamic back-door criterion, the nested-states and memorylessness conditions, and the warning about latent projections creating \"phantom confounding\" are genuinely new and practically useful. The container-logistics simulation makes the stakes concrete: using a state that violates the criteria can produce policies worse than the null policy. Credit where due—this is a thoughtful unification, not a repackaging.\n\nThe soft spot is more than cosmetic. Section 2.2 defines m-separation using walk-blocking: a collider not in the conditioning set blocks a walk. Section 1 uses the standard ancestral definition: a collider blocks only if it is not an ancestor of the conditioning set. These are not equivalent. Example: A<->C<->B with C->D, conditioning on D. Standard m-separation says A and B are connected because C is an ancestor of D; the paper's walk definition declares them separated because C is not in D. The proofs of Lemmas 3 and 4, and hence Lemmas 5, 6, Theorem 7, and Theorem 1, are carried out with the walk notion. Since Assumptions 2 and 3 are stated in the ancestral sense, the implication from graphical separation to conditional independence is not established as written. This is a load-bearing flaw, not a typo. It might be repairable—rework the lemmas with the correct ancestral-blocking calculus—but the current proof chain does not go through.\n\nA secondary weakness is Proposition 2, the global Markov property for d-SWIGs, which is imported from Zhao (2024b), an unreviewed preprint. The main theorem leans on it. The authors should either give a self-contained proof or include a full derivation in the supplement. Also, the simulation is not fully reproducible as shipped: the graph-simulator package is cited without a version and the full code is not included. Minor, but worth asking for.\n\nWho is this for? Researchers working on identification in DTRs and anyone in RL who cares about when off-policy evaluation is legitimate. It deserves a serious referee, not a desk reject. I would send it out with a clear instruction to scrutinize the m-separation equivalence and the dependence on Zhao (2024b). If the authors close the definitional gap, this becomes a solid paper.","headline":"A genuinely useful unification of state selection in off-policy learning, but the proof chain has a load-bearing definitional gap in the m-separation calculus.","tokens_in":27210,"tokens_out":3039,"would_cite":false,"duration_ms":32474,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62M05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper gives graphical identification criteria for off-policy learning: a state set satisfying nested states, memorylessness, and a dynamic back-door condition suffices to identify the distribution under any adaptive policy from data…","keywords":["off-policy learning","state variable selection","acyclic directed mixed graphs","m-separation","dynamic treatment regimes","Markov decision processes","causal identification"],"falsifier":"Run a simulation from a small ADMG with known structural equations that satisfies Assumptions 1-3, record data under a null policy, predict the target-policy distribution from the formula, and compare against direct simulation of the target policy; a mismatch would refute the theorem. Alternatively, exhibit an ADMG satisfying the three assumptions whose d-SWIG distribution violates the global Markov property.","tokens_in":26196,"feed_emoji":"📊","tokens_out":9219,"duration_ms":84598,"temperature":0.7,"pith_summary":"Off-policy learning asks whether the value of a policy of interest can be estimated from data generated under a different policy. This paper claims a precise graphical answer: choose state variables that satisfy a nested-states condition, a memorylessness condition, and a dynamic back-door condition, and the joint distribution of states, actions, and rewards under any adaptive policy is identified by an explicit formula. The result matters because it unifies two literatures that treat the same problem differently: dynamic treatment regimes in statistics and offline reinforcement learning in computer science. It also gives practical guidance for state variable selection in real-world sequential decision problems such as container-logistics pricing. The simulation study shows that a wrong state choice can make a standard reinforcement-learning algorithm converge to a policy worse than the baseline policy.","feed_headline":"Three graph rules make off-policy learning identifiable","feed_subtitle":"The criteria unify dynamic treatment regimes and Markov decision processes, and tell when a policy's value can be estimated.","key_machinery":"The central machinery is the acyclic directed mixed graph (ADMG) with m-separation as its independence criterion, together with the dynamic single-world intervention graph (d-SWIG) representation of policy interventions. m-separation asks whether every path between two variable sets is blocked by a conditioning set, and is the graphical version of conditional independence used in Assumptions 1-3. The d-SWIG $G(g)$ is obtained by splitting each decision vertex into a potential outcome $A_t(g)$, which inherits outgoing edges, and a natural counterfactual $A_t^-(g)$, which inherits incoming edges. Proposition 2 says the distribution of $(V(g), A^-(g))$ is global Markov with respect to $G(g)$; this is what converts the m-separations in dynamic unconfoundedness and memorylessness into the conditional independences of Lemmas 5 and 6. Those two lemmas carry the proof of the recursive equality in Theorem 7, which iterates to give Theorem 1.","core_discovery":"The central claim is that the value of an adaptive policy can be identified from observational data whenever the state variables satisfy three graphical conditions: nested states (Assumption 1), memorylessness (Assumption 2), and the dynamic back-door criterion (Assumption 3), together with dynamic consistency and positivity. Theorem 1 states that under these conditions, and with intermediate rewards contained in states, the joint distribution of all states, actions, and rewards under the target policy $g$ equals $P(r_{T+1} \\mid a_T, s_T) \\prod_{t=1}^T g(a_t \\mid s_t) P(n_t \\mid a_{t-1}, s_{t-1})$, where $P$ is the observational distribution under a null policy. The proof works by recursion: Theorem 7 rewrites the conditional distribution of future variables under a sub-policy using dynamic consistency, unconfoundedness, and the Markov property, so that each step peels off the policy's action density and the observational transition density. This generalizes the static back-door criterion and subsumes the g-computation formula for dynamic treatment regimes.","pith_inferences":["A natural extension, not developed in the paper, is to turn the three graphical conditions into a search procedure that selects a valid state set from candidate variables algorithmically.","Because the identification formula is nonparametric, it gives a template for sequential regression estimators of policy value; whether such estimators inherit gains from the Markov and back-door structure is left open.","The contrast between Assumption 3 and the weaker dynamic unconfoundedness suggests that identification may be policy-dependent: a graph that blocks identification for one policy can still identify another, and exploiting this selectively has not been explored.","The graph conditions also suggest empirical diagnostics: one could test the implied conditional independences on observed data before committing to a state set, and a failure would signal non-identifiability rather than just finite-sample error."],"forward_implications":["When the state set satisfies the three criteria, the value of any adaptive policy can be estimated consistently from data generated by a different policy without additional assumptions beyond dynamic consistency and positivity.","The framework recovers Robins' g-formula for dynamic treatment regimes as a special case, so the dynamic treatment regime identification logic is subsumed by the graph conditions.","For Markov decision processes, the paper makes explicit the randomized-decisions assumption that is usually implicit, and it explains why partially observed MDPs are generally not identified.","When previous actions are excluded from the state variables, verifying only the immediate back-door condition is enough, which simplifies state selection in common MDP-like settings.","The pricing simulation shows that using a state set that violates the criteria can lead a policy-iteration algorithm to learn a policy with negative regret relative to the null policy."],"supporting_citations":[{"why":"Imported as Proposition 2: proves the d-SWIG distribution is global Markov, the step that converts m-separations into conditional independences in the main proof.","marker":"Zhao (2024b)"},{"why":"Introduces single-world intervention graphs, whose dynamic version is used to represent policy interventions and potential outcomes.","marker":"Richardson and Robins (2013)"},{"why":"Introduces m-separation for ADMGs, the graphical independence criterion on which Assumptions 1-3 rest.","marker":"Richardson (2003)"},{"why":"The static back-door criterion that Assumption 3 extends to the dynamic setting.","marker":"Pearl (1993)"},{"why":"The g-computation formula for dynamic treatment regimes, which the paper recovers as a corollary and thereby unifies.","marker":"Robins (1986)"},{"why":"Further development of g-computation and sequential ignorability, used in the comparison with DTR assumptions.","marker":"Robins (1997)"},{"why":"Supplies the standard MDP formalism and factorization whose implicit causal assumptions the paper analyzes.","marker":"Sutton and Barto (2018)"}],"fun_headline_variants":["Graph criteria unify DTRs and MDPs for off-policy learning","Graphical test reveals when offline policy value is identifiable","Three graph rules guarantee off-policy identifiability","Graph approach shows when you can trust off-policy estimates","Graphical conditions for valid off-policy learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the d-SWIG distribution satisfies the global Markov property, a result imported from an unreviewed preprint; if that theorem has hidden conditions, the main identification proof collapses even if the formula is true.","fun_headline_variants_meta":{"raw":{"variants":["Graph criteria unify DTRs and MDPs for off-policy learning","Graphical test reveals when offline policy value is identifiable","Three graph rules guarantee off-policy identifiability","Graph approach shows when you can trust off-policy estimates","Graphical conditions for valid off-policy learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1473,"prompt_tokens":982,"completion_tokens":491,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":409}},"tokens_in":598,"tokens_out":491,"duration_ms":5802,"temperature":1.0,"reasoning_tokens":409,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:41:35.805917+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a simulation from a small ADMG with known structural equations that satisfies Assumptions 1-3, record data under a null policy, predict the target-policy distribution from the formula, and compare against direct simulation of the target policy; a mismatch would refute the theorem. Alternatively, exhibit an ADMG satisfying the three assumptions whose d-SWIG distribution violates the global Markov property.","supporting_citations":[],"review_version":1}