{"id":"4f952340-9ac8-49b8-862a-da1e9e2bde7e","arxiv_id":"2608.10896","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Brownian-bridge self-normalized confidence regions for constant-stepsize TD learning under Markovian sampling, centered on the Richardson-Romberg stationary target or, in a horizon-indexed regime, the projected Bellman solution.","lead":"Temporal-difference (TD) learning is how reinforcement-learning agents update their estimates of how much future reward a state is worth. This paper constructs confidence intervals for those value estimates that remain valid even when the agent sees a single long stream of dependent experience and uses a constant learning rate, without asking the user to pick a bandwidth or batch length.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The fixed-stepsize and horizon-indexed FCLTs both pass through an unproven Poisson equation for an unbounded augmented chain; until Proposition 3 and Eq. (36) are checkable, the pivotal self-normalizer is unsupported.","rationale":"Read in good faith: the paper's goals are clear, and the main theorems are internally coherent. Proposition 10 is a correct transfer from an FCLT to the Brownian-bridge pivotal law; Lemma 9's proof is sound; the online one-pass recursion is a correct O(Ld + qd + q²) implementation; the rate window in Section 6.3 is consistent with Assumption 4. The empirical work is unusually honest: it discloses under-coverage in small-stepsize d=10 FrozenLake cells, reports a post-hoc extension of the horizon, and provides code. The single soft spot is that the technical machinery distinguishing this paper from prior work — the uniform block-contraction bound and the Poisson equation for a chain with an unbounded iterate coordinate — is delegated to an Online Appendix that is not in the reviewed text. The reader's CONDITIONAL verdict is the right calibration; my pass does not move it. The concern is not that the result is wrong, but that the central FCLTs cannot be checked without the missing proofs. I would want Proposition 3 and Eq. (36) made checkable before accepting, and I would want the main text to state explicitly which proof steps use the finite-memory approximation and how the α-dependence of Vα enters Eq. (18).","tokens_in":20303,"tokens_out":21011,"duration_ms":247765,"concrete_test":"Obtain Online Appendix 1 and independently verify: (a) the Lyapunov block-contraction bound, Eq. (11), with constants uniform in the starting state y, using the constants KQ, MQ, and ℓ_Q defined by Eqs. (38)-(39); (b) Proposition 3's Poisson solution for the unbounded augmented chain, including the α-dependence of ||Vα||_{L²(πα)} needed for Eq. (18); (c) the central reduction (36), especially that the boundary term from identity (19) is o_p(√n) under Assumption 4(ii)-(iv). If the appendix cannot be obtained, reconstruct Proposition 3 directly: prove existence of Vα ∈ L²(πα) on Y × R^d under only Assumptions 1-3. Any additional condition needed in that reconstruction would move the verdict from CONDITIONAL to UNVERDICTED.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Both FCLTs in the paper pass through a Poisson-equation step that the submitted text does not prove. Proposition 3 asserts a solution Vα ∈ L²(πα) to Vα − PαVα = φα for the augmented chain Z_t^α = (Y_t, h_{t-1}^α), where h^α is an unbounded, α-slowly-forgetting function of the data chain. Assumption 1's uniform geometric ergodicity applies to Y_t, not to the product chain; the second coordinate is nonlinear and forgets at rate cα, so a classical uniformly ergodic Markov-chain FCLT does not apply. The paper says a finite-memory approximation and projective summability are used, but these arguments are in Online Appendix 1, which is not part of the distributed text. The martingale-difference decomposition (17), the covariance formula (18), and the horizon-indexed reduction (36) all depend on this proposition; Theorem 5, Corollary 6, and Corollary 8 inherit the dependence. The central claim is therefore conditional on a checkable proof of Proposition 3 and of Theorem 1's uniform Lyapunov bound (11). This is a missing-support concern, not an observed contradiction; the visible self-normalization transfer (Lemma 9 and Proposition 10) is standard and appears correct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops asymptotic inference for constant-stepsize linear temporal-difference learning under Markovian sampling. It constructs a stationary fixed-stepsize recursion, proves a fixed-stepsize functional central limit theorem whose covariance retains the multiplicative TD noise, derives a joint FCLT for parallel Richardson–Romberg recursions sharing one trajectory, and shows that a Brownian-bridge self-normalizer yields asymptotically pivotal confidence regions for prespecified contrasts without estimating the long-run covariance or choosing a bandwidth or batch length. A horizon-indexed regime, with stepsize constant within each run and decreasing across runs, is proposed for inference on the projected Bellman solution. The main theoretical results in Sections 4 and 6 rest on stability and Poisson-equation arguments whose proofs are deferred to Online Appendix 1, which is not part of the distributed text; the visible Appendix B gives a complete proof of the self-normalization transfer.","tokens_in":20500,"tokens_out":10334,"duration_ms":106101,"significance":"If the deferred proofs are correct, the paper makes a substantial contribution: it provides a one-pass, memory-bounded confidence-region method for constant-stepsize TD that avoids the notoriously difficult long-run covariance estimation step, and it clarifies that the inferential center is the RR stationary target at fixed stepsize and the projected Bellman solution only under the separate horizon-indexed rate conditions. The visible part of the paper is strong: Lemma 9 and Proposition 10 are proved carefully, the related-work table is informative, the numerical section reports honest finite-sample diagnostics including under-coverage and initialization sensitivity, and a reproducibility repository is provided. The main weakness is verifiability: Theorem 1, Proposition 3, and the horizon-indexed reduction (36) are load-bearing and are not proved in the reviewed text. I found no circularity in the target definition: the paper explicitly defines θ_RR,α as the stationary mean of the RR recursion and claims coverage for that object at fixed stepsize, while using external bias bounds and rate conditions to move to θ_* in the horizon-indexed regime.","major_comments":[{"comment":"The load-bearing technical inputs—Theorem 1's uniform block-contraction bound (11), Proposition 3's Poisson solution (15)–(17), and the horizon-indexed reduction (36)—are proved only in Online Appendix 1, which is not distributed with the reviewed text. I was able to verify Appendix B, Lemma 9, and Proposition 10, but not the stability and Poisson-equation steps on which Theorems 4, 5, 7 and their corollaries depend. Theorem 1's uniformity in the starting state y and in α is used to construct the stationary recursion and to transfer limits from stationary starts to fixed deterministic initializations; Proposition 3 supplies the covariance formula (18); and (36) is the step that makes the multiplicative component and the TD boundary negligible in the horizon-indexed regime. This is a missing-support issue for the central claim, not a presentation detail.","section":"§1.2, §4.1, §6.2"},{"comment":"Proposition 3 asserts an L2 solution Vα to Vα − PαVα = φα for the augmented chain Z_t^α = (Y_t, h_{t−1}^α), where h^α is unbounded and forgets at rate cα. Uniform geometric ergodicity of Y_t alone does not imply the Poisson equation for the product chain, so the finite-memory approximation and projective summability argument must be supplied. Since equation (17), the martingale-difference covariance (18), and the fixed-stepsize FCLT all pass through this proposition, the current text cannot support Theorem 4 without that argument.","section":"§4.2, Proposition 3"},{"comment":"The central reduction (36) is asserted with proof deferred to Online Appendix 1. It is exactly the statement that the centered multiplicative component ζ^RR_{t,n}, the TD boundary term, and the residual RR target shift are uniformly o_p(√n), and it identifies the limiting additive covariance Ω0 in (34). The rate assumptions in Assumption 4(i)–(iii) do not by themselves visibly deliver the displayed uniform-in-r Lp statement; the proof must show the projective maximal controls. Without this reduction, Theorem 7 and Corollary 8 are unsupported.","section":"§6.2, Eq. (36)"},{"comment":"The RR bias bound (21), which is essential for the horizon-indexed inference for θ*, is imported from Huo et al. [2026] and translated in Online Appendix 1. The main text gives no statement of the bias expansion's order, the dependence of its remainder constant on the fixed cancellation order, or the exact stepsize threshold under which it holds. Since the fixed-stepsize results do not need (21) but Corollary 8 does, the paper should either state the expansion in the main text with its constants and thresholds or include the proof in the distributed version.","section":"§4.3, Eq. (21)"}],"minor_comments":[{"comment":"The claim E[φ_t^α] = 0 is correct but deserves a one-line derivation using stationarity of the TD recursion and E[b_t] = Aθ*; as written the reader must reconstruct the algebra.","section":"§4.2, Eq. (13)"},{"comment":"The rate window is stated with strict inequalities, but Table 5 labels ν = 1/2 as the 'upper boundary' even though Assumption 4(iv) requires α_n√n → ∞; the table should state explicitly that ν = 1/2 is outside the admissible regime.","section":"§6.3 and Table 5"},{"comment":"The FrozenLake reset convention—after a terminal transition the process resets to the designated start state while the terminal feature remains in the TD update—needs a formal description of the induced transition kernel to justify treating the resulting sequence as a stationary Markov chain for the asymptotic analysis.","section":"§7.1"},{"comment":"The sentence 'Both convergences above remain valid pointwise...' is ambiguous: if it means finite-dimensional convergence only, that is weaker than what Corollary 6 needs, and if it means weak convergence in D, the word 'pointwise' is misleading. The statement should be clarified.","section":"§4.3, Theorem 5"},{"comment":"The one-pass implementation is clearly explained, but the paper does not report the simulated Brownian critical values κ_{q,1−η} used in the experiments; including a small table or the simulation code in the main text or Online Appendix 2 would make the procedure fully reproducible from the paper alone.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The main text alone is not refereable in the usual sense because the central stability and Poisson-equation proofs are in an absent Online Appendix 1. I would be willing to review the complete version with that appendix. My assessment matches the conditional reader's take: the visible self-normalization transfer is correct, the target definitions are coherent, and I found no internal contradiction; the sole obstacle is verifiability of the deferred technical arguments. If the appendix delivers the stated claims, I expect the paper to be publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper gives a genuinely new recipe: constant-stepsize linear TD under Markov sampling, with Richardson–Romberg recursions sharing one trajectory, and a Brownian-bridge self-normalizer that yields pivotal confidence regions without estimating the long-run covariance. Second, the main inference results are conditional on a block of technical proofs—Theorem 1's Lyapunov block contraction, Proposition 3's Poisson equation for the augmented chain, and the central reduction (36)—all deferred to Online Appendix 1, which is not in the distributed text. I cannot check those steps, and neither can a referee. The visible part of the argument, especially the self-normalization transfer in Appendix B, is standard and correct.\n\nWhat is new and good: the separation between the fixed-stepsize stationary target and the projected Bellman solution is clean, and the horizon-indexed rate window is explicit. The one-pass algorithm with memory independent of the horizon is a real practical win. The experiments are unusually honest: they report the FrozenLake small-stepsize undercoverage, label the horizon extension as post-hoc sensitivity analysis, and share code and data. The pilot selection of the batch-mean rule is disclosed. Credit where due.\n\nSoft spots, in proportion. The missing appendix is the main one. The augmented chain Z_t^α has an unbounded second coordinate that forgets at rate cα; uniform geometric ergodicity of the data chain does not automatically transfer, so Proposition 3 is genuinely load-bearing. If that Lyapunov bound fails at the stated uniformity, the FCLTs and the pivotal inference all collapse. This is missing support, not an observed contradiction, but it is exactly the step that should be in the main text or a reliably available supplement. The horizon-indexed theory inherits the issue through the unproven reduction (36). Minor: the deterministic-start rate condition (31), α_n√n → ∞, makes the exponent window tighter than the stationary version, but the paper states this clearly.\n\nWho this is for: researchers doing statistical inference for reinforcement learning, especially those wanting tuning-free confidence regions from one pass of TD. The full proof of the main theorem is not in the submitted packet, which makes the verdict conditional, but this is a condition, not a fatal flaw.\n\nRecommendation: yes, send it to peer review, but get the online appendix into the referees' hands first. If Theorem 1 and Proposition 3 check out, this is a strong paper.","headline":"Solid, genuinely new inference recipe for constant-stepsize TD—but the load-bearing Poisson-equation proof is only in an unavailable appendix, so treat the FCLTs as conditional.","tokens_in":21123,"tokens_out":3029,"would_cite":true,"duration_ms":28728,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L20","60F17","62F25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Fixed-stepsize TD learning under Markovian sampling admits asymptotically pivotal confidence regions for prespecified state-value contrasts without estimating the long-run covariance, via a Brownian-bridge self-normalizer.","keywords":["temporal-difference learning","self-normalization","Markovian sampling","constant stepsize","functional central limit theorem","Richardson–Romberg extrapolation","Brownian bridge","policy evaluation"],"falsifier":"Numerically evaluate the uniform block-contraction inequality (11) on a slow-mixing finite-state Markov chain across many starting states $y$ and admissible stepsizes $\\alpha$; if any pair violates the stated exponential decay, the stationary-recursion construction and therefore the self-normalized coverage claim collapse.","tokens_in":20016,"feed_emoji":"📊","tokens_out":17686,"duration_ms":142107,"temperature":0.7,"pith_summary":"Constant-stepsize temporal-difference (TD) learning is a streaming method for policy evaluation, but the usual asymptotic theory has not delivered confidence statements from a single Markov trajectory without long-run covariance estimation. This paper tries to prove that, for linear TD with a fixed stepsize, a Brownian-bridge self-normalizer built from the observed partial-sum path makes inference pivotal: for any prespecified linear contrast of state values, the statistic converges to a universal Brownian functional, with no long-run covariance estimator, bandwidth, or batch-length selector. At a fixed stepsize the inferential center is the Richardson–Romberg stationary target $\\theta_{RR,\\alpha}$ rather than the projected Bellman solution $\\theta_*$; a separate horizon-indexed design, with stepsize constant within each run and decreasing across runs, shifts the center to $\\theta_*$ under explicit rate conditions. A sympathetic reader cares because one-pass, memory-bounded confidence regions would make streaming policy evaluation practical without fragile tuning.","feed_headline":"No covariance estimator needed for TD confidence regions","feed_subtitle":"A Brownian-bridge statistic makes fixed-stepsize TD inference pivotal along a single Markov trajectory.","key_machinery":"The load-bearing object is the Brownian-bridge self-normalizer $\\hat V^{SN}_n = n^{-2}\\sum_{s=1}^n (S_s - s\\bar\\vartheta_n)(S_s - s\\bar\\vartheta_n)^\\top$: it is the sample analogue of the integrated squared Brownian bridge, and because the unknown long-run covariance $\\Omega$ multiplies both the endpoint and the bridge in the weak limit, it cancels from $T^{SN}_n = n(R\\bar\\vartheta_n - c)^\\top (R\\hat V^{SN}_n R^\\top)^{-1}(R\\bar\\vartheta_n - c)$. The proof chain that carries the claim is a sequence of reductions: Theorem 1's Lyapunov block contraction shows exponential forgetting uniform in the chain's starting state; Theorem 2's pullback stationary solution supplies the exact stationary law and the innovation $\\phi^\\alpha_t = \\xi_t - \\tilde A_t h^\\alpha_{t-1} - \\bar A \\bar h^\\alpha$; the augmented-chain Poisson equation (Proposition 3) splits that innovation into a martingale difference plus a telescoping remainder that determines $\\Omega_\\alpha$; and the horizon-indexed reduction (36) shows that inside the Assumption 4 rate window the RR partial sums converge to $\\bar A^{-1}$ times the additive Markov-noise partial sums. Proposition 10 then transfers any such FCLT to the same pivotal quadratic form without estimating the covariance.","core_discovery":"On the paper's own terms, the discovery is a distributional limit theorem plus a cancellation device. For a fixed stepsize $\\alpha$, the centered partial sums of the stationary TD recursion converge to the Brownian motion limit $n^{-1/2}\\sum_{t=1}^{\\lfloor nr \\rfloor}(\\theta^{\\alpha,\\circ}_t - \\theta^\\alpha) \\Rightarrow \\Omega_\\alpha^{1/2} W_d(r)$, where the innovation process $\\phi^\\alpha_t = \\xi_t - \\tilde A_t h^\\alpha_{t-1} - \\bar A \\bar h^\\alpha$ keeps the multiplicative term induced by the random TD matrix in the first-order path law. For $L$ Richardson–Romberg recursions driven by the same trajectory, the joint partial-sum process converges with cross-level covariance blocks, and the Brownian-bridge self-normalizer $\\hat V^{SN}_n$ cancels the unknown covariance in the quadratic form, giving the pivotal law $W_q(1)^\\top [\\int \\bar W_q \\bar W_q^\\top]^{-1} W_q(1)$. The paper claims this yields asymptotically valid confidence regions for $R\\theta_{RR,\\alpha}$ at fixed stepsize and, under the horizon-indexed rate conditions in Assumption 4, for $R\\theta_*$ as the residual RR target shift, multiplicative remainder, and initialization effect become root-$n$ negligible.","pith_inferences":["Beyond the paper, the same Brownian-bridge transfer should apply to any constant-stepsize stochastic-approximation recursion with a Hurwitz mean matrix and a stationary pullback solution, making the method a template for linear LSA inference rather than a TD-specific device.","The paper fixes the contrast $R$ before the run; tracking several contrasts simultaneously would amount to keeping one self-normalizer per contrast, though simultaneous coverage would need a stronger argument than pointwise asymptotic pivotality.","The lower endpoint of the rate window pushes toward higher RR cancellation order to reach $\\theta_*$ at practical horizons, but larger signed weights may inflate finite-sample variance; the experiments' pattern of slower stabilization in smaller-stepsize, higher-dimensional cells suggests the practical window may be narrower than the sufficient proof window."],"forward_implications":["At a fixed stepsize, confidence regions for the RR stationary target $R\\theta_{RR,\\alpha}$ have asymptotic coverage $1-\\eta$ with no long-run covariance estimator, bandwidth, or batch-length selection.","The same Brownian-bridge critical values cover the projected Bellman solution $R\\theta_*$ under the horizon-indexed rate window $\\alpha_n = n^{-\\nu}$ with $1/(2(q_{RR}+1)) < \\nu < \\min\\{1/2, 1-2/p\\}$.","The one-pass recursion keeps memory independent of the trajectory length, with $O(Ld+qd+q^2)$ persistent storage for a dense contrast and a reported $0.180$ MB online state at $d=10$, $n=10^6$.","First-order RR with nodes $(1,2)$ and weights $(2,-1)$ cuts the target shift from roughly $1.75\\times 10^{-2}$ to $4.9\\times 10^{-4}$ in the evaluated Garnet designs, with interval length essentially unchanged.","In the horizon-indexed regime the leading stochastic term is additive Markov noise $\\xi_t = b_t - A_t\\theta_*$, so the long-run covariance reduces to $\\bar{A}^{-1}\\Sigma_\\xi \\bar{A}^{-\\top}$."],"supporting_citations":[{"why":"Defines the projected Bellman fixed-point equation and the linear TD parameter $\\theta_*$ that the horizon-indexed regime targets.","marker":"Tsitsiklis and Van Roy [1997]"},{"why":"Supplies the same-trajectory Richardson–Romberg construction and the constant-stepsize Markovian LSA inference baselines that the paper extends.","marker":"Huo et al. [2024]"},{"why":"Provides the finite-order stationary-bias expansion behind the RR bias bound (21) used to control the residual target shift.","marker":"Huo et al. [2026]"},{"why":"Gives the Markovian random-matrix stability result and the moment bound used in the horizon-indexed analysis.","marker":"Durmus et al. [2021]"},{"why":"Introduces the Brownian-bridge self-normalizer whose covariance cancellation is the basis of the pivotal statistic.","marker":"Kiefer et al. [2000]"},{"why":"Develops self-normalized confidence intervals for time series from bridge functionals, the template for $\\hat V^{SN}_n$.","marker":"Shao [2010]"},{"why":"Provides the random-scaling online inference construction that motivates the one-pass memory-bounded implementation.","marker":"Lee et al. [2022]"},{"why":"Establishes the FCLT and random-scaling inference under independent sampling that the paper adapts to Markovian sampling and RR combinations.","marker":"Xie and Zhang [2022]"},{"why":"Develops constant-stepsize path theory and RR extrapolation for SGD whose decomposition informs the same-trajectory RR FCLT.","marker":"Li et al. [2024]"}],"fun_headline_variants":["Self-normalized TD confidence, no covariance estimate","Brownian bridge pivots fixed-stepsize TD inference","One-pass TD confidence with constant memory","Covariance-free pivotal regions for TD learning","No tuning parameters for TD confidence sets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the claim that the TD recursion forgets its starting point exponentially fast, uniformly over every starting state of the Markov chain and every small fixed stepsize; if that claim fails, the stationary solution and all downstream confidence regions collapse.","fun_headline_variants_meta":{"raw":{"variants":["Self-normalized TD confidence, no covariance estimate","Brownian bridge pivots fixed-stepsize TD inference","One-pass TD confidence with constant memory","Covariance-free pivotal regions for TD learning","No tuning parameters for TD confidence sets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000705,"raw_usage":{"total_tokens":3236,"prompt_tokens":1063,"completion_tokens":2173,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":2105}},"tokens_in":679,"tokens_out":2173,"duration_ms":15094,"temperature":1.0,"reasoning_tokens":2105,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:47:04.817401+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Numerically evaluate the uniform block-contraction inequality (11) on a slow-mixing finite-state Markov chain across many starting states $y$ and admissible stepsizes $\\alpha$; if any pair violates the stated exponential decay, the stationary-recursion construction and therefore the self-normalized coverage claim collapse.","supporting_citations":[],"review_version":1}