{"id":"9c59fe85-0f73-40b3-8e99-249019fde131","arxiv_id":"2607.18128","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In exploratory equilibria, low-temperature policy susceptibility grows polynomially on causal chains but exponentially on positive feedback cycles, with Lambert-W critical temperature τ* = βT/W(βT√n).","lead":"This paper works out how small errors in rewards and dynamics get amplified as entropy-smoothed equilibrium policies in time-inconsistent control become more deterministic, showing that feedback loops can turn modest problems into exponentially sensitive ones. An RL theorist would care about the exact critical temperature τ* ≈ 2βT/log n, which sets how much smoothing is needed for root-n statistical estimation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exponential lower bound in Thm 3.4 rests on cone conditions (3.12)–(3.14) verified only for the constant-coefficient bounded-diffusion model; the e^{C/τ} amplification and 1/log n boundary are not established for general state-dependent EEHJB.","rationale":"The reader's weakest-assumption identification is exactly the load-bearing issue: Thm 3.4's lower bound is conditional on (3.12)–(3.14), and those conditions are verified only for the explicit bounded diffusion. My reading of the paper confirms that the authors are transparent about this—they explicitly state the aligned-sign requirement, demonstrate a negative-feedback counterexample, and reserve the e^{C/τ^2} versus e^{C/τ} question as open. However, the broad phrasing of the central claim ('positive feedback cycles produce exponential growth', 'closing one positive cycle changes the boundary to 1/log n') can easily be read as a general EEHJB phenomenon rather than a theorem about a specific model class plus an abstract conditional result. That is a scope risk, not an internal inconsistency: the explicit model is solved exactly, the algebra checks out, and the block-decomposition arguments are sound. Since the reader already assigned CONDITIONAL and flagged this assumption, my stress test does not change the verdict; it sharpens the reason the conditionality is needed and proposes a concrete computational test that would reveal whether the state-dependent gap is real or merely apparent. The missing reproducibility URL and commit hash is a practical weakness but secondary to the mathematical scope concern, and I do not weight it into the verdict change.","tokens_in":30766,"tokens_out":12713,"duration_ms":132201,"concrete_test":"Construct a one-dimensional state-dependent positive-feedback instance of the §4 EEHJB, e.g. take A=[−1,1], b(t,x,a)=a x, r=0, σ≡1, and a terminal reward F(x)=x; linearize about the tied equilibrium at ξ=0 and compute the susceptibility χ_τ(0) numerically for τ=2^{-k}, k=1,…,8 by solving the Volterra influence equation with the same right-endpoint discretization as §6 but a fine grid. Fit log∥χ_τ(0)∥ against 1/τ and against 1/τ^2. If the 1/τ slope is not linear with the predicted constant, or if the 1/τ^2 fit dominates, then the cone conditions (3.14) fail in a natural state-dependent model and the general e^{C/τ} claim does not transfer beyond the explicit state-independent realizations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative central claim—positive feedback cycles convert polynomial susceptibility into exponential e^{C/τ} growth and move the root-n boundary to order 1/log n—is proven at the abstract level only under the aligned-mode cone hypotheses (3.12)–(3.14). These hypotheses are not derived from the general well-posedness assumptions (4.1)–(4.2), and the paper verifies them only for the state-independent bounded-diffusion model in §5.2 (where the causal order is α=1 and the Perron vector ψ(t)=e^{νt}r supplies (3.14)). For the parabolic EEHJB of §4 the natural causal order is α=1/2, coming from the gradient bound ∥D_x P_0_{t,s}f∥ ≤ C_0(s−t)^{-1/2} in (4.11); no ψ, cone, or κ_- satisfying (3.12)–(3.14) is constructed in that setting. The scalar example with A=−κ∫ satisfies the order-one kernel bound used in the upper estimate but gives τ^{-1}e^{-κ(T−t)/τ}, so the cone/sign condition is essential, not cosmetic. Consequently the headline cycle-induced exponential amplification is a sharp theorem for an explicit model class plus a conditional abstract statement, not a demonstrated general law for state-dependent EEHJB. The manuscript itself lists the e^{C/τ^2} versus e^{C/τ} question as open, but the gap here is more specific: even e^{C/τ} lower bounds are not shown outside the exact models.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the sensitivity of entropy-regularized equilibria in time-inconsistent stochastic control to perturbations of the model parameters, with emphasis on low-temperature amplification. The core objects are a backward Volterra resolvent for the derivative of the equilibrium policy and a graph-theoretic decomposition of the causal influence operator. Under an abstract fractional-causal-influence assumption (Assumption 3.2), the paper proves an upper bound of Mittag-Leffler type; under additional cone/order hypotheses (3.12)-(3.14) it proves a matching lower bound. The block decomposition yields a dichotomy: on acyclic influence graphs the susceptibility is polynomial in 1/tau, whereas a positive cycle can produce exponential growth. The paper then proves fixed-temperature C^2 differentiability of the EEHJB equilibrium branch and a function-valued delta method, with an e^{C/tau^2} general stability certificate. Two explicit models - an affine two-action model and a bounded trigonometric diffusion - realize the rates, including a Lambert-W critical temperature for root-n reward noise and a discrete-time mesh-stiffness threshold N tau^2 -> infinity along a cyclic Perron mode. The numerical section reports reproducible experiments with deterministic verification.","tokens_in":31081,"tokens_out":24711,"duration_ms":188473,"significance":"If the results hold, they identify a new and concrete mechanism: the topology of causal feedback cycles, not just the 1/tau Gibbs factor, controls the conditioning of exploratory equilibria at low temperature. The path-cycle distinction is sharp and quantitatively explicit (Theta(tau^{-(L+1)}) vs e^{C/tau}), with a matching nonlinear selection phenomenon in a bounded uniformly elliptic model. The paper is unusually careful: the abstract theorems are stated with explicit hypotheses, the exact models are solved in closed form, the proofs are self-contained relative to stated assumptions, and the numerical code is archived and deterministic, with no fitted constants. The main limitation is that the exponential lower bound is conditional on cone conditions verified only for the exact models; this does not invalidate the exact-model results but restricts the generality of the advertised cycle-induced amplification.","major_comments":[{"comment":"The exponential lower bound (3.15) and the resolution boundary (Cor 3.5, Cor 3.7) are conditional on the cone/order conditions (3.12)–(3.14). These are assumed, not derived from the EEHJB hypotheses of Section 4; the only verification is for the constant-coefficient bounded-diffusion model in §5.2 (Perron cone with ψ(t)=e^{νt}r). For the state-dependent parabolic setting of §4 the natural causal order is α=1/2 from (4.11), and no ψ, cone, or κ_- is constructed. The scalar example following Thm 3.3 (A=-κ∫) shows that without the sign condition the response is damped, so the cone condition is essential. Consequently the paper establishes the e^{C/τ} amplification and the 1/log n boundary for an explicit model class plus a conditional abstract statement, not for general state-dependent EEHJB. Please either verify (3.12)–(3.14) for a nontrivial state-dependent class, or state this limitation","section":"§3, Thm 3.4–Cor 3.7; §4, Assump. 4.1–4.2; §5.2"},{"comment":"The abstract's sentence 'Closing one positive cycle changes the root-n linear-response boundary from a power law to order 1/log n' is not qualified by the aligned-mode/cone assumptions. Since §7 itself leaves open whether a state-dependent model attains even the e^{C/τ} lower rate (the general parabolic upper bound is e^{C/τ²}), the unqualified phrasing overstates the proven scope. Please amend the abstract and the analogous introduction sentence so that the exact models and the conditional abstract theorem are presented as the proven statements.","section":"Abstract; §7"}],"minor_comments":[{"comment":"The 'asymptotic solution' of the canonical balance contains an extra log α in the second-order term. Since only the {1+o(1)} form is claimed, please label it as a leading-order asymptotic solution rather than the exact solution of the displayed equation.","section":"§3, Eq. (3.18)"},{"comment":"The actionwise reward r(y,s,x,a) is not written explicitly; the integrand uses the policy-averaged m^π. State explicitly that r(y,s,x,a)=β(x−y)a, so that the averaged term in (5.3) is the mean reward under π.","section":"§5.1, Eq. (5.3)"},{"comment":"In the DAG part, the assumption N≥L appears only in the text after the proposition. Move it into the statement for clarity.","section":"§6, Prop. 6.1"},{"comment":"'EXPLORA TOR Y' contains a stray space in the title; please check the typesetting.","section":"Title page"},{"comment":"Specify that the 16.2% and 1.32% errors at n=1014 refer to τ_{n,1} and τ_{n,2} in (5.13).","section":"§6.1 (numerics)"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This is a technically strong manuscript with a clear exact-model core. My recommendation of major revision is driven by the gap between the unqualified abstract claims and the conditional status of the exponential lower bounds; I believe a scope revision plus either a small state-dependent verification or explicit limitation statement would make the paper acceptable. The supplementary material is substantial and the code reproducibility is a plus."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real contribution and deserves a serious referee. The new content is concrete: temperature-explicit susceptibility bounds through a backward Volterra-Mittag-Leffler resolvent, the path-cycle dichotomy, a function-valued delta method, and an exact Lambert-W transition scale. The derivations are self-contained relative to stated assumptions; I spot-checked the affine susceptibility, the sign in the coupled-mode exponential, and the Lambert-W boundary, and they check out. No fitted constants or circularity. The negative-feedback example shows the authors know the positivity conditions matter, which is a good sign. The main soft spot is exactly the one flagged in the stress test. The exponential lower bound and the 1/log n boundary require the cone/alignment conditions (3.12)-(3.14). Those are verified for the bounded-diffusion and affine models, but not derived from the general EEHJB hypotheses in Section 4. This is not a fatal flaw: the paper openly states the e^{C/tau^2} versus e^{C/tau} question as open, and the exact models are sharp. But it does mean the headline claim about positive feedback cycles is a theorem for a specific model class plus a conditional abstract statement, not a demonstrated general law for state-dependent EEHJB. The scalar negative-feedback example shows the sign condition is essential, not cosmetic. The second issue is smaller but real: the numerical section promises a reproducibility archive but gives no URL or commit hash, so the numerical claims cannot be independently checked from the preprint. That should be fixed before publication. Who is this for? Researchers working on entropy-regularized RL, time-inconsistent control, and stability of equilibria. A reader who wants a general parabolic lower bound will be disappointed; a reader who wants a careful, useful theory with explicit models will get a lot. I would send this to peer review and would expect the main theorems to survive, with revision needed to make the scope of the cone-condition results unambiguous. I would bring it to a reading group and would cite it if I worked in this area.","headline":"A serious, honestly hedged theory paper: real new results on temperature-explicit susceptibility for EEHJB, with the main soft spot being that exponential amplification is proven for exact models plus a conditional abstract theorem, not for general state-dependent EEHJB.","tokens_in":698,"tokens_out":1607,"would_cite":true,"duration_ms":30860,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93E20","49L20","68T05","60H30","65M12","62M05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Entropy-regularized equilibria become exponentially sensitive to model error when the control graph has a positive feedback cycle, changing the statistical resolution boundary from a power law to order 1/log n.","keywords":["time-inconsistent stochastic control","entropy regularization","exploratory equilibrium HJB equation","Volterra operator","Mittag-Leffler function","feedback cycles","statistical conditioning","Lambert W function"],"falsifier":"In the exact affine model with β=T=1 and ξ_n = Z/√n, integrate (5.5)–(5.7) at the Lambert-W scale τ*_n = βT/W(βT√n) for n ≥ 10^8; Theorem 5.2 predicts m(0) has a nondegenerate random limit with mean about 0.52, while m(t) for fixed t>0 converges to 0. A distribution converging to a point mass, or a nonvanishing later-time response, would refute it.","tokens_in":30548,"feed_emoji":"🔄","tokens_out":9261,"duration_ms":94713,"temperature":0.7,"pith_summary":"This paper asks how errors in learned rewards and dynamics propagate through entropy-regularized equilibria of time-inconsistent stochastic control, where each temporal self optimizes against the future policies of later selves. It establishes that the derivative of the equilibrium policy with respect to model parameters is governed by a backward Volterra resolvent whose low-temperature growth is decided by the causal graph of the system: on a directed acyclic graph the sensitivity is a power of 1/τ, while a single positive feedback cycle can produce exponential growth e^{C/τ}. If true, cooling an entropy-regularized system with feedback cycles is dramatically more fragile than the local Gibbs factor 1/τ would suggest, and the statistical boundary for root-n estimation drops from a power law to order 1/log n. The paper also proves that at fixed temperature the equilibrium branch is twice differentiable in finite-dimensional model parameters, yielding a function-valued delta method, and it constructs explicit models—an affine two-action model and a bounded uniformly elliptic diffusion—that realize both regimes exactly. The same machinery shows that along a cyclic Perron mode, right-endpoint time discretization is relatively consistent exactly when N τ^2 → ∞.","feed_headline":"Closing one cycle flips error scaling from power law to 1/log n","feed_subtitle":"A positive feedback cycle makes low-temperature control equilibria exponentially fragile: the root-n error boundary drops to 1/log n.","key_machinery":"The load-bearing object is the causal Volterra resolvent: the policy tangent solves u = τ^{-1}S(c + Ku), where S is the centered Gibbs covariance operator and K is the future-policy-to-current-score derivative. The influence kernel is measured by a fractional integral I^α_{T-} of order α ∈ (0,2], so the resolvent series is bounded by the Mittag-Leffler function E_α(κ(T-t)^α/τ). In block form, powers of A = SK count directed walks: an acyclic graph satisfies A^{L+1}=0, terminating the series after the longest path, while a positive cycle contributes the factor E_β(g(T-t)^β/τ^q). This same machinery powers the fixed-temperature delta method.","core_discovery":"The paper claims that the derivative of an entropy-regularized equilibrium policy with respect to model parameters is governed by a backward Volterra resolvent u = τ^{-1}S(c + Ku), and that its low-temperature size is set by the causal graph of the system. On a directed acyclic graph the response is exactly polynomial, Θ(τ^{-(L+1)}); a positive feedback cycle produces an exponential factor E_β(g(T-t)^β/τ^q) under an aligned cone condition. In the bounded uniformly elliptic model the susceptibility matrix is χ_τ(t) = (v/τ) exp{(-νI + vK/τ)(T-t)}, and the paper shows that closing one positive cycle changes the root-n linear-response boundary from a power law to order 1/log n, with an exact Lam","pith_inferences":["If the exponential amplification holds generally, then entropy-annealing in cyclic control environments will not achieve root-n accuracy: n^{-1/2} model errors, harmless on a DAG, leave order-one policy errors on a cycle. A practical corollary is to estimate the causal return structure before cooling.","The Lambert-W scale implies that for a cyclic system, achieving a fixed statistical precision at temperature τ requires a sample size exponential in βT/τ—a quantitative prediction testable in tabular or linear-quadratic experiments.","The lower bound's dependence on the invariant-cone condition (assumed, not derived) suggests that negative feedback can mask cycle amplification; characterising the largest class of state-dependent models in which the exponential rate is actually attained remains an open problem, and the paper's e^{C/τ^2} general bound hints the true rate may be even worse in some models."],"forward_implications":["On a directed acyclic graph of longest path L, the equilibrium susceptibility is Θ(τ^{-(L+1)}); along a positive cycle it is at least of exponential order, so closing one edge can change the small-temperature scaling.","The root-n statistical resolution boundary for learned-model noise is of order (log n)^{-α} in general and order 1/log n when a positive cycle with α=1 is present.","At fixed τ>0, the local equilibrium branch is twice differentiable in finite-dimensional model parameters, so function-valued delta-method inference for equilibrium policies is justified.","In the bounded diffusion model, right-endpoint time discretization along a cyclic Perron mode is relatively consistent exactly when Nτ^2→∞; for a DAG, N→∞ suffices without coupling to τ.","The affine model has an exact phase transition at τ*_n = βT/W(βT√n) ~ 2βT/log n, at which the initial policy has a nondegenerate random limit while every fixed later-time policy converges to the reference mixture."],"fun_headline_variants":["One feedback cycle flips error scaling to 1/log n","Closing a positive cycle makes root-n errors scale as 1/log n","Feedback cycle changes error boundary from power law to 1/log n","A single cycle closes and error scaling becomes 1/log n"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The exponential amplification claim rests on an assumed 'aligned positive mode'—a cone preserved by the feedback operator with a temperature-independent test path—which the paper verifies only for special bounded-diffusion models, so without that condition negative feedback could damp the response and no exponential growth need occur.","fun_headline_variants_meta":{"raw":{"variants":["One feedback cycle flips error scaling to 1/log n","Closing a positive cycle makes root-n errors scale as 1/log n","Feedback cycle changes error boundary from power law to 1/log n","A single cycle closes and error scaling becomes 1/log n"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001009,"raw_usage":{"total_tokens":4102,"prompt_tokens":745,"completion_tokens":3357,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":3283}},"tokens_in":489,"tokens_out":3357,"duration_ms":23803,"temperature":1.0,"reasoning_tokens":3283,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:56:21.147568+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the exact affine model with β=T=1 and ξ_n = Z/√n, integrate (5.5)–(5.7) at the Lambert-W scale τ*_n = βT/W(βT√n) for n ≥ 10^8; Theorem 5.2 predicts m(0) has a nondegenerate random limit with mean about 0.52, while m(t) for fixed t>0 converges to 0. A distribution converging to a point mass, or a nonvanishing later-time response, would refute it.","supporting_citations":[],"review_version":1}