{"id":"b2cbb0e2-a0b2-4afd-818e-4d93bd40a8c2","arxiv_id":"2608.05009","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Weak-form LASSO estimates Heston drift, diffusion, and leverage coefficients from price paths, with accurate diffusion and leverage recovery in simulation and proxy-sensitive negative leverage in market data.","lead":"This paper extends weak-form SINDy recovery to Heston stochastic-volatility models, estimating drift, diffusion, and leverage from a single price path. On simulated paths it recovers volatility-of-volatility and leverage with median errors under 2 percent, while market data give proxy-sensitive but mostly negative leverage estimates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The asymptotic theory is mis-normalized: with A_jk=ΣKΘΔt and B_j=ΣKΔv (Eqs. 18–20), the ergodic limits are (1/T)A→∫KΘπ and (1/T)B→∫Kbπ, not the (1/N) limits stated in Theorem 2.2; hence the centered statistic in Theorem 2.5 does not converge for Δt≠1.","rationale":"Good faith reading: the paper is carefully scoped, the synthetic design is informative, and the limitations section is candid about proxy attenuation, sign sensitivity, and the absence of formal inference. The central experimental claim—accurate recovery of ξ, ρ, and ρξ from 30 latent-state paths—is plausible and internally consistent with the in-fill/long-span distinction in Proposition 2.4. The load-bearing theoretical support for that claim, however, contains a normalization error. The reader flagged Theorem 2.5 only as a proof sketch and placed the weakest assumption on H2 library completeness; I agree that H2 is scope-limiting, but the more immediately checkable defect is the factor-of-Δt error in Theorems 2.2 and 2.5. This does not overturn the synthetic point estimates, which cancel the common Δt factor, so the reader's CONDITIONAL verdict stands; the condition should include correcting the theorem statement and either releasing the reproducibility package with the corrected derivation or computing empirical standard errors via simulation. The paper deserves credit for its explicit restrictions, its negative-control phase, and its honest reporting of proxy-dependent sign evidence; none of the concerns here impugn the authors' integrity.","tokens_in":16474,"tokens_out":18981,"duration_ms":200749,"concrete_test":"Re-derive Theorems 2.2 and 2.5 from the displayed definitions (Eqs. 18–20), replacing the 1/N normalization by 1/T, and verify that √T((1/T)B_j−∫K_j b_v dπ) has limiting variance (1/Δt)Σ_l Cov[K_j(v_0)Δv_0,K_j(v_l)Δv_l] rather than the V_j printed in Eq. 16. If the corrected derivation matches this form, the printed theorem needs a normalization fix; if it instead reproduces Eq. 16 with 1/N, the definitions in the paper are inconsistent and must be reconciled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equations 18–20 define A_jk = Σ_n K_j(v_n)Θ_k(v_n)Δt and B_j^(b) = Σ_n K_j(v_n)Δv_n. Under Assumption H1, (1/N)ΣKΘΔt→Δt∫KΘπ, and (1/N)ΣKΔv→Δt∫Kbπ. The paper states these limits without the Δt factor (Theorem 2.2, Eq. 12) and then builds Theorem 2.5 on that limit, asserting √T((1/N)B_j−\\bar B_j)→N(0,V_j) with V_j=ΣCov[K_j(v_0)Δv_0,K_j(v_l)Δv_l]. With Δt=1/252 (daily data), (1/N)B_j converges to 0.004∫Kbπ instead of ∫Kbπ, so the centered statistic diverges as √T; the variance formula also carries units of Δt and should appear divided by Δt under the correct 1/T normalization. Because the estimator is the ratio A^{-1}B, the Δt factor cancels and the Phase 1 point estimates may be unaffected, but the theorem as printed cannot justify the T^{-1/2} asymptotic-normality claim or Proposition 2.4's rate statement. This is a concrete defect in the paper's statistical foundation, not a reproach of the empirical recovery, and it strengthens the need for a corrected derivation or a real proof before the theory is relied on.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a weak-form Galerkin method for recovering the parameters of a Heston stochastic-volatility model from a single price path. The log-price and variance increments are projected onto Gaussian kernels in variance space, and the four targets—variance increments, squared variance increments, squared price increments, and the return–variance cross-product—are fitted against a shared library {1,v} with a LASSO regression, producing estimates of κ, θ, ξ, ρ, and the leverage slope ρξ. The paper reports synthetic Phase 1 results across 30 independent fine-step daily-observed simulations in which ξ, ρcorr, and ρξ are recovered with median errors under 2% and all seeds below 5%, while κ and θ are less accurate (median errors 12.29% and 6.61%). It then studies robustness to state noise, one-sided variance-proxy smoothing, the S&P 500 in 2007–2010, a 120-path leverage-regime screen, a nonlinear-drift falsification, and a 50-stock Indian panel, concluding that empirical leverage evidence is proxy-sensitive and that the method recovers coefficients within a prespecified Heston library rather than discovering general nonlinear structure.","tokens_in":16806,"tokens_out":14642,"duration_ms":145098,"significance":"If the empirical recovery results are correct, the paper makes a useful contribution: it demonstrates that a single path of daily observations can identify the diffusion and leverage structure of a Heston-type generator, which is otherwise an ill-posed inverse problem for Kramers–Moyal-type estimators. The strengths of the manuscript are its honest and careful experimental design: 30 independent seeds with full error distributions, a fine internal discretization for the synthetic truth, explicit reporting of the drift parameters' poorer recovery, a predeclared filter-selection rule, and an openly reported negative result (Phase 6) showing that LASSO support is unreliable for nonlinear-drift discovery. The reproducibility package with fixed seeds, pipeline settings, and CSV tables is a substantial asset. The main reservation concerns the asymptotic theory, which contains a normalization error that invalidates the printed consistency and CLT statements; the point estimates themselves are not necessarily affected because the Δt factor cancels in the ratio estimator.","major_comments":[{"comment":"Theorem 2.2 (Eq. (12)) states that (1/N)A_jk → ∫K_jΘ_kπ and (1/N)B^(b)_j → ∫K_j b_vπ as T→∞ with Δt fixed. However, from Eqs. (18)–(20), A_jk = Σ_n K_j(v_n)Θ_k(v_n)Δt and B^(b)_j = Σ_n K_j(v_n)Δv_n, so Birkhoff's ergodic theorem gives (1/N)A_jk → Δt∫K_jΘ_kπ and (1/N)B^(b)_j → Δt∫K_j b_vπ, because E[Δv_n|F_n] = b_v(v_n)Δt. The printed limits are therefore missing a factor of Δt. As a consequence, the centered statistic in Theorem 2.5, √T((1/N)B^(b)_j − \\bar B^(b)_j), does not converge to the stated normal distribution; with Δt = 1/252 it diverges unless \\bar B^(b)_j is redefined to include the Δt factor. The correct statements should use 1/T normalization, i.e., (1/T)A_jk → ∫K_jΘ_kπ and √T((1/T)B^(b)_j − \\bar B^(b)_j) → N(0,V_j) for a suitably re-derived V_j. Because the estimator is the ratio A^{-1}B, the Δt factor cancels in the point estimates and the Phase 1 numbers are not invalidated, but the theorem as printed cannot justify the T^{-1/2} asymptotic-normality claim or the rate statements in Proposition 2.4.","section":"§2.4, Theorem 2.2 and Theorem 2.5, Eqs. (12) and (16)"},{"comment":"The finite-step bias identities (13)–(15) are proved by substituting Euler–Maruyama increments (8)–(9) at the observation step Δt. However, the Phase 1 data-generating process is a fine-step Euler simulation (internal step 10^{-4} year) observed daily, and empirical market data are not Euler increments at all. For the true Heston transition, the conditional second moment of Δv_n contains additional O(Δt^2) terms beyond the drift-squared term; for example, the CIR marginal satisfies Var(v_{t+Δt}|v_t) = ξ^2v_tΔt + (ξ^2κ/2)(θ−3v_t)Δt^2 + O(Δt^3), which is not equal to the Euler variance ξ^2v_tΔt. Thus the asserted O(NΔt^2) floor in Theorem 2.1 and the drift-squared correction implemented in Section 2.6 do not exactly represent the bias for the object of inference. The paper should either state explicitly that the theorems concern the Euler-discretized generator (in which case a limit argument is needed to connect to the continuous-time Heston parameters) or derive the bias terms using the exact Heston conditional moments, which are available in closed form.","section":"§2.5, Theorem 2.3; §2.6"},{"comment":"The consistency and normality results are stated and proved for the ordinary least-squares estimator \\hat c = (A^T A)^{-1}A^T B (see the proof of Theorem 2.2), but the estimator actually used throughout the experiments is LASSO with five-fold cross-validated penalty selection and no fitted intercept (Table 1). No theorem accounts for the LASSO shrinkage, the data-dependent penalty, or the selection tolerance, so the asymptotic claims do not apply to the implemented estimator. The paper should either provide a LASSO-specific recovery guarantee (or a stability result under the stated assumptions) or explicitly relegate the theorems to a motivating OLS idealization and present the reported simulations as the evidence for the LASSO variant.","section":"§2.4, proof of Theorem 2.2; Table 1"}],"minor_comments":[{"comment":"The phrase 'one LASSO regression' is imprecise: the four targets are separate regression problems that share a common design matrix and are fitted in the same pipeline. Please rephrase to avoid the implication of a single multi-output regression.","section":"Abstract and §2.6"},{"comment":"The sentence 'The authors would like would like to acknowledge' contains a duplicated 'would like' that should be corrected.","section":"Acknowledgments"},{"comment":"Because 'selected' is defined as an absolute fitted coefficient exceeding the numerical tolerance 10^{-12}, the reported 93.3% false-positive rate is unsurprising for a continuous LASSO solution; consider also reporting a stability-selection or threshold-based selection rate to make the negative conclusion more informative.","section":"§3.6"},{"comment":"The statement that the last term 'contributes approximately 2 Var(η)' should specify that this is the expectation of (Δη_n)^2; the cross term 2Δv_nΔη_n has zero expectation under the stated independence but does not vanish pathwise.","section":"§2.8, Eq. (24)"},{"comment":"The two leverage normalizations are clearly explained, but the text should state explicitly that \\hat ρ_corr is the quantity used for the 'all seeds below 5%' claim in Phase 1, since Table 2 reports ρ_corr while the abstract says ρ.","section":"§2.6, Eqs. (22)–(23)"}],"recommendation":"major_revision","confidential_remarks":"The normalization error in Theorem 2.2/2.5 is a genuine, localized defect that must be corrected before the theory can be relied upon, but it is fixable, and the synthetic evidence for the main empirical claim is strong and transparently reported. The manuscript builds very heavily on the authors' own preprint [14], and the incremental contribution is an extension of that framework to a coupled bivariate system plus an extensive set of stress tests; I would ask the editor to verify that this constitutes sufficient novelty for the target journal. The Phase 6 falsification and the proxy-sensitive Indian results are unusual and valuable in their candor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead arXiv:2608.05009. The paper does something genuinely useful: it extends the scalar weak-form Galerkin/SINDy framework to the coupled Heston pair and shows, on 30-seed synthetic paths, that a single LASSO over four weak targets recovers xi, rho, and rho*xi with median errors under 2% when the latent variance is observed. The negative controls matter: Phase 6 shows the same pipeline cannot reliably discover a quadratic drift term, and Phase 7 shows the empirical leverage sign depends materially on the variance proxy. The authors are honest about scope. They explicitly say the results are coefficient recovery inside a prespecified Heston library, not structure discovery, and they display the Indian non-mean-reverting drift as a failure diagnostic rather than hiding it.\n\nThe most serious problem is in the asymptotic theory, and the stress-test note is correct. Theorem 2.2 states (1/N)A_jk -> ∫KΘπ, but A_jk is defined with a Δt factor, so the limit should be Δt∫KΘπ. Same for B: (1/N)ΣKΔv -> Δt∫Kbπ. The CLT in Theorem 2.5 centers at the wrong quantity, so as written the √T statistic diverges for Δt≠1. The ratio estimator A^{-1}B cancels the Δt, so the synthetic point estimates are probably fine, but the stated convergence rates and CLT do not follow from the given proof. This needs a corrected derivation or a real proof before anyone relies on the theoretical rates.\n\nOtherwise, the empirical work is careful. The 30-seed design with full error distributions is reproducible in principle, though the promised final package is not yet released with a commit hash. The filter-selection rule is hand-set, and the empirical estimates have no formal uncertainty quantification. These are minor relative to the theory issue.\n\nWho is this for? Quantitative finance researchers who need path-based physical-measure generator estimates, and the SINDy/weak-form community. It deserves a serious referee: the empirical contribution is useful and the theory can be fixed. I would send it to review, not desk-reject, but with the expectation that the normalization error be addressed and the proofs tightened.","headline":"Solid, honest extension of weak-form SINDy to the Heston model with a strong synthetic negative-control suite, but the asymptotic theorems have a normalization error that needs correction before the stated CLT is credible.","tokens_in":17370,"tokens_out":3605,"would_cite":false,"duration_ms":35313,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M05","62J07","62P05"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single LASSO regression on weak-form increments recovers the Heston generator, including leverage, from one price path.","keywords":["weak-form regression","stochastic SINDy","Heston model","leverage effect","LASSO","variance proxies","errors in variables","Galerkin projection"],"falsifier":"Replicate the Phase 1 experiment: simulate 30 daily-observed Heston paths with the paper's parameters and internal step, run the published pipeline, and check whether every seed keeps the relative errors of ξ, ρ, and ρξ below 5%; if any of the 30 seeds violates that bound, or if the median errors exceed the paper's reported values, the central accuracy claim fails. A second, theory-specific check: quadruple the simulation horizon to 400 years and verify that the median κ error falls by roughly half relative to the 100-year case, as the paper's 1/√T drift prediction requires.","tokens_in":16247,"feed_emoji":"📈","tokens_out":9056,"duration_ms":81349,"temperature":0.7,"pith_summary":"This paper asks whether the coupled drift, diffusion, and leverage structure of a Heston stochastic volatility model can be recovered directly from an observed price path. Its answer is a qualified yes: by projecting four increment targets—variance drift, squared variance, squared price, and their cross-product—onto Gaussian kernels in variance space, a single LASSO regression jointly estimates the mean-reversion rate, long-run variance, vol-of-vol, and leverage correlation. In 30 simulated daily-observed Heston paths, the diffusion and leverage parameters ξ, ρ, and ρξ come out with median errors under 2% and every seed under 5%, while the drift parameters κ and θ recover more slowly, as the paper's theory predicts. On S&P 500 data from the 2007–2010 crisis the method finds negative leverage, and a 50-stock Indian panel shows the sign is sensitive to the variance proxy. The value is a path-only estimator of the physical-measure generator that does not rely on option-implied surfaces.","feed_headline":"One LASSO fit recovers Heston leverage from a price path","feed_subtitle":"A single weak-form regression estimates vol-of-vol, leverage, and mean reversion without option quotes.","key_machinery":"The carrying mechanism is the coupled spatial weak-form projection: Gaussian kernels Kj(v)=exp(−(v−cj)^2/($2h^{2}$)) placed at 50 centers in variance space are multiplied by four increment targets—variance drift, squared variance, squared price, and the price–variance cross-product—and summed along the path. Because all Heston coefficients of interest are affine functions of v, the same design matrix A_jk = Σ_n K_j(v_n)Θ_k(v_n)Δt with library Θ(v)=[1,v] represents all four targets, so one LASSO regression with cross-validated penalty recovers the coefficients. The shared design exploits the triangular structure of Heston: the variance process is autonomous, so kernels need to be localized only in v. A two-step drift-informed correction subtracts the fitted drift-squared term from the squared-variance target before the diffusion regression, removing the O(Δt²) bias of Theorem 2.3.","core_discovery":"The central claim is that the Heston generator—mean reversion, long-run variance, volatility of variance, and the return–variance correlation that defines leverage—is identifiable from one discretely observed path by solving a single weak-form LASSO problem. The key novelty is the cross-variation target: squaring and cross-multiplying Euler–Maruyama increments, the conditional correlation ρ of the two Brownian shocks enters as a leading-order signal ρξvΔt in the cross-product of price and variance increments, so the fitted cross-diffusion slope is ρξ. Under the exact-Heston assumption (library [1,v] complete), the paper proves unbiasedness, strong consistency, and asymptotic normality of the coupled projection, and corrects a finite-step drift-squared bias in the variance-diffusion target. The simulation evidence is strong for diffusion and leverage: across 30 fine-step, daily-observed paths, every estimate of ξ, ρ, and ρξ stayed below 5% error, with median errors of 0.90%, 0.53%, and 1.80% respectively; κ and θ had median errors of 12.29% and 6.61%, matching the theory that drift estimation improves only with horizon length. Empirical results are more qualified: the S&P 500 fit yields negative leverage, but the sign rate varies widely across five variance proxies in the Indian panel, so the paper claims proxy-sensitive contemporaneous evidence, not robust identification.","pith_inferences":["The same shared-design construction should carry over to any triangular bivariate diffusion with an autonomous state coordinate and affine coefficients, so the approach may generalize beyond Heston to models like the double-CIR or other affine volatility specifications, though that extension is not tested here.","A natural stress test implied by the paper's own noise ablation: use realized variance from high-frequency data as the proxy and check whether the recovered ρξ approaches the latent-state accuracy of Phase 1; the paper's attenuation results predict it should come closer than range-based proxies do.","The null false-selection result suggests that weak-form SINDy studies should report selection frequencies under a null simulation; the paper's 93.3% rate is a concrete benchmark for any future library-expansion claim in this framework.","The S&P leverage estimate of about −0.33 under the Heston normalization could be compared with option-implied leverage from the same period; since the paper explicitly does not claim option-pricing calibration, agreement or disagreement between physical and risk-neutral leverage is left open."],"forward_implications":["A price path alone, with no option quotes, can supply physical-measure estimates of leverage and vol-of-vol, which are the inputs the Heston generator needs for scenario simulation and risk analysis.","The method's in-fill vs long-span split (diffusion and leverage improve with sampling frequency; drift improves with horizon) gives a practical rule: use high-frequency data for ρ and ξ, and long histories for κ and θ.","Because the variance state is unobserved in market data, the recovered coefficients are proxy-dependent summaries; the paper's Indian-panel sign variation is a direct warning that single-proxy leverage estimates should not be read as structural.","Within the exact-Heston library, the method achieves the stated accuracy, so it can serve as a fast calibration check against full likelihood or MCMC procedures for the physical measure.","The high false-selection rate for a quadratic drift term under the null means weak-form LASSO support cannot yet be used to discover new drift structure without a null-calibrated selection rule."],"supporting_citations":[{"why":"supplies the spatial weak-form projection framework that the paper extends to the coupled two-dimensional Heston system.","marker":"[14]"},{"why":"defines the Heston model whose generator the paper recovers.","marker":"[19]"},{"why":"introduces SINDy, the sparse-regression approach on which the LASSO library fitting is based.","marker":"[10]"},{"why":"provides the Garman–Klass volatility proxy used in the empirical state construction.","marker":"[16]"},{"why":"documents the bias and attenuation problems in leverage estimation that motivate the noise analyses.","marker":"[1]"},{"why":"provides the exact simulation scheme used to generate synthetic ground-truth trajectories for the recovery tests.","marker":"[9]"},{"why":"supplies the Euler–Maruyama discretization used for the observed increments and the pipeline simulation.","marker":"[22]"}],"fun_headline_variants":["Single weak-form LASSO extracts Heston leverage","Heston leverage recovered from one price path","Leverage from cross-increments: one LASSO fit","Path-only Heston parameters via weak-form regression","Joint recovery of Heston drift and leverage from path"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the true variance dynamics are exactly Heston-shaped—both drift and diffusion lie in the span of the library {1,v} and every covariance entry is affine in v—so that all four weak targets are representable by the same design matrix; if the real generator has any other term, the recovered coefficients are biased summaries of the generator rather than its true parameters.","fun_headline_variants_meta":{"raw":{"variants":["Single weak-form LASSO extracts Heston leverage","Heston leverage recovered from one price path","Leverage from cross-increments: one LASSO fit","Path-only Heston parameters via weak-form regression","Joint recovery of Heston drift and leverage from path"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1436,"prompt_tokens":1074,"completion_tokens":362,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":690,"completion_tokens_details":{"reasoning_tokens":287}},"tokens_in":690,"tokens_out":362,"duration_ms":3927,"temperature":1.0,"reasoning_tokens":287,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T17:01:16.037689+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replicate the Phase 1 experiment: simulate 30 daily-observed Heston paths with the paper's parameters and internal step, run the published pipeline, and check whether every seed keeps the relative errors of ξ, ρ, and ρξ below 5%; if any of the 30 seeds violates that bound, or if the median errors exceed the paper's reported values, the central accuracy claim fails. A second, theory-specific check: quadruple the simulation horizon to 400 years and verify that the median κ error falls by roughly half relative to the 100-year case, as the paper's 1/√T drift prediction requires.","supporting_citations":[{"cited_title":"Aït-Sahalia, J","cited_arxiv_id":null,"evidence_quote":"documents the bias and attenuation problems in leverage estimation that motivate the noise analyses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the Euler–Maruyama discretization used for the observed increments and the pipeline simulation."}],"review_version":2}