Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

Regularized Generalized Covariance (RGCov) Estimator

T0 review · 3 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper proposes a ridge-regularized Generalized Covariance estimator that stays consistent and asymptotically normal in high dimensions, and recovers the GCov efficiency bound when the shrinkage decays to zero.

desk verdict Useful ridge regularization for GCov with good simulations, but the asymptotic theory is missing a rate condition on δ_T and the identifiability claim is unverified for the paper's own designs. read the letter →

arxiv 2504.18678 v1 pith:N6YE4NI3 submitted 2025-04-25 econ.EM

classification econ.EM MSC 62F1262M1062P20
keywords high-dimensionaltimeseriesGeneralizedCovarianceestimatorridgeregularizationmixedcausal-noncausalVARsemiparametricefficiencynonlinearserialdependencetestportmanteauRNLSD
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the Generalized Covariance (GCov) estimator, a semiparametric method for mixed causal–noncausal vector autoregressions, can be made to work when the covariance matrix it has to invert is high-dimensional. The proposed fix replaces the sample variance matrix $\hat\Gamma_T(0;\theta)$ by the ridge-regularized matrix $\delta_T I + \hat\Gamma_T(0;\theta)$ inside the objective function. The resulting RGCov estimator is consistent and asymptotically normal, and when the shrinkage coefficient decays to zero with the sample size it recovers the GCov limiting distribution and its semiparametric efficiency. The same regularization extends the GCov specification test and the nonlinear serial dependence (NLSD) test, with chi-square limits when $\delta_T\to 0$. Simulations show gains over GCov and diagonal GCov, and an application to green-energy stock prices illustrates the method on a 12-variable system.

What carries the argument

The load-bearing object is the regularized variance matrix $\hat\Gamma_T(0;\theta,\delta)=\delta_T I_K+\hat\Gamma_T(0;\theta)$, whose inverse appears in each term of the objective $L_T(\theta,\delta_T)=\sum_{h=1}^H \operatorname{Tr}[\hat\Gamma_T(h;\theta)\hat\Gamma_T(0;\theta,\delta_T)^{-1}\hat\Gamma_T(h;\theta)'\hat\Gamma_T(0;\theta,\delta_T)^{-1}]$. The ridge term inflates the smallest sample eigenvalues and keeps the inverse numerically stable when $K=Jn$ is large, which is what fails for the unregularized GCov. The schedule $\delta_T\to\delta\ge 0$ controls the limiting variance: fixed $\delta$ yields a sandwich-type covariance matrix, while $\delta_T\to 0$ collapses it to the efficient GCov information matrix. A Sherman–Morrison recursion updates the inverse observation by observation, avoiding repeated high-dimensional inversions in the numerical optimization.

What would settle it

Fit the RGCov estimator to a mixed causal–noncausal VAR using only transformations that are functionally dependent, such as $u$ and $u^2$ with a known relation, and check whether the minimizer is unique across starting values and whether the RNLSD test keeps its nominal size; non-uniqueness or size distortion would show that identifiability, not the ridge term, is carrying the asymptotic claims.

Watch

Extended reading notes

Core claim

The central claim is that the regularized estimator $\hat\theta_T(\delta_T)$ is consistent for $\theta_0$ under standard regularity conditions, and that $\sqrt{T}(\hat\theta_T-\theta_0)$ converges in distribution to $N(0,J(\theta_0,\delta)^{-1}I(\theta_0,\delta)J(\theta_0,\delta)^{-1})$ when $\delta_T\to\delta\ge 0$. When $\delta_T\to 0$, this distribution simplifies to $N(0,J(\theta_0)^{-1})$, the same limit as the unregularized GCov estimator, so the regularized version is asymptotically semiparametrically efficient. The paper also proves that the residual-based RGCov specification test and the regularized NLSD test have asymptotic chi-square distributions with $K^2H-\dim(\theta)$ degrees of freedom when $\delta_T\to 0$, and weighted sums of chi-squares when $\delta$ is fixed. The regularization is confined to the weighting matrix in the objective; the VAR coefficients themselves are not shrunk.

Load-bearing premise

The load-bearing premise is that $\theta_0$ is uniquely identified by the finite restrictions $\Gamma(h;\theta)=0$ for $h=1,\ldots,H$, an assumption the paper states but does not prove for the transformation sets it actually uses; if those transforms fail to pin down $\theta_0$, neither consistency nor either chi-square limit follows no matter how $\delta$ is chosen.

Editorial extensions

If this is right

  • High-dimensional GCov estimation becomes feasible without imposing sparsity on the VAR coefficients, because the only matrix that needs to be inverted is made regular by the ridge term.
  • If $\delta_T$ is chosen to vanish with $T$, inference from RGCov is asymptotically equivalent to inference from GCov, so the regularization can be used as a computational stabilizer without sacrificing semiparametric efficiency.
  • The RGCov residual-based specification test and the RNLSD test extend portmanteau testing to cases with many variables or many nonlinear transformations, with known chi-square degrees of freedom when the shrinkage decays.
  • For fixed $\delta>0$, the null distribution is a weighted sum of chi-squares whose weights are products of eigenvalues of $\Gamma(0)^{-1/2}\Gamma(0,\delta)\Gamma(0)^{-1/2}$, giving a principled testing procedure when the shrinkage is not allowed to vanish.
  • In the empirical application, RGCov identifies a mixed causal–noncausal VAR for twelve green-energy stock series where GCov yields near-zero eigenvalues of $\hat\Gamma(0)$, and the estimated causal and noncausal components are used to build portfolios that outperform the index in cumulative returns.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper documents a bias–variance trade-off in $\delta$ but stops short of a data-driven selector; cross-validating $\delta$ on the objective or on test size is a natural next step that the simulations already support.
  • The statement that $J\ge pn$ transformations are needed for identifiability sits awkwardly with the paper's own $n=15$, $J=2$ simulation, where identification succeeds; a fair reading is that identification depends on the informational content of the transformations, not just their count, and a formal condition would strengthen the result.
  • Because the regularization targets the weighting matrix rather than the coefficients, the method does not produce a sparse VAR; combining RGCov with coefficient shrinkage is an extension the paper does not explore.
  • For moderate samples with fixed $\delta$, a bootstrap calibrated to the weighted chi-square mixture could give better size control than the asymptotic approximation; this is a testable refinement not covered in the paper.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper proposes a regularized version of the Generalized Covariance (GCov) estimator, called RGCov, which replaces the inverse of the sample variance matrix Γ̂_T(0;θ) in the GCov objective by (δI + Γ̂_T(0;θ))^{-1}. The authors claim that, for a deterministic sequence δ_T → δ ≥ 0, the RGCov estimator is consistent and asymptotically normal with a covariance matrix they provide, and that when δ_T → 0 it achieves the same limiting distribution as the GCov estimator and thus semiparametric efficiency. They also extend the GCov specification test and the NLSD test to the regularized setting, obtaining a weighted chi-square limit for δ_T → δ>0 and a chi-square limit for δ_T → 0. Finite-sample performance is examined in two simulation designs (high n and high J), and the method is applied to a mixed causal-noncausal VAR for 12 green energy stock prices.

Significance. If the theoretical claims are correct, the paper offers a practical, computationally feasible way to estimate mixed causal-noncausal VARs in high dimensions, with the Sherman-Morrison updates and the diagonal estimator providing useful computational alternatives. The simulation evidence of strong variance reduction relative to the unregularized GCov is striking. However, the identification condition underlying all consistency results is not verified, and the proofs are presented as sketches that rely heavily on prior work. The paper is therefore a potentially useful contribution, but its central theorems need to be placed on firmer ground.

major comments (3)
  1. [Section 3 / Assumption A.1(iii)] The identifiability condition A.1(iii), which requires that Γ(h;θ)=0 for h=1,...,H has θ0 as its unique solution, is never verified for the transformation sets used in Section 4.1 (J=2), Section 4.2 (J=10) or Section 5 (J=4). The only supporting statement, 'it is necessary to use nonlinear transformations J≥pn to ensure the identifiability of the model' (Section 3), is contradicted by the paper's own simulation in Section 4.1, where n=15, p=1 and J=2 yet identification is reported as successful. Because Propositions 1, 2 and 4 all invoke A.1(iii), this is a load-bearing gap.
  2. [Appendix A, step d)] The expansion of the first-order conditions asserts that the term √T ∂²L_T(θ0,δ)/∂θ∂δ′ (δ_T−δ) is negligible whenever δ_T→δ. This requires √T ∂²L_T/∂θ∂δ′ to be stochastically bounded; the manuscript neither proves this nor states primitive conditions. If this derivative is not O_p(1) after scaling, a rate condition on δ_T is needed. The proof of Proposition 1(ii) and the δ_T→0 efficiency claim are therefore not fully established as written.
  3. [Section 3.2.2, Proposition 3] The proof of Proposition 3 is given only for the RNLSD case (dim θ=0). The general case with dim θ>0, which justifies the degrees of freedom K²H−dimθ used in Proposition 4, is dismissed with 'the general case is similar' and no details are provided. Since the test statistics are a central claimed contribution, a rigorous proof or a precise citation for the estimation effect is needed.
minor comments (7)
  1. [Section 2.1, Eq. (2.2)] The displayed definition of v_t(θ) repeats the same blocks twice; it should list a_j[g_i(ỹ_t;θ)] for i=1,...,n and j=1,...,J once.
  2. [Section 3.1, Definition 1] In (3.2), the shrinkage coefficient in R̂_T^2 is written as δ, but the definition uses δ_T; the argument should be made consistent.
  3. [Section 4.2, Table 4] The caption of Table 4 says 'fourteen inside and one outside' but the DGP in Section 4.2 has two eigenvalues inside and one outside the unit circle.
  4. [Section 4.1, text after Table 1] The text refers to 'Table??' when discussing identification frequencies; this should be Table 2.
  5. [Section 4.1, simulation design] In the δ_T=η/T setting, the largest η values (e.g., η=800 with T=800) imply δ_T=1, so the simulations do not actually explore the δ_T→0 regime of Proposition 2; the finite-sample evidence is therefore primarily about fixed (or slowly varying) δ.
  6. [Section 5] The selection of δ=0.3 via 'a higher distance' of eigenvalues from unity is ad hoc and not grounded in the theoretical results; a data-driven selection rule would be preferable.
  7. [Global] The index name is spelled both 'Rennix' and 'Renixx' at different places; please make it consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the δ→0 case is an explicitly acknowledged algebraic special case of GCov, while the fixed-δ limit theorems are derived directly from the regularized objective.

full rationale

The RGCov estimator is defined by Eqs. (3.1)-(3.3) as the minimizer of a δ-regularized modification of the GCov objective. Proposition 1's consistency proof uses the standard Jennrich argument: the objective converges uniformly to L(θ,δ)=Σ_h Tr[Γ(h;θ)Γ(0;θ,δ)^{-1}Γ(h;θ)'Γ(0;θ,δ)^{-1}], which is zero at θ0 and, by Assumption A.1(iii), has no other zero, so the minimizer converges to θ0. The asymptotic normality result is obtained from a first-order expansion of the first-order conditions and the classical CLT for sample autocovariances of i.i.d. vectors (Chitturi 1976; Hannan 1976), giving the matrices J(θ0,δ) and I(θ0,δ) explicitly. Proposition 3 derives the mixture-of-chi-square limit by writing √T vec Γ̂T(h) as X(h)~N(0,Γ(0)⊗Γ(0)) and diagonalizing Γ(0;δ)^{-1}⊗Γ(0;δ)^{-1}; this is a direct quadratic-form calculation rather than an imported distributional conclusion. When δ=0, Γ(0;θ,0)=Γ(0;θ), so the RGCov objective, estimator, and test statistics coincide algebraically with the GCov versions; the paper states this as a special case ('This includes as a special case δT=δ>0' and 'in the special case δ=0... simplified') rather than relabeling an input as a prediction. The efficiency statement at δ→0 cites Gourieroux and Jasiak (2023), but that citation is to published, parameter-free GCov theory whose assumptions do not include the RGCov result; it is prior evidence, not a circular input. Similarly, Jasiak and Neyazi (2023) is cited for the NLSD test, but the RNLSD limit is derived in Proposition 3 and only reduces to the NLSD chi-square when δ=0, again an algebraic limit. The unverified identifiability condition A.1(iii), and the apparent tension between Section 3's 'J≥pn' statement and Section 4.1's J=2 design, are correctness risks rather than circular reductions, since every stated theorem is conditional on A.1(iii). No equation in the paper reduces a claimed output to a fitted parameter or to a self-citation by construction.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The estimator is a ridge modification of a published estimator, so it inherits most of its assumptions from Gourieroux and Jasiak (2023). The paper adds one tuning parameter delta (or path parameter eta) and an unstated rate condition. No new entities are introduced. The most fragile input is identifiability from a finite set of transformations.

free parameters (2)
  • Shrinkage coefficient delta = delta = 0.3 in application; grid 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0 in simulations
    Added to the diagonal of Gamma(0); all finite-sample results and the asymptotic variance in the fixed-delta case depend on it. The paper discusses cross-validation but derives no formal selection rule.
  • Eta in delta_T = eta/T = eta from 100 to 800 in simulations
    Parameterizes the shrinkage path in the delta_T to 0 case; finite-sample performance varies with eta, and the theory does not specify how to choose it.
assumptions (5)
  • domain assumption Asymptotic identifiability: Gamma(h;theta)=0 for h=1,...,H implies theta=theta_0 (Assumption A.1(iii))
    Needed for consistency and for the null hypothesis of the specification test. Not verified for the transformations used; the paper's dimension statement J>=pn is inconsistent with its own simulations.
  • domain assumption Regularity conditions: compact parameter space, geometric ergodicity, continuous invariant distribution, twice differentiability, stochastic Lipschitz equicontinuity (Assumptions A.1-A.2)
    Standard M-estimator conditions imported from the GCov literature; not checked for transformations such as log|u| and |u|^(1/2) used in simulations.
  • domain assumption Known semiparametric efficiency of the GCov estimator (Gourieroux and Jasiak 2023)
    The delta_T to 0 efficiency claim is inherited from this self-cited published result; accepted as prior literature rather than re-derived.
  • ad hoc to paper Unstated rate condition: sqrt(T)(delta_T - delta) -> 0 for asymptotic normality
    The FOC expansion in Appendix A requires this to drop the second derivative term, but the paper only assumes delta_T -> delta.
  • standard math Limiting Gaussianity of sqrt(T) vec Gamma_hat_T(h) under serial independence (Chitturi 1976, Hannan 1976)
    Used in the proof of Proposition 3 for the test distribution.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Regularized Generalized Covariance (RGCov) Estimator." pith.science (2026). https://pith.science/paper/N6YE4NI3

@misc{pith2026250418678,
  author       = {Pith},
  title        = {Pith review of: Regularized Generalized Covariance (RGCov) Estimator},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N6YE4NI3}},
  note         = {Machine review of arXiv:2504.18678}
}
read the original abstract

We introduce a regularized Generalized Covariance (RGCov) estimator as an extension of the GCov estimator to high dimensional setting that results either from high-dimensional data or a large number of nonlinear transformations used in the objective function. The approach relies on a ridge-type regularization for high-dimensional matrix inversion in the objective function of the GCov. The RGCov estimator is consistent and asymptotically normally distributed. We provide the conditions under which it can reach semiparametric efficiency and discuss the selection of the optimal regularization parameter. We also examine the diagonal GCov estimator, which simplifies the computation of the objective function. The GCov-based specification test, and the test for nonlinear serial dependence (NLSD) are extended to the regularized RGCov specification and RNLSD tests with asymptotic Chi-square distributions. Simulation studies show that the RGCov estimator and the regularized tests perform well in the high dimensional setting. We apply the RGCov to estimate the mixed causal and noncausal VAR model of stock prices of green energy companies.

Figures

Figures reproduced from arXiv: 2504.18678 by the authors.

Figure 1
Figure 1. Bias, variance, and MSE of the (R)GCov estimator (first row) and the diagonal GCov estimator [PITH_FULL_IMAGE:figures/full_fig_p016_1.png] view at source ↗
Figure 2
Figure 2. The graph shows the bias, variance, and MSE of the RGCov estimator under the setting [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 4
Figure 4. The graph shows the bias, variance, and MSE of the RGCov estimator under the setting [PITH_FULL_IMAGE:figures/full_fig_p020_4.png] view at source ↗
Figures from the paper (2 more)
Figure 6
Figure 6. Figure 6: Residuals (a) GCov (b) RGCov The representation theorem introduced by Gourieroux and Jasiak (2017) for mixed pro￾cesses distinguishes between their purely causal and noncausal latent components. For the mixed VAR(1) model with a diagonalizable autoregressive coefficien…
Figure 7
Figure 7. Figure 7: Causal and noncausal components of VAR(1) estimated by RGCov [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nonfundamentalness or missing information ? Evidence from causal-noncausal VARs in macro-finance

    econ.EM 2026-07 conditional novelty 6.0 of 10

    In the Stock–Watson monetary VAR, noncausal roots mostly disappear after factor filtering, and filtered IRFs no longer show the price puzzle.

Reference graph

Works this paper leans on

32 extracted references · 30 canonical work pages · cited by 1 Pith paper

  1. [1]

    Bierens, H. J. (1990). A consistent conditional moment test of functional form. Econometrica: Journal of the Econometric Society\/ , 1443--1458

  2. [2]

    Breidt, F. J., R. A. Davis, K.-S. Lii, and M. Rosenblatt (1991). Maximum likelihood estimation for noncausal autoregressive processes. Journal of Multivariate Analysis\/ 36\/ (2), 175--198

  3. [3]

    Ho, and H

    Chan, K.-S., L.-H. Ho, and H. Tong (2006). A note on time-reversibility of multivariate linear processes. Biometrika\/ 93\/ (1), 221--227

  4. [4]

    Chitturi, R. V. (1976). Distribution of multivariate white noise autocorrelations. Journal of the American Statistical Association\/ 71\/ (353), 223--226

  5. [5]

    Sequential Monte Carlo for Noncausal Processes

    Cubadda, G., F. Giancaterini, and S. Grassi (2025). Sequential monte carlo for noncausal processes. arXiv preprint arXiv:2501.03945\/

  6. [6]

    Giancaterini, A

    Cubadda, G., F. Giancaterini, A. Hecq, and J. Jasiak (2024). Optimization of the generalized covariance estimator in noncausal processes. Statistics and Computing\/ 34\/ (4), 127

  7. [7]

    Cubadda, G. and A. Hecq (2011). Testing for common autocorrelation in data-rich environments. Journal of Forecasting\/ 30\/ (3), 325--335

  8. [8]

    Hecq, and S

    Cubadda, G., A. Hecq, and S. Telg (2019). Detecting co-movements in non-causal time series. Oxford Bulletin of Economics and Statistics\/ 81\/ (3), 697--715

Show all 32 references
  1. [9]

    Hecq, and E

    Cubadda, G., A. Hecq, and E. Voisin (2023). Detecting common bubbles in multivariate mixed causal--noncausal models. Econometrics\/ 11\/ (1), 9

  2. [10]

    Davis, R. A. and L. Song (2020). Noncausal vector ar processes with application to economic time series. Journal of Econometrics\/ 216\/ (1), 246--267

  3. [11]

    Engle, R. F. and C. W. Granger (1987). Co-integration and error correction: representation, estimation, and testing. Econometrica: journal of the Econometric Society\/ , 251--276

  4. [12]

    Engle, R. F. and S. Kozicki (1993). Testing for common features. Journal of Business & Economic Statistics\/ 11\/ (4), 369--380

  5. [13]

    and J.-M

    Fries, S. and J.-M. Zakoian (2019). Mixed causal-noncausal ar processes and the modelling of explosive bubbles. Econometric Theory\/ 35\/ (6), 1234--1270

  6. [14]

    Hencic, and J

    Gourieroux, C., A. Hencic, and J. Jasiak (2021). Forecast performance and bubble analysis in noncausal mar (1, 1) processes. Journal of Forecasting\/ 40\/ (2), 301--326

  7. [15]

    Gourieroux, C. and J. Jasiak (2016). Filtering, prediction and simulation methods for noncausal processes. Journal of Time Series Analysis\/ 37\/ (3), 405--430

  8. [16]

    Gourieroux, C. and J. Jasiak (2017). Noncausal vector autoregressive process: Representation, identification and semi-parametric estimation. Journal of Econometrics\/ 200\/ (1), 118--134

  9. [17]

    Gourieroux, C. and J. Jasiak (2022). Nonlinear forecasts and impulse responses for causal-noncausal (s) var models. arXiv preprint arXiv:2205.09922\/

  10. [18]

    Gourieroux, C. and J. Jasiak (2023). Generalized covariance estimator. Journal of Business & Economic Statistics\/ 41\/ (4), 1315--1327

  11. [19]

    Jasiak, and M

    Gourieroux, C., J. Jasiak, and M. Tong (2021). Convolution-based filtering and forecasting: An application to wti crude oil prices. Journal of Forecasting\/ 40\/ (7), 1230--1244

  12. [20]

    and J.-M

    Gouri \'e roux, C. and J.-M. Zako \"i an (2017). Local explosion modelling by non-causal process. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 79\/ (3), 737--756

  13. [21]

    Hall, M. K. and J. Jasiak (2024). Modelling common bubbles in cryptocurrency prices. Economic Modelling\/ 139 , 106782

  14. [22]

    Hannan, E. J. (1976). The asymptotic distribution of serial covariances. The Annals of Statistics\/ 4\/ (2), 396--399

  15. [23]

    Lieb, and S

    Hecq, A., L. Lieb, and S. Telg (2016). Identification of mixed causal-noncausal models in finite samples. Annals of Economics and Statistics/Annales d' \'E conomie et de Statistique\/ (123/124), 307--331

  16. [24]

    Hecq, A. and E. Voisin (2021). Forecasting bubbles with mixed causal-noncausal autoregressive models. Econometrics and Statistics\/ 20 , 29--45

  17. [25]

    Hencic, A. and C. Gouri \'e roux (2015). Noncausal autoregressive model in application to bitcoin/usd exchange rates. Econometrics of risk\/ 583 , 17--40

  18. [26]

    Jasiak, J. and A. M. Neyazi (2023). Gcov-based portmanteau test. arXiv preprint arXiv:2312.05373\/

  19. [27]

    Lanne, M. and P. Saikkonen (2011). Noncausal autoregressions for economic time series. Journal of Time Series Econometrics\/ 3\/ (3), Article 2

  20. [28]

    Lanne, M. and P. Saikkonen (2013). Noncausal vector autoregression. Econometric Theory\/ 29\/ (3), 447--481

  21. [29]

    Lof, M. and H. Nyberg (2017). Noncausality and the commodity currency hypothesis. Energy Economics\/ 65 , 424--433

  22. [30]

    Sherman, J. and W. J. Morrison (1949). Adjustment of an inverse matrix corresponding to a change in one element of a given matrix. In Annals of Mathematical Statistics , Volume 20, pp.\ 317--317

  23. [31]

    Sherman, J. and W. J. Morrison (1950). Adjustment of an inverse matrix corresponding to a change in one element of a given matrix. The Annals of Mathematical Statistics\/ 21\/ (1), 124--127

  24. [32]

    Swensen, A. R. (2022). On causal and non-causal cointegrated vector autoregressive time series. Journal of Time Series Analysis\/ 43\/ (2), 178--196

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.