REVIEW 3 major objections 8 minor 35 references
Approximate Risk Minimization Over Shrinking-Thresholding Rules in Normal Mean Estimation
T0 review · 3 major / 8 minor · reviewed 2026-07-08 · glm-5.2
Pith's one-line read One risk criterion unifies shrinkage and thresholding
desk verdict Unified shrinkage-thresholding framework via approximate risk minimization; consistency theory rests on an unverified Tweedie accuracy assumption that may fail in sparse regimes read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The GEST class (equation 6) parameterized by (t, c); the approximate risk criterion (Proposition 2) replacing unknown theta with Tweedie plug-in z+l'(z); the optimality conditions (Theorem 4.2) yielding implicit equations for t and c; the sieve-based consistency proof (Theorem 6.1); the CMLE fixed-point construction for correlated settings (equation 25); the inverse power-thresholding operator inducing the penalized regression representation (Theorem 9.1).
What would settle it
If the Tweedie plug-in z + l'(z) fails to approximate theta on average — for instance, if the empirical signal distribution G_d does not converge, or if the score estimator is not uniformly consistent — then the approximate risk criterion diverges from the oracle risk, and Theorem 6.1's consistency claim becomes vacuous. One could construct a counterexample with an oscillating or non-convergent signal distribution where the score estimate is systematically biased.
Extended reading notes
Core claim
The key structural finding is that the shape of the empirical marginal score function l'(z) determines where on the shrinkage-to-thresholding spectrum the optimal estimator lies. When the marginal is Gaussian, the score is linear, the optimality conditions force c=0, and NOMAD reduces to James-Stein-type pure shrinkage. When the marginal has Laplace-like tails, the score magnitude is approximately constant, the optimality conditions force c=1, and NOMAD reduces to lasso soft-thresholding. For intermediate marginal shapes, the optimizer selects 0<c<1, producing a hybrid that shrinks large coordinates mildly and thresholds small ones. This means the dense-versus-sparse regime choice, which is
Load-bearing premise
The entire approximate risk criterion rests on the oracle Tweedie accuracy condition (Assumption C, equation 18): that the empirical Bayes posterior mean z + l'(z) approximates the unknown theta well enough on average across coordinates. If this plug-in is not accurate — which depends on the signal distribution converging and the score estimator being uniformly consistent — the approximate risk may not track the true oracle risk, and the consistency guarantees become empty.
Editorial extensions
If this is right
- Practitioners could use NOMAD as a single procedure that automatically adapts between dense-signal shrinkage and sparse-signal thresholding without pre-specifying a sparsity level or choosing between James-Stein and lasso.
- The penalized regression representation with data-adaptive penalty shape (controlled by c) suggests a family of penalties interpolating between ridge (quadratic) and lasso (absolute value), with the shape selected by the data rather than tuned externally.
- Wavelet denoising could benefit from level-dependent NOMAD, since coarse-scale coefficients (often dense) would trigger shrinkage-dominant behavior while fine-scale coefficients (often sparse) would trigger thresholding-dominant behavior, all within one unified procedure.
- The correlated extension via conditional MLE provides a principled way to threshold conditionally-adjusted coordinates, which could reduce spurious cross-talk in high-dimensional regression with correlated designs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript develops a unified approximate-risk-minimization framework for shrinkage-thresholding estimation in normal mean problems. The author introduces the GEST (general shrinkage-thresholding) class, which nests the MLE, James-Stein, ridge, and lasso as special cases through functional parameters t(|z|) and c(|z|). The oracle quadratic risk is expressed as a functional over this class via Stein's identity, and a feasible approximate risk criterion is constructed by replacing the unknown theta with the Tweedie posterior mean z + l'(z), where l'(z) is the marginal score estimated nonparametrically. The resulting estimator, NOMAD, is defined as the minimizer of this approximate risk over a spline sieve. The paper establishes oracle-risk consistency (Theorems 6.1, 6.2, 8.1, 8.2) under regularity conditions, shows that classical estimators arise as special cases under specific score structures (Section 5), extends the framework to correlated normal means via MLE-based and conditional-MLE-based constructions (Section 8), derives an equivalent penalized regression representation with a data-adaptive penalty (Section 9), and provides simulation studies across canonical, correlated, and regression settings (Section 10).
Significance. The paper's ambition is commendable: it provides a single risk-based framework that unifies shrinkage (James-Stein/ridge) and thresholding (lasso/soft-thresholding) as special cases of a two-parameter functional family, with the balance selected data-adaptively via the empirical marginal score. The recovery of classical estimators from specific score shapes (Section 5) is an elegant conceptual contribution. The penalized regression representation (Theorem 9.1) with a data-adaptive penalty shape is a potentially interesting bridge between normal mean theory and regression. The simulation studies are reasonably comprehensive. However, the significance is substantially diminished by the fact that all proofs are deferred to an unavailable supplement, making independent verification of the central theoretical claims impossible at this stage. The framework also introduces several named entities (GEST, NOMAD, power thresholding operator) whose novelty relative to existing shrinkage-thresholding families in the literature (e.g., Antoniadis-Fan, Johnstone-Silverman) is not clearly articulated.
major comments (3)
- §6.1, Assumption C, Eq. (18): The oracle Tweedie accuracy condition requires (1/d) sum_i E[(theta_i - z_i - l'_d(z_i))^2] -> 0. This is the load-bearing assumption connecting the approximate risk to the oracle risk. The manuscript does not verify this condition in any signal regime, nor does it discuss when it may fail. In sparse regimes (e.g., spike-and-slab with small pi), the marginal f_d is close to N(0,1), the score l'_d(z) is near zero for most z, and z + l'_d(z) approximates z itself, which has risk d. The condition (18) may fail precisely in the sparse regime where the method claims to adapt to lasso-type behavior. The paper should either (a) provide sufficient conditions on G_d under which (18) holds, (b) verify it for specific signal classes (e.g., sparse sequences with bounded sparsity), or (c) discuss the regime where it fails and the implications for the consistency theory.
- All proofs of the main results (Propositions 1-2, Theorems 4.1-4.2, 6.1-6.2, 8.1-8.4, 9.1) are deferred to Supplement [22], which is not available in this submission. Without the supplement, the correctness of the risk functional in Proposition 1, the optimality conditions in Theorems 4.1-4.2, and the consistency arguments in Theorems 6.1 and 8.1-8.2 cannot be verified. This is a fundamental barrier to assessment. The supplement must be provided for the review to be complete.
- §3.1, Proposition 1: The risk functional H involves terms with |z_i|^{c(|z|)} in the denominator and factors of |z_i|^{2-2c(|z|)}. When c(|z|) > 0 and |z_i| -> 0, these terms may diverge. The integrability conditions are mentioned ('the integrability conditions in Proposition 1 hold') but never stated explicitly. Since the GEST class includes c=1 (lasso), the behavior near |z_i|=0 needs careful treatment. The paper should state the integrability conditions explicitly and verify that they are satisfied for the special cases in Section 5.
minor comments (8)
- §2.2, Eq. (6): The notation |ZZZ| for the vector of absolute values is introduced but the Hadamard product symbol is used without explicit definition until later. Clarification would help.
- §4.2, Eqs. (12)-(13): The optimality conditions involve sums over i of terms depending on the full vector through c(|z|), but c is treated as a function of the summary statistics s_c(z) in Section 6.1. The relationship between the functional form in Section 4 and the sieve parametrization in Section 6 should be clarified.
- §5.2, Proposition 4: The claim that the NOMAD estimator reduces to James-Stein under Gaussian marginals requires the empirical-Bayes substitution of tau^{-2} by (d-2)/||z||^2. This substitution is a separate step from the approximate-risk minimization. The proposition should clarify that the reduction is to the shrinkage form, with the JS-specific constant arising from a post-processing step.
- §8.3: The SURE formula for the CMLE-based construction involves the matrix Psi whose definition uses indicator functions for nonzero components. The differentiability of the positive-part operator at zero is a standard subtlety; the paper should note how this is handled.
- §9.3, Theorem 9.1, Eq. (31): The penalty involves an integral with upper limit (X'X)^{1/2}_{ii} beta_i / sigma, which couples the penalty to the design through the diagonal of the precision matrix. This differs from standard coordinate-separable penalties. The remark should clarify whether this design-dependence is essential or an artifact of the reduction.
- §10: No code or reproducible computational details are provided. Given the complexity of the NOMAD procedure (score estimation, Newton iteration for c, fixed-point iteration for CMLE), code availability would strengthen the submission.
- The MSC2020 subject classifications (00X00) and keywords are placeholder codes. These should be filled in with appropriate classifications.
- §7.3, Figure 2: The y-axis scales differ across the four panels, making cross-signal comparison difficult. Consider using a common scale or clearly labeling the different scales.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive report. The referee raises three major comments: (1) the oracle Tweedie accuracy condition (Assumption C, Eq. 18) is not verified in any signal regime and may fail in sparse settings; (2) all proofs are deferred to an unavailable supplement; and (3) the integrability conditions in Proposition 1 are not stated explicitly and the behavior near |z_i|=0 needs careful treatment. We address each point below.
read point-by-point responses
-
Referee: §6.1, Assumption C, Eq. (18): The oracle Tweedie accuracy condition requires (1/d) sum_i E[(theta_i - z_i - l'_d(z_i))^2] -> 0. This is the load-bearing assumption connecting the approximate risk to the oracle risk. The manuscript does not verify this condition in any signal regime, nor does it discuss when it may fail. In sparse regimes (e.g., spike-and-slab with small pi), the marginal f_d is close to N(0,1), the score l'_d(z) is near zero for most z, and z + l'_d(z) approximates z itself, which has risk d. The condition (18) may fail precisely in the sparse regime where the method claims to adapt to lasso-type behavior. The paper should either (a) provide sufficient conditions on G_d under which (18) holds, (b) verify it for specific signal classes (e.g., sparse sequences with bounded sparsity), or (c) discuss the regime where it fails and the implications for the consistency theory.
Authors: The referee raises an important point about the load-bearing nature of Assumption C, Eq. (18). We agree that the manuscript does not currently provide sufficient conditions or verification for specific signal classes, and that the sparse regime warrants careful discussion. We address the referee's three suggested remedies in turn. (a) Sufficient conditions: In the revised manuscript, we will add a new subsection (or a dedicated remark in Section 6.1) providing sufficient conditions on G_d under which (18) holds. Specifically, if G_d converges to G in Wasserstein distance and the marginal density f(z) = integral phi(z-theta) dG(theta) has a score l'(z) satisfying a Lipschitz condition and moderate growth, then (18) follows from standard empirical Bayes posterior mean consistency results (cf. Efron, 2011; Brown and Greenshtein, 2009). The key requirement is that G_d is not degenerate at zero in the limit, i.e., the limiting signal distribution G has positive variance or at least is not a point mass at zero. (b) Verification for specific signal classes: We will verify (18) explicitly for two signal classes. First, for dense signals where G_d converges to a non-degenerate distribution (e.g., N(0, tau^2) with tau > 0), the Tweedie posterior mean z + l'_d(z) is a consistent estimator of theta_i in the sense of (18) under standard density estimation rates. Second, for sparse signals with bounded sparsity pi_d -> pi in (0,1), the marginal is a two-component mixture and the score is well-defined and estimable; (18) holds provided the non-null component has bounded second moment and the density estimator achieves the uniform consistency rate stated in Assumption C. (c) The degenerate sparse regime: The referee correctly identifies a genuine limitation. If pi_d -> 0 (vanishing sp) revision: no
Circularity Check
No significant circularity: the derivation is a plug-in approximation with externally verifiable assumptions, not a self-definitional chain.
full rationale
The paper's central derivation chain is structurally sound and not circular. The approximate risk criterion (Proposition 2) replaces the unknown θ in the oracle risk (Proposition 1) with the Tweedie plug-in z + l'(z), where l'(z) is the score of the marginal density estimated from data. This is a standard plug-in approximation, not a circular definition: the oracle risk depends on θ, the approximate risk depends on l'(z), and these are distinct objects. The optimality conditions (Theorem 4.2) are derived by calculus of variations on the approximate risk functional and yield solutions for (t, c) that depend on the empirical score l'(z), not on the risk being minimized. The consistency result (Theorem 6.1) requires Assumption C (eq. 18), which is an explicit regularity condition on the accuracy of the Tweedie posterior mean — it is stated as an assumption, not derived from the paper's own results. The reductions to MLE, James-Stein, and lasso (Propositions 3-5) follow by substituting specific score functions into the optimality conditions and checking that the solutions match known estimators; these are algebraic verifications, not fitted-input-as-prediction patterns. The self-citation [22] is to the author's own supplement for deferred proofs, which is standard practice; the supplement is not invoked to forbid alternatives or to smuggle in an ansatz. The penalized regression representation (Theorem 9.1) is derived from the inverse of the shrinkage-thresholding operator, not from assuming a penalty form. No step in the chain reduces to its inputs by construction. The concern about whether Assumption C holds in sparse regimes is a correctness risk for the assumptions, not a circularity in the derivation itself. The paper is self-contained against external benchmarks (simulations compare against MLE, JS+, lasso, ridge, SURE, VisuShrink, SureShrink), and the theoretical results are stated under explicit, falsifiable conditions rather than being forced by self-citation chains or definitional equivalence. Score 1 reflects one minor self-citation to the supplement for proofs, which is not load-bearing for the logical structure of the arguments as presented in the main text.
Assumptions & free parameters
free parameters (4)
- t(|z|) =
data-driven via eq. 12
- c(|z|) =
data-driven via eq. 13/14
- Spline basis dimensions K_{t,d}, K_{c,d} =
must satisfy m_d/d -> 0
- Score estimator bandwidth/type =
not specified
assumptions (8)
- domain assumption Gaussian noise model: Z ~ N(theta, I_d) or N(theta, Omega^{-1})
- standard math Quadratic loss for risk evaluation
- standard math Tweedie's formula (Lemma 3.1): E(theta|z) = z + l'(z)
- domain assumption Oracle Tweedie accuracy (eq. 18): (1/d) sum E[(theta_i - z_i - l'_d(z_i))^2] -> 0
- domain assumption Sub-exponential signal tails (eq. 17)
- domain assumption Precision matrix eigenvalues bounded away from 0 and infinity
- domain assumption CMLE fixed-point map is a contraction (||A||_op < 1)
- ad hoc to paper Identifiability condition for functional consistency
invented entities (3)
-
GEST class (General shrinkage-thresholding estimator)
independent evidence
-
NOMAD estimator
independent evidence
-
Power thresholding operator P_{t,c} and its inverse
independent evidence
Cite this review
Pith. "Pith review of Approximate Risk Minimization Over Shrinking-Thresholding Rules in Normal Mean Estimation." pith.science (2026). https://pith.science/paper/QYRS2SDD
@misc{pith2026260706367,
author = {Pith},
title = {Pith review of: Approximate Risk Minimization Over Shrinking-Thresholding Rules in Normal Mean Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/QYRS2SDD}},
note = {Machine review of arXiv:2607.06367}
}
read the original abstract
We develop an approximate risk minimization framework for shrinkage-thresholding estimation in normal mean problems. In the canonical multivariate normal mean model, we introduce a general functional class of estimators that contains classical shrinkage and thresholding behavior, including James-Stein-type and lasso-type rules. We express quadratic risk as a functional over this class, derive optimality conditions for both oracle risk and data-driven approximate risk minimization, and construct a feasible approximate risk criterion from the observed data when the oracle risk is unavailable. The resulting estimator, NOMAD, is obtained by minimizing this approximate risk over the proposed class. For the canonical model, we develop an approximate risk minimization theory that includes optimizer characterization, sieve-based consistency under regularity conditions, and approximate-risk inequalities relative to benchmark procedures in the admissible class. We then extend the framework to multivariate normal mean estimation with correlated observations, develop both MLE-based and conditional MLE-based constructions, and establish consistency results under regularity conditions. We further apply the framework to linear regression and derive an equivalent penalized regression representation in which the shrinkage-thresholding map induces a data-adaptive penalty, recovering ridge-type and lasso-type behavior as special cases or limiting forms. The results provide a unified risk-based framework for shrinkage, thresholding, and regularization across canonical and correlated normal mean estimation and linear regression.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[22]
Approximate Risk Minimization over Shrinkage–Thresholding Rules in Normal Mean Estimation
JIANG, W. Supplement to "Approximate Risk Minimization over Shrinkage–Thresholding Rules in Normal Mean Estimation"
-
[1]
ABRAMOVICH, F., BENJAMINI, Y., DONOHO, D. L. and JOHNSTONE, I. M. (2006). Adapting to Unknown Sparsity by Controlling the False Discovery Rate.The Annals of Statistics34584–653. https://doi.org/ 10.1214/009053606000000074
-
[2]
ALDRICH, J. (1997). RA Fisher and the making of maximum likelihood 1912-1922.Statistical Science12 162–176
work page 1997
-
[3]
ANTONIADIS, A. and FAN, J. (2001). Regularization of Wavelet Approximations.Journal of the American Statistical Association96939–967. https://doi.org/10.1198/016214501753208942
-
[4]
AUDIBERT, J.-Y. and CATONI, O. (2011). Robust Linear Least Squares Regression.The Annals of Statistics 392766–2794. https://doi.org/10.1214/11-AOS918
-
[5]
AVELLA-MEDINA, M. (2017). Influence Functions for Penalized M-Estimators.Bernoulli233178–3196. https://doi.org/10.3150/16-BEJ841
-
[6]
BARANCHIK, A. J. (1964).Multiple regression and estimation of the mean of a multivariate normal distri- bution51. Department of Statistics, Stanford University Stanford, CA
work page 1964
-
[7]
BERGER, J. O. (1976). Admissible Minimax Estimation of a Multivariate Normal Mean with Arbitrary Quadratic Loss.The Annals of Statistics4223–226. https://doi.org/10.1214/aos/1176343356
Show all 35 references
-
[8]
BESAG, J. (1974). Spatial Interaction and the Statistical Analysis of Lattice Systems.Journal of the Royal Statistical Society. Series B (Methodological)36192–236. APPROXIMATE RISK MINIMIZATION IN NORMAL MEAN ESTIMATION25
1974
-
[9]
T., LIU, W
CAI, T. T., LIU, W. and LUO, X. (2011). A Constrainedℓ 1 Minimization Approach to Sparse Precision Matrix Estimation.Journal of the American Statistical Association106594–607. https://doi.org/10. 1198/jasa.2011.tm10155
2011
-
[10]
DONOHO, D. L. and JOHNSTONE, I. M. (1994). Ideal Spatial Adaptation by Wavelet Shrinkage.Biometrika 81425–455. https://doi.org/10.1093/biomet/81.3.425
1994 doi
-
[11]
DONOHO, D. L. and JOHNSTONE, I. M. (1995). Adapting to Unknown Smoothness via Wavelet Shrinkage. Journal of the American Statistical Association901200–1224. https://doi.org/10.1080/01621459. 1995.10476626
1995 doi
-
[12]
EFRON, B. (2011). Tweedie’s formula and selection bias.Journal of the American Statistical Association 1061602–1614
2011
-
[13]
and MORRIS, C
EFRON, B. and MORRIS, C. (1973). Stein’s Estimation Rule and Its Competitors—An Empirical Bayes Approach.Journal of the American Statistical Association68117–130. https://doi.org/10.1080/ 01621459.1973.10481350
1973
-
[14]
ELDAR, Y. C. (2009). Generalized SURE for Exponential Families: Applications to Regularization.IEEE Transactions on Signal Processing57471–481. https://doi.org/10.1109/TSP.2008.2008212
2009 doi
-
[15]
and LI, R
FAN, J. and LI, R. (2001). Variable Selection via Nonconcave Penalized Likelihood and Its Oracle Properties.Journal of the American Statistical Association961348–1360. https://doi.org/10.1198/ 016214501753382273
2001
-
[16]
FRANK, I. E. and FRIEDMAN, J. H. (1993). A Statistical View of Some Chemometrics Regression Tools. Technometrics35109–148. https://doi.org/10.1080/00401706.1993.10485033
1993 doi
-
[17]
FU, W. J. (1998). Penalized Regressions: The Bridge Versus the Lasso.Journal of Computational and Graphical Statistics7397–416. https://doi.org/10.1080/10618600.1998.10474784
1998 doi
-
[18]
GEORGE, E. I. and MCCULLOCH, R. E. (1993). Variable Selection via Gibbs Sampling.Journal of the American Statistical Association88881–889. https://doi.org/10.1080/01621459.1993.10476353
1993 doi
-
[19]
HANSEN, B. E. (2016). The risk of James–Stein and Lasso shrinkage.Econometric Reviews351456–1470
2016
-
[20]
HOERL, A. E. and KENNARD, R. W. (1970). Ridge Regression: Biased Estimation for Nonorthogonal Problems.Technometrics1255–67. https://doi.org/10.1080/00401706.1970.10488634
1970 doi
-
[21]
and STEIN, C
JAMES, W. and STEIN, C. (1961). Estimation with quadratic loss.Proc. Fourth Berkeley Symp. Math. Statist. Prob.1361–379
1961
-
[23]
JOHNSTONE, I. M. and SILVERMAN, B. W. (2005). Empirical Bayes Selection of Wavelet Thresholds.The Annals of Statistics331700–1752. https://doi.org/10.1214/009053605000000345
2005 doi
-
[24]
D., SUN, D
LEE, J. D., SUN, D. L., SUN, Y. and TAYLOR, J. E. (2016). Exact Post-Selection Inference, with Applica- tion to the Lasso.The Annals of Statistics44907–927. https://doi.org/10.1214/15-AOS1371
2016 doi
-
[25]
MALLAT, S. (1989). A Theory for Multiresolution Signal Decomposition: The Wavelet Representation. IEEE Transactions on Pattern Analysis and Machine Intelligence11674–693. https://doi.org/10.1109/ 34.192463
1989
-
[26]
and BÜHLMANN, P
MEINSHAUSEN, N. and BÜHLMANN, P. (2006). High-Dimensional Graphs and Variable Selection with the Lasso.The Annals of Statistics341436–1462. https://doi.org/10.1214/009053606000000281
2006 doi
-
[27]
MITCHELL, T. J. and BEAUCHAMP, J. J. (1988). Bayesian Variable Selection in Linear Regression.Jour- nal of the American Statistical Association831023–1032. https://doi.org/10.1080/01621459.1988. 10478694
1988 doi
-
[28]
and CASELLA, G
PARK, T. and CASELLA, G. (2008). The Bayesian Lasso.Journal of the American Statistical Association 103681–686. https://doi.org/10.1198/016214508000000337
2008 doi
-
[29]
ROBBINS, H. E. (1992). An empirical Bayes approach to statistics. InBreakthroughs in Statistics: Founda- tions and basic theory388–394. Springer
1992
-
[30]
STEIN, C. (1956). Inadmissibility of the Usual Estimator for the Mean of a Multivariate Normal Distribu- tion. InProceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability1 197–206. University of California Press, Berkeley, CA
1956
-
[31]
STEIN, C. M. (1981). Estimation of the mean of a multivariate normal distribution.The Annals of Statistics 1135–1151
1981
-
[32]
and CANDÈS, E
SU, W., BOGDAN, M. and CANDÈS, E. J. (2017). False Discoveries Occur Early on the Lasso Path.The Annals of Statistics452133–2150. https://doi.org/10.1214/16-AOS1521
2017 doi
-
[33]
TIBSHIRANI, R. (1996). Regression shrinkage and selection via the lasso.Journal of the Royal Statistical Society Series B: Statistical Methodology58267–288
1996
-
[34]
and QU, A
WANG, L., ZHOU, J. and QU, A. (2012). Penalized Generalized Estimating Equations for High- Dimensional Longitudinal Data Analysis.Biometrics68353–360. https://doi.org/10.1111/j. 1541-0420.2011.01678.x
2012 doi
-
[35]
and ROEDER, K
WASSERMAN, L. and ROEDER, K. (2009). High-Dimensional Variable Selection.The Annals of Statistics 372178–2201. https://doi.org/10.1214/08-AOS646
2009 doi
Reviewed July 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.