REVIEW 3 major objections 6 minor 1 cited by
Minimum Sliced Distance Estimation in a Class of Nonregular Econometric Models
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper proves that minimum sliced distance estimators are root-T asymptotically normal in structural econometric models with parameter-dependent supports, giving practitioners standard Wald inference where maximum likelihood yields…
desk verdict Sliced distance estimation for nonregular models is a good idea, but the printed asymptotic covariance is wrong for dψ>1 and needs a serious fix before the paper's central claim holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the weighted sliced L2 discrepancy between data and model: for each projection direction $u$, the distance integrates squared differences of cumulative distribution functions (sliced Cramér distance) or quantile functions (sliced Wasserstein distance) over $s$ with a weight $w(s)$, then averages over directions. The argument is carried by showing the criterion is quadratically approximated around $\psi_0$: $\hat{S}(\psi) \approx \text{constant} - 2(\psi-\psi_0)^\top A_T/\sqrt{T} + (\psi-\psi_0)^\top B_T(\psi-\psi_0)$, with the remainder controlled by a norm-differentiability condition. That condition permits a finite number of kinks in the model-induced function at the parameter-dependent boundary, which is exactly where likelihood-based theory fails, and the remainder analysis for conditional models routes through degenerate U-statistics.
What would settle it
Take the univariate one-sided uniform model $Y\sim U[0,\psi_0]$ with $\psi_0=2$ and the integrable but unbounded weight $w(s)=|s-2|^{-1/2}$ on $[0,3]$, normalized to integrate to one. Compute the minimum sliced Cramér distance estimator on many simulated samples of size $T=10{,}000$ and compare the empirical distribution of $\sqrt{T}(\hat\psi_T-\psi_0)$ to a normal distribution. Lemma 4.2's verification of Condition 4.2 uses the fact that integrals over a shrinking boundary-crossing interval shrink linearly in the interval length; for this weight the integral over $[\psi_0,\psi]$ is proportional to $\sqrt{\psi-\psi_0}$, so the $T$-scaled remainder term need not vanish. A non-normal sampling distribution or severely distorted Wald coverage for this weight would show that asymptotic normality does not hold for the full class of integrable weights.
Extended reading notes
Core claim
The paper's central discovery is Theorem 3.2: under Assumptions 2.1 and 3.1–3.6, any estimator that minimizes the weighted sliced L2 distance between the empirical measure of the data and the model-induced measure satisfies $\sqrt{T}(\hat{\psi}_T-\psi_0) \xrightarrow{d} N(0, B_0^{-1}\Omega_0 B_0^{-1})$, where $B_0$ is the integrated outer product of the derivative of the model-induced distribution and $\Omega_0$ is the asymptotic variance of a functional central limit term. Proposition 4.1 then verifies the high-level assumptions for the minimum sliced Cramér distance estimator in conditional models with one-sided parameter-dependent support, under primitive smoothness conditions on the conditional CDF and on the boundary function. The key point is that these conditions do not require the conditional density to be bounded away from zero at the boundary or to have a jump there, so the estimator is asymptotically normal whether or not the likelihood would have a non-normal limit.
Load-bearing premise
The argument depends on the model's distribution function being nearly quadratic in the parameter except at the one point where the support boundary moves, and on the weight function's mass near that boundary shrinking away fast enough; the paper only assumes the weight is integrable, which may not be enough when the weight is unbounded.
Editorial extensions
If this is right
- Wald-type confidence intervals and t-tests become available for structural parameters in one-sided and two-sided parameter-dependent support models, eliminating the dichotomy where likelihood-based inference must switch between normal and non-normal limit theory.
- The estimator is a one-step procedure: it does not require an auxiliary regression model or simulated samples from it, unlike indirect inference.
- Because the high-level theorem covers any sliced L2 distance satisfying the assumptions, both the sliced Cramér distance and the sliced Wasserstein distance variants inherit the same asymptotic normality.
- In the independent private-value procurement auction model, the normal approximation is accurate with 100–200 observations and comparable to indirect inference with a well-chosen starting value.
Reading between the lines
- A natural next question is whether the integrability assumption on $w$ can be strengthened to an explicit shrinking-mass condition; the boundary-interval argument in Lemma 4.2 suggests that unbounded weights may require it.
- The same quadratic-approximation proof should transfer to other non-smooth econometric structures, such as censored regressions, kinked regression functions, or threshold models, wherever the model-induced CDF has bounded second derivatives away from a low-dimensional boundary.
- The choice of projection directions is a finite-sample tuning issue the theory does not address: with many projections the estimator approximates the true sliced distance, but the auction simulation uses 100 directions and Adam optimization, so the practical recipe trails the theory.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes minimum sliced distance (MSD) estimation, covering minimum sliced Wasserstein distance (MSWD) and minimum sliced Cramér distance (MSCD), for structural econometric models with possibly parameter-dependent support. The main theoretical result, Theorem 3.2, establishes √T-consistency and asymptotic normality of the MSD estimator under high-level assumptions (Assumptions 3.1–3.6), following the quadratic approximation approach of Andrews (1999) and Pollard (1980). Proposition 4.1 then claims that for conditional MSCD models, primitive smoothness conditions (Conditions 4.1–4.2 and related differentiability assumptions) verify the high-level assumptions, so the estimator is consistent and asymptotically normal regardless of whether the conditional density jumps at the parameter-dependent support boundary. The paper verifies the conditions explicitly for one-sided and two-sided uniform models and presents a simulation study for an auction model.
Significance. If the claims are correct, the paper offers a genuine practical advance: a one-step estimator that avoids auxiliary models and yields standard Wald inference in nonregular structural models. The high-level theorem is clean and follows a well-established template; the exact norm-differentiability computations for the uniform examples are a useful check; and the auction simulation demonstrates good finite-sample normal approximation. However, the covariance formula in Proposition 4.1 (and its analogue in Theorem 3.2) is dimensionally wrong for dψ>1, and the displayed V0 block formula does not match the CLT summand in Lemma 4.6. Because the paper's advertised contribution is 'simple inference', these errors are central and must be corrected before the paper can be accepted.
major comments (3)
- [Theorem 3.2 and Proposition 4.1] The asymptotic covariance expression Ω0=(e1',−e1')V0(e1;−e1) with e1 a dψ-vector of ones is a scalar whenever dψ>1, whereas B0^{-1}Ω0B0^{-1} must be a dψ×dψ matrix. In the auction example dψ=2, so the displayed formula cannot deliver standard errors for the two components of ψ0. The correct selector should be a block matrix such as (I_{dψ}, −I_{dψ}), not a pair of vectors of ones. This affects the main inference claim, not just a special case.
- [Proposition 4.1, V0 display] The Kronecker factor in the displayed V0 is D(s;u,ψ0)D(t;u,ψ0)^⊤, attaching both derivative factors to the same projection u and placing the first projection's threshold in the second slot. The CLT summand in Lemma 4.6 has a first component indexed by u with threshold t and a second component indexed by v with threshold s, so the (1,2) covariance block must be ∫∫∫∫ A12(t,s;u,v) ⊗ D(t;u,ψ0)D(s;v,ψ0)^⊤ w(t)w(s) dt ds dς(u)dς(v). As printed, the covariance would not agree with the actual limiting distribution even after the e1 issue is fixed.
- [Lemma 4.2 and Appendix D.3] The proof of Lemma 4.2 bounds the boundary-crossing contribution to ∫|R_t|^2 w by C∥ψ−ψ0∥^2 |u1|(|g(X_i,ψ)−g(X_i,ψ0)|+2Cτ_T) and calls it O(τ_T^3). This implicitly requires that the integral of the weight w over an interval of length O(τ_T) is O(τ_T). The stated assumption is only that w is integrable; for unbounded integrable weights, the integral over a shrinking interval is o(1) but not necessarily O(τ_T). The verification of Condition 4.2 is therefore incomplete for the class of weights allowed in Section 4. The gap is repairable (one can use boundedness of x^2/(1+x)^2 and vanishing L1 mass), but the proof as written does not establish the stated claim.
minor comments (6)
- [Throughout] The paper mixes √n and √T (e.g., Theorem 3.2, Proposition 4.1); all statements should use √T consistently.
- [Section 2.2.1] There is a typo: 'Wassserstein' should be 'Wasserstein'.
- [Appendix D.3] The word 'seond-order' should be 'second-order'.
- [Lemma D.1] In the displayed expression for E[I(u⊤Z_t ≤ s)|X_t, ψ], the second case should be for u1 = 0, not u2 < 0, and the third case should be for u1 < 0.
- [Verification of Assumption 3.4(ii)] In the final displayed line, G(t;u,ψ0) should be G(s;u,ψ0).
- [References] The citation 'van de Vaart' should be 'van der Vaart'.
Circularity Check
No significant circularity: the asymptotic normality results are derived by verifying explicit high-level conditions rather than by fitting or by self-referential argument.
full rationale
The paper's central claims (Theorem 3.2 and Proposition 4.1) are asymptotic normality results for the proposed minimum sliced distance estimator. These are derived from explicit assumptions: a quadratic approximation of the sample objective (Assumption 3.3 / Condition 4.2), a stochastic equicontinuity-type remainder condition, a CLT on the gradient functional (Assumption 3.5, verified in Lemma 4.6 via standard empirical-process CLTs), and positive definiteness of the limit Hessian B0. Each component is verified from primitive smoothness conditions on the conditional CDF and boundary function, not from a fitted parameter or from the estimator's own limiting behavior assumed in advance. The high-level Assumption 3.5 does state a CLT, but that is a standard division of labor in extremum-estimation proofs; the paper then verifies it for the conditional model, so the conclusion is not identical to an input by construction. No constant is fitted to a subset of data and then renamed a prediction, and no auxiliary model is used to manufacture the claimed simple inference. The only self-reference is the footnote saying the paper is a revised version of part of Park's dissertation (Park [2022]); that citation is purely biographical and is not load-bearing evidence for any theorem. The verification of Condition 4.2 in Lemma 4.2 uses bounded second derivatives of the conditional CDF and Lipschitz boundary behavior; whether the stated conditions are strong enough is a correctness question, not a circularity question. Similarly, the dimensional and index issues in the displayed covariance matrix in Proposition 4.1 are internal consistency concerns, not circularity. Accordingly, no circular step is identified, and the paper is self-contained against external benchmarks for the purpose of this pass.
Assumptions & free parameters
assumptions (5)
- domain assumption Assumption 2.1: random sample from a parametric distribution or parametric conditional distribution.
- ad hoc to paper Condition 4.1: pointwise Lipschitz continuity of the conditional CDF F(y|x,psi) in psi with an integrable envelope M(y,x).
- ad hoc to paper Condition 4.2: population norm-differentiability with remainder bound T times the integral of (R(s;u,psi,psi0))^2 w over (1 + norm(sqrt(T)(psi-psi0)))^2 equals o(1).
- ad hoc to paper Integrability of the weight function w(s) and additional implicit control of w on small intervals.
- standard math Standard empirical process and degenerate U-statistic results: P-Donsker classes, Sherman 1994 Corollary 8, Newey 1991 Corollary 4.1, Briol et al. 2019 Lemma 4.
Cite this review
Pith. "Pith review of Minimum Sliced Distance Estimation in a Class of Nonregular Econometric Models." pith.science (2026). https://pith.science/paper/WI3HUDXZ
@misc{pith2026241205621,
author = {Pith},
title = {Pith review of: Minimum Sliced Distance Estimation in a Class of Nonregular Econometric Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/WI3HUDXZ}},
note = {Machine review of arXiv:2412.05621}
}
read the original abstract
This paper proposes minimum sliced distance estimation in structural econometric models with possibly parameter-dependent supports. In contrast to likelihood-based estimation, we show that under mild regularity conditions, the minimum sliced distance estimator is asymptotically normally distributed leading to simple inference regardless of the presence/absence of parameter dependent supports. We illustrate the performance of our estimator on an auction model.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
A sliced Wasserstein and diffusion approach to random coefficient models
A sliced-Wasserstein and k-nearest-neighbor minimum-distance estimator for the distribution of random coefficients β is consistent with polynomial-in-dimension computation, while its diffusion and causal extensions re...
Reference graph
Works this paper leans on
-
[1]
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savare. Gradient Flows . Birkh \" a user-Verlag, 2008. ISBN 9783764373092. doi:10.1007/b137080
doi:10.1007/b137080 2008
-
[2]
Estimation When a Parameter is on a Boundary
Donald W K Andrews. Estimation When a Parameter is on a Boundary . Econometrica, 67 0 (6): 0 1341--1383, 1999. ISSN 0012-9682. doi:10.1111/1468-0262.00082
-
[3]
On Parameter Estimation with the Wasserstein Distance
Espen Bernton, Pierre E Jacob, Mathieu Gerber, and Christian P Robert. On Parameter Estimation with the Wasserstein Distance . Information and Inference: A Journal of the IMA, 8 0 (4): 0 657--676, 2019. ISSN 2049-8764. doi:10.1093/imaiai/iaz003
-
[4]
One-Dimensional Empirical Measures, Order Statistics, and Kantorovich Transport Distances
Sergey Bobkov and Michel Ledoux. One-Dimensional Empirical Measures, Order Statistics, and Kantorovich Transport Distances . Memoirs of the American Mathematical Society, 261 0 (1259): 0 0, 2019. doi:10.1090/memo/1259
-
[5]
Sliced and Radon Wasserstein Barycenters of Measures
Nicolas Bonneel, Julien Rabin, Gabriel Peyr \' e , and Hanspeter Pfister. Sliced and Radon Wasserstein Barycenters of Measures . Journal of Mathematical Imaging and Vision, 51 0 (1), 1 2015. ISSN 0924-9907. doi:10.1007/s10851-014-0506-3
-
[6]
Audra J. Bowlus, Nicholas M. Kiefer, and George R. Neumann. Equilibrium search models and the transition from school to work . International Economic Review, 42 0 (2), 2001. ISSN 00206598. doi:10.1111/1468-2354.00112
-
[7]
Francois-Xavier Briol, Alessandro Barp, Andrew B. Duncan, and Mark Girolami. Statistical Inference for Generative Models with Maximum Mean Discrepancy . 6 2019
work page 2019
-
[8]
Likelihood Estimation and Inference in a Class of Nonregular Econometric Models
Victor Chernozhukov and Han Hong. Likelihood Estimation and Inference in a Class of Nonregular Econometric Models . Econometrica, 72 0 (5): 0 1445--1480, 2004. doi:10.1111/j.1468-0262.2004.00540.x
arXiv 2004
Show all 29 references
-
[9]
On the Composition of Elementary Errors: Second Paper: Statistical Applications
Harald Cram \' e r. On the Composition of Elementary Errors: Second Paper: Statistical Applications . Scandinavian Actuarial Journal, 1928 0 (1): 0 141--180, January 1928. ISSN 1651-2030. doi:10.1080/03461238.1928.10416872
1928
-
[10]
Quantile Processes with Statistical Applications
Mikl \' o s Cs \" o rg o . Quantile Processes with Statistical Applications . Society for Industrial and Applied Mathematics, 1 1983. ISBN 978-0-89871-185-1. doi:10.1137/1.9781611970289
1983 doi
-
[11]
Generative Modeling Using the Sliced Wasserstein Distance
Ishan Deshpande, Ziyu Zhang, and Alexander Schwing. Generative Modeling Using the Sliced Wasserstein Distance . In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2018. doi:10.1109/CVPR.2018.00367
2018
-
[12]
Superconsistent Estimation and Inference in Structural Econometric Models Using Extreme Order Statistics
Stephen G Donald and Harry J Paarsch. Superconsistent Estimation and Inference in Structural Econometric Models Using Extreme Order Statistics . Journal of Econometrics, 109 0 (2): 0 305--340, 2002. doi:10.1016/s0304-4076(02)00116-1
2002 doi
-
[13]
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets . In Z Ghahramani, M Welling, C Cortes, N Lawrence, and K Q Weinberger, editors, Advances in Neural Information Processi...
2014
-
[14]
Asymptotic Efficiency in Parametric Structural Models with Parameter-Dependent Support
Keisuke Hirano and Jack R Porter. Asymptotic Efficiency in Parametric Structural Models with Parameter-Dependent Support . Econometrica, 71 0 (5): 0 1307--1338, 2003. doi:10.1111/1468-0262.00451
2003
-
[15]
An Adversarial Approach to Structural Estimation
Tetsuya Kaji, Elena Manresa, and Guillaume Pouliot. An Adversarial Approach to Structural Estimation . SSRN Electronic Journal, 2020. ISSN 1556-5068. doi:10.2139/ssrn.3706365
2020 doi
-
[16]
Indirect Inference in Structural Econometric Models
Tong Li. Indirect Inference in Structural Econometric Models . Journal of Econometrics, 157 0 (1): 0 120--128, 7 2010. ISSN 03044076. doi:10.1016/j.jeconom.2009.10.027
2010 doi
-
[17]
Statistical and Topological Properties of Sliced Probability Divergences
Kimia Nadjahi, Alain Durmus, Lénaïc Chizat, Soheil Kolouri, Shahin Shahrampour, and Umut S im s ekli. Statistical and Topological Properties of Sliced Probability Divergences . 3 2020 a
2020
-
[18]
Asymptotic Guarantees for Learning Generative Models with the Sliced-Wasserstein Distance
Kimia Nadjahi, Alain Durmus, Umut S im s ekli, and Roland Badeau. Asymptotic Guarantees for Learning Generative Models with the Sliced-Wasserstein Distance . 6 2020 b
2020
-
[19]
Whitney K. Newey. Uniform Convergence in Probability and Stochastic Equicontinuity . Econometrica, 59 0 (4), 7 1991. ISSN 00129682. doi:10.2307/2938179
1991 doi
-
[20]
Deciding between the Common and Private Value Paradigms in Empirical Models of Auctions
Harry J Paarsch. Deciding between the Common and Private Value Paradigms in Empirical Models of Auctions . Journal of Econometrics, 51 0 (1-2): 0 191--215, 10 1992. doi:10.1016/0304-4076(92)90035-p
1992 doi
-
[21]
Non-Likelihood Based Methods for the Estimation and Inference in Econometric Models , 2022
Hyeonseok Park. Non-Likelihood Based Methods for the Estimation and Inference in Econometric Models , 2022
2022
-
[22]
The Minimum Distance Method of Testing
D Pollard. The Minimum Distance Method of Testing . Metrika, 27 0 (1): 0 43--70, 1980. ISSN 0026-1335. doi:10.1007/bf01893576
1980 doi
-
[23]
Optimal Transport for Applied Mathematicians , volume 87 of Progress in Nonlinear Differential Equations and Their Applications
Filippo Santambrogio. Optimal Transport for Applied Mathematicians , volume 87 of Progress in Nonlinear Differential Equations and Their Applications. Springer International Publishing, Cham, 2015. ISBN 978-3-319-20827-5. doi:10.1007/978-3-319-20828-2
2015 doi
-
[24]
Robert P. Sherman. Maximal Inequalities for Degenerate U -Processes with Applications to Optimization Estimators . The Annals of Statistics, 22 0 (1), 3 1994. ISSN 0090-5364. doi:10.1214/aos/1176325377
1994
-
[25]
Empirical Processes with Applications to Statistics
Galen R Shorack and Jon A Wellner. Empirical Processes with Applications to Statistics . Society for Industrial and Applied Mathematics, 10 2009. ISBN 0898716845. doi:10.1137/1.9780898719017
2009 doi
-
[26]
Sz \' e kely and Maria L
G \' a bor J. Sz \' e kely and Maria L. Rizzo. The Energy of Data . Annual Review of Statistics and Its Application, 4 0 (1): 0 447--479, March 2017. ISSN 2326-831X. doi:10.1146/annurev-statistics-060116-054026
2017 doi
-
[27]
van de Vaart
Aad W. van de Vaart. Asymptotic Statistics . Cambridge University Press, 10 1998. ISBN 9780521496032. doi:10.1017/CBO9780511802256
1998 doi
-
[28]
van der Vaart and Jon A
Aad W. van der Vaart and Jon A. Wellner. Weak Convergence and Empirical Processes . Springer Series in Statistics. Springer New York, New York, NY, 1996. ISBN 978-1-4757-2547-6. doi:10.1007/978-1-4757-2545-2
1996 doi
-
[29]
Ishaq Bhatti
Li Xing Zhu, Kai Tai Fang, and M. Ishaq Bhatti. On Estimated Projection Pursuit-Type Cr \' a mer-Von Mises Statistics . Journal of Multivariate Analysis, 63 0 (1), 1997. ISSN 0047259X. doi:10.1006/jmva.1997.1673
1997
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.