{"id":"72b10eb0-2f18-4271-bf2d-137217123a27","arxiv_id":"2412.05621","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Minimum sliced distance estimators are consistent and asymptotically normal in nonregular econometric models with parameter-dependent supports, unlike maximum likelihood estimators.","lead":"This paper proposes minimum sliced distance estimators for structural econometric models whose data support depends on the parameters, a setting where maximum likelihood often has non-normal asymptotics. It proves these estimators are asymptotically normal, so standard Wald inference works, and illustrates them on an auction model.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 4.1's displayed asymptotic covariance is dimensionally invalid for dψ>1 and pairs derivative factors with the wrong projection index, undermining the paper's Wald-inference claim.","rationale":"The core idea of the paper—using a sliced L2 distance to bypass the non-normal asymptotics of MLE in parameter-dependent support models—is plausible, and the simple uniform examples confirm that the remainder near the kink is O(δ^3), so the quadratic approximation can hold. I looked first at the reader's weakest assumption, the behavior of an integrable weight on shrinking boundary-crossing intervals in Lemma 4.2. That concern does not, on reflection, sink the proof: for w∈L1, local L1 masses over intervals of length O(δ) converge to zero uniformly in the interval center, and in Condition 4.2 the factor Tδ^2 in the numerator is controlled by (1+√Tδ)^2, so the term is at most a vanishing local L1 mass. The proof's displayed O(τ_T^3) is sloppy—it drops w and misstates the scaling—but the claimed implication can be recovered from integrability alone. What cannot be recovered by a simple continuity argument is the covariance formula in Proposition 4.1. The theorem is the paper's inference result, and as printed it is dimensionally invalid for dψ>1 and pairs derivatives with the wrong projection index. This is an internal inconsistency, not a disagreement with prior consensus. It is fixable, so the appropriate verdict remains conditional rather than reject. The concrete check—re-deriving V0 from Lemma 4.6 and then checking Wald coverage in the auction design—would settle whether the corrected covariance supports the paper's 'simple inference' claim.","tokens_in":32303,"tokens_out":18236,"duration_ms":171877,"concrete_test":"Re-derive V0 from Lemma 4.6: write the two components of the summand and compute their covariance, then compare with Proposition 4.1. Check whether the (1,2) block equals ∫∫∫∫ A12(t,s;u,v) D(t;u)D(s;v)^⊤ w(t)w(s)dtds dς(u)dς(v). If it does, the printed factor D(s;u)D(t;u)^⊤ and the e1 combination are wrong. Then, in the Section 5 auction design (dψ=2), compute 95% Wald intervals at T=200 using the corrected covariance and report coverage; coverage near 0.95 would confirm the fix, while coverage far from nominal would show the inference claim needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 4.1, the result that delivers the paper's advertised simple inference, states Ω0=(e1',−e1')V0(e1;−e1) with e1=(1,...,1)' a dψ-vector. For dψ>1 this is a 1×1 scalar, while B0^{-1}Ω0B0^{-1} must be dψ×dψ; in the auction example dψ=2, so the displayed covariance matrix is not a valid covariance matrix for the limiting normal law. The problem is not just the e1/identity confusion. In the displayed V0, the block A(t,s;u,v) is indexed by two different projection directions u and v, but the Kronecker factor is written D(s;u,ψ0)D(t;u,ψ0)^⊤, attaching both derivative factors to the first projection u and swapping the threshold arguments. An independent derivation from Lemma 4.6's summand gives covariance blocks with D(t;u)D(s;v)^⊤: the (1,2) block is ∫∫∫∫ A12(t,s;u,v)⊗D(t;u)D(s;v)^⊤ w(t)w(s)dtds dς(u)dς(v). As printed, the formula would not produce correct Wald standard errors even after the e1 issue is fixed. This is internally inconsistent rather than a matter of disputed regularity conditions, and it directly affects the central 'simple inference' contribution. The reader's separate concern about unbounded w in Lemma 4.2 is less damaging: for w∈L1, ∫_{c−δ}^{c+δ}w=o(1) uniformly in c, and the factor x^2/(1+x)^2 with x=√Tδ is bounded by 1, so only vanishing of the local L1 mass is needed, not linear shrinkage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes minimum sliced distance (MSD) estimation, covering minimum sliced Wasserstein distance (MSWD) and minimum sliced Cramér distance (MSCD), for structural econometric models with possibly parameter-dependent support. The main theoretical result, Theorem 3.2, establishes √T-consistency and asymptotic normality of the MSD estimator under high-level assumptions (Assumptions 3.1–3.6), following the quadratic approximation approach of Andrews (1999) and Pollard (1980). Proposition 4.1 then claims that for conditional MSCD models, primitive smoothness conditions (Conditions 4.1–4.2 and related differentiability assumptions) verify the high-level assumptions, so the estimator is consistent and asymptotically normal regardless of whether the conditional density jumps at the parameter-dependent support boundary. The paper verifies the conditions explicitly for one-sided and two-sided uniform models and presents a simulation study for an auction model.","tokens_in":32742,"tokens_out":7289,"duration_ms":64816,"significance":"If the claims are correct, the paper offers a genuine practical advance: a one-step estimator that avoids auxiliary models and yields standard Wald inference in nonregular structural models. The high-level theorem is clean and follows a well-established template; the exact norm-differentiability computations for the uniform examples are a useful check; and the auction simulation demonstrates good finite-sample normal approximation. However, the covariance formula in Proposition 4.1 (and its analogue in Theorem 3.2) is dimensionally wrong for dψ>1, and the displayed V0 block formula does not match the CLT summand in Lemma 4.6. Because the paper's advertised contribution is 'simple inference', these errors are central and must be corrected before the paper can be accepted.","major_comments":[{"comment":"The asymptotic covariance expression Ω0=(e1',−e1')V0(e1;−e1) with e1 a dψ-vector of ones is a scalar whenever dψ>1, whereas B0^{-1}Ω0B0^{-1} must be a dψ×dψ matrix. In the auction example dψ=2, so the displayed formula cannot deliver standard errors for the two components of ψ0. The correct selector should be a block matrix such as (I_{dψ}, −I_{dψ}), not a pair of vectors of ones. This affects the main inference claim, not just a special case.","section":"Theorem 3.2 and Proposition 4.1"},{"comment":"The Kronecker factor in the displayed V0 is D(s;u,ψ0)D(t;u,ψ0)^⊤, attaching both derivative factors to the same projection u and placing the first projection's threshold in the second slot. The CLT summand in Lemma 4.6 has a first component indexed by u with threshold t and a second component indexed by v with threshold s, so the (1,2) covariance block must be ∫∫∫∫ A12(t,s;u,v) ⊗ D(t;u,ψ0)D(s;v,ψ0)^⊤ w(t)w(s) dt ds dς(u)dς(v). As printed, the covariance would not agree with the actual limiting distribution even after the e1 issue is fixed.","section":"Proposition 4.1, V0 display"},{"comment":"The proof of Lemma 4.2 bounds the boundary-crossing contribution to ∫|R_t|^2 w by C∥ψ−ψ0∥^2 |u1|(|g(X_i,ψ)−g(X_i,ψ0)|+2Cτ_T) and calls it O(τ_T^3). This implicitly requires that the integral of the weight w over an interval of length O(τ_T) is O(τ_T). The stated assumption is only that w is integrable; for unbounded integrable weights, the integral over a shrinking interval is o(1) but not necessarily O(τ_T). The verification of Condition 4.2 is therefore incomplete for the class of weights allowed in Section 4. The gap is repairable (one can use boundedness of x^2/(1+x)^2 and vanishing L1 mass), but the proof as written does not establish the stated claim.","section":"Lemma 4.2 and Appendix D.3"}],"minor_comments":[{"comment":"The paper mixes √n and √T (e.g., Theorem 3.2, Proposition 4.1); all statements should use √T consistently.","section":"Throughout"},{"comment":"There is a typo: 'Wassserstein' should be 'Wasserstein'.","section":"Section 2.2.1"},{"comment":"The word 'seond-order' should be 'second-order'.","section":"Appendix D.3"},{"comment":"In the displayed expression for E[I(u⊤Z_t ≤ s)|X_t, ψ], the second case should be for u1 = 0, not u2 < 0, and the third case should be for u1 < 0.","section":"Lemma D.1"},{"comment":"In the final displayed line, G(t;u,ψ0) should be G(s;u,ψ0).","section":"Verification of Assumption 3.4(ii)"},{"comment":"The citation 'van de Vaart' should be 'van der Vaart'.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The e1/identity and V0 cross-index problems are substantial but appear to be localizable errors in the derivation; they do not suggest the method itself is invalid, and the simulation evidence is unaffected because the reported QQ plots use Monte Carlo standard deviations rather than plug-in standard errors. I would recommend a major revision with a recomputation of the covariance formula and a tightening of the weight assumptions in Lemma 4.2. There is no indication of citation or novelty problems; the paper is a genuine methodological contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis paper is worth reading, but do not trust the displayed covariance matrix. Fan and Park propose minimum sliced distance estimation for structural models with parameter-dependent supports and claim standard sqrt(T) asymptotic normality, which would be a real simplification over the non-normal MLE limits in Chernozhukov-Hong and over indirect inference, since no auxiliary model is needed. The high-level theorem follows the Andrews/Pollard template cleanly, and the exact computations for the one- and two-sided uniform models check out. The simulation on the Paarsch auction model is suggestive, though it reports only QQ plots and histograms, no coverage rates.\n\nThe soft spot is not the regularity conditions; it is the main formula. Theorem 3.2 and Proposition 4.1 state Omega0 = (e1', -e1') V0 (e1; -e1) with e1 a dpsi-vector of ones. For dpsi > 1 that is a scalar, and the limiting covariance B0^{-1} Omega0 B0^{-1} has to be dpsi x dpsi. The auction example has dpsi = 2, so the advertised Wald inference is not delivered. Worse, the V0 blocks appear to pair the wrong projection indices: the (1,2) block should be A12(t,s;u,v) (x) D(t;u)D(s;v)^T, but the paper has D(s;u)D(t;u)^T. This is an internal inconsistency, not a matter of interpretation, and it affects the central 'simple inference' claim.\n\nThe reader worried about unbounded weight functions in Lemma 4.2. On reading the proof, I think that concern is minor: for w in L1 the integral over the shrinking boundary-crossing interval vanishes by absolute continuity, and the T factor is absorbed by the (1 + sqrt(T)||psi-psi0||)^2 denominator. Integrability is enough.\n\nThe paper deserves a serious referee. The method is new in this econometric context and the high-level framework is useful, but the covariance formula has to be corrected and verified before the result can be used. I would send it out, with a clear instruction to re-derive V0 and Omega0 in the multivariate case.","headline":"Sliced distance estimation for nonregular models is a good idea, but the printed asymptotic covariance is wrong for dψ>1 and needs a serious fix before the paper's central claim holds.","tokens_in":33179,"tokens_out":7290,"would_cite":true,"duration_ms":61362,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62G20","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that minimum sliced distance estimators are root-T asymptotically normal in structural econometric models with parameter-dependent supports, giving practitioners standard Wald inference where maximum likelihood yields…","keywords":["minimum sliced distance estimation","sliced Cramér distance","sliced Wasserstein distance","parameter-dependent support","nonregular econometric models","asymptotic normality","auction model","indirect inference"],"falsifier":"Take the univariate one-sided uniform model $Y\\sim U[0,\\psi_0]$ with $\\psi_0=2$ and the integrable but unbounded weight $w(s)=|s-2|^{-1/2}$ on $[0,3]$, normalized to integrate to one. Compute the minimum sliced Cramér distance estimator on many simulated samples of size $T=10{,}000$ and compare the empirical distribution of $\\sqrt{T}(\\hat\\psi_T-\\psi_0)$ to a normal distribution. Lemma 4.2's verification of Condition 4.2 uses the fact that integrals over a shrinking boundary-crossing interval shrink linearly in the interval length; for this weight the integral over $[\\psi_0,\\psi]$ is proportional to $\\sqrt{\\psi-\\psi_0}$, so the $T$-scaled remainder term need not vanish. A non-normal sampling distribution or severely distorted Wald coverage for this weight would show that asymptotic normality does not hold for the full class of integrable weights.","tokens_in":32070,"feed_emoji":"📊","tokens_out":12185,"duration_ms":106909,"temperature":0.7,"pith_summary":"This paper proposes estimating structural parameters by minimizing a sliced L2 distance between the empirical measure of one-dimensional projections of the data and the distribution those projections would have under the structural model. The central claim is that, unlike maximum likelihood, the resulting minimum sliced distance estimator is asymptotically normal in models whose support depends on the unknown parameter, so standard Wald inference is available. The asymptotic normality is proved for a general high-level class of sliced distances and then verified for the minimum sliced Cramér distance estimator in conditional models, including one-sided and two-sided models where the conditional density may jump at the boundary. A simulation on an auction model with a parameter-dependent winning-bid support shows the normal approximation works at sample sizes of 100 and 200.","feed_headline":"One estimator gives standard inference where MLE turns non-normal","feed_subtitle":"The estimator stays root-T normal even when the density jumps at a parameter-dependent support boundary.","key_machinery":"The central object is the weighted sliced L2 discrepancy between data and model: for each projection direction $u$, the distance integrates squared differences of cumulative distribution functions (sliced Cramér distance) or quantile functions (sliced Wasserstein distance) over $s$ with a weight $w(s)$, then averages over directions. The argument is carried by showing the criterion is quadratically approximated around $\\psi_0$: $\\hat{S}(\\psi) \\approx \\text{constant} - 2(\\psi-\\psi_0)^\\top A_T/\\sqrt{T} + (\\psi-\\psi_0)^\\top B_T(\\psi-\\psi_0)$, with the remainder controlled by a norm-differentiability condition. That condition permits a finite number of kinks in the model-induced function at the parameter-dependent boundary, which is exactly where likelihood-based theory fails, and the remainder analysis for conditional models routes through degenerate U-statistics.","core_discovery":"The paper's central discovery is Theorem 3.2: under Assumptions 2.1 and 3.1–3.6, any estimator that minimizes the weighted sliced L2 distance between the empirical measure of the data and the model-induced measure satisfies $\\sqrt{T}(\\hat{\\psi}_T-\\psi_0) \\xrightarrow{d} N(0, B_0^{-1}\\Omega_0 B_0^{-1})$, where $B_0$ is the integrated outer product of the derivative of the model-induced distribution and $\\Omega_0$ is the asymptotic variance of a functional central limit term. Proposition 4.1 then verifies the high-level assumptions for the minimum sliced Cramér distance estimator in conditional models with one-sided parameter-dependent support, under primitive smoothness conditions on the conditional CDF and on the boundary function. The key point is that these conditions do not require the conditional density to be bounded away from zero at the boundary or to have a jump there, so the estimator is asymptotically normal whether or not the likelihood would have a non-normal limit.","pith_inferences":["A natural next question is whether the integrability assumption on $w$ can be strengthened to an explicit shrinking-mass condition; the boundary-interval argument in Lemma 4.2 suggests that unbounded weights may require it.","The same quadratic-approximation proof should transfer to other non-smooth econometric structures, such as censored regressions, kinked regression functions, or threshold models, wherever the model-induced CDF has bounded second derivatives away from a low-dimensional boundary.","The choice of projection directions is a finite-sample tuning issue the theory does not address: with many projections the estimator approximates the true sliced distance, but the auction simulation uses 100 directions and Adam optimization, so the practical recipe trails the theory."],"forward_implications":["Wald-type confidence intervals and t-tests become available for structural parameters in one-sided and two-sided parameter-dependent support models, eliminating the dichotomy where likelihood-based inference must switch between normal and non-normal limit theory.","The estimator is a one-step procedure: it does not require an auxiliary regression model or simulated samples from it, unlike indirect inference.","Because the high-level theorem covers any sliced L2 distance satisfying the assumptions, both the sliced Cramér distance and the sliced Wasserstein distance variants inherit the same asymptotic normality.","In the independent private-value procurement auction model, the normal approximation is accurate with 100–200 observations and comparable to indirect inference with a well-chosen starting value."],"supporting_citations":[{"why":"Documents the non-normal likelihood asymptotics in one-sided and two-sided parameter-dependent support models and provides the motivating models this paper targets.","marker":"Chernozhukov and Hong [2004]"},{"why":"Supplies the one-sided model and the boundary positivity condition that the paper's primitive conditions deliberately avoid.","marker":"Hirano and Porter [2003]"},{"why":"Proposes the indirect inference estimator that serves as the comparison baseline and whose auxiliary-model dependence this paper removes.","marker":"Li [2010]"},{"why":"Provides the quadratic-approximation framework used to prove consistency and asymptotic normality of the extremum estimator.","marker":"Andrews [1999]"},{"why":"Supplies the minimum-distance estimation theory from which the proof strategy descends.","marker":"Pollard [1980]"},{"why":"Establishes asymptotics for minimum Wasserstein distance estimation, the univariate basis for the sliced variants considered here.","marker":"Bernton et al. [2019]"},{"why":"Analyzes sliced Wasserstein estimation and its non-normal asymptotics, the setting the paper's normal theory extends.","marker":"Nadjahi et al. [2020b]"},{"why":"Gives maximal inequalities for degenerate U-processes used to control the remainder terms in the conditional-model verification.","marker":"Sherman [1994]"},{"why":"Contributes the Lipschitz-kernel lemma used in the degenerate U-statistic argument for Assumption 3.1 (ii).","marker":"Briol et al. [2019]"},{"why":"Provides uniform-convergence and stochastic equicontinuity results used to verify uniform consistency of the estimated criterion.","marker":"Newey [1991]"}],"fun_headline_variants":["Normal limit even when density jumps at support boundary","Root-T normal inference without smooth density","Minimum sliced distance normal even when MLE fails","Simple asymptotics for models with moving supports","One estimator, normal limit, no smooth density required"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the model's distribution function being nearly quadratic in the parameter except at the one point where the support boundary moves, and on the weight function's mass near that boundary shrinking away fast enough; the paper only assumes the weight is integrable, which may not be enough when the weight is unbounded.","fun_headline_variants_meta":{"raw":{"variants":["Normal limit even when density jumps at support boundary","Root-T normal inference without smooth density","Minimum sliced distance normal even when MLE fails","Simple asymptotics for models with moving supports","One estimator, normal limit, no smooth density required"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001401,"raw_usage":{"total_tokens":5591,"prompt_tokens":800,"completion_tokens":4791,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":416,"completion_tokens_details":{"reasoning_tokens":4722}},"tokens_in":416,"tokens_out":4791,"duration_ms":31622,"temperature":1.0,"reasoning_tokens":4722,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:32:17.367105+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the univariate one-sided uniform model $Y\\sim U[0,\\psi_0]$ with $\\psi_0=2$ and the integrable but unbounded weight $w(s)=|s-2|^{-1/2}$ on $[0,3]$, normalized to integrate to one. Compute the minimum sliced Cramér distance estimator on many simulated samples of size $T=10{,}000$ and compare the empirical distribution of $\\sqrt{T}(\\hat\\psi_T-\\psi_0)$ to a normal distribution. Lemma 4.2's verification of Condition 4.2 uses the fact that integrals over a shrinking boundary-crossing interval shrink linearly in the interval length; for this weight the integral over $[\\psi_0,\\psi]$ is proportional to $\\sqrt{\\psi-\\psi_0}$, so the $T$-scaled remainder term need not vanish. A non-normal sampling distribution or severely distorted Wald coverage for this weight would show that asymptotic normality does not hold for the full class of integrable weights.","supporting_citations":[{"cited_title":"Asymptotic Efficiency in Parametric Structural Models with Parameter-Dependent Support","cited_arxiv_id":null,"evidence_quote":"Supplies the one-sided model and the boundary positivity condition that the paper's primitive conditions deliberately avoid."},{"cited_title":"Estimation When a Parameter is on a Boundary","cited_arxiv_id":null,"evidence_quote":"Provides the quadratic-approximation framework used to prove consistency and asymptotic normality of the extremum estimator."},{"cited_title":"The Minimum Distance Method of Testing","cited_arxiv_id":null,"evidence_quote":"Supplies the minimum-distance estimation theory from which the proof strategy descends."},{"cited_title":"Duncan, and Mark Girolami","cited_arxiv_id":null,"evidence_quote":"Contributes the Lipschitz-kernel lemma used in the degenerate U-statistic argument for Assumption 3.1 (ii)."}],"review_version":1}