REVIEW 3 major objections 5 minor 1 cited by
Distribution Regression with Censored Selection
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Using a censored selection rule and a single binary instrument, the paper proves that wage–work-hours sorting is point identified at every hours threshold, and uses this to decompose the UK gender wage gap by worker type.
desk verdict A useful censored-selection extension of CFL with a real computation-theory gap in the smoothed Step 3 that needs attention before the inference results are fully trustworthy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the local Gaussian representation (LGR) of the joint CDF of the latent variables, $F_{S^*,Y^*}(s,y)=\Phi_2(\Phi^{-1}(F_{S^*}(s)), \Phi^{-1}(F_{Y^*}(y)); \rho(s,y))$, with local correlation parameter $\rho(s,y)$ measuring local dependence. The censored selection rule $S=\max(S^*,0)$ observable together with $Y=Y^*$ whenever $S>0$ turns the problem into one of recovering $\rho(s,y)$ at every threshold from the distribution of $(S,Y,Z)$. The load-bearing identity is equation (4), $\Pr(0<S\le s, Y\le y\mid Z=z) = \Phi_2(\mu_z(s), \nu(y); \rho_z(s,y)) - \Phi_2(\mu_z(s_0), \nu(y); \rho(s_0,y))$, whose left side is observed and whose right side is strictly increasing in $\rho_z(s,y)$, yielding point identification from a binary instrument. The estimator is a three-step procedure: probit regressions for each selection margin $\mu_z(s)$, a probit with sample-selection correction for $(\nu(y), \rho(s_0,y))$, and a bivariate probit for each remaining sorting parameter $\rho_z(s,y)$.
What would settle it
Using the observed data, one can solve the two-equation system at $s_0$ for each wage level $y$; if no solution exists with $\rho(s_0,y)\in[-1,1]$ for some $y$, the exclusion restrictions are rejected. If the instrument has more than two values, a minimum-distance test of the overidentifying restrictions directly tests Assumption 1; additionally, the estimated $\rho_z(s,y)$ from equation (4) must keep all implied joint probabilities in $[0,1]$ across thresholds, which the paper's Remark 1 shows can fail numerically—checking whether this happens at the estimated parameters, not just at starting values, would falsify the model.
Extended reading notes
Core claim
Under the local Gaussian representation, the joint distribution of the latent selection variable $S^*$ and latent outcome $Y^*$ is written at every point $(s,y)$ as a bivariate normal CDF $\Phi_2(\mu(s), \nu(y); \rho(s,y))$, where the local correlation $\rho(s,y)$ is the sorting parameter that governs the sign and strength of selection. With the censored selection rule $S=\max(S^*,0)$ and $Y=Y^*$ if $S>0$, and a binary instrument $Z$ satisfying non-degeneracy, relevance, outcome exclusion, and sorting exclusion at $s_0$, the paper proves that the selection margins $\mu_z(s)$ are identified by the selection probabilities, the pair $(\nu(y), \rho(s_0,y))$ is the unique solution of a two-equation system at $s_0$, and then every other threshold $s\neq s_0$ yields $\rho_z(s,y)$ uniquely from equation (4) because $\Phi_2$ is strictly increasing in $\rho$. The paper calls the resulting model censored distribution regression, proves this identification as Theorem 1, provides a functional central limit theorem for the three-step estimator as Theorem 2, and bootstrap-uniform confidence bands as Theorem 3. On UK work-hours and wage data, it shows that sorting into full-time and overtime work is heterogeneous across gender, marital status, time, and wage quantile, patterns that a binary employment selection rule cannot reveal.
Load-bearing premise
The load-bearing premise is the sorting exclusion restriction: at the censoring point (zero weekly hours), the local correlation between latent desired hours and offered wage is the same for both values of the out-of-work benefit instrument once covariates are controlled; if this fails, the outcome distribution and every sorting parameter are unidentified and all application estimates inherit the bias.
Editorial extensions
If this is right
- Researchers who currently dichotomize censored selection variables—employment, program participation, unemployment duration—can recover the full selection-sorting function at every threshold using the same binary-instrument exclusion assumptions required by standard Heckman-type models.
- Wage gaps can be decomposed by worker type: the paper's UK application separates composition, wage structure, hours structure, and hours-wage sorting, finding that hours structure and sorting narrow the low-quantile gender gap and widen the high-quantile gap for full-time workers, and that selection behavior explains most of the small overtime wage gap.
- The model covers continuous, discrete, and mixed outcomes, extending distributional analysis beyond the mean or median and beyond Gaussian errors, unlike quantile selection models that require continuous outcomes.
- Uniform confidence bands for the sorting function obtained by multiplier bootstrap allow testing functional hypotheses such as sorting being zero, non-negative, or constant across wage quantiles.
- Because the exclusion restrictions are local to a single threshold $s_0$, the same design applies at any censoring or policy cutoff, such as the 34- and 40-hour thresholds used to define full-time and overtime work.
Reading between the lines
- The identification logic cascades: once $(\nu(y), \rho(s_0,y))$ is identified, each additional threshold contributes its own monotone equation, so adding a threshold costs only one more bivariate probit; this suggests a general multi-threshold selection design for settings like disability severity bins or loan-to-value cutoffs.
- If the sorting exclusion at $s_0$ fails, the parameters are partially identified; bounding the local correlation would propagate bounds to $\rho(s,y)$ at all thresholds, yielding a sensitivity analysis that the paper does not develop.
- The smoothing fix in Remark 1 for negative predicted probabilities implies the likelihood can be ill-behaved when selection is strong at nearby thresholds; a testable robustness check is to vary the smoothing threshold $\tau$ and report whether the estimated sorting function changes.
- The application's finding that selection effects move the gender gap in opposite directions at low and high quantiles for full-time workers implies the binary selection model would report a sign of selection that is wrong at the top; comparing censored-DR and binary-DR estimates on the same UK data is a direct check.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a semiparametric distribution regression model with a censored selection rule, extending the binary-selection distribution regression of Chernozhukov, Fernández-Val, and Luo (CFL) to settings where the selection variable is censored rather than binary. The model is built on the local Gaussian representation (LGR) of the joint distribution of latent selection S* and latent outcome Y*, and identification is achieved through exclusion restrictions: an outcome exclusion restriction and a local sorting exclusion restriction at a point s0. The main theoretical contribution is Theorem 1, which proves point identification of the local sorting parameter ρz(s,y) for s≠s0 as the unique solution to equation (4); identification of the remaining parameters (ν(y),ρ(s0,y)) is imported from CFL. The authors propose a three-step estimation algorithm: probit for the selection margins, selection-corrected bivariate probit for the outcome margin and sorting at s0, and bivariate probit for the sorting parameter at each (s,y). They state a functional central limit theorem and a multiplier-bootstrap uniform inference procedure for the sorting function, and they apply the method to UK data to estimate selection sorting into full-time and overtime work and to decompose gender wage gaps by worker type.
Significance. If the theorems are correct, this is a useful and timely extension that lets researchers estimate selection sorting as a function of both the outcome and the level of the censored selection variable, rather than only a binary employment indicator. The clean monotonicity argument behind Theorem 1 is a genuine strength, and the paper provides explicit score and Hessian expressions together with a multiplier-bootstrap algorithm, which is valuable for applied work. The LGR is a representation rather than a testable restriction, so I do not see the circularity concern raised by the reader as an internal inconsistency; the substantive content comes from the stated exclusion restrictions. The empirical application illustrates the new objects and reports decomposition results that are interpretable and policy-relevant. However, two technical issues must be resolved before the reported confidence bands can be taken at face value: the implemented smoothing in Step 3 is outside the asymptotic theory, and the proof of the FCLT does not establish a key uniform lower bound on the denominators appearing in the scores and Hessians.
major comments (3)
- [Section 3.3, Remark 1; Section 3.4; Appendix B] The implemented Step 3 in Remark 1 replaces every model probability p by f(p) whenever p<τ, and f'(p) is not equal to 1 in that region. The asymptotic theory in Section 3.4 and Appendix B is derived for the maximizer of the unmodified likelihood L3(ρsy,ηsy), whose score S3sy in (10) uses denominators A1–A4. If any observation has a predicted probability below τ at the true parameters, at the final estimates, or along the bootstrap draws, the implemented estimator solves different first-order conditions, and the influence function (14), the matrices H3sy and J3sy, and the bootstrap bands no longer describe the reported estimator. The paper reports no value of τ, no diagnostic on how many observations are affected at the final estimates or across the 500 bootstrap repetitions, and the simulation remark concerns only negative probabilities rather than positive probabilities in (0,τ). Since the sorting function and its uniform confidence bands are the paper's main claimed contribution, this is a load-bearing gap between computation and theory. The authors should either set τ=0, or extend the asymptotic theory to cover the transformed objective and provide diagnostics showing that the set of affected observations is negligible, or show sensitivity of the empirical conclusions to τ.
- [Appendix A, Step 2; Assumption 2; equation (10)] The proof of Theorem 2 requires uniform boundedness of the quantities (Ã1,Ã2,Ã3,Ã4) used as denominators in the scores and Hessians. The verification in Step 2 asserts that these are bounded uniformly, but no lower bound is established. Assumption 2 only imposes upper bounds on conditional densities and compactness of parameter and support sets; it does not rule out A1=Φ2(Z'μs,X'νy;g(Z'ρsy)) or A3=Φ2(Z'μ0,X'νy;g(X'ρ0y))−Φ2(Z'μs,X'νy;g(Z'ρsy)) approaching zero as z approaches the boundary of its support or as y approaches the boundary of Y. Without a uniform lower bound of the form inf_{z∈Z1} min_j A_j > c > 0 on SY, the quantities H3sy, J3sy, and the influence function ψ3sy are not well-defined, and the FCLT in Theorem 2 is not established. This needs to be stated as an assumption or proved from the existing assumptions.
- [Section 2.2, Assumption 1(4)] The sorting exclusion restriction ρz(s0,y)=ρ(s0,y) is the key identifying assumption for (ν(y),ρ(s0,y)), and through equation (4) it also underpins identification of every ρz(s,y) for s≠s0. With a binary instrument the restriction is not testable, and the paper's justification—that the widely used HSM satisfies the analogous restriction—is a plausibility argument rather than direct evidence. If this assumption fails, the sorting estimates and all wage decompositions in Section 4 inherit the bias. The paper should provide a sensitivity analysis (for example, estimates under alternative choices of s0 or under a model that relaxes the restriction), discuss what is partially identified without it, or report an overidentification check if more than two values of Z are available. This is a limitation rather than an internal inconsistency, but it is load-bearing for the empirical conclusions.
minor comments (5)
- [Section 1, page 3] There is a typo: 'Fisher trasnformation' should read 'Fisher transformation'.
- [Section 3.1, equation (5)] The notation is confusing: z'ν(y) is used in the outcome equation but then z'ν(y)=x'ν(y) is given as the exclusion restriction. The authors should define z=(x',z1')' explicitly and state which coefficients are set to zero under the exclusion restrictions.
- [Section 3.4, Theorem 2] In the display following Theorem 2, the symbol ';Zρsy' appears where a weak-convergence arrow (⇝ or ⇒) is intended; this should be corrected.
- [Section 4.2, Figure 3] The x-axis labels in Figure 3 include the R expression 'seq(0.1, 0.9, 0.01)'; the axis should simply be labeled 'Wage quantile index'.
- [Section 4.3.1 and Appendix C] The text refers to 'Figures 9 and 10 in the Appendix C', but Appendix C contains Figures 10 and 11; the cross-reference is incorrect.
Circularity Check
No circularity: the censored-selection DR identification is a genuine monotonicity inversion and the sorting parameters are estimated by likelihood, not read back from fitted constants.
full rationale
The paper's new identification result (Theorem 1) proves uniqueness of rho_z(s,y) by strict monotonicity of Phi2 in its correlation parameter (Appendix A.1), after the marginals and rho(s0,y) are identified; the observed probability in (4) is data and the parameter is solved by inversion, so no output equation reduces to an input equation. The LGR (Lemma 1 from CFL) is a parameter-free mathematical identity - for any joint CDF value inside the Frechet bounds there is a Gaussian-copula rho reproducing it - so citing it is not circular. The paper explicitly imports the binary-selection identification of (nu(y), rho(s0,y)) from CFL (Section 2.2), a prior work with overlapping authorship; however, CFL's assumptions do not include the present target (sorting at s>s0), the new Step 3 likelihood and FCLT are derived in the paper, and the self-citation is cumulative support rather than a premise equaling the conclusion. The empirical sorting functions and decompositions are maximum-likelihood estimates of model functionals, not predictions forced by fitted constants. Remark 1's smoothing transformation may create a gap between the implemented Step-3 objective and the asymptotic theory, but that is a correctness risk, not circularity.
Assumptions & free parameters
free parameters (3)
- Work-hours thresholds (34, 40) =
34 and 40 hours
- Smoothing threshold tau with epsilon = tau/2 =
not reported
- Exclusion point s0 =
0
assumptions (4)
- domain assumption Assumption 1: existence of a binary instrument Z1 satisfying outcome exclusion nu_z(y) = nu(y) and sorting exclusion rho_z(s0,y) = rho(s0,y), plus relevance and non-degeneracy at s0.
- domain assumption Equation (5): the conditional joint distribution belongs to the BDR family Phi2(-z'mu(s), -x'nu(y); g(z'rho(s,y))).
- standard math Lemma 1 (LGR): any joint CDF can be represented pointwise as a bivariate standard Gaussian CDF with local correlation rho(s,y).
- domain assumption Assumption 2: iid sampling, compact supports, smooth conditional densities, unique interior maximizers, and non-singular expected Hessians.
Cite this review
Pith. "Pith review of Distribution Regression with Censored Selection." pith.science (2026). https://pith.science/paper/4Q7FGGTI
@misc{pith2026250510814,
author = {Pith},
title = {Pith review of: Distribution Regression with Censored Selection},
year = {2026},
howpublished = {\url{https://pith.science/paper/4Q7FGGTI}},
note = {Machine review of arXiv:2505.10814}
}
read the original abstract
We develop a distribution regression model with a censored selection rule, offering a semi-parametric generalization of the Heckman selection model. Our approach applies to the entire distribution, extending beyond the mean or median, accommodates non-Gaussian error structures, and allows for heterogeneous effects of covariates on both the selection and outcome distributions. By employing a censored selection rule, our model can uncover richer selection patterns according to both outcome and selection variables, compared to the binary selection case. We analyze identification, estimation, and inference of model functionals such as sorting parameters and distributions purged of sample selection. An application to labor supply using data from the UK reveals different selection patterns into full-time and overtime work across gender, marital status, and time. Additionally, decompositions of wage distributions by gender show that selection effects contribute to a decrease in the observed gender wage gap at low quantiles and an increase in the gap at high quantiles for full-time workers. The observed gender wage gap among overtime workers is smaller, which may be driven by different selection behaviors into overtime work across genders.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
A Pairwise Differencing Distribution Regression Approach for Network Models
A conditional maximum likelihood estimator for distribution regression in dyadic networks with two-way fixed effects is developed, with joint inference across thresholds.
Reference graph
Works this paper leans on
-
[1]
Amemiya, T. (1978): The estimation of a simultaneous equation generalized probit model, Econometrica: Journal of the Econometric Society, 1193--1205
work page 1978
-
[2]
--- -.1pt --- -.1pt --- (1979): The estimation of a simultaneous-equation Tobit model, International economic review, 169--181
work page 1979
-
[3]
Arellano, M. and S. Bonhomme (2017): Quantile selection models with an application to understanding changes in wage inequality, Econometrica, 85, 1--28
work page 2017
-
[4]
Bertrand, M., C. Goldin, and L. F. Katz (2010): Dynamics of the gender gap for young professionals in the financial and corporate sectors, American economic journal: applied economics, 2, 228--255
work page 2010
-
[5]
Blank, R. M. (1990): Are Part-Time Jobs Bad Jobs? in A Future of Lousy Jobs? The Changing Structure of U.S. Wages, ed. by G. Burtless, Washington, DC: Brookings Institution Press, 123--155
work page 1990
-
[6]
Blau, F. D. and L. M. Kahn (2017): The gender wage gap: Extent, trends, and explanations, Journal of economic literature, 55, 789--865
work page 2017
-
[7]
Blinder, A. S. (1973): Wage discrimination: reduced form and structural estimates, Journal of Human resources, 436--455
work page 1973
-
[8]
Blundell, R., M. Costa-Dias, D. Goll, and C. Meghir (2021): Wages, experience, and training of women over the life cycle, Journal of Labor Economics, 39, S275--S315
work page 2021
Show all 36 references
-
[9]
Gosling, H
Blundell, R., A. Gosling, H. Ichimura, and C. Meghir (2007): Changes in the distribution of male and female wages accounting for employment composition using bounds, Econometrica, 75, 323--363
2007
-
[10]
Reed, and T
Blundell, R., H. Reed, and T. M. Stoker (2003): Interpreting aggregate wage growth: The role of labor market participation, American Economic Review, 93, 1114--1131
2003
-
[11]
(2012): Involuntary part-time workers in Britain: evidence from the labour force survey, Industrial Relations Journal, 43, 242--259
Cam, S. (2012): Involuntary part-time workers in Britain: evidence from the labour force survey, Industrial Relations Journal, 43, 242--259
2012
-
[12]
(1997): Semiparametric estimation of the Type-3 Tobit model, Journal of Econometrics, 80, 1--34
Chen, S. (1997): Semiparametric estimation of the Type-3 Tobit model, Journal of Econometrics, 80, 1--34
1997
-
[13]
Chen, S., N. Liu, H. Zhang, and Y. Zhou (2024): Estimation of wage inequality in the UK by quantile regression with censored selection, Journal of Econometrics, 105733
2024
-
[14]
Fern \'a ndez-Val, and S
Chernozhukov, V., I. Fern \'a ndez-Val, and S. Luo (2023): Distribution regression with sample selection and UK wage decomposition, Tech. rep., cemmap working paper
2023
-
[15]
Fern\'andez-Val, and S
Chernozhukov, V., I. Fern\'andez-Val, and S. Luo (2025): Distribution Regression with Sample Selection and UK Wage Decomposition, 1978-2013. [data collection], Tech. Rep. SN: 9355, Office for National Statistics, Institute for Fiscal Studies, [original data producer(s)]
2025
-
[16]
Fern \'a ndez-Val, and B
Chernozhukov, V., I. Fern \'a ndez-Val, and B. Melly (2013): Inference on counterfactual distributions, Econometrica, 81, 2205--2268
2013
-
[17]
(2018): Who chooses part-time work and why, Monthly Lab
Dunn, M. (2018): Who chooses part-time work and why, Monthly Lab. Rev., 141, 1
2018
-
[18]
van Vuuren, and F
Fern \'a ndez-Val, I., A. van Vuuren, and F. Vella (2021): Nonseparable sample selection models with censored selection rules, Journal of Econometrics, 105088
2021
-
[19]
Gale, D. and H. Nikaido (1965): The Jacobian matrix and global univalence of mappings, Mathematische Annalen, 159, 81--93
1965
-
[20]
Kampelmann, and F
Garnero, A., S. Kampelmann, and F. Rycx (2014): Part-time work, wages, and productivity: evidence from Belgian matched panel data, ILR Review, 67, 926--954
2014
-
[21]
Gin \'e , E. and J. Zinn (1984): Some limit theorems for empirical processes, The Annals of Probability, 929--989
1984
-
[22]
(2014): A grand gender convergence: Its last chapter, American economic review, 104, 1091--1119
Goldin, C. (2014): A grand gender convergence: Its last chapter, American economic review, 104, 1091--1119
2014
-
[23]
Machin, and C
Gosling, A., S. Machin, and C. Meghir (2000): The changing distribution of male wages in the UK, The Review of Economic Studies, 67, 635--666
2000
-
[24]
(1974): Shadow prices, market wages, and labor supply, Econometrica: journal of the econometric society, 679--694
Heckman, J. (1974): Shadow prices, market wages, and labor supply, Econometrica: journal of the econometric society, 679--694
1974
-
[25]
--- -.1pt --- -.1pt --- (1979): Sample selection bias as a specification error, Econometrica
1979
-
[26]
Hirsch, B. T. (2005): Why do part-time workers earn less? The role of worker and job skills, ILR Review, 58, 525--551
2005
-
[27]
Honore, B. E., E. Kyriazidou, and C. Udry (1997): Estimation of type 3 tobit models using symmetric trimming and pairwise comparisons, Journal of econometrics, 76, 107--128
1997
-
[28]
Kitagawa, E. M. (1955): Components of a difference between two rates, Journal of the american statistical association, 50, 1168--1194
1955
-
[29]
Lee, M.-j. and F. Vella (2006): A semi-parametric estimator for censored selection models with endogeneity, Journal of Econometrics, 130, 235--252
2006
-
[30]
Maasoumi, E. and L. Wang (2019): The gender gap between earnings distributions, Journal of Political Economy, 127, 2438--2504
2019
-
[31]
Mulligan, C. B. and Y. Rubinstein (2008): Selection, investment, and women's relative wages over time, The Quarterly Journal of Economics, 123, 1061--1110
2008
-
[32]
Noonan, M. C., M. E. Corcoran, and P. N. Courant (2005): Pay differences among the highly trained: Cohort differences in the sex gap in lawyers' earnings, Social forces, 84, 853--872
2005
-
[33]
(1973): Male-female wage differentials in urban labor markets, International economic review, 693--709
Oaxaca, R. (1973): Male-female wage differentials in urban labor markets, International economic review, 693--709
1973
-
[34]
van der Vaart, A. W. (1998): Asymptotic statistics, Cambridge university press
1998
-
[35]
van der Vaart, A. W. and J. A. Wellner (1996): Weak Convergence and Empirical Processes, Springer Series in Statistics
1996
-
[36]
(1993): A simple estimator for simultaneous models with censored endogenous regressors, International Economic Review, 441--457
Vella, F. (1993): A simple estimator for simultaneous models with censored endogenous regressors, International Economic Review, 441--457
1993
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.