{"id":"e6f1fde1-50cf-4fe2-99c5-196af24c1542","arxiv_id":"2505.10814","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A distribution regression model with censored selection identifies and estimates selection sorting at each work-hours threshold, and a UK application shows different sorting into full-time and overtime work across gender, marital status, and time.","lead":"This paper extends distribution regression with sample selection from a binary rule to a censored work-hours rule, allowing researchers to estimate how selection into part-time, full-time, and overtime work varies with wages. It supplies identification, a three-step estimator, uniform inference, and a UK gender wage-gap decomposition, but the empirical findings rest on strong and partly untestable exclusion restrictions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Implemented Step 3 smoothing is outside the asymptotic theory, so Theorems 2–3 may not cover the reported estimator.","rationale":"I considered Assumption 1(4), the sorting exclusion restriction, which is indeed load-bearing because failure breaks identification of ν(y) and ρ(s0,y). However, it is an explicit and standard identifying assumption that the paper acknowledges as controversial, and it is not an internal inconsistency. The sharper, more checkable issue is the mismatch between the implemented Step 3 objective with the f-transformation and the asymptotic theory for the unmodified likelihood: the FCLT and bootstrap FCLT are proven for a different estimator unless the affected set is asymptotically negligible. The reader flagged the smoothing transformation in the rationale but selected sorting exclusion as the weakest assumption; my read is that the implementation-theory mismatch is the more concrete load-bearing concern. The identification theorem itself appears correct given Assumption 1, since uniqueness follows from strict monotonicity of Φ2 in ρ. I therefore leave the verdict unchanged: acceptance should remain conditional on reporting τ, documenting the share of affected observations, and either aligning the implementation with the theory or extending the proofs to the transformed objective.","tokens_in":30554,"tokens_out":11182,"duration_ms":124938,"concrete_test":"Re-run the reported specification (or a simulation calibrated to the UK sample) with the transformation deactivated, replacing f(p) by p and using constrained optimization so that all four probability terms stay non-negative, and recompute ρ(s,y) and the confidence bands. Also record the share of observations with p < τ at the final estimates and across bootstrap samples. If that share is not o(n^{-1/2}), or if the sorting estimates and bands change materially, Theorems 2–3 do not cover the implemented estimator and the theory must be re-derived with f in the objective; if the share is zero at the final estimates for all (s,y), the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Algorithm 1 Step 3 as implemented in Remark 1 replaces every model probability p in the Step-3 objective with f(p) whenever p < τ, and f'(p) ≠ 1 in that region. The asymptotic theory in Section 3.4 and Appendix B is derived for the maximizer of the unmodified likelihood L3(ρsy, ηsy), whose score uses denominators A1–A4 in (10). If any observation has a predicted probability below τ at the true parameters, at the final estimates, or along the bootstrap path, the implemented estimator solves different first-order conditions and the influence function (14), H3sy, J3sy, and Σρ in Appendix B no longer describe it. The paper reports no value of τ, no diagnostic on how many observations are affected at the final estimates or across the 500 bootstrap draws, and the simulation remark only concerns negative probabilities, not positive probabilities below τ. Since ρsy and its uniform confidence bands are the main claimed contribution, this is an internal gap between computation and theory, not a tuning detail.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a semiparametric distribution regression model with a censored selection rule, extending the binary-selection distribution regression of Chernozhukov, Fernández-Val, and Luo (CFL) to settings where the selection variable is censored rather than binary. The model is built on the local Gaussian representation (LGR) of the joint distribution of latent selection S* and latent outcome Y*, and identification is achieved through exclusion restrictions: an outcome exclusion restriction and a local sorting exclusion restriction at a point s0. The main theoretical contribution is Theorem 1, which proves point identification of the local sorting parameter ρz(s,y) for s≠s0 as the unique solution to equation (4); identification of the remaining parameters (ν(y),ρ(s0,y)) is imported from CFL. The authors propose a three-step estimation algorithm: probit for the selection margins, selection-corrected bivariate probit for the outcome margin and sorting at s0, and bivariate probit for the sorting parameter at each (s,y). They state a functional central limit theorem and a multiplier-bootstrap uniform inference procedure for the sorting function, and they apply the method to UK data to estimate selection sorting into full-time and overtime work and to decompose gender wage gaps by worker type.","tokens_in":30730,"tokens_out":7392,"duration_ms":77436,"significance":"If the theorems are correct, this is a useful and timely extension that lets researchers estimate selection sorting as a function of both the outcome and the level of the censored selection variable, rather than only a binary employment indicator. The clean monotonicity argument behind Theorem 1 is a genuine strength, and the paper provides explicit score and Hessian expressions together with a multiplier-bootstrap algorithm, which is valuable for applied work. The LGR is a representation rather than a testable restriction, so I do not see the circularity concern raised by the reader as an internal inconsistency; the substantive content comes from the stated exclusion restrictions. The empirical application illustrates the new objects and reports decomposition results that are interpretable and policy-relevant. However, two technical issues must be resolved before the reported confidence bands can be taken at face value: the implemented smoothing in Step 3 is outside the asymptotic theory, and the proof of the FCLT does not establish a key uniform lower bound on the denominators appearing in the scores and Hessians.","major_comments":[{"comment":"The implemented Step 3 in Remark 1 replaces every model probability p by f(p) whenever p<τ, and f'(p) is not equal to 1 in that region. The asymptotic theory in Section 3.4 and Appendix B is derived for the maximizer of the unmodified likelihood L3(ρsy,ηsy), whose score S3sy in (10) uses denominators A1–A4. If any observation has a predicted probability below τ at the true parameters, at the final estimates, or along the bootstrap draws, the implemented estimator solves different first-order conditions, and the influence function (14), the matrices H3sy and J3sy, and the bootstrap bands no longer describe the reported estimator. The paper reports no value of τ, no diagnostic on how many observations are affected at the final estimates or across the 500 bootstrap repetitions, and the simulation remark concerns only negative probabilities rather than positive probabilities in (0,τ). Since the sorting function and its uniform confidence bands are the paper's main claimed contribution, this is a load-bearing gap between computation and theory. The authors should either set τ=0, or extend the asymptotic theory to cover the transformed objective and provide diagnostics showing that the set of affected observations is negligible, or show sensitivity of the empirical conclusions to τ.","section":"Section 3.3, Remark 1; Section 3.4; Appendix B"},{"comment":"The proof of Theorem 2 requires uniform boundedness of the quantities (Ã1,Ã2,Ã3,Ã4) used as denominators in the scores and Hessians. The verification in Step 2 asserts that these are bounded uniformly, but no lower bound is established. Assumption 2 only imposes upper bounds on conditional densities and compactness of parameter and support sets; it does not rule out A1=Φ2(Z'μs,X'νy;g(Z'ρsy)) or A3=Φ2(Z'μ0,X'νy;g(X'ρ0y))−Φ2(Z'μs,X'νy;g(Z'ρsy)) approaching zero as z approaches the boundary of its support or as y approaches the boundary of Y. Without a uniform lower bound of the form inf_{z∈Z1} min_j A_j > c > 0 on SY, the quantities H3sy, J3sy, and the influence function ψ3sy are not well-defined, and the FCLT in Theorem 2 is not established. This needs to be stated as an assumption or proved from the existing assumptions.","section":"Appendix A, Step 2; Assumption 2; equation (10)"},{"comment":"The sorting exclusion restriction ρz(s0,y)=ρ(s0,y) is the key identifying assumption for (ν(y),ρ(s0,y)), and through equation (4) it also underpins identification of every ρz(s,y) for s≠s0. With a binary instrument the restriction is not testable, and the paper's justification—that the widely used HSM satisfies the analogous restriction—is a plausibility argument rather than direct evidence. If this assumption fails, the sorting estimates and all wage decompositions in Section 4 inherit the bias. The paper should provide a sensitivity analysis (for example, estimates under alternative choices of s0 or under a model that relaxes the restriction), discuss what is partially identified without it, or report an overidentification check if more than two values of Z are available. This is a limitation rather than an internal inconsistency, but it is load-bearing for the empirical conclusions.","section":"Section 2.2, Assumption 1(4)"}],"minor_comments":[{"comment":"There is a typo: 'Fisher trasnformation' should read 'Fisher transformation'.","section":"Section 1, page 3"},{"comment":"The notation is confusing: z'ν(y) is used in the outcome equation but then z'ν(y)=x'ν(y) is given as the exclusion restriction. The authors should define z=(x',z1')' explicitly and state which coefficients are set to zero under the exclusion restrictions.","section":"Section 3.1, equation (5)"},{"comment":"In the display following Theorem 2, the symbol ';Zρsy' appears where a weak-convergence arrow (⇝ or ⇒) is intended; this should be corrected.","section":"Section 3.4, Theorem 2"},{"comment":"The x-axis labels in Figure 3 include the R expression 'seq(0.1, 0.9, 0.01)'; the axis should simply be labeled 'Wage quantile index'.","section":"Section 4.2, Figure 3"},{"comment":"The text refers to 'Figures 9 and 10 in the Appendix C', but Appendix C contains Figures 10 and 11; the cross-reference is incorrect.","section":"Section 4.3.1 and Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The manuscript relies heavily on CFL for the identification of (ν(y),ρ(s0,y)) and for the Z-process inference template; because CFL is a self-cited companion with overlapping authorship, the editor may wish to confirm that the relevant CFL results are published or otherwise in a form the journal can rely on. The computational-theory gap in Remark 1 is the main technical obstacle, and it is fixable within the manuscript's scope, so I do not recommend rejection. I do not see a circularity problem: the LGR is a representation, and the sorting parameter is estimated by maximum likelihood rather than imposed from a fitted constant."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The censored-selection extension of CFL's distribution regression is a genuine and useful advance, and the empirical sorting-by-hours results will interest labor people. The identification theorem for rho_z(s,y) is a clean monotonicity argument, and the paper is admirably clear that censoring does not itself aid identification. The three-step algorithm is a natural generalization, and the decomposition by worker type is a nice payoff from modeling the intensive margin.\n\nThe soft spots are real but mostly fixable. The main issue is the smoothing in Remark 1. The asymptotic theory in Section 3.4 and Appendix B is derived for the maximizer of the unmodified likelihood L3, but the implemented estimator replaces every model probability below tau with f(p). If any observation has predicted probability below tau at the true parameters or along the bootstrap path, the influence function (14) no longer describes the estimator. The paper reports no value of tau and no diagnostics on how many observations are affected at the final estimates or across the 500 bootstrap draws. The claim that negative probabilities only occur far from the truth does not cover positive probabilities below tau, which can occur at the true parameters for some covariate cells. Since rho(s,y) and its uniform bands are the paper's main contribution, this is an internal gap between computation and theory, not a tuning detail. A referee should ask for the smoothing threshold, the fraction of affected observations, and either a proof that the smoothing is asymptotically negligible or a theory that covers the smoothed estimator.\n\nSecond, the sorting exclusion restriction at s0=0 is strong and partly untestable. The authors acknowledge this and defend it by analogy to HSM, which is fair but not fully convincing. The empirical sorting results and wage decompositions inherit this assumption.\n\nThird, there is no replication code or archive. The data are public, but code would help.\n\nWho is this for? Applied microeconomists studying selection at the intensive margin and econometricians working on distribution regression. It deserves a serious referee. The extension is solid, the empirical findings are interesting, and the smoothing gap is fixable. I would only cite the estimator after the authors close that gap or provide convincing diagnostics.","headline":"A useful censored-selection extension of CFL with a real computation-theory gap in the smoothed Step 3 that needs attention before the inference results are fully trustworthy.","tokens_in":31291,"tokens_out":2474,"would_cite":false,"duration_ms":25386,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P20","91B82","62G05","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Using a censored selection rule and a single binary instrument, the paper proves that wage–work-hours sorting is point identified at every hours threshold, and uses this to decompose the UK gender wage gap by worker type.","keywords":["distribution regression","censored selection","sample selection","local Gaussian representation","sorting parameter","gender wage gap","work hours","Heckman selection model"],"falsifier":"Using the observed data, one can solve the two-equation system at $s_0$ for each wage level $y$; if no solution exists with $\\rho(s_0,y)\\in[-1,1]$ for some $y$, the exclusion restrictions are rejected. If the instrument has more than two values, a minimum-distance test of the overidentifying restrictions directly tests Assumption 1; additionally, the estimated $\\rho_z(s,y)$ from equation (4) must keep all implied joint probabilities in $[0,1]$ across thresholds, which the paper's Remark 1 shows can fail numerically—checking whether this happens at the estimated parameters, not just at starting values, would falsify the model.","tokens_in":30306,"feed_emoji":"📊","tokens_out":10494,"duration_ms":87152,"temperature":0.7,"pith_summary":"This paper claims that sample-selection problems—wages observed only for people who work—can be handled by a distribution regression model in which the selection rule is a censored continuous variable such as weekly work hours rather than a binary employment indicator. The central result is that, with a binary instrument satisfying exclusion restrictions at a single threshold (the censoring point), the entire sorting function that describes local dependence between latent desired hours and latent offered wage is point identified at every hours threshold. This matters because it turns a censored selection variable, which researchers routinely dichotomize, into a source of information about who selects into part-time, full-time, or overtime work and how that selection varies across the wage distribution. The paper also delivers a three-step estimator with multiplier-bootstrap uniform inference and an application to UK data showing that selection patterns differ sharply by gender, marital status, and worker type and that these selection effects shape the observed gender wage gap.","feed_headline":"Censored hours reveal hidden selection into full-time and overtime work","feed_subtitle":"One binary instrument now identifies wage–hour sorting at every hours threshold, exposing gendered selection patterns.","key_machinery":"The central object is the local Gaussian representation (LGR) of the joint CDF of the latent variables, $F_{S^*,Y^*}(s,y)=\\Phi_2(\\Phi^{-1}(F_{S^*}(s)), \\Phi^{-1}(F_{Y^*}(y)); \\rho(s,y))$, with local correlation parameter $\\rho(s,y)$ measuring local dependence. The censored selection rule $S=\\max(S^*,0)$ observable together with $Y=Y^*$ whenever $S>0$ turns the problem into one of recovering $\\rho(s,y)$ at every threshold from the distribution of $(S,Y,Z)$. The load-bearing identity is equation (4), $\\Pr(0<S\\le s, Y\\le y\\mid Z=z) = \\Phi_2(\\mu_z(s), \\nu(y); \\rho_z(s,y)) - \\Phi_2(\\mu_z(s_0), \\nu(y); \\rho(s_0,y))$, whose left side is observed and whose right side is strictly increasing in $\\rho_z(s,y)$, yielding point identification from a binary instrument. The estimator is a three-step procedure: probit regressions for each selection margin $\\mu_z(s)$, a probit with sample-selection correction for $(\\nu(y), \\rho(s_0,y))$, and a bivariate probit for each remaining sorting parameter $\\rho_z(s,y)$.","core_discovery":"Under the local Gaussian representation, the joint distribution of the latent selection variable $S^*$ and latent outcome $Y^*$ is written at every point $(s,y)$ as a bivariate normal CDF $\\Phi_2(\\mu(s), \\nu(y); \\rho(s,y))$, where the local correlation $\\rho(s,y)$ is the sorting parameter that governs the sign and strength of selection. With the censored selection rule $S=\\max(S^*,0)$ and $Y=Y^*$ if $S>0$, and a binary instrument $Z$ satisfying non-degeneracy, relevance, outcome exclusion, and sorting exclusion at $s_0$, the paper proves that the selection margins $\\mu_z(s)$ are identified by the selection probabilities, the pair $(\\nu(y), \\rho(s_0,y))$ is the unique solution of a two-equation system at $s_0$, and then every other threshold $s\\neq s_0$ yields $\\rho_z(s,y)$ uniquely from equation (4) because $\\Phi_2$ is strictly increasing in $\\rho$. The paper calls the resulting model censored distribution regression, proves this identification as Theorem 1, provides a functional central limit theorem for the three-step estimator as Theorem 2, and bootstrap-uniform confidence bands as Theorem 3. On UK work-hours and wage data, it shows that sorting into full-time and overtime work is heterogeneous across gender, marital status, time, and wage quantile, patterns that a binary employment selection rule cannot reveal.","pith_inferences":["The identification logic cascades: once $(\\nu(y), \\rho(s_0,y))$ is identified, each additional threshold contributes its own monotone equation, so adding a threshold costs only one more bivariate probit; this suggests a general multi-threshold selection design for settings like disability severity bins or loan-to-value cutoffs.","If the sorting exclusion at $s_0$ fails, the parameters are partially identified; bounding the local correlation would propagate bounds to $\\rho(s,y)$ at all thresholds, yielding a sensitivity analysis that the paper does not develop.","The smoothing fix in Remark 1 for negative predicted probabilities implies the likelihood can be ill-behaved when selection is strong at nearby thresholds; a testable robustness check is to vary the smoothing threshold $\\tau$ and report whether the estimated sorting function changes.","The application's finding that selection effects move the gender gap in opposite directions at low and high quantiles for full-time workers implies the binary selection model would report a sign of selection that is wrong at the top; comparing censored-DR and binary-DR estimates on the same UK data is a direct check."],"forward_implications":["Researchers who currently dichotomize censored selection variables—employment, program participation, unemployment duration—can recover the full selection-sorting function at every threshold using the same binary-instrument exclusion assumptions required by standard Heckman-type models.","Wage gaps can be decomposed by worker type: the paper's UK application separates composition, wage structure, hours structure, and hours-wage sorting, finding that hours structure and sorting narrow the low-quantile gender gap and widen the high-quantile gap for full-time workers, and that selection behavior explains most of the small overtime wage gap.","The model covers continuous, discrete, and mixed outcomes, extending distributional analysis beyond the mean or median and beyond Gaussian errors, unlike quantile selection models that require continuous outcomes.","Uniform confidence bands for the sorting function obtained by multiplier bootstrap allow testing functional hypotheses such as sorting being zero, non-negative, or constant across wage quantiles.","Because the exclusion restrictions are local to a single threshold $s_0$, the same design applies at any censoring or policy cutoff, such as the 34- and 40-hour thresholds used to define full-time and overtime work."],"supporting_citations":[{"why":"Supplies the distribution regression with binary selection and the local Gaussian representation that this paper extends to a censored selection rule.","marker":"Chernozhukov et al. (2023)"},{"why":"Introduces the classical sample-selection model with censored selection and provides the baseline parametric structure that the paper generalizes.","marker":"Heckman (1974)"},{"why":"Global univalence theorem used to guarantee the nonlinear system at $s_0$ has a unique solution for $(\\nu(y), \\rho(s_0,y))$.","marker":"Gale and Nikaido (1965)"},{"why":"Provides the Z-process framework and functional delta method used to prove the functional central limit theorems and uniform inference.","marker":"Chernozhukov et al. (2013)"},{"why":"Supplies the empirical-process and bootstrap-consistency definitions used in Theorems 2 and 3.","marker":"van der Vaart and Wellner (1996)"},{"why":"Constructs the UK data and the out-of-work benefit instrument that identifies the selection equations in the application.","marker":"Blundell et al. (2003)"},{"why":"Extends the benefit-instrument construction and bounds approach that the application relies on for the outcome exclusion restriction.","marker":"Blundell et al. (2007)"},{"why":"Shows nonparametric identification in censored selection models and motivates the censored rule used here.","marker":"Fernández-Val et al. (2021)"}],"fun_headline_variants":["Censored selection model uncovers sorting at every hours threshold","Binary instrument identifies local sorting in censored regression","Local Gaussian copula reveals gendered selection into work hours","Wage-hours sorting identified for full-time and overtime work"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the sorting exclusion restriction: at the censoring point (zero weekly hours), the local correlation between latent desired hours and offered wage is the same for both values of the out-of-work benefit instrument once covariates are controlled; if this fails, the outcome distribution and every sorting parameter are unidentified and all application estimates inherit the bias.","fun_headline_variants_meta":{"raw":{"variants":["Censored selection model uncovers sorting at every hours threshold","Binary instrument identifies local sorting in censored regression","Local Gaussian copula reveals gendered selection into work hours","Wage-hours sorting identified for full-time and overtime work"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000157,"raw_usage":{"total_tokens":1264,"prompt_tokens":1029,"completion_tokens":235,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":170}},"tokens_in":645,"tokens_out":235,"duration_ms":2705,"temperature":1.0,"reasoning_tokens":170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:03:09.718138+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Using the observed data, one can solve the two-equation system at $s_0$ for each wage level $y$; if no solution exists with $\\rho(s_0,y)\\in[-1,1]$ for some $y$, the exclusion restrictions are rejected. If the instrument has more than two values, a minimum-distance test of the overidentifying restrictions directly tests Assumption 1; additionally, the estimated $\\rho_z(s,y)$ from equation (4) must keep all implied joint probabilities in $[0,1]$ across thresholds, which the paper's Remark 1 shows can fail numerically—checking whether this happens at the estimated parameters, not just at starting values, would falsify the model.","supporting_citations":[{"cited_title":"Fern \\'a ndez-Val, and S","cited_arxiv_id":null,"evidence_quote":"Supplies the distribution regression with binary selection and the local Gaussian representation that this paper extends to a censored selection rule."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Global univalence theorem used to guarantee the nonlinear system at $s_0$ has a unique solution for $(\\nu(y), \\rho(s_0,y))$."},{"cited_title":"Fern \\'a ndez-Val, and B","cited_arxiv_id":null,"evidence_quote":"Provides the Z-process framework and functional delta method used to prove the functional central limit theorems and uniform inference."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the empirical-process and bootstrap-consistency definitions used in Theorems 2 and 3."},{"cited_title":"Reed, and T","cited_arxiv_id":null,"evidence_quote":"Constructs the UK data and the out-of-work benefit instrument that identifies the selection equations in the application."},{"cited_title":"Gosling, H","cited_arxiv_id":null,"evidence_quote":"Extends the benefit-instrument construction and bounds approach that the application relies on for the outcome exclusion restriction."},{"cited_title":"van Vuuren, and F","cited_arxiv_id":null,"evidence_quote":"Shows nonparametric identification in censored selection models and motivates the censored rule used here."}],"review_version":1}