{"id":"5716c233-60e0-46db-a4dd-2190b4d318cc","arxiv_id":"2511.18550","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Selective conditional inference gives Wald tests on estimated latent panel groups a truncated chi-square null distribution, valid even without group separation.","lead":"This paper develops a way to run valid statistical tests on panel data after the data have first been split into latent groups by clustering. It matters because standard tests that ignore the uncertainty from the clustering step can reject true null hypotheses almost all the time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The exact truncated-χ² result is conditioned on a single algorithmic path, but the implementation selects the minimum-objective run over many random starts; the argmin event is not in A_TSK/A_PCR, so the theorem does not cover the implemented procedure.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing gap: the exact conditional distribution is derived for a single algorithmic path, while the implementation selects the path with minimum objective over many random starts. I agree that this is the most serious threat to the paper's central claim. The theorem's conditioning set A_TSK/A_PCR includes the full iteration path but not the argmin selection event that determines which path is reported. Because the argmin is a function of the data through the objective function values, conditioning only on the selected path omits an informative selection event. The paper's own Appendix A acknowledges the multi-start argmin procedure, so this is not an artifact of the review pipeline. The issue is addressable—one could augment the conditioning set with the random starts and the event that the chosen run has minimal objective, or prove that the argmin selection is asymptotically negligible—but the current theorems do not provide exact validity for the implemented estimator. The simulations use estimated variances and dependent errors, so they do not fill this gap either. Thus the verdict CONDITIONAL is appropriate: the theoretical logic appears viable, but the exactness claim must be restricted to the single-start, known-variance setting until the argmin selection is incorporated or proven harmless.","tokens_in":33571,"tokens_out":6533,"duration_ms":79926,"concrete_test":"Under the exact assumptions of Theorem 1 (Gaussian homoskedastic errors with known variance, no serial/spatial dependence, DGP1 with no group separation), implement the conditional TSK/PCR test in two ways: (i) using a single fixed random initialization and the paper's truncation-set formulas; (ii) using the actual multi-start implementation that selects the run with minimum objective and then applying the same formulas to the selected run's path. Compare empirical rejection frequencies conditional on the selected group path at nominal 5%. If the single-start version controls size but the multi-start argmin version over-rejects, the concern is confirmed; if both control size, the argmin selection may be benign in this setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central exactness claim (Theorems 1 and 2) conditions on A_TSK = ∩_{m=0}^M {γ^{(m)}_D = γ^{(m)}_d} (and analogously for PCR), i.e., the full iteration path of a single run of Algorithm 1 or 2. However, the actual implementation, as stated in Appendix A, 'utilize[s] a large number of random initial values and select[s] the final estimates that yield the minimum value of the objective function.' The argmin-over-starts event is a data-dependent selection event: the chosen path is the one with the smallest objective, which is informative about the data and about the Wald statistic. This event is not included in the conditioning sets A_TSK or A_PCR. Consequently, the truncated-χ² distribution derived for a fixed path does not apply to the estimator actually used in the simulations and applications. This is not a purely technical gap: when group separation is weak or absent, the minimum-objective run tends to select partitions with larger between-group separation, which is precisely the selection bias the method is designed to remove. The paper offers no theorem or formal argument showing that the argmin selection can be ignored, and the law-of-total-probability argument in the proof does not cover it because the selected path is not independent of the test statistic conditional on the observed path.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops selective conditional inference methods for linear panel data models with latent group structure. For the Two-Step K-means (TSK) and Panel Clusterwise Regression (PCR) estimators, it derives Wald-type test statistics for general linear hypotheses Rα(γ_D)=r and claims that, conditional on the estimated group path and nuisance components, these statistics follow truncated χ² distributions under Gaussian errors with known variance (Theorems 1 and 2). The appendices derive the truncation sets as unions of quadratic inequalities. Monte Carlo simulations and two empirical applications illustrate the method and argue that it controls size even when group separation fails.","tokens_in":33811,"tokens_out":6706,"duration_ms":76204,"significance":"If the stated exactness held for the implemented procedure, this would be a valuable contribution: it extends polyhedral selective inference from simple clustering settings to panel data estimators, handles general linear restrictions, and addresses group non-separation. The algebraic decompositions in Section 4.2 (Eq. (8)-(9)) and the truncation-set formulas in Propositions A.1-A.2 are nontrivial and generalize earlier work. The simulation evidence is extensive and suggests the method works well in practice. However, the paper's central exactness claim is currently tied to an idealized version of the algorithm and to known variances, while the implementation and simulations depart from these assumptions without a bridging theorem.","major_comments":[{"comment":"The implemented algorithm uses 'a large number of random initial values and select[s] the final estimates that yield the minimum value of the objective function' (Appendix A). The conditioning set A_TSK in Theorem 1 contains only the iteration path of one run, γ^(m)_D = γ^(m)_d for m=0,...,M, together with direction and nuisance components. The argmin-over-starts event is not included in A_TSK (or A_PCR). Since the identity of the minimum-objective run is data-dependent, the selected path is not a function of the data alone in the way the proof in Appendix B.3 assumes; the event that a particular path is the argmin is informative about the Wald statistic. The statement in Appendix A that random initialization needs no S^(0) conditioning addresses only the non-data-dependence of the initial labels, not the selection across starts. Thus Theorems 1-2 do not cover the estimator actually used","section":"§4.2, Theorem 1 (Eq. 10), and Appendix A"},{"comment":"Theorems 1 and 2 assume σ² and Σ are known. Section 4.4 then introduces estimated variances: the Pesaran estimator for TSK and the Driscoll-Kraay estimator for PCR, and all simulations and applications use these estimated variances. Moreover, the simulation errors are non-Gaussian (t(6) in the second half), serially correlated, and cross-sectionally dependent. No theorem or proposition establishes that the truncated-χ² result remains valid when the variance is estimated, nor that any asymptotic version holds under these general error distributions. The abstract's claim that the tests are 'asymptotically valid under general error distributions' is therefore unsupported. At minimum, the paper needs an asymptotic statement for estimated variances or a formal justification for why the known-variance conditional distribution can be used with plug-in variance estimators.","section":"§4.4 and Section 6; abstract"},{"comment":"The exact finite-sample theorems require fixed non-random regressors with identical second moments: 'Σ is nonsingular' and 'Σ = I_K' or PT_{t=1} X_it X_it' = Σ for every i (Section 4.3, Theorem 2). In the Monte Carlo design, the regressors are stochastic AR(1) processes with spatial correlation, so these assumptions are not satisfied even approximately in finite samples. While the simulations may be intended as a robustness check, the paper does not state this explicitly or provide any theoretical link. This is a gap between the formal claims and the numerical evidence: the simulations do not directly verify the exact theorem, and the asymptotic claim that would cover them is not proved.","section":"§4.2-4.3 and Section 6"}],"minor_comments":[{"comment":"Theorem 1's truncation set S_TSK includes m=0 (the initial assignments), while Appendix A says that under random initialization 'there is no need to consider S^(0)_TSK and S^(0)_PCR since the group assignments are not data-dependent.' This is inconsistent: either the conditioning set includes the initial labels or it does not. Please clarify.","section":"§4.2 and Appendix A"},{"comment":"In the definition of A_PCR, the text says 'w_T SK(bγd) that of W_T SK(bγD)' but should presumably be 'w_PCR' and 'W_PCR'. Also, Theorem 2 states 'H_PCR(eγD)|A_PCR' but the statistic uses bγD, not eγD.","section":"§4.3"},{"comment":"The text says the simulations use T∈{50,100}, but Tables 2-4 report rows for T=20 and T=50. Please reconcile the reported sample sizes.","section":"Section 6 and Tables 2-4"},{"comment":"Taylor and Tibshirani (2015a) and (2015b) appear to be the same article. If one is intended to be a different paper or a separate note, please correct the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is promising and the central idea is attractive, but the current version has a substantial gap between the theoretical exactness claim and the implemented algorithm: the argmin-over-random-starts selection is not conditioned on, and the variance estimation and non-Gaussian/dependent errors used in practice are not covered by any theorem. A major revision that either changes the implementation to match the theorem, extends the conditioning set, or proves asymptotic validity for the implemented procedure would be needed before publication. I do not see this as a fatal flaw; the algebraic machinery appears sound and the empirical results are suggestive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper is a real contribution: it extends polyhedral selective inference to TSK and PCR estimators in panel latent-group models and handles arbitrary R, not just pairwise homogeneity. The decomposition in (8)-(9) is genuinely more general than the Chen-Witten/Gao/Chen-Gao/Yun-He special cases, and the PCR conditioning set is new. The appendix proofs are detailed and, under the stated Gaussian/known-variance assumptions, the conditional truncated chi-square logic holds together. The related-literature discussion is fair, including the contemporaneous Wan et al. work.\n\nSecond, there is a load-bearing gap between the theorems and the implemented procedure. Theorems 1 and 2 condition on the full iteration path of a single run. Appendix A says the implementation uses many random initializations and keeps the run with minimum objective. That argmin event is data-dependent and is not in A_TSK or A_PCR, so no theorem covers the estimator actually used in the simulations and applications. I think the stress-test note is right: when separation is weak, the minimum-objective run tends to pick more separated partitions, which is exactly the bias the method is supposed to remove. This is fixable in principle—condition on the argmin event or prove it is ignorable—but it is not done here.\n\nOther soft spots, in proportion. The abstract promises asymptotic validity under general errors and estimated variances, but no theorem delivers that; Section 4.4 gives estimators and the simulations are encouraging, but that is not the same. The conditional GFE procedure is used in the simulations but only sketched in Section 5.2 with no theorem. The simulation text says T is in {50,100} while the tables report T=20 and 50; that looks like a leftover inconsistency. None of these kill the core idea. The central conditional-distribution argument is plausible and the proofs are broadly coherent; the gaps are in coverage of the claims, not in the algebra.\n\nWho should read it: applied econometricians doing post-clustering inference, and methodologists working on selective inference for clustering. It deserves a serious referee. My recommendation is to send it out, and ask the authors to either restrict the exactness claims to the single-start known-variance setting or add formal results for the multi-start argmin and estimated variance. The paper would be stronger if the Monte Carlo and applications were brought into line with the theorems.","headline":"Genuine extension of polyhedral selective inference to latent-group panels with arbitrary linear restrictions, but exactness is proven for a single random start while the implementation selects the minimum over many starts.","tokens_in":34371,"tokens_out":2428,"would_cite":true,"duration_ms":26340,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F12","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Wald tests for latent group panels stay valid when groups blur","keywords":["latent group panel","selective inference","post-clustering inference","K-means","Wald test","truncated chi-squared","group non-separation","panel data"],"falsifier":"Generate panels under a homogeneous DGP with no true group separation, run the implemented multi-start TSK and PCR procedures with argmin selection over many random initializations, and test a true null at the 5% level; if the rejection rate stays at 5% across large samples, the size guarantee extends to the implemented estimator, while if it exceeds the nominal level, the gap between the theorem and the implementation is real.","tokens_in":33352,"feed_emoji":"📊","tokens_out":3159,"duration_ms":35281,"temperature":0.7,"pith_summary":"This paper tackles the double-dipping problem in panel data: when the same data are used to estimate latent groups and then to test hypotheses about group-specific coefficients, standard Wald tests badly over-reject. The authors show that conditioning the test on the estimated group structure—and on a few nuisance components—yields an exact truncated chi-squared null distribution, so the test controls the selective Type I error even when the groups are not well separated in the population. The method applies to general linear restrictions, not just homogeneity tests, and works for both a two-step K-means estimator and a panel clusterwise regression estimator that estimates groups and coefficients jointly. Even when group separation does hold, the finite-sample size control improves on conventional asymptotic tests. Simulations with serial and cross-sectional dependence confirm the theory, and two empirical applications show that apparent heterogeneity often becomes insignificant once selection is accounted for.","feed_headline":"Exact selective tests for latent groups when separation fails","feed_subtitle":"Conditioning on the estimated partition gives truncated chi-squared nulls and fixes severe over-rejection in naive tests.","key_machinery":"The polyhedral conditioning method: the data are decomposed into a component along the direction of the test statistic and an independent nuisance component, so that the Wald statistic depends only on a scalar perturbation phi. Conditioning on the full iteration path of the clustering algorithm—each assignment step for each unit—defines a truncation set for phi characterized by quadratic inequalities, and the null distribution is a chi-squared truncated to that set. This set can be computed analytically for both the two-step K-means algorithm and the panel clusterwise regression algorithm.","core_discovery":"Under Gaussian homoskedastic errors with known variance, the Wald statistics for the two-step K-means and panel clusterwise regression estimators, conditional on the estimated group path and on nuisance components, follow truncated chi-squared distributions. This makes the tests exactly valid in finite samples: they control the selective Type I error rate—the probability of rejecting a true null given that the estimated groups equal the observed ones—even when group separation fails, meaning the number of groups is overspecified or the groups are not distinguishable in the population. The result holds for arbitrary linear restrictions on the group-specific coefficients, not only for homogene","pith_inferences":["The exact size guarantee in the theorems conditions on the full iteration path from a single random start, but the implemented procedures use many random starts and select the run with the minimum objective; that argmin selection event is not part of the conditioning set, so the implemented estimator's size guarantee is not directly covered.","The same conditioning template could extend to other clustering algorithms whose assignment rules admit polyhedral descriptions, such as hierarchical or spectral clustering, provided the iteration path can be characterized.","The independence lemmas rely on known variance and exact Gaussianity; with estimated variances or general error distributions, the truncated chi-squared result is only asymptotic, and the paper's simulations suggest the approximation needs reasonably large T to hold.","The approach points toward a broader principle for post-selection inference: conditioning on the algorithm's trajectory rather than just its final output can make the truncation set tractable while still delivering unconditional selective error control."],"forward_implications":["Tests remain valid without group separation, covering overspecified numbers of groups and groups that are homogeneous on some coefficients.","Even when groups are separated, the conditional tests give better finite-sample size control than conventional asymptotic Wald tests that ignore group-selection uncertainty.","Inverting the conditional tests yields selective confidence sets with coverage valid conditional on the estimated group structure.","The framework handles arbitrary linear restrictions, so it can test homogeneity of subsets of coefficients or across selected groups, not just full separation.","In empirical settings, the method can overturn naive evidence of heterogeneity: point estimates may differ across groups while formal selective tests fail to reject equality."],"fun_headline_variants":["Exact selective tests for latent groups even without separation","Panel group models? Robust inference when separation fails","Conditional inference for latent group panel models: exact tests","Selective testing for panel groups: valid even when groups overlap","Fix over-rejection in latent group inference with selective tests"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The theorems condition on the full iteration path of a single random start, but the implemented procedures instead use many random starts and select the one with the minimum objective; that argmin selection event is not part of the conditioning set.","fun_headline_variants_meta":{"raw":{"variants":["Exact selective tests for latent groups even without separation","Panel group models? Robust inference when separation fails","Conditional inference for latent group panel models: exact tests","Selective testing for panel groups: valid even when groups overlap","Fix over-rejection in latent group inference with selective tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000555,"raw_usage":{"total_tokens":2470,"prompt_tokens":725,"completion_tokens":1745,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":1666}},"tokens_in":469,"tokens_out":1745,"duration_ms":12652,"temperature":1.0,"reasoning_tokens":1666,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T20:41:54.665886+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate panels under a homogeneous DGP with no true group separation, run the implemented multi-start TSK and PCR procedures with argmin selection over many random initializations, and test a true null at the 5% level; if the rejection rate stays at 5% across large samples, the size guarantee extends to the implemented estimator, while if it exceeds the nominal level, the gap between the theorem and the implementation is real.","supporting_citations":[],"review_version":1}