{"id":"b003ff2e-3bb7-4095-af79-82c1b018b98a","arxiv_id":"2507.14621","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A selective inference test for clustered equal predictive ability that is valid when clusters are estimated by panel k-means.","lead":"This paper develops a way to test whether two forecasting methods perform equally well across groups of countries or firms, when those groups are not known in advance and must be inferred from the data. The authors' method corrects for the statistical bias that arises when the same data are used both to form the groups and to test them, and it is shown to control false positives in simulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selective p-value validity depends on independence of D and ΠZ; Lemma S.7's Euclidean-orthogonality argument fails under general cross-sectional dependence, so the truncated-χ calibration is unproven.","rationale":"The reader's weakest assumption focused on multiple random initializations, an acknowledged conjecture. I identify a more fundamental gap: the selective p-value's validity requires independence of the test statistic from the conditioning nuisance, and Lemma S.7 establishes this only under spherical covariance. Under the paper's own 'general CD' claim, heterogeneous factor loadings break that independence, so the truncated-χ calibration is not proven even for a single deterministic initialization. This is a load-bearing concern because Theorem 3(a) inherits the validity of the pairwise selective p-values. The paper's simulations use equal factor loadings, so they cannot detect the failure. A targeted simulation with loadings split by sign would isolate the issue. The appropriate verdict remains CONDITIONAL, matching the reader, because the paper is promising but the core proof needs either an additional assumption (e.g., factor loadings uncorrelated with cluster contrasts) or a covariance-whitened selective inference derivation.","tokens_in":43413,"tokens_out":21164,"duration_ms":255355,"concrete_test":"Simulate under H0 with K=2 known, N=80, T=200, and Z_{it} = λ_i F_t + e_{it}, where F_t and e_{it} are iid N(0,1) and λ_i = 1 for i ≤ N/2, λ_i = -1 for i > N/2, so Kmeans separates units by loading. Run the selective inference p-value from Proposition 1 with a single fixed initialization (so multiple starts are irrelevant) and compute rejection frequency at q=0.05 over 1000 replications. If the rejection rate exceeds 0.10, the independence assumption in Lemma S.7 fails under strong cross-sectional dependence; if it stays near 0.05, the concern is mitigated in this design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 1 and Theorem 1(a) rest on Lemma S.7, which asserts that Π_{k,g}Z, D_{k,g}(C), and J_Z are asymptotically pairwise independent. The proof of Lemma S.7 claims that Π_{k,g}ν_{k,g}=0 and 'properties of the matrix normal distribution' give independence of Z'ν_{k,g} and Π_{k,g}Z. This is only valid when vec(Z) has covariance proportional to the identity. Assumptions G1–G3 allow arbitrary autocorrelation and cross-sectional dependence, including common factors with heterogeneous loadings. In that case the Euclidean orthogonal complement and the contrast are correlated through the common factors, and a CLT for the scalar D (Lemma S.6) does not imply asymptotic independence from the high-dimensional projection used in the conditioning event. Without this independence, the conditional distribution of D given the clustering event and ΠZ = Πz is not the unconditional χ_P law truncated to T; the truncation set (14) ignores the dependence of selection on ΠZ, so the p-value in Proposition 1 may not control the selective Type I error. This gap affects the central claim even with a single deterministic initialization. The Monte Carlo design in Section 5 uses a common factor with the same loading λ for every unit, so the cluster-means contrast cancels the factor and the problematic dependence does not arise; the paper's claim of validity under 'arbitrary forms and strengths of cross-sectional dependence' is therefore not exercised. The paper does not flag this as an open assumption, unlike the multiple-initializations issue in OA Section D.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a selective-inference framework for testing clustered equal predictive ability (C-EPA) when clusters are unknown and estimated by Panel Kmeans. The authors construct a Wald-type root statistic for pairwise equality of cluster means, claim that its asymptotic distribution conditional on the estimated clustering is a truncated χ variable, characterize the truncation set through quadratic inequalities, and combine the resulting pairwise p-values with an overall-EPA p-value using a dependence-robust merging function. Monte Carlo simulations compare the selective test with predetermined-cluster, naive, and split-sample tests, and an exchange-rate forecasting application illustrates the method.","tokens_in":43746,"tokens_out":10469,"duration_ms":131901,"significance":"If the central calibration result is correct, the paper makes a useful contribution: it offers a full-sample alternative to sample splitting for post-clustering forecast evaluation, accommodates conditioning variables, and is designed to be robust to autocorrelation and both weak and strong cross-sectional dependence. The p-value combination strategy is well motivated and the simulation study is extensive; the availability of replication code and R packages is a further strength. However, the validity of the selective p-values rests on an independence claim in Lemma S.7 that is not established under the paper's general dependence assumptions, and the implementation's reliance on random initializations is acknowledged as an unresolved issue. These are load-bearing gaps for the main theorem.","major_comments":[{"comment":"The proof claims that Z'ν_{k,g} and Π_{k,g}Z are independent because Π_{k,g}ν_{k,g}=0 and by 'properties of the matrix normal distribution.' This is valid only when vec(Z) has scalar covariance, or more generally when ν_{k,g} is an eigenvector of Cov(Z). Assumptions G1–G3 allow arbitrary autocorrelation and cross-sectional dependence. For example, if Z_{it}=λ_i f_t+v_{it} with heterogeneous loadings λ_i, the scalar Z'ν_{k,g} contains the factor average weighted by the cluster loading difference, while Π_{k,g}Z retains the factor component because the factor vector is not orthogonal to the complement unless λ_i is constant across units. The two quantities are then correlated, so the conditional law of D_{k,g} given Π_{k,g}Z=π is not the unconditional χ_P law truncated to the set T. This invalidates the derivation of the p-value in Proposition 1 and the proof of Theorem 1(a) under the stated general dependence assumptions. The Monte Carlo design in Section 5.1 uses a common factor with the same loading λ for all units, so the contrast D_{k,g} is orthogonal to the factor and the problematic dependence is not exercised.","section":"Online Appendix H.3, Lemma S.7"},{"comment":"The conditioning event in Definition 1 and the truncation set T in (14) require that the Panel Kmeans output, including the sequence of assignments at every iteration, be a deterministic function of the data. Algorithm 1 as stated is not deterministic: it starts from random initializations, and Section 5.1 reports that all tests use 10 random initializations. Section D explicitly states that 'does using multiple initializations require additional conditioning? Intuitively, no' and that 'a formal treatment is left for future work, but our simulations support this conjecture.' With random starts, the realized output is not a deterministic function of Z, so the event ∩_{i=1}^N {k_i(Z)=k_i(z)} is not well-defined for the implemented procedure. This is a load-bearing gap because the reported simulations and the provided R package rely on multiple initializations, yet the theoretical guarantee in Theorem 1(a) applies only to a deterministic clustering map.","section":"Section 3.1, Definition 1; Section I, Algorithm 1; Online Appendix D"},{"comment":"The proof asserts that after noting a linear combination inherits mixing properties, 'Corollary 2.2 of Phillips & Durlauf (1986) applies, yielding a multivariate invariance principle.' That citation concerns time-series invariance principles for integrated processes and does not by itself cover N^{1-ε}-scaled cross-sectional averages under strong cross-sectional dependence, which is the case ε=1 highlighted in the paper. The conditions needed for the Gaussian limit with the stated normalizing matrix, especially under factor structures with nonzero average loadings, are not verified. Please provide a self-contained argument or a precise theorem that covers the strong-dependence case, or state the additional assumptions required for Lemma 1(b).","section":"Online Appendix H.1, Lemma 1(b)"}],"minor_comments":[{"comment":"The statement 'lim sup p(F_SI,r)≤q' appears to omit the probability operator; it should read 'lim sup P[p(F_SI,r)≤q]≤q'.","section":"Theorem 3(a)"},{"comment":"The phrase 'for all k,g∈{2,...,K}, k≠g' should presumably be 'for all k,g∈{1,...,K}, k≠g'.","section":"Online Appendix H.3, Lemma S.6"},{"comment":"The notation for the covariance estimator is inconsistent: equation (12) uses S-hat^{-1/2}_{k,g}(C) in J_Z while the main text and equation (13) use Σ-hat^{-1/2}_{k,g}(C); please use a single symbol.","section":"Equations (12)–(13)"},{"comment":"The input line 'Z=(Z'_{11},Z'_{12}...,Z')′' appears truncated; it should be 'Z=(Z'_{11},Z'_{12},...,Z'_{NT})′'.","section":"Section I, Algorithm 1"},{"comment":"There is a typo: 'Galatararay University' should be 'Galatasaray University'.","section":"Acknowledgements"}],"recommendation":"major_revision","confidential_remarks":"The main blocker is Lemma S.7: the selective-p-value calibration is not proven under the general dependence structure advertised in the paper. The simulations sidestep the issue by using a common factor with identical loadings. I would encourage the editor to request a substantive revision in which the authors either prove the required independence under a well-specified class of dependence structures, restrict the theory to a setting where it holds (e.g., Gaussian or conditionally spherical data), or redesign the conditioning so that the nuisance statistic is independent of the test statistic by construction. The self-acknowledged open point about multiple random initializations in Online Appendix D should also be resolved before the method is claimed to be valid for the implemented algorithm."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about, and a serious candidate for peer review, but the central proof has a gap that the authors have not acknowledged. The extension of selective inference from iid k-means to panel k-means is genuinely new, and the C-EPA testing framework is useful. They also ship working code and replication data, which is more than most.\n\nThe weak spot is Lemma S.7. The proof claims ΠZ, the test statistic, and the direction are asymptotically independent because Πν=0 and 'properties of the matrix normal distribution.' That step requires the covariance of vec(Z) to be scalar. Under G1–G3, which allow arbitrary autocorrelation and cross-sectional dependence, the projection and the contrast can be correlated through common factors. Without that independence, the truncation set in (14) is not the right conditioning, and the p-value in Proposition 1 may not control selective Type I error. I do not see this as a trivial technicality; it is the load-bearing step for Theorem 1 and Theorem 3.\n\nThe simulations don't help here. The Monte Carlo DGP uses a common factor with the same loading on every unit, so the contrast between cluster means cancels the factor. That design avoids the problematic dependence rather than testing it. The paper claims validity under 'arbitrary forms and strengths of cross-sectional dependence,' but the evidence doesn't cover the case where the Lemma S.7 argument fails.\n\nThe multiple-initialization issue is real and the authors are honest about it: they say a formal treatment is future work, and the proof of Proposition S.2 relies on a unique output. In practice they run 10 or 10000 starts. If the clustering depends on the seed, the conditioning event in Definition 1 is not a function of the data, so the selective p-value is not well-defined. They hand-wave that it controls error uniformly over initial partitions, but that is not a proof.\n\nEverything else—the p-value merging, the OS variance estimator, the empirical application—is solid and clearly presented. The paper deserves a serious referee. My recommendation: send it to peer review, but ask for a major revision that either proves Lemma S.7 under the stated dependence assumptions or narrows the claims, and that addresses the initialization conditioning. As it stands, the central result is plausible but not proven.","headline":"The selective-inference extension to panel k-means is genuinely new and the package is a plus, but the key independence lemma is not proven under general cross-sectional dependence and the simulations dodge that case.","tokens_in":44260,"tokens_out":3539,"would_cite":false,"duration_ms":47373,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A selective inference procedure makes testing clustered equal predictive ability valid when clusters are estimated from the data.","keywords":["equal predictive ability","panel data","selective inference","Panel Kmeans","post-selection inference","p-value combination","forecast evaluation","clustered heterogeneity"],"falsifier":"Rejection rates above the nominal level in a Monte Carlo design where the null holds and the same data set is clustered from many random starting points, with different seeds producing different final clusterings, would directly contradict the claimed uniform Type I error control. A concrete version: simulate the paper's AR(1) panel with \\((\\psi_1,\\psi_2,\\psi_3)=(0,0,0)\\), run Algorithm 1 with two seeds that yield different assignment paths, and check whether the selective p-values for the same pair of clusters differ or whether the empirical rejection rate at \\(q=0.05\\) exceeds 0.05.","tokens_in":43216,"feed_emoji":"📈","tokens_out":10388,"duration_ms":99487,"temperature":0.7,"pith_summary":"The paper aims to make it possible to test whether two forecasts are equally good on average within each of several clusters, when the clusters themselves are unknown and estimated from the same data by Panel Kmeans. That setup creates a double-dipping problem, and the paper argues that the right fix is selective conditional inference rather than sample splitting. The authors derive the asymptotic distribution of the square root of a Wald statistic for pairwise cluster equality conditional on the full sequence of cluster assignments produced by the algorithm, and show it is a truncated \\(\\chi_P\\) variable with a truncation set characterized by quadratic inequalities. They combine the resulting p-values for all pairs and for overall equal predictive ability using a merging function that controls family-wise Type I error under arbitrary dependence. The claimed consequence is that a researcher can test clustered equal predictive ability with unknown clusters and obtain asymptotically valid inference, supported by simulations and an exchange-rate forecasting application.","feed_headline":"Selective inference tames double-dipping in clustered forecast tests","feed_subtitle":"The test conditions on the full clustering path and merges pairwise p-values, keeping false positives under control.","key_machinery":"The key object is the selective p-value \\(p[D_{k,g}(\\hat C)]=1-F_{\\chi_P}[d_{k,g}(\\hat C);\\mathcal T]\\), where \\(F_{\\chi_P}(\\cdot;\\mathcal T)\\) is the CDF of a \\(\\chi_P\\) variable truncated to the set \\(\\mathcal T\\). The data are decomposed as \\(Z=\\hat\\Pi_{k,g}Z+D_{k,g}(\\hat C)$T^{{-1/2}}$(\\hat\\nu_{k,g}/\\|\\hat\\nu_{k,g}\\|_2)\\hat J_Z'\\hat\\Sigma_{k,g}^{1/2}(\\hat C)\\), separating the component that sets the statistic from components held fixed by conditioning. The conditioning event includes the cluster assignment at every iteration \\($k_i^{{(m)}}$(Z)=$k_i^{{(m)}}$(z)\\), which makes the truncation region a nested intersection of quadratic inequalities in the scalar \\(\\phi\\) that indexes the perturbation \\(z(\\phi)\\). This construction converts an intractable conditional distribution into a one-dimensional truncated chi-square calculation, and the same p-values are then aggregated with the generalized-mean merging function \\(F_{SI,r}\\).","core_discovery":"The central claim is that post-clustering inference on forecast loss differentials does not require sample splitting or known cluster memberships. Conditional on the event that every unit is assigned to the same cluster at every iteration of the Panel Kmeans algorithm, and conditional on nuisance directions, \\(D_{k,g}(\\hat C)\\), the square root of the pairwise Wald statistic, converges to a \\(\\chi_P\\) distribution truncated to the set \\(\\mathcal T=\\{\\phi\\ge 0:\\cap_{m=1}^M\\cap_{i=1}^N \\{$k_i^{{(m)}}$(z(\\phi))=$k_i^{{(m)}}$(z)\\}\\}\\). The perturbation \\(z(\\phi)\\) moves the two estimated clusters toward or away from each other along the test direction, so the truncation set records which perturbations keep the clustering path unchanged. These inequalities are quadratic in \\(\\phi\\), making the selective p-value computable. Combining the \\(n_p\\) pairwise p-values with the O-EPA p-value through the calibrated M-family merging function yields \\(F_{SI,r}\\), and the paper's Theorem 3 states that under the null \\(\\limsup P[p(F_{SI,r})\\le q]\\le q\\) for every \\(q\\in(0,1)\\). Under well-separated alternatives with the true number of clusters, the test is consistent.","pith_inferences":["Editorial inference: the same conditioning-on-the-algorithm-path idea should transfer to other iterative clustering algorithms, such as hierarchical or spectral clustering, with the quadratic truncation sets replaced by the appropriate assignment rules.","Editorial inference: the paper's conjecture that multiple random initializations preserve validity is untested in the formal theory; a proof would likely require conditioning on the full initialization path or establishing uniqueness of the algorithm's output with probability one.","Editorial inference: the robustness simulation in the appendix suggests the selective test, unlike split-sample tests, keeps size control when cluster means change halfway through the sample, which points toward a formal theory of selective inference under structural breaks.","Editorial inference: the merging-function step is portable; any panel forecast comparison producing several dependent p-values on the same data could reuse the calibrated generalized-mean combination to preserve family-wise error control."],"forward_implications":["A researcher can test clustered equal predictive ability with unknown clusters and obtain asymptotically valid inference without discarding data in a sample split.","The tests remain valid under general autocorrelation and cross-sectional dependence, including strong common-factor dependence, because the variance estimator is HAC-based.","When the number of clusters is chosen by the paper's information criterion, no additional conditioning is needed for the selective p-values to remain valid, provided the clustering output is unique.","If the alternative is true and the clusters are well separated, the combined test rejects with probability tending to one.","In the exchange-rate application, the selective test detects cluster-specific gains for machine-learning models such as XGBoost that an aggregate O-EPA test alone would not reveal."],"supporting_citations":[{"why":"Supplies the selective inference construction for k-means clustering that the paper generalizes to Panel Kmeans and dependent panel data.","marker":"Chen & Witten (2023)"},{"why":"Defines the Panel Kmeans estimator whose iterative assignment path is the conditioning event.","marker":"Bonhomme & Manresa (2015)"},{"why":"Documents the failure of sample splitting and provides the selective p-value framework for clustering that motivates the approach.","marker":"Gao et al. (2024)"},{"why":"Gives the M-family merging function and calibration constant used to combine the dependent p-values.","marker":"Vovk & Wang (2020)"},{"why":"Provides the orthonormal-series HAC variance estimator and the F-distribution limit used for the O-EPA test.","marker":"Sun (2013)"},{"why":"Supplies the split-sample homogeneity test that is the main comparator and the post-clustering validity warning.","marker":"Patton & Weller (2023)"},{"why":"Provides the conditional predictive ability framework and testing functions that justify the moment conditions.","marker":"Giacomini & White (2006)"},{"why":"Establishes the polyhedral post-selection inference method that underpins the truncated chi-square distribution.","marker":"Lee et al. (2016)"}],"fun_headline_variants":["Clustered forecast tests get selective inference fix","No sample splitting for post-clustering forecast tests","Hypothesis tests that respect unknown clusters","Truncated chi-square p-values for data-driven clusters","Double-dipping in forecast clusters? Test with selective p"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Panel Kmeans output, including the sequence of cluster assignments at every iteration, is a deterministic function of the data; the paper leaves a formal treatment of multiple random initializations to future work.","fun_headline_variants_meta":{"raw":{"variants":["Clustered forecast tests get selective inference fix","No sample splitting for post-clustering forecast tests","Hypothesis tests that respect unknown clusters","Truncated chi-square p-values for data-driven clusters","Double-dipping in forecast clusters? Test with selective p"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1295,"prompt_tokens":956,"completion_tokens":339,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":276}},"tokens_in":572,"tokens_out":339,"duration_ms":4741,"temperature":1.0,"reasoning_tokens":276,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:52:57.851323+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rejection rates above the nominal level in a Monte Carlo design where the null holds and the same data set is clustered from many random starting points, with different seeds producing different final clusterings, would directly contradict the claimed uniform Type I error control. A concrete version: simulate the paper's AR(1) panel with \\((\\psi_1,\\psi_2,\\psi_3)=(0,0,0)\\), run Algorithm 1 with two seeds that yield different assignment paths, and check whether the selective p-values for the same pair of clusters differ or whether the empirical rejection rate at \\(q=0.05\\) exceeds 0.05.","supporting_citations":[],"review_version":1}