{"id":"c128db4a-de0d-4731-8ca9-dfda7fc85afd","arxiv_id":"2607.18545","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A data-adaptive exponential randomization scheme yields conditionally valid confidence intervals for top-k winners, with selection quality close to standard top-k and shorter intervals than polyhedral methods.","lead":"This paper proposes a randomized post-selection inference method that corrects the winner's curse when researchers report effects for top-k winners chosen from the same data. It produces confidence intervals valid conditional on the realized set of winners, while keeping selection quality close to ordinary top-k selection and shorter intervals than existing conditional methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Plug-in covariance in non-Gaussian applications is not covered by Theorem 5.1, so the headline nonparametric guarantee is unproved.","rationale":"The reader's weakest_assumption is exactly the gap I find: Theorem 5.1 fixes the covariance while applications plug in estimates. I agree with the CONDITIONAL verdict. I give credit where it is due: Theorem 3.5 is exact and the randomized construction is elegant; the regret calibration in Corollary 2.1 is clean. The concern is not that the plug-in step is false, only that it is unstated and unproved while carrying the abstract's claim of broad nonparametric applicability. The proof of Theorem 5.1 shows why this is nontrivial: the conditional likelihood-ratio weight F and the reconstruction map r both depend on eΣ; the fixed-eΣ standardization and derivative bounds do not automatically transfer to random eΣ. A delta-method heuristic is not enough because the selection event is a discrete conditional event and F appears in the denominator. Since the paper's contribution is precisely this flexibility, the missing plug-in theorem is load-bearing. The proposed test—conditional coverage in the non-Gaussian designs or an analytic extension—would settle whether the gap is benign.","tokens_in":35737,"tokens_out":8972,"duration_ms":105990,"concrete_test":"Re-run the Binomial and feature-importance simulations (Section 7, Figures 3 and 5) with B≥5000, reporting per-winner conditional coverage on {j∈E^o} using the implemented plug-in \\Σ; a rate below 0.95−~1.2% in the weak-separation regime would show the gap matters. Analytically, prove the missing extension: with \\Σ_n−Σ=o_p(1), show the plug-in pivot converges to the fixed-Σ pivot conditionally on {E_n=E^o}; if this requires an additional condition linking \\Σ_n to the selection event, Theorem 5.1 is not the guarantee used in practice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The exact Gaussian construction (Theorem 3.5) appears sound, but the asymptotic claim that carries the nonparametric applications is narrower than presented. Theorem 5.1 states the pivot with the true covariance blocks σ²_{j0} and Γ_{j0} fixed, and the proof—the standardization ζ_n in (22), the transfer of the ALR in Proposition B.1, and the derivative bounds in Proposition B.6—treats eΣ as a fixed matrix. Section 5's opening promises cases where Σ is consistently estimable, yet no theorem or proof supplies the plug-in version. The applications are not merely formal: Section C explicitly runs the Binomial design with \\Σ=Diag(\\π_j(1-\\π_j)), and the feature-importance and BTD implementations likewise use estimated covariance/Fisher information. Replacing fixed Σ by a consistent estimator inside a pivot is non-routine here: the selection probability Λ_{E^o} and the conditioning statistic T^⊥ both depend on Σ through the reconstruction map r, so the conditional law can shift with \\Σ. Without a uniform-in-Σ argument or a rate condition, the claimed assumption-lean, nonparametric conditional coverage is an extrapolation, not a proved result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a randomized selection rule based on a data-adaptive exponential mechanism for top-k winners, followed by conditional inference on the selected effects. In the exact Gaussian model, it defines a pivot that is the conditional CDF of the selected coordinate given the selection event and a nuisance-sufficient statistic, and proves exact conditional uniformity (Theorem 3.5). A tuning-free temperature is calibrated through a finite-sample regret bound (Corollary 2.1), and an O(pk) dynamic program is given for sampling and weight evaluation. For non-Gaussian applications, the paper states an asymptotic version (Theorem 5.1) for asymptotically linear selection statistics, and applies it to binomial A/B/n trials, Bradley–Terry–Davidson rankings, and nonparametric feature importance. Simulations compare the method to polyhedral inference, data splitting, and zoom correction on marginal coverage and interval length.","tokens_in":36094,"tokens_out":6073,"duration_ms":70067,"significance":"If the results hold as stated, the paper makes a substantive contribution: it provides exact conditional inference for winners without characterizing the selection event, a regret-based calibration of the randomization level, and efficient implementation. The Gaussian pivot argument is clean and self-contained, and the code is publicly available. The regret bound is simple and interpretable. However, the strongest advertised contribution—assumption-lean nonparametric conditional validity—rests on Theorem 5.1, and the theorem as stated does not cover the plug-in covariance estimators used in every non-Gaussian application and simulation. This is a load-bearing gap that must be addressed before the broad claims are supported.","major_comments":[{"comment":"Theorem 5.1 is stated for the pivot with the true covariance blocks σ²_{jo} and Γ_{jo} held fixed; see Definition B.4 and eq. (28). The proof, including the standardization ζ_n in (22), Proposition B.1, and the Lindeberg bounds in Corollary B.1, treats eΣ as a fixed matrix. Section 5's introduction promises cases where Σ is consistently estimable, but no theorem or proof supplies the plug-in version. This matters because Sections 6 and C use estimated covariances in all non-Gaussian applications: the binomial design uses \\hatΣ=Diag{\\hatπ_j(1-\\hatπ_j)}, and the BTD and feature-importance implementations use estimated Fisher information / covariance. Replacing Σ by \\hatΣ changes Λ_{E^o} and T^⊥ through the reconstruction map r, so the conditional law can shift. A uniform-in-Σ argument or explicit rate conditions are required. Without this, the nonparametric conditional-coverage guarantee i","section":"Section 5, Theorem 5.1"},{"comment":"The non-Gaussian simulations are the only empirical evidence for the asymptotic conditional guarantee, yet they report only marginal coverage. The method's advertised advantage over marginal baselines such as Zoom Correction is precisely conditional validity. The conditional-coverage demonstration in Figure 1(d) is for the exact Gaussian model, not for the asymptotically linear applications of Theorem 5.1. The paper should report conditional coverage for the binomial, BTD, and feature-importance designs—for example, averaged over the realized selected sets as implied by Corollary 3.1—to support the central claim that the nonparametric extension provides conditional validity.","section":"Section 7, Figures 3–5"},{"comment":"The verification of Assumption 1(i) is incomplete in the BTD and feature-importance proofs. In Proposition B.10, the remainder R_n is bounded by δ_n‖Z_n‖ + (K/√n)‖U_n‖² with δ_n = o_p(1) random, and the proof asserts sup_n E[exp{4C_R‖R_n‖}] < ∞ from sub-Gaussianity of Z_n and U_n. This does not follow without a rate condition on δ_n or an additional argument bounding the exponential moment of the product. Proposition B.11 makes a similar leap, treating boundedness of exponential moments of the two displayed terms as sufficient for the remainder. These steps are needed for the ALR condition and should be made rigorous or replaced by explicit sufficient conditions.","section":"Appendix B.5, Propositions B.10/B.11"}],"minor_comments":[{"comment":"The sentence 'Here the term If no hold-out dataset is available...' is incomplete; it should introduce the cross-fitting alternative properly.","section":"Section 6.3"},{"comment":"The captions of Figures 2–5 say 'selection equality' where 'selection quality' is meant.","section":"Figure captions"},{"comment":"The notation Λ_{E^o}(u1,u2) is used before the standardized version F(v) is introduced in Appendix B; a short pointer in Section 3 would improve readability.","section":"Section 3.3, Definition 3.1"},{"comment":"Assumption 1(i) defines C_R via c_f^{(1)} which is only fully defined in Proposition B.6 in Appendix B. This is a minor organizational issue but would benefit from a forward reference near Assumption 1.","section":"Section 5, Assumption 1"}],"recommendation":"major_revision","confidential_remarks":"The plug-in covariance gap is substantial but potentially fixable within the manuscript's scope; the exact Gaussian pivot, the regret calibration, and the implementation are solid. If the authors can supply a rigorous theorem for plug-in covariance under explicit rate conditions—or alternatively restrict the asymptotic claim to known covariance—the paper could be acceptable. I do not see this as a reject, because the core mechanism is sound and the missing piece is a clearly identifiable extension rather than an internal contradiction. A second concern for the editor: the non-Gaussian simulations should be revised to evaluate conditional coverage, since the current empirical section sidesteps the paper's central guarantee."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the thing to know: this is a real contribution to conditional post-selection inference, and the exact Gaussian part is rock solid. The pivot in Theorem 3.5 is the conditional CDF under the exponential mechanism, so validity follows from the probability integral transform—no fitting, no circularity. The regret-budget calibration (Corollary 2.1) is a nice way to set the temperature, and the O(pk) dynamic programming for sampling and weights makes the method actually usable. Credit where due: the exponential mechanism idea is inherited from Bakshi–Panigrahi and Wu et al., but the adaptive scaling, regret analysis, and the asymptotic framework are new.\n\nThe soft spot is real and it is exactly where the stress-test note lands. Theorem 5.1 is stated with the true covariance blocks fixed, but every non-Gaussian application plugs in an estimated covariance (binomial \\hat\\Sigma, empirical Fisher information, etc.). Replacing \\Sigma by a consistent estimator inside the pivot is not a routine Slutsky step, because the selection weights and the conditioning statistic both depend on \\Sigma through the reconstruction map. The paper offers no uniform-in-\\Sigma argument or rate condition. So the broad nonparametric guarantee claimed in the abstract and Section 5 is, as written, an extrapolation. I don't think the proof is doomed—the machinery looks plausible—but the theorem as stated does not cover the applications.\n\nThe other issue is empirical. The non-Gaussian simulations report marginal coverage and interval length, but not the conditional coverage that is the entire point of the method. Figure 1 does show conditional coverage, but only in the Gaussian decoy design. If you are claiming conditional validity in binomial and BTD settings, you should be able to demonstrate it. Their absence leaves the central empirical claim unexamined.\n\nThe citation pattern is fine; the earlier exponential-mechanism papers are credited, and the new pieces are genuinely new. The exact Gaussian construction alone is enough to warrant peer review. I'd send this to a strong statistics journal and ask the referees to demand either a plug-in theorem with conditions or an explicit limitation, plus conditional-coverage simulations for the non-Gaussian designs. It's a good paper with a specific, fixable gap.","headline":"Elegant exact Gaussian pivot and regret-calibrated randomization, but the non-Gaussian asymptotics rest on an unproved plug-in step and the central conditional guarantee goes unchecked in simulations.","tokens_in":36482,"tokens_out":2598,"would_cite":true,"duration_ms":29392,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a data-adaptive exponential randomization scheme for selecting top-k winners yields inference with exact conditional coverage in Gaussian models and asymptotically exact conditional coverage for asymptotically linear","keywords":["post-selection inference","conditional coverage","winner's curse","exponential mechanism","randomized selection","top-k selection","asymptotically linear statistics","nonparametric inference"],"falsifier":"In a binomial A/B/n design with small n, apply the method with plug-in covariance and measure, over many replications, the empirical coverage of the 95% interval conditional on a fixed realized winner set E^o; if coverage does not converge to 0.95 as n grows, the nonparametric reach claimed in the paper is unsupported.","tokens_in":35697,"feed_emoji":"🎯","tokens_out":5955,"duration_ms":62405,"temperature":0.7,"pith_summary":"Researchers often report estimates for the best-performing treatments, models, or features, but naive estimates suffer from the winner's curse: they are systematically too optimistic. This paper proposes a randomized selection rule based on an exponential mechanism with a data-adaptive temperature, so the selected set stays close to the standard top-k set. Conditioning on the realized winners and on nuisance statistics, the selection probability has a closed form, which yields a pivot whose distribution is exactly uniform in Gaussian models and asymptotically uniform for asymptotically linear statistics. Consequently, the method gives per-winner conditional coverage at the nominal level, with confidence intervals that are shorter and more stable than those from the polyhedral method and comparable to data splitting without sacrificing selection quality. The approach covers nonparametric settings such as binary-outcome trials, paired-comparison rankings, and black-box feature importance.","feed_headline":"Softmax selection yields conditionally valid winner intervals","feed_subtitle":"A regret-tuned randomization keeps top-k selection quality and gives shorter intervals, beyond Gaussian data.","key_machinery":"The key object is the exponential mechanism over all k-subsets of [p], which selects the winner set E with probability proportional to exp{s_E(T)/(tau * sigma_s(T))}, where s_E is an additive score, sigma_s is the dispersion of scores across subsets, and tau is a temperature parameter. Scaling by sigma_s makes tau scale-free, and the paper derives a tuning-free choice tau = (log |E_k|)^{-1} q that bounds the expected standardized regret by q. The selection weight Lambda_{E^o}(u1, u2) equals the conditional selection probability P(E=E^o | T = r(u1,u2)), and because the mechanism is a softmax, this weight is available in closed form. The conditional density of the target statistic T_{j0} given","core_discovery":"The central claim, stated as Theorem 3.5, is that for any winner j0 in the observed selected set E^o, the pivot Pivot_{mu_{j0}}^{E^o}(T_{j0}; T^perp_{j0}, sigma^2_{j0}) is exactly Uniform(0,1) conditional on {E = E^o} in the Gaussian model T ~ N(mu, Sigma). Corollary 3.1 turns this into exact 1-alpha conditional coverage for the interval obtained by inverting the pivot, along with exact p-values. Theorem 5.1 extends the claim to selection statistics with an asymptotic linear representation: the same pivot, computed with the true covariance blocks sigma^2_{j0} and Gamma_{j0}, converges in distribution to Uniform(0,1) conditional on the realized selection. The method therefore provides conditi","pith_inferences":["If the plug-in covariance estimator preserves the asymptotic uniformity of the pivot—which the paper assumes in its applications but does not prove—the method would provide fully rigorous nonparametric conditional inference; currently that step is a gap between Theorem 5.1 and the implemented procedure.","The closed-form selection weight suggests the same conditioning strategy could be applied to other randomized selection rules, such as additive noise or bootstrap perturbation, whenever the conditional selection probability is tractable.","A natural extension is to construct simultaneous conditional intervals for the full vector of selected winners, or to test comparisons such as winner versus runner-up, by forming joint pivots from the shared conditioning set.","The finite-sample behavior with estimated covariance is an empirical question; simulations in the paper show nominal coverage, but a formal higher-order analysis or a Berry–Esseen-type bound would clarify when the approximation is reliable."],"forward_implications":["Intervals for each selected winner are conditionally valid at level 1-alpha, meaning they answer the question 'how large is this particular winner's true effect?' even though the winner was chosen from the same data.","The method applies beyond Gaussian data to any setting where the selection statistics are asymptotically linear, including binomial A/B/n trials, Bradley–Terry–Davidson rankings, and feature-importance measures from black-box models.","The temperature parameter is set automatically from a user-specified regret budget q, so the randomized rule's selection quality is provably close to the top-k rule.","Conditional intervals adapt to the strength of the signal: they become shorter as the gap between the top-k and the rest grows, unlike fixed-width marginal intervals.","The same pivot gives valid p-values and a bias-corrected point estimate for each selected winner."],"fun_headline_variants":["Randomized selection yields shorter valid winner intervals","Conditional inference for winners with shorter intervals","Adaptive randomization gives conditionally valid winner estimates","Flexible winner inference with exact conditional coverage","Softmax selection shortens winner intervals with valid coverage"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The asymptotic validity of the pivot is proven only when the true covariance of the selection statistics is used; the paper's applications plug in estimated covariances, and no theorem shows the plug-in pivot retains its conditional uniform limit.","fun_headline_variants_meta":{"raw":{"variants":["Randomized selection yields shorter valid winner intervals","Conditional inference for winners with shorter intervals","Adaptive randomization gives conditionally valid winner estimates","Flexible winner inference with exact conditional coverage","Softmax selection shortens winner intervals with valid coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1272,"prompt_tokens":697,"completion_tokens":575,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":516}},"tokens_in":441,"tokens_out":575,"duration_ms":6259,"temperature":1.0,"reasoning_tokens":516,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:05:06.995021+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a binomial A/B/n design with small n, apply the method with plug-in covariance and measure, over many replications, the empirical coverage of the 95% interval conditional on a fixed realized winner set E^o; if coverage does not converge to 0.95 as n grows, the nonparametric reach claimed in the paper is unsupported.","supporting_citations":[],"review_version":1}