REVIEW 3 major objections 5 minor 23 references
Sampling models for selective inference
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read If selection never makes an outcome impossible, conditioning can proceed on data after selection.
desk verdict Useful conceptual contribution to selective inference, but Proposition 2's proof has a concrete Jacobian error and an unproved existence claim; Proposition 1 is solid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the selection function p(y) together with the normalized selective density f_S(y;θ)=f(y;θ)p(y)/φ(θ), and with it the conditional selection probability φ(ψ;a)=Eθ[p(T,A)|a]. The arguments turn on two ancillary notions: G-ancillarity, in which A is complete for the nuisance parameter for each fixed interest parameter, and M~-ancillarity, in which the family of marginal densities provides a perfect fit—an approximate mode—for every observation, after some smooth reparametrization. These definitions were built for nuisance-parameter settings where classical ancillarity is too strong; the paper shows they are the ones that survive the change to the selective measure, with the factor φ(ψ;a) doing the work of reweighting completeness and mode conditions.
What would settle it
Find a selection rule with φ(ψ;a)=0 for some pair (ψ,a), for instance in the winners problem with Y_i ~ N(θ_i,1) where selection of I=1 occurs only when Y_2,...,Y_m lie in a fixed interval, and simulate the selective distribution to check whether Y_2,...,Y_m carry information about θ_1 after conditioning; Proposition 1 predicts that ancillarity can fail in exactly such cases.
Extended reading notes
Core claim
Under a selection mechanism with known selection function p(y), the selective distribution has density f_S(y;θ)=f(y;θ)p(y)/φ(θ). The paper's central discovery is that two notions of ancillarity that allow the ancillary statistic to depend on the parameter of interest are invariant under this change of measure, provided φ(ψ;a)>0 for every (ψ,a). Proposition 1 shows that A is G-ancillary—complete for the nuisance parameter χ at each fixed ψ—in the selective model if and only if it was in the non-selective model. Proposition 2 establishes the same invariance for M~-ancillarity, a variant of Barndorff-Nielsen's M-ancillarity in which a perfect fit is required only up to a smooth bijection. Consequently, conditioning on A after selection is justified by the Conditionality Principle exactly in the cases where it was justified before, and the paper applies this to the winners problem, where A is the data of non-selected groups.
Load-bearing premise
The load-bearing premise is that the conditional selection probability φ(ψ;a) is positive for every possible value of the ancillary statistic and, in Proposition 2, that a smooth bijection with a prescribed Jacobian exists.
Editorial extensions
If this is right
- In the winners problem with normal data, conditioning on the nonselected observations A=(Y2,...,Ym) is justified: A remains G- and M~-ancillary after selection, so inference on the winner's parameter can proceed from the simpler selective model that conditions on the nonselected data.
- For Gaussian linear regression with a fixed design and a projection parameter ψ_s, the ancillary direction A=P^T Y is preserved after selection, and T|A given selection is a truncated one-dimensional Gaussian, yielding exact post-selection inference for polyhedral selection events.
- The equivalence is strict: if φ(ψ;a)>0 fails, ancillarity can fail after selection, so positivity of the conditional selection probability is a needed hypothesis for the preservation theorem.
- When the selection stage processed information through the conditional distribution T|A, the selective density should be normalized by φ(ψ;a) rather than φ(θ), avoiding artificial dependence of the ancillary statistic on the parameter.
- The preservation results apply to both complete and mode-based ancillary structures, so the Conditionality Principle can be invoked in selective inference whenever the selection probability is positive on the whole ancillary space.
Reading between the lines
- An extension the paper leaves implicit: selection rules with zero selection probability on some ancillary values are common in deterministic algorithms, such as LASSO with a fixed penalty, which excludes an affine set with probability one in parts of the space; for those rules the preservation theorems do not directly apply, and conditioning on the ancillary statistic may introduce information abo
- A testable further question is whether a version of Proposition 2 holds without the smoothness and existence assumptions on the bijection h2, perhaps by replacing M~-ancillarity with a purely measure-theoretic definition whose preservation does not require a prescribed Jacobian.
- The paper's two principles suggest that in practice the choice between the random- and fixed-parameter regimes of Bayesian selective inference, or between models (1) and (2) in the winners problem, should be guided by whether the selection probability factors through an ancillary statistic; one could simulate both models in a selected-mean problem and compare conditional coverages to see where the
- If the preservation theorems extend to other weak ancillarity notions such as S-ancillarity or B-ancillarity, the same conditioning arguments would justify more complex post-selection models beyond the normal examples treated here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies sampling models for selective inference, focusing on the choice of conditioning strategy after selection. It contrasts the Conditionality Principle with the 'condition on selection' paradigm and identifies two guiding principles: conditioning on statistics that were used at the selection stage, and using refined notions of ancillarity in the selective model. The formal results are Proposition 1, which states that G-ancillarity (Godambe's completeness-based ancillarity) is preserved after conditioning on selection when the selection probability is positive everywhere, and Proposition 2, which states the same for a newly introduced M~-ancillarity. These results are used to justify conditioning on statistics such as the nonselected data in the winners problem and on the design matrix in certain regression settings. The paper also discusses fixed versus random design regression and the difference between fixed-parameter and random-parameter Bayesian regimes.
Significance. The paper addresses an important and timely question: whether post-selection conditioning can be grounded in classical principles such as ancillarity and the Conditionality Principle. Proposition 1 is cleanly proved and is a useful contribution. The conceptual discussion, including the distinction between conditioning on selection in the full generative model versus in a model that conditions on ancillary information first, is valuable. However, the new M~-ancillarity preservation result (Proposition 2) is not established as stated: the proof contains a concrete change-of-variable error and relies on an unproved existence assertion for a transformation with prescribed Jacobian. Since Proposition 2 is the paper's main new technical claim, the central result of the paper is not currently airtight. The definition of M~-ancillarity is also very flexible, which makes the preservation claim less informative than it may first appear; the paper should state clearly what additional content Proposition 2 provides beyond that definitional flexibility.
major comments (3)
- [Section 4, Proposition 2 proof] The proof contains a change-of-variable error that invalidates the mode-preservation argument as written. After defining Z1 = h1(A) and Z2 = h2(Z1), the selective density of Z1 is f_{S,Z1}(z1) = φ(ψ; h1^{-1}(z1))/φ(θ) f_{Z1}(z1). If h2 is chosen so that |det Dh2(z1)| = φ(ψ; h1^{-1}(z1)) |det Dh1^{-1}(z1)|, then |det Dh2^{-1}(z2)| = 1/[φ(ψ; h1^{-1}(h2^{-1}(z2))) |det Dh1^{-1}(h2^{-1}(z2))|]. Substituting gives f_{S,Z2}(z2) = f_{Z1}(h2^{-1}(z2)) / [φ(θ) |det Dh1^{-1}(h2^{-1}(z2))|], which is not proportional to f_{Z1}(h2^{-1}(z2)) unless |det Dh1^{-1}| is constant. Thus the statement that the density of Z2 given selection is proportional to f_{Z1}(h2^{-1}(z2)) is false as written. The Jacobian of h2 must be chosen to cancel only the selection probability φ(ψ; h1^{-1}(z1)), not the additional factor |det Dh1^{-1}(z1)|, for the mode-preservation inference to go through.
- [Section 4, Proposition 2 proof] Even with the Jacobian corrected, the proof asserts the existence of a smooth bijection h2 with a prescribed absolute Jacobian determinant on an open subset of a Euclidean space. In dimension at least two this is a nontrivial result that requires conditions from the theory of prescribed Jacobians (e.g., Moser's theorem or Dacorogna--Moser-type results), such as strict positivity, regularity, and integral compatibility. The paper neither states nor proves such conditions, nor does it give a reference. Consequently Proposition 2 is not substantiated. The authors should either provide a proof of existence under explicit conditions, restrict the proposition to settings where h2 can be constructed explicitly, or revise the claim to a conditional statement.
- [Section 4, Propositions 1 and 2] The theorems assume φ(ψ; a) > 0 for all possible (ψ, a), but the paper does not discuss how restrictive this is. Many common selection mechanisms use indicator selection functions, for which φ(ψ; a) may be zero on some ancillary sets (e.g., thresholding rules with p(y)=1{u(y)≤α} when the conditional probability of selection is not always positive). The paper should state which examples in Sections 3 and 4 satisfy the assumption and whether the results extend, possibly with modified conclusions, when the positivity condition fails.
minor comments (5)
- [Section 4, Proposition 2 statement] The statement contains a typo: 'if and only of it is' should be 'if and only if it is'.
- [Section 4, Definition 1] The definition of M~-ancillarity would benefit from explicitly stating the regularity of the bijection h and the domain/codomain, and from clarifying that 'perfect fit' refers to Equation (3) applied to the transformed density.
- [Section 4, proof of Proposition 2] The phrase 'The first part of the proof is the same as in Proposition 1' is terse; spelling out that minimal sufficiency and the χ-free conditional distribution carry over would improve readability.
- [Section 3, Example 3] The two displayed formulas for fS(n1, y1, ..., yn; θ) are visually identical except for the normalizer; it would help to label the models or explicitly describe which one conditions on selection after observing n1.
- [Section 5, Discussion] The discussion could mention the positivity assumption as a limitation and indicate directions for relaxing it, rather than only noting that some ancillarity notions are not preserved.
Circularity Check
Proposition 2's proof reduces the new M~-ancillarity preservation claim to an unproved and algebraically incorrect Jacobian construction, while Proposition 1 remains independent.
-
self definitional
[Section 4, Definition 1 and Proposition 2 proof]
"h2 is a bijection with absolute Jacobian determinant φ(ψ; h−1 1 (z1))|J h−1 1 (z1)|, where |J h−1 1 (z1)| denotes the absolute Jacobian determinant of h−1 1 (z1). By construction, the density of Z2 given selection is proportional to fZ1 (h−1 2 (z1); ψ, χ)."
The proof postulates a smooth bijection h2 whose Jacobian is chosen to absorb the selection factor φ(ψ; ·), so that the selective-model M~-ancillarity appears to follow by construction from the nonselective one. But the algebra does not close: with q1(z1)=|det Dh1^{-1}(z1)|, the selective density of Z1 is φ(ψ;h1^{-1}(z1)) f_{Z1}(z1)/φ(θ), and dividing by the posited Jacobian φ(ψ;h1^{-1}(z1)) q1(z1) leaves the residual factor 1/q1(z1), so the density of Z2 given selection is proportional to f_{Z1}/q1, not to f_{Z1}. Even if the Jacobian were corrected, the paper gives no proof that a global smooth bijection with the required prescribed Jacobian exists; in dimension at least two such existence is a nontrivial condition.
full rationale
The paper's Proposition 1 on G-ancillarity is a genuine, self-contained derivation: the proof uses only the form of the selective density and completeness under the assumption φ(ψ;a)>0, and its equivalence argument is algebraic and does not presuppose the conclusion. The same is true of the general conditioning framework in Sections 2-3, which stands independently of the M~-ancillarity result. The self-citations to García Rasines and Young (2021, 2022, 2023) are used for context, power comparisons, and randomized selection methodology, not as the load-bearing justification of the ancillarity preservation theorems, so they do not constitute circularity. The circularity concern is concentrated in Proposition 2: the new M~-ancillarity definition is introduced with maximal flexibility ('there exists a smooth bijection h'), and the proof then attempts to make the selective version true by positing another bijection h2 with a Jacobian that cancels the selection probability. This is a construction-by-fiat rather than a derivation, and the stated Jacobian in fact fails to yield the claimed proportionality unless an extra constant-Jacobian condition holds. Because the existence of h2 is asserted without proof and is essentially the content of the theorem, the paper's new ancillarity-preservation result reduces by construction to an unproved existence postulate. On that basis the overall circularity score is 5: partial circularity in the paper's newest technical claim, while the more classical Proposition 1 and the paper's broader selective-inference framework retain independent content.
Assumptions & free parameters
assumptions (6)
- domain assumption There exists a minimal sufficient statistic (T,A) whose conditional distribution T|A is free of the nuisance parameter χ.
- domain assumption The selective distribution is f_S(y;θ) ∝ f(y;θ)p(y), the conditional approach to selective inference.
- domain assumption φ(ψ;a)>0 for all possible (ψ,a).
- ad hoc to paper A smooth bijection h2 exists with absolute Jacobian determinant φ(ψ;h1^{-1}(z1))|J h1^{-1}(z1)|.
- domain assumption The Conditionality Principle is a valid normative requirement for inference.
- standard math All distributions are dominated by Lebesgue or counting measure and densities are sufficiently smooth for change-of-variable arguments.
invented entities (1)
-
M~-ancillarity (Definition 1)
Cite this review
Pith. "Pith review of Sampling models for selective inference." pith.science (2026). https://pith.science/paper/THRZT3PD
@misc{pith2026250202213,
author = {Pith},
title = {Pith review of: Sampling models for selective inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/THRZT3PD}},
note = {Machine review of arXiv:2502.02213}
}
read the original abstract
This paper explores the challenges of constructing suitable inferential models in scenarios where the parameter of interest is determined in light of the data, such as regression after variable selection. Two compelling arguments for conditioning converge in this context, whose interplay can introduce ambiguity in the choice of conditioning strategy: the Conditionality Principle, from classical statistics, and the `condition on selection' paradigm, central to selective inference. We discuss two general principles that can be employed to resolve this ambiguity in some recurrent contexts. The first one refers to the consideration of how information is processed at the selection stage. The second one concerns an exploration of ancillarity in the presence of selection. We demonstrate that certain notions of ancillarity are preserved after conditioning on the selection event, supporting the application of the Conditionality Principle. We illustrate these concepts through examples and provide guidance on the adequate inferential approach in some common scenarios.
Reference graph
Works this paper leans on
-
[9]
Conditional likelihood and unconditional optimum estimating equations. Biometrika 63 (2): 277–284. https://doi.org/10.2307/2335620 . Godambe, V.P
-
[13]
Data fission: Splitting a single data point. JASA 0 (0): 1–12. https://doi.org/10.1080/01621459.2023. 2270748 . Lockhart, R., J.E. Taylor, R. Tibshirani, and R. Tibshirani
-
[15]
‘A significance test for forward stepwise model selection’. arXiv:1405.3920v1. Mandel, M. and Y. Rinott
-
[19]
Bootstrapping and sample splitting for high-dimensional, assumption-lean inference. Ann. Stat. 47(6): 3438–3469. https: //doi.org/10.1214/18-AOS1784 . Rosenthal, R
-
[21]
Selective inference with a randomized response. Ann. Stat. 46 (2): 679–710. https://doi.org/10.1214/17-AOS1564 . Tibshirani, R., A. Rinaldo, R. Tibshirani, and L. Wasserman
-
[22]
Uniform asymp- totic inference and the bootstrap after model selection. Ann. Stat. 46 (3): 1255–1287. https://doi.org/10.1214/17-AOS1584 . Yekutieli, D
- [32]
- [865]
Show all 23 references
-
[1925]
Theory of statistical estimation. Proc. Camb. Phil. Soc. 22 (5): 700–725. https://doi.org/10.1017/S0305004100009580 . Fisher, R.A. 1934a. The effect of methods of ascertainment upon the estimation of frequencies. Ann. Eugen. 6: 13–25. https://doi.org/10.1111/j.1469-1809.1934. ...
-
[1962]
On the foundations of statistical inference. J. Am. Stat. Assoc. 57: 269–326. https://doi.org/10.1080/01621459.1962.10480660 . Brown, L.D
1962
-
[1976]
Biometrika 63 (3): 567–571
Noninformation. Biometrika 63 (3): 567–571. https: //doi.org/10.1093/biomet/63.3.567 . Barndorff-Nielsen, O.E. and D.R. Cox
-
[1977]
Conditional confidence statements and confidence estimators. J. Am. Stat. Assoc. 72 (360a): 789–808. https://doi.org/10.1080/01621459.1977.10479956 . Kuffner, T.A. and G.A. Young
1977
-
[1979]
Psychol Bull
The file drawer problem and tolerance for null results. Psychol Bull. 86 (3): 638–641. https://doi.org/10.1037/0033-2909.86.3.638 . Tian, X. and J.E. Taylor
-
[1980]
Biometrika 67 (1): 155–162
On sufficiency and ancillarity in the presence of a nuisance parameter. Biometrika 67 (1): 155–162. https://doi.org/10.2307/2335328 . 15 Jørgensen, B
-
[1995]
Controlling the false discovery rate: A practical and powerful approach to multiple testing. J. R. Statist. Soc. B (Methodologi- cal) 57 (1): 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x . Berger, J.O
1995
-
[2012]
Adjusted Bayesian inference for selected parameters. J. R. Statist. Soc. B 74 (3): 515–541. https://doi.org/10.1111/j.1467-9868.2011.01016.x . Young, G.A. and R.L. Smith
2011
-
[2014]
A significance test for the lasso. Ann. Stat. 42 (2): 413–468. https://doi.org/10.1214/13-AOS1175 . Loftus, J.R. and J.E. Taylor
-
[2017]
arXiv:1410.2597v4
‘Optimal inference after model selection’. arXiv:1410.2597v4. Garc ´ ıa Rasines, D. and G. Young
-
[2018]
Electron
Scalable methods for Bayesian selective inference. Electron. J. Stat. 12: 2355–2400. https://doi.org/10.1214/18-EJS1452 . 16 Panigrahi, S., J.E. Taylor, and A. Weinstein
-
[2019]
arXiv:1901.09973v1
Inference after black box selection. arXiv:1901.09973v1. Neufeld, A., A. Dharamshi, L. Gao, and D. Witten
1901 arXiv
-
[2021]
arXiv:2008.04584v2
Bayesian selective inference: non-informative priors. arXiv:2008.04584v2. Garc ´ ıa Rasines, D. and G.A. Young
2008 arXiv
-
[2023]
Biometrika 110 (3): 597–614
Splitting strategies for post-selection inference. Biometrika 110 (3): 597–614. https://doi.org/10.1093/biomet/asac070 . Godambe, V.P
- [2824]
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.