Pith. sign in

REVIEW 3 major objections 5 minor 23 references

Sampling models for selective inference

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read If selection never makes an outcome impossible, conditioning can proceed on data after selection.

desk verdict Useful conceptual contribution to selective inference, but Proposition 2's proof has a concrete Jacobian error and an unproved existence claim; Proposition 1 is solid. read the letter →

arxiv 2502.02213 v2 pith:THRZT3PD submitted 2025-02-04 math.ST stat.TH

classification math.STstat.TH MSC 62A0162B05
keywords selectiveinferencepost-selectionconditionalityprincipleancillaritynuisanceparametersselectionfunctionwinnersproblemM-ancillarity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

In selective inference, the parameter of interest is chosen after seeing the data, so classical probabilities and the 'condition on selection' imperative can point to different conditioning strategies. This paper argues that the ambiguity can be resolved in two ways: by tracking how information was processed at the selection stage, and by asking which statistics remain ancillary after selection. Its central results are preservation theorems: if the selection probability is positive for every possible value of the ancillary statistic, then two refined notions of ancillarity—G-ancillarity and M~-ancillarity—hold after conditioning on selection exactly when they held before. A sympathetic reader would take this as support for the Conditionality Principle in selective inference: provided selection cannot rule out any ancillary value, one may condition on statistics such as the nonselected data in the winners problem without making them informative about the selected parameter. The payoff is that post-selection inference can legitimately use the simpler conditional model in the winners problem, which is easier to compute and more robust, without sacrificing frequentist validity.

What carries the argument

The load-bearing object is the selection function p(y) together with the normalized selective density f_S(y;θ)=f(y;θ)p(y)/φ(θ), and with it the conditional selection probability φ(ψ;a)=Eθ[p(T,A)|a]. The arguments turn on two ancillary notions: G-ancillarity, in which A is complete for the nuisance parameter for each fixed interest parameter, and M~-ancillarity, in which the family of marginal densities provides a perfect fit—an approximate mode—for every observation, after some smooth reparametrization. These definitions were built for nuisance-parameter settings where classical ancillarity is too strong; the paper shows they are the ones that survive the change to the selective measure, with the factor φ(ψ;a) doing the work of reweighting completeness and mode conditions.

What would settle it

Find a selection rule with φ(ψ;a)=0 for some pair (ψ,a), for instance in the winners problem with Y_i ~ N(θ_i,1) where selection of I=1 occurs only when Y_2,...,Y_m lie in a fixed interval, and simulate the selective distribution to check whether Y_2,...,Y_m carry information about θ_1 after conditioning; Proposition 1 predicts that ancillarity can fail in exactly such cases.

Watch

Extended reading notes

Core claim

Under a selection mechanism with known selection function p(y), the selective distribution has density f_S(y;θ)=f(y;θ)p(y)/φ(θ). The paper's central discovery is that two notions of ancillarity that allow the ancillary statistic to depend on the parameter of interest are invariant under this change of measure, provided φ(ψ;a)>0 for every (ψ,a). Proposition 1 shows that A is G-ancillary—complete for the nuisance parameter χ at each fixed ψ—in the selective model if and only if it was in the non-selective model. Proposition 2 establishes the same invariance for M~-ancillarity, a variant of Barndorff-Nielsen's M-ancillarity in which a perfect fit is required only up to a smooth bijection. Consequently, conditioning on A after selection is justified by the Conditionality Principle exactly in the cases where it was justified before, and the paper applies this to the winners problem, where A is the data of non-selected groups.

Load-bearing premise

The load-bearing premise is that the conditional selection probability φ(ψ;a) is positive for every possible value of the ancillary statistic and, in Proposition 2, that a smooth bijection with a prescribed Jacobian exists.

Editorial extensions

If this is right

  • In the winners problem with normal data, conditioning on the nonselected observations A=(Y2,...,Ym) is justified: A remains G- and M~-ancillary after selection, so inference on the winner's parameter can proceed from the simpler selective model that conditions on the nonselected data.
  • For Gaussian linear regression with a fixed design and a projection parameter ψ_s, the ancillary direction A=P^T Y is preserved after selection, and T|A given selection is a truncated one-dimensional Gaussian, yielding exact post-selection inference for polyhedral selection events.
  • The equivalence is strict: if φ(ψ;a)>0 fails, ancillarity can fail after selection, so positivity of the conditional selection probability is a needed hypothesis for the preservation theorem.
  • When the selection stage processed information through the conditional distribution T|A, the selective density should be normalized by φ(ψ;a) rather than φ(θ), avoiding artificial dependence of the ancillary statistic on the parameter.
  • The preservation results apply to both complete and mode-based ancillary structures, so the Conditionality Principle can be invoked in selective inference whenever the selection probability is positive on the whole ancillary space.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An extension the paper leaves implicit: selection rules with zero selection probability on some ancillary values are common in deterministic algorithms, such as LASSO with a fixed penalty, which excludes an affine set with probability one in parts of the space; for those rules the preservation theorems do not directly apply, and conditioning on the ancillary statistic may introduce information abo
  • A testable further question is whether a version of Proposition 2 holds without the smoothness and existence assumptions on the bijection h2, perhaps by replacing M~-ancillarity with a purely measure-theoretic definition whose preservation does not require a prescribed Jacobian.
  • The paper's two principles suggest that in practice the choice between the random- and fixed-parameter regimes of Bayesian selective inference, or between models (1) and (2) in the winners problem, should be guided by whether the selection probability factors through an ancillary statistic; one could simulate both models in a selected-mean problem and compare conditional coverages to see where the
  • If the preservation theorems extend to other weak ancillarity notions such as S-ancillarity or B-ancillarity, the same conditioning arguments would justify more complex post-selection models beyond the normal examples treated here.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies sampling models for selective inference, focusing on the choice of conditioning strategy after selection. It contrasts the Conditionality Principle with the 'condition on selection' paradigm and identifies two guiding principles: conditioning on statistics that were used at the selection stage, and using refined notions of ancillarity in the selective model. The formal results are Proposition 1, which states that G-ancillarity (Godambe's completeness-based ancillarity) is preserved after conditioning on selection when the selection probability is positive everywhere, and Proposition 2, which states the same for a newly introduced M~-ancillarity. These results are used to justify conditioning on statistics such as the nonselected data in the winners problem and on the design matrix in certain regression settings. The paper also discusses fixed versus random design regression and the difference between fixed-parameter and random-parameter Bayesian regimes.

Significance. The paper addresses an important and timely question: whether post-selection conditioning can be grounded in classical principles such as ancillarity and the Conditionality Principle. Proposition 1 is cleanly proved and is a useful contribution. The conceptual discussion, including the distinction between conditioning on selection in the full generative model versus in a model that conditions on ancillary information first, is valuable. However, the new M~-ancillarity preservation result (Proposition 2) is not established as stated: the proof contains a concrete change-of-variable error and relies on an unproved existence assertion for a transformation with prescribed Jacobian. Since Proposition 2 is the paper's main new technical claim, the central result of the paper is not currently airtight. The definition of M~-ancillarity is also very flexible, which makes the preservation claim less informative than it may first appear; the paper should state clearly what additional content Proposition 2 provides beyond that definitional flexibility.

major comments (3)
  1. [Section 4, Proposition 2 proof] The proof contains a change-of-variable error that invalidates the mode-preservation argument as written. After defining Z1 = h1(A) and Z2 = h2(Z1), the selective density of Z1 is f_{S,Z1}(z1) = φ(ψ; h1^{-1}(z1))/φ(θ) f_{Z1}(z1). If h2 is chosen so that |det Dh2(z1)| = φ(ψ; h1^{-1}(z1)) |det Dh1^{-1}(z1)|, then |det Dh2^{-1}(z2)| = 1/[φ(ψ; h1^{-1}(h2^{-1}(z2))) |det Dh1^{-1}(h2^{-1}(z2))|]. Substituting gives f_{S,Z2}(z2) = f_{Z1}(h2^{-1}(z2)) / [φ(θ) |det Dh1^{-1}(h2^{-1}(z2))|], which is not proportional to f_{Z1}(h2^{-1}(z2)) unless |det Dh1^{-1}| is constant. Thus the statement that the density of Z2 given selection is proportional to f_{Z1}(h2^{-1}(z2)) is false as written. The Jacobian of h2 must be chosen to cancel only the selection probability φ(ψ; h1^{-1}(z1)), not the additional factor |det Dh1^{-1}(z1)|, for the mode-preservation inference to go through.
  2. [Section 4, Proposition 2 proof] Even with the Jacobian corrected, the proof asserts the existence of a smooth bijection h2 with a prescribed absolute Jacobian determinant on an open subset of a Euclidean space. In dimension at least two this is a nontrivial result that requires conditions from the theory of prescribed Jacobians (e.g., Moser's theorem or Dacorogna--Moser-type results), such as strict positivity, regularity, and integral compatibility. The paper neither states nor proves such conditions, nor does it give a reference. Consequently Proposition 2 is not substantiated. The authors should either provide a proof of existence under explicit conditions, restrict the proposition to settings where h2 can be constructed explicitly, or revise the claim to a conditional statement.
  3. [Section 4, Propositions 1 and 2] The theorems assume φ(ψ; a) > 0 for all possible (ψ, a), but the paper does not discuss how restrictive this is. Many common selection mechanisms use indicator selection functions, for which φ(ψ; a) may be zero on some ancillary sets (e.g., thresholding rules with p(y)=1{u(y)≤α} when the conditional probability of selection is not always positive). The paper should state which examples in Sections 3 and 4 satisfy the assumption and whether the results extend, possibly with modified conclusions, when the positivity condition fails.
minor comments (5)
  1. [Section 4, Proposition 2 statement] The statement contains a typo: 'if and only of it is' should be 'if and only if it is'.
  2. [Section 4, Definition 1] The definition of M~-ancillarity would benefit from explicitly stating the regularity of the bijection h and the domain/codomain, and from clarifying that 'perfect fit' refers to Equation (3) applied to the transformed density.
  3. [Section 4, proof of Proposition 2] The phrase 'The first part of the proof is the same as in Proposition 1' is terse; spelling out that minimal sufficiency and the χ-free conditional distribution carry over would improve readability.
  4. [Section 3, Example 3] The two displayed formulas for fS(n1, y1, ..., yn; θ) are visually identical except for the normalizer; it would help to label the models or explicitly describe which one conditions on selection after observing n1.
  5. [Section 5, Discussion] The discussion could mention the positivity assumption as a limitation and indicate directions for relaxing it, rather than only noting that some ancillarity notions are not preserved.

Circularity Check

1 steps flagged · score 5.0 of 10

Proposition 2's proof reduces the new M~-ancillarity preservation claim to an unproved and algebraically incorrect Jacobian construction, while Proposition 1 remains independent.

  1. self definitional [Section 4, Definition 1 and Proposition 2 proof]
    "h2 is a bijection with absolute Jacobian determinant φ(ψ; h−1 1 (z1))|J h−1 1 (z1)|, where |J h−1 1 (z1)| denotes the absolute Jacobian determinant of h−1 1 (z1). By construction, the density of Z2 given selection is proportional to fZ1 (h−1 2 (z1); ψ, χ)."

    The proof postulates a smooth bijection h2 whose Jacobian is chosen to absorb the selection factor φ(ψ; ·), so that the selective-model M~-ancillarity appears to follow by construction from the nonselective one. But the algebra does not close: with q1(z1)=|det Dh1^{-1}(z1)|, the selective density of Z1 is φ(ψ;h1^{-1}(z1)) f_{Z1}(z1)/φ(θ), and dividing by the posited Jacobian φ(ψ;h1^{-1}(z1)) q1(z1) leaves the residual factor 1/q1(z1), so the density of Z2 given selection is proportional to f_{Z1}/q1, not to f_{Z1}. Even if the Jacobian were corrected, the paper gives no proof that a global smooth bijection with the required prescribed Jacobian exists; in dimension at least two such existence is a nontrivial condition.

full rationale

The paper's Proposition 1 on G-ancillarity is a genuine, self-contained derivation: the proof uses only the form of the selective density and completeness under the assumption φ(ψ;a)>0, and its equivalence argument is algebraic and does not presuppose the conclusion. The same is true of the general conditioning framework in Sections 2-3, which stands independently of the M~-ancillarity result. The self-citations to García Rasines and Young (2021, 2022, 2023) are used for context, power comparisons, and randomized selection methodology, not as the load-bearing justification of the ancillarity preservation theorems, so they do not constitute circularity. The circularity concern is concentrated in Proposition 2: the new M~-ancillarity definition is introduced with maximal flexibility ('there exists a smooth bijection h'), and the proof then attempts to make the selective version true by positing another bijection h2 with a Jacobian that cancels the selection probability. This is a construction-by-fiat rather than a derivation, and the stated Jacobian in fact fails to yield the claimed proportionality unless an extra constant-Jacobian condition holds. Because the existence of h2 is asserted without proof and is essentially the content of the theorem, the paper's new ancillarity-preservation result reduces by construction to an unproved existence postulate. On that basis the overall circularity score is 5: partial circularity in the paper's newest technical claim, while the more classical Proposition 1 and the paper's broader selective-inference framework retain independent content.

Assumptions & free parameters 0 free parameters · 6 assumptions · 1 invented entities

The paper is theoretical: no data are fitted, so there are no free parameters. The load-bearing axioms are the adopted conditional sampling model, a minimal-sufficiency and ancillarity structure, the positivity assumption φ(ψ;a)>0, and an unproved existence of the transformation h2 in Proposition 2. The paper also relies on the normative Conditionality Principle.

assumptions (6)
  • domain assumption There exists a minimal sufficient statistic (T,A) whose conditional distribution T|A is free of the nuisance parameter χ.
    Invoked in Section 4 to set up Propositions 1 and 2; holds for exponential-family and location-model settings but not generally.
  • domain assumption The selective distribution is f_S(y;θ) ∝ f(y;θ)p(y), the conditional approach to selective inference.
    Adopted from Section 2; the paper works within this frequentist framework and does not justify it against randomized or unconditional alternatives.
  • domain assumption φ(ψ;a)>0 for all possible (ψ,a).
    Stated before Propositions 1 and 2; the proofs cancel this factor and use completeness or mode arguments that fail if some a has zero selection probability.
  • ad hoc to paper A smooth bijection h2 exists with absolute Jacobian determinant φ(ψ;h1^{-1}(z1))|J h1^{-1}(z1)|.
    Asserted in the proof of Proposition 2 without proof or stated regularity conditions; central to the M~-ancillarity preservation claim.
  • domain assumption The Conditionality Principle is a valid normative requirement for inference.
    Used throughout Sections 1 to 3 to justify conditioning on ancillary or near-ancillary statistics; it is a foundational principle, not derived in the paper.
  • standard math All distributions are dominated by Lebesgue or counting measure and densities are sufficiently smooth for change-of-variable arguments.
    State in the Introduction; standard regularity for the parametric framework.
invented entities (1)
  • M~-ancillarity (Definition 1)
    purpose: A modified notion of M-ancillarity allowing an arbitrary smooth bijection; designed so that ancillarity is preserved after conditioning on selection and can justify conditioning on statistics such as nonselected data.
    It is a new definition introduced in this paper. It has no separate empirical content, and its preservation result in Proposition 2 follows in large part from the built-in freedom to choose the transformation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sampling models for selective inference." pith.science (2026). https://pith.science/paper/THRZT3PD

@misc{pith2026250202213,
  author       = {Pith},
  title        = {Pith review of: Sampling models for selective inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/THRZT3PD}},
  note         = {Machine review of arXiv:2502.02213}
}
read the original abstract

This paper explores the challenges of constructing suitable inferential models in scenarios where the parameter of interest is determined in light of the data, such as regression after variable selection. Two compelling arguments for conditioning converge in this context, whose interplay can introduce ambiguity in the choice of conditioning strategy: the Conditionality Principle, from classical statistics, and the `condition on selection' paradigm, central to selective inference. We discuss two general principles that can be employed to resolve this ambiguity in some recurrent contexts. The first one refers to the consideration of how information is processed at the selection stage. The second one concerns an exploration of ancillarity in the presence of selection. We demonstrate that certain notions of ancillarity are preserved after conditioning on the selection event, supporting the application of the Conditionality Principle. We illustrate these concepts through examples and provide guidance on the adequate inferential approach in some common scenarios.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 12 canonical work pages

  1. [9]

    Biometrika 63 (2): 277–284

    Conditional likelihood and unconditional optimum estimating equations. Biometrika 63 (2): 277–284. https://doi.org/10.2307/2335620 . Godambe, V.P

  2. [13]

    JASA 0 (0): 1–12

    Data fission: Splitting a single data point. JASA 0 (0): 1–12. https://doi.org/10.1080/01621459.2023. 2270748 . Lockhart, R., J.E. Taylor, R. Tibshirani, and R. Tibshirani

  3. [15]

    arXiv:1405.3920v1

    ‘A significance test for forward stepwise model selection’. arXiv:1405.3920v1. Mandel, M. and Y. Rinott

  4. [19]

    Bootstrapping and sample splitting for high-dimensional, assumption-lean inference. Ann. Stat. 47(6): 3438–3469. https: //doi.org/10.1214/18-AOS1784 . Rosenthal, R

  5. [21]

    Selective inference with a randomized response. Ann. Stat. 46 (2): 679–710. https://doi.org/10.1214/17-AOS1564 . Tibshirani, R., A. Rinaldo, R. Tibshirani, and L. Wasserman

  6. [22]

    Uniform asymp- totic inference and the bootstrap after model selection. Ann. Stat. 46 (3): 1255–1287. https://doi.org/10.1214/17-AOS1584 . Yekutieli, D

  7. [32]

    Birnbaum, A

    https://doi.org/10.1214/ss/1056397485 . Birnbaum, A

  8. [865]

    Kiefer, J

    https://doi.org/10.1214/aos/1176343584 . Kiefer, J

Show all 23 references
  1. [1925]

    Theory of statistical estimation. Proc. Camb. Phil. Soc. 22 (5): 700–725. https://doi.org/10.1017/S0305004100009580 . Fisher, R.A. 1934a. The effect of methods of ascertainment upon the estimation of frequencies. Ann. Eugen. 6: 13–25. https://doi.org/10.1111/j.1469-1809.1934. ...

  2. [1962]

    On the foundations of statistical inference. J. Am. Stat. Assoc. 57: 269–326. https://doi.org/10.1080/01621459.1962.10480660 . Brown, L.D

  3. [1976]

    Biometrika 63 (3): 567–571

    Noninformation. Biometrika 63 (3): 567–571. https: //doi.org/10.1093/biomet/63.3.567 . Barndorff-Nielsen, O.E. and D.R. Cox

  4. [1977]

    Conditional confidence statements and confidence estimators. J. Am. Stat. Assoc. 72 (360a): 789–808. https://doi.org/10.1080/01621459.1977.10479956 . Kuffner, T.A. and G.A. Young

  5. [1979]

    Psychol Bull

    The file drawer problem and tolerance for null results. Psychol Bull. 86 (3): 638–641. https://doi.org/10.1037/0033-2909.86.3.638 . Tian, X. and J.E. Taylor

  6. [1980]

    Biometrika 67 (1): 155–162

    On sufficiency and ancillarity in the presence of a nuisance parameter. Biometrika 67 (1): 155–162. https://doi.org/10.2307/2335328 . 15 Jørgensen, B

  7. [1995]

    Controlling the false discovery rate: A practical and powerful approach to multiple testing. J. R. Statist. Soc. B (Methodologi- cal) 57 (1): 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x . Berger, J.O

  8. [2012]

    Adjusted Bayesian inference for selected parameters. J. R. Statist. Soc. B 74 (3): 515–541. https://doi.org/10.1111/j.1467-9868.2011.01016.x . Young, G.A. and R.L. Smith

  9. [2014]

    A significance test for the lasso. Ann. Stat. 42 (2): 413–468. https://doi.org/10.1214/13-AOS1175 . Loftus, J.R. and J.E. Taylor

  10. [2017]

    arXiv:1410.2597v4

    ‘Optimal inference after model selection’. arXiv:1410.2597v4. Garc ´ ıa Rasines, D. and G. Young

  11. [2018]

    Electron

    Scalable methods for Bayesian selective inference. Electron. J. Stat. 12: 2355–2400. https://doi.org/10.1214/18-EJS1452 . 16 Panigrahi, S., J.E. Taylor, and A. Weinstein

  12. [2019]

    arXiv:1901.09973v1

    Inference after black box selection. arXiv:1901.09973v1. Neufeld, A., A. Dharamshi, L. Gao, and D. Witten

  13. [2021]

    arXiv:2008.04584v2

    Bayesian selective inference: non-informative priors. arXiv:2008.04584v2. Garc ´ ıa Rasines, D. and G.A. Young

  14. [2023]

    Biometrika 110 (3): 597–614

    Splitting strategies for post-selection inference. Biometrika 110 (3): 597–614. https://doi.org/10.1093/biomet/asac070 . Godambe, V.P

  15. [2824]

    Rao, C.R

    https: //doi.org/10.1214/21-AOS2057 . Rao, C.R

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.