Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Exoplaneteers Keep Overestimating Sigma Significances

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Bayes-factor 'sigma' claims are ceilings, not exact values.

desk verdict The paper's central claim inverts the Sellke inequality: converted sigmas are lower bounds, not upper bounds, so the 'overestimation' warning is backwards. read the letter →

arxiv 2506.05392 v2 pith:LIW57Z25 submitted 2025-06-03 astro-ph.IM astro-ph.EP

classification astro-ph.IMastro-ph.EP
keywords BayesfactorsigmasignificanceSellkeinequalitydetectionexoplanetatmospheresmodelcomparisonK2-18bupperbound
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a common astronomical practice—converting Bayes factors into frequentist sigma values by inverting the Sellke et al. (2001) formula—systematically overstates detection confidence. Because the Sellke inequality is an upper bound on the Bayes factor for a given p-value, reversing it yields the most optimistic possible sigma for a given Bayes factor, with the true significance generally lower. The authors trace the practice to Benneke & Seager (2013) and show that a prominent recent '3 sigma' detection claim would more honestly be read as 'less than 3 sigma.' They urge the community to stop converting to sigmas and to report Bayes factors directly.

What carries the argument

The load-bearing object is the Sellke et al. (2001) inequality $B_{01} \geq -e\, p\log p$, which lower-bounds the Bayes factor of a precise null against a composite alternative in terms of the p-value, under assumptions of a univariate, monotonic and continuous likelihood ratio and a proper prior. Flipping this to $B_{10} \leq -1/(e\, p\log p)$ and numerically inverting through $p = \mathrm{erfc}[n_\sigma/\sqrt{2}]$ yields an upper bound on the number of sigmas; the inversion is the mechanism that turns a conservative bound into an optimistic $\sigma$ score. The argument's force comes from tracking the direction of this inequality through every step of the conversion.

What would settle it

Find a concrete retrieval-based model comparison where the likelihood ratio is non-monotonic or multivariate (violating the Sellke assumptions), compute the true frequentist significance by simulation, and show it exceeds the sigma value obtained by inverting the Sellke formula; that would break the claim that the conversion is always an upper bound. Alternatively, identify any published exoplanet detection whose actual calibrated significance is higher than the Sellke-derived ceiling.

Watch

Extended reading notes

Core claim

The central discovery is that the inversion of the Sellke et al. (2001) lower bound $B_{01} \geq -e\, p\log p$, as popularized in exoplanet spectroscopy by Benneke & Seager (2013), has been widely misread as an equality. Because the original inequality states that the null-to-alternative Bayes factor is at least a certain function of the p-value, the alternative-to-null Bayes factor $B_{10}$ is at most the reciprocal. Numerically inverting that upper bound to obtain $n_\sigma$ therefore returns the largest $\sigma$ value consistent with the Bayes factor, and the true frequentist significance will in general be smaller. The paper demonstrates the inflation with examples and applies the correction to a recent DMS/DMDS detection in K2-18 b, concluding that the practice produces overstated significance claims and should be abandoned in favor of reporting Bayes factors.

Load-bearing premise

The argument assumes that the Bayes factors produced by real atmospheric retrieval model comparisons satisfy the Sellke conditions—a precise null, a composite alternative, a univariate problem, and a monotonic continuous likelihood ratio—so that the inequality direction literally bounds the reported sigma values.

Editorial extensions

If this is right

  • Reported detection significances obtained by inverting the Sellke formula become ceilings: a Bayes factor of 21 corresponds to at most 3.0 sigma, and could be 2.8 sigma or less.
  • Prominent claims such as the 3-sigma DMS/DMDS evidence in K2-18 b should be rephrased as 'less than 3-sigma significance' if Bayes factors are quoted as the evidence.
  • The community should report Bayes factors (or odds ratios) directly, avoiding the inflation and false precision inherent in sigma conversion.
  • Alternative conversion schemes (two-tailed, one-tailed, Kass & Raftery, Schmidt et al.) are more conservative, but none is fully satisfactory; the paper argues the conversion exercise itself is ill-advised.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The inflation mechanism is not limited to exoplanet retrievals: any field that inverts the Sellke bound and quotes sigma values inherits the same optimistic bias, so papers using the 'Benneke & Seager scale' elsewhere in astronomy may be systematically overclaiming.
  • A practical fix would be to quote the Sellke-derived sigma as an explicit upper limit (e.g., '$\leq 3.0\,\sigma$') or to accompany the conversion with a calibrated simulation that estimates the true significance for the specific model comparison at hand.
  • Moving to odds ratios in public communication might actually improve public understanding, since non-specialists often grasp '17 to 1 odds' more readily than a Gaussian tail probability.
  • The same inequality-direction problem may affect other Bayesian-to-frequentist calibration tools, such as the Kass & Raftery approximation when used outside its nested and large-sample regime.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This manuscript argues that the common practice in exoplanet atmospheric studies of inverting the Sellke et al. (2001) bound to convert Bayes factors into frequentist sigma values yields an upper limit on the true significance, so that reported detections such as the 3-σ DMS claim in Madhusudhan et al. (2025) should be read as "less than 3-σ." The authors trace the practice to Benneke & Seager (2013), provide a worked example (B=21 → “at most 3.0σ”), and recommend abandoning the conversion altogether in favor of reporting Bayes factors directly.

Significance. If the paper's central claim were correct, it would provide a useful caution against a widespread practice in exoplanet atmospheric retrieval and would re-interpret several high-profile detection significances. The broader recommendation to prefer Bayes factors is reasonable and likely to be uncontroversial. However, the central mathematical claim is wrong: the inequality manipulation in Section 2 is reversed. The paper therefore does not provide the correction to community practice that it promises. The manuscript also fails to verify the Sellke et al. conditions for retrieval-based Bayes factors, but the direction error already invalidates the main conclusion.

major comments (2)
  1. [Section 2, Eq. (3)] The inversion direction is backwards. The Sellke bound is B10 ≤ U(p) with U(p) = -1/(e p log p). On the branch p < 1/e, U(p) is strictly decreasing in p: dU/dp = (log p + 1)/(e p^2 (log p)^2) < 0. Therefore B10 ≤ U(p) is equivalent to p ≤ U^{-1}(B10), which, since p = erfc(n_σ/√2) decreases with n_σ, gives n_σ ≥ n_σ(U^{-1}(B10)). The converted value is thus a lower bound on the frequentist significance, not an upper bound. The paper's example in Section 2 (“a Bayes factor of 21 corresponds, at most, to a 3.0σ”) is therefore reversed; the correct statement is that the significance is at least 3.0σ. A concrete counterexample: for H0: X~N(0,1) and H1: θ~N(0,1.92) with X|θ~N(θ,1), the observation x=3.3 gives two-tailed p=9.7×10^{-4} (3.3σ) and Bayes factor B10≈21. The Sellke bound gives U(9.7×10^{-4})≈54.7 ≥ 21, so the inversion holds, but the actual p-value corresponds to a higher (3.3σ), not lower, significance. This directly contradicts the manuscript's claim in Section 2 that “the true number of sigmas will, in general, be less.”
  2. [Section 3, Madhusudhan et al. example] The re-interpretation of the 3-σ DMS claim rests on the same inversion error. Under the correct direction, the reported Bayes factors of 17.5–68.0 in Table 2 of Madhusudhan et al. (2025) would imply p-values at most the thresholds obtained by inverting Eq. (3), hence significances at least 2.9–3.0σ; the phrase “at less than 3-σ significance” is not supported. The critique of the press release interpretation (0.3% probability by chance) is likewise misdirected, since the correct reading is that the p-value is no larger than that value. If the Sellke conditions fail for retrieval-based model comparison, the bound may not apply at all; but the paper does not verify those conditions for the specific retrieval settings of Benneke & Seager (2013) or Madhusudhan et al. (2025). Thus the paper's central warning about overestimation is unsupported in either case.
minor comments (4)
  1. [Section 2, Eq. (1)] The notation “− \(\exp p \log p\)” is at best ambiguous and likely a typographical error for \(-1/(e\,p\log p)\); as written, \(-e^{p\log p} = -p^p\) is negative and cannot be a lower bound. Please clarify the intended expression.
  2. [Section 2, text after Eq. (3)] The phrase “Equation (10) presents an upper bound on the Bayes factor; that means that a Bayes factor of for examples B10 = 21 corresponds, at most, to a 3.0 σ” is internally contradictory even under the paper's own reading, because an upper bound on B10 says nothing about an upper bound on n_σ unless monotonicity is established; the monotonicity actually gives the opposite direction.
  3. [Keywords and acknowledgments] There are minor typographical issues: “regatrding” in the acknowledgments, “Jeffrey’s scale” should be “Jeffreys scale”, and the keywords “The Princess Bride — Bayesian Blues” are unconventional and may not be suitable for a formal journal.
  4. [Figure 2] The figure is helpful but the caption does not state whether the curves assume nested models, equal prior odds, or which branch of the inverse; please add a brief description of the computational definitions underlying each curve.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claim rests on the external Sellke et al. bound, and the paper's self-citations are historical or stylistic rather than load-bearing.

full rationale

The paper's derivation chain begins with the Sellke et al. (2001) theorem, which is quoted as an external result with its assumptions explicitly stated. The inversion argument is a direct application of that theorem, not a parameter fitted to the paper's own outputs, nor a definition that presupposes the conclusion. The only self-references are Benneke & Seager (2013), cited as the historical source of the community conversion practice, and Kipping (2025), cited as a stylistic comparison regarding citation tracking; neither carries the logical weight of the central claim. The Madhusudhan et al. (2025) example is an external application used to illustrate the issue, not a fitted input. No equation in the paper reduces by construction to its own inputs, and no load-bearing premise is justified solely by a self-citation. Even if the direction of the inequality were disputed, that would be a mathematical correctness concern rather than a circularity concern.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's argument depends on the external Sellke theorem, the standard Gaussian p-to-sigma relation, the unity prior odds convention, and the unverified applicability of Sellke conditions to retrieval Bayes factors. No free parameters or invented entities appear.

assumptions (4)
  • standard math Sellke et al. (2001) lower bound B01 >= -e p log p under the stated conditions.
    External theorem, cited and quoted; not derived in this note and not fitted to data.
  • standard math Gaussian p-value to sigma relation p = erfc[n_sigma/sqrt(2)].
    Standard definition of a two-sided normal sigma value, used to translate p-values into sigmas.
  • domain assumption Prior odds Pr(H1)/Pr(H0) are taken as unity in the detection examples.
    The note follows the common convention that reported Bayes factors equal posterior odds, as in Benneke and Seager (2013) and the Madhusudhan et al. (2025) example.
  • domain assumption Bayes factors from exoplanet retrievals can be compared with the Sellke setup.
    Section 3 applies the Sellke bound to retrieval-based Bayes factors such as Madhusudhan et al. (2025) without verifying the theorem conditions, which is the main caveat of the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exoplaneteers Keep Overestimating Sigma Significances." pith.science (2026). https://pith.science/paper/LIW57Z25

@misc{pith2026250605392,
  author       = {Pith},
  title        = {Pith review of: Exoplaneteers Keep Overestimating Sigma Significances},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LIW57Z25}},
  note         = {Machine review of arXiv:2506.05392}
}
abstract

Astronomers, and in particular exoplaneteers, have a curious habit of expressing Bayes factors as frequentist sigma values. This is of course completely unnecessary and arguably rather ill-advised. Regardless, the practice is common - especially in the detection claims of chemical species within exoplanet atmospheres. The current canonical conversion strategy stems from a statistics paper from Sellke et al. (2001), who derived an upper bound on the Bayes factor between the test and null hypotheses, as a function of the $p$-value (or number of sigmas, $n_{\sigma}$). A common practice within the exoplanet atmosphere community is to numerically invert this formula, going from a Bayes factor to $n_\sigma$. This goes back to Benneke & Seager (2013) -- a highly cited paper that introduced Bayesian model comparison as a means of inferring the presence of specific chemical species -- in an attempt to calibrate the Bayes factors from their technique for a community that in 2013 was more familiar with frequentist sigma significances. However, as originally noted by Sellke et al. (2001), the conversion only provides an upper limit on $n_\sigma$, with the true value generally being lower. This can result in inflations of claimed detection significances, and this note strongly urges the community to stop converting to $n_\sigma$ at all and simply stick with Bayes factors.

Figures

Figures reproduced from arXiv: 2506.05392 by the authors.

Figure 1
Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Five schemes for converting Bayes factors into sigmas. The Sellke et al. (2001) scheme produces the most optimistic values and should be understood as the ceiling. 2 I also note that Trotta (2008) allude to this idea in their Section 4.5 [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The KPF SURFS-UP Survey I: Transmission Spectroscopy of WASP-76 b

    astro-ph.EP 2025-11 conditional novelty 6.0 of 10

    KPF transmission spectroscopy of WASP-76 b detects Fe I with a clear ingress-egress asymmetry but no measurable asymmetry for Na I or Ca II, supporting altitude-dependent circulation.

Reference graph

Works this paper leans on

17 extracted references · 3 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter doi edition editor eprint howpublished institution journal key month number organization pages publisher school series title misctitle type volume year version url label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts ...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION format.url url empty "" new.block "" url * "" * if FUNCTION format.eprint eprint empty "" archivePrefix empty "" archivePrefix "arXiv" = new.block " " eprint * " " * new.block " " eprint * " " * if if if FUNCTION format.doi doi empty "" " " doi * " " * if FUNCTION format.pid doi empty eprint empty ur...

  3. [3]

    !1A Qa

    thebibliography [1] 20pt to REFERENCES 6pt =0pt -12pt 10pt plus 3pt =0pt =0pt =1pt plus 1pt =0pt =0pt -12pt =13pt plus 1pt =20pt =13pt plus 1pt \@M =10000 =-1.0em =0pt =0pt 0pt =0pt =1.0em @enumiv\@empty 10000 10000 `\.\@m \@noitemerr \@latex@warning Empty `thebibliography' environment \@ifnextchar \@reference \@latexerr Missing key on reference command E...

  4. [4]

    2013, , 778, 153, 10.1088/0004-637X/778/2/153

    Benneke , B., & Seager , S. 2013, , 778, 153, 10.1088/0004-637X/778/2/153

  5. [5]

    M., et al

    Chatrchyan , S., Khachatryan , V., Sirunyan , A. M., et al. 2012, Physics Letters B, 716, 30, 10.1016/j.physletb.2012.08.021

  6. [6]

    M., Speagle , J

    Eadie , G. M., Speagle , J. S., Cisewski-Kehe , J., et al. 2023, arXiv e-prints, arXiv:2302.04703, 10.48550/arXiv.2302.04703

  7. [7]

    2007, , 382, 1859, 10.1111/j.1365-2966.2007.12707.x

    Gordon , C., & Trotta , R. 2007, , 382, 1859, 10.1111/j.1365-2966.2007.12707.x

  8. [8]

    Hubbard, R., & Lindsay, R. M. 2008, Theory & Psychology, 18, 69, 10.1177/0959354307086923

Show all 17 references
  1. [9]

    1939, Theory of Probability

    Jeffreys , H. 1939, Theory of Probability

  2. [10]

    1995, Journal of the American Statistical Association, 90, 773, 10.1080/01621459.1995.10476572

    Kass , R., & Raftery , A. 1995, Journal of the American Statistical Association, 90, 773, 10.1080/01621459.1995.10476572

  3. [11]

    2025, arXiv e-prints, arXiv:2504.13238, 10.48550/arXiv.2504.13238

    Kipping , D. 2025, arXiv e-prints, arXiv:2504.13238, 10.48550/arXiv.2504.13238

  4. [12]

    2025, , 983, L40, 10.3847/2041-8213/adc1c8

    Madhusudhan , N., Constantinou , S., Holmberg , M., et al. 2025, , 983, L40, 10.3847/2041-8213/adc1c8

  5. [13]

    2014, , 506, 150, 10.1038/506150a

    Nuzzo , R. 2014, , 506, 150, 10.1038/506150a

  6. [14]

    R., et al

    Radica , M., Taylor , J., Wakeford , H. R., et al. 2025, , 538, 1853, 10.1093/mnras/staf402

  7. [15]

    P., MacDonald , R

    Schmidt , S. P., MacDonald , R. J., Tsai , S.-M., et al. 2025, arXiv e-prints, arXiv:2501.18477, 10.48550/arXiv.2501.18477

  8. [16]

    J., & Berger, J

    Sellke, T., Bayarri, M. J., & Berger, J. O. 2001, The American Statistician, 55, 62. http://www.jstor.org/stable/2685531

  9. [17]

    2008, Contemporary Physics, 49, 71, 10.1080/00107510802066753

    Trotta , R. 2008, Contemporary Physics, 49, 71, 10.1080/00107510802066753

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.