REVIEW 2 major objections 4 minor 1 cited by
Exoplaneteers Keep Overestimating Sigma Significances
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Bayes-factor 'sigma' claims are ceilings, not exact values.
desk verdict The paper's central claim inverts the Sellke inequality: converted sigmas are lower bounds, not upper bounds, so the 'overestimation' warning is backwards. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Sellke et al. (2001) inequality $B_{01} \geq -e\, p\log p$, which lower-bounds the Bayes factor of a precise null against a composite alternative in terms of the p-value, under assumptions of a univariate, monotonic and continuous likelihood ratio and a proper prior. Flipping this to $B_{10} \leq -1/(e\, p\log p)$ and numerically inverting through $p = \mathrm{erfc}[n_\sigma/\sqrt{2}]$ yields an upper bound on the number of sigmas; the inversion is the mechanism that turns a conservative bound into an optimistic $\sigma$ score. The argument's force comes from tracking the direction of this inequality through every step of the conversion.
What would settle it
Find a concrete retrieval-based model comparison where the likelihood ratio is non-monotonic or multivariate (violating the Sellke assumptions), compute the true frequentist significance by simulation, and show it exceeds the sigma value obtained by inverting the Sellke formula; that would break the claim that the conversion is always an upper bound. Alternatively, identify any published exoplanet detection whose actual calibrated significance is higher than the Sellke-derived ceiling.
Extended reading notes
Core claim
The central discovery is that the inversion of the Sellke et al. (2001) lower bound $B_{01} \geq -e\, p\log p$, as popularized in exoplanet spectroscopy by Benneke & Seager (2013), has been widely misread as an equality. Because the original inequality states that the null-to-alternative Bayes factor is at least a certain function of the p-value, the alternative-to-null Bayes factor $B_{10}$ is at most the reciprocal. Numerically inverting that upper bound to obtain $n_\sigma$ therefore returns the largest $\sigma$ value consistent with the Bayes factor, and the true frequentist significance will in general be smaller. The paper demonstrates the inflation with examples and applies the correction to a recent DMS/DMDS detection in K2-18 b, concluding that the practice produces overstated significance claims and should be abandoned in favor of reporting Bayes factors.
Load-bearing premise
The argument assumes that the Bayes factors produced by real atmospheric retrieval model comparisons satisfy the Sellke conditions—a precise null, a composite alternative, a univariate problem, and a monotonic continuous likelihood ratio—so that the inequality direction literally bounds the reported sigma values.
Editorial extensions
If this is right
- Reported detection significances obtained by inverting the Sellke formula become ceilings: a Bayes factor of 21 corresponds to at most 3.0 sigma, and could be 2.8 sigma or less.
- Prominent claims such as the 3-sigma DMS/DMDS evidence in K2-18 b should be rephrased as 'less than 3-sigma significance' if Bayes factors are quoted as the evidence.
- The community should report Bayes factors (or odds ratios) directly, avoiding the inflation and false precision inherent in sigma conversion.
- Alternative conversion schemes (two-tailed, one-tailed, Kass & Raftery, Schmidt et al.) are more conservative, but none is fully satisfactory; the paper argues the conversion exercise itself is ill-advised.
Reading between the lines
- The inflation mechanism is not limited to exoplanet retrievals: any field that inverts the Sellke bound and quotes sigma values inherits the same optimistic bias, so papers using the 'Benneke & Seager scale' elsewhere in astronomy may be systematically overclaiming.
- A practical fix would be to quote the Sellke-derived sigma as an explicit upper limit (e.g., '$\leq 3.0\,\sigma$') or to accompany the conversion with a calibrated simulation that estimates the true significance for the specific model comparison at hand.
- Moving to odds ratios in public communication might actually improve public understanding, since non-specialists often grasp '17 to 1 odds' more readily than a Gaussian tail probability.
- The same inequality-direction problem may affect other Bayesian-to-frequentist calibration tools, such as the Kass & Raftery approximation when used outside its nested and large-sample regime.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript argues that the common practice in exoplanet atmospheric studies of inverting the Sellke et al. (2001) bound to convert Bayes factors into frequentist sigma values yields an upper limit on the true significance, so that reported detections such as the 3-σ DMS claim in Madhusudhan et al. (2025) should be read as "less than 3-σ." The authors trace the practice to Benneke & Seager (2013), provide a worked example (B=21 → “at most 3.0σ”), and recommend abandoning the conversion altogether in favor of reporting Bayes factors directly.
Significance. If the paper's central claim were correct, it would provide a useful caution against a widespread practice in exoplanet atmospheric retrieval and would re-interpret several high-profile detection significances. The broader recommendation to prefer Bayes factors is reasonable and likely to be uncontroversial. However, the central mathematical claim is wrong: the inequality manipulation in Section 2 is reversed. The paper therefore does not provide the correction to community practice that it promises. The manuscript also fails to verify the Sellke et al. conditions for retrieval-based Bayes factors, but the direction error already invalidates the main conclusion.
major comments (2)
- [Section 2, Eq. (3)] The inversion direction is backwards. The Sellke bound is B10 ≤ U(p) with U(p) = -1/(e p log p). On the branch p < 1/e, U(p) is strictly decreasing in p: dU/dp = (log p + 1)/(e p^2 (log p)^2) < 0. Therefore B10 ≤ U(p) is equivalent to p ≤ U^{-1}(B10), which, since p = erfc(n_σ/√2) decreases with n_σ, gives n_σ ≥ n_σ(U^{-1}(B10)). The converted value is thus a lower bound on the frequentist significance, not an upper bound. The paper's example in Section 2 (“a Bayes factor of 21 corresponds, at most, to a 3.0σ”) is therefore reversed; the correct statement is that the significance is at least 3.0σ. A concrete counterexample: for H0: X~N(0,1) and H1: θ~N(0,1.92) with X|θ~N(θ,1), the observation x=3.3 gives two-tailed p=9.7×10^{-4} (3.3σ) and Bayes factor B10≈21. The Sellke bound gives U(9.7×10^{-4})≈54.7 ≥ 21, so the inversion holds, but the actual p-value corresponds to a higher (3.3σ), not lower, significance. This directly contradicts the manuscript's claim in Section 2 that “the true number of sigmas will, in general, be less.”
- [Section 3, Madhusudhan et al. example] The re-interpretation of the 3-σ DMS claim rests on the same inversion error. Under the correct direction, the reported Bayes factors of 17.5–68.0 in Table 2 of Madhusudhan et al. (2025) would imply p-values at most the thresholds obtained by inverting Eq. (3), hence significances at least 2.9–3.0σ; the phrase “at less than 3-σ significance” is not supported. The critique of the press release interpretation (0.3% probability by chance) is likewise misdirected, since the correct reading is that the p-value is no larger than that value. If the Sellke conditions fail for retrieval-based model comparison, the bound may not apply at all; but the paper does not verify those conditions for the specific retrieval settings of Benneke & Seager (2013) or Madhusudhan et al. (2025). Thus the paper's central warning about overestimation is unsupported in either case.
minor comments (4)
- [Section 2, Eq. (1)] The notation “− \(\exp p \log p\)” is at best ambiguous and likely a typographical error for \(-1/(e\,p\log p)\); as written, \(-e^{p\log p} = -p^p\) is negative and cannot be a lower bound. Please clarify the intended expression.
- [Section 2, text after Eq. (3)] The phrase “Equation (10) presents an upper bound on the Bayes factor; that means that a Bayes factor of for examples B10 = 21 corresponds, at most, to a 3.0 σ” is internally contradictory even under the paper's own reading, because an upper bound on B10 says nothing about an upper bound on n_σ unless monotonicity is established; the monotonicity actually gives the opposite direction.
- [Keywords and acknowledgments] There are minor typographical issues: “regatrding” in the acknowledgments, “Jeffrey’s scale” should be “Jeffreys scale”, and the keywords “The Princess Bride — Bayesian Blues” are unconventional and may not be suitable for a formal journal.
- [Figure 2] The figure is helpful but the caption does not state whether the curves assume nested models, equal prior odds, or which branch of the inverse; please add a brief description of the computational definitions underlying each curve.
Circularity Check
No significant circularity: the central claim rests on the external Sellke et al. bound, and the paper's self-citations are historical or stylistic rather than load-bearing.
full rationale
The paper's derivation chain begins with the Sellke et al. (2001) theorem, which is quoted as an external result with its assumptions explicitly stated. The inversion argument is a direct application of that theorem, not a parameter fitted to the paper's own outputs, nor a definition that presupposes the conclusion. The only self-references are Benneke & Seager (2013), cited as the historical source of the community conversion practice, and Kipping (2025), cited as a stylistic comparison regarding citation tracking; neither carries the logical weight of the central claim. The Madhusudhan et al. (2025) example is an external application used to illustrate the issue, not a fitted input. No equation in the paper reduces by construction to its own inputs, and no load-bearing premise is justified solely by a self-citation. Even if the direction of the inequality were disputed, that would be a mathematical correctness concern rather than a circularity concern.
Assumptions & free parameters
assumptions (4)
- standard math Sellke et al. (2001) lower bound B01 >= -e p log p under the stated conditions.
- standard math Gaussian p-value to sigma relation p = erfc[n_sigma/sqrt(2)].
- domain assumption Prior odds Pr(H1)/Pr(H0) are taken as unity in the detection examples.
- domain assumption Bayes factors from exoplanet retrievals can be compared with the Sellke setup.
Cite this review
Pith. "Pith review of Exoplaneteers Keep Overestimating Sigma Significances." pith.science (2026). https://pith.science/paper/LIW57Z25
@misc{pith2026250605392,
author = {Pith},
title = {Pith review of: Exoplaneteers Keep Overestimating Sigma Significances},
year = {2026},
howpublished = {\url{https://pith.science/paper/LIW57Z25}},
note = {Machine review of arXiv:2506.05392}
}
abstract
Astronomers, and in particular exoplaneteers, have a curious habit of expressing Bayes factors as frequentist sigma values. This is of course completely unnecessary and arguably rather ill-advised. Regardless, the practice is common - especially in the detection claims of chemical species within exoplanet atmospheres. The current canonical conversion strategy stems from a statistics paper from Sellke et al. (2001), who derived an upper bound on the Bayes factor between the test and null hypotheses, as a function of the $p$-value (or number of sigmas, $n_{\sigma}$). A common practice within the exoplanet atmosphere community is to numerically invert this formula, going from a Bayes factor to $n_\sigma$. This goes back to Benneke & Seager (2013) -- a highly cited paper that introduced Bayesian model comparison as a means of inferring the presence of specific chemical species -- in an attempt to calibrate the Bayes factors from their technique for a community that in 2013 was more familiar with frequentist sigma significances. However, as originally noted by Sellke et al. (2001), the conversion only provides an upper limit on $n_\sigma$, with the true value generally being lower. This can result in inflations of claimed detection significances, and this note strongly urges the community to stop converting to $n_\sigma$ at all and simply stick with Bayes factors.
Figures
Forward citations
Cited by 1 Pith paper
-
The KPF SURFS-UP Survey I: Transmission Spectroscopy of WASP-76 b
KPF transmission spectroscopy of WASP-76 b detects Fe I with a clear ingress-egress asymmetry but no measurable asymmetry for Na I or Ca II, supporting altitude-dependent circulation.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter doi edition editor eprint howpublished institution journal key month number organization pages publisher school series title misctitle type volume year version url label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts ...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION format.url url empty "" new.block "" url * "" * if FUNCTION format.eprint eprint empty "" archivePrefix empty "" archivePrefix "arXiv" = new.block " " eprint * " " * new.block " " eprint * " " * if if if FUNCTION format.doi doi empty "" " " doi * " " * if FUNCTION format.pid doi empty eprint empty ur...
-
[3]
!1A Qa
thebibliography [1] 20pt to REFERENCES 6pt =0pt -12pt 10pt plus 3pt =0pt =0pt =1pt plus 1pt =0pt =0pt -12pt =13pt plus 1pt =20pt =13pt plus 1pt \@M =10000 =-1.0em =0pt =0pt 0pt =0pt =1.0em @enumiv\@empty 10000 10000 `\.\@m \@noitemerr \@latex@warning Empty `thebibliography' environment \@ifnextchar \@reference \@latexerr Missing key on reference command E...
2021
-
[4]
2013, , 778, 153, 10.1088/0004-637X/778/2/153
Benneke , B., & Seager , S. 2013, , 778, 153, 10.1088/0004-637X/778/2/153
-
[5]
Chatrchyan , S., Khachatryan , V., Sirunyan , A. M., et al. 2012, Physics Letters B, 716, 30, 10.1016/j.physletb.2012.08.021
-
[6]
Eadie , G. M., Speagle , J. S., Cisewski-Kehe , J., et al. 2023, arXiv e-prints, arXiv:2302.04703, 10.48550/arXiv.2302.04703
-
[7]
2007, , 382, 1859, 10.1111/j.1365-2966.2007.12707.x
Gordon , C., & Trotta , R. 2007, , 382, 1859, 10.1111/j.1365-2966.2007.12707.x
arXiv 2007
-
[8]
Hubbard, R., & Lindsay, R. M. 2008, Theory & Psychology, 18, 69, 10.1177/0959354307086923
Show all 17 references
-
[9]
1939, Theory of Probability
Jeffreys , H. 1939, Theory of Probability
1939
-
[10]
1995, Journal of the American Statistical Association, 90, 773, 10.1080/01621459.1995.10476572
Kass , R., & Raftery , A. 1995, Journal of the American Statistical Association, 90, 773, 10.1080/01621459.1995.10476572
1995
- [11]
-
[12]
2025, , 983, L40, 10.3847/2041-8213/adc1c8
Madhusudhan , N., Constantinou , S., Holmberg , M., et al. 2025, , 983, L40, 10.3847/2041-8213/adc1c8
2025 doi
- [13]
-
[14]
R., et al
Radica , M., Taylor , J., Wakeford , H. R., et al. 2025, , 538, 1853, 10.1093/mnras/staf402
2025 doi
- [15]
-
[16]
J., & Berger, J
Sellke, T., Bayarri, M. J., & Berger, J. O. 2001, The American Statistician, 55, 62. http://www.jstor.org/stable/2685531
2001
-
[17]
2008, Contemporary Physics, 49, 71, 10.1080/00107510802066753
Trotta , R. 2008, Contemporary Physics, 49, 71, 10.1080/00107510802066753
2008 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.