{"id":"e82c5b2f-3731-421c-8894-0a0d56f2c1e4","arxiv_id":"1908.02631","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Gaussian process regression gives better-calibrated uncertainties for hot Jupiter dayside temperatures than error-weighted averaging or linear interpolation, and produces a twelve-planet catalogue with credible error bars.","lead":"This paper uses Gaussian process regression to estimate hot Jupiter dayside temperatures from sparse infrared eclipse measurements, and argues it produces more reliable error bars than simpler averaging or interpolation. It also warns that temperature estimates from only two Spitzer bands can be biased by up to 20 percent.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uncertainty calibration is not independently validated: the signal variance is fixed after inspecting z-scores on the same simulated spectra, and those spectra exclude thermal inversions and clouds, so the 68% coverage claim is an extrapolation.","rationale":"I agree with the reader's conditional assessment. The weakest assumption is correctly identified: the calibration transfer is not established, and there is a partially circular selection of the signal-variance hyperparameter. I would not move the verdict away from CONDITIONAL because the paper is transparent about the limited simulation suite and the GP method is a reasonable new application; the catalogue is useful, and the IRAC-only warning is a substantive result. But the 68% coverage statement should be presented as conditional on the cloud-free, non-inverted spectral family, and the benchmark should be re-run with a proper train/test split and with inversion/cloud models before the uncertainty calibration claim is taken at face value. The reader's weakest_assumption and my concern are essentially the same, so I mark agreement as agree and verdict as unchanged.","tokens_in":18156,"tokens_out":6760,"duration_ms":74529,"concrete_test":"Use the published code to run a controlled re-validation: (a) split the 324 model spectra into a calibration set and a held-out test set, choose log-signal variance from the 37,000 training fits (or from the calibration set) before computing any z-scores, and report z-score mean/std on the held-out set; (b) add a set of Pyrat Bay models with stratospheric TiO/VO thermal inversions and a set with parameterized cloud opacity, then recompute the GP coverage for WFC3 plus IRAC1 plus IRAC2 on the augmented suite. If held-out z-score standard deviations exceed about 1.1, or if coverage on the inversion/cloud models is substantially below 68%, the central calibration claim should be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark in Sec. 3.1 is presented as evidence that the GP returns unbiased temperatures with calibrated uncertainties, but the GP's key free parameter was tuned on that same benchmark. Sec. 2.1.2 states that the log-signal variance is fixed at -4 (14%) and then notes that this choice 'also becomes strongly motivated following our analysis in Section 3.1: with this hyperparameter, we retrieve statistically-appropriate distributions of effective temperature estimates.' The z-score distributions reported in Sec. 3.1 are computed from the same 324 model spectra and 97,200 realizations that motivated this choice. Selecting sigma^2 on the validation metric means the near-N(0,1) z-scores are a consistency check, not an independent test; the same logic would let any flexible method pass if its hyperparameters are tuned to the test set. This circularity is load-bearing because the stronger-than-EWM/LI claim rests on the z-score comparison. A separate transfer problem compounds it: the simulation suite contains only cloud-free, non-inverted, radiative-equilibrium models (Sec. 2.2.1; acknowledged in Sec. 3.1), so the fixed 14% amplitude and the EWM prior are never tested against atmospheres with thermal inversions or patchy clouds. For such targets, the catalogue statement that 68% of true temperatures fall in the reported 1-sigma intervals is therefore an assumption, not a measured coverage.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Gaussian-process regression method for estimating dayside effective temperatures of hot Jupiters from sparse secondary-eclipse photometry, specifically WFC3 white-light data plus warm Spitzer IRAC channels 1 and 2. The GP uses a squared-exponential kernel whose length scale is fixed from the HITEMP water spectrum and whose signal variance is fixed at 14%, with an inverse-error-weighted-mean prior. The authors benchmark the method against the error-weighted mean and linear interpolation on 97,200 simulated data sets generated from 324 cloud-free, radiative-equilibrium Pyrat Bay models, using Test/Teff ratios and z-scores as metrics. They find GP z-score distributions close to N(0,1) across SNR regimes while EWM and LI produce biased z-scores, and they caution that using only IRAC channels is significantly biased. The method is then applied to twelve hot Jupiters, yielding effective temperatures with uncertainties from 66 to 136 K, and the authors assert that 68% of the catalogue true temperatures will fall within the reported 1-sigma intervals.","tokens_in":18369,"tokens_out":7342,"duration_ms":86218,"significance":"If the calibration claims hold, this is a useful and fast model-independent estimator for sparse exoplanet eclipse data, and it provides a uniform catalogue of effective temperatures. The paper is transparent about its methods, provides public code, uses a large and reproducible simulation suite, and gives an honest treatment of the IRAC-only limitation. The main weakness is that the uncertainty calibration is not independently validated: the fixed signal variance is selected in part to make the same benchmark's z-scores look calibrated, and the simulation suite contains only non-inverted, cloud-free models, so the transfer of the 68% coverage statement to real planets is an extrapolation rather than a measured property.","major_comments":[{"comment":"The fixed log-signal variance (-4, 14%) is explicitly justified in Section 2.1.2 by the z-score behavior obtained in Section 3.1, and Section 3.1 uses the same 97,200 simulated data sets to validate the method. This makes the reported near-N(0,1) z-score distributions a consistency check rather than an independent test, and the headline comparison of GP against EWM and LI is therefore compromised as evidence for superior uncertainty estimation. I request a hold-out or cross-validated benchmark, or a sensitivity analysis showing that the conclusions are robust to signal variance over a physically plausible range, and a corresponding rephrasing of the validation claims.","section":"Section 2.1.2 and Section 3.1"},{"comment":"The simulation suite contains only cloud-free, non-inverted, radiative-equilibrium models, and Section 3.1 explicitly acknowledges the absence of thermal inversions and clouds. The physical argument for a 14% signal variance in Section 2.1.2 is based on a skin-temperature bound that assumes a non-inverted temperature profile. Consequently, the Section 3.2 assertion that 68% of catalogue true temperatures will fall within the reported 1-sigma intervals is not a measured coverage for real hot Jupiters, which may have thermal inversions or patchy clouds. Please either add simulations with inversions and clouds or replace that assertion with a clearly conditional statement about coverage under the simulation assumptions.","section":"Sections 2.2.1, 3.1, and 3.2"},{"comment":"The temperature estimates are repeatedly described as \"model-independent,\" but the GP result depends on the fixed kernel hyperparameters, the chosen mean function, and the normalization scheme. I recommend either defining the term carefully or describing the estimates as GP-prior-based or empirically calibrated, so that readers are not misled about the role of the adopted assumptions in the catalogue values.","section":"Abstract, Section 2.1, and Table 1"}],"minor_comments":[{"comment":"In the WASP-103 row, the IRAC channel 1 uncertainty appears as \"±0.38\" without a leading zero; if this is intended to be ±0.038, please correct the table.","section":"Table 1"},{"comment":"The text cites \"Pyrat Bay (Cubillos et al., in prep.),\" but the reference list contains only Cubillos (2016) and Blecic (2016); please provide the appropriate in-preparation citation or revise the text.","section":"Section 2.2.1"},{"comment":"The integral in Equation (5) is missing a closing parenthesis in the integrand; please rewrite it with unambiguous notation for the wavelength limits.","section":"Equation (5)"},{"comment":"The database is referred to variously as \"exoplanets.org,\" \"Exoplanets Data Explorer,\" and \"exoplanet.org\"; please use one consistent name.","section":"Throughout"},{"comment":"The sentence \"This choice is consistent with theoretical expectations, as we have discussed\" appears before the skin-layer discussion that follows; consider reordering or adding a pointer so the discussion is not introduced after its conclusion.","section":"Section 2.1.2"}],"recommendation":"major_revision","confidential_remarks":"The authors are commendably transparent that the signal-variance choice is motivated by the z-score behavior in Section 3.1, but this transparency does not remove the circularity: the benchmark on the same simulated data cannot independently validate the uncertainty calibration. The paper is publishable after a validation redesign or, at minimum, after adding a hold-out test and restricting the coverage claim to the simulation assumptions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing to know: this is a genuinely new use of Gaussian process regression for estimating dayside effective temperatures from sparse secondary-eclipse photometry, and it comes with a useful uniform catalogue of twelve hot Jupiters. The catch is real but not disqualifying: the uncertainty calibration is not independently validated, because the signal variance was fixed in part after looking at z-scores on the same simulated spectra used for the validation, and those spectra exclude thermal inversions and clouds.\n\nWhat the paper does well: the z-score framework is the right metric for comparing estimators that report both a value and an uncertainty. The three-SNR grouping is sane. In the simulated regime, the conclusion that EWM and LI underestimate the undersampling error while the GP does not is convincing. The 20% uncertainty warning for IRAC-only observations is a genuinely useful practical result, and the authors are honest that it is a worst-case scenario. The archival catalogue fills a real gap; the temperatures are physically reasonable and the quoted errors are derived in a transparent, reproducible way. The code is public. The relevant literature is covered, and the self-citations are to methods they are explicitly benchmarking, which is fair.\n\nWhere it gets soft: Section 2.1.2 says the log-signal variance is fixed at -4 and then notes that this choice 'also becomes strongly motivated following our analysis in Section 3.1'—the same Section 3.1 that reports the z-scores. That makes the near-N(0,1) z-scores a consistency check rather than an independent test. It is not fatal: the length scale comes from external HITEMP data, and the 14% amplitude has physical motivation (the skin-layer argument), but the 68% coverage claim for the catalogue inherits the circularity. Second, the simulation suite has no thermal inversions and no clouds, so the fixed 14% amplitude and the EWM prior are never tested against those spectral shapes. For a real planet with a strong inversion, the reported 1-sigma intervals may not have 68% coverage. The authors acknowledge the inversion/cloud omissions when discussing the IRAC-only bias, but not fully when asserting the catalogue's coverage. Minor: the GitHub code is not pinned to a commit hash and hyperparameter uncertainty is not propagated; the latter is defensible given fixed hyperparameters, but worth stating.\n\nWho it's for: exoplanet atmosphere people who need effective temperatures from sparse data. It deserves a serious referee; the method is plausible, the catalogue is useful, and the validation can be strengthened by adding inversion/cloud simulations and by choosing sigma^2 without reference to the benchmark. I would engage with it, and I'd want the calibration issue addressed before relying on the exact quoted intervals.","headline":"A useful new application of GP regression to sparse secondary-eclipse photometry, with a nice catalogue, but the uncertainty calibration is partly circular because the signal variance was tuned on the same simulated data used for validation.","tokens_in":18937,"tokens_out":3518,"would_cite":true,"duration_ms":36048,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Gaussian process regression can recover unbiased dayside effective temperatures of hot Jupiters from just three broad-band eclipse measurements, with uncertainties that behave like true 68% confidence intervals.","keywords":["hot Jupiters","dayside effective temperature","Gaussian process regression","secondary eclipses","Spitzer IRAC","HST WFC3","model-independent temperature estimation","radiative-equilibrium model spectra"],"falsifier":"Take a planet with both sparse WFC3+IRAC eclipse measurements and a full JWST secondary-eclipse spectrum, integrate the full spectrum to obtain the true bolometric dayside temperature, and check whether the truth falls inside the quoted 1σ interval; if it does for fewer than about 68% of a sample of such planets, the coverage claim is falsified.","tokens_in":17875,"feed_emoji":"🪐","tokens_out":11646,"duration_ms":108068,"temperature":0.7,"pith_summary":"The paper develops a model-independent way to estimate a hot Jupiter's dayside effective temperature from sparse secondary-eclipse measurements, using Gaussian process regression to interpolate brightness temperature across wavelength and to propagate both measurement noise and spectrum undersampling. Using 97,200 simulated data sets built from radiative-equilibrium hot Jupiter models, the authors show that when observations include white-light HST WFC3 plus the two warm Spitzer IRAC channels, the GP method returns unbiased effective temperatures with uncertainties that are neither systematically too large nor too small. The same simulations show that using only IRAC 3.6 and 4.5 µm data biases all estimators low by up to 20% at 1σ, because those bands probe the cooler upper atmosphere. Applied to the twelve hot Jupiters with published WFC3 and IRAC eclipse depths, the method yields dayside effective temperatures with 1σ uncertainties from ±66 K to ±136 K, and the paper asserts that the true temperature will fall inside each quoted interval 68% of the time.","feed_headline":"Three eclipse measurements recover hot-Jupiter dayside temperatures","feed_subtitle":"A Gaussian process turns sparse WFC3 and Spitzer eclipses into dayside temperatures with honest error bars.","key_machinery":"The carrying object is Gaussian process regression with the squared-exponential covariance kernel $k(r)=\\sigma^2 \\exp(-r^2/(2l^2))$, applied to brightness-temperature spectra converted from wavelength to frequency. The hyperparameters are fixed rather than fit to the sparse target data: the log length scale is set to $\\ln(l^2)=-8.55$ (about $1.4\\times10^{13}$ Hz, i.e., 0.19 µm at 2 µm), chosen from the low-resolution structure of water opacity, and the log signal variance is set to $\\ln(\\sigma^2)=-4$ (14% of the normalized brightness temperature), chosen from a 37-planet training sample. The GP uses a constant mean function equal to the inverse-error-weighted mean brightness temperature, so it behaves like the error-weighted mean method far from observed points but inflates uncertainty where the spectrum is undersampled. This uncertainty inflation is the key mechanism that the simpler estimators lack.","core_discovery":"The central claim is that a Gaussian process with a squared-exponential kernel, a correlation length scale fixed by water opacity, and a signal variance of 14% can recover unbiased dayside effective temperatures from as few as three broad-band measurements, with uncertainty estimates that are statistically accurate in all signal-to-noise regimes. On 97,200 simulated data sets, the GP produces z-scores—how far an estimate sits from the true temperature in units of its quoted uncertainty—centered near zero with standard deviations near one, whereas the error-weighted mean and linear-interpolation methods produce z-score spreads larger than one, meaning they underestimate the total error, especially at high signal-to-noise where undersampling dominates. The paper also establishes a limitation: with only the 3.6 and 4.5 µm IRAC bands, effective temperatures are systematically underestimated and known to no better than about 20% at 1σ, because those bands form in the cooler upper atmosphere and the model suite contains no thermal inversions.","pith_inferences":["If the GP is applied to other band combinations, the fixed signal variance of 14% should be retrained: observations that resolve finer spectral structure would likely favour a shorter length scale and a smaller amplitude, otherwise the quoted uncertainties may become too conservative.","Injecting thermal-inversion models into the simulation suite is a direct stress test; it would likely show that the IRAC-only low-temperature bias shrinks or reverses, and it would reveal how much of the claimed 68% coverage depends on the no-inversion assumption.","The z-score validation used in this paper could usefully become a standard check for any future empirical temperature estimator, since it exposes underestimation of uncertainty that average accuracy alone does not."],"forward_implications":["Dayside effective temperatures with reliable uncertainties can be obtained from just three broad-band eclipse measurements (WFC3 plus IRAC channels 1 and 2), so planets without full spectra no longer require a retrieval to get a trustworthy temperature.","IRAC-only 3.6 and 4.5 µm eclipse data should not be used to quote effective temperatures with precision better than about 20% at 1σ, regardless of estimator.","The error-weighted mean method, if used, should switch from inverse-variance weighting to inverse-error weighting to reduce outlier influence.","The twelve published temperatures, with 1σ uncertainties between ±66 K and ±136 K, constitute a testable prediction that upcoming space-based spectra will confirm or refute."],"supporting_citations":[{"why":"Supplies the linear-interpolation method benchmarked here and the 5% two-band systematic-error estimate that the IRAC-only result revises.","marker":"Cowan & Agol 2011"},{"why":"Defines the error-weighted-mean estimator whose inverse-variance weighting the paper replaces with inverse-error weighting.","marker":"Schwartz & Cowan 2015"},{"why":"Provides the Gaussian process regression formalism that the method adapts to sparse spectra.","marker":"Rasmussen & Williams 2006"},{"why":"The exoplanet secondary-eclipse catalogue supplies the 37 training planets for the signal-variance hyperparameter and the system parameters for the archival analysis.","marker":"Han et al. 2014"},{"why":"The water opacity line list sets the GP length scale from the low-resolution structure of water opacity.","marker":"Rothman et al. 2010"},{"why":"Its iterative radiative-equilibrium procedure underpins the model spectra used to generate the 97,200 simulated observations.","marker":"Malik et al. 2017"},{"why":"Motivates the inflated-uncertainty simulation scenario and documents how underconstrained retrievals can be outlier-sensitive.","marker":"Hansen et al. 2014"},{"why":"Supplies the reanalysed WASP-12 b WFC3 eclipse depth used in the archival catalogue.","marker":"Stevenson et al. 2014b"}],"fun_headline_variants":["Gaussian process turns three eclipses into hot-Jupiter temperatures","Three photometric points suffice for hot-Jupiter dayside temperatures","GP recovers hot-Jupiter dayside temperatures with honest error bars","Model-independent dayside temperatures from just three eclipses","Unbiased hot-Jupiter dayside temperatures from three-band photometry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's claimed 68% coverage depends on real hot Jupiter spectra resembling the simulated suite: cloud-free, without thermal inversions, and with brightness-temperature variability around the assumed 14%.","fun_headline_variants_meta":{"raw":{"variants":["Gaussian process turns three eclipses into hot-Jupiter temperatures","Three photometric points suffice for hot-Jupiter dayside temperatures","GP recovers hot-Jupiter dayside temperatures with honest error bars","Model-independent dayside temperatures from just three eclipses","Unbiased hot-Jupiter dayside temperatures from three-band photometry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000821,"raw_usage":{"total_tokens":3602,"prompt_tokens":966,"completion_tokens":2636,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":2550}},"tokens_in":582,"tokens_out":2636,"duration_ms":18673,"temperature":1.0,"reasoning_tokens":2550,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:39:35.592026+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a planet with both sparse WFC3+IRAC eclipse measurements and a full JWST secondary-eclipse spectrum, integrate the full spectrum to obtain the true bolometric dayside temperature, and check whether the truth falls inside the quoted 1σ interval; if it does for fewer than about 68% of a sample of such planets, the coverage claim is falsified.","supporting_citations":[{"cited_title":"J., Schwartz J","cited_arxiv_id":null,"evidence_quote":"Motivates the inflated-uncertainty simulation scenario and documents how underconstrained retrievals can be outlier-sensitive."}],"review_version":1}