{"id":"547f4de9-1181-4871-abcf-454d6502f9d4","arxiv_id":"2608.07128","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Pyrat Bay 2.0 adds equilibrium chemistry, radiative equilibrium, and vertically varying abundance retrieval, and simulations indicate JWST-quality spectra can recover such variations.","lead":"This paper describes version 2.0 of the open-source Pyrat Bay framework for modeling exoplanet atmospheres, adding new chemistry, radiative equilibrium, and retrieval tools. It validates the new modules against existing codes and uses simulated JWST observations of a warm Jupiter to show that ignoring altitude-dependent gas abundances can bias retrieved compositions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recoverability claim rests on a cloud-free closed loop in which the forward model and retrieval share the same opacities; real WASP-69b has haze evidence, so the 3.5-sigma-bias result may not generalize to JWST data.","rationale":"The paper is a solid methods contribution: chemcat is benchmarked against GGchem and FastChem, the radiative-equilibrium module is compared with HELIOS, and the code and cross-section repository are openly available. The non-isobaric VMR parameterization is useful, and the closed-loop experiment is carefully executed with realistic noise simulations and MultiNest posterior sampling. My concern is not an internal inconsistency in the software; it is the external validity of the headline scientific claim. The reader's weakest assumption already identifies the closed-loop, cloud-free limitation, and I agree with that assessment. I would add only a sharper framing: because the forward and retrieval models share the same opacity sources, the closed-loop agreement is partly a consistency check, not a test of the fidelity of the physical inputs. The proposed concrete test, injecting an independent forward model plus a Mie haze into the otherwise identical retrieval setup, would directly show whether the recoverability result and the constant-VMR bias survive a realistic departure from the idealized model. Since the paper itself acknowledges the idealized conditions and the reader's CONDITIONAL verdict already accounts for this, I do not recommend moving the verdict; I would keep it CONDITIONAL, with the abstract-level generality of the JWST claim remaining the main caveat.","tokens_in":26678,"tokens_out":14632,"duration_ms":150776,"concrete_test":"Construct a new WASP-69b truth model using an independent forward code (e.g., HELIOS or a photochemical-kinetic model) with the same bulk parameters (3x solar, C/O=0.59) but including a Mie haze with slant optical depth around 1 in the optical and a gray cloud near 1 mbar; simulate Gen TSO noise for NIRISS, NIRSpec, and MIRI at R=100; then run the identical constant and non-isobaric CH4 retrievals from Section 3.2. If the non-isobaric CH4 profile does not contain the true input VMR within 1 sigma, or if the H2O/CO bias of the constant model shifts by more than 1 sigma, the central claim is specific to cloud-free equilibrium models rather than established for real JWST-quality spectra.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim, that JWST-quality data can accurately recover vertical abundance variations and that constant-VMR retrievals are biased, is demonstrated only with synthetic spectra generated by the same Pyrat Bay/chemcat framework whose assumptions the retrievals test. Section 3.2 explicitly concedes 'somewhat idealized, cloud-free conditions,' yet Section 3.1 notes that WASP-69b's optical and near-infrared observations are 'characteristic of non-gray hazes.' Because the forward and retrieval models share the same CH4 line lists, sampling resolution, and continuum opacities, any common-mode opacity error is invisible in the closed loop and can either mask or mimic a vertical CH4 gradient. The quoted H2O/CO biases (about 3.5 sigma), the inferred 18x-solar metallicity, and the large Bayes factors (lnB=35.7-89.1) are therefore self-consistency statistics rather than demonstrated predictions for real JWST observations. If a Mie haze or gray cloud sits within the pressure range probed by the 3.3 and 8.0 micron CH4 features, the effective pressure differential between those bands shrinks, which could collapse the vertical constraint and change or erase the constant-VMR bias. The paper's caveat in Section 3.2 mitigates overstatement, but the abstract drops that caveat and states the JWST generalization without it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents version 2.0 of the open-source Pyrat Bay exoplanet-atmosphere modeling framework. New capabilities include the standalone thermochemical-equilibrium package chemcat; a 1D radiative-equilibrium solver with mixing-length convection; MultiNest nested sampling; depth-offset and error-scaling parameters for multi-epoch observations; isotopic-ratio fitting; H- opacity; a transit-light-source correction; and a free-chemistry parameterization that allows non-isobaric (sloped and capped) VMR profiles. The chemistry and radiative-equilibrium modules are benchmarked against GGchem, FastChem, and HELIOS, with abundance agreement of about 0.05-0.1 dex, temperature-profile differences of 10-50 K, and NIR emission spectra agreeing within 1-10%. The paper then presents a retrieval case study on simulated JWST-quality transmission spectra of WASP-69b (NIRISS/SOSS, NIRSpec/G395H, MIRI/LRS) generated from a 3x-solar radiative-thermochemical-equilibrium forward model with enhanced NH3 and SO2. Retrievals with a constant CH4 VMR cannot simultaneously fit the CH4 features probing different pressure levels, and produce H2O and CO abundances about 3.5 sigma above the input values (inferred metallicity about 18x solar instead of 3x solar), whereas a non-isobaric CH4 parameterization recovers the input profile, improves the reduced chi-squared from 1.56/1.61 to 1.09/1.14, and is strongly favored by the evidence (ln B = 35.7 and 89.1).","tokens_in":26972,"tokens_out":20633,"duration_ms":169589,"significance":"The framework validation is the paper's strongest asset: chemcat and the radiative-equilibrium solver are checked against independent external codes (GGchem, FastChem, HELIOS) rather than against the framework itself, with quantified agreement, and the code, pre-computed cross sections for 38 species, and the reproducibility compendium (Zenodo records cited in Sections 2.4.2 and the Data Availability statement) are openly available. The non-isobaric VMR parameterization is a genuinely useful bridge between constant free-chemistry and full equilibrium-chemistry assumptions, and the WASP-69b closed-loop experiment is a clear, explicit control study with quantitative diagnostics (reduced chi-squared and Bayes factors) demonstrating how vertical CH4 structure can bias constant-VMR retrievals. If the claims hold, Pyrat Bay 2.0 should be a widely used community tool for JWST-era analysis. The main limitation is scope: the headline recoverability result is demonstrated for one species, one idealized cloud-free forward model, and one closed loop sharing the retrieval's opacity inputs, so the abstract's generalization to actual JWST observations is stronger than the evidence presented.","major_comments":[{"comment":"The abstract's central finding — 'vertical abundance variations in planetary atmospheres can be accurately recovered using data of JWST quality' — is broader than what the paper demonstrates. Section 3 is a single closed-loop simulation for one species (CH4) under 'somewhat idealized, cloud-free conditions' (Section 3.2), using a WASP-69b forward model that omits the non-gray haze that Section 3.1 itself cites as characteristic of this target's observed optical and near-infrared spectrum. The conclusions ('with simulated transit observations') are properly qualified, but the abstract drops those qualifiers and asserts a general result about JWST data. Because a real cloud or haze inside the probed pressure window (10^-2 to 10^-5 bar) would reduce the pressure contrast between the CH4 features and could weaken or remove the vertical constraint and the quoted 3.5-sigma biases, I recommend either carrying the 'idealized, cloud-free, simulated' qualifiers into the abstract, or adding a robustness test (e.g., a gray-cloud or Mie-haze forward model, or a perturbed CH4 opacity treatment) that demonstrates the result still holds.","section":"Abstract; §3.2"},{"comment":"The forward model and the retrieval share the same CH4 line list (Yurchenko et al. 2024a, Table 2), the same R = 25,000 cross-section sampling (Section 2.4.2), and the same continuum, alkali, and collision-induced absorption opacities. Any common-mode opacity error is therefore invisible to the closed loop, so the non-isobaric retrieval's 'accurate recovery' is a self-consistency check of the new VMR parameterization given the adopted opacity set, not an external validation of the recovered vertical profile. I do not regard this as a flaw of the experimental design — an injection test is the correct way to validate a new parameterization — and the retrieval's parametric Madhusudhan-Seager temperature profile (Section 3.2) differs from the forward model's self-consistent radiative-equilibrium profile, which partly mitigates the shared-input concern. However, the manuscript should state explicitly that the recoverability result is conditional on the adopted opacities, and, if feasible, test sensitivity by perturbing the CH4 line-wing or sampling treatment in the forward model.","section":"§3.2; §2.4.2"},{"comment":"The quantitative headline numbers — H2O and CO retrieved about 3.5 sigma above the true values, and a metallicity of about 18x solar instead of 3x solar — require a scalar definition of the 'true' abundance for species whose input VMR varies with pressure (Fig. 7 middle panel; black dashed curves in Fig. 8). Table A1 reports posterior medians and 68% intervals but no comparison metric (e.g., transit-geometry-weighted mean, pressure-averaged VMR, or value at a reference pressure within the 10^-2 to 10^-5 bar window). Please state the metric used for each quoted deviation so that the 'within 1 sigma' and '3.5 sigma' claims are checkable; without it, the central bias claim cannot be precisely evaluated.","section":"§3.2; Table A1"}],"minor_comments":[{"comment":"The sentence 'The ‘equilibrium-chemistry’ parameterization is a second approach commonly used used for exoplanet atmospheric retrievals' contains a duplicated 'used', which should be removed.","section":"§2.3.3"},{"comment":"The text 'we will explore the posibility to incorporate additional thermodynamic databases' contains the typo 'posibility'; it should be 'possibility'.","section":"§2.1.2"},{"comment":"The caption 'Equilibrium tempertatures were derived from the system parameters' contains the typo 'tempertatures'; it should be 'temperatures'.","section":"Table 1 caption"},{"comment":"The non-isobaric retrievals end at reduced chi-squared values of 1.09 and 1.14, still somewhat above unity; a sentence explaining the residual tension (e.g., the parametric Madhusudhan-Seager temperature profile or the constant-VMR treatment of species such as CO2 or H2S) would help readers interpret the model comparison.","section":"§3.2"},{"comment":"The equilibrium-chemistry, hybrid, and isotopic-ratio retrieval modes are implemented and described but not exercised or validated in Section 3; a sentence stating which features are externally benchmarked, which are demonstrated by the closed-loop test, and which remain untested would set accurate expectations for users of the new framework.","section":"§2.3.2-§2.3.5"},{"comment":"The caption's clause 'The panels for CH4 shows the abundance profile as a function of pressure' should read 'show the abundance profiles', since multiple retrieved profiles are displayed.","section":"Figure 8 caption"}],"recommendation":"major_revision","confidential_remarks":"To the editor: this is a solid software/methods contribution whose external benchmarks (GGchem, FastChem, HELIOS) and open reproducibility give it value independent of the retrieval case study. The reviewer-level disagreement is about the weight of the abstract's recoverability claim: the body is properly qualified, but the abstract generalizes beyond the demonstrated idealized, cloud-free, opacity-consistent closed loop. That gap is fixable either by re-scoping the abstract or by adding a haze/opacity robustness test, so I recommend major revision rather than rejection. A secondary point is that several advertised capabilities (equilibrium-chemistry retrieval, hybrid mode, isotopic-ratio fitting) are described but not demonstrated; requesting a short validation-status statement should resolve it. I have no concerns about citation practices or fit with MNRAS."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real methods paper, not a hype job. The new pieces—chemcat, radiative-equilibrium profiles, MultiNest integration, non-isobaric VMR parameterization, H- opacity, transit light source modeling, depth offsets and noise scaling—are what the abstract says they are. I trust the chemistry module: agreement with GGchem to about 0.1 dex and with FastChem to 0.05 dex on shared networks is credible, and the HELIOS radiative-equilibrium comparison (10–50 K, 1–10% spectra) is the right kind of external validation. The code, cross-section tables, and data compendium are all actually shipped, and that buys real goodwill.\n\nThe non-isobaric VMR model is the most interesting contribution. It is simple, cheap, and lets a retrieval move from a constant VMR to slanted, quenched, or capped profiles without leaving the free-chemistry framework. The WASP-69b closed-loop experiment is a sensible first test, and the chi-squared improvement from 1.56/1.61 to 1.09/1.14 with large Bayes factors is internally consistent. The claim that constant-VMR retrievals get biased by about 3.5 sigma in H2O and CO when the truth has a vertical CH4 gradient is a fair demonstration of the model's value.\n\nSoft spots, in proportion. First, the abstract's statement that vertical abundance variations can be accurately recovered using data of JWST quality is broader than what was shown. The simulations are cloud-free, generated with the same forward model and opacities the retrievals assume, and WASP-69b itself shows haze evidence in real data. The paper does flag the idealized conditions in Section 3.2, so this is an overstatement in the abstract rather than a hidden flaw, but it should be fixed. A Mie haze or gray cloud near the probed pressures could shrink the pressure lever arm between the CH4 bands and weaken the vertical constraint. Second, the benchmark comparisons carry no explicit uncertainties, and convergence of the radiative-equilibrium iterations and MultiNest runs is asserted more than demonstrated. Minor, but a referee should ask for details. Third, the simulated spectra are closed-loop; that is fine for a code demo, and the authors present it that way, but readers should not treat the recovered metallicity as a measurement of WASP-69b.\n\nThe citation pattern looks clean; the self-citations to Pyrat Bay and chemcat are appropriate. I disagree with any suggestion that the paper is circular in a damaging way: the chemistry and radiative-transfer modules are checked against independent codes, and the retrieval demo is explicitly a control experiment.\n\nWho this is for: people building or using retrieval codes for JWST transmission spectra, and anyone who wants a flexible non-isobaric abundance parameterization. It deserves a serious referee. I would send it out, with a request to soften the abstract and add a few convergence and uncertainty details. I would probably cite the non-isobaric VMR model in a retrieval paper this year.","headline":"A genuinely useful, well-benchmarked code upgrade whose headline JWST recoverability claim is demonstrated only in an idealized closed loop; worth refereeing, with the abstract needing to carry the caveat.","tokens_in":27576,"tokens_out":2553,"would_cite":true,"duration_ms":24072,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulated JWST observations show that vertical abundance variations in exoplanet atmospheres can be recovered, while constant-abundance retrievals bias H2O and CO by about 3.5 sigma and infer a metallicity of roughly 18 times solar…","keywords":["exoplanet atmospheres","atmospheric retrieval","thermochemical equilibrium","radiative equilibrium","transmission spectroscopy","James Webb Space Telescope","non-isobaric abundances","nested sampling"],"falsifier":"Run the same constant-VMR versus non-isobaric CH4 retrieval comparison on synthetic spectra that include a gray cloud deck near 10 mbar or a three-dimensional circulation-consistent temperature structure; if the non-isobaric model then fails to recover the input abundances within 1 sigma, or if real WASP-69b spectra do not prefer the vertical model, the claim that JWST can recover vertical abundance variations would be undercut.","tokens_in":26426,"feed_emoji":"🪐","tokens_out":6509,"duration_ms":55267,"temperature":0.7,"pith_summary":"This paper argues that JWST-quality transmission spectra can reveal how gas abundances change with altitude in exoplanet atmospheres, and that the standard retrieval assumption of constant-with-altitude abundances can badly mislead. The authors upgrade the open-source Pyrat Bay framework with a fast standalone thermochemical-equilibrium package (chemcat), self-consistent radiative-equilibrium temperature profiles, nested-sampling retrievals, and a parametric non-isobaric volume-mixing-ratio model that lets a species' abundance slant or step with pressure. In a closed-loop WASP-69b test, a methane profile that varies with altitude reproduces the simulated JWST data and recovers input abundances within 1 sigma, while the constant-VMR retrieval overestimates H2O and CO by about 3.5 sigma and returns about 18 times solar metallicity instead of 3 times. The paper concludes that vertical abundance structure is within reach of JWST and must be modeled to avoid biased abundance constraints.","feed_headline":"Flat-abundance retrievals skew JWST results by 3.5 sigma","feed_subtitle":"Allowing methane to vary with altitude recovers true exoplanet abundances that constant mixing ratios miss.","key_machinery":"The load-bearing tool is the non-isobaric VMR parameterization: a slanted log VMR versus log pressure line, capped between minimum and maximum abundance values, with one to four free parameters depending on how much structure the data support. With a single free parameter it reduces to the standard constant-abundance model, which enables nested-sampling model comparison between the two. The other machinery carrying the argument is chemcat, a Gibbs free-energy minimizer with a Newton-Raphson and Lagrange-multiplier solver that computes equilibrium abundances fast enough for retrievals, coupled to a two-stream radiative-equilibrium iteration that self-consistently updates temperature and chemistry. The WASP-69b demonstration uses the slanted CH4 profile to show what the constant model misses.","core_discovery":"The central claim is that vertical abundance variations in a gas-giant atmosphere are detectable and accurately recoverable with data of JWST quality, and that retrievals that neglect these variations can be substantially biased. In the WASP-69b demonstration, CH4 is the most abundant trace gas with a strongly pressure-dependent VMR; its absorption bands at 3.3 and 3.9 \\textmu m probe different pressure levels, so a constant CH4 profile cannot fit both simultaneously. That mismatch propagates through abundance correlations, inflating H2O and CO by about 3.5 $\\sigma$ and producing a metallicity estimate of about 18 times solar rather than the true 3 times solar. Replacing constant CH4 with the slanted non-isobaric model recovers the true abundances within 1 $\\sigma$ and improves reduced $\\chi^2$ from 1.56/1.61 to 1.09/1.14, with Bayes factors (\\ln B) of 35.7 and 89.1 favoring the vertical model. The authors also validate chemcat and the radiative-equilibrium module against GGchem, FastChem, and HELIOS, and present cross-section tables, isotopic-ratio fitting, transit light source corrections, and data-offset and noise-scaling models as part of the framework.","pith_inferences":["If this result transfers to real JWST spectra, published gas-giant retrievals that used constant-VMR free chemistry may need reanalysis; inferred metallicities could shift once vertical gradients are allowed.","Recovered vertical abundance slopes could serve as diagnostics of quench pressures, photochemical destruction, or vertical mixing, linking retrieval outputs to atmospheric dynamics.","Because the simulated dataset is cloud-free and one-dimensional, real observations with clouds, hazes, or terminator inhomogeneities may require joint cloud-plus-vertical-abundance retrieval before the claimed 1-sigma recovery is seen.","A direct test would be to run the same constant-versus-non-isobaric CH4 comparison on actual WASP-69b JWST data and check whether the vertical model is preferred and whether the resulting H2O and CO abundances agree with independent constraints."],"forward_implications":["JWST retrievals that assume constant volume mixing ratios for species with vertical gradients will systematically overestimate H2O and CO abundances and inferred metallicity.","Non-isobaric VMR or equilibrium-chemistry parameterizations should be preferred for gas giants with JWST-quality data, and Bayes-factor model comparison can identify when vertical structure is statistically required.","Combining NIRISS/SOSS with NIRSpec and MIRI extends the pressure range probed by transmission spectra and tightens constraints on vertical abundance profiles.","Species whose abundances increase with altitude, like CO2 in this model, are less prone to altitude-induced bias in transmission geometry, so biases are species-specific rather than global."],"supporting_citations":[{"why":"Original Pyrat Bay modeling and retrieval framework that version 2.0 upgrades; supplies the opacity, atmospheric, and spectral machinery.","marker":"Cubillos & Blecic 2021"},{"why":"TEA thermochemical-equilibrium code from which chemcat's capabilities are derived and which chemcat replaces for speed.","marker":"Blecic et al. 2016"},{"why":"GGchem, the benchmark for the thermochemical-equilibrium comparison, yielding agreement within about 0.1 dex for most species.","marker":"Woitke et al. 2018"},{"why":"FastChem, the second benchmark for chemcat neutral and ionic equilibrium abundances, matching within 0.05 dex.","marker":"Stock et al. 2018"},{"why":"HELIOS radiative-equilibrium scheme and two-stream solver that Pyrat Bay's radiative equilibrium follows and is benchmarked against, with profile differences of 10-50 K.","marker":"Malik et al. 2017"},{"why":"MultiNest nested-sampling algorithm used for posterior sampling and Bayesian evidence computation in the retrievals.","marker":"Feroz et al. 2009"},{"why":"Gen TSO simulator used to generate the simulated JWST NIRISS, NIRSpec, and MIRI transmission observations.","marker":"Cubillos 2024"},{"why":"Earlier simulations that motivate the need for non-isobaric abundances in atmospheric retrievals.","marker":"Changeat et al. 2019"}],"fun_headline_variants":["Altitude-varying methane shifts exoplanet abundance fits","Constant CH4 profiles bias H2O and CO by 3.5 sigma","Pyrat Bay 2.0 recovers true exoplanet abundances from JWST data","Vertical chemistry models fix exoplanet retrieval biases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole demonstration assumes that the simulated JWST spectra are generated by the same idealized, cloud-free, one-dimensional radiative-thermochemical model that the retrievals test, so real atmospheres with clouds, three-dimensional structure, or opacity errors could behave differently.","fun_headline_variants_meta":{"raw":{"variants":["Altitude-varying methane shifts exoplanet abundance fits","Constant CH4 profiles bias H2O and CO by 3.5 sigma","Pyrat Bay 2.0 recovers true exoplanet abundances from JWST data","Vertical chemistry models fix exoplanet retrieval biases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000874,"raw_usage":{"total_tokens":3854,"prompt_tokens":1091,"completion_tokens":2763,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":707,"completion_tokens_details":{"reasoning_tokens":2688}},"tokens_in":707,"tokens_out":2763,"duration_ms":18358,"temperature":1.0,"reasoning_tokens":2688,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:18:09.552224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same constant-VMR versus non-isobaric CH4 retrieval comparison on synthetic spectra that include a gray cloud deck near 10 mbar or a three-dimensional circulation-consistent temperature structure; if the non-isobaric model then fails to recover the input abundances within 1 sigma, or if real WASP-69b spectra do not prefer the vertical model, the claim that JWST can recover vertical abundance variations would be undercut.","supporting_citations":[],"review_version":1}