{"id":"d33577dc-7681-45fc-b458-b84e5e18170c","arxiv_id":"2411.14056","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using 3D radiative MHD simulations of G2V, K0V, and M0V stars, this paper finds that 1D mixing-length models fail for K and M dwarf spot spectra but work for Sun-like stars.","lead":"This paper computes the first starspot spectra from 3D magnetic simulations of Sun-like, K, and M dwarf stars. It shows the standard 1D cool-star approximation misrepresents spot light for K and M dwarfs, directly affecting how exoplanet atmospheres are read.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never varies the mixing-length parameter α in the 1D RE models, so the claimed failure of mixing-length theory for M0V/K0V could be a calibration artifact rather than a structural limitation.","rationale":"Read in good faith: the paper is a strong and valuable first step, and the diagnostic in Section 5.2—showing that horizontal structures are not the main issue and that vertical stratification is—is well designed. The published 1D MHD atmospheres and spectra are useful community resources. The load-bearing weakness is not the MURaM setup per se, although that is also a fair concern; it is that the comparison to 1D RE models never explores the main free parameter of the very approximation being criticized. Since the solar-calibrated mixing-length parameter need not transfer to M dwarfs, the observed mismatch could be reduced or removed by recalibration. The proposed α-grid test is inexpensive, uses the paper's own tools and published tables, and would directly decide whether the failure is structural or parametric. This does not change the overall verdict (conditional acceptance), but it sharpens the condition under which the claim should be accepted.","tokens_in":15766,"tokens_out":4354,"duration_ms":45828,"concrete_test":"Recompute the M0V (and K0V) umbral, penumbral, and quiet-region RE models with MPS-ATLAS for a grid of mixing-length parameters, e.g. α = 0.5, 0.8, 1.0, 1.25, 1.6, 2.0, 2.5, and optionally a nonlocal MLT prescription, keeping Teff matched as in Section 5.1. Calculate the relative contrast differences (Equation 3) longward of 500 nm for each case, and compare the resulting T(P) stratifications to the published '1D MHD' atmospheres in Tables 2 and 3. If some α reduces the M0V umbral/penumbral error below ~10%, the paper's attribution to a structural failure of mixing-length theory needs revision; if all α values retain errors near 50%, the central conclusion is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the comparison in Section 5.1 between 3D MHD contrasts and 1D RE contrasts. The RE models use mixing-length convection with a fixed solar-style calibration, and the only free parameter varied is effective temperature (ΔTeff = ±50, ±100 K). Section 5.2 and Figure 8 show that the RE temperature stratification is steeper than the horizontally averaged MHD atmosphere, and the paper attributes this to failure of mixing-length theory. But mixing-length theory has a free parameter α (and alternative formulations); increasing α increases convective efficiency and flattens the superadiabatic gradient. In M0V stars, where convection is important throughout the photosphere (Beeck et al. 2013; Panja et al. 2020), the RE model's temperature gradient is expected to be highly sensitive to α. Without testing whether any plausible α or MLT variant can reproduce the MHD stratification and contrasts, the conclusion that '1D models fail' is overstated: it may be only the specific default MLT calibration that fails. This is load-bearing because the abstract's 50% error for M0V is the paper's headline quantitative result, and the proposed fix (recalibrated 1D models) would materially change the practical message for exoplanet contamination studies.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the first computations of starspot spectra based on 3D radiative MHD simulations (MURaM) for G2V, K0V, and M0V stars, with spectra synthesized by ray-by-ray transfer using MPS-ATLAS under LTE. The central comparison is between umbral and penumbral flux contrasts from these MHD cubes and contrasts from 1D radiative-equilibrium (RE) models with mixing-length convection, matched in effective temperature and also with Delta_Teff = +/-50, +/-100 K. The authors report that the 1D RE models reproduce the G2V contrasts well but fail for K0V and M0V, with errors growing toward later spectral types, and they use horizontal averaging to argue that horizontal inhomogeneities are not the cause. They attribute the discrepancy to a steeper vertical temperature stratification in the RE models caused by a failure of mixing-length theory, and they provide machine-readable 1D MHD atmospheres and spot spectra as data products.","tokens_in":16030,"tokens_out":8400,"duration_ms":90718,"significance":"If the conclusions hold, the paper has substantial practical importance: standard 1D cool-star spot models are widely used to correct exoplanet transmission spectra and to model stellar variability, and the paper implies that such corrections are systematically inaccurate for K and M dwarfs. The work is significant as a first application of 3D radiative MHD to starspot spectra, and it contains a valuable internal consistency check, namely that horizontally averaged MHD atmospheres reproduce the full 3D MHD contrasts, isolating the vertical stratification as the source of the discrepancy. The G2V result is reassuringly consistent with the known success of 1D solar variability models, and the comparison with PHOENIX models in Appendix A provides a useful cross-check. The main caveats are that the MHD simulations are treated as the reference truth despite relying on an ad hoc boundary condition and on a simulation description deferred to another paper, and that the 1D comparison has so far been made with a single, not fully specified mixing-length calibration.","major_comments":[{"comment":"The attribution of the M0V/K0V discrepancy to a structural failure of mixing-length theory is not yet established, because the MPS-ATLAS RE models are computed with one mixing-length formulation and the calibration parameter alpha is neither quoted nor varied. A larger alpha increases convective efficiency and flattens the superadiabatic gradient, so at least part of the steeper RE stratification seen in Fig. 8 could be a calibration artifact rather than a fundamental limitation of 1D mixing-length models. Please state the value of alpha used in Section 2.3, perform a scan over a plausible range of alpha (or test an alternative 1D convection prescription) for the M0V and K0V umbra and penumbra, and report whether any reasonable calibration reproduces the horizontally averaged MHD stratification and contrasts. This is load-bearing because the abstract's 'about 50%' error statement and the conclusion that '3D MHD modelling is necessary' are stronger than what a single-calibration comparison can support. The PHOENIX comparison in Appendix A partly mitigates the concern by showing that a second 1D code also fails, but it does not replace a direct sensitivity test.","section":"§5.2, Fig. 8"},{"comment":"The 3D MHD simulations are the reference against which the 1D models are judged, but their realism is not sufficiently documented in this manuscript. The spot setup is referred to a paper in preparation (Bhatia et al. 2024), and the penumbra is maintained through an ad hoc top boundary condition in which the magnetic field is made three times more horizontal than a potential field. These choices can directly affect the vertical temperature stratification and the penumbral structure, which are precisely the quantities that control the comparison in Section 5. Please include the essential setup details and a sensitivity check of the umbral and penumbral contrasts to the boundary-condition parameter, or provide a citable description of the simulations. Without this, the conclusion that 1D models fail for K0V and M0V stars is conditional on the realism of these particular MURaM simulations.","section":"§2.1, Table 1"},{"comment":"The paper's quantitative claim for G2V, namely errors of less than 2% for the umbra and less than 10% for the penumbra, is not accompanied by uncertainty estimates for the MHD contrasts. Appendix D and Fig. 12 show that individual simulation cubes differ from the mean by up to roughly +/-10%, which is comparable to the claimed penumbral accuracy. Please provide cube-to-cube uncertainty estimates for the MHD contrasts and state whether the G2V penumbral agreement remains significant once this variability is taken into account. This is needed to support the asymmetry between the G2V and the K0V/M0V conclusions.","section":"§5.1, Appendix D"}],"minor_comments":[{"comment":"The wavelength range is quoted inconsistently: the abstract says 250-6000 nm, Section 2.2 says 200-6000 nm, and the normalization integrals in Equations (1) and (2) are said to cover 300-6000 nm. Please make these numbers consistent.","section":"Abstract, §2.2, §5.1"},{"comment":"The abstract's headline number, 'errors longward of 500 nm of about 50%' for M0V, is not explicitly derived or located in the text or a figure. Please point the reader to the relevant computation or add the value to the discussion of Figures 4-5.","section":"Abstract, §5.1"},{"comment":"Equations (1) and (2) define contrasts that are normalized by construction, so the term 'absolute contrast' is potentially misleading. The unnormalized quantity used in Equation (3) should be given a distinct name, such as 'unnormalized absolute contrast', to avoid confusion.","section":"§5.1, Eqs. (1)-(3)"},{"comment":"The umbral and penumbral mask thresholds at 400 nm are chosen separately for each spectral type and directly set the effective temperatures of the features used to construct the RE models. A sensitivity test of the derived contrasts and effective temperatures to these thresholds would strengthen the quantitative comparison.","section":"§3"},{"comment":"The paper does not state the mixing-length parameters (for example alpha and the ratio of mixing length to pressure scale height) used in the MPS-ATLAS RE models. These should be listed so that the claimed failure of the mixing-length approximation can be reproduced and tested.","section":"§2.3"},{"comment":"The LTE assumption is stated but not discussed. Given the cool temperatures and strong molecular features in M0V spots, a brief justification of LTE for the broad-band fluxes and contrasts considered here would be useful.","section":"§2.2"},{"comment":"The reference to Bhatia et al. (2024, in prep.) should be updated with a preprint number or DOI, or the relevant details should be summarized in the text.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and presents a genuinely new calculation with useful data products. The main reasons for major revision are the untested mixing-length calibration and the insufficiently documented MHD setup; both are load-bearing for the central claim. I do not see grounds for rejection, as the issues are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you should know: this gives the community something it didn't have — the first starspot spectra from 3D radiative MHD simulations for G2V, K0V, and M0V, plus machine-readable 1D MHD atmospheres for quiet regions, penumbra, and umbra for K0V and M0V. That alone is worth having. The comparison with standard 1D radiative-equilibrium models is the real news: for G2V the standard approach is fine, which makes sense and matches SATIRE's success, but for K0V, and especially M0V, it misses the wavelength dependence of spot contrast by a lot — roughly 50% longward of 500 nm in the M0V case. The internal check with horizontally averaged MHD atmospheres is a good piece of diagnostics: it shows horizontal inhomogeneities are not the problem; the vertical stratification is.\n\nThe soft spots are real but not fatal. First, the MURaM spot setup lives in Bhatia et al. (2024, in prep.), including the cylindrical flux-tube setup and the ad hoc top boundary condition that makes the field three times more horizontal to preserve a penumbra. That top boundary is a modeling crutch; if it biases the magnetic structure and hence the temperature stratification in the penumbra, the quantitative contrast for penumbrae could shift. The LTE assumption is softer; it is standard for this wavelength range and mostly affects lines, not continuum contrasts.\n\nSecond, the cube-selection in Appendix D is qualitative (\"do not deviate significantly from the mean\"), and Figure 12 shows a few percent scatter, so this is minor.\n\nThird, there are no error bars on the contrast curves; for a simulation paper that is acceptable but worth noting.\n\nNow the stress-test concern about not varying alpha. It is partially right. The paper shows that the specific 1D RE models (MPS-ATLAS with standard MLT) do not reproduce the MHD contrasts, and the temperature gradient in the RE model is steeper than the averaged MHD atmosphere. But they never vary the mixing-length parameter alpha in the 1D models. So the claim \"mixing-length theory is inaccurate\" is a bit too broad: it could be that only the standard calibration fails. That said, for the practical audience the message survives — the 1D models everyone actually uses (ATLAS, PHOENIX-style) are the ones tested, and the paper shows they are off. The headline 50% number is for those standard models. The fix they offer is not a re-tuned MLT anyway; it is the tabulated MHD atmospheres. So I would call this a moderate wording issue, not a load-bearing flaw.\n\nWho is this for: exoplanet transmission spectroscopy folks modelling spot contamination, and stellar activity people. They should read it and probably take the tables. As a referee: yes, send it out; it deserves serious review. The key things to push on are the companion paper details and an explicit statement that alpha was not varied, so the MLT conclusion is framed as \"standard MLT models\" rather than \"MLT in general.\"","headline":"Strong new resource: first 3D MHD spot spectra for G2V/K0V/M0V and a mostly convincing case that standard 1D spot models fail for cool dwarfs, though 'mixing-length failure' is a bit too sweeping without testing alpha.","tokens_in":16646,"tokens_out":3454,"would_cite":true,"duration_ms":33524,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Starspot spectra from 3D magnetohydrodynamic simulations show standard 1D cooler-star models fail for K- and M-type dwarfs, with roughly 50% errors longward of 500 nm for M0V, while remaining adequate for Sun-like G2V stars.","keywords":["starspots","stellar magnetic fields","radiative magnetohydrodynamics","M dwarf atmospheres","mixing length theory","stellar contamination","exoplanet transmission spectroscopy","spectral synthesis"],"falsifier":"Observe a spot-crossing event on an M0 dwarf with wavelength-resolved spectroscopy (the paper itself points to TOI-3884 as a candidate and to HST/JWST as capable instruments): if the measured wavelength dependence of the umbral and penumbral contrast follows the 1D radiative-equilibrium prediction rather than the 3D MHD prediction, the central claim would be refuted. A cheaper calculation-based test would be to rerun the M0 spot with a potential-field top boundary condition and see whether the ~50% discrepancy survives.","tokens_in":15587,"feed_emoji":"⭐","tokens_out":14220,"duration_ms":112009,"temperature":0.7,"pith_summary":"Stars' dark spots change the spectrum of the light we measure, and exoplanet observations are routinely corrected by treating a spot as a cooler, non-magnetic star represented by a one-dimensional (1D) radiative-equilibrium atmosphere with mixing-length convection. This paper replaces that approximation with self-consistent 3D radiative magnetohydrodynamic (MHD) simulations of spots on G2V, K0V, and M0V stars and compares the resulting spectra to those from the 1D models. The paper's central finding is that the 1D models fail for K0V and M0V stars: for M0V the umbral and penumbral flux contrast longward of 500 nm is off by about 50%, while for G2V the errors are below 2% (umbra) and 10% (penumbra). The authors attribute the failure to mixing-length theory being inaccurate when convection and radiation transport energy together over a broad range of heights, which is the case in cooler stars but not in solar-type stars. If correct, current spot-contamination corrections for K- and M-dwarf transmission spectra are systematically biased, and 3D MHD spot spectra are needed for those stars.","feed_headline":"50% error: standard 1D starspot models fail on M dwarfs","feed_subtitle":"Mixing-length spot models are off by 50% on M dwarfs, threatening exoplanet water detections.","key_machinery":"The central comparison is the wavelength-dependent flux contrast of a magnetic feature (umbra, penumbra, or combined spot) against quiet-star flux, computed two ways: from 3D MHD simulation cubes produced by MURaM via ray-by-ray radiative transfer in LTE with MPS-ATLAS, and from 1D radiative-equilibrium model atmospheres that treat convection with the Böhm-Vitense mixing-length approximation. The MURaM setup uses cylindrical flux tubes and an ad hoc top boundary condition that makes the magnetic field three times more horizontal than a potential field in order to sustain a penumbra. The decisive diagnostic is the vertical temperature stratification as a function of pressure: horizontally averaging the 3D MHD columns preserves the 3D contrasts, whereas the 1D RE atmospheres develop steeper temperature gradients because mixing-length theory mishandles convection when convection and radiation both carry significant energy over a wide height range.","core_discovery":"The paper presents the first starspot spectra computed self-consistently from 3D radiative MHD simulations of spots on G2V, K0V, and M0V stars, using the MURaM code for the atmospheric structure and the MPS-ATLAS code for ray-by-ray LTE radiative transfer at a resolving power of about 500 from 250 nm to 6000 nm. Comparing the wavelength-dependent flux contrast of umbra, penumbra, and combined spot relative to the quiet star with contrasts from 1D radiative-equilibrium (RE) models at the same effective temperature, the authors find that the 1D approximation cannot reproduce the 3D contrasts for K0V and M0V stars: for M0V the relative error longward of 500 nm is about 50% for both umbral and penumbral contrast, and for K0V the RE models miss the water-band structure between 2 and 3 microns in the penumbral contrast and the slope of the umbral contrast beyond 4 microns. For G2V the 1D models work well, with errors below 2% for umbrae and below 10% for penumbrae longward of 500 nm. The paper attributes the discrepancy to the failure of mixing-length convection rather than to horizontal substructures, because horizontally averaged 3D MHD atmospheres reproduce the 3D contrasts while 1D RE atmospheres have steeper vertical temperature gradients.","pith_inferences":["If the paper is right, a likely consequence the authors leave implicit is that water detections or upper limits in transmission spectra of planets around M dwarfs could be spuriously created or erased by unocculted spots, because the ~50% contrast error sits in the 2-3 micron region where the paper's own opacity test shows H2O dominates the spot contrast.","The paper only simulates M0V and notes that later M dwarfs need updated molecular opacities; extending the same MURaM/MPS-ATLAS pipeline to cooler M subtypes is a natural next step, and the mixing-length failure could plausibly grow larger there since convection penetrates the photosphere even more deeply.","The horizontally averaged '1D MHD' atmospheres could be embedded in fast retrieval and stellar-variability codes to capture most of the 3D spectral signal at 1D cost; a natural test would be to compare line profiles from the averaged models against full 3D synthesis for stronger lines, since the paper only validates them at low resolution.","Combining spot MHD models with the already-studied facular MHD models from the same code family would yield a full active-region spectral model, allowing contamination corrections that treat spots and faculae on equal footing rather than spot-only corrections."],"forward_implications":["For M0V stars, the umbral and penumbral flux contrast longward of 500 nm is in error by roughly 50% when 1D radiative-equilibrium models are used, so exoplanet transmission-spectrum corrections built on cooler-star spot spectra are systematically biased at exactly the wavelengths used for molecular-band analysis.","For K0V stars, 1D RE models fail to reproduce the water-band structure in the penumbral contrast between 2 and 3 microns and the slope of the umbral contrast beyond 4 microns, and adjusting the spot effective temperature by ±100 K cannot fix the mismatch.","For G2V stars, 1D RE models remain adequate for umbral and penumbral contrasts, with errors below 2% and 10% respectively longward of 500 nm, so existing solar-type variability models built on 1D spots are not invalidated by this result.","Even for G2V, a single 1D model cannot represent the spectrum of an entire spot (umbra plus penumbra), so two-component spot models with distinct umbral and penumbral temperatures remain necessary.","The discrepancy is driven by the vertical temperature stratification rather than by horizontal inhomogeneities, which means the horizontally averaged 1D MHD atmospheres published for K0V and M0V can serve as a practical substitute for full 3D MHD in low-resolution spectral work."],"supporting_citations":[{"why":"Supplies the MURaM code that produces the 3D MHD spot atmospheres.","marker":"Vögler et al. 2005"},{"why":"Supplies the MPS-ATLAS code used for ray-by-ray spectral synthesis and for the 1D radiative-equilibrium models in the comparison.","marker":"Witzke et al. 2021"},{"why":"Provides the spot simulation setup that the cylindrical flux-tube simulations extend, and explains why spot contrast decreases toward cooler stars.","marker":"Panja et al. 2020"},{"why":"Contributes the two-spot and isolated-spot procedure used to form and preserve a penumbra in the simulations.","marker":"Rempel 2012"},{"why":"Provides the four-group multi-group radiative transfer scheme and the analysis showing that convection and radiation are both important in M-dwarf photospheres.","marker":"Beeck et al. 2013"},{"why":"Defines the mixing-length approximation that the 1D radiative-equilibrium models rely on and whose failure is the paper's identified cause of the discrepancy.","marker":"Böhm-Vitense 1958"},{"why":"Documents the standard practice of representing spots by cooler 1D stellar models in exoplanet transmission-spectroscopy contamination corrections, the baseline the paper challenges.","marker":"Rackham et al. 2023"},{"why":"Showed that 1D radiative-equilibrium models fail for facular contrast in 3D MHD simulations, motivating the analogous test for spots here.","marker":"Witzke et al. 2022"},{"why":"Provides the grid of 1D radiative-equilibrium models from MPS-ATLAS used to construct the comparison models at matched effective temperatures.","marker":"Kostogryz et al. 2023"}],"fun_headline_variants":["First 3D MHD starspot spectra show 50% error on M dwarfs","M-dwarf starspots: 1D models off by half, 3D simulations reveal","Mixing-length spot models fail on M dwarfs: 50% flux contrast error","3D MHD starspot spectra expose mixing-length theory's limits","Starspot spectra from 3D MHD: M dwarfs break 1D models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the MURaM spot simulations faithfully represent real starspots, since the penumbra is sustained by an ad hoc top boundary condition and the spectra are synthesized assuming local thermodynamic equilibrium; if the simulated temperature stratification is unrealistic, the 1D models' failure for K and M dwarfs could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["First 3D MHD starspot spectra show 50% error on M dwarfs","M-dwarf starspots: 1D models off by half, 3D simulations reveal","Mixing-length spot models fail on M dwarfs: 50% flux contrast error","3D MHD starspot spectra expose mixing-length theory's limits","Starspot spectra from 3D MHD: M dwarfs break 1D models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000805,"raw_usage":{"total_tokens":3655,"prompt_tokens":1185,"completion_tokens":2470,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":801,"completion_tokens_details":{"reasoning_tokens":2359}},"tokens_in":801,"tokens_out":2470,"duration_ms":15242,"temperature":1.0,"reasoning_tokens":2359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:35:09.553839+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Observe a spot-crossing event on an M0 dwarf with wavelength-resolved spectroscopy (the paper itself points to TOI-3884 as a candidate and to HST/JWST as capable instruments): if the measured wavelength dependence of the umbral and penumbral contrast follows the 1D radiative-equilibrium prediction rather than the 3D MHD prediction, the central claim would be refuted. A cheaper calculation-based test would be to rerun the M0 spot with a potential-field top boundary condition and see whether the ~50% discrepancy survives.","supporting_citations":[],"review_version":1}