{"id":"80fa4c3b-11ac-4041-9b1e-60d860804ab5","arxiv_id":"2502.00187","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"At low tungsten concentrations, SPARC's full-field H-mode is predicted to sustain fusion gain above 5, and an 8-tesla mode is predicted to reach breakeven-level gain above 1.","lead":"Computer simulations map how uncertain assumptions change the predicted fusion performance of SPARC, a compact fusion experiment. They find that the full-field design could consistently reach fusion gain above 5 at low tungsten levels, and that a reduced-field mode could reach breakeven, helping plan early operations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Q>5 consistency claim hinges on the EPED-NN pedestal being unbiased, yet the paper itself labels EPED an upper limit and the sensitivity scan stops at ptop/pEPED=0.8; a systematic pedestal overprediction would push many cases below Q=5.","rationale":"The reader's weakest assumption is the EPED-NN pedestal model, and my independent reading arrives at the same point. The central claim is that Q>5 is consistently achieved at 11 MW below a certain W concentration; because the EPED-NN sets the pedestal top pressure that seeds the entire core profile, and because the paper's own sensitivity scan covers only ptop/pEPED from 0.8 to 1.2, the claim is only robust if EPED-NN is approximately unbiased. The paper's admission that EPED is likely an upper limit directly undercuts that assumption. A 20% overprediction would place real pedestals at the bottom edge of the scanned range; the roughly linear dependence in Figure 3 then predicts many cases with Q<5. This is not an internal inconsistency, but it is a correctness risk tied to an unpublished, unvalidated surrogate. The reader's CONDITIONAL verdict already captures this risk, so I do not change the verdict. The proposed test - extending the scan to lower ptop/pEPED - is a direct way to see whether the headline claim survives the paper's own caveat; if it does not, the abstract should be softened to reflect that Q>5 is contingent on the EPED pedestal being realized, rather than consistently achieved across the assumed input ranges.","tokens_in":21799,"tokens_out":4823,"duration_ms":48681,"concrete_test":"Re-run the PRD 11 MW low-W database with ptop/pEPED sampled uniformly in [0.5,1.0] (for example, 50 additional converged cases at fW=1.5e-5), and count the fraction with Q>5. If that fraction drops below 90%, the 'consistently Q>5' claim does not survive the paper's own stated upper-limit caveat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 states that 'the ELMy H-mode pedestal predicted by EPED is likely to be an upper limit for safe operation,' but the headline 'Q>5 consistently achieved at 11 MW' is established from simulations whose pedestal top pressure is set by the EPED-NN and varied only within ptop/pEPED in [0.8,1.2]. Figure 3 shows Q rising roughly linearly with ptop/pEPED, so a systematic overprediction of only 20-30% by EPED would place real cases at or below the lower edge of the scanned range, where the linear trend suggests Q < 5 for a substantial fraction of the database. The EPED-NN is trained on 2000 unpublished EPED simulations with no released validation set, so the sign and magnitude of any bias cannot be assessed from the manuscript. Because the central performance claim scales directly with pedestal height, this is the most load-bearing uncertainty; the converged-subset caveat is secondary because the low-W filter largely removes the radiative-collapse cases.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an integrated modeling study of two SPARC H-mode scenarios: the 12.2 T Primary Reference Discharge and a reduced-field 8 T H-mode. Using ASTRA with TGLF-SAT2 core transport and a neural-network surrogate of EPED for the pedestal, the authors build databases by randomly permuting four input assumptions (W concentration, DT fuel fraction, ion-to-electron temperature ratio at the pedestal top, and pedestal pressure relative to the EPED prediction). They then perform scans in ICRH power and pedestal density. The central claims are that, for the 12.2 T PRD, Q>5 is consistently achieved at 11 MW auxiliary power once W concentration is below a threshold, and Q>2 is assured at higher input power; for the 8 T scenario, a Q>1 breakeven-relevant window exists at low W concentration, with Q=1.4 at fG=0.46 and 25 MW. The paper also presents statistical turbulence-spectrum analyses showing ITG dominance and marginal ETG activity.","tokens_in":21948,"tokens_out":3266,"duration_ms":33700,"significance":"If the underlying pedestal model is unbiased, this is a valuable result for SPARC planning: it provides a quantitative mapping of the performance variability induced by realistic input uncertainties, and it identifies pedestal pressure and Ti/Te at the pedestal top as the dominant sensitivities. The work is also a useful methodological demonstration of medium-fidelity integrated modeling with random input sampling and statistical output analysis, including explicit treatment of radiative collapse and L-H power balance. The strengths are the large database (284 PRD and 192 8T simulations), the self-consistent coupling of core transport, pedestal, and radiation, and the honest reporting of convergence statistics and model limitations. However, the headline performance claims rest on an unpublished EPED-trained neural network, on simulation convergence that is systematically biased toward low-W and high-pedestal cases, and on L-H power scalings with acknowledged wide error bars. These issues do not invalidate the sensitivity study, but they do affect the strength of the absolute Q conclusions.","major_comments":[{"comment":"","section":"Sec. 2 and Sec. 4 (Fig. 3)"},{"comment":"","section":"Sec. 4 (Fig. 1) and Sec. 5 (Fig. 15)"},{"comment":"","section":"Sec. 4.2 and Abstract"}],"minor_comments":[{"comment":"","section":"Sec. 3"},{"comment":"","section":"Table 3"},{"comment":"","section":"Sec. 4.2"},{"comment":"","section":"Sec. 5.1"},{"comment":"","section":"Sec. 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful sensitivity study and the authors are transparent about many caveats, but the central quantitative claims are more conditional than the abstract suggests. The reliance on an unpublished EPED-NN, the biased convergence statistics, and the use of L-H scalings with wide error bars collectively mean that the Q>5 claim is not yet established to the level implied by 'consistently achieved.' I would encourage the editor to request a revised version that either adds validation/benchmarking of the pedestal model, restricts the headline claims to the converged, H-mode-sustained subset with uncertainty bounds, or both. The paper's methodology and database are valuable and within the scope of the journal, so rejection is not warranted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Sanjay — quick take on Muraca et al. The paper is a genuinely useful systematic sensitivity study of SPARC H-mode performance, and it's more honest about its own limits than most. The new content is the database itself, the ranking of input assumptions (ptop/pEPED and Ti/Te at the pedestal are the big levers), and the 8 T operational window at fG=0.46 and 25 MW showing Q≈1.4. The TGLF spectrum analysis is thorough, and the density and power scans are logically laid out. I found no circular derivation: the four varied inputs are not fitted to Q, and the EPED-NN is a surrogate for an external pedestal model. The self-citation is heavy but appropriate given the prior SPARC modeling by this group.\n\nThe soft spot is the one the stress-test flags, and it's load-bearing. The whole Q>5 consistency claim at 11 MW rests on the EPED-NN pedestal height. The paper itself says EPED is 'likely to be an upper limit for safe operation,' and the scan only varies ptop/pEPED down to 0.8. Figure 3 shows Q roughly linear in ptop/pEPED, so a 20-30% systematic overprediction would push a substantial fraction of the database below Q=5. The NN is trained on 2000 unpublished EPED runs, no validation set is shipped, so you can't assess the bias from the manuscript. That's an addressable weakness, not a dealbreaker, but it should be fixed before publication: release the validation set, or at least show the EPED-NN against existing pedestal data.\n\nThe non-convergence filter is a secondary issue. 26% of PRD and 46% of 8T runs didn't converge, and the survivors favor low W and high pressure. The paper is transparent about this, and the low-W filter for the Q>5 claim does remove most radiative collapse cases, so it's not disqualifying. The L-H scalings have wide error bars and the paper acknowledges that.\n\nIf I were refereeing, I'd ask for the EPED-NN validation material and a robustness check on the pedestal bias before accepting. But this is a serious paper that deserves referee time. I'd cite it for the 8T window and the parameter ranking.","headline":"Honest and useful sensitivity study, but the Q>5 claim is only as good as the unpublished EPED-NN pedestal, which the paper itself flags as an upper limit.","tokens_in":22614,"tokens_out":6305,"would_cite":true,"duration_ms":41017,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Modeling finds SPARC 12.2 T H-modes sustain Q>5 when tungsten stays low.","keywords":["SPARC","H-mode","fusion gain","pedestal model","tungsten concentration","transport modeling","sensitivity study","breakeven"],"falsifier":"Measure the fusion gain and pedestal top pressure in a SPARC H-mode pulse at 12.2 T, 11 MW of ICRH, fG≈0.37, and low tungsten concentration. If the measured Q falls below 5, or if the achieved pedestal top pressure is more than about 20% below the EPED-NN prediction, the paper's consistency claim is falsified.","tokens_in":21503,"feed_emoji":"🔥","tokens_out":8906,"duration_ms":81299,"temperature":0.7,"pith_summary":"The paper builds a large database of integrated transport simulations for two SPARC H-mode scenarios and asks how much the predicted fusion gain varies when four uncertain inputs are moved within realistic ranges. It finds that, for the 12.2 T primary reference discharge at 11 MW of auxiliary heating, essentially every converged low-tungsten case reaches Q>5, and all converged cases exceed Q=2; increasing heating power to 25 MW keeps the plasma in H-mode and preserves Q>2 even under pessimistic radiation assumptions. For the 8 T scenario, a low-tungsten window with Q>1 exists, with a chosen operational point at fG=0.46 and 25 MW giving Q about 1.4, making it a candidate for early breakeven operation. The two strongest performance drivers are the pedestal top pressure relative to the EPED prediction and the ratio of ion to electron temperature at the pedestal top; tungsten concentration mainly decides whether the plasma radiates away too much power and collapses. The practical point is that SPARC's predicted performance is sensitive enough to modeling assumptions that sensitivity studies should accompany any single-point prediction.","feed_headline":"SPARC 12T H-modes sustain Q>5 when tungsten stays low","feed_subtitle":"A 284-run sensitivity database maps an 8T breakeven window at Q≈1.4 for early SPARC operations.","key_machinery":"The arguments are carried by an integrated modeling chain: ASTRA evolves the stationary flat-top profiles; TGLF with the SAT2 saturation rule and electromagnetic effects supplies core turbulent fluxes; a neural network trained on EPED peeling-ballooning simulations sets the pedestal height and width self-consistently; and Kadomtsev sawtooth mixing plus precomputed ICRH deposition profiles fix the remaining profile physics. The database is produced by randomly and uniformly permuting four uncertain inputs—tungsten fraction f_W, DT fraction f_DT, ion-to-electron temperature ratio at the pedestal top T_i,top/T_e,top, and pedestal pressure ratio p_top/p_EPED—and then scanning auxiliary power and pedestal density around the reference points. The key diagnostic quantities are Q=P_fus/(P_aux+P_Ohm) and the H-mode sustainment ratio f_LH=P_sep/P_LH.","core_discovery":"The central claim is that SPARC's nominal 12.2 T H-mode has a robust burning-plasma operating window: with 11 MW of ICRH, fG=0.37, and tungsten fraction below the radiative-collapse threshold, the converged database gives Q=P_fus/(P_aux+P_Ohm)>5 consistently, and every converged point in the initial 284-simulation database lies above Q=2. Raising the auxiliary power to 25 MW does not change the fusion power, because the core profiles are stiff, but it restores P_sep above the empirical L-H thresholds, and the converged points then all satisfy Q>3. The paper also claims that the 8 T H-mode, with H-minority heating, 25 MW, and fG=0.46, reaches Q about 1.4 in the overlap region where Q>1 and sustained H-mode both hold, giving a plausible early-operation breakeven scenario. The database shows Q increases roughly linearly with p_top/p_EPED and with T_i,top/T_e,top, while f_W acts mainly as a kill switch through radiation and H-mode loss rather than as a continuous performance scaler.","pith_inferences":["Because Q scales roughly linearly with p_top/p_EPED, the real margin for the 8 T breakeven point is thin: if the achieved pedestal is even about 10% below the EPED value, Q=1.4 could drop below 1.","The paper's consistency statement covers converged simulations only; high-tungsten cases radiatively collapse and never reach a steady solution, so the Q>5 window is conditional on keeping tungsten below the collapse threshold, not just on the four varied parameters.","The same sensitivity method could become an operational planning tool: given real-time estimates of tungsten concentration and pedestal pressure, database maps could predict whether SPARC stays above Q=5 or Q=1 without launching new transport simulations.","The disagreement between the two L-H scalings on the 8 T scenario suggests that the ion-electron heat partition and radiation set the effective L-H threshold; a dedicated measurement of the ion power at the separatrix in early 8 T pulses would constrain which scaling applies."],"forward_implications":["At 12.2 T, 11 MW of ICRH, and fG=0.37, all converged low-tungsten simulations exceed Q=2 and almost all exceed Q=5, implying a substantial burning-plasma window for the nominal SPARC discharge.","Raising ICRH power to 25 MW does not increase fusion power, because the core profiles are stiff, but it moves P_sep/P_LH upward so H-mode is sustained even at high tungsten; all converged cases stay above Q=3.","Increasing pedestal density raises fusion power and Q at fixed auxiliary power, because the whole profile shifts upward, but it also raises the L-H threshold, so H-mode sustainment becomes less certain.","For the 8 T scenario, Q>1 and sustained H-mode overlap only in a narrow region; the selected point at fG=0.46 and 25 MW gives Q about 1.4, identifying a candidate early-operation breakeven scenario.","The turbulence is consistently ITG-dominated with marginal high-k electron-scale activity; increased density raises the high-k growth rate but not the electron heat flux, so the core transport picture is stable across the database."],"supporting_citations":[{"why":"Defines the SPARC reference discharge and the POPCON-identified operational point (11 MW, fG=0.37) whose Q=11 prediction this study tests.","marker":"[11]"},{"why":"Supplies the TGLF SAT2 quasi-linear core transport model, with electromagnetic effects and three species, used for every database run.","marker":"[12]"},{"why":"Provides the EPED peeling-ballooning pedestal model that the neural network surrogate is trained on.","marker":"[13]"},{"why":"Provides the ASTRA integrated transport framework that evolves profiles to flux-matched flat-top solutions.","marker":"[15]"},{"why":"Establishes the methodology for training a neural network on EPED results to predict pedestal height and width quickly.","marker":"[29]"},{"why":"Gives one of the two L-H power threshold scalings used to judge whether each case sustains H-mode.","marker":"[46]"},{"why":"Gives the alternative L-H threshold scaling used for comparison, especially for the 8 T scenario.","marker":"[47]"},{"why":"Provides the high-fidelity gyrokinetic results used to benchmark the TGLF density peaking in the primary reference discharge.","marker":"[58]"}],"fun_headline_variants":["SPARC 12T H-modes: low W yields Q>5 consistently","Sensitivity scan: SPARC Q>5 at 12T if tungsten low","8T SPARC reaches Q~1.4, a potential breakeven","284-run database maps SPARC burning plasma window","SPARC Q>5 robust at 12T with clean plasma"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire prediction hangs on the EPED-trained neural network giving the correct pedestal height, and because the paper itself cautions that EPED is likely an upper limit for safe operation, a real SPARC pedestal noticeably below the prediction would erase the Q>5 margin even with low tungsten.","fun_headline_variants_meta":{"raw":{"variants":["SPARC 12T H-modes: low W yields Q>5 consistently","Sensitivity scan: SPARC Q>5 at 12T if tungsten low","8T SPARC reaches Q~1.4, a potential breakeven","284-run database maps SPARC burning plasma window","SPARC Q>5 robust at 12T with clean plasma"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0003,"raw_usage":{"total_tokens":1830,"prompt_tokens":1142,"completion_tokens":688,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":758,"completion_tokens_details":{"reasoning_tokens":592}},"tokens_in":758,"tokens_out":688,"duration_ms":6981,"temperature":1.0,"reasoning_tokens":592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T19:52:54.380328+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the fusion gain and pedestal top pressure in a SPARC H-mode pulse at 12.2 T, 11 MW of ICRH, fG≈0.37, and low tungsten concentration. If the measured Q falls below 5, or if the achieved pedestal top pressure is more than about 20% below the EPED-NN prediction, the paper's consistency claim is falsified.","supporting_citations":[{"cited_title":"The weak linear relationship with both the variables suggests that Qe,low−k is impacted by TEM activity","cited_arxiv_id":null,"evidence_quote":"Supplies the TGLF SAT2 quasi-linear core transport model, with electromagnetic effects and three species, used for every database run."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the EPED peeling-ballooning pedestal model that the neural network surrogate is trained on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ASTRA integrated transport framework that evolves profiles to flux-matched flat-top solutions."},{"cited_title":"of Physics Publ) ISBN 9780750305891","cited_arxiv_id":null,"evidence_quote":"Establishes the methodology for training a neural network on EPED results to predict pedestal height and width quickly."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the alternative L-H threshold scaling used for comparison, especially for the 8 T scenario."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the high-fidelity gyrokinetic results used to benchmark the TGLF density peaking in the primary reference discharge."}],"review_version":1}