{"id":"dd8a320a-8f46-4cd8-b449-312c4549ead8","arxiv_id":"2509.04359","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In predictive TRANSP simulations of 37 NSTX discharges, MMM matches measured Te and Ti profiles more closely and more consistently than TGLF, whose errors vary strongly with beta.","lead":"This paper compares two reduced turbulence models, MMM and TGLF, in time-dependent TRANSP simulations of 37 high-performing NSTX discharges, finding that MMM reproduces measured electron and ion temperature profiles more consistently than TGLF. It matters because spherical tokamak reactor design relies on reduced models, and the study identifies beta-dependent biases that could distort NSTX-U scenario predictions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MMM-over-TGLF ranking is confounded by solver choice: TRANSP/PT SOLVER results may not reflect intrinsic model performance.","rationale":"The reader's weakest assumption is exactly the solver dependence: the comparison is performed only in PT SOLVER with fixed density/rotation and a boundary at rho=0.7, and the paper's own FUSE results show TGLF can perform much better in a different solver. This is the most load-bearing concern because the abstract's comparative claim and the cost-benefit conclusion ('despite TGLF requiring orders of magnitude greater computational cost') depend on the TRANSP-specific ranking. If a different solver removes or reverses TGLF's deficit, the conclusion that MMM is the more reliable reduced transport model for NSTX-like plasmas does not generalize. The paper acknowledges this in Sec. VII but still leads with the unqualified MMM-over-TGLF statement in the abstract. A secondary concern is MMM's ETG submodel having been calibrated on NSTX data (Sec. I), which weakens the fairness of the comparison but does not invalidate the empirical ranking within this dataset. The concrete test proposed—running both models in a flux-matching solver—would directly resolve whether the ranking is intrinsic to the models or an artifact of the solver implementation. Since the reader already assigned CONDITIONAL based on this concern, my stress-test does not change the verdict.","tokens_in":44813,"tokens_out":5370,"duration_ms":51380,"concrete_test":"Run both MMM and TGLF in the same time-slice flux-matching transport solver (e.g., TGYRO or FUSE) on a representative subset (or all 37) of the NSTX discharges, using identical experimental profile fits, boundary conditions, neoclassical transport (NCLASS or NEO), and grid. If the median RMSE for Te and Ti from TGLF becomes comparable to or lower than MMM, the TRANSP/PT SOLVER-based ranking is solver-dependent and the central claim must be re-scoped to 'within PT SOLVER'. If MMM retains a clear RMSE advantage, the ranking is robust to solver choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that MMM 'more consistently agrees' with NSTX observations than TGLF is made from simulations that couple each turbulence model to a single time-dependent transport solver (TRANSP's PT SOLVER) with density, rotation, and boundary profiles fixed to experimental fits (Sec. II A). The paper itself demonstrates in Sec. V B that TGLF surrogate models evaluated in the FUSE time-slice flux-matching solver achieve far lower RMSE (Te 20%, Ti 16% for TGLF-NN SAT2; Te 13%, Ti 9% for GKNN SAT2) than the TRANSP TGLF results (Te 46%, Ti 25% for EM TGLF SAT0). This improvement conflates changes in solver, turbulence settings, and self-consistent density/rotation prediction, so it does not isolate the turbulence model. But it directly shows that 'TGLF is much worse than MMM' is not a property of TGLF alone. Since the abstract's cost-benefit conclusion rests on this ranking, and Sec. VII explicitly warns against concluding TGLF is inadequate, the headline result is solver-dependent. The fixed rho=0.7 boundary and fixed density/rotation may further bias the comparison by constraining the profiles differently for each model. Without a controlled comparison of MMM and TGLF in the same flux-matching solver, the 'MMM is more reliable' claim is not established for reduced transport models generally.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale validation of two reduced turbulent transport models, MMM and TGLF, in time-dependent predictive TRANSP simulations of 37 high-performing NSTX discharges. Only the electron and ion temperature profiles are evolved inside ρ=0.7; density, rotation, and the edge temperature are fixed to experimental fits, and fast-ion confinement is classical. The central quantitative result is that MMM overpredicts the median Te and Ti profiles by 28% and 27% RMSE, whereas electromagnetic TGLF SAT0 overpredicts Te by 46% with larger variance and underpredicts Ti by 25%, with a pronounced β-dependent Ti flattening. Electrostatic TGLF SAT1 is substantially worse. The authors also present TGLF neural-net surrogates evaluated in the FUSE time-slice flux matcher, which achieve much better agreement with experiment (Te 13–20%, Ti 9–16% RMSE) than the TRANSP TGLF runs, and they explicitly warn that TGLF should not be judged inadequate on the TRANSP results alone. The paper closes with a summary, a discussion comparing with DIII-D, MAST-U, KSTAR, and JET studies, and a list of suggested future work.","tokens_in":45118,"tokens_out":3768,"duration_ms":38441,"significance":"If the central claim holds, the paper provides a practically important benchmark: within the TRANSP/PT SOLVER environment, MMM is the more reliable reduced transport model for high-β NSTX-like plasmas, at orders-of-magnitude lower computational cost. The study uses a large, well-analyzed discharge database, clear profile-error metrics, aggregate statistics, and explicit treatment of model-setting sensitivity. The inclusion of TGLF surrogate models with a different flux-matching solver is a valuable honesty check that demonstrates solver dependence. The main risk is that the headline comparison conflates the transport model with the solver and the fixed-profile setup, and that experimental/profile-fit errors are not quantified. These issues are acknowledged in the text (Sec. II C and Sec. VII), but they are load-bearing for the abstract's general claim that MMM is 'more reliable' than TGLF. With additional controlled comparisons and uncertainty quantification, the paper would be a significant contribution to reduced-model validation for spherical tokamaks.","major_comments":[{"comment":"The paper states explicitly that 'measurement error in the underlying experimental data and possible systematic errors introduced when creating profile fits ... is not quantified in this work.' This is load-bearing because the headline comparison (MMM Te RMSE 28% vs EM TGLF 46%, Table I) is presented as a quantitative ranking. NSTX profile fits can have substantial systematic uncertainty, especially in the core and near the boundary. Without error bars on the experimental profiles, or at least a sensitivity analysis using perturbed fits, the reader cannot tell whether the reported differences are statistically meaningful. Please add an uncertainty estimate or temper the quantitative claims.","section":"Sec. II C and Table I"},{"comment":"The central claim that MMM 'more consistently agrees' with NSTX observations than TGLF is established only within the TRANSP/PT SOLVER setup, where density, rotation, and the ρ=0.7 boundary are fixed to experimental fits. The paper's own Sec. V B shows that TGLF surrogate models in the FUSE flux-matching solver yield Te RMSE 20% and Ti RMSE 16% (TGLF-NN SAT2), much lower than the TRANSP TGLF values 46% and 25%. This demonstrates that the model ranking is solver-dependent and does not reflect an intrinsic property of TGLF. The discussion in Sec. VII correctly warns against concluding TGLF is inadequate, but the abstract and title still make a general claim. A controlled comparison of MMM and TGLF within the same flux-matching solver, or a systematic study of the PT SOLVER iteration and boundary-condition effects, is needed to support the general statement. Without it, the headline should","section":"Secs. II A, V B, VII"},{"comment":"MMM's ETG submodel is stated to have been calibrated against NSTX data (Refs. 70 and 73), and the validation database in this paper is also NSTX. No statement is made about the disjointness of calibration and validation sets, or about sensitivity of the MMM predictions to the calibrated constants. This creates a fairness concern when comparing MMM against TGLF on the same device. Additionally, the chosen TGLF settings (EM SAT0 and ES SAT1) are motivated by prior studies on MAST-U and NSTX, so the comparison may implicitly tune TGLF as well. Please provide evidence that the MMM-over-TGLF ranking is not dominated by this calibration overlap, or explicitly discuss the limitation.","section":"Sec. I and database selection"},{"comment":"The computational-cost comparison is presented as a major advantage of MMM ('orders of magnitude lower cost'). However, Table II reports CPU hours per simulated second in PT SOLVER, which includes the number of Newton iterations and the convergence behavior of the coupled system, not the intrinsic cost of the turbulence models. The wall-clock factor of 64 for TGLF parallelism is also not included in the abstract. The authors do give per-iteration timings, which is helpful, but the cost-benefit conclusion should separate model cost from solver-iteration cost and should acknowledge that surrogate models (Sec. V B) erase much of the cost difference. Please clarify the scope of the cost claim.","section":"Sec. IV, Table II"}],"minor_comments":[{"comment":"Typographical errors: 'sytematically' and 'signficant' should be corrected.","section":"Sec. I"},{"comment":"'TRANSP simulations simulations' is duplicated.","section":"Fig. 9 caption"},{"comment":"The sentence about two outlier discharges with unphysically large energy confinement time is vague; please report which discharges and how the outliers were excluded from statistics.","section":"Sec. III C"},{"comment":"The data archive address is given as '(placeholder)'. This must be replaced with a working DOI or repository link before publication.","section":"Sec. IX"},{"comment":"Typo 'DIIII-D' should be 'DIII-D'.","section":"Sec. VII"},{"comment":"The metric σ in Eq. (3c) is a normalized RMS error, not an RMS error in physical units. Calling it 'RMSE' in Table I is acceptable, but the axes in Fig. 12 and Fig. 3 should make the normalization explicit to avoid confusion with absolute temperature errors.","section":"Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually transparent about its limitations, which is commendable. The main concern is that the abstract and title make a general model-comparison claim that the body of the paper carefully conditions to a specific solver and setup. I would recommend asking the authors to either add a same-solver MMM/TGLF comparison in a flux-matching framework (even for a subset of discharges) or to qualify the central claim in the title/abstract. Also, the unquantified experimental uncertainty and the NSTX-calibration overlap for MMM are genuine fairness issues that should be addressed explicitly. The paper is otherwise rigorous in its statistics and discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a careful read if you work on ST integrated modeling. The new thing here is a 37-discharge, time-dependent TRANSP comparison of MMM against TGLF for NSTX, with a clear cost figure: MMM sits around 0.7 CPU-hours per simulated second versus thousands for TGLF. The headline result—MMM overpredicts Te by 28% median RMSE, EM TGLF by 46% and with much larger variance—is believable and useful for anyone needing a default reduced model in TRANSP for NSTX-U scenario prep. The analysis of TGLF's beta-dependent Ti flattening is also substantive, not a side remark.\n\nThe paper is honest in ways that matter. It explicitly says in Sec. VII that you should not conclude TGLF is inadequate for spherical tokamaks, and it shows in Sec. V B that TGLF surrogate models in a flux-matching solver (FUSE) do far better than TGLF inside PT SOLVER—Te RMSE 20% versus 46%. So the MMM-over-TGLF ranking is real for this solver and setup, but it is not a clean statement about intrinsic model quality. The abstract slightly oversells by stating the ranking without the solver caveat up front; a reader might miss the FUSE part and take away a false model-level conclusion. I'd suggest the authors add one sentence in the abstract.\n\nMain soft spots, in order: (1) The ETG submodel of MMM was calibrated on NSTX data, and the validation database is NSTX, with no statement on overlap; that is an in-sample advantage and should be quantified or bounded. (2) Experimental/profile-fit errors are explicitly not quantified, which weakens the RMSE significance for a comparison that is often within a few tens of percent. (3) The data archive is a placeholder; the run IDs are given, but a real archive should be required before publication. (4) The neural net results are in-sample evaluations, though they are clearly labeled as preliminary.\n\nMy bottom line: the central claim holds up within its scope, the paper is a legitimate benchmark with honest limitations, and it belongs in peer review. I would send it to a referee, require the data release and a response on the calibration overlap, and ask the authors to tighten the abstract. It deserves to be in the literature and will get cited.","headline":"Solid, honestly-scoped TRANSP benchmark: MMM beats TGLF in PT SOLVER on 37 NSTX shots and costs far less, but don't read the abstract as a model-vs-model verdict; the paper's own FUSE results and section VII provide the needed caveat.","tokens_in":45735,"tokens_out":1630,"would_cite":true,"duration_ms":21427,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["52.55.Fa"],"model":"deepseek-v4-flash","headline":"Time-dependent TRANSP simulations of 37 high-performing NSTX discharges show the Multi-Mode Model reproduces electron and ion temperature profiles more consistently than the Trapped Gyro-Landau Fluid model, at orders-of-magnitude lower comp","keywords":["tokamak transport","spherical tokamak","NSTX","reduced transport models","TGLF","Multi-Mode Model","temperature profile prediction","electromagnetic turbulence"],"falsifier":"Take the same set of NSTX discharges and run both MMM and electromagnetic TGLF in a time-slice flux-matching solver, as the paper did for TGLF surrogates, and compare median RMSEs; if TGLF's median errors drop below MMM's roughly 28/27% across the full database, the paper's claimed ordering is a property of PT SOLVER rather than of the turbulence models.","tokens_in":44698,"feed_emoji":"⚛️","tokens_out":5786,"duration_ms":54533,"temperature":0.7,"pith_summary":"This paper tests whether two mature reduced turbulent transport models, MMM and TGLF, can predict electron and ion temperature profiles in high-performing spherical tokamak plasmas when run inside time-dependent TRANSP simulations. It claims that MMM agrees with NSTX observations more consistently than TGLF, with median profile errors around 28% for Te and 27% for Ti, while TGLF overpredicts Te by 46% and underpredicts Ti by 25%, with a strong and experimentally conflicting beta-dependent flattening of Ti. The paper also shows that TGLF's agreement improves substantially when run in a time-slice flux-matching solver with neural-net surrogates instead of TRANSP's time-dependent solver, indicating that the ranking is partly tied to the simulation method. The practical stakes are which reduced model can be trusted for scenario development and reactor design in the spherical tokamak regime, where TGLF is far more expensive than MMM.","feed_headline":"MMM beats TGLF on NSTX profiles at a fraction of the cost","feed_subtitle":"MMM median errors are ~28%; TGLF overpredicts Te by 46% and flattens Ti at high beta.","key_machinery":"The central object is the comparison of two reduced turbulent transport models coupled to TRANSP's implicit time-dependent solver PT SOLVER. MMM combines submodels for ITG/TEM/KBM, ETG, microtearing modes, and drift-resistive inertial ballooning modes, while TGLF is a trapped gyro-Landau fluid model with quasilinear saturation rules; both include electromagnetic effects and E×B shear. The load-bearing comparison is quantified by root-mean-square profile errors, relative offset errors, peaking factors, and correlations with beta, and is complemented by a time-slice flux-matching solver with neural-net surrogates to separate model error from solver error.","core_discovery":"For a large, well-analyzed set of high-performing NSTX discharges, predictive TRANSP simulations using MMM produce temperature profiles closer to experiment than electromagnetic TGLF with SAT0, and far closer than electrostatic TGLF with SAT1. Median RMSEs are 28% for Te and 27% for Ti with MMM, versus 46% and 25% with electromagnetic TGLF; electrostatic TGLF overpredicts Te by 93%. TGLF's predictions have a strong beta dependence: as beta rises, TGLF predicts lower Te and progressively flatter Ti profiles, in conflict with NSTX data, while MMM shows weaker and more consistent trends. Stored energy is predicted to about 18% median error by both MMM and electromagnetic TGLF. The paper additio","pith_inferences":["I read MMM's advantage as partly a calibration effect: MMM's ETG submodel was previously fitted to NSTX data, so part of its 28/27% performance may encode device-specific knowledge that TGLF lacks.","The striking gap between TRANSP and flux-matching TGLF results suggests that validation studies should report solver-induced error separately, or they risk attributing a solver artifact to the physics model.","A like-for-like comparison could be made by training neural-net surrogates for MMM on the same NSTX time slices and running them in the same flux-matching solver used for TGLF, which would isolate the turbulence-model contribution to the ranking.","If TGLF's beta-dependent Ti flattening persists across improved saturation rules, it may point to a missing stabilization mechanism in high-beta spherical tokamaks that the current MMM submodel set captures."],"forward_implications":["MMM becomes a practical reduced model for full-pulse, non-inductive scenario development on NSTX-U, with characteristic Te and Ti errors near 28% and stored energy within about 18%.","TGLF-based time-dependent TRANSP predictions for high-beta spherical tokamaks should be treated cautiously, since its beta-dependent Ti flattening and large Te error variance can produce opposite errors depending on beta.","Electrostatic TGLF is not suitable for NSTX-class plasmas; the factor-of-two Te overprediction shows electromagnetic fluctuations must be retained in reduced models at high beta.","Conclusions about turbulence-model quality are entangled with the solver choice: the same TGLF physics agrees much better in a time-slice flux-matching framework than in PT SOLVER.","TGLF's orders-of-magnitude higher CPU cost makes it impractical for full-pulse integrated modeling, unless fast surrogate models are embedded into a time-dependent transport solver."],"supporting_citations":[{"why":"Defines TGLF, the trapped gyro-Landau fluid model and its quasilinear saturation rules, which is one of the two models being tested.","marker":"[46, 47]"},{"why":"Defines the Multi-Mode Model, MMM, with its combined submodels for ITG/TEM/KBM, ETG, MTM, and drift-resistive ballooning modes.","marker":"[67, 68]"},{"why":"Describes the PT SOLVER implicit transport solver in TRANSP, the time-dependent predictive framework used for both models.","marker":"[83-85]"},{"why":"A large database study that motivated the choice of electromagnetic TGLF SAT0 as the primary TGLF configuration.","marker":"[59]"},{"why":"Time-slice flux-matching analysis with electrostatic TGLF SAT1 on low-beta NSTX discharges, providing the basis for the electrostatic TGLF configuration.","marker":"[60]"},{"why":"A large-database validation of time-dependent TGLF predictions on a conventional tokamak, serving as the baseline for comparing NSTX results.","marker":"[66]"},{"why":"Describes the TGYRO time-slice flux-matching transport solver used in companion time-slice analyses.","marker":"[96]"},{"why":"Introduces neural network surrogate models for TGLF that enable the CPU-time study and the broad time-slice survey of TGLF settings outside TRANSP.","marker":"[110, 111]"}],"fun_headline_variants":["MMM tops TGLF on NSTX temperature profiles at low cost","Cheap model outperforms costly TGLF on NSTX profiles","High-beta TGLF flops; MMM matches NSTX temperatures","Simple model beats sophisticated TGLF for NSTX plasmas","MMM's modest errors beat TGLF's beta-dependent misses"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The ranking rests on PT SOLVER's time-dependent setup, where only Te and Ti are evolved inside rho=0.7 while density, rotation, magnetic equilibrium, and fast-ion confinement are fixed to experimental or classical inputs; the paper itself shows TGLF's agreement improves markedly when a time-slice flux-matching solver replaces that setup, so the ranking may reflect solver choice as much as model quality.","fun_headline_variants_meta":{"raw":{"variants":["MMM tops TGLF on NSTX temperature profiles at low cost","Cheap model outperforms costly TGLF on NSTX profiles","High-beta TGLF flops; MMM matches NSTX temperatures","Simple model beats sophisticated TGLF for NSTX plasmas","MMM's modest errors beat TGLF's beta-dependent misses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1227,"prompt_tokens":945,"completion_tokens":282,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":689,"completion_tokens_details":{"reasoning_tokens":186}},"tokens_in":689,"tokens_out":282,"duration_ms":3144,"temperature":1.0,"reasoning_tokens":186,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:13:25.054770+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same set of NSTX discharges and run both MMM and electromagnetic TGLF in a time-slice flux-matching solver, as the paper did for TGLF surrogates, and compare median RMSEs; if TGLF's median errors drop below MMM's roughly 28/27% across the full database, the paper's claimed ordering is a property of PT SOLVER rather than of the turbulence models.","supporting_citations":[{"cited_title":"Candy, C","cited_arxiv_id":null,"evidence_quote":"Describes the TGYRO time-slice flux-matching transport solver used in companion time-slice analyses."}],"review_version":1}