{"id":"127c247e-8fc9-4e35-b672-a56d184f84cb","arxiv_id":"2507.16960","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A DeepONet surrogate trained on Vlasov-Poisson simulations reproduces electric field energy evolution of Landau damping for unseen temperatures with relative L2 errors near 0.5% to 0.8%.","lead":"This paper trains a Deep Operator Network to predict how the electric field energy of a plasma evolves during Landau damping, using data from kinetic Vlasov-Poisson simulations. It shows low interpolation errors for unseen temperature values in single-mode and five-mode setups, including a nonlinear regime.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aggregate L2 error norm is dominated by early linear decay and may not substantiate the claimed nonlinear-regime accuracy; time-windowed errors are required.","rationale":"The paper's narrow interpolation claim, predicting electric field energy from temperature for fixed initial perturbations, is plausible and the reported full-interval errors are consistent with that claim. However, the central claim explicitly includes accurate capture of nonlinear Landau damping, and the metric used to support that claim, Eq. 9, is dominated by the linear phase because E(t) decays exponentially. This is a measurement issue rather than an attack on the reference data or the smoothness of the mapping. The reader's weakest assumption concerned Gkeyll accuracy and the sufficiency of temperature as the only varied parameter; this concern is related but distinct, so I mark partial agreement. A concrete time-windowed error analysis would settle whether the nonlinear-regime claim holds. Since the reader's verdict is already conditional and this concern adds a specific technical condition rather than overturning the central interpolation result, the verdict remains unchanged: the paper should be accepted only after the authors provide late-time or pointwise-relative error statistics. I do not see a reason to move to accept, reject, or unverified based on this single concern, because the aggregate metric could in principle be masking a real failure and the authors have not released the artifacts needed for independent verification.","tokens_in":9204,"tokens_out":4261,"duration_ms":53985,"concrete_test":"Recompute the relative L2 errors for the five-mode case in disjoint time windows, e.g., t in [0,10] and [10,40], and for the single-mode case in t in [0,5] and [5,20], using the same formula as Eq. 9 but restricted to each window. Also plot the pointwise relative error |E(t)-G(T)(t)|/E(t) for the test samples shown in Figures 3 and 5. If the late-time-window errors exceed roughly 10% while the full-interval errors remain below 1%, the claim that the nonlinear regime is accurately captured is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative evidence for nonlinear accuracy is the relative L2 error defined in Eq. 9, computed over the full time interval. Because E(t) decays exponentially, both the numerator and denominator of Eq. 9 are dominated by the first few omega_pe^-1 of the linear phase. For the single-mode case with k lambda_e = 0.35, the damping rate is substantial, so by t = 20 the energy has fallen by many orders of magnitude; the five-mode case extends to t = 40 with similarly strong decay. Late-time nonlinear features such as trapped-particle oscillations and phase-space-hole-driven energy exchange therefore contribute almost nothing to the reported mean errors near 1%. This means the headline claim that DeepONets accurately capture the nonlinear regime is not actually tested by the tabulated metric. The plotted log-scale comparisons in Figures 3 and 5 may look good even if late-time relative errors are large, because a log plot compresses small absolute differences. This is not a question of Gkeyll accuracy or smoothness of the temperature-to-energy map; it is a property of the chosen error norm. The conclusion explicitly claims success in both linear and nonlinear regimes, so the paper should report errors that isolate the nonlinear phase.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript constructs DeepONet surrogates that map a scalar temperature parameter T in [0.5, 1.5] to the electric-field-energy time series E(t) for Vlasov-Poisson Landau damping. The training data are generated with the Gkeyll continuum Vlasov code for two configurations: a single mode (k lambda_e = 0.35, A = 0.05) evolved to t = 20 omega_pe^{-1}, and a five-mode superposition (Table 1) evolved to t = 40 omega_pe^{-1}. The authors report mean relative L2 errors (Eq. 9) below 1% on held-out temperature samples and interpret this as evidence that DeepONets capture both linear and nonlinear Landau damping. The paper also reports loss convergence and an error-versus-training-set-size study for the five-mode case.","tokens_in":9459,"tokens_out":7079,"duration_ms":78397,"significance":"If the accuracy claim were supported, this would be a useful proof-of-concept for replacing repeated kinetic Vlasov simulations with a cheap surrogate for parameter studies. The paper's strengths are its use of first-principles kinetic reference data, a clean train/test split, and a systematic study of training-set size. The caveat is that the reported error metric is dominated by the early linear phase, so the headline claim about nonlinear-regime accuracy is not yet substantiated; the scope of the 'generalization' claim also exceeds the experiments. With phase-resolved error reporting and more cautious wording, the result would be a solid, if narrow, contribution to surrogate modeling of kinetic plasma physics.","major_comments":[{"comment":"The relative L2 error is computed over the full time interval, but because E(t) decays substantially in the linear phase, both the numerator and denominator of Eq. (9) are dominated by the early, larger-amplitude portion of the time series. Late-time trapped-particle oscillations and phase-space-hole dynamics therefore contribute almost nothing to the reported mean errors. The claim in Section 5 that the model 'accurately captures the evolution of electric field energy in both linear and nonlinear regimes' is thus not tested by the tabulated metric. Please report time-windowed relative L2 errors restricted to the nonlinear epoch (e.g., after the linear decay has saturated), or normalized errors computed separately for the linear and nonlinear phases, and include pointwise error time series so the nonlinear phase is visible.","section":"Section 4, Eq. (9) and Tables 3-4"},{"comment":"The statement that the model generalizes 'across varying initial conditions and perturbations' overstates what is demonstrated. The only varying parameter is temperature in [0.5, 1.5]; the initial perturbation amplitudes and wavenumbers are fixed (Table 1 and Eq. 8). The evidence supports interpolation over temperature for this specific fixed perturbation set, not generalization over perturbations. Please either qualify the conclusion and abstract or add experiments that vary A_i and k_i.","section":"Section 5 (Conclusion) and Abstract"},{"comment":"The claimed 'significant speedup compared to the conventional numerical solver' is not quantified: no Gkeyll wall-clock time is reported for the test dataset, and no baseline surrogate (e.g., a linear-damping analytic model, principal-component regression, or a standard MLP) is compared. Without a baseline, the reader cannot tell whether the DeepONet architecture is essential to the reported accuracy or whether the small errors simply reflect a smooth, low-dimensional temperature-to-E(t) map. Please add a baseline comparison and report Gkeyll runtimes for the same test cases.","section":"Section 4.2, final paragraph"}],"minor_comments":[{"comment":"Please specify the Gkeyll numerical configuration (domain length, velocity-grid resolution, time step, boundary conditions) and provide a convergence check for the five-mode nonlinear case, since the accuracy of the reference data is the basis for the reported errors.","section":"Section 3.2"},{"comment":"The branch input is described as 'temperature values,' but temperature is a scalar parameter rather than a function over a spatial domain; please clarify how the branch sensors are defined (e.g., a constant function sampled at m points) so that the operator-learning formulation is precise.","section":"Section 2 and Figure 1"},{"comment":"The operator N in N(T,E) = 0 is not defined; please state explicitly that it denotes the Vlasov-Poisson residual operator, and define the function spaces for T and E.","section":"Eq. (1)"},{"comment":"Unlike Table 3, Table 4 reports only mean errors for the five-mode case; please add standard deviations or min/max ranges for each training-set size, since mean errors alone may hide poor per-sample performance in the nonlinear regime.","section":"Table 4"},{"comment":"Because the vertical axes are logarithmic, late-time differences are visually compressed; please add linear-scale insets or error-versus-time panels for the nonlinear phase so the reader can assess the late-time agreement directly.","section":"Figures 3 and 5"},{"comment":"Please correct minor language issues: 'notable for by their ability' in Section 5, 'permitivity' in Section 3.1, and 'dynamical evolution' style repetitions.","section":"Throughout"},{"comment":"For reproducibility, please state the number of epochs, batch size, and hardware details for the reported per-epoch training time, and make the trained models and datasets available or specify a repository.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope, and the underlying train/test procedure appears sound. My main concern is that the headline nonlinearity claim is not supported by the reported error metric; this is fixable with time-windowed errors and more cautious language, so I would not reject on that basis. The revision should focus on evidence rather than framing, and the authors should avoid overclaiming generalization beyond temperature interpolation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, narrowly scoped demonstration that a DeepONet can fit the electric field energy trajectory of 1D electrostatic Landau damping as a function of temperature, with test errors near 1%. The honest contribution is extending ML surrogates for this problem from the linear-only PINN of Qin et al. to a five-mode case that shows nonlinear energy oscillations. That is real, but smaller than the abstract suggests.\n\nWhat works well: the data generation is solid (Gkeyll Vlasov-Poisson), the training curves and error tables are standard, and the convergence trend with more training samples is sensible. The visual agreement in the linear phase is convincing, and the interpolation over held-out temperatures appears to hold.\n\nSoft spots, in order of importance. First, the reported relative L2 error (Eq. 9) is dominated by the early linear decay, because the energy falls by orders of magnitude. Late-time nonlinear features contribute almost nothing to the norm. So the claim that the surrogate 'accurately captures' the nonlinear regime is not actually supported by the numbers; the plots may look good on a log scale even if late-time relative errors are large. The paper should report time-windowed errors, e.g., t > 20 for the five-mode case, or errors normalized per segment. Second, the input is a scalar temperature, not a function, so calling this an operator network is a stretch. A simple MLP or even a polynomial baseline might perform similarly, and no baseline is tested. Third, the generalization language in the abstract and conclusion overstates what is varied: only temperature changes, while perturbation amplitudes and wavenumbers stay fixed. Finally, no code or data are released, so the exact numbers can't be checked.\n\nNone of these are fatal. The core interpolation claim is probably fine. But the paper as written overreaches in its conclusions, and the nonlinear accuracy is asserted rather than measured.\n\nWho is this for? Plasma physicists thinking about ML surrogates for kinetic processes, and method developers looking for a clean benchmark application. It deserves a serious referee, though I'd expect a request for revision rather than acceptance as is.","headline":"A clean but narrow interpolation study: DeepONets fit Landau damping energy over temperature, but the headline nonlinear-accuracy claim is not backed by the reported error metric.","tokens_in":9964,"tokens_out":2557,"would_cite":false,"duration_ms":41814,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Deep Operator Network trained on Vlasov-Poisson data can act as a surrogate for Landau damping, predicting electric field energy time series with sub-percent mean relative L2 errors in both linear and nonlinear regimes.","keywords":["Landau damping","Deep operator networks","operator learning","Vlasov-Poisson equations","surrogate modeling","electric field energy","nonlinear plasma dynamics","kinetic simulation"],"falsifier":"Take the trained single-mode model and evaluate it at a temperature outside the training range, say $T=0.3$ or $T=2.0$, or at a different perturbation amplitude with the same wavenumber; compare the prediction with a fresh Vlasov-Poisson run at that parameter. If the relative L2 error rises far above the reported test mean of 0.0083, the surrogate is interpolating its training domain rather than having learned the damping law itself.","tokens_in":9032,"feed_emoji":"⚡","tokens_out":8171,"duration_ms":81480,"temperature":0.7,"pith_summary":"Landau damping is a basic collisionless plasma process whose kinetic simulations are computationally expensive. This paper argues that a trained Deep Operator Network can serve as a surrogate: given a plasma temperature, it returns the time evolution of the electric field energy without solving the Vlasov-Poisson equations. The claim is demonstrated for a single damped mode and for a five-mode mixture whose large perturbations push the system into nonlinear Landau damping, with electron phase-space holes and oscillatory energy exchange. If the claim holds, temperature parameter studies that currently require a kinetic run per sample could be replaced by a single trained model that evaluates new cases in milliseconds at mean relative L2 errors near or below one percent.","feed_headline":"Neural operator predicts plasma damping with sub-1% error","feed_subtitle":"Trained on kinetic simulation data, it maps temperature to electric field energy in linear and nonlinear regimes.","key_machinery":"The central object is the DeepONet $G_\\omega$, an operator network with a branch sub-network that encodes the temperature input $T$ at sensor points and a trunk sub-network that encodes time coordinates $t_j$; the outputs combine to approximate $E(t)=\\int |E_x(x,t)|^2\\,dx$, the electric field energy. Training minimizes the mean squared error between the network output and reference time series, using Adam with an exponentially decaying learning rate, tanh activations, and fully connected hidden layers of width 200 and depth 6. The reference data are generated by a continuum Vlasov-Poisson solver with fixed perturbation modes (Table 1), and the trained network is evaluated on temperatures in $[0.5,1.5]$ not seen during training. The operator-learning construction is what allows a single model to predict entire energy time series for new temperature values without re-running a kinetic simulation.","core_discovery":"The central claim is that the operator mapping temperature to electric field energy is learnable by DeepONets with accuracy comparable to fully kinetic first-principles simulations. In the single-mode case, 200 training samples produce a mean relative L2 error of 0.0078 on training data and 0.0083 on unseen test data. In the five-mode case, scaling the training set from 50 to 800 samples lowers the test error from 0.0312 to 0.0043; with 400 training samples the test error is 0.0049. The five-mode setup includes nonlinear Landau damping, so the paper is establishing that a surrogate can follow not just the monotonic linear decay but the oscillatory, phase-space-hole-driven energy evolution that follows large initial perturbations.","pith_inferences":["Beyond the paper, a natural next test is to feed the branch network a parameter vector containing temperatures, wavenumbers, and amplitudes; the same architecture should handle it if the energy functional depends smoothly on those parameters, but that remains to be shown.","Beyond the paper, the reported errors are in-range interpolation errors; out-of-range temperatures or initial conditions outside the fixed mode set could break the accuracy, so the conservative reading is that the surrogate is a fast interpolator on the explored parameter domain.","Beyond the paper, the operator formulation is not tied to Landau damping, so the same training recipe could transfer to other kinetic plasma phenomena, such as ion-acoustic waves or driven instabilities, provided reference data that resolve those dynamics exist."],"forward_implications":["A trained model predicts electric field energy time series for previously unseen temperatures in $[0.5,1.5]$ with mean relative L2 errors below roughly one percent in both single- and five-mode tests.","Evaluating 100 five-mode test cases took 0.00148 seconds on one GPU, so a temperature sweep that would require many kinetic simulations can be performed nearly instantly once training is done.","The five-mode runs show that the surrogate preserves the nonlinear signature of Landau damping, including oscillatory energy exchange associated with electron phase-space holes, rather than only reproducing monotonic linear decay.","Test error in the five-mode case drops from 0.0312 with 50 training samples to 0.0043 with 800 samples, indicating that accuracy improves steadily as more reference data are supplied."],"supporting_citations":[{"why":"Introduced DeepONets, the operator-learning architecture whose branch and trunk networks are used here to map temperature to electric field energy.","marker":"L. Lu et al. 2021"},{"why":"Provides the continuum Vlasov-Poisson solver that generated the reference data used for training and testing.","marker":"J. Juno et al. 2018"},{"why":"States the universal approximation theorem that justifies representing the solution operator by a neural network.","marker":"K. Hornik et al. 1989"},{"why":"Earlier PINN-based model of linear Landau damping that this work extends into the five-mode nonlinear regime.","marker":"Y. Qin et al. 2023"},{"why":"Documents the electron phase-space holes that characterize nonlinear Landau damping, the physical structure the surrogate must capture in the five-mode case.","marker":"Z. Huang et al. 2025"}],"fun_headline_variants":["DeepONet models Landau damping with sub-1% error","Neural operator emulates plasma damping across regimes","Surrogate plasma damping via DeepONet, 0.8% test error"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the reference Vlasov-Poisson data faithfully represent Landau damping, and that with the fixed perturbation modes and amplitudes, the scalar temperature in $[0.5,1.5]$ determines the electric field energy time series smoothly enough for the network to interpolate it.","fun_headline_variants_meta":{"raw":{"variants":["DeepONet models Landau damping with sub-1% error","Neural operator emulates plasma damping across regimes","Surrogate plasma damping via DeepONet, 0.8% test error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000442,"raw_usage":{"total_tokens":2177,"prompt_tokens":818,"completion_tokens":1359,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":434,"completion_tokens_details":{"reasoning_tokens":1301}},"tokens_in":434,"tokens_out":1359,"duration_ms":11986,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:59:19.113427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained single-mode model and evaluate it at a temperature outside the training range, say $T=0.3$ or $T=2.0$, or at a different perturbation amplitude with the same wavenumber; compare the prediction with a fresh Vlasov-Poisson run at that parameter. If the relative L2 error rises far above the reported test mean of 0.0083, the surrogate is interpolating its training domain rather than having learned the damping law itself.","supporting_citations":[],"review_version":1}