{"id":"95ae9e8c-5480-4719-b6a9-27f3e6228698","arxiv_id":"2504.17375","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Deep reparameterization with a CNN and pretraining-based reference embedding gives the most accurate inversion, improves robustness to noise and sparse acquisition, and a backbone-branch variant reduces crosstalk in multiparameter elastic FWI.","lead":"A geophysics team tested whether representing the subsurface velocity model as the output of a neural network, instead of optimizing the model directly, improves seismic full waveform inversion. Their benchmark on synthetic models suggests a CNN with warm-up initialization is the most accurate and more robust to noise and sparse receivers than standard FWI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central comparison is made only from warm starts where the pretraining target is the Gaussian-smoothed true model; if the initial model contains large-scale errors, the claimed CNN-vp advantages and robustness may not transfer.","rationale":"The reader's weakest assumption correctly identifies the warm-start regime as the key external-validity gap. I considered two alternative concerns: hyperparameters are selected on the test set, and the multiparameter Marmousi2 results are qualitative rather than quantitative. Both are genuine reporting weaknesses, but they are less load-bearing than the warm-start issue because fixing them would not establish whether the method retains its advantage when the initial model is realistically imperfect. The warm-start assumption affects every major quantitative claim in the paper, including the robustness results under noise and sparse acquisition, since those experiments also start from the smoothed true model. The paper is internally coherent and the experiments are consistent with its claims within the tested regime, so the concern does not justify rejection; it does justify conditioning the general practical conclusions on having a reliable low-wavenumber initial model.","tokens_in":16587,"tokens_out":6017,"duration_ms":72790,"concrete_test":"Repeat the Marmousi2, Overthrust, and Foothill ablations using initial models with controlled low-wavenumber errors instead of the Gaussian-smoothed true model—for example, add a +5% background velocity ramp, shift the smoothing center laterally, or use a simplified travel-time tomography result—while keeping the pretraining of Eq. 11 otherwise identical. Compare CNN-vp against traditional FWI in MAPE, SSIM, and SNR. If the CNN-vp advantage shrinks or reverses as the initial-model error grows, the headline claim is warm-start-specific.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-A states that for each benchmark the initial velocity model is obtained by Gaussian smoothing the true model, and Eq. 11 pretrains the network to reproduce exactly that smoothed reference. All headline quantitative results—CNN-vp versus traditional FWI, the architecture comparison in Table II, and the noise/sparsity robustness curves in Section III-C—are therefore obtained in a regime where the starting model already contains the correct large-scale structure of the answer. In field applications the initial model comes from travel-time tomography or similar and will contain low-wavenumber errors, including incorrect velocity trends and misplaced slow/fast regions. Because pretraining hard-wires the reference into the network weights, a wrong reference becomes a wrong prior that the subsequent inversion may not overcome; the paper provides no experiment with any initial model that has realistic large-scale error. This is not an internal inconsistency, but it is the weakest load-bearing assumption behind the central claim that the method offers practical design guidance and reduced reliance on accurate initialization.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes DR-FWI, a deep reparameterization framework for full waveform inversion in which subsurface parameters (e.g., vp, vs, rho) are represented by the output of a neural network and the network weights are optimized against the seismic data misfit. It compares three network architectures (CNN, MLP, U-Net) and two reference-embedding strategies (pretraining-based vs. perturbation-based), and reports that a CNN with pretraining-based initialization (CNN-vp) outperforms traditional FWI and the other architecture/strategy combinations on the Marmousi2, Overthrust, and Foothill models. The paper also presents robustness experiments under noise and sparse acquisition, a spectral-bias analysis of the inversion dynamics, and a backbone-branch architecture for multiparameter joint inversion. The authors conclude that DR-FWI provides practical design guidance for network-based FWI, reduces acquisition requirements, and mitigates multiparameter crosstalk.","tokens_in":16746,"tokens_out":5700,"duration_ms":59604,"significance":"If the claims hold, this work provides a systematic and useful comparison of reparameterization architectures and embedding strategies, and it demonstrates a plausible extension to multiparameter FWI. The paper is thorough in its ablation scope, uses the open-source ADFWI toolbox, and reports a variety of quantitative metrics (MAPE, SSIM, SNR). The spectral-bias discussion is a valuable interpretive addition. However, the benchmark methodology has three important weaknesses: the best network configuration is selected on the test model itself, the traditional FWI baseline is unregularized, and the single-parameter experiments are performed only in a warm-start regime where the initial model is a smoothed version of the true model. These issues make the reported quantitative advantages optimistic and limit the strength of the practical design guidance that can be drawn from the current experiments.","major_comments":[{"comment":"The architecture comparison is compromised by selecting the best network configuration on the test model. The text states, 'we perform a network architecture search for each architecture and select the best-performing configuration for comparison,' and the same Marmousi2, Overthrust, and Foothill models are then used to report the final results in Table II and Figs. 4-5. This is selection on the test set: each reported number is the best among 8 to 16 random or hyperparameter variants, which inflates measured accuracy and makes the ranking between architectures an upper-bound comparison rather than an unbiased benchmark. Please use a validation model or hold-out split for configuration selection, or report the full distributions rather than only the best run.","section":"Section III-B"},{"comment":"The single-parameter experiments are performed only in a warm-start regime. For each model, the initial velocity model is obtained by Gaussian smoothing of the true model, and Eq. (11) pretrains the network to reproduce exactly that smoothed reference. Thus the pretraining target and the starting model contain the correct large-scale structure of the answer, and all reported improvements over traditional FWI are demonstrated only from such warm starts. The paper provides no experiment with an initial model containing low-wavenumber errors (e.g., incorrect velocity trends or misplaced structures), which is the situation encountered in field practice. To support the claimed practical guidance and real-world applicability, please add experiments with erroneous initial models (for example, from travel-time tomography or strongly perturbed smooth models) and show that the inversion can correct, or at least copes with, an incorrect reference.","section":"Section III-A and Eq. (11)"},{"comment":"The traditional FWI baseline uses no regularization, multiscale continuation, or similar safeguards, even though the paper attributes the advantage of DR-FWI largely to implicit regularization. This makes it unclear how much of the reported improvement comes from the network reparameterization itself and how much from the absence of standard regularization in the baseline. Please add a conventional regularized or multiscale baseline (e.g., total-variation regularization or frequency continuation) to the comparisons in Table II and the robustness curves in Figs. 6-7, or justify why the unregularized baseline is the appropriate reference.","section":"Section III-A and Table II"},{"comment":"All quantitative results appear to be single runs per configuration. The shaded regions in Figs. 5-7 show variability across network hyperparameters, not across random initializations or optimizer stochasticity. Since the pretraining and the subsequent inversion are stochastic, the reported differences (e.g., CNN-vp versus MLP-vp, or DR-FWI versus traditional FWI) may not be statistically stable. Please report the mean and standard deviation over multiple independent random seeds for the headline comparisons, and state whether the differences are significant.","section":"Table II and Figs. 6-7"}],"minor_comments":[{"comment":"The update rule mk+1 = mk + alpha_k * partial L / partial mk has the wrong sign for gradient descent; as written it performs ascent on the objective unless partial L / partial mk is defined as the negative gradient. Please correct the equation or clarify the sign convention.","section":"Section II-B, Eq. (5)"},{"comment":"The high-frequency ratio threshold rc = 2.375 is introduced without derivation or sensitivity analysis. Please define the frequency units and report how the qualitative conclusions of Fig. 10 depend on the choice of rc.","section":"Section IV-A, Eq. (17)"},{"comment":"The Introduction criticizes the absence of open-source implementations, but no code is released for the proposed DR-FWI framework itself; the paper only cites the ADFWI toolbox. Please provide a public repository with the network architectures and training/inversion scripts.","section":"Introduction and Limitations"},{"comment":"There are several minor grammatical errors and typos, such as 'supported by the the Shanghai' in the footnote and 'this results highlight' in Section III-D. A careful proofread is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central idea is sound and the multiparameter extension is interesting, but the benchmark methodology needs substantial strengthening before publication. In particular, the selection of network configurations on the test model and the lack of a regularized baseline are issues that the authors must address; the warm-start-only nature of the single-parameter experiments limits the practical conclusions. The novelty relative to He & Wang (2021) and Zhu et al. (2021) should be clarified explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis paper is a genuinely useful empirical benchmark of deep reparameterization for FWI. It systematically compares CNN, U-Net, and MLP parameterizations under two reference-embedding strategies across three synthetic models, and reports a consistent winner (CNN plus pretraining). The sparse-acquisition and noise-robustness study is new and practically relevant, and the backbone-branch multiparameter architecture is a sensible extension that shows real promise on the anomaly model. The writing is clear, and the use of the open-source ADFWI framework is a plus for reproducibility.\n\nThat said, the paper's practical claims outrun the evidence in a few specific ways. The most important is the warm-start setup: every test initializes from a Gaussian-smoothed version of the true model, and the pretraining step (Eq. 11) fits the network to exactly that smoothed reference. The comparison is therefore only in the regime where the large-scale structure is already known. The stress-test note is right. In a field setting where the initial model has misplaced low-wavenumber features, a wrong pretraining target becomes a wrong prior, and we have no experiment showing the method recovers from that. This doesn't undermine the internal comparison, but it does mean the 'practical design guidance' claim in the abstract goes beyond what is shown.\n\nSecond, the hyperparameter search selects the best network configuration on the test model itself (Section III-B). That likely inflates the reported accuracy numbers, and it weakens the claim that the chosen architecture is generically superior. There are also no repeated-run error bars, so we can't tell how much of the margin is noise. The traditional FWI baseline is unfavorably bare—no multiscale continuation or regularization—so the comparison is not apples-to-apples. On the multiparameter side, the Marmousi2 results are only shown as images, with no quantitative metrics, unlike the anomaly model. The spectral-bias analysis is a plausible post hoc interpretation rather than a derivation, and the threshold rc=2.375 is hand-picked, but it's clearly labeled as analysis and doesn't carry the main claims.\n\nNone of these are fatal. The central empirical finding—CNN-vp consistently beats plain FWI and the other variants—holds up within the warm-start regime, and the robustness trends look believable. The paper deserves a serious peer review, but it needs major revision: error bars, exact hyperparameters, a weaker-initial-model experiment, metrics for the multiparameter Marmousi2, and ideally code release.\n\nI'd bring it to reading group and would cite it, but I'd cite it as a benchmark study with caveats, not as proof that DR-FWI works from poor starting models.","headline":"A useful benchmark study of deep reparameterization for FWI, but warm-start initialization and test-set hyperparameter selection mean the practical claims outrun the evidence.","tokens_in":17324,"tokens_out":2703,"would_cite":true,"duration_ms":25543,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A CNN that represents the velocity model improves full waveform inversion accuracy and robustness.","keywords":["deep reparameterization","full waveform inversion","network architecture search","spectral bias","sparse acquisition","multiparameter inversion","crosstalk mitigation","seismic imaging"],"falsifier":"Run CNN-vp and conventional FWI on a synthetic model (e.g., a salt-inclusion model) where the initial model is obtained by smoothing a deliberately wrong large-scale structure, such as a shifted salt body, rather than the true model. If, under the paper's noise and sparsity settings, CNN-vp no longer beats conventional FWI in MAPE or SSIM—or if several CNN architectures with the pretraining warm-up collapse—the claim that reparameterization robustly improves FWI would be refuted. A second check targets the mechanism: on a different model, measure the early-iteration high-frequency energy ratio of conventional FWI; if it does not show the reported abrupt surge, the spectral-bias account needs revision.","tokens_in":16353,"feed_emoji":"🌊","tokens_out":5318,"duration_ms":46467,"temperature":0.7,"pith_summary":"The paper argues that representing the subsurface velocity model as the output of a deep neural network, rather than optimizing pixel values directly, makes full waveform inversion more accurate, stable, and tolerant of noisy or sparse data. It systematically benchmarks three network families—U-Net, CNN, and MLP—combined with two ways of injecting a starting model, and finds that a simple CNN with a pretraining-based warm-up works best across three benchmark models. The same reparameterization, in a backbone–branch form, is extended to joint inversion of P-wave velocity, S-wave velocity, and density, and is reported to suppress crosstalk artifacts. The paper also proposes a mechanism: the network's spectral bias imposes a progressive low-to-high frequency learning schedule that mimics multiscale FWI. If these claims hold, the practical consequences are design rules for network-based FWI and reduced data-acquisition costs.","feed_headline":"Shallow CNN prior sharpens seismic inversion","feed_subtitle":"Pretraining the network on a smoothed start model cuts error and survives noisy, sparse data.","key_machinery":"The load-bearing object is the deep reparameterization map $m = \\mathcal{N}(\\theta|I_r)$, where a generative network's weights are optimized by the FWI objective. The specific variant that carries the result is CNN-vp: a shallow CNN that takes a fixed random array, passes it through a fully connected layer and convolutional layers, and is warmed up by pretraining to output the Gaussian-smoothed initial model before inversion. In the multiparameter extension, the machinery is a backbone-branch CNN in which a shared backbone extracts common structure and separate branches, each followed by normalization to its physical scale, produce vp, vs, and rho. The paper attributes the method's behaviour to spectral bias—the network learns low-frequency components first, enforcing an implicit progressive multi-scale regularization.","core_discovery":"On its own terms, the central claim is that deep reparameterization—replacing the model $m$ by a generator network output $m = \\mathcal{N}(\\theta|I_r)$ and optimizing the network weights $\\theta$—materially improves FWI, and that the best configuration is a multi-layer CNN pretrained to reproduce the initial smoothed velocity model before physics-driven inversion. The paper reports that this CNN-vp configuration beats traditional FWI, U-Net, and MLP variants on the Marmousi2, Overthrust, and Foothill models in MAPE, SSIM, and SNR, holds up under Gaussian noise up to $6\\sigma_0$, and retains an advantage at extreme sparsity (e.g., 2 sources or 20 receivers). For multiparameter FWI, a shared-backbone plus branch CNN with per-parameter normalization and denormalization is claimed to nearly eliminate crosstalk on an anomaly model and substantially improve vp/vs/rho recovery on Marmousi2. The proposed explanation is spectral: conventional FWI injects high-frequency energy abruptly, while the network first dominates low frequencies and expands bandwidth gradually, reducing cycle-skipping.","pith_inferences":["If the spectral-bias mechanism is the real driver, then network architectures with stronger low-frequency bias should further improve robustness; this can be tested by comparing different activation functions, depths, or Fourier-feature embeddings within the same benchmark.","The pretraining warm-up effectively injects the entire initial model as a prior, so the method's edge over conventional FWI may shrink when the initial model is poor; a direct experiment would be to replace the Gaussian-smoothed true model with a biased or outdated velocity model and measure the performance gap.","The backbone-branch success suggests a route to multiphysics joint inversion (e.g., elastic plus electromagnetic or gravity), where parameter-specific branches could be governed by different forward operators while the backbone shares structure.","Because the network input is fixed and random, the method yields a deterministic map per initialization; ensembling over multiple random inputs could provide cheap uncertainty estimates for the inverted models."],"forward_implications":["A simple CNN with pretraining-based warm-up can be adopted as a plug-in reparameterization module in existing FWI workflows, with architecture guidance that shallower CNNs outperform U-Net and MLP for velocity reconstruction.","Pretraining-based initial-model embedding is consistently better than direct perturbation superposition across all tested architectures, so new DR-FWI designs should prefer the warm-up strategy.","DR-FWI retains accuracy at acquisition levels where conventional FWI degrades (e.g., 10 sources by 20 receivers), implying reduced field acquisition requirements for similar image quality.","In multiparameter elastic inversion, the backbone-branch structure lets vp, vs, and rho be inverted jointly with less crosstalk, without explicit Hessian-based corrections.","The spectral-bias explanation predicts that DR-FWI behaves like an adaptive multiscale method, which could reduce the need for manual frequency-continuation strategies."],"supporting_citations":[{"why":"Supplies the two-phase pretraining-based reparameterization scheme that the CNN-vp strategy builds on.","marker":"[21]"},{"why":"Supplies the perturbation-learning reparameterization variant used as the Delta-vp comparison.","marker":"[22]"},{"why":"Supplies the implicit FWI baseline using coordinate-mapping networks that motivates the method's design.","marker":"[23]"},{"why":"Supplies the automatic-differentiation FWI workflow that enables unified GPU computation in the framework.","marker":"[24]"},{"why":"Supplies the autoencoder-based multiparameter crosstalk framework that the backbone-branch design extends.","marker":"[29]"},{"why":"Supplies the forward modeling toolbox used for acoustic and elastic wave propagation in all experiments.","marker":"[34]"},{"why":"Supplies the global-correlation misfit function used to measure waveform discrepancy throughout the inversions.","marker":"[36]"},{"why":"Supplies the deep image prior concept that motivates representing the model as a network output.","marker":"[41]"},{"why":"Supplies the reference-guided DIP initialization strategy that the pretraining-based embedding refines.","marker":"[43]"}],"fun_headline_variants":["CNN pretraining keys to accurate, noise-robust FWI","Deep reparameterization: fewer shots, same resolution","Spectral bias guides CNN-FWI to avoid cycle-skipping","Backbone-branch CNN slashes cross-talk in multiparameter FWI","Benchmarking deep priors for full waveform inversion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmarks start every inversion from an initial model obtained by Gaussian-smoothing the true model, and the pretraining step fits the network to that smoothed reference, so the demonstrated superiority over conventional FWI is only shown in a warm-start regime where the large-scale structure of the answer is already known.","fun_headline_variants_meta":{"raw":{"variants":["CNN pretraining keys to accurate, noise-robust FWI","Deep reparameterization: fewer shots, same resolution","Spectral bias guides CNN-FWI to avoid cycle-skipping","Backbone-branch CNN slashes cross-talk in multiparameter FWI","Benchmarking deep priors for full waveform inversion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000419,"raw_usage":{"total_tokens":2226,"prompt_tokens":1084,"completion_tokens":1142,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":700,"completion_tokens_details":{"reasoning_tokens":1065}},"tokens_in":700,"tokens_out":1142,"duration_ms":11101,"temperature":1.0,"reasoning_tokens":1065,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:41:33.750708+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run CNN-vp and conventional FWI on a synthetic model (e.g., a salt-inclusion model) where the initial model is obtained by smoothing a deliberately wrong large-scale structure, such as a shifted salt body, rather than the true model. If, under the paper's noise and sparsity settings, CNN-vp no longer beats conventional FWI in MAPE or SSIM—or if several CNN architectures with the pretraining warm-up collapse—the claim that reparameterization robustly improves FWI would be refuted. A second check targets the mechanism: on a different model, measure the early-iteration high-frequency energy ratio of conventional FWI; if it does not show the reported abrupt surge, the spectral-bias account needs revision.","supporting_citations":[{"cited_title":"Integrating deep neural networks with full-waveform inversion: Reparameterization, regularization, and uncertainty quantification,","cited_arxiv_id":null,"evidence_quote":"Supplies the perturbation-learning reparameterization variant used as the Delta-vp comparison."},{"cited_title":"Implicit seismic full wave- form inversion with deep neural representation,","cited_arxiv_id":null,"evidence_quote":"Supplies the implicit FWI baseline using coordinate-mapping networks that motivates the method's design."},{"cited_title":"Automatic differentiation-based full waveform inversion with flexible workflows,","cited_arxiv_id":null,"evidence_quote":"Supplies the automatic-differentiation FWI workflow that enables unified GPU computation in the framework."},{"cited_title":"Elastic-adjointnet: A physics-guided deep au- toencoder to overcome crosstalk effects in multiparameter full-waveform inversion,","cited_arxiv_id":null,"evidence_quote":"Supplies the autoencoder-based multiparameter crosstalk framework that the backbone-branch design extends."},{"cited_title":"Liufeng2317/adfwi: Zenodo,","cited_arxiv_id":null,"evidence_quote":"Supplies the forward modeling toolbox used for acoustic and elastic wave propagation in all experiments."},{"cited_title":"Application of multi-source waveform inversion to marine streamer data using the global correlation norm,","cited_arxiv_id":null,"evidence_quote":"Supplies the global-correlation misfit function used to measure waveform discrepancy throughout the inversions."},{"cited_title":"Deep image prior,","cited_arxiv_id":null,"evidence_quote":"Supplies the deep image prior concept that motivates representing the model as a network output."},{"cited_title":"Reference-driven compressed sensing mr image reconstruction using deep convolutional neural networks without pre-training,","cited_arxiv_id":null,"evidence_quote":"Supplies the reference-guided DIP initialization strategy that the pretraining-based embedding refines."}],"review_version":1}