{"id":"f7b2b0e5-aab3-44fa-a107-7750a17ee1ab","arxiv_id":"2607.28535","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Geometry-invariant physics encoding plus a conformal residual ensemble and held-out-shot audit cuts FWI error ~38% on synthetics and restores ~0.9 coverage zero-shot on Marmousi-2 under shift.","lead":"A laptop-scale seismic inversion pipeline maps variable surveys into fixed physics-derived model images, then applies a calibrated neural residual corrector. It cuts velocity error versus classical FWI and keeps uncertainty intervals honest under geometry change, noise, and wavelet error on synthetics and zero-shot Marmousi-2.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Audit’s C(τ) peak is only a heuristic proxy for model-space inflation; Table 2 already shows large τ errors and coverage outside the abstract’s 0.89–0.91 band when noise or aperture violate the stated regime.","rationale":"The reader correctly isolated the audit identifiability assumption (Eqs. 12–15, §3.4/§5.6) as the soft spot under the strongest claim. I sharpen it with the concrete Table 2 discrepancies the abstract smooths over and with the missing link between data-space peak under assumed physics and model-space marginal coverage, but I do not find a deeper internal contradiction: GIPE transfer, residual gains vs classical prior, and budget-matched baseline gaps are adequately supported inside the 2-D acoustic synthetic+Marmousi regime. Code-not-yet-shipped and purely synthetic primary evidence remain external cautions already priced into CONDITIONAL. No verdict move is warranted; the same conditions the authors document (long paths, model error above noise floor, fixed audit config for peak-height comparisons) remain the acceptance gates.","tokens_in":18306,"tokens_out":694,"duration_ms":46752,"concrete_test":"On the full-line Marmousi reference, re-run the audit (defaults M=12, δ=0.005) three ways: (a) held-out set = only the four nearest-offset shots, (b) only the four farthest-offset shots, (c) replace 60 m correlation in ξ_i by white noise and by 200 m. Record τ_audit−τ_oracle and Cov.aud. If any case yields |τ error|>0.5 or Cov.aud outside [0.85,0.95], the C(τ)→model-space mapping is materially fragile beyond the paper’s favorable spine.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The deployment half of the strongest claim—that the label-free held-out-shot audit restores near-nominal pixel-wise coverage under shift—rests on Eqs. 12–14: τ_audit is chosen where data-space coverage C(τ) of shots simulated under the *assumed* forward operator peaks. That construction is only guaranteed to track the oracle model-space inflation if (i) the errors that need widening are visible in those held-out residuals, (ii) C(τ) has an identifiable interior peak, and (iii) the 60 m correlated perturbations span the true error structure. The paper’s own full-line spine already violates the tidy abstract range: at SNR 2, τ_audit=8.0 vs oracle 4.24 and Cov.aud=0.971; at SNR 32, Cov.aud=0.859 (Table 2). On 4 km crops, C(τ) plateaus and the plateau-right rule runs to the grid edge (τ=12), so the procedure only over-covers (§5.6, Table S2, Fig. 12). Eq. 15 is a post-hoc length criterion, not a pre-deployment certificate. Thus the audit is an empirically useful recalibrator inside the documented long-path, model-error-dominated regime, not a general ground-truth-free coverage guarantee.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes a laptop-scale residual FWI pipeline that never feeds shot gathers to the network. Variable acquisitions are mapped into a fixed ten-channel model-space tensor (starting model, ADMM prior, two misfit gradients, six fast-marching illumination/wavenumber maps); a six-member heterogeneous ensemble then predicts a calibrated residual on that prior. Physics-conditioned (Mondrian) conformal prediction supplies finite-sample pixel-wise marginal coverage on the calibration distribution, and a label-free held-out-shot physics audit rescales interval width under domain shift by maximizing data-space coverage C(τ) of unused shots. On a 1000-model corpus spanning six acquisition families the ensemble cuts RMSE 38% relative to its classical prior and shows no geometry-specific gap on three never-trained families. Zero-shot on full-line Marmousi-2 it improves 354→304 m/s; raw coverage collapses to ~0.42 and the audit repairs it under noise, wavelet error, shot decimation, and an elastic–acoustic mismatch stress test, with peak height flagging physics mismatch. Budget-matched gather-based baselines underperform off their training geometry or everywhere.","tokens_in":18668,"tokens_out":1710,"duration_ms":38825,"significance":"Acquisition fragility and uncalibrated uncertainty remain the main barriers between DL-FWI demos and tools practitioners would trust. The geometry-invariant physics encoding is a clean architectural answer to the first problem; the combination of conformal calibration on synthetics with a wave-equation audit for label-free recalibration is a concrete, falsifiable answer to the second. Strengths that raise the bar for the field include: a seeded 1000-instance corpus with held-out acquisition families, budget-matched gather baselines (fixed and source-conditioned), two ablations tied to specific claims, eleven full-line corruption conditions plus elastic mismatch, explicit audit failure modes on short-offset crops, and a complete experimental matrix runnable in ~90 h on one laptop with planned Zenodo release. If the audit regime is stated carefully, this is a useful, reproducible contribution to calibrated residual FWI rather than another uncalibrated end-to-end network.","major_comments":[{"comment":"Abstract and §7 claim the audit “restores coverage to 0.89–0.91 across eleven corruption conditions.” Table 2 contradicts that band: SNR 2 yields Cov.aud = 0.971 (τ_audit = 8.0 vs oracle 4.24); SNR 32 yields 0.859; the 32-shot row has no audit by construction; elastic mismatch (§5.4) repairs only to 0.83. The body (§5.5–5.6) already documents these cases and the short-offset plateau failure. The abstract/conclusion range should be revised to match Table 2 (e.g., “typically 0.86–0.91, with documented overshoot when the noise floor dominates and under-repair under physics mismatch”) so the central deployment claim is not overstated.","section":"Abstract; Table 2; §5.5; §7"},{"comment":"The deployment half of the strongest claim rests on Eqs. (12)–(14): maximizing held-out data-space coverage C(τ) under the assumed forward operator is taken as a proxy for the model-space inflation that restores marginal coverage. §5.6 and Eq. (15) correctly show that C(τ) is identifiable only when paths are long enough and model error dominates the predictive band; on 4 km crops the curve plateaus and the plateau-right rule runs to the grid edge (Table S2, Fig. 12), producing only conservative over-coverage. This is an empirically useful recalibrator inside a documented regime, not a general ground-truth-free coverage guarantee. The abstract and conclusions should state the aperture/noise preconditions with the same clarity as the Discussion, and preferably give a checkable pre-deployment diagnostic (e.g., pre-arrival noise vs. predicted data residual, or a minimum path-length criterion","section":"§3.4 Eqs. (12)–(14); §5.6; Eq. (15); Table S2"},{"comment":"§5.3 reports that both budget-matched gather baselines (InversionNet-style fixed geometry and a 35M source-conditioned FiLM/Fourier-DeepONet-style net) land at or worse than the classical prior. The paper appropriately limits the claim to “at the matched 600-instance budget.” To keep the comparison load-bearing rather than a straw man, the main text should state the training protocol more explicitly (same augmentation, output grid, early-stopping, and that neither baseline received the GIPE channels or the ADMM prior as input) and note whether a residual-on-prior or multi-scale gather baseline was considered. Without that, readers may discount the geometry-invariance claim as an artifact of weak baselines rather than of representation.","section":"§5.3; Figs. 9–10; Table S5"}],"minor_comments":[{"comment":"Fig. 4 and §5.1 report Spearman 0.76 and AUSE 0.118 for σ ranking error; add the corresponding curves or a one-sentence definition of the sparsification protocol so the number is interpretable without the supplement.","section":"§5.1; Fig. 4"},{"comment":"Eq. (8) sets λ = 0.1 in normalized units with no sensitivity. A one-row ablation or brief note on whether λ was tuned on validation would help, given the Discussion’s own remark that the residual scale caps deep-section gains.","section":"§3.3 Eq. (8); §6"},{"comment":"Mondrian strata are “illumination quartiles and depth halves” (§3.3). State how many calibration pixels fall in the worst stratum and whether empty/near-empty strata are pooled; finite-sample quantile validity is sensitive to stratum size.","section":"§3.3"},{"comment":"Self-citation to Kumar & Tripathi (2026) is listed as “Under review.” Clarify what is inherited (ADMM prior, reweighted-ℓ1) versus new (GIPE, ensemble, conformal+audit, pre-stack velocity) so novelty is unambiguous.","section":"§1; References"},{"comment":"Table 1 and §4.3: “PyTorch 2.13” is likely a typo (current public releases are 2.x with different minor numbering); correct for reproducibility.","section":"§4.3; Table 1"},{"comment":"Fig. 3 caption and Eq. (6): channel order and units (normalized vs. physical) are not fully specified; a short table or colorbar units would help re-implementation.","section":"Fig. 3; Eq. (6)"},{"comment":"Minor prose: “Simut˙ e” encoding in the Fichtner references; “Deepak Kumara” vs. “Kumar” in the author line; ensure consistent author spelling before production.","section":"Title block; References"}],"recommendation":"minor_revision","confidential_remarks":"The experimental matrix and honesty about audit failure modes are above average for this subfield; I would not reject over the abstract’s coverage band if the authors tighten the wording. Fit for a methods-oriented geophysics or computational-imaging venue is good. The 2026 self-citation and arXiv dateline are slightly unusual but not disqualifying if the impedance paper is clearly scoped as upstream."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful laptop-scale residual FWI pipeline that actually moves the needle on two practitioner pain points—acquisition fragility and out-of-distribution uncertainty—without pretending to replace classical inversion.\n\nWhat is new is the combination, not the parts. They never feed gathers. Variable surveys collapse into a fixed ten-channel model-space tensor (start model, ADMM prior, two gradients, six fast-marching coverage maps). A small heterogeneous residual ensemble gets Mondrian conformal calibration on illumination/depth strata, then a held-out-shot wave-equation audit rescales interval width from data-space coverage C(τ) with no velocity truth. On 1000 synthetics spanning six acquisition families (three unseen), they cut prior RMSE ~38% with no geometry gap. Zero-shot full Marmousi-2 goes 354→304 m/s; raw coverage collapses to ~0.42 and the audit mostly repairs it. Budget-matched gather baselines, fixed and source-conditioned, lose badly. Ablations are honest: coverage channels are not load-bearing for mean accuracy, and fixed-acquisition training still transfers—so the encoding, not the aug, carries the claim. They also document when the audit fails (short crops plateau; SNR 2 overshoots).\n\nSoft spots in proportion. Absolute gains are modest and residual-capped where the prior is blank. Primary evidence is 2D acoustic synthetic plus one public model; code is promised, not shipped. The stress-test note is right that C(τ) is not a certificate: Table 2 already shows τ_audit=8 vs oracle 4.24 and Cov 0.971 at SNR 2, and 0.859 at SNR 32—outside the abstract’s tidy 0.89–0.91 band. Peak height as mismatch flag is useful inside their long-path, model-error-dominated regime, not general. Math and citations look standard and fair; circularity is low.\n\nWho it is for: people who ship FWI or scientific-ML inverse methods and care about geometry transfer and deployable intervals. Worth a serious referee. I would engage, cite the GIPE+audit pattern if I work this area, and bring it to reading group for the evaluation discipline as much as the architecture.","headline":"Solid practice-facing FWI+ML stack: geometry transfer via physics encoding is real; the label-free audit helps on long lines but is a regime-limited heuristic, not a coverage guarantee.","tokens_in":19325,"tokens_out":576,"would_cite":true,"duration_ms":18778,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A physics-encoded residual network transfers across unseen seismic acquisitions and a label-free audit restores trustworthy uncertainty under real distribution shift.","keywords":["full-waveform inversion","deep learning","uncertainty quantification","conformal prediction","acquisition geometry","Marmousi","physics encoding","residual correction"],"falsifier":"On a real marine line with well logs (for example Viking Graben), check whether the audited intervals achieve near-nominal coverage at the boreholes after zero-shot application, and whether the audit’s peak height still drops under known wavelet or elastic mismatch while staying high under pure noise or shot decimation.","tokens_in":19130,"feed_emoji":"🌍","tokens_out":1026,"duration_ms":21078,"temperature":0.7,"pith_summary":"Networks that turn raw seismic shot gathers into velocity models usually memorize the survey layout they saw in training, and their uncertainty estimates fail when the data leave that distribution. This paper argues that both failures can be fixed without replacing classical full-waveform inversion. Variable acquisitions are first mapped by classical physics operators into a fixed ten-channel model-space tensor—starting model, regularized inversion, misfit gradients, and illumination maps—so the network never sees gathers and never encodes geometry in its weights. A small heterogeneous ensemble then learns only a residual correction to that physics prior and is calibrated by physics-conditioned conformal prediction. When the target leaves the calibration distribution, a held-out-shot physics audit rescales the intervals by simulating unused shots through the predictive samples; no ground truth is required. On a thousand synthetic models the method cuts error 38 percent relative to its classical prior and shows no measurable gap on never-seen geometries; zero-shot on the full 17 km Marmousi-2 line it lowers error from 354 to 304 m/s and the audit restores near-nominal coverage across noise, wavelet error, and shot decimation while its peak height flags physics mismatch.","feed_headline":"Seismic AI transfers across unseen surveys; audit fixes its uncertainty","feed_subtitle":"Physics encoding plus held-out-shot recalibration restores coverage on full Marmousi-2 without ground truth","key_machinery":"Geometry-invariant physics encoding (GIPE): a fixed ten-channel model-space tensor built only from classical operators (starting model, ADMM-regularized FWI, two misfit gradients, six fast-marching illumination/wavenumber maps). The network is only a calibrated residual corrector on this prior; a held-out-shot physics audit then rescales interval width where simulated data-space coverage peaks, without ground truth.","core_discovery":"A geometry-invariant physics encoding plus a residual ensemble with physics-conditioned conformal calibration transfers across unseen acquisition families without measurable degradation, and a label-free held-out-shot physics audit restores near-nominal pixel-wise coverage (about 0.89–0.91) under domain shift on full-line Marmousi-2 while the peak of data-space coverage cleanly separates physics mismatch from benign corruptions.","pith_inferences":["Any inverse problem that already owns a reliable forward operator and can reserve a few observations could adopt the same held-out physics audit instead of relying on network uncertainty alone.","Iterating the encode–correct cycle—feeding the corrected model back as a new prior—could lift the residual ceiling in deep, poorly illuminated zones where a single pass saturates.","The short-offset plateau of the audit curve supplies a practical pre-check: if paths are shorter than roughly a few kilometres at exploration frequencies, expect only conservative over-coverage rather than tight recalibration."],"forward_implications":["Acquisition geometry no longer has to be fixed or explicitly conditioned for a learned FWI corrector to transfer.","Uncertainty statements from synthetic conformal calibration can be transported to field data without labels by using the wave equation on held-out shots.","Physics mismatch (wrong wavelet, acoustic-vs-elastic) becomes detectable from the depressed ceiling of data-space coverage rather than from ground-truth error.","Budget-matched gather-to-model networks, even with source conditioning, are sample-inefficient relative to the physics-encoded residual route at laptop scale.","The same encode–correct–audit pattern is dimension-agnostic and can be tried in 3-D once classical-chain cost is managed."],"fun_headline_variants":["Physics encoding transfers residual FWI across unseen surveys","Ensemble cuts error 38% then audits coverage without labels","Zero-shot Marmousi-2: audit restores 0.89–0.91 coverage","Geometry-invariant priors beat gather-trained baselines off-survey","Held-out-shot physics audit separates mismatch from noise"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The audit works only when held-out-shot data-space coverage has a clear peak that correctly marks how much to widen the model intervals—something that requires long enough propagation paths and model error that sits above the noise floor.","fun_headline_variants_meta":{"raw":{"variants":["Physics encoding transfers residual FWI across unseen surveys","Ensemble cuts error 38% then audits coverage without labels","Zero-shot Marmousi-2: audit restores 0.89–0.91 coverage","Geometry-invariant priors beat gather-trained baselines off-survey","Held-out-shot physics audit separates mismatch from noise"]},"model":"grok-4.5","effort":"low","cost_usd":0.005196,"raw_usage":{"total_tokens":1511,"prompt_tokens":913,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":51964000,"prompt_tokens_details":{"text_tokens":913,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":526,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":913,"tokens_out":72,"duration_ms":8986,"temperature":1.0,"reasoning_tokens":526,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T04:21:07.151513+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a real marine line with well logs (for example Viking Graben), check whether the audited intervals achieve near-nominal coverage at the boreholes after zero-shot application, and whether the audit’s peak height still drops under known wavelet or elastic mismatch while staying high under pure noise or shot decimation.","supporting_citations":[],"review_version":1}