{"id":"25b4755a-f6bd-4ee5-8a9d-7f59a5b8f062","arxiv_id":"2502.07171","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Augmenting the training data of a CO2 storage digital shadow with ten Brie rock physics exponents improves plume reconstruction when the true rock physics is unknown.","lead":"This paper tests whether training a machine-learning 'digital shadow' for CO2 storage monitoring with many different rock physics models makes it more reliable when the true rock physics is unknown. In a synthetic North Sea reservoir test, the augmented model tracks the CO2 plume better than the non-augmented version, though the evidence is qualitative and single-case.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'out-of-distribution' test in Fig. 1b/c draws the rock-physics exponent from the same U(1,10) distribution used to train the augmented model, so the demonstrated gain is in-support averaging rather than robustness to model misspecification; no quantitative metrics support the qualitative claim.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing concern: the augmented training distribution and the test distribution coincide, so the reported 'out-of-distribution' improvement is actually an in-distribution interpolation result. This is not an internal inconsistency, because the paper's own conclusion carefully limits the claim to the Brie family with unspecified exponent, but the experimental design conflates an unknown parameter value with an unknown physical model. The absence of quantitative error and uncertainty metrics compounds the problem by making the qualitative improvement impossible to assess independently. A targeted held-out evaluation with exponents outside [1,10] and, ideally, a non-Brie forward model would settle whether the method generalizes beyond the training support. Since the reader's CONDITIONAL verdict already requires quantitative evaluation and a true OOD test, my stress test does not change that verdict.","tokens_in":4490,"tokens_out":4877,"duration_ms":46949,"concrete_test":"Generate held-out test observations with e values outside the training support (e.g., e=0.5 and e=15) and, if feasible, with a non-Brie mixing model; also include held-out e values inside [1,10] as a control. For each test set, compute mean L2 saturation error and a calibration score (e.g., CRPS or coverage) for the augmented and non-augmented CNFs. If the augmented model's advantage shrinks or reverses for e outside [1,10], the central claim reduces to in-support averaging, not robustness to misspecified rock physics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise of the robustness claim is that the unknown true saturation-mixing behavior lies inside the family used for augmentation. Section 3.2 defines training observations with Rk~p(R), e~U(1,10), and Section 4 labels Figure 1b/c as an 'unknown rock physics model.' But the test observation is generated from the same Brie family and the same exponent range as the training data. For the augmented model, that is not an out-of-distribution test; it is an in-distribution sample from the marginal p(y)=∫p(y|e)p(e)de. The CNF is therefore performing Bayesian model averaging over e, not robust inference under a misspecified forward model. The paper's own Conclusion restricts the claim to 'within the family of Brie saturation models where the exponent e is not specified,' so the result is an interpolation statement; it offers no evidence for robustness to an e outside [1,10] or to a different mixing model (e.g., Voigt/Reuss bounds). A second, compounding gap is that no quantitative errors or uncertainty metrics are reported for Figure 1, so 'closer' and 'reduced uncertainty' are only qualitative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes making a Digital Shadow data-assimilation framework for CO2 storage monitoring more robust to rock-physics model uncertainty by augmenting the training set of a conditional normalizing flow with multiple Brie saturation models, with exponent e drawn from U(1,10). The authors train the augmented and non-augmented CNFs on synthetic forecast/observation pairs from a 2D Compass-based model, then compare their posterior plume estimates on a synthetic test in which the seismic observation is generated with an 'unknown' rock physics model. They report qualitatively that augmentation brings the conditional mean closer to ground truth and reduces error and uncertainty.","tokens_in":4686,"tokens_out":3092,"duration_ms":30646,"significance":"If the effect is real, ensemble augmentation over rock-physics models is a simple, practical way to hedge against misspecified saturation mixing models in learned seismic monitoring, and the paper builds on a reproducible pipeline of open-source tools (JutulDarcy.jl, JUDI.jl, InvertibleNetworks.jl). The authors are explicit that the range of protection is the Brie family with unspecified exponent. However, the current evidence is only qualitative, and the 'out-of-distribution' test is not outside the training support of the augmented model, so the central robustness claim is not yet established.","major_comments":[{"comment":"The test labeled 'unknown rock physics model' in Figure 1b/c is generated from the same Brie family and the same U(1,10) exponent range used to augment the training data in Section 3.2. For the augmented model, this is an in-distribution sample from the marginal p(y)=∫p(y|e)p(e)de, so Figure 1c demonstrates Bayesian model averaging over the training prior on e, not robustness to an out-of-distribution or misspecified rock physics model. To support the OOD claim, the authors should either remove the 'out-of-distribution' label or add truly OOD tests, such as e values outside [1,10] or a different mixing model (e.g., Voigt/Reuss bounds or patchy versus uniform saturation).","section":"Section 3.2 and Section 4, Figure 1"},{"comment":"No quantitative error, uncertainty, or statistical coverage metrics are reported; 'closer to the GT' and 'reduced uncertainty' are supported only by visual inspection of three panels. The authors should report numerical measures such as RMSE, mean absolute error, energy score, or interval coverage, and repeat the experiment over multiple flow realizations, noise draws, and test exponents to show that the improvement is consistent rather than a single cherry-picked realization.","section":"Section 4, Figure 1"},{"comment":"The paper states that the Brie exponent is uniformly sampled from e~U(1,10), but then says 10 different exponents are used to augment the data tenfold. This is a discrete set, not a continuous uniform sample, and the number of distinct rock-physics variants is an important design choice. A sensitivity analysis varying the number of exponent values (e.g., 3, 10, 30) would clarify whether the observed robustness gain is stable or depends heavily on this choice.","section":"Section 3.2"}],"minor_comments":[{"comment":"The notation is inconsistent: the text defines the network weights as phi and the objective uses f_phi, while the surrounding sentence refers to the Jacobian of f_theta; please unify the symbol.","section":"Equation (3)"},{"comment":"The caption says 'at k=1' without explaining whether results at other timesteps are similar; please state whether this is representative or add the corresponding panels for other timesteps.","section":"Figure 1"},{"comment":"The injection rate is written as '0.0500 m3/s' and other quantities use plain text; use proper superscript formatting for units throughout.","section":"Section 3.1"},{"comment":"The reference 'Panel on Climate Change) 2018' has a formatting error and should be corrected to the standard IPCC citation.","section":"Introduction/References"},{"comment":"The conclusion mentions augmentation of 'fluid-flow properties', but the augmentation described in Section 3.2 only varies the rock physics model, not the flow simulation; please align the wording with the actual procedure.","section":"Section 5, Conclusions"}],"recommendation":"major_revision","confidential_remarks":"The core idea is worthwhile and the use of established open-source simulation and inversion tools is a strength, but the current evaluation is too weak: the OOD experiment is in-distribution for the augmented model and no quantitative metrics are provided. I encourage the authors to add an actual OOD test and numerical benchmarks; with those additions the paper could become suitable for publication. The reference list is adequate, though the paper would benefit from citing work on Bayesian model averaging and on robust inverse problems under model misspecification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is narrow but real: the authors show, on one synthetic case, that training their CNF-based Digital Shadow on seismic images generated with several Brie saturation exponents (e~U(1,10)) produces better looking posterior means and lower qualitative uncertainty than a shadow trained on a single exponent, when the test image comes from a different exponent in that range. That is a sensible extension of their earlier Digital Shadow work, and the augmentation cost is cheap because the expensive flow simulations are reused. The paper is honest in its conclusion that the robustness is only demonstrated within the Brie family, and it uses open-source tools throughout, which helps reproducibility. The soft spots are exactly where the stress-test lands. The Figure 1b/c setup is labeled 'unknown rock physics model,' but for the augmented model the test exponent is drawn from the same U(1,10) range used in training. That is in-support Bayesian model averaging over e, not out-of-distribution generalization to a misspecified forward model. Calling it 'out-of-distribution' oversells the result. The second problem is the evidence itself: no quantitative error or uncertainty metrics, no error bars, no repeated experiments, and the non-augmented baseline is not described beyond 'the correct rock physics model.' So the central claim is plausible but under-supported. The authors should report numbers (e.g., NRMSE or log-likelihood for the posterior mean/samples), show what happens for e outside [1,10] or for a different mixing model like Voigt/Reuss bounds, and clarify that the augmented model is marginalizing over the training distribution of e, not magically robust to anything outside it. This is a short preprint that reads like an extended abstract. The direction matters for GCS monitoring, and the result, if confirmed with quantitative evidence, would be a useful incremental contribution. As it stands, the paper is not publishable as-is, but it deserves a serious referee rather than a desk reject, because the idea is sensible and the missing evidence is obtainable with modest additional work. Tell the authors to add metrics and a true OOD test; then it becomes a solid workshop or short-conference paper.","headline":"A plausible but under-evidenced augmentation trick for Digital Shadow CO2 monitoring; the 'out-of-distribution' test is actually in-distribution for the augmented model.","tokens_in":596,"tokens_out":709,"would_cite":false,"duration_ms":19855,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When a CO2 plume monitor is trained on ten rock physics models instead of one, its forecasts stay accurate even when the true saturation-mixing law is unknown, within the trained family.","keywords":["CO2 storage monitoring","Digital Shadow","rock physics uncertainty","Brie saturation model","conditional normalizing flows","time-lapse seismic","ensemble augmentation","Bayesian data assimilation"],"falsifier":"Run the same augmented digital shadow on a test seismic image generated with a rock physics law that is deliberately outside the Brie family, for example a different effective-medium mixing model, or with a Brie exponent above 10. If the augmented shadow's posterior mean error and uncertainty are no better than the non-augmented baseline, the paper's claim of robustness to unknown rock physics models is refuted.","tokens_in":4242,"feed_emoji":"🌍","tokens_out":6880,"duration_ms":55307,"temperature":0.7,"pith_summary":"This paper argues that a CO2 plume monitoring system built on machine-learning data assimilation can be made insensitive to a key modeling guess: the rock physics law that converts reservoir fluid saturation into measurable seismic properties. The authors train a normalizing-flow digital shadow on a forecast ensemble in which each simulated plume is paired with seismic images generated under ten different Brie saturation models, with the mixing exponent drawn uniformly from 1 to 10. They report that when the true rock physics model used to create a test seismic image is unknown, this augmented training keeps the posterior mean close to the true plume and lowers both error and uncertainty, whereas a shadow trained on one rock physics model degrades. The practical stake is that geological CO2 storage sites have poorly constrained rock physics, and safe monitoring needs forecasts that do not silently fail when the assumed model is wrong.","feed_headline":"Adding 10 rock-physics models keeps CO2 plume forecasts accurate","feed_subtitle":"When the true mixing law is unknown, the augmented digital shadow cuts both error and uncertainty.","key_machinery":"The load-bearing object is the conditional normalizing flow as a neural posterior density estimator, trained on simulation pairs of plume state and seismic observation. The augmentation mechanism is the Brie saturation model, a rock physics mixing law whose exponent controls how CO2 saturation maps to the effective bulk modulus of the pore fluid, interpolating between uniform and patchy fluid distributions; drawing the exponent uniformly from 1 to 10 and generating one seismic image per exponent multiplies the forecast ensemble tenfold. This teaches the network a family of observation operators instead of a single assumed model, so the learned posterior marginalizes over rock physics uncertainty. The seismic observations are produced by simulating wave propagation and imaging with colored Gaussian noise added before migration.","core_discovery":"The central claim is that data augmentation over rock physics models makes the Digital Shadow robust to misspecification of the saturation-mixing law. In the paper's formulation, the observation operator draws a rock physics model from the Brie family with exponent uniformly sampled from [1,10], and each CO2 flow simulation is converted into seismic images for ten such exponents, producing 1280 training pairs from 128 flow simulations. A conditional normalizing flow is then trained to map seismic observations to posterior plume states. At inference, when the seismic observation is generated with an unknown (but within-family) exponent, the augmented shadow yields a conditional mean closer to the ground-truth plume and reduced uncertainty compared with the non-augmented shadow trained on a single rock physics model. The authors frame this as mitigating the negative effects of incorrect rock physics assumptions and improving the reliability of CO2 storage monitoring.","pith_inferences":["Because the training distribution explicitly randomizes the Brie exponent, the learned posterior is effectively a mixture over rock physics models; this makes the observed uncertainty reduction a form of Bayesian model averaging, an interpretation the paper does not spell out.","The paper's out-of-distribution test is still in-distribution with respect to the augmented training family, so the result is interpolation; a stronger test would use a mixing law absent from the Brie family.","The same augmentation idea should carry over to other uncertain components of the observation operator, such as wavelets, noise statistics, or anisotropy parameters; the paper only randomizes the saturation exponent.","One could extend the flow to output a posterior over the rock physics exponent itself, turning an assumed nuisance parameter into a monitored quantity that flags when the true model leaves the trained family."],"forward_implications":["An augmented digital shadow remains accurate when the true rock physics exponent is unknown but within the trained range [1,10], where the non-augmented shadow's mean deviates from the true plume.","Posterior uncertainty is lower in the unknown-exponent case, giving operators narrower and more truthful confidence bands.","The augmentation multiplies the effective training set tenfold from the same 128 flow simulations, so the additional cost is seismic simulations rather than new flow simulations.","The approach inherits the amortized character of the shadow: after training, inference on a new seismic survey is fast, which matters for repeated monitoring."],"supporting_citations":[{"why":"Defines the uncertainty-aware Digital Shadow and the conditional normalizing flow training objective that this paper modifies by augmenting the ensemble.","marker":"Gahlot, Orozco, et al. 2024"},{"why":"Supplies the Brie saturation model family and the uniform-versus-patchy terminology used for the augmentation.","marker":"Avseth, Mukerji, and Mavko 2010"},{"why":"Provides the open-source multi-phase flow simulator that produces the 128 ensemble CO2 plume forecasts.","marker":"Møyner, Bruer, and Yin 2023"},{"why":"Provides the symbolic seismic modeling framework used to generate shot records and images.","marker":"Witte et al. 2019"},{"why":"Provides the seismic imaging implementation used to turn acoustic property changes into observations.","marker":"Louboutin et al. 2023"},{"why":"Provides the normalizing flow library used to train the conditional posterior density estimator.","marker":"Orozco et al. 2024"},{"why":"Provides the probabilistic full-waveform inversion that initializes the permeability distribution for the plume ensemble.","marker":"Yin et al. 2024"},{"why":"Provides the synthetic Compass Earth model from which the 2D test subsurface is derived.","marker":"E. Jones et al. 2012"}],"fun_headline_variants":["Augmented rock physics shrinks CO2 plume forecast errors","Diverse rock models boost CO2 monitoring accuracy","10 rock-physics variants cut CO2 plume uncertainty","Robust CO2 tracking with multiple rock physics models","Mixing rock physics improves CO2 storage forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole robustness result rests on the true rock physics behavior being representable by a Brie saturation model with exponent between 1 and 10; if the real mixing law falls outside that family or range, the augmented training gives no guarantee.","fun_headline_variants_meta":{"raw":{"variants":["Augmented rock physics shrinks CO2 plume forecast errors","Diverse rock models boost CO2 monitoring accuracy","10 rock-physics variants cut CO2 plume uncertainty","Robust CO2 tracking with multiple rock physics models","Mixing rock physics improves CO2 storage forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000533,"raw_usage":{"total_tokens":2539,"prompt_tokens":896,"completion_tokens":1643,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1568}},"tokens_in":512,"tokens_out":1643,"duration_ms":10240,"temperature":1.0,"reasoning_tokens":1568,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T13:34:20.217729+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same augmented digital shadow on a test seismic image generated with a rock physics law that is deliberately outside the Brie family, for example a different effective-medium mixing model, or with a Brie exponent above 10. If the augmented shadow's posterior mean error and uncertainty are no better than the non-augmented baseline, the paper's claim of robustness to unknown rock physics models is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Brie saturation model family and the uniform-versus-patchy terminology used for the augmentation."}],"review_version":1}