{"id":"1467e04f-201d-4903-b94c-125016440d1e","arxiv_id":"2506.00214","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A diffusion-based framework (Diff-SPORT) reconstructs urban turbulent flows from sparse sensors and ranks sensor locations via Shapley values, outperforming existing reconstruction and placement baselines on a simulated building flow.","lead":"Diff-SPORT is a machine-learning framework that reconstructs turbulent wind patterns around buildings from a small number of sensors and then chooses where to put those sensors. It combines a diffusion model, a Bayesian inference step, and a game-theory attribution method, and the authors show it outperforms current baselines on a simulated urban flow.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-geometry, single-Reynolds, 2D validation cannot support the zero-shot 'foundation model' claim for urban flows; a held-out geometry test is required.","rationale":"The paper is a competent application of diffusion priors to sparse reconstruction and sensor placement, and the in-distribution comparison against PiGDM and QR-pivoting is a reasonable starting point. However, the headline contribution is the 'zero-shot, modular alternative for urban flow monitoring,' and the evidence for that is a single canonical geometry at a single Reynolds number, evaluated on the tail of the same simulation. That is a mismatch between claim and evidence. The reader's weakest_assumption identifies exactly this gap. I agree. A second, compounding issue is temporal dependence: with 26,000 snapshots over 130 convective time units, the last 5% spans 6.5 convective units; if the shedding period is O(5-10) convective units, the test set may be heavily correlated with the training set, further inflating apparent generalization. But the dominant concern is external validity. The proposed test—zero-shot application to a different geometry—would directly settle whether p_theta(Psi) transfers. If it does not, the framework is still a useful proof of concept, but the abstract and conclusions must be tempered to avoid overclaiming. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":15509,"tokens_out":5311,"duration_ms":56026,"concrete_test":"Take the already-trained Diff-SPORT checkpoint and, with no fine-tuning, apply MAP-GA to reconstruct DNS or LES data of a different urban-like geometry (e.g., two square cylinders in tandem or a staggered cube array) at a comparable Reynolds number, using the same 15% feasible mask and MSE metric. If the out-of-distribution reconstruction error is within, say, 2x of the in-distribution error, the foundation-model claim gains support; if it is substantially larger, the paper must be reframed as single-case validation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Diff-SPORT is a zero-shot, modular alternative for urban flow monitoring rests on the diffusion prior p_theta(Psi) being representative of urban flow distributions. This is never tested. The prior is trained on a single DNS of flow around a wall-mounted square cylinder at Re_h=2000 (Methods, 'Numerical simulation and flow description'), on one 2D mid-span slice (z/h=0), and evaluated on the last 5% of the same simulation. No held-out geometry, Reynolds number, or 3D configuration is considered. The Discussion itself lists 'improve generalization to out-of-distribution flows' and 'training foundation models' as future work, which contradicts the abstract's foundation-model framing. Moreover, the 5% test span (6.5 convective time units) may be shorter than the vortex-shedding period, so even the in-distribution reconstruction numbers may not reflect statistically independent samples. Without an out-of-distribution test, the practical urban-monitoring claims are unsupported regardless of how well MAP-GA and SHAP perform on the training distribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Diff-SPORT, a three-stage framework: a DDPM trained on DNS data of flow around a wall-mounted square cylinder at Re_h=2000, a MAP gradient-ascent (MAP-GA) scheme for sparse reconstruction from masked velocity fluctuations, and a SHAP-based sensor-placement method with a modified coalition weighting kernel. The authors evaluate unconditional generation statistics, conditional reconstruction against ΠGDM and unconditional DDPM, and sensor placement against QR-pivoting and random placement. They report that MAP-GA outperforms ΠGDM by roughly a factor of three and that SHAP-based placement outperforms random placement and matches or slightly exceeds QR-pivoting, especially at low sensor counts. The paper frames the pre-trained diffusion prior as a zero-shot, modular 'foundation model' for urban flow monitoring.","tokens_in":15745,"tokens_out":3912,"duration_ms":43493,"significance":"If the central claims hold, the framework is a useful modular contribution: a single pre-trained diffusion prior, used without retraining, can serve both sparse reconstruction for arbitrary masks and interpretable sensor placement. The unconditional generation results, including second-order statistics and PDFs, provide genuinely supporting evidence for the quality of the prior. The comparison with external baselines (ΠGDM, QR-pivoting, random placement) and the use of multiple MAP-GA runs are strengths. The main significance is therefore conditional: the methodology is promising and the in-distribution evidence is substantial, but the paper's broad 'urban environment' and 'foundation model' claims are not supported by the presently tested single-geometry, single-Reynolds-number, two-dimensional setting.","major_comments":[{"comment":"The generalization claim is load-bearing and unsupported. The diffusion prior is trained on one DNS of flow around a wall-mounted square cylinder at Re_h=2000, on a single 2D mid-span slice, and evaluated on the last 5% of the same simulation. The Discussion itself states that future work should 'improve generalization to out-of-distribution flows' and 'training foundation models', which directly contradicts the abstract's and Introduction's characterization of Diff-SPORT as a 'zero-shot alternative' and the diffusion model as a 'foundation model'. At minimum, the claims should be restricted to in-distribution reconstruction for the canonical case, or an out-of-distribution test (different geometry, Reynolds number, or three-dimensional configuration) should be added.","section":"Discussion and conclusions; Methods (Numerical simulation and flow description)"},{"comment":"The evaluation mode in Figure 4(c) is not achievable in deployment. Selecting, for each test field, the MAP-GA run with the lowest reconstruction error requires access to the ground-truth field; the paper's statement that this 'best selection strategy can be utilized in practical scenarios, leveraging the temporal error evolution curves' does not explain how the error curve would be known without ground truth. This mode removes the stochasticity of MAP-GA from the comparison and can bias the ranking of placement methods. The deployment-relevant comparison is panel (d), which includes run-to-run variability; the claims that SHAP 'matches or exceeds' QR-pivoting should be based on panel (d) or on a selection rule that does not use ground-truth error.","section":"Fig. 4(c)-(d), Optimal sensor placement"},{"comment":"The SHAP kernel range [k_min, k_max] is fitted to the data rather than prescribed or validated on independent data. The text states that the range is 'determined empirically from random baseline experiments' and that 'optimal performance typically observed in the 3.6-9% pixels range, as seen in figure 4(c)-(d)'. Since the same figures are used to report the final SHAP versus random comparison, the sensor-placement evaluation is not fully independent of the choice of this hyperparameter. The authors should either fix the kernel range from a validation split, report sensitivity of the conclusions to this range, or provide a principled selection criterion.","section":"Methods (Shapley values for optimal sensor placement), Eq. (15)"},{"comment":"The measurement model and the practical deployment scenario are mismatched. The paper subtracts the time-averaged mean field and reconstructs only the fluctuation tensor Ψ(x,y,t), with the mean field 'considered to be known from the flow statistics'. In a real urban deployment, sensors measure the total velocity, and the local mean wind is generally not known from a prior DNS of the same flow. This assumption is central to the practical urban-monitoring claims. The paper should state clearly that the method reconstructs fluctuations conditioned on a known mean, or it should demonstrate how the mean is obtained in deployment without access to the simulation statistics.","section":"Results (Overview), Eq. (1), Eq. (6)"}],"minor_comments":[{"comment":"The phrase 'overbars denote ensemble averages in time' is internally inconsistent; an ensemble average is not a time average. The authors likely mean a time average under the assumption of statistical stationarity, and this should be reworded.","section":"Methods (Equation (1))"},{"comment":"The claim that 'MAP-GA outperforms ΠGDM by nearly a factor of three in terms of accuracy and precision' is stated without numerical support. Reporting the mean and standard deviation of the error metric in a table would make the factor-of-three claim verifiable.","section":"Results (Sparse reconstruction, Figure 3)"},{"comment":"There is a duplicated word in the sentence introducing α_t ('where where α_t = ...'). This should be corrected.","section":"Methods (Equation (3b))"},{"comment":"The random baseline uses seven masks per sensor count, but the error envelopes in Figure 4 do not appear to include confidence intervals or statistical significance tests. Adding error bars or a paired comparison would strengthen the claim that SHAP 'consistently outperforms' random placement.","section":"Figure 4 and Methods (Random baseline)"},{"comment":"The manuscript does not include a data or code availability statement. For a methods paper centered on a computational pipeline, providing access to the trained models or code would substantially aid reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper's technical core is sound in-distribution, but the framing considerably oversells the evidence. The 'foundation model' and 'urban environments' language in the title and abstract should be tempered unless out-of-distribution validation is added. The best-of-MAP-GA selection in Figure 4(c) should not be presented as a deployable strategy. I would not recommend rejection because the methodology itself is a reasonable contribution and the in-distribution comparisons are informative; however, the required revisions affect the central claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a solid, incremental extension of the authors' prior work. They already used a diffusion prior for sensor placement and sparse reconstruction in their CTR 2024 proceedings, and the MAP-GA optimizer comes from their WACV 2025 paper. What is new here is the packaging — MAP-GA plus a SHAP-based attribution layer with a custom kernel — evaluated on a 2D mid-span slice of a wall-mounted square cylinder at Re_h=2000. The unconditional generation statistics (Reynolds stresses, PDFs) are convincing, and the reconstruction comparison against PiGDM shows a genuine, large improvement, roughly a factor of three. The SHAP versus random/QR comparisons look fair as far as I can tell, and the paper is honest about several limitations in the Discussion.\n\nThe soft spots are real. First, Fig. 4c reports the best MAP-GA run per test field. That selection is not achievable in deployment because you do not have ground truth to pick the winner; the paper's appeal to \"temporal error evolution curves\" is hand-wavy. The companion panel with mean over runs is more honest, but the abstract leans on the best-run numbers. Second, the SHAP kernel range is fit to random-placement experiments, and no sensitivity analysis is shown, so it is hard to know how robust the sensor rankings are to that choice. Third, and most important, the \"foundation model\" / \"zero-shot urban monitoring\" claim is not supported by the evidence. The prior is trained on one geometry at one Reynolds number, on one 2D slice, with the mean subtracted. The test set is the last 5% of the same simulation, which amounts to only 6.5 convective time units — possibly less than one vortex-shedding period, so even the in-distribution numbers may not be statistically stable. The paper itself lists \"improve generalization to out-of-distribution flows\" and \"training foundation models\" as future work, undercutting the abstract.\n\nNone of this is fatal if the paper is framed as a methodological demonstration on a canonical case. The practical claims need a held-out geometry or Reynolds test, and the evaluation should avoid best-run selection and report variability. No code or data release is a minor additional concern, but the core method is described well enough to reimplement.\n\nBottom line: worth a serious referee. The authors should be asked to temper the framing, add at least one out-of-distribution test, and ideally release artifacts.","headline":"Competent incremental extension of the authors' own diffusion-based reconstruction work, but the 'foundation model for urban flows' framing is not supported by the single-geometry, 2D validation.","tokens_in":16301,"tokens_out":3471,"would_cite":false,"duration_ms":35530,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["47.27.-i"],"model":"deepseek-v4-flash","headline":"Diff-SPORT reconstructs turbulent city winds from sparse, deployable sensors using a single diffusion prior.","keywords":["diffusion models","sparse reconstruction","optimal sensor placement","turbulent flow","urban flow","Shapley values","maximum a posteriori estimation","zero-shot inference"],"falsifier":"Run the same trained Diff-SPORT prior on a held-out second geometry, a different Reynolds number, or the full 3D volume instead of the 2D mid-span slice; if the mean reconstruction error rises sharply relative to the in-distribution baseline, the foundation-model claim is not supported.","tokens_in":15315,"feed_emoji":"🌬️","tokens_out":5645,"duration_ms":57374,"temperature":0.7,"pith_summary":"Diff-SPORT aims to show that one diffusion model trained on high-fidelity simulation data can handle both halves of the urban-flow monitoring problem: reconstructing the full instantaneous velocity-fluctuation field from a sparse, practically feasible sensor mask, and deciding where those sensors should be placed. The paper argues that its MAP-GA reconstruction, which optimizes the posterior directly through the diffusion prior's gradients, is roughly three times more accurate and stable than the score-based ΠGDM baseline, and that a Shapley-value attribution with a custom coalition kernel selects sensor subregions that outperform random placement and match QR-pivoting. If correct, a single pretrained model would serve multiple monitoring tasks without retraining, at inference speeds compatible with near-real-time use. The demonstration uses a 2D mid-span slice of flow around a wall-mounted square cylinder at Reynolds number 2000, with the time-mean field subtracted beforehand.","feed_headline":"Diffusion model rebuilds turbulent city winds from a 15% sensor mask","feed_subtitle":"One pretrained prior reconstructs flows and ranks sensor spots, beating random placement.","key_machinery":"The load-bearing object is the pretrained DDPM acting as a probabilistic surrogate for the flow-field distribution, with two wrappers around it. MAP-GA treats the reverse diffusion chain as a deterministic map $f_\\theta$ from noise $\\Psi_T$ to a clean field $\\Psi_0$ and maximizes $\\log p(f_\\theta(\\Psi_T)|S)$ by gradient ascent, approximating the Jacobian with the Tweedie-type denoiser $E(\\Psi_0|\\Psi_\\tau)$ at each of 20 diffusion steps with 50 ascent iterations per step. For sensor placement, kernelSHAP evaluates a value function $v(C)=-\\mathrm{MSE}(\\Psi_{\\mathrm{DNS}},\\Psi_0^{(C)})$ over sensor coalitions $C$, using a modified weighting kernel $\\pi_{\\mathrm{mod}}$ that zeros out coalitions outside an empirically chosen 3.6\\% to 9\\% pixel range to avoid misleading attributions from very sparse or very dense coalitions.","core_discovery":"The central claim is that a variance-preserving DDPM trained on 26,000 mean-subtracted snapshots of streamwise and vertical velocity fluctuations acts as a reusable probabilistic prior for urban-scale turbulent flow. Unconditionally, it generates samples that match DNS in second-order statistics and full probability distributions; conditionally, MAP-GA reconstructs instantaneous fields from as little as 15% of the domain, using only near-ground and obstacle-adjacent regions and no wake sensors, with errors concentrated in the wake yet low overall. MAP-GA replaces the learned score with a deterministic denoiser map and performs gradient ascent on the posterior, yielding a sharper error distribution and roughly three times better accuracy and precision than ΠGDM. The same prior is then used inside kernelSHAP with a custom coalition kernel, where the value function is negative reconstruction MSE for each sensor coalition; thresholding the resulting importance map gives spatially coherent sensor subregions that beat random placement and are competitive with QR-pivoting, especially at tight sensor budgets.","pith_inferences":["Editorial inference: because the mean field is subtracted before training, real deployments must supply the time-averaged flow separately; Diff-SPORT's current results assume that mean is known, which the paper does not quantify.","Editorial inference: the SHAP value function uses ground-truth DNS fields to compute MSE, but in live monitoring no ground truth exists; a surrogate error metric would be needed for practical attribution, and the paper does not test ranking stability under noisy or approximate errors.","Editorial inference: the framework is demonstrated on one 2D slice; applying the same prior to a 3D volume or to temporal sequences would test whether the learned distribution captures spanwise and time correlations, which the current evaluation does not cover.","Editorial inference: a natural extension is to reuse the same prior with different forward operators (point sensors, partial fields, or noisy measurements), since zero-shot MAP-GA is operator-agnostic; the paper only tests noiseless masked-region inpainting."],"forward_implications":["With a 15% coverage baseline mask and no wake sensors, MAP-GA reconstructs instantaneous velocity fields with errors mostly confined to the wake and low overall magnitude.","SHAP-generated sensor placements are spatially coherent, support flexible sensor budgets, and outperform random placement while matching QR-pivoting under tight budgets.","A single trained diffusion model handles unconditional generation, sparse reconstruction, and sensor attribution without retraining, making the pipeline modular and zero-shot.","At roughly 12 seconds per snapshot on one A100 GPU, with batched and parallel inference, the approach is compatible with near-real-time urban monitoring."],"supporting_citations":[{"why":"Supplies the DDPM training objective and reverse-process parameterization that define the generative prior.","marker":"[39]"},{"why":"Supplies the MAP-GA algorithm that the paper adapts for sparse reconstruction and benchmarks.","marker":"[35]"},{"why":"The pseudoinverse-guided diffusion baseline that MAP-GA is compared against and outperforms.","marker":"[27]"},{"why":"Supplies kernelSHAP, which the paper adapts with a custom coalition kernel for sensor attribution.","marker":"[38]"},{"why":"Provides the DNS dataset of the wall-mounted square cylinder at Re_h=2000 used for training and testing.","marker":"[42]"},{"why":"The QR-pivoting baseline used for sensor placement comparison.","marker":"[28]"}],"fun_headline_variants":["Diffusion prior turns 15% sensor data into full urban wind fields","One pretrained model reconstructs flows and picks optimal sensor spots","Generative model speeds up urban wind monitoring with smart sensor placement","Zero-shot diffusion reconstructs urban winds and ranks sensor placements","Fast, interpretable sensor placement via diffusion model for urban flows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The diffusion prior is trained only on a 2D mid-span slice of a single wall-mounted square-cylinder flow at Reynolds number 2000, and the paper treats it as a general urban-flow foundation model, so if this prior does not transfer to other geometries, Reynolds numbers, or 3D fields, the practical claims degrade.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion prior turns 15% sensor data into full urban wind fields","One pretrained model reconstructs flows and picks optimal sensor spots","Generative model speeds up urban wind monitoring with smart sensor placement","Zero-shot diffusion reconstructs urban winds and ranks sensor placements","Fast, interpretable sensor placement via diffusion model for urban flows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1577,"prompt_tokens":890,"completion_tokens":687,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":600}},"tokens_in":506,"tokens_out":687,"duration_ms":7018,"temperature":1.0,"reasoning_tokens":600,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:09:36.719098+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same trained Diff-SPORT prior on a held-out second geometry, a different Reynolds number, or the full 3D volume instead of the 2D mid-span slice; if the mean reconstruction error rises sharply relative to the in-distribution baseline, the foundation-model claim is not supported.","supporting_citations":[{"cited_title":"Manohar, B","cited_arxiv_id":null,"evidence_quote":"The QR-pivoting baseline used for sensor placement comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DDPM training objective and reverse-process parameterization that define the generative prior."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MAP-GA algorithm that the paper adapts for sparse reconstruction and benchmarks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The pseudoinverse-guided diffusion baseline that MAP-GA is compared against and outperforms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies kernelSHAP, which the paper adapts with a custom coalition kernel for sensor attribution."},{"cited_title":"Mart ´ ınez-S´ anchez, E","cited_arxiv_id":null,"evidence_quote":"Provides the DNS dataset of the wall-mounted square cylinder at Re_h=2000 used for training and testing."}],"review_version":1}