{"id":"cb2afb6a-a8c9-4473-9e10-5749dbc69fd9","arxiv_id":"2507.05143","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A generalized Wasserstein-2 loss with a differentiable surrogate lets stochastic neural networks reconstruct random field models whose outputs combine continuous and categorical variables.","lead":"The authors propose a generalized Wasserstein-2 distance for training stochastic neural networks when the target output mixes continuous and categorical components. They prove a universal approximation result and demonstrate the loss on classification, mixed-distribution reconstruction, and noisy gene-regulation dynamics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2.1 rests on an unstated black-box approximation theorem from the authors' prior preprint, and the proof explicitly assumes away the regularity conditions on D; until those conditions are verified, the central approximation claim is only conditional.","rationale":"The manuscript has a clear, actionable proposal: replace the categorical metric by a bounded truncated-quadratic cost, smooth the categorical components, and inherit a universal approximation theorem. The experiments are nontrivial and the sensitivity tests are a genuine plus. I read the metric definition charitably: despite the notation in (2.1) and (2.7), the quantity \\hat W_2 can be interpreted as W2 for the metric D(y,\\hat y)=sqrt(λ∑(y_i-\\hat y_i)^2 + ∑ min(2|y_j-\\hat y_j|,1)^2), so the triangle inequality used in Appendix B is not an obvious logical error. The central weakness is not an internal inconsistency I can exhibit; it is an unstated dependence. Appendix A does not prove the SNN approximation theorem; it takes it from [34, Appendix H] and explicitly assumes away the conditions on D. Since Theorem 2.1 is the paper's headline theoretical contribution, the validity of the paper rests on a black-box from a preprint that is not even cited with a precise statement. That warrants keeping the CONDITIONAL verdict: the approximation claim is plausible but unverified as written. A concrete check -- listing the hypotheses of [34, Appendix H] and verifying them for f_{ε,x} -- would either close the gap or reveal a genuine counterexample.","tokens_in":21351,"tokens_out":18262,"duration_ms":220713,"concrete_test":"Obtain [34, Appendix H], list the theorem's exact hypotheses, and check them for f_{ε,x} in (A.2) with D as used in Theorem 2.1. In particular, verify (i) the regularity conditions on D and ν, (ii) f_{ε,x} is in the required Sobolev/uniform-continuity classes uniformly in x, and (iii) the SNN approximation constant is independent of ε in the needed way for the ε→0 step in (A.13). If any condition fails, construct a target random field satisfying Assumption 2.1 whose smoothed f_{ε,x} violates it; that would refute the proof as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 2.1, the paper's central theoretical claim, is a reduction to an external universal approximation theorem for continuous random fields, [34, Appendix H], which is never stated in this manuscript. After smoothing the categorical coordinates in (A.2), the proof verifies the smoothness/Lipschitz conditions it can see from Assumption 2.1 and then applies [34, Appendix H] to obtain (A.9). The appendix adds: '[34, Appendix H] also imposes some technical regularity conditions on the bounded set D for x in Eq. (1.1). For simplicity, we assume those conditions hold here.' This is the load-bearing step. If [34, Appendix H] requires more than Assumption 2.1 -- for example a specific geometry for D, a product or absolutely continuous sampling measure ν, or control of f_{ε,x} in a function class that is not inherited from f_x for every ε -- then (A.9) is unproved and Theorem 2.1 collapses to a conditional statement. No independent proof of the external theorem is supplied, and the theorem is not reproduced, so a reader cannot check the reduction. This gap, not the surrogate loss or the unverified Lipschitz assumptions in the experiments, is the primary obstacle to the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a generalized Wasserstein-2 distance for random field models whose outputs contain both continuous and categorical components. It defines a generalized squared W2 loss, proves a universal approximation theorem for stochastic neural networks under this distance (Theorem 2.1) and a generalization error bound for the empirical local loss (Theorem 2.2), then introduces a differentiable surrogate loss and tests the method on classification, mixed continuous/categorical regression (abalone), and a gene-regulatory ODE/Markov jump system.","tokens_in":21619,"tokens_out":13970,"duration_ms":155310,"significance":"If the theoretical results are made fully rigorous, the paper offers a useful principled extension of local squared Wasserstein training to categorical and mixed outputs, a setting where the usual W2 distance is not directly applicable. The differentiable surrogate with a detach-based rounding trick is a practical contribution, and the three diverse UQ experiments plus the sensitivity analyses give useful empirical evidence. The main limitation is that the central approximation theorem currently rests on an unstated external theorem and a smoothing construction that is not fully justified, so the theoretical support does not yet match the strength of the claims.","major_comments":[{"comment":"The proof of Theorem 2.1 reduces the approximation of the smoothed measure f_{ε,x} to the universal approximation theorem of [34, Appendix H], but that theorem is neither stated nor proved in this manuscript, and the text explicitly says 'For simplicity, we assume those conditions hold here' concerning technical regularity conditions on D. Since [34] is a preprint and its hypotheses are not checkable from the present paper, Theorem 2.1 is conditional at its load-bearing step. Please state the external theorem in full and either prove it or verify all of its hypotheses, or replace the reduction with a self-contained proof.","section":"Appendix A, Eq. (A.9)"},{"comment":"The definition of f_{ε,x} in Eq. (A.2) is not a probability density as written: it appears to depend on a single discrete y with no summation over the categorical support, and it is not normalized. Moreover, the convolution with φ_ε acts only on the categorical coordinates, so the claimed smoothness of f_{ε,x} in all d coordinates does not follow from the uniform continuity of f_x in the continuous components stated in Assumption 2.1.7. This affects the verification of Eqs. (A.4) and (A.9). Please define f_{ε,x} as a genuine convolution in all coordinates (for example, summing over y_2 and mollifying the continuous coordinates as well) and prove that the mixed-derivative bounds in Assumption 2.1 are inherited.","section":"Appendix A, Eqs. (A.2)-(A.4)"},{"comment":"The proof of Theorem 2.2 applies the triangle inequality to \\hat W_2, but the cost defined in Eq. (2.1) is not shown to make \\hat W_2 a metric and it is not a norm: for a one-dimensional continuous component, \\|(0.8,0,\\dots,0)\\| = 1, while 2\\|(0.4,0,\\dots,0)\\| = 1.6, so homogeneity fails. The paper calls Eq. (2.1) a norm and claims the coefficient 4 ensures this, but that claim is false. A lemma establishing that the generalized distance satisfies the triangle inequality (or a different argument) is needed before Eq. (B.7) can be used.","section":"Section 2, Eqs. (2.1)-(2.2); Appendix B, Eq. (B.7)"},{"comment":"The loss actually minimized in all experiments is the differentiable surrogate |\\cdot|_1 with the detach/straight-through estimator in Eq. (2.25), while Theorems 2.1 and 2.2 are stated for the generalized \\hat W_2 built from Eq. (2.1). The equality between |y-\\hat y|_1 and \\|y-\\hat y\\| holds only when both variables are rounded to integers, which is not the training phase. No argument shows that minimizing the surrogate controls \\hat W_2 or that the detach operation preserves the objective. Please either prove a relation between the surrogate and \\hat W_2, or explicitly present Eq. (3.1) as a heuristic whose validation is empirical.","section":"Section 2.3, Eqs. (2.23)-(2.26)"}],"minor_comments":[{"comment":"The notation \\hat\\delta_{y_j,0} has one argument, but Eqs. (2.14)-(2.15) and (D.1) use a two-argument form \\hat\\delta_{y_i,\\hat y_i}; please define this two-argument version explicitly.","section":"Section 2, Eq. (2.2)"},{"comment":"The notation \\|\\cdot\\|_2 is used both for the generalized quantity in Eq. (2.1) and for the Euclidean ℓ2 norm in Eq. (2.7), which is confusing; please use distinct symbols.","section":"Section 2, Eq. (2.1)"},{"comment":"Assumptions 2.1.3 and 2.1.4 require the conditional distributions to be uniformly Lipschitz in x under the generalized W2 distance, but the experiments provide no diagnostic or discussion of whether this regularity condition is plausible for the tested models.","section":"Section 3, Examples 3.1-3.4"},{"comment":"All experimental results appear to be single-run and no error bars, confidence intervals, or multiple-seed variability are reported, which makes it difficult to judge whether the observed differences between methods are significant.","section":"Section 3, Figures 2-4 and Table 2"},{"comment":"The sentence 'For predicting the categorical sex variable on the testing set' appears to be a copy-paste from Example 3.3; in Example 3.2 the target is a multilabel binary vector, not a sex variable.","section":"Section 3.2, paragraph after Eq. (3.3)"},{"comment":"The detach operation in round_1 is a straight-through estimator and deserves a more explicit explanation, since it is central to the practical training procedure but is mentioned only in one sentence.","section":"Section 2.3, Eq. (2.25)"}],"recommendation":"major_revision","confidential_remarks":"The heavy reliance on the authors' own prior preprint [34] is not itself disqualifying, but because the central theorem is reduced to an unstated and unverified external result, the manuscript should not be accepted in its current form. The authors should also clarify the status of the surrogate loss relative to the theory; if it is meant as a heuristic, the paper should say so explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a plausible extension of the authors' earlier local squared Wasserstein-2 framework to mixed continuous/categorical random variables, but the main approximation theorem leans on an unstated external result and the proof kicks the regularity conditions down the road. The experiments are promising but thin.\n\nWhat's genuinely new: the generalized distance in Eq. (2.1) with its hybrid pseudo-norm, the differentiable surrogate in (2.23)-(2.25), and the application to targets with both continuous and discrete components. The gene toggle example (Example 3.4) is a nice real-world stress test, and the classification and abalone experiments show competitive performance against sensible baselines. The surrogate loss is a pragmatic fix for the non-differentiable delta.\n\nThe paper does what it claims well enough to deserve a real referee, but there are real soft spots. Theorem 2.1 reduces to [34, Appendix H] after smoothing, and the appendix literally says 'For simplicity, we assume those conditions hold here' about the regularity of D. That is the load-bearing step, and a reader cannot check the reduction because the external theorem is never stated. The practical loss is a surrogate, and no theorem ties it to the distance in the approximation or generalization statements. The experiments are single-run with no error bars, and the Lipschitz assumptions in Assumption 2.1.3-4 are not verified in any example.\n\nThese are addressable. I'd send it to review expecting heavy revision: state the external theorem, verify or weaken its hypotheses, add error bars and release code, and either prove a surrogate-distance gap or qualify the claim. The core idea — smoothing categorical components and using a generalized W2 loss — is sensible and likely useful for uncertainty quantification in mixed-type prediction.\n\nWho's it for: people working on stochastic neural network approximation or Wasserstein training for UQ. Not a paradigm shift, but a solid extension that should be reviewed properly.","headline":"A plausible extension of local W2 training to mixed continuous/categorical random variables, but the central approximation theorem depends on an unstated prior theorem and hand-waved regularity conditions; worth a serious referee who expects revision.","tokens_in":22121,"tokens_out":2033,"would_cite":false,"duration_ms":21875,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60A05","68Q87","65C99"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that stochastic neural networks can approximate any sufficiently regular random field with mixed continuous and categorical outputs to arbitrary accuracy under a generalized Wasserstein-2 distance, and trains them with a…","keywords":["Wasserstein distance","uncertainty quantification","random field reconstruction","mixed random variable","stochastic neural network","universal approximation","local squared Wasserstein-2 loss"],"falsifier":"Take a mixed random field with bounded outcomes whose categorical class probabilities jump discontinuously at a threshold in $x$, so the uniform Lipschitz condition in Eq. (2.8) fails, and train an SNN with the proposed local squared loss; if the learned distributions still converge to the target under $\\hat W_2$, the Lipschitz assumption is unnecessary, and if they do not, the theorem's stated scope is confirmed. A complementary check is to verify the technical regularity conditions on $D$ named in [34, Appendix H] for a concrete domain used in the experiments; if those conditions are violated, the proof of Theorem 2.1 is incomplete as written.","tokens_in":21133,"feed_emoji":"🎲","tokens_out":9136,"duration_ms":94142,"temperature":0.7,"pith_summary":"This paper proposes a way to train stochastic neural networks so that they reproduce the full conditional distribution of a random field model whose output mixes continuous and categorical components, not just its mean. The central theoretical claim is a universal approximation result: under boundedness, Lipschitz, and regularity conditions, an SNN can make the integrated squared generalized Wasserstein-2 distance from its output distribution to the true field smaller than any prescribed positive tolerance. To make this usable from finite data, the paper defines a generalized Wasserstein-2 distance whose categorical part is a bounded discrepancy, introduces a differentiable surrogate so gradients flow through discrete targets, and proves a finite-sample bound showing that a local empirical version of the loss converges to the true distance. In experiments on classification, a mixed abalone dataset, and a stochastic gene-toggle dynamical system, the trained SNN matches or improves on comparison methods for uncertainty quantification.","feed_headline":"Generalized Wasserstein-2 loss trains SNNs on mixed random fields","feed_subtitle":"One network reconstructs continuous uncertainty and categorical class probabilities at once, with an approximation proof and finite-sample…","key_machinery":"The load-bearing object is a modified norm on the mixed vector, $\\|y\\|^2=\\lambda\\sum_{i=1}^{d_1} y_i^2 + \\sum_{j=d_1+1}^d \\hat\\delta_{y_j,0}$, where $\\hat\\delta_{y_j,0}=4y_j^2$ for $|y_j|\\le 1/2$ and $1$ otherwise. This makes categorical mismatches contribute a bounded, almost quadratic penalty instead of an unbounded Euclidean distance, and it defines the generalized Wasserstein-2 distance $\\hat W_2$ between conditional distributions. The approximation proof smooths the categorical components with a mollifier so that the known SNN universal approximation theorem for continuous random fields applies, then transfers the resulting bound back to the original mixed measure by triangle inequalities. The training side uses a local empirical version of $\\hat W_2^2$, with a differentiable surrogate for the categorical discrepancy, quadratic when prediction is near the label and cosine-based when it is far, so that gradient descent can train through discrete targets.","core_discovery":"On its own terms, the paper's central discovery is Theorem 2.1: for any random field $y_x=y(x;\\omega)$ whose conditional distributions $f_x$ are uniformly bounded, uniformly Lipschitz in $x$ under the generalized Wasserstein-2 metric, and satisfy the smoothness conditions in Assumption 2.1, and for any $\\epsilon_1>0$, there exists a stochastic neural network whose output distribution $\\hat f_x$ satisfies $\\int_D \\hat W_2^2(f_x,\\hat f_x)\\,\\nu(dx)\\le \\epsilon_1$. Continuous outputs are rounded at test time to deliver categorical predictions, while training uses a surrogate of the categorical discrepancy that keeps gradients nonzero. The paper also proves Theorem 2.2, a finite-sample generalization bound showing that the local empirical loss approaches the true squared distance at a rate controlled by $1/\\sqrt N$, the neighborhood size $\\delta$, and a dimension-dependent empirical-measure term. These two results underwrite the practical recipe: minimize the generalized local squared Wasserstein-2 loss over the means and variances of the stochastic network weights.","pith_inferences":["The categorical penalty saturates at 1 for well-separated labels, so the loss effectively ignores how different two wrong categories are; permuting the category codes should leave training unchanged, a testable consequence the paper does not run.","The mollification step is a general transfer principle: any universal approximation theorem for continuous random fields should lift to mixed random fields under the same assumptions, so the proof strategy could be reused for other stochastic network architectures without redoing the argument.","The bound in Theorem 2.2 has three competing terms depending on the neighborhood size $\\delta$; balancing $1/\\sqrt N$, $h(N(x,\\delta),d)$, and $L\\delta$ would give a principled rule for choosing $\\delta$, which the paper leaves to heuristic search.","The method only needs local empirical distributions of outputs, so it should extend to online or streaming settings where minibatches are refreshed during training, though the paper does not analyze this regime."],"forward_implications":["If Theorem 2.1 is correct, SNNs are a universal model class for mixed-output random fields: no separate continuous and categorical architectures are needed for distributional reconstruction.","The local squared loss provides a finite-data objective whose error relative to the true distance is controlled by three explicit terms, sample size, local neighborhood statistics, and Lipschitz regularity, so users can predict when reconstruction will be accurate.","Because the loss is differentiable and local, it can be used as a drop-in replacement for mean squared error or cross-entropy in uncertainty-quantification pipelines for mixed targets.","The learned per-input randomness comes from sampling stochastic weights, so repeated forward passes at a fixed $x$ give a Monte Carlo estimate of the full conditional distribution.","For multivariate categorical targets, encoding them as a single categorical variable improves both accuracy and runtime, matching the dimension dependence in the bound."],"supporting_citations":[{"why":"Supplies the black-box universal approximation theorem for SNNs approximating continuous random fields, which the proof of Theorem 2.1 invokes after smoothing the categorical components.","marker":"[34]"},{"why":"Provides the local squared Wasserstein-2 loss formulation and the generalization-bound template that Theorem 2.2 extends to mixed variables.","marker":"[33]"},{"why":"Gives the Wasserstein triangle inequality used in the proof of Theorem 2.2 to decompose the distance between true and empirical local measures.","marker":"[4]"},{"why":"Provides the empirical-measure convergence rate in Wasserstein distance used to bound the local empirical error term $h(N(x,\\delta),d)$.","marker":"[6]"},{"why":"Supplies the gene toggle model, an ODE-Markov jump system, used in Example 3.4 to validate reconstruction of coupled continuous and categorical dynamics.","marker":"[13]"},{"why":"Provides the abalone dataset with continuous measurements and mixed targets, rings and sex, used in Example 3.3.","marker":"[20]"}],"fun_headline_variants":["Generalized Wasserstein-2 loss trains SNNs on mixed fields","Provable random field approximation with a Wasserstein-2 network loss","One stochastic net reconstructs mixed continuous and categorical fields","Wasserstein-2 loss with proven bounds for mixed random field models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the regularity package in Assumption 2.1, above all the uniform Lipschitz continuity of the true conditional distributions in $x$ and the technical conditions on the domain $D$ that the proof inherits from an earlier continuous-field approximation theorem and simply assumes hold, so that the black-box universal approximation result can be applied.","fun_headline_variants_meta":{"raw":{"variants":["Generalized Wasserstein-2 loss trains SNNs on mixed fields","Provable random field approximation with a Wasserstein-2 network loss","One stochastic net reconstructs mixed continuous and categorical fields","Wasserstein-2 loss with proven bounds for mixed random field models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3511,"prompt_tokens":876,"completion_tokens":2635,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":2562}},"tokens_in":492,"tokens_out":2635,"duration_ms":21571,"temperature":1.0,"reasoning_tokens":2562,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:32:24.657213+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a mixed random field with bounded outcomes whose categorical class probabilities jump discontinuously at a threshold in $x$, so the uniform Lipschitz condition in Eq. (2.8) fails, and train an SNN with the proposed local squared loss; if the learned distributions still converge to the target under $\\hat W_2$, the Lipschitz assumption is unnecessary, and if they do not, the theorem's stated scope is confirmed. A complementary check is to verify the technical regularity conditions on $D$ named in [34, Appendix H] for a concrete domain used in the experiments; if those conditions are violated, the proof of Theorem 2.1 is incomplete as written.","supporting_citations":[{"cited_title":"Clement and W","cited_arxiv_id":null,"evidence_quote":"Gives the Wasserstein triangle inequality used in the proof of Theorem 2.2 to decompose the distance between true and empirical local measures."},{"cited_title":"Fournier and A","cited_arxiv_id":null,"evidence_quote":"Provides the empirical-measure convergence rate in Wasserstein distance used to bound the local empirical error term $h(N(x,\\delta),d)$."},{"cited_title":"Jaruszewicz and T","cited_arxiv_id":null,"evidence_quote":"Supplies the gene toggle model, an ODE-Markov jump system, used in Example 3.4 to validate reconstruction of coupled continuous and categorical dynamics."}],"review_version":1}