{"id":"a76b6a7e-01d8-42ca-a44f-339d631e99ec","arxiv_id":"2607.21542","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Zero-flow two-sample test (ZF2ST) derives a test statistic from the midpoint conditional displacement of paired samples, learned on one split and evaluated on another, with valid type-I error control and strong power on structured alternatives.","lead":"A new two-sample test uses the average displacement between paired points from two distributions, evaluated at midpoints, as a directional witness of difference. The method learns a vector field on one data split and checks its alignment on a fresh split, yielding a calibrated test that is powerful on structured high-dimensional shifts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ZFD identification rests on unproved, self-cited zero-flow criterion; if the criterion fails, the test cannot detect certain alternatives.","rationale":"After reading the paper, I find the same load-bearing concern as the reader: the zero-flow criterion (Proposition 2.1) is imported from a self-cited preprint and is not proved in this manuscript. This criterion is essential for the identification of ZFD and for the consistency result, since a failure would make the population alignment zero for all witnesses under some alternatives, leaving the test powerless. The concern is genuine and not manufactured; although I believe the criterion may be true (via a characteristic-function argument), the paper itself does not demonstrate it, and the reader cannot verify the central claim without an external reference. This warrants a conditional verdict until the proof is supplied. Other issues (missing WiTS baseline, no code) are secondary and do not affect correctness of the core logic. Thus I agree with the reader's assessment and recommend no change to the verdict.","tokens_in":15424,"tokens_out":19546,"duration_ms":191723,"concrete_test":"Prove or disprove Proposition 2.1. A promising route: from E[D|M]=0 derive E[D e^{itM}]=0 for all t, which gives φ_P(t)φ_Q'(t) - φ_P'(t)φ_Q(t)=0. Near t=0, φ_Q(t)≠0, so φ_P/φ_Q is constant and equals 1; then the uniqueness of characteristic functions on an interval yields P=Q. Verify this argument handles distributions with atoms and finite second moments. Alternatively, inspect Wang et al. (2026) for a rigorous proof and check for hidden assumptions (e.g., density or support conditions). If the proof is correct, the identification concern is resolved; if it has a counterexample, the paper's central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central identification claim, Proposition 3.1, states ZFD(P,Q)=0 iff P=Q. Its proof (Appendix A.1) invokes Proposition 2.1, the 'zero-flow condition,' which is stated without proof and referenced to the authors' own preprint Wang et al. (2026). The nontrivial direction—vanishing conditional mean E[Y-X|(X+Y)/2=m] implies P=Q—is essential. If this fails (e.g., for distributions with atoms or when characteristic-function zeros create non-uniqueness), then ZFD could vanish for P≠Q, and the population alignment T(P,Q;u)=E<u(M),D> would be zero for every u. Consequently, the test statistic would be centered under the alternative too, giving no power. The paper supplies no independent argument, and Proposition 2.1 is not otherwise derived. This is the linchpin of the paper's validity claims: without it, ZFD is not a valid discrepancy and Proposition 4.2's consistency guarantee becomes vacuous. Under H0, the null centering of the Gaussian calibration only needs the trivial direction (P=Q implies v=0), but the identification and power claims require the converse. The paper's own experiments cannot resolve this because they only cover specific alternatives.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-sample test based on the zero-flow criterion. For independent X~P and Y~Q, with midpoint M=(X+Y)/2 and displacement D=Y-X, the population midpoint velocity v(m)=E[D|M=m] is used to define the Zero-Flow Discrepancy ZFD(P,Q)=E||v(M)||^2. The paper claims that ZFD is a valid discrepancy (zero iff P=Q) and develops ZF2ST, a sample-split testing procedure: a witness field u is learned on one half of the data and evaluated on the other half through a studentized alignment statistic T(P,Q;u)=E<u(M),D>. Conditional on the training split, the held-out alignment scores are i.i.d., yielding an asymptotic normal null distribution; under the alternative, power is characterized by a signal-to-noise ratio. A paired sign-flip calibration is also proposed and shown to control type-I error exactly. Experiments on Gaussian-mixture, Blob, and MNIST benchmarks compare two witness-learning variants (regression and max-SNR) against MMD and classifier-based baselines, with ZF2ST performing well on structured high-dimensional alternatives.","tokens_in":15780,"tokens_out":17399,"duration_ms":154760,"significance":"If the zero-flow criterion is valid, ZF2ST provides an interpretable vector-valued witness for two-sample testing and benefits from a clean sample-splitting construction that decouples witness learning from null calibration. The exact sign-flip calibration is a useful finite-sample guarantee, and the SNR-based power analysis gives practical guidance for witness learning. The experiments are reasonably comprehensive and suggest that the method can outperform strong baselines on structured alternatives. However, the paper's core identification claim rests on an unproved, self-cited theorem, and the consistency result is conditional on an oracle property of the learned witness. These gaps currently limit the theoretical foundation of the method.","major_comments":[{"comment":"The identification of ZFD (Prop 3.1) depends entirely on the zero-flow criterion (Prop 2.1), which is stated without proof and cited to a self-published preprint (Wang et al., 2026). The proof in Appendix A.1 simply invokes this theorem. The converse direction—v≡0 implies P=Q—is nontrivial and is what guarantees power against every alternative. If the criterion fails for some distributions, ZFD can vanish for P≠Q and the consistency claim in Prop 4.2 becomes empty. Please provide a self-contained proof of Prop 2.1 (or a detailed verification) in an appendix; as written, the paper does not establish the central identification claim.","section":"Section 2, Proposition 2.1; Section 3, Proposition 3.1"},{"comment":"The consistency result is conditional on the learned witness u_tr having T(P,Q;u_tr)>0. This condition is not guaranteed by the proposed learning rules (8) and (9) and is not a property of the alternative alone. The statement 'under any fixed alternative H1 for which T...=c′>0' is an oracle conditional power analysis, not a consistency theorem for the procedure. To substantiate the claim that ZF2ST is consistent, the authors need to prove that with high probability the learned witness yields positive alignment when P≠Q (e.g., under universal approximation and bounded optimization error), or explicitly label Prop 4.2 as an oracle result. Section 7 also lists convergence of witness learning as future work, confirming this gap.","section":"Section 4.3, Proposition 4.2"},{"comment":"The abstract and contributions state 'We prove the validity of ZFD.' Given that Proposition 3.1's identification is imported from the unproved Proposition 2.1, this statement overstates what is established. Please rephrase to attribute the zero-flow criterion to Wang et al. (2026) and make the conditional nature of the theoretical results transparent.","section":"Abstract and Section 1"}],"minor_comments":[{"comment":"The statistic bZZF is a t-statistic. The text states that 'for sufficiently large Nte, Theorem 4.1 gives the asymptotic right-tail p-value.' It should be clarified that this Gaussian calibration is an approximation and that the experiments correctly use the exact sign-flip calibration.","section":"Section 4.2, Eq. (5)"},{"comment":"The power approximation (7) is derived informally by replacing bσ_te with σ in the signal term. While the direct consistency proof in Appendix A.4 is rigorous, (7) should be clearly labeled as an approximation, not an asymptotic equality.","section":"Section 4.3, Eq. (7)"},{"comment":"In the HDGM dimension experiment, only one sample size (n=3000) is used. It would be informative to report results at a smaller sample size as dimension increases, since the paper's motivation is high-dimensional alternatives where power is typically limited.","section":"Section 5, Figure 1"},{"comment":"Wang et al. (2026) is cited as an arXiv preprint. Since Proposition 2.1 is load-bearing, the authors should ensure the preprint is publicly available and, ideally, include a version identifier or DOI.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The central identification theorem (Prop 2.1) comes from the authors' own unpublished preprint and is not proved in the manuscript. Given its centrality, I strongly recommend requiring a self-contained proof or a thorough independent verification of this result before publication. The conditional nature of the consistency claim (Prop 4.2) should also be addressed, as the paper currently overstates the theoretical guarantees. The empirical and methodological core is otherwise sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper proposes ZF2ST, a two-sample test built on the zero-flow criterion: for independent X~P, Y~Q, the midpoint velocity v(m)=E[Y-X|(X+Y)/2=m] vanishes iff P=Q. The test learns a vector-valued witness on one data split and evaluates its alignment with held-out pairs, using a studentized statistic and an exact sign-flip calibration. The sample-splitting device is adapted from WiTS, but the vector-valued midpoint witness is a genuine new twist, and the dual representation (Prop 4.1) plus the SNR power analysis (Prop 4.2) provide a clean framework for witness design. The sign-flip calibration is exact under the null, not just asymptotically, and the experiments are honestly reported, including a stated limitation on local covariance changes.\n\nThe load-bearing soft spot is the identification result (ZFD=0 iff P=Q), which is imported without proof from the authors' own preprint (Wang et al., 2026). The trivial direction P=Q => v=0 is obvious. The converse is not, and the consistency and power guarantees here all condition on a learned witness having positive alignment—which is vacuous if the zero-flow criterion can fail for some distributions. The paper says it proves the validity of ZFD, but what it actually does is appeal to an unproved, self-cited theorem. I am not claiming the theorem is false; it may well hold for benign distributions. But it needs a proof in this paper or a citation to an externally verified result. The authors' experiments cannot resolve this by themselves.\n\nTwo smaller issues: WiTS, the natural baseline for a sample-split witness test, does not appear in any of the experiments; and no code is provided, making the empirical comparisons hard to reproduce.\n\nOverall, this is a well-designed test with principled null control and encouraging evidence of power on structured high-dimensional differences. The missing proof and the missing baseline are the main gaps. I would send it to peer review, with the referees asked to focus on the zero-flow identification theorem.","headline":"A well-designed sample-split test whose identification rests on an unproved self-cited theorem—worth refereeing, but the zero-flow condition needs to be proved or independently verified.","tokens_in":16200,"tokens_out":5234,"would_cite":true,"duration_ms":48990,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G20","62H15"],"pacs":[],"model":"deepseek-v4-flash","headline":"Two distributions can be distinguished by the midpoint-conditional mean displacement; a sample-split studentized alignment test on this field is asymptotically normal under the null and consistent under alternatives with positive signal.","keywords":["two-sample testing","zero-flow discrepancy","midpoint velocity","witness learning","sample splitting","sign-flip calibration","flow matching","hypothesis testing"],"falsifier":"Search for two distinct distributions P and Q with finite second moments for which E[Y−X | (X+Y)/2 = m] = 0 holds for every midpoint m. A single such pair would make the zero-flow discrepancy zero despite P≠Q, breaking the identification that the test depends on; the search could begin with atomic distributions or distributions whose characteristic functions vanish on a set.","tokens_in":15366,"feed_emoji":"🎯","tokens_out":8381,"duration_ms":71319,"temperature":0.7,"pith_summary":"The paper is trying to establish that the zero-flow criterion — the midpoint-conditional mean displacement between paired samples — is both a valid population measure of distributional difference and a workable basis for a practical two-sample test. It defines the Zero-Flow Discrepancy, proves it equals zero exactly when the two distributions are identical, and then builds ZF2ST, a sample-splitting test that learns a vector witness on a training split and evaluates its alignment with held-out displacements. The central claim is that the held-out studentized statistic is asymptotically standard normal under the null for any fixed witness, and that the test is consistent whenever the learned witness has positive average alignment with the true midpoint velocity, so flexible neural-network witnesses can be used without breaking type-I error control. A sympathetic reader would care because this offers a way to convert the representational power of learned flow fields into statistically calibrated tests for distribution shift, model criticism, and dataset comparison.","feed_headline":"Midpoint flow field powers a new two-sample test","feed_subtitle":"Learn a vector witness on one split, test on another: the held-out statistic stays valid and catches structured shifts.","key_machinery":"The central object is the midpoint velocity field v(m)=E[Y−X | (X+Y)/2=m], the expected displacement of paired samples given their midpoint; the zero-flow criterion says this field vanishes everywhere exactly when P=Q. The test evaluates a learned vector field u (the witness) through the signed alignment T(P,Q;u)=E[⟨u(M),D⟩], which is zero under the null for any fixed u. By splitting the data, learning u on a training fold and evaluating on an independent test fold, the per-pair alignment scores are i.i.d. conditionally on u, so the studentized mean obeys a standard normal CLT. The power-optimized witness objective maximizes the ratio of alignment to its standard deviation, and sign-flip ran","core_discovery":"On its own terms, the paper's central discovery is that the vector field of expected midpoint displacement—the zero-flow velocity v(m)=E[Y−X | (X+Y)/2=m]—vanishes identically if and only if the two source distributions coincide. The paper defines the Zero-Flow Discrepancy as the expected squared norm of this field, proves a scale-invariant dual representation in which ZFD is the supremum over witness fields of squared alignment divided by witness norm, and then constructs ZF2ST. The test learns a witness field on a training split, forms fixed pairs on a held-out split, and studentizes the mean of per-pair alignment scores. Conditional on the training data, the witness is fixed, so under the","pith_inferences":["The zero-flow criterion, if it holds broadly, suggests a deeper connection between midpoint conditional moments and distributional equality; this might yield other discrepancy measures based on higher-order conditional moments or conditional characteristic functions.","One could extend the sample-splitting scheme beyond midpoints to time-dependent flow velocities, potentially giving a family of flow-based tests that trade off local sensitivity and sample complexity.","The paper's own experiments suggest a limitation: for highly localized within-mode changes, random cross-pairing can dilute the signal; a variant that adaptively selects informative pairs or uses multiple pairings during evaluation might recover power at the cost of more complex dependence structure.","Because the null centering does not depend on witness accuracy, the same framework could be used for distributional comparison in high-dimensional settings where the witness is a deep feature representation, without needing a separate calibration dataset beyond the split."],"forward_implications":["If ZF2ST is correct, two-sample testing can be done with flexible neural-network witnesses without risk of biased null calibration, provided witness learning is confined to a separate split.","The zero-flow discrepancy is a valid population divergence: it is nonnegative, symmetric, and zero exactly when the distributions are equal, so it can be reported as an interpretable measure of mismatch.","The test is consistent for any fixed alternative with positive signal-to-noise ratio, so power approaches one as the evaluation sample grows whenever the learned witness is positively aligned with the true midpoint velocity.","The scale-invariant dual form implies only the direction of the witness matters for the population discrepancy, not its magnitude, which justifies direction-focused training objectives.","The paired sign-flip procedure gives finite-sample type-I error control under the null without retraining or re-evaluating the witness."],"fun_headline_variants":["Zero-flow two-sample test: learn witness, then test","Split-witness test for distributional shifts","Zero-flow discrepancy detects structured differences","Midpoint-flow field powers valid two-sample test"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The test's validity rests entirely on the zero-flow theorem—that a zero conditional mean displacement at every midpoint forces the two distributions to be equal—which this paper imports from an earlier preprint and does not prove here; if that theorem fails for some distributions, the null centering and consistency claims would fail with it.","fun_headline_variants_meta":{"raw":{"variants":["Zero-flow two-sample test: learn witness, then test","Split-witness test for distributional shifts","Zero-flow discrepancy detects structured differences","Midpoint-flow field powers valid two-sample test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000436,"raw_usage":{"total_tokens":2022,"prompt_tokens":678,"completion_tokens":1344,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":422,"completion_tokens_details":{"reasoning_tokens":1296}},"tokens_in":422,"tokens_out":1344,"duration_ms":10328,"temperature":1.0,"reasoning_tokens":1296,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:06:46.015319+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search for two distinct distributions P and Q with finite second moments for which E[Y−X | (X+Y)/2 = m] = 0 holds for every midpoint m. A single such pair would make the zero-flow discrepancy zero despite P≠Q, breaking the identification that the test depends on; the search could begin with atomic distributions or distributions whose characteristic functions vanish on a set.","supporting_citations":[],"review_version":1}