Pith. sign in

REVIEW 3 cited by

Refereeing the Referees: Evaluating Two-Sample Tests for Validating Generators in Precision Sciences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16336 v1 pith:S4UFDMSE submitted 2024-09-24 stat.ML cs.LGhep-phstat.AP

classification stat.MLcs.LGhep-phstat.AP
keywords testsmetricscomputationaldatasetdistanceevaluateevaluatinggaussians
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We propose a robust methodology to evaluate the performance and computational efficiency of non-parametric two-sample tests, specifically designed for high-dimensional generative models in scientific applications such as in particle physics. The study focuses on tests built from univariate integral probability measures: the sliced Wasserstein distance and the mean of the Kolmogorov-Smirnov statistics, already discussed in the literature, and the novel sliced Kolmogorov-Smirnov statistic. These metrics can be evaluated in parallel, allowing for fast and reliable estimates of their distribution under the null hypothesis. We also compare these metrics with the recently proposed unbiased Fr\'echet Gaussian Distance and the unbiased quadratic Maximum Mean Discrepancy, computed with a quartic polynomial kernel. We evaluate the proposed tests on various distributions, focusing on their sensitivity to deformations parameterized by a single parameter $\epsilon$. Our experiments include correlated Gaussians and mixtures of Gaussians in 5, 20, and 100 dimensions, and a particle physics dataset of gluon jets from the JetNet dataset, considering both jet- and particle-level features. Our results demonstrate that one-dimensional-based tests provide a level of sensitivity comparable to other multivariate metrics, but with significantly lower computational cost, making them ideal for evaluating generative models in high-dimensional settings. This methodology offers an efficient, standardized tool for model comparison and can serve as a benchmark for more advanced tests, including machine-learning-based approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Standard Model structure from LHC data with Riemannian flow matching

    hep-ph 2026-07 conditional novelty 7.0 of 10

    ShellFlow, a Riemannian flow-matching transformer fed only on-shell and invariant-mass priors and ~8×10^8 recorded ATLAS events, reproduces the SM's dilepton resonances, Weinberg angle, and top/W mass peaks in a singl...

  2. MCBench: A Benchmark Suite for Monte Carlo Sampling Algorithms

    stat.CO 2025-01 conditional novelty 6.0 of 10

    MCBench is a modular Julia benchmark suite that scores Monte Carlo samplers by comparing their samples against IID reference samples using metrics like sliced Wasserstein distance and maximum mean discrepancy.

  3. The fundamental limit of jet tagging: Beyond top jets

    hep-ph 2026-07 conditional novelty 4.0 of 10

    Using generative-model likelihood ratios, the authors estimate that modern taggers nearly reach the model-defined optimal limit for W, Z, and H-to-gg jets, while the top-jet gap remains large.

Pith tools