Pith. sign in

REVIEW 3 major objections 4 minor 5 cited by

A Neural-Network Extraction of Unpolarised Transverse-Momentum-Dependent Distributions

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Neural networks outperform traditional parametrisations in extracting unpolarised TMD distributions from Drell-Yan data.

desk verdict First NN-based TMD extraction, with a real chi2 improvement over MAP22, but the headline 'outperform' claim needs a model-comparison test before it fully lands. read the letter →

arxiv 2502.04166 v1 pith:DT2WNUP7 submitted 2025-02-06 hep-ph hep-ex

classification hep-phhep-ex
keywords transverse-momentum-dependentdistributionsneuralnetworksDrell-YannonperturbativeQCDN3LLresummationpartonmachinelearningunpolarisedTMD
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper presents the first extraction of unpolarised transverse-momentum-dependent (TMD) quark distributions where the nonperturbative part is parametrised by a neural network instead of a fixed functional form. Fitting the same Drell-Yan data at N3LL accuracy, the network achieves a total reduced chi-squared of 0.97 versus 1.28 for the previously used 12-parameter form, with the ATLAS subset improving from 3.51 to 1.38. The authors interpret this as evidence that neural networks capture information in the data that rigid parametrisations miss. They frame the result as a proof of concept for machine-learning-based TMD extraction.

What carries the argument

The engine is the factorised nonperturbative model f_NP(x, b_T; ζ) = NN(x, b_T)/NN(x, 0) · exp(−g_2 $b_T^{2}$ log(ζ/$Q_0^{2}$)/2). Dividing by NN(x,0) enforces the perturbative limit f_NP → 1 as b_T → 0; the exponential carries the nonperturbative part of the rapidity evolution; and the neural network itself, with two inputs (x and b_T), ten hidden nodes, and the smooth activation σ(z) = (1 + z/(1+|z|))/2, supplies a flexible shape for the intrinsic transverse momentum. The 42 total parameters are fitted with analytic gradients, Monte-Carlo replicas propagate experimental uncertainties, and a cross-validation split determines the stopping point to curb overfitting.

What would settle it

Re-run the fit on a random half of the data, freeze the parameters, and compute the chi-square on the other half; if the neural network does not beat the 12-parameter form on that held-out half, the superiority claim fails.

Watch

Extended reading notes

Core claim

The central claim is that the nonperturbative function f_NP entering the TMD cross section can be represented by a small neural network (41 weights plus one evolution parameter) and that this representation fits the Drell-Yan data better than the previous 12-parameter exponential/Gaussian form. The improvement is not only in the global chi-squared: the network reduces the correlated-shift contribution, produces smaller cross-section uncertainty bands, and brings the most precise data set (ATLAS) from a poor 3.51 to a good 1.38 reduced chi-squared. The authors conclude that current data encode complexity beyond the reach of traditional functional forms, and that neural-network TMD extraction is feasible at N3LL accuracy.

Load-bearing premise

The central assumption is that the network's lower chi-square reflects genuine modelling power rather than the advantage of having 30 extra free parameters, since the paper reports no held-out evaluation or model-comparison statistic beyond the training/validation split.

Editorial extensions

If this is right

  • TMD extractions no longer need to commit to a fixed functional shape for the intrinsic transverse momentum distribution.
  • The same architecture can be extended to flavour-dependent fits and to simultaneous fits of TMD PDFs and TMD fragmentation functions, which the authors identify as the next step.
  • The dramatic improvement on ATLAS data suggests the network resolves shape features that more rigid parametrisations flatten out, particularly at moderate q_T.
  • Smaller correlated shifts in the NN fit imply that systematic uncertainties are absorbed by the flexible model rather than by large pulls on data points.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the chi-squared improvement survives a proper model-comparison statistic (e.g., a held-out evaluation or an information criterion), published functional-form extractions would likely need to be redone with neural networks; the paper stops short of providing such a statistic.
  • The NN's roughly stable uncertainty band out to |k⊥| ≈ 0.6 GeV hints that the data constrain the shape of the intrinsic transverse momentum, not just its width — a testable prediction as more precise low-q_T data arrive.
  • The same machinery could be adapted to polarised TMDs (Sivers, Boer–Mulders), where the functional bias of current parametrisations is even more pronounced.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents a proof-of-concept extraction of unpolarised quark TMD PDFs from Drell-Yan data at N3LL accuracy, using a neural network to parametrise the nonperturbative function fNP in Eq. (7). The NN architecture [2,10,1] has 42 free parameters and the fit is compared with the 12-parameter MAP22 functional form on the same data set. Table I reports a total reduced chi2 of 0.97 for the NN fit versus 1.28 for MAP22, with the ATLAS subset improving from 3.51 to 1.38. The authors conclude that the NN parametrisation outperforms traditional parametrisations and that NNs can better capture the information in the data. The analysis uses Monte Carlo replicas for uncertainty propagation, analytic gradients for the minimisation, and a 50/50 validation split to choose the stopping point of the training.

Significance. If the superiority claim is established, this would be a valuable methodological proof of concept, showing that NN parametrisations can reduce the bias of traditional functional forms in TMD extractions and can scale to future multi-dimensional fits. The paper has clear strengths: the N3LL framework and data set are standard, the NN and MAP22 fits are compared under identical perturbative and data conditions, the gradient computation is analytic, and the MAP collaboration's public tools support reproducibility. The principal weakness is that the headline comparison is in-sample and not controlled for model complexity, so the central claim that NNs 'outperform' traditional parametrisations is not yet supported.

major comments (3)
  1. [Results/Table I; cross-validation paragraph after Eq. (7)] The central claim that the NN parametrisation outperforms traditional parametrisations is not supported by the comparison as presented. Table I reports total reduced chi2 of 0.97 for the NN (42 free parameters) versus 1.28 for MAP22 (12 parameters) evaluated on the full data set, while the 50/50 validation split is used only to choose the stopping point of the minimisation. A lower in-sample chi2 is expected for the more flexible model even if the extra flexibility mostly fits noise, and no held-out chi2, information criterion, or cross-validated prediction error is provided. Quantitatively, the total chi2 improvement of about 150 units is accompanied by 30 additional parameters; with 482 data points a standard BIC correction would add roughly 185 units, so the ranking could reverse. Please provide a model-comparison statistic or a held-out evaluation, and report the distribution of chi2 over the 250 MC replicas.
  2. [Results, data-selection paragraph] The PHENIX data of Ref. [46] are excluded because only two points survive the q_T/Q cut and their description is stated to be 'typically poor for any parametrisation of fNP'. Since the comparison with MAP22 is the sole basis for the headline conclusion, the paper should state whether the exclusion was decided before or after seeing the NN fit, and should show a sensitivity test in which the two PHENIX points are included or the cut is varied. As written, the data selection is justified by fit quality, which could alter the relative ranking of the two parametrisations and therefore needs explicit robustness checks.
  3. [Formalism and parametrisation, Eq. (7)] The paper states that several NN parametrisations were explored and that the [2,10,1] architecture was chosen for this proof of concept, but the exploration is not described and no comparison among architectures is shown. If the reported architecture was selected using the same data, the quoted chi2 is an optimistic estimate of predictive performance. At minimum, please report the architecture search protocol, the set of architectures considered, the selection criterion, and whether the MAP22 comparison was made after the architecture selection.
minor comments (4)
  1. [Fig. 1] The ATLAS panel axis label contains the rendering artifact 'GeV□1'; this should read 'GeV^-1'.
  2. [Section II heading] The heading 'F ormalism' has a spurious space and should be corrected to 'Formalism'.
  3. [References] Ref. [47] is cited as a website URL; please provide a versioned citation or an arXiv identifier so that the MAP22 fit used for the comparison can be reproduced exactly.
  4. [Cross-validation paragraph after Eq. (7)] For reproducibility, please specify how the validation/training split is constructed, including whether the split is at the level of data points or experiments, and how the random seed is chosen.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the NN-based fNP is fitted to external Drell-Yan data and compared with MAP22 on the same data.

full rationale

The derivation chain is not circular. The parameters of fNP in Eq. (7) are free parameters fitted by minimising a chi^2 against external Drell-Yan cross-section data through Eqs. (1)-(6); the extracted TMD PDFs are then compared with the same data and with a competing parametrisation (MAP22) evaluated on the same footing. No fitted parameter is relabelled as a prediction, and no result is assumed in its own input: the experimental data are external, and the perturbative ingredients (MSHT2020 PDFs, N3LL coefficients, the b* prescription) are standard inputs from previous literature, not consequences of this fit. The comparison chi^2 values in Table I are in-sample fit qualities, so the 'outperform' statement would be strengthened by a held-out or information-criterion test; that is a statistical-comparison caveat, not a circularity. Self-citations to earlier MAP papers provide the benchmark functional form and the b* regulator but are not load-bearing: the benchmark is recomputed here and the regulator is a conventional choice. Therefore no self-definitional, fitted-input-as-prediction, or self-citation-reduction step is present.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No new physical entities are introduced. The analysis rests on standard TMD factorization, an external PDF set, and a chosen 42-parameter functional form; the main epistemic costs are the assumed separability and flavour independence of fNP, the b* regulator, and reliance on external perturbative inputs.

free parameters (3)
  • NN parameters (41 weights and biases) = not reported in paper
    Weights and biases of the [2,10,1] network in Eq. (7); fitted to Drell-Yan data to shape the intrinsic nonperturbative part fNP.
  • g2 = not reported in paper
    Coefficient of the b_T^2 log(zeta/Q0^2) term in Eq. (7); fitted to data and controls the rapidity-evolution component of fNP.
  • Q0 = 1 GeV
    Reference scale for the rapidity-evolution logarithm in Eq. (7); chosen by hand, not fitted. It changes the split between the NN intrinsic part and the evolution exponential, though the same choice is used in both compared fits.
assumptions (5)
  • domain assumption TMD factorization for Drell-Yan at low transverse momentum (Eq. 1)
    The cross section is written as a convolution of TMD PDFs with a perturbative hard factor under the condition q_T << Q, enforced by the qT/Q < 0.2 cut.
  • domain assumption b* prescription in Eqs. (4)-(5) with b_max = 2 e^-gamma_E GeV^-1
    Needed to avoid the Landau pole at large b_T; the associated power corrections are absorbed into fNP, so the extraction depends on this regulator.
  • ad hoc to paper Separable form of fNP in Eq. (7), NN(x,b_T)/NN(x,0) times an evolution exponential
    Assumes that intrinsic transverse-momentum effects and rapidity-evolution nonperturbative effects factor cleanly; this structure is chosen by the authors and is not derived.
  • ad hoc to paper Flavour independence of the intrinsic transverse momentum
    The same NN and g2 are used for all active quark flavours; the paper states this flavour dependence is neglected.
  • domain assumption MSHT2020 collinear PDFs and N3LL perturbative ingredients are reliable inputs
    The matching coefficient C and evolution are computed from external N3LL inputs and the MSHT2020 PDF set; any bias in these inputs propagates into the extracted fNP.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Neural-Network Extraction of Unpolarised Transverse-Momentum-Dependent Distributions." pith.science (2026). https://pith.science/paper/DT2WNUP7

@misc{pith2026250204166,
  author       = {Pith},
  title        = {Pith review of: A Neural-Network Extraction of Unpolarised Transverse-Momentum-Dependent Distributions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DT2WNUP7}},
  note         = {Machine review of arXiv:2502.04166}
}
read the original abstract

We present the first extraction of transverse-momentum-dependent distributions of unpolarised quarks from experimental Drell-Yan data using neural networks to parametrise their nonperturbative part. We show that neural networks outperform traditional parametrisations providing a more accurate description of data. This work establishes the feasibility of using neural networks to explore the multi-dimensional partonic structure of hadrons and paves the way for more accurate determinations based on machine-learning techniques.

Figures

Figures reproduced from arXiv: 2502.04166 by the authors.

Figure 2
Figure 2. FIG. 2: The unpolarised TMD PDF of the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 1
Figure 1. FIG. 1: Comparison between experimental data (black dots) [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. First constraints on the nonperturbative gluon Collins-Soper kernel

    hep-lat 2026-07 conditional novelty 7.0 of 10

    First lattice-QCD constraints on the nonperturbative gluon Collins-Soper kernel are obtained at near-physical pion mass with uNNLL LaMET matching on a single a=0.15 fm ensemble.

  2. Soft background fields at next-to-leading power in transverse momentum dependent SIDIS with jets

    hep-ph 2025-07 conditional novelty 7.0 of 10

    New next-to-leading-power factorization for SIDIS with jets from a background-field method with explicit soft modes, including operator-level definitions of twist-3 TMDs free of rapidity and endpoint divergences.

  3. Symbolic Extraction of Non-Perturbative Transverse-Momentum-Dependent Distributions from Drell-Yan Data

    hep-ph 2026-07 conditional novelty 6.0 of 10

    Symbolic regression on a factorized NN fit to 482 Drell–Yan points yields a 9-constant analytical non-perturbative TMD with χ²/ndf≈1.04 and a retained x–b_T cross term.

  4. New insights from the flavor dependence of quark transverse momentum distributions in the pion

    hep-ph 2025-09 conditional novelty 6.0 of 10

    First flavor-dependent extraction of unpolarized quark TMDs in the pion, finding a wider transverse momentum tail for valence d quarks than for sea quarks, with large uncertainties.

  5. Parton Distribution Functions and their Generalizations

    hep-ph 2025-07 unverdicted novelty 1.0 of 10

    A textbook-style introduction to the definitions, QCD properties, experimental access, and current phenomenological status of PDFs and their generalizations.

Reference graph

Works this paper leans on

47 extracted references · 43 canonical work pages · cited by 5 Pith papers

  1. [46]

    Aidala et al

    C. Aidala et al. (PHENIX), Phys. Rev. D 99, 072003 (2019)

  2. [1]

    Bacchetta, F

    A. Bacchetta, F. Delcarro, C. Pisano, M. Radici, and A. Signori, JHEP 06, 081 (2017), [Erratum: JHEP 06, 051 (2019)]

  3. [2]

    Scimemi and A

    I. Scimemi and A. Vladimirov, Eur. Phys. J. C 78, 89 (2018)

  4. [3]

    Bertone, I

    V. Bertone, I. Scimemi, and A. Vladimirov, JHEP 06, 028 (2019)

  5. [4]

    Scimemi and A

    I. Scimemi and A. Vladimirov, JHEP 06, 137 (2020)

  6. [5]

    Bacchetta, V

    A. Bacchetta, V. Bertone, C. Bissolotti, G. Bozzi, F. Del- carro, F. Piacenza, and M. Radici, JHEP 07, 117 (2020)

  7. [6]

    M. Bury, F. Hautmann, S. Leal-Gomez, I. Scimemi, A. Vladimirov, and P. Zurita, JHEP 10, 118 (2022)

  8. [7]

    Bacchetta, V

    A. Bacchetta, V. Bertone, C. Bissolotti, G. Bozzi, M. Cerutti, F. Piacenza, M. Radici, and A. Signori (MAP (Multi-dimensional Analyses of Partonic distributions)), JHEP 10, 127 (2022)

Show all 47 references
  1. [8]

    V. Moos, I. Scimemi, A. Vladimirov, and P. Zurita, JHEP 05, 036 (2024)

  2. [9]

    Bacchetta, V

    A. Bacchetta, V. Bertone, C. Bissolotti, G. Bozzi, M. Cerutti, F. Delcarro, M. Radici, L. Rossi, and A. Sig- nori (MAP), JHEP 08, 232 (2024)

  3. [10]

    ˇCui´ c, K

    M. ˇCui´ c, K. Kumeriˇ cki, and A. Sch¨ afer, Phys. Rev. Lett. 125, 232005 (2020)

  4. [11]

    Abdul Khalek, V

    R. Abdul Khalek, V. Bertone, A. Khoudli, and E. R. Nocera (MAP (Multi-dimensional Analyses of Partonic distributions)), Phys. Lett. B 834, 137456 (2022)

  5. [12]

    I. P. Fernando and D. Keller, Phys. Rev. D 108, 054007 6 (2023)

  6. [13]

    Bertone, A

    V. Bertone, A. Chiefa, and E. R. Nocera (MAP) (2024), arXiv:2404.04712 [hep-ph]

  7. [14]

    R. D. Ball et al. (NNPDF), Eur. Phys. J. C 84, 659 (2024)

  8. [15]

    Collins, Foundations of Perturbative QCD , vol

    J. Collins, Foundations of Perturbative QCD , vol. 32 of Cambridge Monographs on Particle Physics, Nuclear Physics and Cosmology (Cambridge University Press, 2023), ISBN 978-1-009-40184-5, 978-1-009-40183-8, 978- 1-009-40182-1

  9. [16]

    Cerutti, L

    M. Cerutti, L. Rossi, S. Venturini, A. Bacchetta, V. Bertone, C. Bissolotti, and M. Radici (MAP (Multi- dimensional Analyses of Partonic distributions)), Phys. Rev. D 107, 014014 (2023)

  10. [17]

    Catani, M

    S. Catani, M. L. Mangano, P. Nason, and L. Trentadue, Nucl. Phys. B 478, 273 (1996)

  11. [18]

    Kulesza, G

    A. Kulesza, G. F. Sterman, and W. Vogelsang, Phys. Rev. D 66, 014011 (2002)

  12. [19]

    Laenen, G

    E. Laenen, G. F. Sterman, and W. Vogelsang, Phys. Rev. Lett. 84, 4296 (2000)

  13. [20]

    Kulesza, G

    A. Kulesza, G. F. Sterman, and W. Vogelsang, Phys. Rev. D 69, 014012 (2004)

  14. [21]

    Agarwal, K

    S. Agarwal, K. Mierle, and T. C. S. Team, Ceres Solver (2023), https://github.com/ceres-solver/ceres- solver

  15. [22]

    Abdul Khalek and V

    R. Abdul Khalek and V. Bertone (2020), arXiv:2005.07039 [physics.comp-ph]

  16. [23]

    R. D. Ball et al. (NNPDF), Eur. Phys. J. C 77, 663 (2017)

  17. [24]

    R. D. Ball et al. (NNPDF), Eur. Phys. J. C 82, 428 (2022)

  18. [25]

    Del Debbio, S

    L. Del Debbio, S. Forte, J. I. Latorre, A. Piccione, and J. Rojo (NNPDF), JHEP 03, 039 (2007)

  19. [26]

    Moreno et al., Phys

    G. Moreno et al., Phys. Rev. D 43, 2815 (1991)

  20. [27]

    A. S. Ito et al., Phys. Rev. D 23, 604 (1981)

  21. [28]

    P. L. McGaughey et al. (E772), Phys. Rev. D 50, 3038 (1994), [Erratum: Phys.Rev.D 60, 119903 (1999)]

  22. [29]

    Affolder et al

    T. Affolder et al. (CDF), Phys. Rev. Lett. 84, 845 (2000)

  23. [30]

    Aaltonen et al

    T. Aaltonen et al. (CDF), Phys. Rev. D 86, 052010 (2012)

  24. [31]

    Abbott et al

    B. Abbott et al. (D0), Phys. Rev. D 61, 032004 (2000)

  25. [32]

    V. M. Abazov et al. (D0), Phys. Rev. Lett. 100, 102002 (2008)

  26. [33]

    V. M. Abazov et al. (D0), Phys. Lett. B 693, 522 (2010)

  27. [34]

    Collaboration, Phys

    S. Collaboration, Phys. Lett. B 854, 138715 (2024)

  28. [35]

    Aaij et al

    R. Aaij et al. (LHCb), JHEP 08, 039 (2015)

  29. [36]

    Aaij et al

    R. Aaij et al. (LHCb), JHEP 01, 155 (2016)

  30. [37]

    Aaij et al

    R. Aaij et al. (LHCb), JHEP 09, 136 (2016)

  31. [38]

    Chatrchyan et al

    S. Chatrchyan et al. (CMS), Phys. Rev. D 85, 032002 (2012)

  32. [39]

    Khachatryan et al

    V. Khachatryan et al. (CMS), JHEP 02, 096 (2017)

  33. [40]

    A. M. Sirunyan et al. (CMS), JHEP 12, 061 (2019)

  34. [41]

    Aad et al

    G. Aad et al. (ATLAS), JHEP 09, 145 (2014)

  35. [42]

    Aad et al

    G. Aad et al. (ATLAS), Eur. Phys. J. C 76, 291 (2016)

  36. [43]

    Aad et al

    G. Aad et al. (ATLAS), Eur. Phys. J. C 80, 616 (2020)

  37. [44]

    Bailey, T

    S. Bailey, T. Cridge, L. A. Harland-Lang, A. D. Martin, and R. S. Thorne, Eur. Phys. J. C 81, 341 (2021)

  38. [45]

    Buckley, J

    A. Buckley, J. Ferrando, S. Lloyd, K. Nordstr¨ om, B. Page, M. R¨ ufenacht, M. Sch¨ onherr, and G. Watt, Eur. Phys. J. C 75, 132 (2015)

  39. [47]

    TMD fits, https://mapcollaboration.github.io

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.