Pith. sign in

REVIEW 3 major objections 5 minor 105 references

Transformer networks for Heavy flavor jet tagging

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A compact cross-attention mixer matches transformer-level top tagging at roughly one-twentieth the training time.

desk verdict Useful review of transformer-based jet taggers, but the headline '20x faster' speedup over ParT is not backed by controlled benchmarking and the manuscript has draft placeholders. read the letter →

arxiv 2411.11519 v1 pith:AMTTLVT7 submitted 2024-11-18 hep-ph

classification hep-ph
keywords jettaggingheavyflavortransformernetworkscross-attentionMLP-mixersubjetsparticlecloudstop
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the accuracy of large attention-based transformers for heavy-flavor jet tagging can be reached by a much smaller network once physics structure is built into the architecture. The authors review image-, graph-, and particle-cloud representations of jets, then present the Cross-Attention-Mixer (CA-Mixer), which processes all jet constituents with two small MLP layers and adds a cross-attention block whose keys and values come from subjets reconstructed at a smaller clustering radius. On the standard boosted-top versus QCD benchmark, CA-Mixer reaches AUC 0.9859 with 86K parameters, matching the 2.14M-parameter Particle Transformer at 0.9858 and ParticleNet, while training in roughly 33 seconds per epoch versus 612 seconds. The claim is that injecting the jet's prong structure through subjet cross-attention substitutes for much of the learned capacity of a transformer, pointing toward cheaper and more interpretable taggers for LHC-style analyses.

What carries the argument

The central object is the Cross-Attention-Mixer (CA-Mixer) network, a permutation-invariant architecture in which an MLP-mixer processes the unordered set of jet constituents while a cross-attention block reads subjet information. Each mixer layer applies two shared MLPs, one mixing particle tokens and one mixing feature channels, preserving permutation invariance; the same constituents are reclustered into subjets with a jet algorithm at a smaller radius, and multi-head cross-attention uses those subjets as keys and values while the mixed constituent features act as queries. A skip connection preserves the original particle coordinates, and global max pooling feeds a final classifier. The mechanism's job is to inject the physically known prong structure of boosted heavy particles into a network that is too small to rediscover it from data alone.

What would settle it

Run Particle Transformer and CA-Mixer on the same GPU with the same framework, batch size, and number of epochs on the boosted-top dataset and compare wall-clock training time; if the per-epoch ratio is far from 20, the headline speed claim fails.

Watch

Extended reading notes

Core claim

The paper's central discovery is that cross-attention between whole-jet and subjet-level inputs lets a compact permutation-invariant network reproduce transformer-level top-tagging performance. The CA-Mixer replaces the heavy self-attention stack of a particle transformer with an MLP-mixer that mixes particle tokens and feature channels, then applies multi-head attention in which the jet constituents act as queries and reclustered subjets provide the keys and values. This encodes the factorized structure of QCD parton showers, correlating resolved prongs with the hadron distribution around them, so the network does not need millions of parameters to learn prong structure from scratch. The reported result on the boosted-top dataset is an AUC of 0.9859 and a background rejection of 416 ± 5 at 50% signal efficiency, with 86.03K parameters and 33 seconds per epoch, against Particle Transformer's 0.9858, 413 ± 16, 2.14M parameters, and 612 seconds.

Load-bearing premise

The speed comparison assumes that the published per-epoch training times for the other networks were obtained under comparable hardware, framework, and optimization conditions; the paper specifies the GPU only for its own CA-Mixer runs.

Editorial extensions

If this is right

  • Top-tagging accuracy at the level of a 2.14M-parameter particle transformer can be obtained with an 86K-parameter network when subjet cross-attention is added, so the large self-attention stack is not strictly necessary for this task.
  • Per-epoch training time drops by about a factor of 20 on the stated GPU (33 seconds versus 612 seconds), which makes iterative retraining, hyperparameter scans, and model development much cheaper.
  • The CKA representation-similarity analysis shows the cross-attention layer carries information distinct from the mixer MLPs, supporting the view that subjets contribute complementary physics structure rather than extra capacity.
  • The same cross-attention idea extends beyond single jets to event-level classification, as already demonstrated for two-Higgs final states, so the mechanism is not limited to top tagging.
  • The CA-Mixer keeps permutation invariance by construction, so it inherits the main advantage of particle-cloud representations without the computational cost of pairwise self-attention over all constituents.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the training-time comparison in Table I were repeated under strictly identical hardware, software, and optimization budgets, the factor-of-20 advantage would probably shrink but would not vanish, because the parameter-count difference is intrinsic to the designs; a controlled benchmark is the clean test.
  • Subjet cross-attention should transfer most strongly to other multi-pronged signatures such as boosted W, Z, and Higgs tagging, and the architecture's gain over a plain mixer should be largest for jets with well-separated prongs.
  • A promising next variant is combining the subjet cross-attention with Lorentz-equivariant aggregators of the PELICAN type; the paper itself notes Lorentz equivariance as a possible further improvement, and that combination could be checked for whether it closes the small AUC gap to PELICAN.
  • The interpretability analysis yields a testable diagnostic: ablating the cross-attention block should hurt multi-pronged top jets more than QCD jets, which would confirm that prong structure, not raw capacity, carries the performance.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper is a review of machine-learning methods for heavy-flavor (boosted top, W/Z/H) jet tagging, with a focus on transformer architectures. It summarizes data representations (jet images, graphs, particle clouds), describes the attention mechanism and its particle-transformer extension, and discusses architectures that incorporate physics structure: PELICAN (Lorentz invariance), LundNet (IRC-safe clustering), and the authors' CA-Mixer, which uses cross-attention between jet constituents and reconstructed subjets. Sec. IV.B compares top-tagging performance on the standard boosted-top dataset, reporting that CA-Mixer achieves AUC 0.9859 with 86K parameters and 33 s/epoch, matching ParticleNET and ParT (AUC 0.9858, 2.14M parameters, 612 s/epoch) while training 'approximately 20 times faster.' Sec. V reviews interpretability tools (CKA, attention maps, Grad-CAM) and includes a CKA figure for CA-Mixer taken from Ref. [24].

Significance. If the speed comparison were properly supported, the paper would make a useful contribution: it calls attention to a compact 86K-parameter architecture whose AUC is statistically indistinguishable from much larger transformer taggers on the standard top-tagging benchmark, and it frames this in the context of physics-motivated network design. The review portions are generally accurate: the descriptions of jet-image, graph, and particle-cloud representations are consistent with the cited literature, and the attention equations in Sec. II are standard. The paper also gives credit to the original sources and makes a concrete, falsifiable performance table rather than stopping at qualitative claims. However, the paper provides no independent verification of the CA-Mixer numbers (they are quoted from the authors' own Ref. [24]), no convergence curves or total training times, and no normalized timing benchmark for the '20 times faster' claim. The paper also contains a large undigested draft block, which suggests the manuscript is not in final form.

major comments (3)
  1. [Sec. IV.B, Table I] The central speedup claim ('trains approximately 20 times faster') is not supported as stated. The table caption says results for JEDI-net, PFN, PCT, LorentzNet, ParticleNET, PELICAN, and ParT are 'quoted from their published results,' while only CA-Mixer's 33 s/epoch is clearly attributed to the authors' measurement on an NVIDIA RTX A6000. Nothing in the caption or text demonstrates that ParT's 612 s/epoch was obtained on the same GPU, with the same software framework, mixed-precision settings, data-loading pipeline, and batch size; the caption specifies batch size 1024 but does not state whether the quoted ParT run used it. In addition, per-epoch time is not the same as training cost to a fixed target AUC unless the two models require the same number of epochs, and no convergence curves or total training times are given. Please either rerun ParT (and ideally ParticleNET) under identical local conditions and report epochs to the quoted AUC, or remove the '20 times faster' claim and limit Sec. IV.B to the AUC/parameter comparison, which is what the quoted evidence supports.
  2. [Sec. IV.A (after the LundNet paragraph)] The manuscript contains a large undigested draft block beginning 'REPEAT & Global MAX pooling ...' and continuing through a duplicate title/author block dated March 18, 2024, a list of 'Needed plots,' and a second 'FIG. 3' caption. This is not a presentation typo: it indicates that the LaTeX source contains leftover material from an earlier draft. The block must be completely removed and the surrounding text and figure numbering re-checked so that the published version contains exactly one title block, one abstract, and one coherent set of figures.
  3. [Sec. IV.A (CA-Mixer description)] The description of CA-Mixer is too incomplete for a paper whose title and Sec. IV.B highlight this network. The text refers to Ref. [24] for details but does not specify how subjets are formed (which reclustering algorithm, which radius Rcut, how many subjets are kept), how the cross-attention queries, keys, and values are constructed from the constituent and subjet sets (Eqs. (2)-(6) define generic attention, not the specific cross-attention shown in Fig. 3), or the mixer-layer dimensions and hyperparameters behind the quoted 86K parameter count. Without this information, the reported AUC and the parameter comparison in Table I cannot be checked. Please add a short but complete architectural specification, or state explicitly that all CA-Mixer numbers are taken from Ref. [24] and not reproduced independently here.
minor comments (5)
  1. [Throughout] There are numerous typos and misspellings that should be corrected: 'Coliider' for 'Collider' in Sec. I, 'psudo-rapidity' for 'pseudo-rapidity' and 'marge' for 'merge' in Sec. IV.A, and 'pretaining' for 'pretraining' in the caption of Table I.
  2. [Sec. V.B] The attention-map discussion says each element of alpha_ij in Eq. (3) is the attention from particle i to particle j, but in the CA-Mixer model the attention is cross-attention between constituents and subjets. The text should generalize the description or clarify that Eq. (3) describes the self-attention case.
  3. [Table I] The table would benefit from explicit definitions: 'Rej 50%' should state that it is the background rejection at 50% signal efficiency (or the inverse), and the entries with '--' should be explained as not reported in the original publications rather than left ambiguous.
  4. [Sec. V] The text says 'we explore different methods for interpreting network decision-making' and 'we will apply some of the methods,' but the section only reproduces one CKA figure from Ref. [24] and gives no new Grad-CAM or attention-map results for CA-Mixer. The wording should be adjusted to reflect the review nature of the contribution.
  5. [Abstract and Sec. I] The abstract and introduction describe the paper as a review, yet Sec. IV.B contains an original comparative speed claim. Please add a sentence clarifying that the CA-Mixer numbers and interpretability figures are taken from the authors' prior work (Ref. [24]) and that the new element in this paper is the synthesis and comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a review whose quantitative claims are cited from prior published work, including the authors' own Ref. [24], and are not derived from or defined in terms of the present manuscript.

full rationale

The manuscript is an explicit review of machine-learning jet-tagging methods. It does not fit parameters and then rename them as predictions, and no equation in the paper is constructed so that an output equals an input by definition. The CA-Mixer architecture, its AUC, parameter count, and per-epoch training time are introduced with the phrase 'In Ref. [24], we have proposed...' and Table I is captioned 'Taken from Ref. [24]', with all other network results 'quoted from their published results'. This is self-citation, but it is not load-bearing circularity: Ref. [24] is a separate, externally available publication with its own stated assumptions and measurements, and the comparison rows against ParticleNET, ParT, PELICAN, and other networks are independent external benchmarks. The paper makes no uniqueness claim and invokes no theorem from the same authors to forbid alternative architectures. The physics motivation via QCD factorization is a design rationale, not a derived prediction. The only substantive concern is benchmark comparability of per-epoch timings quoted from different papers, which is a correctness and normalization issue, not a circularity. Accordingly, no circular step meets the evidentiary bar.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This review presents no new derivation or model; its claims rest on the validity of the cited benchmark dataset, the comparability of published performance numbers, and the completeness of the CA-Mixer description in Ref. [24].

assumptions (3)
  • domain assumption The benchmark top tagging dataset from Ref. [6] (Pythia8 plus Delphes, pT in [550,650] GeV, |eta| < 2) provides a valid standardized testbed for comparing jet taggers.
    The paper bases its performance comparisons entirely on this dataset and on published results from different papers, which the review treats as directly comparable (Sec. IV.B).
  • domain assumption Results from different papers (JEDI-net, PFN, PCT, LorentzNet, ParticleNET, PELICAN, ParT, CA-Mixer) are comparable despite differing implementations, training procedures, and measurement environments.
    Table I merges published numbers and training times; the paper does not control for hardware, framework, or hyperparameter tuning differences.
  • domain assumption The CA-Mixer architecture, described only schematically here, is exactly as specified in Ref. [24].
    The review relies on Ref. [24] for all network details and performance numbers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Transformer networks for Heavy flavor jet tagging." pith.science (2026). https://pith.science/paper/AMTTLVT7

@misc{pith2026241111519,
  author       = {Pith},
  title        = {Pith review of: Transformer networks for Heavy flavor jet tagging},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AMTTLVT7}},
  note         = {Machine review of arXiv:2411.11519}
}
read the original abstract

In this article, we review recent machine learning methods used in challenging particle identification of heavy-boosted particles at high-energy colliders. Our primary focus is on attention-based Transformer networks. We report the performance of state-of-the-art deep learning networks and further improvement coming from the modification of networks based on physics insights. Additionally, we discuss interpretable methods to understand network decision-making, which are crucial when employing highly complex and deep networks.

Figures

Figures reproduced from arXiv: 2411.11519 by the authors.

Figure 2
Figure 2. FIG. 2. Image for top jet (left) and QCD jet (right) after the [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 1
Figure 1. FIG. 1. schematical figure of top quark decay into b and W [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 3
Figure 3. FIG. 3. A Schematical figure of the Cross Atention-Mixer [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: FIG. 4. Taken from Ref. [ [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

105 extracted references · 9 canonical work pages

  1. [24]

    J. M. Munoz, I. Batatia and C. Ortner, Boost invariant polynomials for efficient jet tagging , Mach. Learn. Sci. Tech. 3 (2022) 04LT05 [2207.08272]

  2. [1]

    J. M. Butterworth, A. R. Davison, M. Rubin and G. P. Salam, Jet substructure as a new Higgs search channel at the LHC , Phys. Rev. Lett. 100 (2008) 242001 [0802.2470]

  3. [2]

    minimum spanning tree + hierarchical dendogram (on how the hdbscan work)

  4. [3]

    four plots for subjets clustering for the top case

  5. [4]

    ROC for varying radius from 0.1 to 0.5 using CA

  6. [5]

    plot for the input subjets for the top and qcd jets

  7. [6]

    plot for the cross attention

  8. [7]

    Foundation of Machine Learning Physics

    ROCs for every thing. PACS numb ers: I. INTRODUCTION N x (Mixer layer) Input dataset: Particle constitution dim: (100 , 7) Global Max Pooling MLP Class Mixer layer Layer NormEmbedding T MLP-1 MLP-1 Shared T Layer Norm Cross-attention Sub-jets Particles Features Particles Features Particles Particles Features Particles Particles MLP-2 MLP-2 Shared Skip-con...

Show all 105 references
  1. [8]

    Abdesselam et al., Boosted Objects: A Probe of Beyond the Standard Model Physics , Eur

    A. Abdesselam et al., Boosted Objects: A Probe of Beyond the Standard Model Physics , Eur. Phys. J. C 71 (2011) 1661 [ 1012.5412]

  2. [9]

    Altheimer et al., Jet Substructure at the Tevatron and LHC: New Results, New Tools, New Benchmarks , J

    A. Altheimer et al., Jet Substructure at the Tevatron and LHC: New Results, New Tools, New Benchmarks , J. Phys. G 39 (2012) 063001 [ 1201.0008]

  3. [10]

    Altheimer et al., Boosted Objects and Jet Substructure at the LHC

    A. Altheimer et al., Boosted Objects and Jet Substructure at the LHC. Report of BOOST2012, held at IFIC Valencia, 23rd-27th of July 2012 , Eur. Phys. J. C 74 (2014) 2792 [ 1311.2708]

  4. [11]

    L. G. Almeida, M. Backovi´ c, M. Cliche, S. J. Lee and M. Perelstein, Playing Tag with ANN: Boosted Top Identification with Pattern Recognition, JHEP 07 (2015) 086 [ 1501.05968]

  5. [12]

    Butter, G

    A. Butter, G. Kasieczka, T. Plehn and M. Russell, Deep-learned Top Tagging with a Lorentz Layer , SciPost Phys. 5 (2018) 028 [ 1707.08966]

  6. [13]

    Kasieczka, T

    G. Kasieczka, T. Plehn, M. Russell and T. Schell, Deep-learning Top Taggers or The End of QCD? , JHEP 05 (2017) 006 [ 1701.08784]

  7. [14]

    Louppe, K

    G. Louppe, K. Cho, C. Becot and K. Cranmer, QCD-Aware Recursive Neural Networks for Jet Physics , JHEP 01 (2019) 057 [ 1702.00748]

  8. [15]

    Butter et al., The Machine Learning landscape of top taggers, SciPost Phys

    A. Butter et al., The Machine Learning landscape of top taggers, SciPost Phys. 7 (2019) 014 [ 1902.09914]

  9. [16]

    Chakraborty, S

    A. Chakraborty, S. H. Lim, M. M. Nojiri and M. Takeuchi, Neural Network-based Top Tagger with Two-Point Energy Correlations and Geometry of Soft Emissions, JHEP 07 (2020) 111 [ 2003.11787]

  10. [17]

    Bhattacharya, M

    S. Bhattacharya, M. Guchait and A. H. Vijay, Boosted top quark tagging and polarization measurement using machine learning, Phys. Rev. D 105 (2022) 042005 [2010.11778]

  11. [18]

    Ju and B

    X. Ju and B. Nachman, Supervised Jet Clustering with Graph Neural Networks for Lorentz Boosted Bosons , Phys. Rev. D 102 (2020) 075014 [ 2008.06064]

  12. [19]

    F. A. Dreyer and H. Qu, Jet tagging in the Lund plane with graph networks , JHEP 03 (2021) 052 [2012.08526]

  13. [20]

    Tannenwald, C

    B. Tannenwald, C. Neu, A. Li, G. Buehlmann, A. Cuddeback, L. Hatfield et al., Benchmarking Machine Learning Techniques with Di-Higgs Production at the LHC , 2009.06754

  14. [21]

    F. A. Dreyer, R. Grabarczyk and P. F. Monni, Leveraging universality of jet taggers through transfer learning, Eur. Phys. J. C 82 (2022) 564 [ 2203.06210]

  15. [22]

    Hammad, S

    A. Hammad, S. Khalil and S. Moretti, Search for mono-Higgs signals in bb ¯ final states using deep neural networks, Phys. Rev. D 107 (2023) 075027 [2208.10133]

  16. [23]

    Ahmed, A

    I. Ahmed, A. Zada, M. Waqas and M. U. Ashraf, Application of deep learning in top pair and single top quark production at the LHC , Eur. Phys. J. Plus 138 (2023) 795 [ 2203.12871]

  17. [25]

    He and D

    M. He and D. Wang, Quark/gluon discrimination and top tagging with dual attention transformer , Eur. Phys. J. C 83 (2023) 1116 [ 2307.04723]

  18. [26]

    J. A. Aguilar-Saavedra, E. Arganda, F. R. Joaquim, R. M. Sand´ a Seoane and J. F. Seabra,Gradient Boosting MUST taggers for highly-boosted jets , 2305.04957

  19. [27]

    Athanasakos, A

    D. Athanasakos, A. J. Larkoski, J. Mulligan, M. P losko´ n and F. Ringer, Is infrared-collinear safe information all you need for jet classification? , 2305.08979

  20. [28]

    Grossi, M

    M. Grossi, M. Incudini, M. Pellen and G. Pelliccioli, Amplitude-assisted tagging of longitudinally polarised bosons using wide neural networks , Eur. Phys. J. C 83 (2023) 759 [ 2306.07726]

  21. [29]

    Hammad, P

    A. Hammad, P. Ko, C.-T. Lu and M. Park, Exploring Exotic Decays of the Higgs Boson to Multi-Photons at the LHC via Multimodal Learning Approaches , 2405.18834. 10

  22. [30]

    Hammad and M

    A. Hammad and M. M. Nojiri, Streamlined jet tagging network assisted by jet prong structure , JHEP 06 (2024) 176 [2404.14677]

  23. [31]

    CMS collaboration, A. M. Sirunyan et al., Identification of heavy-flavour jets with the CMS detector in pp collisions at 13 TeV , JINST 13 (2018) P05011 [1712.07158]

  24. [32]

    ATLAS collaboration, Identification of Jets Containing b-Hadrons with Recurrent Neural Networks at the ATLAS Experiment,

  25. [33]

    CMS collaboration, A. M. Sirunyan et al., Identification of heavy, energetic, hadronically decaying particles using machine-learning techniques, JINST 15 (2020) P06005 [2004.08262]

  26. [34]

    Aad et al., Fast b-tagging at the high-level trigger of the ATLAS experiment in LHC Run 3, JINST 18 (2023) P11006 [ 2306.09738]

    ATLAS collaboration, G. Aad et al., Fast b-tagging at the high-level trigger of the ATLAS experiment in LHC Run 3, JINST 18 (2023) P11006 [ 2306.09738]

  27. [35]

    Andrews et al., End-to-end jet classification of boosted top quarks with the CMS open data , EPJ Web Conf

    M. Andrews et al., End-to-end jet classification of boosted top quarks with the CMS open data , EPJ Web Conf. 251 (2021) 04030 [ 2104.14659]

  28. [36]

    Keicher, Machine Learning in Top Physics in the ATLAS and CMS Collaborations , in 15th International Workshop on Top Quark Physics , 1, 2023, 2301.09534

    P. Keicher, Machine Learning in Top Physics in the ATLAS and CMS Collaborations , in 15th International Workshop on Top Quark Physics , 1, 2023, 2301.09534

  29. [37]

    Baroˇ n, J

    P. Baroˇ n, J. Kvita, R. Pˇ r ´ ıvara, J. Tomeˇ cek and R. Vod´ ak,Application of Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes , in 16th International Workshop on Top Quark Physics , 10, 2023, 2310.13009

  30. [38]

    Hammad, S

    A. Hammad, S. Moretti and M. Nojiri, Multi-scale cross-attention transformer encoder for event classification, JHEP 03 (2024) 144 [ 2401.00452]

  31. [39]

    Esmail, A

    W. Esmail, A. Hammad and S. Moretti, Sharpening the A → Z(∗)h signature of the Type-II 2HDM at the LHC through advanced Machine Learning, JHEP 11 (2023) 020 [2305.13781]

  32. [40]

    Datta, A

    K. Datta, A. Larkoski and B. Nachman, Automating the Construction of Jet Observables with Machine Learning , Phys. Rev. D 100 (2019) 095016 [ 1902.07180]

  33. [41]

    Chakraborty, S

    A. Chakraborty, S. H. Lim and M. M. Nojiri, Interpretable deep learning for two-prong jet classification with jet spectra , JHEP 07 (2019) 135 [1904.02092]

  34. [42]

    Kim and A

    T. Kim and A. Martin, A W ± polarization analyzer from Deep Neural Networks , 2102.05124

  35. [43]

    Subba and R

    A. Subba and R. K. Singh, Role of polarizations and spin-spin correlations of W’s in e-e+ →W-W+ at s=250 GeV to probe anomalous W-W+Z/ γ couplings, Phys. Rev. D 107 (2023) 073004 [ 2212.12973]

  36. [44]

    Bogatskiy, T

    A. Bogatskiy, T. Hoffman, D. W. Miller, J. T. Offermann and X. Liu, Explainable equivariant neural networks for particle physics: PELICAN , JHEP 03 (2024) 113 [ 2307.16506]

  37. [45]

    S. Akar, T. J. Boettcher, S. Carl, H. F. Schreiner, M. D. Sokoloff, M. Stahl et al., An updated hybrid deep learning algorithm for identifying and locating primary vertices, 2007.01023

  38. [46]

    Shlomi, S

    J. Shlomi, S. Ganguly, E. Gross, K. Cranmer, Y. Lipman, H. Serviansky et al., Secondary vertex finding in jets with neural networks , Eur. Phys. J. C 81 (2021) 540 [ 2008.02831]

  39. [47]

    K. Goto, T. Suehara, T. Yoshioka, M. Kurata, H. Nagahara, Y. Nakashima et al., Development of a vertex finding algorithm using Recurrent Neural Network, Nucl. Instrum. Meth. A 1047 (2023) 167836 [2101.11906]

  40. [48]

    Guiang et al., Improving tracking algorithms with machine learning: a case for line-segment tracking at the High Luminosity LHC , in Connecting The Dots 2023, 3, 2024, 2403.13166

    J. Guiang et al., Improving tracking algorithms with machine learning: a case for line-segment tracking at the High Luminosity LHC , in Connecting The Dots 2023, 3, 2024, 2403.13166

  41. [49]

    Erdmann, A tagger for strange jets based on tracking information using long short-term memory , JINST 15 (2020) P01021 [ 1907.07505]

    J. Erdmann, A tagger for strange jets based on tracking information using long short-term memory , JINST 15 (2020) P01021 [ 1907.07505]

  42. [50]

    Nakai, D

    Y. Nakai, D. Shih and S. Thomas, Strange Jet Tagging, 2003.09517

  43. [51]

    Erdmann, O

    J. Erdmann, O. Nackenhorst and S. V. Zeißner, Maximum performance of strange-jet tagging at hadron colliders, JINST 16 (2021) P08039 [ 2011.10736]

  44. [52]

    P. T. Komiske, E. M. Metodiev and M. D. Schwartz, Deep learning in color: towards automated quark/gluon jet discrimination , JHEP 01 (2017) 110 [ 1612.01551]

  45. [53]

    Cheng, Recursive Neural Networks in Quark/Gluon Tagging, Comput

    T. Cheng, Recursive Neural Networks in Quark/Gluon Tagging, Comput. Softw. Big Sci. 2 (2018) 3 [1711.02633]

  46. [54]

    Abbas, A

    M. Abbas, A. Khan, A. S. Qureshi and M. W. Khan, Extracting Signals of Higgs Boson From Background Noise Using Deep Neural Networks , 2010.08201

  47. [55]

    CMS collaboration, A. Tumasyan et al., Search for Higgs Boson and Observation of Z Boson through their Decay into a Charm Quark-Antiquark Pair in Boosted Topologies in Proton-Proton Collisions at s=13 TeV , Phys. Rev. Lett. 131 (2023) 041801 [ 2211.14181]

  48. [56]

    Zhang, J

    Z. Zhang, J. Liu, J. Hu, Q. Wang and U.-G. Meißner, Revealing the nature of hidden charm pentaquarks with machine learning, Sci. Bull. 68 (2023) 981 [2301.05364]

  49. [57]

    Goswami, S

    K. Goswami, S. Prasad, N. Mallick, R. Sahoo and G. B. Mohanty, A machine learning-based study of open-charm hadrons in proton-proton collisions at the Large Hadron Collider , 2404.09839

  50. [58]

    H. Qu, C. Li and S. Qian, Particle Transformer for Jet Tagging, 2202.03772

  51. [59]

    Beauchesne, Z.-E

    H. Beauchesne, Z.-E. Chen and C.-W. Chiang, Improving the performance of weak supervision searches using transfer and meta-learning , JHEP 02 (2024) 138 [2312.06152]

  52. [60]

    Cogan, M

    J. Cogan, M. Kagan, E. Strauss and A. Schwarztman, Jet-Images: Computer Vision Inspired Techniques for Jet Tagging, JHEP 02 (2015) 118 [ 1407.5675]

  53. [61]

    de Oliveira, M

    L. de Oliveira, M. Kagan, L. Mackey, B. Nachman and A. Schwartzman, Jet-images — deep learning edition , JHEP 07 (2016) 069 [ 1511.05190]

  54. [62]

    Barnard, E

    J. Barnard, E. N. Dawe, M. J. Dolan and N. Rajcic, Parton Shower Uncertainties in Jet Substructure Analyses with Deep Neural Networks , Phys. Rev. D 95 (2017) 014018 [ 1609.00607]

  55. [63]

    P. T. Komiske, E. M. Metodiev, B. Nachman and M. D. Schwartz, Learning to classify from impure samples with high-dimensional data , Phys. Rev. D 98 (2018) 011502 [1801.10158]

  56. [64]

    J. S. H. Lee, I. Park, I. J. Watson and S. Yang, Quark-Gluon Jet Discrimination Using Convolutional Neural Networks, J. Korean Phys. Soc. 74 (2019) 219 [2012.02531]

  57. [65]

    Collado, K

    J. Collado, K. Bauer, E. Witkowski, T. Faucett, D. Whiteson and P. Baldi, Learning to isolate muons , JHEP 21 (2020) 200 [ 2102.02278]. 11

  58. [66]

    Li and H

    J. Li and H. Sun, An Attention Based Neural Network for Jet Tagging , 2009.00170

  59. [67]

    J. Li, T. Li and F.-Z. Xu, Reconstructing boosted Higgs jets from event image segmentation , JHEP 04 (2021) 156 [2008.13529]

  60. [68]

    Filipek, S.-C

    J. Filipek, S.-C. Hsu, J. Kruper, K. Mohan and B. Nachman, Identifying the Quantum Properties of Hadronic Resonances using Machine Learning , 2105.04582

  61. [69]

    T. Han, I. M. Lewis, H. Liu, Z. Liu and X. Wang, A guide to diagnosing colored resonances at hadron colliders, JHEP 08 (2023) 173 [ 2306.00079]

  62. [70]

    Kheddar, Y

    H. Kheddar, Y. Himeur, A. Amira and R. Soualah, Image Classification in High-Energy Physics: A Comprehensive Survey of Applications to Jet Analysis , 2403.11934

  63. [71]

    Abdughani, J

    M. Abdughani, J. Ren, L. Wu and J. M. Yang, Probing stop pair production at the LHC with graph neural networks, JHEP 08 (2019) 055 [ 1807.09088]

  64. [72]

    E. A. Moreno, O. Cerri, J. M. Duarte, H. B. Newman, T. Q. Nguyen, A. Periwal et al., JEDI-net: a jet identification algorithm based on interaction networks , Eur. Phys. J. C 80 (2020) 58 [ 1908.05318]

  65. [73]

    Bernreuther, T

    E. Bernreuther, T. Finke, F. Kahlhoefer, M. Kr¨ amer and A. M¨ uck,Casting a graph net to catch dark showers, SciPost Phys. 10 (2021) 046 [ 2006.08639]

  66. [74]

    Iiyama et al., Distance-Weighted Graph Neural Networks on FPGAs for Real-Time Particle Reconstruction in High Energy Physics , Front

    Y. Iiyama et al., Distance-Weighted Graph Neural Networks on FPGAs for Real-Time Particle Reconstruction in High Energy Physics , Front. Big Data 3 (2020) 598927 [ 2008.03601]

  67. [75]

    Heintz et al., Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs , in 34th Conference on Neural Information Processing Systems , 11, 2020, 2012.01563

    A. Heintz et al., Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs , in 34th Conference on Neural Information Processing Systems , 11, 2020, 2012.01563

  68. [76]

    Exa.TrkX collaboration, X. Ju et al., Graph Neural Networks for Particle Reconstruction in High Energy Physics detectors, in 33rd Annual Conference on Neural Information Processing Systems, 3, 2020, 2003.11603

  69. [77]

    J. Pata, J. Duarte, J.-R. Vlimant, M. Pierini and M. Spiropulu, MLPF: Efficient machine-learned particle-flow reconstruction using graph neural networks , Eur. Phys. J. C 81 (2021) 381 [ 2101.08578]

  70. [78]

    Verma and S

    Y. Verma and S. Jena, Jet characterization in Heavy Ion Collisions by QCD-Aware Graph Neural Networks , 2103.14906

  71. [79]

    Atkinson, A

    O. Atkinson, A. Bhardwaj, S. Brown, C. Englert, D. J. Miller and P. Stylianou, Improved constraints on effective top quark interactions using edge convolution networks, JHEP 04 (2022) 137 [ 2111.01838]

  72. [80]

    S. Gong, Q. Meng, J. Zhang, H. Qu, C. Li, S. Qian et al., An efficient Lorentz equivariant graph neural network for jet tagging , JHEP 07 (2022) 030 [2201.08187]

  73. [81]

    F. Ma, F. Liu and W. Li, Jet tagging algorithm of graph network with Haar pooling message passing , Phys. Rev. D 108 (2023) 072007 [ 2210.13869]

  74. [82]

    Bogatskiy, T

    A. Bogatskiy, T. Hoffman, D. W. Miller and J. T. Offermann, PELICAN: Permutation Equivariant and Lorentz Invariant or Covariant Aggregator Network for Particle Physics , 2211.00454

  75. [83]

    Builtjes, S

    L. Builtjes, S. Caron, P. Moskvitina, C. Nellist, R. R. de Austri, R. Verheyen et al., Attention to the strengths of physical interactions: Transformer and graph-based event classification for particle physics experiments , 2211.05143

  76. [84]

    F. A. Di Bello et al., Reconstructing particles in jets using set transformer and hypergraph prediction networks, Eur. Phys. J. C 83 (2023) 596 [ 2212.01328]

  77. [85]

    Mokhtar, R

    F. Mokhtar, R. Kansal and J. Duarte, Do graph neural networks learn traditional jet substructure? , in 36th Conference on Neural Information Processing Systems: Workshop on Machine Learning and the Physical Sciences, 11, 2022, 2211.09912

  78. [86]

    Huang, X

    A. Huang, X. Ju, J. Lyons, D. Murnane, M. Pettee and L. Reed, Heterogeneous Graph Neural Network for identifying hadronically decayed tau leptons at the High Luminosity LHC , JINST 18 (2023) P07001 [2301.00501]

  79. [87]

    ATLAS collaboration, A. Duperrin, Flavour tagging with graph neural networks with the ATLAS detector , in 30th International Workshop on Deep-Inelastic Scattering and Related Subjects , 6, 2023, 2306.04415

  80. [88]

    Konar, V

    P. Konar, V. S. Ngairangbam and M. Spannowsky, Hypergraphs in LHC phenomenology — the next frontier of IRC-safe feature extraction , JHEP 01 (2024) 113 [2309.17351]

  81. [89]

    S. R. Qasim, J. Kieseler, Y. Iiyama and M. Pierini, Learning representations of irregular particle-detector geometry with distance-weighted graph networks , Eur. Phys. J. C 79 (2019) 608 [ 1902.07987]

  82. [90]

    C. R. Qi, H. Su, K. Mo and L. J. Guibas, Pointnet: Deep learning on point sets for 3d classification and segmentation, in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 652–660, 2017

  83. [91]

    Zaheer, S

    M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Poczos, R. R. Salakhutdinov and A. J. Smola, Deep sets, Advances in neural information processing systems 30 (2017)

  84. [92]

    P. T. Komiske, E. M. Metodiev and J. Thaler, Energy Flow Networks: Deep Sets for Particle Jets , JHEP 01 (2019) 121 [ 1810.05165]

  85. [93]

    Zhang, V

    L. Zhang, V. Tozzo, J. Higgins and R. Ranganath, Set norm and equivariant skip connections: Putting the deep in deep sets , in International Conference on Machine Learning, pp. 26559–26574, PMLR, 2022

  86. [94]

    Qu and L

    H. Qu and L. Gouskos, ParticleNet: Jet Tagging via Particle Clouds , Phys. Rev. D 101 (2020) 056019 [1902.08570]

  87. [95]

    Shmakov, M

    A. Shmakov, M. J. Fenton, T.-W. Ho, S.-C. Hsu, D. Whiteson and P. Baldi, SPANet: Generalized permutationless set assignment for particle physics using symmetry preserving attention , SciPost Phys. 12 (2022) 178 [ 2106.03898]

  88. [96]

    Finke, M

    T. Finke, M. Kr¨ amer, A. M¨ uck and J. T¨ onshoff, Learning the language of QCD jets with transformers , JHEP 06 (2023) 184 [ 2303.07364]

  89. [97]

    Mikuni and F

    V. Mikuni and F. Canelli, Point cloud transformers applied to collider physics , Mach. Learn. Sci. Tech. 2 (2021) 035027 [ 2102.05073]

  90. [98]

    Spinner, V

    J. Spinner, V. Bres´ o, P. de Haan, T. Plehn, J. Thaler and J. Brehmer, Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics , 2405.14806

  91. [99]

    Bhardwaj, C

    A. Bhardwaj, C. Englert, W. Naskar, V. S. Ngairangbam and M. Spannowsky, Equivariant, safe and sensitive — graph networks for new physics , JHEP 07 (2024) 245 [ 2402.12449]

  92. [100]

    van Beekveld et al., A new standard for the logarithmic accuracy of parton showers , 2406.02661

    M. van Beekveld et al., A new standard for the logarithmic accuracy of parton showers , 2406.02661. 12

  93. [101]

    Bierlich et al., A comprehensive guide to the physics and usage of PYTHIA 8.3 , SciPost Phys

    C. Bierlich et al., A comprehensive guide to the physics and usage of PYTHIA 8.3 , SciPost Phys. Codeb. 2022 (2022) 8 [ 2203.11601]

  94. [102]

    de Favereau, C

    DELPHES 3collaboration, J. de Favereau, C. Delaere, P. Demin, A. Giammanco, V. Lema ˆ ıtre, A. Mertens et al., DELPHES 3, A modular framework for fast simulation of a generic collider experiment , JHEP 02 (2014) 057 [ 1307.6346]

  95. [103]

    Kornblith, M

    S. Kornblith, M. Norouzi, H. Lee and G. Hinton, Similarity of neural network representations revisited , in International conference on machine learning , pp. 3519–3529, PMLR, 2019

  96. [104]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh and D. Batra, Grad-cam: visual explanations from deep networks via gradient-based localization , International journal of computer vision 128 (2020) 336

  97. [105]

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva and A. Torralba, Learning deep features for discriminative localization, in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 2921–2929, 2016

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.