REVIEW 3 major objections 5 minor 105 references
Transformer networks for Heavy flavor jet tagging
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A compact cross-attention mixer matches transformer-level top tagging at roughly one-twentieth the training time.
desk verdict Useful review of transformer-based jet taggers, but the headline '20x faster' speedup over ParT is not backed by controlled benchmarking and the manuscript has draft placeholders. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Cross-Attention-Mixer (CA-Mixer) network, a permutation-invariant architecture in which an MLP-mixer processes the unordered set of jet constituents while a cross-attention block reads subjet information. Each mixer layer applies two shared MLPs, one mixing particle tokens and one mixing feature channels, preserving permutation invariance; the same constituents are reclustered into subjets with a jet algorithm at a smaller radius, and multi-head cross-attention uses those subjets as keys and values while the mixed constituent features act as queries. A skip connection preserves the original particle coordinates, and global max pooling feeds a final classifier. The mechanism's job is to inject the physically known prong structure of boosted heavy particles into a network that is too small to rediscover it from data alone.
What would settle it
Run Particle Transformer and CA-Mixer on the same GPU with the same framework, batch size, and number of epochs on the boosted-top dataset and compare wall-clock training time; if the per-epoch ratio is far from 20, the headline speed claim fails.
Extended reading notes
Core claim
The paper's central discovery is that cross-attention between whole-jet and subjet-level inputs lets a compact permutation-invariant network reproduce transformer-level top-tagging performance. The CA-Mixer replaces the heavy self-attention stack of a particle transformer with an MLP-mixer that mixes particle tokens and feature channels, then applies multi-head attention in which the jet constituents act as queries and reclustered subjets provide the keys and values. This encodes the factorized structure of QCD parton showers, correlating resolved prongs with the hadron distribution around them, so the network does not need millions of parameters to learn prong structure from scratch. The reported result on the boosted-top dataset is an AUC of 0.9859 and a background rejection of 416 ± 5 at 50% signal efficiency, with 86.03K parameters and 33 seconds per epoch, against Particle Transformer's 0.9858, 413 ± 16, 2.14M parameters, and 612 seconds.
Load-bearing premise
The speed comparison assumes that the published per-epoch training times for the other networks were obtained under comparable hardware, framework, and optimization conditions; the paper specifies the GPU only for its own CA-Mixer runs.
Editorial extensions
If this is right
- Top-tagging accuracy at the level of a 2.14M-parameter particle transformer can be obtained with an 86K-parameter network when subjet cross-attention is added, so the large self-attention stack is not strictly necessary for this task.
- Per-epoch training time drops by about a factor of 20 on the stated GPU (33 seconds versus 612 seconds), which makes iterative retraining, hyperparameter scans, and model development much cheaper.
- The CKA representation-similarity analysis shows the cross-attention layer carries information distinct from the mixer MLPs, supporting the view that subjets contribute complementary physics structure rather than extra capacity.
- The same cross-attention idea extends beyond single jets to event-level classification, as already demonstrated for two-Higgs final states, so the mechanism is not limited to top tagging.
- The CA-Mixer keeps permutation invariance by construction, so it inherits the main advantage of particle-cloud representations without the computational cost of pairwise self-attention over all constituents.
Reading between the lines
- If the training-time comparison in Table I were repeated under strictly identical hardware, software, and optimization budgets, the factor-of-20 advantage would probably shrink but would not vanish, because the parameter-count difference is intrinsic to the designs; a controlled benchmark is the clean test.
- Subjet cross-attention should transfer most strongly to other multi-pronged signatures such as boosted W, Z, and Higgs tagging, and the architecture's gain over a plain mixer should be largest for jets with well-separated prongs.
- A promising next variant is combining the subjet cross-attention with Lorentz-equivariant aggregators of the PELICAN type; the paper itself notes Lorentz equivariance as a possible further improvement, and that combination could be checked for whether it closes the small AUC gap to PELICAN.
- The interpretability analysis yields a testable diagnostic: ablating the cross-attention block should hurt multi-pronged top jets more than QCD jets, which would confirm that prong structure, not raw capacity, carries the performance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a review of machine-learning methods for heavy-flavor (boosted top, W/Z/H) jet tagging, with a focus on transformer architectures. It summarizes data representations (jet images, graphs, particle clouds), describes the attention mechanism and its particle-transformer extension, and discusses architectures that incorporate physics structure: PELICAN (Lorentz invariance), LundNet (IRC-safe clustering), and the authors' CA-Mixer, which uses cross-attention between jet constituents and reconstructed subjets. Sec. IV.B compares top-tagging performance on the standard boosted-top dataset, reporting that CA-Mixer achieves AUC 0.9859 with 86K parameters and 33 s/epoch, matching ParticleNET and ParT (AUC 0.9858, 2.14M parameters, 612 s/epoch) while training 'approximately 20 times faster.' Sec. V reviews interpretability tools (CKA, attention maps, Grad-CAM) and includes a CKA figure for CA-Mixer taken from Ref. [24].
Significance. If the speed comparison were properly supported, the paper would make a useful contribution: it calls attention to a compact 86K-parameter architecture whose AUC is statistically indistinguishable from much larger transformer taggers on the standard top-tagging benchmark, and it frames this in the context of physics-motivated network design. The review portions are generally accurate: the descriptions of jet-image, graph, and particle-cloud representations are consistent with the cited literature, and the attention equations in Sec. II are standard. The paper also gives credit to the original sources and makes a concrete, falsifiable performance table rather than stopping at qualitative claims. However, the paper provides no independent verification of the CA-Mixer numbers (they are quoted from the authors' own Ref. [24]), no convergence curves or total training times, and no normalized timing benchmark for the '20 times faster' claim. The paper also contains a large undigested draft block, which suggests the manuscript is not in final form.
major comments (3)
- [Sec. IV.B, Table I] The central speedup claim ('trains approximately 20 times faster') is not supported as stated. The table caption says results for JEDI-net, PFN, PCT, LorentzNet, ParticleNET, PELICAN, and ParT are 'quoted from their published results,' while only CA-Mixer's 33 s/epoch is clearly attributed to the authors' measurement on an NVIDIA RTX A6000. Nothing in the caption or text demonstrates that ParT's 612 s/epoch was obtained on the same GPU, with the same software framework, mixed-precision settings, data-loading pipeline, and batch size; the caption specifies batch size 1024 but does not state whether the quoted ParT run used it. In addition, per-epoch time is not the same as training cost to a fixed target AUC unless the two models require the same number of epochs, and no convergence curves or total training times are given. Please either rerun ParT (and ideally ParticleNET) under identical local conditions and report epochs to the quoted AUC, or remove the '20 times faster' claim and limit Sec. IV.B to the AUC/parameter comparison, which is what the quoted evidence supports.
- [Sec. IV.A (after the LundNet paragraph)] The manuscript contains a large undigested draft block beginning 'REPEAT & Global MAX pooling ...' and continuing through a duplicate title/author block dated March 18, 2024, a list of 'Needed plots,' and a second 'FIG. 3' caption. This is not a presentation typo: it indicates that the LaTeX source contains leftover material from an earlier draft. The block must be completely removed and the surrounding text and figure numbering re-checked so that the published version contains exactly one title block, one abstract, and one coherent set of figures.
- [Sec. IV.A (CA-Mixer description)] The description of CA-Mixer is too incomplete for a paper whose title and Sec. IV.B highlight this network. The text refers to Ref. [24] for details but does not specify how subjets are formed (which reclustering algorithm, which radius Rcut, how many subjets are kept), how the cross-attention queries, keys, and values are constructed from the constituent and subjet sets (Eqs. (2)-(6) define generic attention, not the specific cross-attention shown in Fig. 3), or the mixer-layer dimensions and hyperparameters behind the quoted 86K parameter count. Without this information, the reported AUC and the parameter comparison in Table I cannot be checked. Please add a short but complete architectural specification, or state explicitly that all CA-Mixer numbers are taken from Ref. [24] and not reproduced independently here.
minor comments (5)
- [Throughout] There are numerous typos and misspellings that should be corrected: 'Coliider' for 'Collider' in Sec. I, 'psudo-rapidity' for 'pseudo-rapidity' and 'marge' for 'merge' in Sec. IV.A, and 'pretaining' for 'pretraining' in the caption of Table I.
- [Sec. V.B] The attention-map discussion says each element of alpha_ij in Eq. (3) is the attention from particle i to particle j, but in the CA-Mixer model the attention is cross-attention between constituents and subjets. The text should generalize the description or clarify that Eq. (3) describes the self-attention case.
- [Table I] The table would benefit from explicit definitions: 'Rej 50%' should state that it is the background rejection at 50% signal efficiency (or the inverse), and the entries with '--' should be explained as not reported in the original publications rather than left ambiguous.
- [Sec. V] The text says 'we explore different methods for interpreting network decision-making' and 'we will apply some of the methods,' but the section only reproduces one CKA figure from Ref. [24] and gives no new Grad-CAM or attention-map results for CA-Mixer. The wording should be adjusted to reflect the review nature of the contribution.
- [Abstract and Sec. I] The abstract and introduction describe the paper as a review, yet Sec. IV.B contains an original comparative speed claim. Please add a sentence clarifying that the CA-Mixer numbers and interpretability figures are taken from the authors' prior work (Ref. [24]) and that the new element in this paper is the synthesis and comparison.
Circularity Check
No significant circularity: the paper is a review whose quantitative claims are cited from prior published work, including the authors' own Ref. [24], and are not derived from or defined in terms of the present manuscript.
full rationale
The manuscript is an explicit review of machine-learning jet-tagging methods. It does not fit parameters and then rename them as predictions, and no equation in the paper is constructed so that an output equals an input by definition. The CA-Mixer architecture, its AUC, parameter count, and per-epoch training time are introduced with the phrase 'In Ref. [24], we have proposed...' and Table I is captioned 'Taken from Ref. [24]', with all other network results 'quoted from their published results'. This is self-citation, but it is not load-bearing circularity: Ref. [24] is a separate, externally available publication with its own stated assumptions and measurements, and the comparison rows against ParticleNET, ParT, PELICAN, and other networks are independent external benchmarks. The paper makes no uniqueness claim and invokes no theorem from the same authors to forbid alternative architectures. The physics motivation via QCD factorization is a design rationale, not a derived prediction. The only substantive concern is benchmark comparability of per-epoch timings quoted from different papers, which is a correctness and normalization issue, not a circularity. Accordingly, no circular step meets the evidentiary bar.
Assumptions & free parameters
assumptions (3)
- domain assumption The benchmark top tagging dataset from Ref. [6] (Pythia8 plus Delphes, pT in [550,650] GeV, |eta| < 2) provides a valid standardized testbed for comparing jet taggers.
- domain assumption Results from different papers (JEDI-net, PFN, PCT, LorentzNet, ParticleNET, PELICAN, ParT, CA-Mixer) are comparable despite differing implementations, training procedures, and measurement environments.
- domain assumption The CA-Mixer architecture, described only schematically here, is exactly as specified in Ref. [24].
Cite this review
Pith. "Pith review of Transformer networks for Heavy flavor jet tagging." pith.science (2026). https://pith.science/paper/AMTTLVT7
@misc{pith2026241111519,
author = {Pith},
title = {Pith review of: Transformer networks for Heavy flavor jet tagging},
year = {2026},
howpublished = {\url{https://pith.science/paper/AMTTLVT7}},
note = {Machine review of arXiv:2411.11519}
}
read the original abstract
In this article, we review recent machine learning methods used in challenging particle identification of heavy-boosted particles at high-energy colliders. Our primary focus is on attention-based Transformer networks. We report the performance of state-of-the-art deep learning networks and further improvement coming from the modification of networks based on physics insights. Additionally, we discuss interpretable methods to understand network decision-making, which are crucial when employing highly complex and deep networks.
Figures
Reference graph
Works this paper leans on
-
[24]
J. M. Munoz, I. Batatia and C. Ortner, Boost invariant polynomials for efficient jet tagging , Mach. Learn. Sci. Tech. 3 (2022) 04LT05 [2207.08272]
arXiv 2022
-
[1]
J. M. Butterworth, A. R. Davison, M. Rubin and G. P. Salam, Jet substructure as a new Higgs search channel at the LHC , Phys. Rev. Lett. 100 (2008) 242001 [0802.2470]
arXiv 2008
-
[2]
minimum spanning tree + hierarchical dendogram (on how the hdbscan work)
-
[3]
four plots for subjets clustering for the top case
-
[4]
ROC for varying radius from 0.1 to 0.5 using CA
-
[5]
plot for the input subjets for the top and qcd jets
-
[6]
plot for the cross attention
-
[7]
Foundation of Machine Learning Physics
ROCs for every thing. PACS numb ers: I. INTRODUCTION N x (Mixer layer) Input dataset: Particle constitution dim: (100 , 7) Global Max Pooling MLP Class Mixer layer Layer NormEmbedding T MLP-1 MLP-1 Shared T Layer Norm Cross-attention Sub-jets Particles Features Particles Features Particles Particles Features Particles Particles MLP-2 MLP-2 Shared Skip-con...
Show all 105 references
-
[8]
Abdesselam et al., Boosted Objects: A Probe of Beyond the Standard Model Physics , Eur
A. Abdesselam et al., Boosted Objects: A Probe of Beyond the Standard Model Physics , Eur. Phys. J. C 71 (2011) 1661 [ 1012.5412]
2011 arXiv
-
[9]
Altheimer et al., Jet Substructure at the Tevatron and LHC: New Results, New Tools, New Benchmarks , J
A. Altheimer et al., Jet Substructure at the Tevatron and LHC: New Results, New Tools, New Benchmarks , J. Phys. G 39 (2012) 063001 [ 1201.0008]
2012 arXiv
-
[10]
Altheimer et al., Boosted Objects and Jet Substructure at the LHC
A. Altheimer et al., Boosted Objects and Jet Substructure at the LHC. Report of BOOST2012, held at IFIC Valencia, 23rd-27th of July 2012 , Eur. Phys. J. C 74 (2014) 2792 [ 1311.2708]
2014 arXiv
-
[11]
L. G. Almeida, M. Backovi´ c, M. Cliche, S. J. Lee and M. Perelstein, Playing Tag with ANN: Boosted Top Identification with Pattern Recognition, JHEP 07 (2015) 086 [ 1501.05968]
2015 arXiv
-
[12]
Butter, G
A. Butter, G. Kasieczka, T. Plehn and M. Russell, Deep-learned Top Tagging with a Lorentz Layer , SciPost Phys. 5 (2018) 028 [ 1707.08966]
2018 arXiv
-
[13]
Kasieczka, T
G. Kasieczka, T. Plehn, M. Russell and T. Schell, Deep-learning Top Taggers or The End of QCD? , JHEP 05 (2017) 006 [ 1701.08784]
2017 arXiv
-
[14]
Louppe, K
G. Louppe, K. Cho, C. Becot and K. Cranmer, QCD-Aware Recursive Neural Networks for Jet Physics , JHEP 01 (2019) 057 [ 1702.00748]
2019 arXiv
-
[15]
Butter et al., The Machine Learning landscape of top taggers, SciPost Phys
A. Butter et al., The Machine Learning landscape of top taggers, SciPost Phys. 7 (2019) 014 [ 1902.09914]
2019 arXiv
-
[16]
Chakraborty, S
A. Chakraborty, S. H. Lim, M. M. Nojiri and M. Takeuchi, Neural Network-based Top Tagger with Two-Point Energy Correlations and Geometry of Soft Emissions, JHEP 07 (2020) 111 [ 2003.11787]
2020 arXiv
-
[17]
Bhattacharya, M
S. Bhattacharya, M. Guchait and A. H. Vijay, Boosted top quark tagging and polarization measurement using machine learning, Phys. Rev. D 105 (2022) 042005 [2010.11778]
2022 arXiv
-
[18]
Ju and B
X. Ju and B. Nachman, Supervised Jet Clustering with Graph Neural Networks for Lorentz Boosted Bosons , Phys. Rev. D 102 (2020) 075014 [ 2008.06064]
2020 arXiv
-
[19]
F. A. Dreyer and H. Qu, Jet tagging in the Lund plane with graph networks , JHEP 03 (2021) 052 [2012.08526]
2021 arXiv
-
[20]
Tannenwald, C
B. Tannenwald, C. Neu, A. Li, G. Buehlmann, A. Cuddeback, L. Hatfield et al., Benchmarking Machine Learning Techniques with Di-Higgs Production at the LHC , 2009.06754
2009 arXiv
-
[21]
F. A. Dreyer, R. Grabarczyk and P. F. Monni, Leveraging universality of jet taggers through transfer learning, Eur. Phys. J. C 82 (2022) 564 [ 2203.06210]
2022 arXiv
-
[22]
Hammad, S
A. Hammad, S. Khalil and S. Moretti, Search for mono-Higgs signals in bb ¯ final states using deep neural networks, Phys. Rev. D 107 (2023) 075027 [2208.10133]
2023 arXiv
-
[23]
Ahmed, A
I. Ahmed, A. Zada, M. Waqas and M. U. Ashraf, Application of deep learning in top pair and single top quark production at the LHC , Eur. Phys. J. Plus 138 (2023) 795 [ 2203.12871]
2023 arXiv
-
[25]
He and D
M. He and D. Wang, Quark/gluon discrimination and top tagging with dual attention transformer , Eur. Phys. J. C 83 (2023) 1116 [ 2307.04723]
2023 arXiv
-
[26]
J. A. Aguilar-Saavedra, E. Arganda, F. R. Joaquim, R. M. Sand´ a Seoane and J. F. Seabra,Gradient Boosting MUST taggers for highly-boosted jets , 2305.04957
-
[27]
Athanasakos, A
D. Athanasakos, A. J. Larkoski, J. Mulligan, M. P losko´ n and F. Ringer, Is infrared-collinear safe information all you need for jet classification? , 2305.08979
-
[28]
Grossi, M
M. Grossi, M. Incudini, M. Pellen and G. Pelliccioli, Amplitude-assisted tagging of longitudinally polarised bosons using wide neural networks , Eur. Phys. J. C 83 (2023) 759 [ 2306.07726]
2023 arXiv
-
[29]
Hammad, P
A. Hammad, P. Ko, C.-T. Lu and M. Park, Exploring Exotic Decays of the Higgs Boson to Multi-Photons at the LHC via Multimodal Learning Approaches , 2405.18834. 10
-
[30]
Hammad and M
A. Hammad and M. M. Nojiri, Streamlined jet tagging network assisted by jet prong structure , JHEP 06 (2024) 176 [2404.14677]
2024 arXiv
-
[31]
CMS collaboration, A. M. Sirunyan et al., Identification of heavy-flavour jets with the CMS detector in pp collisions at 13 TeV , JINST 13 (2018) P05011 [1712.07158]
2018 arXiv
-
[32]
ATLAS collaboration, Identification of Jets Containing b-Hadrons with Recurrent Neural Networks at the ATLAS Experiment,
-
[33]
CMS collaboration, A. M. Sirunyan et al., Identification of heavy, energetic, hadronically decaying particles using machine-learning techniques, JINST 15 (2020) P06005 [2004.08262]
2020 arXiv
-
[34]
Aad et al., Fast b-tagging at the high-level trigger of the ATLAS experiment in LHC Run 3, JINST 18 (2023) P11006 [ 2306.09738]
ATLAS collaboration, G. Aad et al., Fast b-tagging at the high-level trigger of the ATLAS experiment in LHC Run 3, JINST 18 (2023) P11006 [ 2306.09738]
2023 arXiv
-
[35]
Andrews et al., End-to-end jet classification of boosted top quarks with the CMS open data , EPJ Web Conf
M. Andrews et al., End-to-end jet classification of boosted top quarks with the CMS open data , EPJ Web Conf. 251 (2021) 04030 [ 2104.14659]
2021 arXiv
-
[36]
Keicher, Machine Learning in Top Physics in the ATLAS and CMS Collaborations , in 15th International Workshop on Top Quark Physics , 1, 2023, 2301.09534
P. Keicher, Machine Learning in Top Physics in the ATLAS and CMS Collaborations , in 15th International Workshop on Top Quark Physics , 1, 2023, 2301.09534
2023 arXiv
-
[37]
Baroˇ n, J
P. Baroˇ n, J. Kvita, R. Pˇ r ´ ıvara, J. Tomeˇ cek and R. Vod´ ak,Application of Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes , in 16th International Workshop on Top Quark Physics , 10, 2023, 2310.13009
2023 arXiv
-
[38]
Hammad, S
A. Hammad, S. Moretti and M. Nojiri, Multi-scale cross-attention transformer encoder for event classification, JHEP 03 (2024) 144 [ 2401.00452]
2024 arXiv
-
[39]
Esmail, A
W. Esmail, A. Hammad and S. Moretti, Sharpening the A → Z(∗)h signature of the Type-II 2HDM at the LHC through advanced Machine Learning, JHEP 11 (2023) 020 [2305.13781]
2023 arXiv
-
[40]
Datta, A
K. Datta, A. Larkoski and B. Nachman, Automating the Construction of Jet Observables with Machine Learning , Phys. Rev. D 100 (2019) 095016 [ 1902.07180]
2019 arXiv
-
[41]
Chakraborty, S
A. Chakraborty, S. H. Lim and M. M. Nojiri, Interpretable deep learning for two-prong jet classification with jet spectra , JHEP 07 (2019) 135 [1904.02092]
2019 arXiv
-
[42]
Kim and A
T. Kim and A. Martin, A W ± polarization analyzer from Deep Neural Networks , 2102.05124
-
[43]
Subba and R
A. Subba and R. K. Singh, Role of polarizations and spin-spin correlations of W’s in e-e+ →W-W+ at s=250 GeV to probe anomalous W-W+Z/ γ couplings, Phys. Rev. D 107 (2023) 073004 [ 2212.12973]
2023 arXiv
-
[44]
Bogatskiy, T
A. Bogatskiy, T. Hoffman, D. W. Miller, J. T. Offermann and X. Liu, Explainable equivariant neural networks for particle physics: PELICAN , JHEP 03 (2024) 113 [ 2307.16506]
2024 arXiv
-
[45]
S. Akar, T. J. Boettcher, S. Carl, H. F. Schreiner, M. D. Sokoloff, M. Stahl et al., An updated hybrid deep learning algorithm for identifying and locating primary vertices, 2007.01023
2007 arXiv
-
[46]
Shlomi, S
J. Shlomi, S. Ganguly, E. Gross, K. Cranmer, Y. Lipman, H. Serviansky et al., Secondary vertex finding in jets with neural networks , Eur. Phys. J. C 81 (2021) 540 [ 2008.02831]
2021 arXiv
-
[47]
K. Goto, T. Suehara, T. Yoshioka, M. Kurata, H. Nagahara, Y. Nakashima et al., Development of a vertex finding algorithm using Recurrent Neural Network, Nucl. Instrum. Meth. A 1047 (2023) 167836 [2101.11906]
2023 arXiv
-
[48]
Guiang et al., Improving tracking algorithms with machine learning: a case for line-segment tracking at the High Luminosity LHC , in Connecting The Dots 2023, 3, 2024, 2403.13166
J. Guiang et al., Improving tracking algorithms with machine learning: a case for line-segment tracking at the High Luminosity LHC , in Connecting The Dots 2023, 3, 2024, 2403.13166
2023 arXiv
-
[49]
Erdmann, A tagger for strange jets based on tracking information using long short-term memory , JINST 15 (2020) P01021 [ 1907.07505]
J. Erdmann, A tagger for strange jets based on tracking information using long short-term memory , JINST 15 (2020) P01021 [ 1907.07505]
2020 arXiv
- [50]
-
[51]
Erdmann, O
J. Erdmann, O. Nackenhorst and S. V. Zeißner, Maximum performance of strange-jet tagging at hadron colliders, JINST 16 (2021) P08039 [ 2011.10736]
2021 arXiv
-
[52]
P. T. Komiske, E. M. Metodiev and M. D. Schwartz, Deep learning in color: towards automated quark/gluon jet discrimination , JHEP 01 (2017) 110 [ 1612.01551]
2017 arXiv
-
[53]
Cheng, Recursive Neural Networks in Quark/Gluon Tagging, Comput
T. Cheng, Recursive Neural Networks in Quark/Gluon Tagging, Comput. Softw. Big Sci. 2 (2018) 3 [1711.02633]
2018 arXiv
-
[54]
Abbas, A
M. Abbas, A. Khan, A. S. Qureshi and M. W. Khan, Extracting Signals of Higgs Boson From Background Noise Using Deep Neural Networks , 2010.08201
2010 arXiv
-
[55]
CMS collaboration, A. Tumasyan et al., Search for Higgs Boson and Observation of Z Boson through their Decay into a Charm Quark-Antiquark Pair in Boosted Topologies in Proton-Proton Collisions at s=13 TeV , Phys. Rev. Lett. 131 (2023) 041801 [ 2211.14181]
2023 arXiv
-
[56]
Zhang, J
Z. Zhang, J. Liu, J. Hu, Q. Wang and U.-G. Meißner, Revealing the nature of hidden charm pentaquarks with machine learning, Sci. Bull. 68 (2023) 981 [2301.05364]
2023 arXiv
-
[57]
Goswami, S
K. Goswami, S. Prasad, N. Mallick, R. Sahoo and G. B. Mohanty, A machine learning-based study of open-charm hadrons in proton-proton collisions at the Large Hadron Collider , 2404.09839
-
[58]
H. Qu, C. Li and S. Qian, Particle Transformer for Jet Tagging, 2202.03772
-
[59]
Beauchesne, Z.-E
H. Beauchesne, Z.-E. Chen and C.-W. Chiang, Improving the performance of weak supervision searches using transfer and meta-learning , JHEP 02 (2024) 138 [2312.06152]
2024 arXiv
-
[60]
Cogan, M
J. Cogan, M. Kagan, E. Strauss and A. Schwarztman, Jet-Images: Computer Vision Inspired Techniques for Jet Tagging, JHEP 02 (2015) 118 [ 1407.5675]
2015 arXiv
-
[61]
de Oliveira, M
L. de Oliveira, M. Kagan, L. Mackey, B. Nachman and A. Schwartzman, Jet-images — deep learning edition , JHEP 07 (2016) 069 [ 1511.05190]
2016 arXiv
-
[62]
Barnard, E
J. Barnard, E. N. Dawe, M. J. Dolan and N. Rajcic, Parton Shower Uncertainties in Jet Substructure Analyses with Deep Neural Networks , Phys. Rev. D 95 (2017) 014018 [ 1609.00607]
2017 arXiv
-
[63]
P. T. Komiske, E. M. Metodiev, B. Nachman and M. D. Schwartz, Learning to classify from impure samples with high-dimensional data , Phys. Rev. D 98 (2018) 011502 [1801.10158]
2018 arXiv
-
[64]
J. S. H. Lee, I. Park, I. J. Watson and S. Yang, Quark-Gluon Jet Discrimination Using Convolutional Neural Networks, J. Korean Phys. Soc. 74 (2019) 219 [2012.02531]
2019 arXiv
-
[65]
Collado, K
J. Collado, K. Bauer, E. Witkowski, T. Faucett, D. Whiteson and P. Baldi, Learning to isolate muons , JHEP 21 (2020) 200 [ 2102.02278]. 11
2020 arXiv
-
[66]
Li and H
J. Li and H. Sun, An Attention Based Neural Network for Jet Tagging , 2009.00170
2009 arXiv
-
[67]
J. Li, T. Li and F.-Z. Xu, Reconstructing boosted Higgs jets from event image segmentation , JHEP 04 (2021) 156 [2008.13529]
2021 arXiv
-
[68]
Filipek, S.-C
J. Filipek, S.-C. Hsu, J. Kruper, K. Mohan and B. Nachman, Identifying the Quantum Properties of Hadronic Resonances using Machine Learning , 2105.04582
-
[69]
T. Han, I. M. Lewis, H. Liu, Z. Liu and X. Wang, A guide to diagnosing colored resonances at hadron colliders, JHEP 08 (2023) 173 [ 2306.00079]
2023 arXiv
-
[70]
Kheddar, Y
H. Kheddar, Y. Himeur, A. Amira and R. Soualah, Image Classification in High-Energy Physics: A Comprehensive Survey of Applications to Jet Analysis , 2403.11934
-
[71]
Abdughani, J
M. Abdughani, J. Ren, L. Wu and J. M. Yang, Probing stop pair production at the LHC with graph neural networks, JHEP 08 (2019) 055 [ 1807.09088]
2019 arXiv
-
[72]
E. A. Moreno, O. Cerri, J. M. Duarte, H. B. Newman, T. Q. Nguyen, A. Periwal et al., JEDI-net: a jet identification algorithm based on interaction networks , Eur. Phys. J. C 80 (2020) 58 [ 1908.05318]
2020 arXiv
-
[73]
Bernreuther, T
E. Bernreuther, T. Finke, F. Kahlhoefer, M. Kr¨ amer and A. M¨ uck,Casting a graph net to catch dark showers, SciPost Phys. 10 (2021) 046 [ 2006.08639]
2021 arXiv
-
[74]
Iiyama et al., Distance-Weighted Graph Neural Networks on FPGAs for Real-Time Particle Reconstruction in High Energy Physics , Front
Y. Iiyama et al., Distance-Weighted Graph Neural Networks on FPGAs for Real-Time Particle Reconstruction in High Energy Physics , Front. Big Data 3 (2020) 598927 [ 2008.03601]
2020 arXiv
-
[75]
Heintz et al., Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs , in 34th Conference on Neural Information Processing Systems , 11, 2020, 2012.01563
A. Heintz et al., Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs , in 34th Conference on Neural Information Processing Systems , 11, 2020, 2012.01563
2020 arXiv
-
[76]
Exa.TrkX collaboration, X. Ju et al., Graph Neural Networks for Particle Reconstruction in High Energy Physics detectors, in 33rd Annual Conference on Neural Information Processing Systems, 3, 2020, 2003.11603
2020 arXiv
-
[77]
J. Pata, J. Duarte, J.-R. Vlimant, M. Pierini and M. Spiropulu, MLPF: Efficient machine-learned particle-flow reconstruction using graph neural networks , Eur. Phys. J. C 81 (2021) 381 [ 2101.08578]
2021 arXiv
-
[78]
Verma and S
Y. Verma and S. Jena, Jet characterization in Heavy Ion Collisions by QCD-Aware Graph Neural Networks , 2103.14906
-
[79]
Atkinson, A
O. Atkinson, A. Bhardwaj, S. Brown, C. Englert, D. J. Miller and P. Stylianou, Improved constraints on effective top quark interactions using edge convolution networks, JHEP 04 (2022) 137 [ 2111.01838]
2022 arXiv
-
[80]
S. Gong, Q. Meng, J. Zhang, H. Qu, C. Li, S. Qian et al., An efficient Lorentz equivariant graph neural network for jet tagging , JHEP 07 (2022) 030 [2201.08187]
2022 arXiv
-
[81]
F. Ma, F. Liu and W. Li, Jet tagging algorithm of graph network with Haar pooling message passing , Phys. Rev. D 108 (2023) 072007 [ 2210.13869]
2023 arXiv
-
[82]
Bogatskiy, T
A. Bogatskiy, T. Hoffman, D. W. Miller and J. T. Offermann, PELICAN: Permutation Equivariant and Lorentz Invariant or Covariant Aggregator Network for Particle Physics , 2211.00454
-
[83]
Builtjes, S
L. Builtjes, S. Caron, P. Moskvitina, C. Nellist, R. R. de Austri, R. Verheyen et al., Attention to the strengths of physical interactions: Transformer and graph-based event classification for particle physics experiments , 2211.05143
-
[84]
F. A. Di Bello et al., Reconstructing particles in jets using set transformer and hypergraph prediction networks, Eur. Phys. J. C 83 (2023) 596 [ 2212.01328]
2023 arXiv
-
[85]
Mokhtar, R
F. Mokhtar, R. Kansal and J. Duarte, Do graph neural networks learn traditional jet substructure? , in 36th Conference on Neural Information Processing Systems: Workshop on Machine Learning and the Physical Sciences, 11, 2022, 2211.09912
2022 arXiv
-
[86]
Huang, X
A. Huang, X. Ju, J. Lyons, D. Murnane, M. Pettee and L. Reed, Heterogeneous Graph Neural Network for identifying hadronically decayed tau leptons at the High Luminosity LHC , JINST 18 (2023) P07001 [2301.00501]
2023 arXiv
-
[87]
ATLAS collaboration, A. Duperrin, Flavour tagging with graph neural networks with the ATLAS detector , in 30th International Workshop on Deep-Inelastic Scattering and Related Subjects , 6, 2023, 2306.04415
2023 arXiv
-
[88]
Konar, V
P. Konar, V. S. Ngairangbam and M. Spannowsky, Hypergraphs in LHC phenomenology — the next frontier of IRC-safe feature extraction , JHEP 01 (2024) 113 [2309.17351]
2024 arXiv
-
[89]
S. R. Qasim, J. Kieseler, Y. Iiyama and M. Pierini, Learning representations of irregular particle-detector geometry with distance-weighted graph networks , Eur. Phys. J. C 79 (2019) 608 [ 1902.07987]
2019 arXiv
-
[90]
C. R. Qi, H. Su, K. Mo and L. J. Guibas, Pointnet: Deep learning on point sets for 3d classification and segmentation, in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 652–660, 2017
2017
-
[91]
Zaheer, S
M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Poczos, R. R. Salakhutdinov and A. J. Smola, Deep sets, Advances in neural information processing systems 30 (2017)
2017
-
[92]
P. T. Komiske, E. M. Metodiev and J. Thaler, Energy Flow Networks: Deep Sets for Particle Jets , JHEP 01 (2019) 121 [ 1810.05165]
2019 arXiv
-
[93]
Zhang, V
L. Zhang, V. Tozzo, J. Higgins and R. Ranganath, Set norm and equivariant skip connections: Putting the deep in deep sets , in International Conference on Machine Learning, pp. 26559–26574, PMLR, 2022
2022
-
[94]
Qu and L
H. Qu and L. Gouskos, ParticleNet: Jet Tagging via Particle Clouds , Phys. Rev. D 101 (2020) 056019 [1902.08570]
2020 arXiv
-
[95]
Shmakov, M
A. Shmakov, M. J. Fenton, T.-W. Ho, S.-C. Hsu, D. Whiteson and P. Baldi, SPANet: Generalized permutationless set assignment for particle physics using symmetry preserving attention , SciPost Phys. 12 (2022) 178 [ 2106.03898]
2022 arXiv
-
[96]
Finke, M
T. Finke, M. Kr¨ amer, A. M¨ uck and J. T¨ onshoff, Learning the language of QCD jets with transformers , JHEP 06 (2023) 184 [ 2303.07364]
2023 arXiv
-
[97]
Mikuni and F
V. Mikuni and F. Canelli, Point cloud transformers applied to collider physics , Mach. Learn. Sci. Tech. 2 (2021) 035027 [ 2102.05073]
2021 arXiv
-
[98]
Spinner, V
J. Spinner, V. Bres´ o, P. de Haan, T. Plehn, J. Thaler and J. Brehmer, Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics , 2405.14806
-
[99]
Bhardwaj, C
A. Bhardwaj, C. Englert, W. Naskar, V. S. Ngairangbam and M. Spannowsky, Equivariant, safe and sensitive — graph networks for new physics , JHEP 07 (2024) 245 [ 2402.12449]
2024 arXiv
-
[100]
van Beekveld et al., A new standard for the logarithmic accuracy of parton showers , 2406.02661
M. van Beekveld et al., A new standard for the logarithmic accuracy of parton showers , 2406.02661. 12
-
[101]
Bierlich et al., A comprehensive guide to the physics and usage of PYTHIA 8.3 , SciPost Phys
C. Bierlich et al., A comprehensive guide to the physics and usage of PYTHIA 8.3 , SciPost Phys. Codeb. 2022 (2022) 8 [ 2203.11601]
2022 arXiv
-
[102]
de Favereau, C
DELPHES 3collaboration, J. de Favereau, C. Delaere, P. Demin, A. Giammanco, V. Lema ˆ ıtre, A. Mertens et al., DELPHES 3, A modular framework for fast simulation of a generic collider experiment , JHEP 02 (2014) 057 [ 1307.6346]
2014 arXiv
-
[103]
Kornblith, M
S. Kornblith, M. Norouzi, H. Lee and G. Hinton, Similarity of neural network representations revisited , in International conference on machine learning , pp. 3519–3529, PMLR, 2019
2019
-
[104]
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh and D. Batra, Grad-cam: visual explanations from deep networks via gradient-based localization , International journal of computer vision 128 (2020) 336
2020
-
[105]
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva and A. Torralba, Learning deep features for discriminative localization, in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 2921–2929, 2016
2016
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.