REVIEW 5 major objections 7 minor 115 references
A multi-category classifier on mixed jet samples is bounded by a simplex whose vertices define operational jet flavors and recover their mixing fractions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 05:47 UTC pith:KWHKCW6B
load-bearing objection Clean multi-topic generalization of operational jet flavor with a real pipeline; the physics demo is useful but leans on an untested universality assumption in tag-and-probe. the 5 major comments →
Simplex Demixing: Disentangling Multiple Light-Flavor Jets at Colliders
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
At the categorical cross-entropy minimum, a classifier trained on M mixtures that are convex combinations of T ≤ M mutually irreducible topics has convex hull equal to a (T−1)-simplex. The T vertices of that simplex determine the mixing fractions (up to permutation) whenever the fraction matrix has full column rank. The paper turns this geometry into a three-stage learning procedure—learn, shape, prune—called simplex demixing, and shows that it recovers multiple light-flavor operational topics from dijet mixtures.
What carries the argument
Simplex demixing: a classifier is parameterized so its outputs live in a learnable simplex inside the probability simplex; an edge-length loss pulls the vertices onto the data cloud and an L1 weight prunes excess vertices, after which the vertices invert to the mixing-fraction matrix.
Load-bearing premise
Each latent flavor must occupy a nonempty pure region in the measured jet features so the simplex vertices are actually reached in finite data; rare flavors and limited particle identification can leave those regions empty.
What would settle it
Train the demixer on the same tag-and-probe dijet mixtures but with particle-ID features removed or with substantially lower statistics; if the seven-vertex geometry collapses or the strong one-to-one match to down/up/strange/anti-strange/gluon disappears, the claimed identifiability fails.
If this is right
- Mixing fractions read from simplex vertices let one reconstruct any jet observable’s distribution for each operational light flavor without parton labels.
- The same geometric procedure applies to any continuous or set-valued features, not only jets, whenever mixed samples hide mutually irreducible topics.
- With HL-LHC-scale dijet samples and full hadron information, five light flavors are strongly recoverable at the ensemble level even if single-jet tagging remains hard.
- Operational topics extracted from different processes can be compared directly, testing how much “quark” or “gluon” depends on the surrounding event.
Where Pith is reading between the lines
- Choosing the number of topics without domain knowledge will likely need stability or lasso-path criteria already common in sparse feature selection; the paper flags this but does not solve it.
- If π/K separation is weak, strange-related vertices may merge with down-like ones, so the method’s reach is tied to detector PID more tightly than the idealized study shows.
- The tag-and-probe construction still leans on a supervised Monte-Carlo tagger to build mixtures; a fully unsupervised route to diverse mixtures would remove that last label dependence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript generalizes the operational quark/gluon jet definition from two mixtures to M jet samples and T mutually irreducible topics. It proves that, at the categorical cross-entropy optimum, classifier outputs lie in a (T−1)-simplex whose vertices determine the topic mixing fractions, subject to mutual irreducibility and a full-rank fraction matrix. A three-stage learn/shape/prune network implements the idea. Toy Pythia mixtures recover d/u/g structure and show the expected collapse to a line for two topics. In an HL-LHC-like dijet study, a Pythia-trained seven-flavor tagger and eta binning create 14 tag-and-probe mixtures; after enforcing T=7, five topics align strongly with d, u, s, anti-s, and gluon jets and two align weakly with anti-d and anti-u. Topic-weighted distributions of constituent multiplicity, 2-subjettiness, and jet charge generally reproduce the corresponding Pythia distributions.
Significance. If the universality assumptions are validated, this is a significant extension of data-driven jet-flavor definitions beyond quark/gluon separation and a useful bridge between topic modeling and collider measurements. Notable strengths are an explicit theorem with stated rank and mutual-irreducibility conditions, a practical architecture for continuous point-cloud jets, public code and toy data, bootstrap uncertainty estimates, falsifiable simplex geometry, and an unusually candid treatment of rare antiflavors, gluon contamination, and the idealized detector assumptions. At present, however, the physics study remains a proof of concept because its mixtures and topic count rely on Pythia information and its most important conditional-independence assumption is not tested.
major comments (5)
- [§5.2, Eqs. (5.1) and (5.4)] Theorem 1 applies to the physics study only if all 14 mixtures are convex combinations of the same topic distributions. Equation (5.1) assumes that the two jets’ hadron-level features are conditionally independent given their topics, and Eq. (5.4) then drops c(x'), eta, and tau from p(x|t,c(x'),eta,tau). Real dijets retain pT-balance, color, MPI/UE, and shower correlations, and the confidence cut can alter the probe features at fixed topic. If this universality fails, the learned vertices are selection-weighted effective topics and Eq. (5.12) inherits a systematic not visible to the bootstrap. Please test this directly, e.g. compare probe observables at fixed flavor/topic across tags, eta bins, and thresholds, and repeat the demixing under varied selections.
- [§5.1–§5.2; abstract and §1] The 14 mixtures are constructed with a supervised PFN trained on seven Pythia-labeled flavors and a 0.8 confidence cut. Thus the demixing stage is label-free, but the overall procedure is not yet a data-only extraction and can steer the learned simplex toward the generator’s own flavor taxonomy. The agreement in Figs. 7–9 is therefore partly a Pythia/tagger closure test rather than independent discovery. The abstract and introduction should narrow the claim to a tag-assisted proof of concept, and the paper should quantify tagger/generator dependence or explain a concrete deployment path that does not use truth-labeled simulation.
- [§5.3, stage three; Fig. 6] The L1 strength beta is explicitly chosen to keep exactly seven vertices because seven light flavors are expected. Consequently, Fig. 6 establishes that a stable seven-vertex representation can be learned after imposing T=7; it does not independently show that seven topics emerge from the data. Since the seven-topic conclusion is central, please add domain-agnostic evidence—e.g. validation CCE/Ledge versus T, active-vertex stability across bootstrap runs, and comparisons for T=6 and T=8—or state consistently that the result is conditional on the externally supplied topic count.
- [§3.5, Eqs. (3.35)–(3.38)] The equality p(x|t,c)=p_t(x) is asserted because the topics were constructed without the category labels, but absence from training does not imply conditional statistical independence. In general, conditioning on an overlapping truth category c reweights x within topic t. Hence p(t|c) need not be a nonnegative conditional probability even asymptotically; like p(c|t), it is generally a signed linear-overlap coefficient unless an additional screening-off assumption holds. This affects the interpretation of Figs. 4, 7, and 8 and the probability arguments in §5.5. Please state and justify the extra assumption or recast both coefficient matrices as quasi-probability/overlap matrices.
- [§4.2, §5.3, and §5.6] The reported 15%–85% intervals use a fixed-hyperparameter bootstrap from one Pythia sample after one supervised tagger and selection. As the text acknowledges, this omits retuning variance; it also omits vertex-number, threshold, eta-bin, tagger, generator, and broken-factorization systematics. Since Figs. 9–10 assess physical agreement against these bands, they should be labeled as internal statistical intervals and supplemented by at least the leading selection/model variations. In particular, bootstrap resampling cannot reveal a bias common to every resample, such as a violation of Eq. (5.1).
minor comments (7)
- [§5.2, dataset preparation] Please clarify whether the 70%/20%/10% split is performed by event or by jet. Both jets from one dijet enter the pooled probe sample, so a per-jet split could place correlated objects in training and validation/test sets; block bootstrap resampling alone would not remove that leakage.
- [Fig. 6] Only 10 of the 91 possible two-dimensional projections are shown. Please explain the selection criterion and provide either a supplementary full projection grid or a quantitative measure demonstrating that every retained vertex lies on the learned convex hull.
- [Table 2] The fractions are rounded to two decimals, while G_mc enters a pseudoinverse. Please state that unrounded fractions were used and clarify whether the quoted fractions were renormalized after the perfect heavy-flavor exclusion.
- [§4.2 and §5.3] The toy and physics studies use 20 and 34 bootstrap resamples, respectively. The 15% and 85% quantiles are then based on only a few tail samples; please report convergence of the intervals with resample count.
- [§3.4, Eq. (3.27)] The notation t,t' in the edge loss should make clear that the sum is over all M architectural vertices before pruning, even when the eventual physical topic count is smaller.
- [§5.4] Please define quantitative criteria for the terms “strongly identifiable” and “weakly identifiable,” rather than relying only on visual inspection of the coefficient matrices.
- [Code Availability] The code link is welcome. It would improve reproducibility to archive a versioned release, fixed configuration files, random seeds, and the trained supervised tagger used to construct the mixtures.
Circularity Check
Mostly non-circular: Theorem 1 is a self-contained derivation; mild definitional character of operational topics and domain-knowledge choice of T=7 do not force the Pythia-closure results.
specific steps
-
self definitional
[Sec. 2.1 (operational identification); Sec. 5.3–5.4 (T=7 prune)]
"The key assumption of the operational definition of quark and gluon jets is that these two topics should be identified with the "quark" and "gluon" distributions up to permutation [62]. ... We pick β to keep only T=7 active vertices, using our domain knowledge that there should be seven light flavors in the samples."
Operational topics are defined as the mutually irreducible simplex vertices (maximally separable categories), then identified with flavor names by assumption; T is set to the expected number of light flavors rather than selected by a data-only criterion. This makes the topic count and the label–topic dictionary partly definitional. It is mild: the measured p(t|c)/p(c|t) alignments and substructure shapes are still empirical and can (and do) fail for rare flavors.
full rationale
Theorem 1 (Sec. 3.2) derives the (T−1)-simplex geometry and recoverability of F_mt from the categorical cross-entropy stationary point plus mutual irreducibility and full column rank; the proof is internal and does not reduce to a fit or to an unverified self-citation. The operational definition intentionally equates topics with maximally separable (mutually irreducible) categories—this is transparent methodology, not a hidden claim that an independent external label was derived. Validation against Pythia uses separate truth labels via pseudoinverse quasi-probabilities and reports partial failure modes (weak d̄/ū, nonzero κ_qg), which is the opposite of a forced closure. Self-citations to the two-mixture CWoLa/jet-topics papers supply the special case being generalized; the multi-mixture theorem and three-stage architecture are new and proved/specified here. The only mild circularity-adjacent choices are (i) fixing T=7 from prior knowledge of seven light flavors when pruning, so the count of topics is not discovered, and (ii) building the 14 mixtures with a Pythia-supervised tagger, which injects composition diversity aligned with those labels—yet the probe-side demixing and substructure inversion remain nontrivial empirical tests. Assumption risks in the tag-and-probe factorization (Eq. 5.1) affect correctness, not circularity of the derivation chain. Score 2 reflects those minor design choices without elevating them to load-bearing circular reduction.
Axiom & Free-Parameter Ledger
free parameters (6)
- edge-loss weight α =
5e-4 (toy); 5e-5 (physics)
- L1 prune weight β =
≈0.0245 for T=7
- L2 weight γ =
1e-4 (toy default); 0 (physics)
- topic count T =
7
- tagger confidence threshold =
0.8
- η bin boundary =
0.75
axioms (6)
- domain assumption Mixtures are row-stochastic convex combinations of T latent topic distributions (Eq. 3.8).
- domain assumption Topics are mutually irreducible: each has an anchor region where it is positive and all others vanish.
- domain assumption Fraction matrix F has full column rank (vertices linearly independent).
- standard math At δL_CCE=0 the network outputs posterior mixture probabilities (Eq. 3.14).
- ad hoc to paper Tag-and-probe mixtures from a supervised Pythia tagger plus η binning span the same universal operational topics as the dijet ensemble.
- ad hoc to paper Perfect hadron-level particle ID (including π/K/p) and negligible untagged heavy flavor.
invented entities (1)
-
operational topics (simplex vertices as hadron-level jet flavors)
no independent evidence
read the original abstract
Providing a practical and hadron-level definition of multiple jet flavors has been a long-standing challenge in collider physics. Previous work has introduced a data-driven, operational definition of quark and gluon jets, but no robust generalization beyond two jet categories presently exists. To address this, we introduce a machine-learning framework called "simplex demixing'' to extract $T$ jet flavors (or topics in the statistics literature) from $M$ data samples (or mixtures) with minimal constraints. Intuitively, our procedure identifies the maximally separable categories in the data, translating a multi-category classifier on the $M$ mixtures into a bounded geometric object with $T$ vertices. We first demonstrate our procedure on a toy problem to infer the truth-level fractions of down-quark, up-quark, and gluon jets from synthetic mixtures of the three pure samples. We then propose a tag-and-probe strategy to extract multiple light-flavor categories in a more realistic collider setting involving dijet production. As expected, the identifiability of jet flavors depends on their relative abundance in the samples and the hadron-level information available to the classifier architecture. Our work opens the door to data-driven extractions of multiple jet flavor properties at the Large Hadron Collider.
Reference graph
Works this paper leans on
-
[1]
Nilles and K.H
H.P. Nilles and K.H. Streng,Quark - Gluon Separation in Three Jet Events,Phys. Rev. D 23(1981) 1944
1981
-
[2]
Jones,Tests for Determining the Parton Ancestor of a Hadron Jet,Phys
L.M. Jones,Tests for Determining the Parton Ancestor of a Hadron Jet,Phys. Rev. D39 (1989) 2550
1989
-
[3]
Fodor,How to See the Differences Between Quark and Gluon Jets,Phys
Z. Fodor,How to See the Differences Between Quark and Gluon Jets,Phys. Rev. D41 (1990) 1726
1990
-
[4]
Jones,TOWARDS A SYSTEMATIC JET CLASSIFICATION,Phys
L. Jones,TOWARDS A SYSTEMATIC JET CLASSIFICATION,Phys. Rev. D42(1990) 811
1990
-
[5]
Lonnblad, C
L. Lonnblad, C. Peterson and T. Rognvaldsson,Using neural networks to identify jets, Nucl. Phys. B349(1991) 675
1991
-
[6]
Pumplin,How to tell quark jets from gluon jets,Phys
J. Pumplin,How to tell quark jets from gluon jets,Phys. Rev. D44(1991) 2025
1991
-
[7]
J. Gallicchio and M.D. Schwartz,Quark and Gluon Tagging at the LHC,Phys. Rev. Lett. 107(2011) 172001 [1106.3076]
Pith/arXiv arXiv 2011
-
[8]
J. Gallicchio and M.D. Schwartz,Quark and Gluon Jet Substructure,JHEP04(2013) 090 [1211.7038]
Pith/arXiv arXiv 2013
-
[9]
B. Bhattacherjee, S. Mukhopadhyay, M.M. Nojiri, Y. Sakaki and B.R. Webber,Associated jet and subjet rates in light-quark and gluon jet discrimination,JHEP04(2015) 131 [1501.04794]
Pith/arXiv arXiv 2015
-
[10]
D. Ferreira de Lima, P. Petrov, D. Soper and M. Spannowsky,Quark-Gluon tagging with Shower Deconstruction: Unearthing dark matter and Higgs couplings,Phys. Rev. D95 (2017) 034001 [1607.06031]
Pith/arXiv arXiv 2017
-
[11]
B. Bhattacherjee, S. Mukhopadhyay, M.M. Nojiri, Y. Sakaki and B.R. Webber,Quark-gluon discrimination in the search for gluino pair production at the LHC,JHEP01(2017) 044 [1609.08781]
Pith/arXiv arXiv 2017
-
[12]
J. Davighi and P. Harris,Fractal based observables to probe jet substructure of quarks and gluons,Eur. Phys. J. C78(2018) 334 [1703.00914]
Pith/arXiv arXiv 2018
-
[13]
A.J. Larkoski and E.M. Metodiev,A Theory of Quark vs. Gluon Discrimination,JHEP10 (2019) 014 [1906.01639]. [14]CMScollaboration,Search for direct production of supersymmetric partners of the top quark in the all-jets final state in proton-proton collisions at √s= 13TeV,JHEP10(2017) 005 [1707.03316]. – 39 – [15]CMScollaboration,Search for vectorlike light-...
Pith/arXiv arXiv 2019
-
[23]
G.P. Salam,Towards Jetography,Eur. Phys. J. C67(2010) 637 [0906.1833]
Pith/arXiv arXiv 2010
-
[24]
Abdesselam et al.,Boosted Objects: A Probe of Beyond the Standard Model Physics, Eur
A. Abdesselam et al.,Boosted Objects: A Probe of Beyond the Standard Model Physics, Eur. Phys. J. C71(2011) 1661 [1012.5412]
Pith/arXiv arXiv 2011
-
[25]
T. Plehn and M. Spannowsky,Top Tagging,J. Phys. G39(2012) 083001 [1112.4441]
Pith/arXiv arXiv 2012
-
[26]
Altheimer et al.,Jet Substructure at the Tevatron and LHC: New Results, New Tools, New Benchmarks,J
A. Altheimer et al.,Jet Substructure at the Tevatron and LHC: New Results, New Tools, New Benchmarks,J. Phys. G39(2012) 063001 [1201.0008]
Pith/arXiv arXiv 2012
-
[27]
J. Shelton,Jet Substructure, inTheoretical Advanced Study Institute in Elementary Particle Physics: Searching for New Physics at Small and Large Scales, pp. 303–340, 2013, DOI [1302.0260]
Pith/arXiv arXiv 2013
-
[28]
Altheimer et al.,Boosted Objects and Jet Substructure at the LHC
A. Altheimer et al.,Boosted Objects and Jet Substructure at the LHC. Report of BOOST2012, held at IFIC Valencia, 23rd-27th of July 2012,Eur. Phys. J. C74(2014) 2792 [1311.2708]
Pith/arXiv arXiv 2012
-
[29]
Adams et al.,Towards an Understanding of the Correlations in Jet Substructure,Eur
D. Adams et al.,Towards an Understanding of the Correlations in Jet Substructure,Eur. Phys. J. C75(2015) 409 [1504.00679]
Pith/arXiv arXiv 2015
-
[30]
Cacciari,Phenomenological and theoretical developments in jet physics at the LHC,Int
M. Cacciari,Phenomenological and theoretical developments in jet physics at the LHC,Int. J. Mod. Phys. A30(2015) 1546001 [1509.02272]
Pith/arXiv arXiv 2015
-
[31]
Kogler et al.,Jet Substructure at the Large Hadron Collider: Experimental Review,Rev
R. Kogler et al.,Jet Substructure at the Large Hadron Collider: Experimental Review,Rev. Mod. Phys.91(2019) 045003 [1803.06991]
Pith/arXiv arXiv 2019
-
[32]
S. Marzani, G. Soyez and M. Spannowsky,Looking inside jets: an introduction to jet substructure and boosted-object phenomenology, vol. 958, Springer (2019), 10.1007/978-3-030-15709-8, [1901.10342]. – 40 –
Pith/arXiv arXiv 2019
-
[33]
A.J. Larkoski, I. Moult and B. Nachman,Jet Substructure at the Large Hadron Collider: A Review of Recent Advances in Theory and Machine Learning,Phys. Rept.841(2020) 1 [1709.04464]
Pith/arXiv arXiv 2020
-
[34]
Larkoski,QCD masterclass lectures on jet physics and machine learning,Eur
A.J. Larkoski,QCD masterclass lectures on jet physics and machine learning,Eur. Phys. J. C84(2024) 1117 [2407.04897]
Pith/arXiv arXiv 2024
-
[35]
D. Guest, K. Cranmer and D. Whiteson,Deep Learning and its Application to LHC Physics,Ann. Rev. Nucl. Part. Sci.68(2018) 161 [1806.11484]
Pith/arXiv arXiv 2018
-
[36]
Albertsson et al.,Machine Learning in High Energy Physics Community White Paper,J
K. Albertsson et al.,Machine Learning in High Energy Physics Community White Paper,J. Phys. Conf. Ser.1085(2018) 022008 [1807.02876]
Pith/arXiv arXiv 2018
-
[37]
Radovic, M
A. Radovic, M. Williams, D. Rousseau, M. Kagan, D. Bonacorsi, A. Himmel et al.,Machine learning at the energy and intensity frontiers of particle physics,Nature560(2018) 41
2018
-
[38]
G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby et al.,Machine learning and the physical sciences,Rev. Mod. Phys.91(2019) 045002 [1903.10563]
Pith/arXiv arXiv 2019
-
[39]
Bourilkov,Machine and Deep Learning Applications in Particle Physics,Int
D. Bourilkov,Machine and Deep Learning Applications in Particle Physics,Int. J. Mod. Phys. A34(2020) 1930019 [1912.08245]
Pith/arXiv arXiv 2020
-
[40]
Schwartz,Modern Machine Learning and Particle Physics,2103.12226
M.D. Schwartz,Modern Machine Learning and Particle Physics,2103.12226
-
[41]
M. Feickert and B. Nachman,A Living Review of Machine Learning for Particle Physics, 2102.02770
-
[42]
Boehnlein et al.,Colloquium: Machine learning in nuclear physics,Rev
A. Boehnlein et al.,Colloquium: Machine learning in nuclear physics,Rev. Mod. Phys.94 (2022) 031003 [2112.02309]
Pith/arXiv arXiv 2022
-
[43]
Karagiorgi, G
G. Karagiorgi, G. Kasieczka, S. Kravitz, B. Nachman and D. Shih,Machine learning in the search for new fundamental physics,Nature Rev. Phys.4(2022) 399
2022
-
[44]
T. Plehn, A. Butter, B. Dillon, T. Heimel, C. Krause and R. Winterhalder,Modern Machine Learning for LHC Physicists,2211.01421
-
[45]
Bonilla et al.,Jets and Jet Substructure at Future Colliders,Front
J. Bonilla et al.,Jets and Jet Substructure at Future Colliders,Front. in Phys.10(2022) 897719 [2203.07462]
Pith/arXiv arXiv 2022
-
[46]
DeZoort, P.W
G. DeZoort, P.W. Battaglia, C. Biscarat and J.-R. Vlimant,Graph neural networks at the Large Hadron Collider,Nature Rev. Phys.5(2023) 281
2023
-
[47]
K. Zhou, L. Wang, L.-G. Pang and S. Shi,Exploring QCD matter in extreme conditions with Machine Learning,Prog. Part. Nucl. Phys.135(2024) 104084 [2303.15136]
Pith/arXiv arXiv 2024
-
[48]
V. Belis, P. Odagiu and T.K. Aarrestad,Machine learning for anomaly detection in particle physics,Rev. Phys.12(2024) 100091 [2312.14190]
Pith/arXiv arXiv 2024
-
[49]
S. Mondal and L. Mastrolorenzo,Machine learning in high energy physics: a review of heavy-flavor jet tagging at the LHC,Eur. Phys. J. ST233(2024) 2657 [2404.01071]
Pith/arXiv arXiv 2024
-
[50]
P. Gras, S. H¨ oche, D. Kar, A. Larkoski, L. L¨ onnblad, S. Pl¨ atzer et al.,Systematics of quark/gluon tagging,JHEP07(2017) 091 [1704.03878]
Pith/arXiv arXiv 2017
-
[51]
A. Banfi, G.P. Salam and G. Zanderighi,Infrared safe definition of jet flavor,Eur. Phys. J. C47(2006) 113 [hep-ph/0601139]
Pith/arXiv arXiv 2006
-
[52]
A. Buckley and C. Pollard,QCD-aware partonic jet clustering for truth-jet flavour labelling, Eur. Phys. J. C76(2016) 71 [1507.00508]. – 41 –
Pith/arXiv arXiv 2016
-
[53]
S. Caletti, A.J. Larkoski, S. Marzani and D. Reichelt,Practical jet flavour through NNLO, Eur. Phys. J. C82(2022) 632 [2205.01109]
Pith/arXiv arXiv 2022
-
[54]
S. Caletti, A.J. Larkoski, S. Marzani and D. Reichelt,A fragmentation approach to jet flavor,JHEP10(2022) 158 [2205.01117]
Pith/arXiv arXiv 2022
-
[55]
M. Czakon, A. Mitov and R. Poncelet,Infrared-safe flavoured anti-k T jets,JHEP04(2023) 138 [2205.11879]
Pith/arXiv arXiv 2023
-
[56]
R. Gauld, A. Huss and G. Stagnitto,Flavor Identification of Reconstructed Hadronic Jets, Phys. Rev. Lett.130(2023) 161901 [2208.11138]
Pith/arXiv arXiv 2023
-
[57]
F. Caola, R. Grabarczyk, M.L. Hutt, G.P. Salam, L. Scyboz and J. Thaler,Flavored jets with exact anti-kt kinematics and tests of infrared and collinear safety,Phys. Rev. D108 (2023) 094010 [2306.07314]
Pith/arXiv arXiv 2023
-
[58]
Behring et al.,Flavoured jet algorithms: a comparative study,JHEP09(2025) 149 [2506.13449]
A. Behring et al.,Flavoured jet algorithms: a comparative study,JHEP09(2025) 149 [2506.13449]
arXiv 2025
-
[59]
J. Gallicchio and M.D. Schwartz,Pure Samples of Quark and Gluon Jets at the LHC, JHEP10(2011) 103 [1104.1175]
Pith/arXiv arXiv 2011
-
[60]
C. Frye, A.J. Larkoski, M.D. Schwartz and K. Yan,Factorization for groomed jet substructure beyond the next-to-leading logarithm,JHEP07(2016) 064 [1603.09338]
Pith/arXiv arXiv 2016
-
[61]
C. Frye, A.J. Larkoski, M.D. Schwartz and K. Yan,Precision physics with pile-up insensitive observables,1603.06375
-
[62]
P.T. Komiske, E.M. Metodiev and J. Thaler,An operational definition of quark and gluon jets,JHEP11(2018) 059 [1809.01140]
Pith/arXiv arXiv 2018
-
[63]
I.W. Stewart and X. Yao,Pure quark and gluon observables in collinear drop,JHEP09 (2022) 120 [2203.14980]
Pith/arXiv arXiv 2022
-
[64]
E.M. Metodiev, B. Nachman and J. Thaler,Classification without labels: Learning from mixed samples in high energy physics,JHEP10(2017) 174 [1708.02949]
Pith/arXiv arXiv 2017
-
[65]
E.M. Metodiev and J. Thaler,Jet Topics: Disentangling Quarks and Gluons at Colliders, Phys. Rev. Lett.120(2018) 241602 [1802.00008]
Pith/arXiv arXiv 2018
-
[66]
P.T. Komiske, S. Kryhin and J. Thaler,Disentangling quarks and gluons in CMS open data, Phys. Rev. D106(2022) 094021 [2205.04459]
Pith/arXiv arXiv 2022
-
[67]
M.J. Dolan, J. Gargalionis and A. Ore,Quark-versus-gluon tagging in CMS Open Data with CWoLa and TopicFlow,JHEP08(2025) 024 [2312.03434]. [68]ATLAScollaboration,Properties of jet fragmentation using charged particles measured with the ATLAS detector inppcollisions at √s= 13TeV,Phys. Rev. D100(2019) 052011 [1906.09254]. [69]ATLAS, CMScollaboration,Producti...
Pith/arXiv arXiv 2025
-
[72]
J. Brewer, J. Thaler and A.P. Turner,Data-driven quark and gluon jet modification in heavy-ion collisions,Phys. Rev. C103(2021) L021901 [2008.08596]
Pith/arXiv arXiv 2021
-
[73]
Bierlich et al.,A comprehensive guide to the physics and usage of PYTHIA 8.3,SciPost Phys
C. Bierlich et al.,A comprehensive guide to the physics and usage of PYTHIA 8.3,SciPost Phys. Codeb.2022(2022) 8 [2203.11601]
Pith/arXiv arXiv 2022
-
[74]
B.M. Dillon, D.A. Faroughy and J.F. Kamenik,Uncovering latent jet substructure,Phys. Rev. D100(2019) 056002 [1904.04200]
Pith/arXiv arXiv 2019
-
[75]
B.M. Dillon, D.A. Faroughy, J.F. Kamenik and M. Szewc,Learning the latent structure of collider events,JHEP10(2020) 206 [2005.12319]
Pith/arXiv arXiv 2020
-
[76]
E. Alvarez, M. Spannowsky and M. Szewc,Unsupervised Quark/Gluon Jet Tagging With Poissonian Mixture Models,Front. Artif. Intell.5(2022) 852970 [2112.11352]
Pith/arXiv arXiv 2022
-
[77]
M. LeBlanc, B. Nachman and C. Sauer,Going off topics to demix quark and gluon jets in αS extractions,JHEP02(2023) 150 [2206.10642]
Pith/arXiv arXiv 2023
-
[78]
Blei, A.Y
D.M. Blei, A.Y. Ng and M.I. Jordan,Latent dirichlet allocation,J. Mach. Learn. Res.3 (2003) 993–1022
2003
-
[79]
Papadimitriou, H
C.H. Papadimitriou, H. Tamaki, P. Raghavan and S. Vempala,Latent semantic indexing: a probabilistic analysis, inProceedings of the Seventeenth ACM SIGACT-SIGMOD-SIGART Symposium on Principles of Database Systems, PODS ’98, (New York, NY, USA), p. 159–168, Association for Computing Machinery, 1998, DOI
1998
-
[80]
Arora, R
S. Arora, R. Ge and A. Moitra,Learning topic models – going beyond svd, inProceedings of the 2012 IEEE 53rd Annual Symposium on Foundations of Computer Science, FOCS ’12, (USA), p. 1–10, IEEE Computer Society, 2012, DOI
2012
-
[81]
Arora, R
S. Arora, R. Ge, Y. Halpern, D. Mimno, A. Moitra, D. Sontag et al.,A practical algorithm for topic modeling with provable guarantees, inProceedings of the 30th International Conference on Machine Learning, S. Dasgupta and D. McAllester, eds., vol. 28 of Proceedings of Machine Learning Research, (Atlanta, Georgia, USA), pp. 280–288, PMLR, 17–19 Jun, 2013, ...
2013
-
[82]
Bioucas-Dias, A
J.M. Bioucas-Dias, A. Plaza, N. Dobigeon, M. Parente, Q. Du, P. Gader et al.,Hyperspectral unmixing overview: Geometrical, statistical, and sparse regression-based approaches,IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing5(2012) 354
2012
-
[83]
Cutler and L
A. Cutler and L. Breiman,Archetypal analysis,Technometrics36(1994) 338
1994
-
[84]
Eugster and F
M.J. Eugster and F. Leisch,From spider-man to hero—archetypal analysis in r,Journal of Statistical Software30(2009) 1
2009
-
[85]
Suleman,Validation of archetypal analysis, in2017 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), pp
A. Suleman,Validation of archetypal analysis, in2017 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), pp. 1–6, IEEE, 2017
2017
-
[86]
Winter,N-FINDR: an algorithm for fast autonomous spectral end-member determination in hyperspectral data, inImaging Spectrometry V, M.R
M.E. Winter,N-FINDR: an algorithm for fast autonomous spectral end-member determination in hyperspectral data, inImaging Spectrometry V, M.R. Descour and S.S. Shen, eds., vol. 3753, pp. 266 – 275, International Society for Optics and Photonics, SPIE, 1999, DOI
1999
-
[87]
Li and J.M
J. Li and J.M. Bioucas-Dias,Minimum volume simplex analysis: A fast algorithm to unmix hyperspectral data, inIGARSS 2008 - 2008 IEEE International Geoscience and Remote Sensing Symposium, vol. 3, pp. III – 250–III – 253, 2008, DOI. – 43 –
2008
-
[88]
X. Li, T. Liu, B. Han, G. Niu and M. Sugiyama,Provably end-to-end label-noise learning without anchor points, inProceedings of the 38th International Conference on Machine Learning, M. Meila and T. Zhang, eds., vol. 139 ofProceedings of Machine Learning Research, pp. 6403–6413, PMLR, 2021, https://proceedings.mlr.press/v139/li21l.html
2021
-
[89]
Katz-Samuels, G
J. Katz-Samuels, G. Blanchard and C. Scott,Decontamination of mutual contamination models,J. Mach. Learn. Res.20(2019) 1521–1577
2019
-
[90]
Bonnet-Guerrini, J
R. Bonnet-Guerrini, J. Ioannou-Nikolaides, T. Petersen and V. Piuri,Multiclass classification without labels via posterior simplex geometry,to appear
-
[91]
Neyman and E.S
J. Neyman and E.S. Pearson,On the Problem of the Most Efficient Tests of Statistical Hypotheses,Phil. Trans. Roy. Soc. Lond. A231(1933) 289
1933
-
[92]
Scott, G
C. Scott, G. Blanchard and G. Handy,Classification with asymmetric label noise: Consistency and maximal denoising, inProceedings of the 26th Annual Conference on Learning Theory, S. Shalev-Shwartz and I. Steinwart, eds., vol. 30 ofProceedings of Machine Learning Research, (Princeton, NJ, USA), pp. 489–511, PMLR, 12–14 Jun, 2013, https://proceedings.mlr.pr...
2013
-
[93]
Mair and J
S. Mair and J. Sj¨ olund,Archetypal analysis++: Rethinking the initialization strategy, Transactions on Machine Learning Research(2024)
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.