Pith. sign in

REVIEW 4 cited by

Is Tokenization Needed for Masked Particle Modelling?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12589 v2 pith:GPAQLZJY submitted 2024-09-19 hep-ph cs.LG

classification hep-phcs.LG
keywords learningmodelsdatafoundationmaskedmethodsobjectiveparticle
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Explicit or Implicit? Encoding Physics at the Precision Frontier

    hep-ph 2026-03 conditional novelty 6.0 of 10

    On three precision classification tasks — reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection — a Lorentz-equivariant transformer and a pretrained foundation model perform...

  2. Enhancing next token prediction based pre-training for jet foundation models

    hep-ph 2025-12 conditional novelty 6.0 of 10

    Using continuous particle features as input and combining next-token with masked-token pre-training markedly improves classification accuracy of the OmniJet jet foundation model without visibly hurting its generative quality.

  3. Learning Symmetry-Independent Jet Representations via Jet-Based Joint Embedding Predictive Architecture

    hep-ph 2024-12 conditional novelty 6.0 of 10

    J-JEPA pretraining on 1M jets modestly improves top jet tagging versus from-scratch training, but gains are inconsistent for the strongest baseline model.

  4. Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough

    physics.data-an 2026-07 accept novelty 4.0 of 10

    Verification of ML in fundamental physics is essential precisely when models enter statistical modeling, inference, or hypothesis testing, and is bounded by unavoidable inductive bias, sample complexity, and experimen...

Pith tools