{"id":"fe2aee54-75af-4fbb-a3c1-bd0f5bd5815f","arxiv_id":"2412.05333","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"J-JEPA pretraining on 1M jets modestly improves top jet tagging versus from-scratch training, but gains are inconsistent for the strongest baseline model.","lead":"The paper introduces J-JEPA, a self-supervised method that learns jet representations by predicting masked subjet representations from context subjets, without requiring hand-crafted augmentations. Finetuning these representations for top quark tagging improves performance over training from scratch, with the largest gains at 10% of labeled data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 does not uniformly support the central claim: the best AE-SjT-T Cls Attn model is worse after finetuning on full-data rejection (95.47 vs 99.38), and several gains are within error bars.","rationale":"The most load-bearing weakness is not a formal inconsistency but the empirical support for the central claim. The paper's own Table 1 contains counterexamples: the best model configuration performs worse after finetuning on the full dataset, and several other comparisons are statistically indistinguishable. The reader's identified EMA-collapse assumption is plausible but secondary; even if the EMA target prevents collapse, the transfer gain is not consistently observed. A paired-seed significance test on the strongest configuration would settle whether the negative result is real or statistical noise. If it is real, the conclusion should be revised to 'pretraining helps for certain architectures and dataset sizes,' and the paper should discuss why the best architecture does not benefit. The current CONDITIONAL verdict is appropriate; the condition should include this check. The paper does provide code and a reproducible setup, which is independent support, but the central claim needs qualification before acceptance.","tokens_in":6823,"tokens_out":3978,"duration_ms":38066,"concrete_test":"Recompute the full-dataset comparison for AE-SjT-T with class-attention aggregation using paired seeds: initialize the from-scratch and finetuned models from the same seed, train both on the full Top Tagging dataset, and evaluate 1/εB(εS=0.5) and accuracy over at least 10 seeds. If the finetuned model's rejection remains below the from-scratch baseline, the headline claim fails for the strongest configuration and must be qualified to exclude it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in Section 5, is that finetuning a J-JEPA-pretrained model outperforms training from scratch, 'especially for smaller dataset sizes.' The decisive evidence is Table 1, and it does not consistently support that claim. In the best-performing architecture, AE-SjT-T with class-attention aggregation, finetuning on the full Top Tagging dataset yields 1/εB(εS=0.5)=95.47±1.83 versus 99.38±2.80 from scratch, i.e., the pretrained initialization is worse on the primary metric. The same configuration is also slightly worse on 10% accuracy (88.82±0.11 vs 88.84±0.21). Several apparent gains are within one standard deviation (e.g., AE-SjT-T Flatten 10% accuracy 88.94±0.13 vs 88.92±0.15; AE-SjT-T Flatten full rejection 97.52±1.71 vs 97.79±3.90). Thus the claim is an overgeneralization: benefits are configuration- and metric-dependent, and the largest reported improvements occur in weaker architectures. This is not primarily a question of representation collapse; even granting the EMA mechanism, the empirical headline is not established uniformly.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces J-JEPA, a self-supervised pretraining method for particle jet representations. A large-radius jet is reclustered into subjets; a random 30% of subjets are designated as targets, and a context encoder processes the remaining subjets. A predictor, conditioned on the positions of the target subjets, is trained to predict the target-subjet representations produced by a target encoder that is updated by exponential moving average. The authors pretrain on 1M JetClass top and QCD jets, then finetune on 10% or 100% of the Top Tagging dataset, comparing accuracy and background rejection against training the same architecture from scratch. They report that finetuning a pretrained model outperforms from-scratch training, especially with limited labeled data, and they ablate MLP-based versus attention-based subjet embeddings and flattening versus class-attention aggregation.","tokens_in":7056,"tokens_out":5409,"duration_ms":52272,"significance":"If the central empirical claim held, J-JEPA would provide a useful augmentation-free SSL initialization for jet tagging and a step toward a cross-task foundation model. The paper has several strengths: the code is publicly released, the evaluation is a direct comparison against from-scratch training on standard datasets, and the architecture ablations (embedding type and aggregation method) are informative. The paper is honest about reporting standard deviations and trial counts. However, the headline claim is an empirical performance comparison, and the data in Table 1 do not uniformly support it; the current evidence is mixed and in some configurations favors the from-scratch baseline. The contribution is therefore promising but needs a more careful and statistically disciplined statement of what is demonstrated.","major_comments":[{"comment":"The central claim that finetuning a J-JEPA-pretrained model outperforms training from scratch is not uniformly supported by Table 1. In the strongest architecture, AE-SjT-T with class-attention aggregation, the finetuned full-data rejection is 95.47 ± 1.83 versus 99.38 ± 2.80 from scratch, and the 10% accuracy is 88.82 ± 0.11 versus 88.84 ± 0.21, so the pretrained initialization is slightly worse on both of these entries. Several apparent gains are also within one standard deviation, for example AE-SjT-T Flatten 10% accuracy (88.94 ± 0.13 vs 88.92 ± 0.15) and AE-SjT-T Flatten full-data rejection (97.52 ± 1.71 vs 97.79 ± 3.90). I recommend reporting paired significance tests over the five random seeds and revising the claim to identify the configurations and dataset sizes for which the improvement is statistically meaningful.","section":null},{"comment":"The phrase \"especially for smaller dataset sizes\" is not established by the reported results. For the two SjT-T configurations the full-data rejection gains are at least as large as the 10% gains (SjT-T Flatten: 19.36 at full versus 13.17 at 10%; SjT-T Cls Attn: 11.76 at full versus 8.76 at 10%), while for the AE-SjT-T configurations neither the 10% nor the full-data rejection gains are significant. The paper should either demonstrate a data-size interaction statistically or reframe the conclusion as configuration- and metric-dependent rather than a general small-data benefit.","section":null},{"comment":"The paper states that EMA in the target encoder prevents informational collapse and that \"this principle holds true for J-JEPA,\" but no diagnostic is provided to support that claim. Because the pretraining objective is an L2 regression into an EMA-updated representation space, a representation-collapse check (for example an effective-rank measure, a nearest-neighbor probe, or a linear-probe evaluation of the learned subjet embeddings) would substantially strengthen the mechanism invoked for why the method should work. This is not the primary weakness, but it is directly relevant to interpreting the downstream results.","section":null}],"minor_comments":[{"comment":"The caption contains a typo: \"empbedding\" should be \"embedding.\"","section":null},{"comment":"The text \"the full 3 Top Tagging dataset\" should be \"the full Top Tagging dataset\" with the footnote marker placed after the dataset name; the footnote should clarify whether the 785,767 jets are the training set only or include the validation set.","section":null},{"comment":"The finetuning hyperparameters are not reported separately from the pretraining hyperparameters; please provide the finetuning learning rate, schedule, epoch count, and batch size, since these are needed to reproduce the central comparison.","section":null},{"comment":"The \"SEL\" labels in Figure 1 are not defined in the caption or in the text; please define them or remove the unexplained label.","section":null}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-scale empirical study that extends I-JEPA to jets in a reasonable way and ships public code. The central empirical claim is overstated relative to Table 1, and the manuscript would need a revised conclusion and a proper significance analysis before it meets the bar for this journal. I would not reject outright, because the idea is interesting and the comparison framework is useful, but the current evidence is not sufficient to support the headline claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is a reasonable workshop-level contribution, and the method is genuinely new—nobody has applied a joint-embedding predictive architecture to jets before. The code and data pipeline are available, the setup is reproducible, and the authors compare against a from-scratch baseline with five trials, which is more than many SSL papers do. The attention-based subjet embedding and the physical positional encoding (Δη, Δφ, sin(φ/2)) are sensible ideas worth building on.\n\nThe problem is the headline claim. Section 5 says finetuning outperforms from scratch 'especially for smaller dataset sizes.' Table 1 doesn't uniformly support that. The best model, AE-SjT-T with class attention, is worse after finetuning on full-data rejection: 95.47 ± 1.83 vs 99.38 ± 2.80. The same configuration ties on 10% accuracy. Several gains are within one standard deviation. So the claim is overgeneralized. The strongest improvements are in the weaker architectures (SjT-T Flatten rejection: 53.67 vs 40.50 for 10%; 90.06 vs 70.70 for full). That's still a useful result, but it should be reported as 'pretraining helps most for simpler architectures and on rejection,' not as a uniform win.\n\nThe bigger omission is the lack of any comparison to other jet SSL methods. Masked particle modeling, MAE-style reconstruction, and contrastive pretraining have all been applied to jets. I don't need a full benchmark, but one comparison on the same data would tell me whether the JEPA-style objective is adding anything. Without it, I can only conclude the method works, not that it's better.\n\nThe EMA argument for avoiding collapse is standard from I-JEPA, and the paper correctly cites that. But it's only argued, not checked: no probe of the representation space, no measurement of collapse. Minor, since the downstream results show the representations are usable.\n\nThe citation pattern looks fine; no self-citation circularity. The related work is representative.\n\nWho is this for? Anyone building jet foundation models or doing SSL in HEP. It's a short workshop paper, not a definitive study. I'd accept it for a workshop after revision, but I wouldn't publish it as-is in a journal without adding baselines and fixing the overclaim.","headline":"A credible first JEPA for jets with real code, but the central finetuning claim is shakier than the text admits.","tokens_in":7630,"tokens_out":2195,"would_cite":true,"duration_ms":22985,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"J-JEPA, a jet-based joint embedding predictive architecture that trains without hand-crafted augmentations, learns jet representations whose finetuning outperforms training from scratch on top-quark tagging, especially when labeled data…","keywords":["self-supervised learning","jet tagging","joint embedding predictive architecture","masked subjet prediction","top quark tagging","transformers","particle physics","pretraining"],"falsifier":"Measure the effective rank of the target encoder's output on a held-out sample of jets and compare the rank and the finetuned tagging accuracy for a model trained with EMA versus one trained with the target encoder updated directly from the context encoder; if removing EMA leaves both the representation rank and the finetuned accuracy essentially unchanged, the assertion that EMA is crucial for J-JEPA's success is refuted.","tokens_in":6609,"feed_emoji":"⚛️","tokens_out":12576,"duration_ms":104950,"temperature":0.7,"pith_summary":"This paper introduces J-JEPA, a self-supervised method that trains a transformer to predict the internal representations of randomly masked subjets inside a particle jet from the representations of the remaining context subjets, using the subjets' coordinates as joint information. The goal is to learn jet representations that transfer to downstream tasks without hand-crafted augmentations, which normally build task-specific symmetries into the model. The paper demonstrates that finetuning a J-JEPA-pretrained model on top-quark tagging outperforms training the same architecture from scratch, with the largest gains when only 10% of the labeled training data is used. If this holds, J-JEPA offers a route toward a single pretrained jet representation that can be adapted to multiple physics tasks without per-task augmentation design.","feed_headline":"Jet pretraining beats from-scratch tagging on limited labels","feed_subtitle":"A self-supervised J-JEPA predicts masked subjets from context, and finetuning the result outperforms training from scratch.","key_machinery":"The central object is the J-JEPA architecture itself: a context encoder and a target encoder, both subjet transformers (SjTs), plus a smaller predictor transformer. The context encoder processes only the unmasked subjets; the target encoder processes the full jet but only its outputs on the masked subjets serve as prediction targets; the predictor takes the context representations plus a spatial embedding of each target subjet's position (pseudorapidity and azimuth) and must match the target representations under an L2 loss. The target encoder's parameters are an exponential moving average (EMA) of the context encoder's parameters, which the paper identifies as the mechanism that prevents informational collapse. Two further design choices carry the argument: masking is applied at the encoder outputs rather than the inputs, so both encoders see the full jet, and the predictor includes an information bottleneck to force the context representations to be informative.","core_discovery":"The central claim is that a jet-based joint embedding predictive architecture (J-JEPA) learns useful, symmetry-independent jet representations by predicting the latent representations of randomly selected target subjets from context subjet representations, conditioned on the target subjets' positions. The paper shows that finetuning the target encoder of a J-JEPA-pretrained subjet transformer (SjT) for top-quark tagging outperforms training the same SjT from scratch on identical labeled data, in both accuracy and background rejection at 50% signal efficiency. The advantage is larger when the finetuning set is 10% of the labeled top-quark dataset (about 120k jets) than when the full set (785,767 jets after filtering) is used. This result is reported for two subjet-embedding designs (a plain MLP and one using class-attention blocks) and two ways of aggregating subjet representations (flattening and class attention), with standard deviations from five random initializations.","pith_inferences":["A natural test the paper does not run is whether the same pretrained encoder transfers to other jet tasks, such as quark-gluon discrimination or mass regression, with no change to the pretraining recipe.","The paper borrows the EMA-collapse-prevention argument from the image domain without reporting diagnostics of the learned jet representation space, so measuring its effective rank or nearest-neighbor structure would directly test that borrowed premise.","Because the attention-based subjet embedding outperforms the MLP embedding even from scratch, some of the reported gain may be attributable to the architecture rather than to pretraining; an ablation on a large labeled set with a fixed architecture would separate these factors.","Truly symmetry-independent pretraining does not remove the need for symmetry-specific supervision downstream: tasks requiring exact rotational or translational invariance would still need those invariances imposed during finetuning."],"forward_implications":["Finetuning a J-JEPA-pretrained encoder yields higher top-tagging accuracy and background rejection than the same architecture trained from scratch, with the improvement growing as the labeled sample size shrinks.","Because pretraining uses no data augmentations, the same pretrained checkpoints can be reused for downstream tasks that require different symmetries, without redesigning the pretraining stage.","Masking at the encoder-output level means both encoders see the full jet, so the learned representations are driven by full-jet semantics rather than by structural gaps from input masking.","Scaling pretraining to the full 100-million-jet unlabeled dataset is a direct next step that the paper expects to amplify these finetuning gains."],"supporting_citations":[{"why":"Supplies the joint embedding predictive architecture from computer vision that J-JEPA adapts, including the multi-block masking strategy.","marker":"[8]"},{"why":"Establishes the EMA target encoder as a mechanism that prevents representation collapse, which J-JEPA relies on.","marker":"[22]"},{"why":"Defines the vision transformer architecture on which the subjet transformers are based.","marker":"[23]"},{"why":"Provides the class attention blocks used in the attention-based subjet embedding variant.","marker":"[25]"},{"why":"Defines the anti-kT algorithm used to cluster large-radius jets.","marker":"[17]"},{"why":"Defines the Cambridge-Aachen algorithm used to recluster jets into subjets.","marker":"[18, 19]"},{"why":"Provides the jet-clustering software library used to perform the reclustering.","marker":"[20]"},{"why":"Supplies the large unlabeled jet dataset used for pretraining.","marker":"[26]"},{"why":"Supplies the labeled top-quark dataset used for finetuning and evaluation.","marker":"[27]"}],"fun_headline_variants":["J-JEPA jets: self-supervised pretraining beats from-scratch","Predicting subjets pretrains jets for better tagging with fewer labels","Self-supervised jet model beats supervised training on 10% labels","Jet pretraining without augmentations wins with limited labels","J-JEPA: predict subjets, fine-tune jets, outdo from-scratch"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the exponential-moving-average update to the target encoder prevents representation collapse, so that the L2 prediction loss actually drives the model to learn diverse, physically meaningful subjet embeddings rather than a trivial constant output.","fun_headline_variants_meta":{"raw":{"variants":["J-JEPA jets: self-supervised pretraining beats from-scratch","Predicting subjets pretrains jets for better tagging with fewer labels","Self-supervised jet model beats supervised training on 10% labels","Jet pretraining without augmentations wins with limited labels","J-JEPA: predict subjets, fine-tune jets, outdo from-scratch"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000764,"raw_usage":{"total_tokens":3366,"prompt_tokens":901,"completion_tokens":2465,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":2369}},"tokens_in":517,"tokens_out":2465,"duration_ms":17825,"temperature":1.0,"reasoning_tokens":2369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:22:47.820575+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the effective rank of the target encoder's output on a held-out sample of jets and compare the rank and the finetuned tagging accuracy for a model trained with EMA versus one trained with the target encoder updated directly from the context encoder; if removing EMA leaves both the representation rank and the finetuned accuracy essentially unchanged, the assertion that EMA is crucial for J-JEPA's success is refuted.","supporting_citations":[{"cited_title":"Emerging properties in self-supervised vision transformers","cited_arxiv_id":null,"evidence_quote":"Establishes the EMA target encoder as a mechanism that prevents representation collapse, which J-JEPA relies on."},{"cited_title":"JetClass: A Large-Scale Dataset for Deep Learning in Jet Physics","cited_arxiv_id":null,"evidence_quote":"Supplies the large unlabeled jet dataset used for pretraining."}],"review_version":1}