{"id":"f1cbb251-b6a4-42cc-bf64-241d03f02bc2","arxiv_id":"2607.29508","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using generative-model likelihood ratios, the authors estimate that modern taggers nearly reach the model-defined optimal limit for W, Z, and H-to-gg jets, while the top-jet gap remains large.","lead":"This note applies a machine-learning framework for estimating the optimal jet-tagging limit to boosted W, Z, and Higgs-to-two-gluons jets, extending a previous study that only covered top jets. It reports that modern taggers are much closer to the model-defined optimum for these processes than for top jets.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generative model fidelity is unvalidated; the reported jet-dependent gap could be an artifact of differing model accuracy across processes.","rationale":"The paper's claim is explicitly about an 'estimated' optimal limit, and it cites the model-dependence caveat. Nevertheless, the headline comparison across processes is presented as a finding about jet tagging, not merely about the models. The load-bearing step is the transfer from model densities to physical densities. The authors do not validate p_S or p_QCD on real JetClass data, and the only evidence is training loss curves that say nothing about density quality. Since the generative model for top may be less accurate (given more complex substructure), the relative gaps may be biased. This is exactly the concern the reader flagged, and I agree no further independent support is provided. A single check—evaluating the LLR on held-out JetClass data—would distinguish a physical effect from a modeling artifact. Until then, the conditional verdict is appropriate; the claim is plausible but not established.","tokens_in":3298,"tokens_out":7128,"duration_ms":78952,"concrete_test":"Compute the LLR of Eq. (1) on held-out JetClass test events (not synthetic) for each process, and compare the resulting ROC/AUC to the LLR evaluated on the synthetic events. If the AUC on real data is substantially lower than on synthetic data, the generative models are not faithful, and the process-dependent gaps in Fig. 1 are not reliable indicators of physical limits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the gap to the estimated optimal limit is small for W/Z/H and large for top—rests on the assumption that the transformer densities p_S and p_QCD used in Eq. (1) are accurate approximations of the true JetClass/event densities, and that their accuracy is comparable across the four processes. The paper provides no quantitative validation of these densities. Fig. 2 shows only training curves for top and H and says nothing about the quality of the learned densities. Crucially, the authors cite Ref. [4] showing the inferred limit depends on the choice and quality of the generative model. Because top jets contain a three-prong structure and a b-quark, the autoregressive model may approximate the top density less accurately than the relatively smooth W/Z/H densities. If so, the top LLR in Eq. (1) is farther from the physical optimal classifier than the W/Z/H LLRs, artificially inflating the top gap relative to the others. The observed process-dependence could therefore be a statement about the generative models' varying expressiveness, not about jet physics. Without a validation of p_S and p_QCD against held-out JetClass events—e.g., a two-sample test or a comparison of LLR performance on real vs. synthetic data—the headline comparison is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This proceedings note extends the framework of Ref. [1] for estimating the optimal jet-tagging limit from generative models to boosted W→qq′, Z→q̄q, and H→gg jets, in addition to the original top-quark case. Autoregressive transformers are trained on JetClass constituent-level data; after discretization, each class model supplies an explicit density p_S, and the log-likelihood ratio log[p_S(x)/p_QCD(x)] in Eq. (1) is taken as the estimated optimal classifier. A baseline transformer classifier (BTC) is trained on the same synthetic samples, and the ROC curves are compared in Fig. 1. The paper reports that the gap between BTC and the estimated optimal LLR is strongly process-dependent: large for top jets, but substantially reduced and in some cases nearly closed for W, Z, and H→gg jets. The authors are careful to call this an 'estimated' limit, cite Ref. [4] on generative-model dependence, and describe ongoing work on validation, scaling, and interpretation.","tokens_in":3528,"tokens_out":2618,"duration_ms":32118,"significance":"If the central empirical claim survives scrutiny, it is a valuable benchmark for the jet-tagging community. It would show that, for several standard tagging tasks, modern transformer-based classifiers are already close to the information-theoretic limit defined by the generative model, while top tagging remains an outlier. The paper's honesty about the model-dependent character of the limit, its explicit citation of Ref. [4], and its extension to four benchmark processes are strengths. The potential impact, however, is limited by the lack of any validation of the learned densities and by the absence of uncertainties on the reported ROC comparison; the headline process dependence could be a property of the generative models rather than of jet physics.","major_comments":[{"comment":"The central claim—that the gap to the optimal limit is small for W/Z/H and large for top—rests entirely on the fidelity of the transformer densities p_S and p_QCD used in Eq. (1) and on the comparability of their accuracy across processes. No quantitative validation of these densities is provided. Training curves in Fig. 2 show only loss values, not whether the learned densities match held-out JetClass events. Since Ref. [4] demonstrated that the inferred limit depends on the choice and quality of the generative model, and since top jets have a more complex three-prong plus b-quark structure, one cannot exclude that the larger top gap is an artifact of a less accurate generative model for top jets. The paper should include a two-sample test or a comparison of classifier performance on real vs. synthetic events, at least for one signal class and QCD, before drawing the cross-process concl","section":"§2–§3, Eq. (1), Fig. 1"},{"comment":"The ROC comparison has no statistical or systematic uncertainties and no numerical performance table. Without error bars, multiple training seeds, or a measure of training variability, the statements that the gap is 'substantially smaller' and 'nearly closed' for W/Z/H are not quantitatively supported. With 10 million synthetic events, statistical fluctuations on the ROC points may be small, but the BTC is trained on a finite sample and the generative models are stochastic; a table with, e.g., background rejection at a few signal efficiencies, including standard deviations over seeds, is necessary to assess whether the observed jet dependence is significant.","section":"Fig. 1, §3"},{"comment":"The interpretation that W/Z/H jets are 'closer to QCD' morphologically and in stochasticity, thereby making the optimum easier to approach, is presented as a plausible explanation but is not tested. If the central claim is established after addressing the density-fidelity concern, this interpretation would benefit from a quantitative measure of closeness, such as a divergence between learned signal and QCD densities or a direct comparison of the LLR distributions. As written, it is a post hoc narrative, not a result.","section":"§3, last paragraph"}],"minor_comments":[{"comment":"The title says 'fundamental limit' while the text carefully says 'estimated optimal limit.' Given the acknowledged model dependence, the title overstates the result; consider 'model-estimated limit' or similar.","section":"Title"},{"comment":"Representative training curves are shown only for top and H→gg. The reader cannot tell whether W and Z models converged similarly. Adding all four or at least mentioning that they are representative would be useful.","section":"Fig. 2"},{"comment":"The discretization into (40,30,30) bins and the approximate 39k-token vocabulary are stated, but no details are given on the transformer architecture, training hyperparameters, or the generation procedure. For a proceedings note this may be acceptable, but citing the full method or an appendix would improve reproducibility.","section":"§2"},{"comment":"Refs. [5–7] are cited as ongoing validation work, but the manuscript does not state whether any of those methods have been applied to the present models. A sentence clarifying the status would help avoid the impression that validation is deferred indefinitely.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a proceedings note that honestly carries the caveats from Ref. [4]. The main concern is not the authors' hedging but the fact that the headline process-dependence claim is not yet supported by any validation of the learned densities. I think this is fixable within the scope of a proceedings contribution: adding a two-sample test for at least one process, plus a small table with uncertainties, would address the main risk. I would not reject, but the current evidence is insufficient to accept as is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a short proceedings note from the RWTH group applying the Ref. [1] framework to W, Z, and H→gg jets. The new empirical result is that the gap between a transformer tagger and the model-defined log-likelihood-ratio classifier is much smaller for these processes than for top jets, nearly closing in some cases. That is a useful observation within the ML-for-HEP subfield, and the paper is honest about its limitations.\n\nThe paper is well written and transparent. It explicitly cites Ref. [4] on model-dependence and says the interpretation requires care. The training curves in Fig. 2 are a small step toward showing the models are not overfitting, and the authors frame the goal as comparison across jet classes within a common framework, not a definitive physical limit.\n\nThe main weakness is the circularity the reader flagged. Eq. (1) builds the \"optimal\" classifier from the same learned densities that generate the synthetic training data for the baseline classifier, so the gap measures how well a classifier approximates the generative model's own likelihood ratio, not necessarily the true physical optimum. The stress-test note is on point: top jets are structurally more complex (three-prong plus b-quark), and the transformer may approximate them less accurately than W/Z/H. If so, the larger top gap could be an artifact of model expressiveness rather than a physics statement. The paper provides no two-sample validation of p_S or p_QCD against held-out JetClass events, no error bars on the ROC curves, and no numerical performance table. These are addressable—and the authors list statistical validation as ongoing work—but they mean the headline process-dependence is not yet established.\n\nThat said, the paper does not overclaim. The abstract says \"estimated optimal limit\" and the text cautions about interpretation. As a proceedings contribution, it is a reasonable progress report.\n\nI would send it to peer review. It is a legitimate extension of a published framework, the claim is testable, and a referee can push for density validation and error estimates. It is not a strong paper, but it is honest and competent. A serious referee would likely request additional validation before publication, but it is not a desk reject.\n\nI would probably bring it to a reading group for the methodological discussion, but not as a headline paper.","headline":"A concise, honest proceedings note extending the generative-model jet-tagging limit to W/Z/H, but the central process-dependence is only as solid as the unvalidated learned densities.","tokens_in":4050,"tokens_out":2265,"would_cite":true,"duration_ms":24357,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper estimates that current jet taggers are nearly optimal for boosted W, Z, and H→gg jets, while a large gap remains for top jets.","keywords":["jet tagging","likelihood-ratio classifier","autoregressive transformers","density estimation","top jets","W/Z/H jets","QCD jets","optimal classifier limit"],"falsifier":"Train a substantially more expressive or better-validated generative model for W, Z, or H→gg jets and recompute the likelihood-ratio ROC curve; if the estimated optimal curve shifts to significantly higher background rejection at fixed signal efficiency, the near-closed gap reported here is an artifact of generative-model insufficiencies. Alternatively, a two-sample statistical test rejecting the learned densities as inconsistent with an independent simulated sample would invalidate the estimated limit.","tokens_in":3135,"feed_emoji":"🎯","tokens_out":11916,"duration_ms":94295,"temperature":0.7,"pith_summary":"Machine-learning jet taggers tell particle physicists what kind of particle produced a spray of high-energy hadrons. This paper asks how close such taggers are to the best classifier that the data allow, using generative models that assign exact probabilities to synthetic jets. Extending a previous top-jet study, it reports that for boosted W, Z, and H→gg jets, the estimated optimal limit is nearly reached, while top-tagging still shows a sizeable gap. The result is process-dependent: jets with less distinctive substructure leave less room for improvement. The authors caution that the estimated limit depends on the quality of the generative models, so the gap is a model-based estimate rather than a definitive bound.","feed_headline":"Top tagging far from optimal; W, Z, Higgs jets nearly there","feed_subtitle":"The gap to the optimal limit is much smaller for W, Z, and Higgs jets than for top jets, showing where gains remain.","key_machinery":"The central object is the estimated log-likelihood ratio log[p_S(x)/p_QCD(x)], where p_S and p_QCD are the probability densities learned by autoregressive transformers, one per jet class, after discretizing each jet's constituents into tokens. Because the synthetic events have exact model likelihoods, this ratio yields the optimal ROC curve for the model-defined problem. The work compares that curve with a baseline transformer classifier sharing the same backbone but with a classification head, on identical synthetic samples, so the only difference between the estimator and the classifier is the loss used to train them.","core_discovery":"Using autoregressive transformers trained on a large simulated jet sample, the authors construct explicit density estimates for top, W, Z, and H→gg jets and for QCD background. From these densities they form the log-likelihood ratio log p_S(x)/p_QCD(x), which by the fundamental lemma of hypothesis testing is the optimal discriminator, and they compare it with a baseline transformer classifier trained on the same synthetic data. Their central observation is that the gap between the baseline tagger and this estimated optimum is strongly jet-type dependent: for W→qq′, Z→qq̄, and H→gg jets the gap is substantially reduced and in some cases nearly closed, whereas for top jets the previously obser","pith_inferences":["The near-closed gap for W/Z/H could be partly an artifact of the generative models under-representing true jet stochasticity for these classes; a stronger test would validate the learned densities against real data with two-sample tests, and if the LLR curve moves upward, the reported near-optimality is model-limited.","A direct extrapolation: if a generative model were trained with more expressive architectures or more data and the W/Z/H LLR curve remained stable, then the current taggers are effectively at the information limit for the constituent-level representation used here, implying only input-level changes (e.g., particle identification, vertexing) could push further.","The same per-class generative-optimum construction could be applied to other signal-background pairs (e.g., tau vs QCD, or quark vs gluon discrimination), giving a general map of where modern taggers sit relative to their theoretical limits.","The interpretation implies a testable asymmetry: for W/Z/H, doubling the generative model capacity should not materially change the estimated optimal ROC, whereas for top jets it should, if the top gap is due to model capacity rather than intrinsic information."],"forward_implications":["For boosted W, Z, and H→gg tagging, current transformer classifiers are already close to the estimated optimal limit, so large algorithmic improvements are not expected for these channels; remaining gains likely require richer input information.","Top-tagging retains a sizeable gap to the estimated optimum, indicating that more headroom exists for future architecture or training improvements in that channel.","The gap estimate is process-dependent, so benchmarking taggers against process-specific estimated limits is more informative than a single global limit.","The framework allows the separation of information-limited tasks (W/Z/H) from model-limited tasks (top) within a common setup.","Using a common generative-model and dataset setup, the ranking of tagger headroom across jet types is a stable qualitative result even if the absolute limit may shift with better generators."],"fun_headline_variants":["W, Z, Higgs jets near optimal tag limit; top jets lag","Optimal jet tagging nearly reached for W, Z, Higgs, not top","Jet tagging gap shrinks for W, Z, Higgs; top still far","Top jets far from optimal; W, Z, Higgs close to limit","Higgs, W, Z jets near optimal; top far below"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The estimated optimal limit is only as reliable as the learned generative densities: the likelihood-ratio classifier is provably optimal for p_S and p_QCD as modeled, but not guaranteed to be optimal for the true physical jet distributions, so the reported gap is a property of the models, not necessarily of nature.","fun_headline_variants_meta":{"raw":{"variants":["W, Z, Higgs jets near optimal tag limit; top jets lag","Optimal jet tagging nearly reached for W, Z, Higgs, not top","Jet tagging gap shrinks for W, Z, Higgs; top still far","Top jets far from optimal; W, Z, Higgs close to limit","Higgs, W, Z jets near optimal; top far below"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000604,"raw_usage":{"total_tokens":2618,"prompt_tokens":672,"completion_tokens":1946,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":416,"completion_tokens_details":{"reasoning_tokens":1849}},"tokens_in":416,"tokens_out":1946,"duration_ms":11976,"temperature":1.0,"reasoning_tokens":1849,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:35:08.859570+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a substantially more expressive or better-validated generative model for W, Z, or H→gg jets and recompute the likelihood-ratio ROC curve; if the estimated optimal curve shifts to significantly higher background rejection at fixed signal efficiency, the near-closed gap reported here is an artifact of generative-model insufficiencies. Alternatively, a two-sample statistical test rejecting the learned densities as inconsistent with an independent simulated sample would invalidate the estimated limit.","supporting_citations":[],"review_version":1}