{"id":"5694e69c-2e69-4128-915e-5ef0b98ad431","arxiv_id":"2506.17127","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Using a transverse momentum cutoff in kT factorization, the authors fix the soft and semihard mean multiplicities in a double negative binomial fit and find that KNO scaling holds only for narrow rapidity windows and that KNO violation is not caused by the rising semihard fraction.","lead":"The authors split particle production in proton collisions into soft and semihard parts using a momentum cutoff in a QCD factorization model, then fit the two parts with two overlapping statistical curves. The fit uses fewer free parameters than earlier attempts and suggests that the rising semihard fraction does not drive the observed breaking of KNO scaling.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (19) equates theoretical multiplicities to full NBD means while the fit explicitly excludes the first few multiplicity bins; the unmodelled low-n bins carry positive mean, so the 4-parameter reduction is not self-consistent and the extracted alpha and k_total trends may be artifacts.","rationale":"The reader's weakest assumption correctly flags parton-hadron duality, the hand-set Lambda, and the unspecified removal of low-multiplicity bins. My pass sharpens the bin-removal issue into a concrete internal inconsistency: even granting parton-hadron duality and the Lambda choice, Eq. (19) is not a valid reduction because it equates theoretical full averages with the full-distribution means of an NBD expression while the fit explicitly truncates away the low-multiplicity bins that carry a positive contribution to the true mean. The values of lambda in Table II (0.851-1.083) confirm that the fitted expression is not normalized over all n, so the missing probability in the excluded bins contributes to the data mean but is absent from Eq. (18). This is not a disagreement with external consensus but a correctness risk in the derivation of the main constraints. The reader's CONDITIONAL verdict is appropriate if the only issues were untested assumptions, but the inconsistent use of truncated-fit parameters in full-distribution constraints means the central quantitative results cannot be trusted as written. A re-fit that handles the truncation correctly, or a demonstration that the excluded bins have negligible mean multiplicity, is required before the 4-parameter reduction and the KNO-scaling conclusions can be accepted. Hence I recommend moving from CONDITIONAL to REJECT pending this check.","tokens_in":17186,"tokens_out":14078,"duration_ms":152818,"concrete_test":"Reproduce Table II by fitting Eq. (15) to the published CMS and ALICE multiplicity distributions using the exact truncated range, but replace Eq. (19) with the correct truncated-mean constraints: Ns = lambda*alpha*sum_{n>=n0} n*Ps(n) / sum_{n>=n0} [alpha*Ps(n)+(1-alpha)*Psh(n)] and analogously for Nsh, and also run the fully unconstrained 6-parameter fit. Then compare the resulting alpha(sqrt(s)) and k_total(sqrt(s)) trends with Fig. 4. If the 'k_total constant only for |eta|<1.0' and 'alpha falls with sqrt(s)' conclusions disappear or change qualitatively, Eq. (19) as written is the load-bearing error.","verdict_should_be":"REJECT","load_bearing_attack":"The central quantitative step is Eq. (19), which identifies the kT-factorization soft and semihard gluon multiplicities Ns and Nsh with lambda*alpha*<n>_s and lambda*(1-alpha)*<n>_sh. This identification presupposes that the mean of the fitted DNBD expression (15), summed over all multiplicities, equals the true data mean N. But the paper states that 'the first few bins can not be reproduced by any NBD... therefore, these bins were removed from the fitting process.' For a truncated fit, Eq. (18) is the mean of the mathematical expression (15) over n=0..infinity, not the mean of the truncated data; the excluded low-multiplicity bins have n>0 and therefore contribute positive multiplicity to the true average. The normalization parameter lambda only adjusts the total probability in the fitted range, and cannot by itself convert the full-distribution mean into the truncated-data mean. Unless the unmodelled low-n mechanism has exactly zero mean multiplicity, setting N = lambda*[alpha*<n>_s + (1-alpha)*<n>_sh] and then Ns = lambda*alpha*<n>_s, Nsh = lambda*(1-alpha)*<n>_sh imposes constraints inconsistent with the truncated likelihood actually fitted. Since Table II shows lambda values from 0.851 to 1.083, the fit is not normalized over all n, and the missing probability resides in the excluded bins. Consequently the reported alpha(sqrt(s)), ks(sqrt(s)), ksh(sqrt(s)) and k_total(sqrt(s)) are obtained by enforcing an equality that does not follow from the data. The claim that Eq. (19) 'legitimately reduces' the number of parameters from 6 to 4 is therefore not established, and the headline conclusions about the physical meaning of n1 and n2 and the KNO-scaling behavior inherit this flaw. This concern is independent of parton-hadron duality; it is an internal consistency issue in the constraint step.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes to separate soft and semihard particle production in pp collisions using the kT factorization formalism with a transverse momentum cutoff Lambda = 1.4 GeV. The computed average multiplicities Ns and Nsh are then inserted as constraints, via Eq. (19), into a double negative binomial (DNBD) fit of LHC multiplicity distributions, reducing the number of free parameters from six to four. From the resulting fits, the authors extract the energy dependence of the soft fraction alpha, the NBD parameters ks and ksh, and the combination ktotal. They conclude that KNO scaling of the total multiplicity distribution holds only for |eta| < 1.0, where ktotal is energy independent, and that alpha decreases with energy in all pseudorapidity windows.","tokens_in":17637,"tokens_out":8487,"duration_ms":86066,"significance":"If the central identification is valid, the paper gives a physical interpretation to the two NBD components of the empirically successful DNBD model and connects parton-level kT-factorization calculations to measured multiplicity distributions. The algebraic derivations in Sections III and IV are internally consistent, and the reported fits are of high quality, with the claim of parameter reduction from six to four being a concrete and falsifiable proposition. However, the main physical conclusions rest on untested assumptions and on a constraint whose consistency with the actual fitting procedure is questionable; the significance of the results is therefore conditional on resolving these issues.","major_comments":[{"comment":"The constraint (19) is derived from Eq. (18), which is the mean of the full DNBD expression summed over all n. However, the fits explicitly remove the first few multiplicity bins, stating that they cannot be reproduced by any NBD. For a truncated fit, the mean of the fitted distribution over the fitted range is not <n> = lambda[alpha<n>_s + (1-alpha)<n>_sh]; the excluded bins carry positive multiplicity, and lambda only rescales the total probability. Table II reports lambda values between 0.851 and 1.083, so the model is not normalized over all n. Consequently, equating N = <n> in Eq. (19) imposes constraints that are inconsistent with the truncated likelihood actually fitted. The extracted alpha, ks, ksh, and ktotal are therefore not determined by a self-consistent procedure, and the central claim that the kT-factorization multiplicities give physical meaning to the NBD components is not established. The manuscript also does not specify which bins are excluded for each data set, so the fits are not reproducible.","section":"Section III, Eq. (19)"},{"comment":"The separation scale Lambda = 1.4 GeV is fixed by hand, with no sensitivity scan. Because Ns and Nsh enter directly into Eq. (19), all extracted parameters and the conclusion that ktotal is constant only for |eta| < 1.0 depend on this choice. The authors cite [37] for the value, but do not demonstrate that the qualitative conclusions are stable under variations of Lambda in a reasonable range (for example, 1.0 to 2.0 GeV). A sensitivity scan is needed to establish that the reported trends are physical rather than artifacts of the cutoff.","section":"Section V, Lambda = 1.4 GeV"},{"comment":"The identification of the computed gluon multiplicities with the mean multiplicities of the two NBD components of final charged hadrons is assumed in Section II ('we shall assume here that the number of produced partons is equal to the number of hadrons'). This assumption is load-bearing for Eq. (19). The cited test [27] concerns hadron multiplicities within jets, not inclusive pp collisions, and no quantitative uncertainty is assigned to the parton-hadron duality assumption. Without a dedicated justification or an estimated hadronization correction, the constraints in Eq. (19) inherit an unquantified systematic error that propagates into all physical conclusions.","section":"Section II, parton-hadron duality"}],"minor_comments":[{"comment":"The phrase 'can be well understood with in perturbative QCD' contains a typo; it should read 'within perturbative QCD'.","section":"Abstract"},{"comment":"In the first sentence, 'one the observables' should be 'one of the observables'.","section":"Section I"},{"comment":"There are typos in the captions: 'collsion' should be 'collision', and in Fig. 4 'refet' should be 'refer'.","section":"Figure 1 and Figure 4 captions"},{"comment":"The reported chi2/dof values are extremely low, with several below 0.1 (e.g., 0.040 and 0.057). The authors should comment on whether the experimental uncertainties include correlated systematics and whether such low values indicate that the uncertainties are overestimated or that the model is overfitting.","section":"Table II and Section V"},{"comment":"The 5.02 and 13 TeV data are INEL > 0 while the other data are NSD, yet the model does not distinguish event classes. The authors note the resulting break in trend, but a more explicit discussion of the possible systematic effect of mixing event classes would be useful.","section":"Section V and Table II"},{"comment":"The formatting of reference [18] is inconsistent with the other references; please standardize it.","section":"Reference [18]"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant problem and the algebraic framework is coherent. The main concern is the consistency of Eq. (19) with the truncated fitting procedure; this is a load-bearing issue that affects the interpretation of all extracted parameters. A sensitivity scan over Lambda and a re-derivation of the constraint using the truncated distribution (or a fit without the constraint) should be requested. The extremely low chi2/dof values also merit scrutiny regarding the treatment of uncertainties."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this paper has a genuinely good idea—use kT factorization to define soft and semihard multiplicities and feed them into a double NBD fit—but the central constraint Eq. (19) is not self-consistent, and the headline trends in alpha and k_total inherit that problem.\n\nWhat is new: earlier DNBD analyses treated the two component means as free parameters. Here they are computed from the GBW unintegrated gluon distribution with a cutoff Lambda = 1.4 GeV, reducing the fit from six to four parameters. That is a real step. The paper is well-written, the kT-factorization derivation is careful, the inputs are standard, and the reported chi2/dof values are low. The authors explicitly state the parton–hadron duality assumption and cite a recent PYTHIA-based check.\n\nThe soft spot is load-bearing. The fits remove the first few multiplicity bins because no NBD can reproduce them. Eq. (15) is then multiplied by lambda, and Eq. (18) gives the mean of that expression over all n. But the data mean N includes the excluded bins, which have positive multiplicity. Setting N = lambda[alpha<n>_s + (1-alpha)<n>_sh] therefore imposes a constraint that the truncated fit does not satisfy. lambda values in Table II are typically 0.85–1.08, so the missing probability is non-negligible and the lost mean is positive. No rescaling of lambda can fix this, because lambda only changes the overall normalization, not the relationship between the means. This concern is independent of parton–hadron duality; it is an internal mismatch between Eq. (19) and the fitting procedure.\n\nOther issues are minor in comparison: Lambda is fixed by hand with no sensitivity scan, the bin-exclusion criterion is never specified, and no code or data are provided, so the numerical fits cannot be checked. The qualitative difference from ALICE's alpha trend is interesting, but both sets of trends are suspect until the normalization question is resolved.\n\nWho benefits: people working on soft QCD phenomenology, KNO scaling, and the interpretation of the two-NBD components. The idea is worth pursuing, but the current paper does not establish its central claims. I would send it to a serious referee, with the explicit request to address the truncated-mean consistency—e.g., by fitting the full distribution including a low-n model, or by showing that the excluded bins contribute negligibly to the mean.","headline":"A good idea—kT factorization to fix the two NBD means—undone by a truncation inconsistency that makes the extracted alpha and k_total trends unreliable.","tokens_in":18174,"tokens_out":5118,"would_cite":false,"duration_ms":49681,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The two components of LHC multiplicity distributions are soft and semihard gluon multiplicities computed from kT factorization, split by a 1.4 GeV momentum cutoff.","keywords":["multiplicity distributions","kT factorization","soft and semihard components","double negative binomial distribution","KNO scaling","unintegrated gluon distribution","proton-proton collisions","parton-hadron duality"],"falsifier":"Repeat the calculation of Ns and Nsh from Eq. (10) with a different unintegrated gluon distribution, or with the cutoff moved from 1.4 GeV to, say, 1.0 or 2.0 GeV, and refit the same multiplicity data. If the extracted trends, particularly alpha falling with energy and ktotal remaining constant only for |eta| < 1.0, shift by more than the quoted uncertainties, then the physical meaning assigned to the DNBD components is not stable enough to support the conclusions.","tokens_in":16973,"feed_emoji":"⚛️","tokens_out":7392,"duration_ms":70828,"temperature":0.7,"pith_summary":"The paper aims to give a physical meaning to the two negative-binomial components, n1 and n2, that have long been used to fit charged-particle multiplicity distributions. Using the kT factorization formalism for gluon production and a cutoff at transverse momentum Lambda = 1.4 GeV, it computes the average soft and semihard multiplicities directly, then feeds these numbers into the double negative binomial fit through Eq. (19). Doing so reduces the fit from six to four free parameters while describing LHC pp data from 0.9 to 13 TeV across several pseudorapidity windows. The key physical conclusions are that the fitted soft-event fraction alpha falls with energy in every window, that KNO scaling holds for the total distribution only when the combined parameter ktotal is energy independent, which happens only for |eta| < 1.0, and that KNO violation is therefore not correlated with alpha. If correct, this connects an empirical two-component fit to a perturbative-QCD calculation and sharpens what the fit parameters mean.","feed_headline":"Gluon-count calculation fixes soft and semihard event split","feed_subtitle":"One transverse-momentum cutoff fixes the two parts of LHC multiplicity curves and cuts fit parameters from six to four.","key_machinery":"The load-bearing object is the kT-factorization gluon production formula (10), with a dipole-model unintegrated gluon distribution supplying the gluon densities, and the separation scale Lambda inserted as a cutoff on the transverse momentum integral that splits the total mean multiplicity into Ns and Nsh. The identity that carries the argument is Eq. (19), which identifies the DNBD combinations lambda alpha <n>s and lambda(1 - alpha)<n>sh with the computed Ns and Nsh; this is what reduces the fit from six to four parameters and ties every extracted trend (alpha, ks, ksh, and ktotal) to a QCD calculation. The combined parameter ktotal, defined in Eq. (25) through the DNBD variance, is the diagnostic for KNO scaling: KNO holds when ktotal is independent of collision energy.","core_discovery":"On the paper's own terms: the mean multiplicities of the two negative binomial components used in double-NBD fits to LHC multiplicity distributions are not free phenomenological inputs. They are the soft and semihard gluon multiplicities obtained from the kT-factorization expression (10), in which the integral over the produced gluon's transverse momentum is cut at Lambda = 1.4 GeV. Equation (19), together with parton-hadron duality, then tells the fit that lambda alpha <n>s = Ns and lambda(1 - alpha)<n>sh = Nsh, where Ns and Nsh are computed. With these constraints, the remaining four parameters (lambda, alpha, ks, ksh) fit the measured distributions, and the extracted energy dependence shows a clear boundary: for |eta| < 1.0 the parameters ks, ksh and the combined ktotal are energy-independent and KNO scaling holds; for wider windows they fall with energy and KNO scaling is violated. The paper concludes that the previously used components n1 and n2 are actually soft and semihard multiplicities, that they approximately satisfy <nsh> ~ 3<ns>, and that the violation of KNO scaling is not tied to the soft-event fraction alpha but to the rapidity-window dependence of Ns, Nsh, ks, and ksh.","pith_inferences":["If the identification is right, the separation scale Lambda becomes a physical parameter: one should be able to locate it independently from the transverse momentum at which final-state hadron production switches from a non-perturbative to a perturbative signature, rather than fixing it at 1.4 GeV by hand.","A direct test would be to repeat the calculation with other unintegrated gluon distributions and see whether the predicted Ns and Nsh, and thus the fitted alpha and ktotal trends, remain within the quoted errors; if the boundary at |eta| ~ 1.0 moves with the choice of gluon distribution, the boundary is not a property of the data.","Because alpha falls with energy while KNO violation appears only in wide windows, one could check the predicted semihard fraction against an independent observable such as the rate of mini-jets or the transverse-momentum spectrum; these should track the computed Nsh/(Ns+Nsh) rather than the fitted alpha."],"forward_implications":["The two components n1 and n2 in double negative binomial fits are identified as soft and semihard mean multiplicities that can be computed from a kT-factorization calculation, giving a physical meaning to a formerly empirical decomposition.","The constrained fit uses four free parameters instead of six and still reproduces the LHC multiplicity distributions from 0.9 TeV to 13 TeV with low chi2/dof values.","The soft-event fraction alpha decreases with collision energy in every pseudorapidity window, so the growth of semihard events alone does not explain KNO violation; the violation instead comes from the energy and window dependence of the soft/semihard mean multiplicities and NBD parameters.","KNO scaling of the total multiplicity holds precisely when ktotal is energy-independent, which the fit finds only for |eta| < 1.0; in those windows ks and ksh are also energy-independent.","The energy trends of ks and ksh do not match any of the three scenarios proposed by earlier two-component studies, indicating a change of dynamics between narrow central windows and wide pseudorapidity windows."],"supporting_citations":[{"why":"Supplies the kT-factorization cross-section formula from which the gluon production expression (10) is obtained.","marker":"[18, 19]"},{"why":"Supplies the dipole-model unintegrated gluon distribution used for the gluon densities in Eq. (10).","marker":"[21, 22]"},{"why":"Provides the approximated form of the kT-factorization integral that reduces it to a product of integrated gluon densities.","marker":"[20]"},{"why":"Introduced the two-component soft/semihard double negative binomial model that the paper constrains.","marker":"[28]"},{"why":"ALICE multiplicity-distribution data sets used in the fits.","marker":"[7, 15, 29]"},{"why":"CMS multiplicity-distribution data used in the fits.","marker":"[39]"},{"why":"Provides a simulation-based check that parton and hadron multiplicities are approximately equal, supporting parton-hadron duality.","marker":"[27]"},{"why":"Defines KNO scaling, whose validity is diagnosed through ktotal.","marker":"[14]"},{"why":"Justifies the choice Lambda = 1.4 GeV as the soft/semihard separation scale.","marker":"[37]"}],"fun_headline_variants":["kT cutoff separates soft and semihard event counts","Gluon kT cut fixes soft-semihard multiplicity split","One gluon scale pins soft and semihard production rates","kT factorization determines soft and semihard means","Soft-semihard events set by gluon transverse momentum cut"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument rests on assuming each computed gluon becomes one final hadron and that the cutoff at 1.4 GeV is the true soft/semihard boundary; if either is wrong, the fitted parameters do not have the physical meaning claimed.","fun_headline_variants_meta":{"raw":{"variants":["kT cutoff separates soft and semihard event counts","Gluon kT cut fixes soft-semihard multiplicity split","One gluon scale pins soft and semihard production rates","kT factorization determines soft and semihard means","Soft-semihard events set by gluon transverse momentum cut"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1604,"prompt_tokens":1045,"completion_tokens":559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":477}},"tokens_in":661,"tokens_out":559,"duration_ms":5582,"temperature":1.0,"reasoning_tokens":477,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:11:31.207163+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the calculation of Ns and Nsh from Eq. (10) with a different unintegrated gluon distribution, or with the cutoff moved from 1.4 GeV to, say, 1.0 or 2.0 GeV, and refit the same multiplicity data. If the extracted trends, particularly alpha falling with energy and ktotal remaining constant only for |eta| < 1.0, shift by more than the quoted uncertainties, then the physical meaning assigned to the DNBD components is not stable enough to support the conclusions.","supporting_citations":[{"cited_title":"Negative binomial multiplicity distribution in proton-proton collisions in limited pseudorapidity intervals at LHC up to sqrt (s) = 7 TeV and the clan model","cited_arxiv_id":"1202.4221","evidence_quote":"Provides the approximated form of the kT-factorization integral that reduces it to a product of integrated gluon densities."},{"cited_title":"Pseudo-rapidity Distributions of Charged Hadrons in pp and pA Collisions at the LHC","cited_arxiv_id":"1305.5106","evidence_quote":"Introduced the two-component soft/semihard double negative binomial model that the paper constrains."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CMS multiplicity-distribution data used in the fits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the choice Lambda = 1.4 GeV as the soft/semihard separation scale."}],"review_version":2}