REVIEW 3 major objections 4 minor 33 references
Decorrelation of neural networks from particle lifetimes in the LHCb topological $b$ trigger
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Adding a quadratic moment-decomposition penalty to LHCb's topological b-trigger neural networks flattens their efficiency for long-lived b-hadrons without losing classification power.
desk verdict Useful simulation study of lifetime decorrelation for the LHCb topological b trigger, but the claimed MoDe superiority is not statistically supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The MoDe penalty expands the cumulative distribution of network scores in bins of lifetime into Legendre polynomials and penalises the difference between the actual distribution and its polynomial fit; with order l=2 and monotonicity constraints it allows the sharp efficiency turn-on at small lifetimes while enforcing flatness at large lifetimes. DisCo uses distance correlation between score and lifetime as a penalty. The comparison is made by fitting linear slopes to the efficiency as a function of true lifetime beyond 4 ps.
What would settle it
Take a real Run 3 data sample where the topological b trigger did not fire, e.g., selected by a muon trigger, reconstruct B+ to J/psi K+ candidates, and measure the trigger efficiency as a function of reconstructed B+ decay time for tau>4 ps; if the fitted slope is significantly more negative than the predicted -11 ns^-1, the MoDe-lambda=0.2 claim is falsified.
Extended reading notes
Core claim
The central claim is that the decay-time bias introduced by pile-up-related track misassociation can be removed by adding a quadratic penalty based on Legendre decomposition of the classifier score distribution as a function of lifetime. Measured on simulated B+ to J/psi K+ and B0 to D- pi+ events passing HLT2, the MoDe-penalised models (lambda=0.2) reduce the negative slope of trigger efficiency versus true lifetime for tau>4 ps from about -25 ns^-1 to about -11 ns^-1 (2-body) and from about -22 ns^-1 to about -7 ns^-1 (3-body, consistent with zero), while preserving precision, recall and ROC AUC. The paper concludes MoDe is better suited than DisCo for lifetime decorrelation in the topolog
Load-bearing premise
Everything is demonstrated in simulated events; if the simulation does not reproduce the Run 3 rate of multiple simultaneous collisions and their track-association properties, the observed flattening may not hold for real data.
Editorial extensions
If this is right
- Time-dependent CP and mixing analyses using the topological b trigger will no longer lose statistical sensitivity from a sharp drop in acceptance for b hadrons living beyond about 4 ps.
- The optimal MoDe working point (lambda=0.2) can be adopted for both 2- and 3-body trigger models without retuning, as it preserves precision and ROC AUC.
- The identified PV-misassociation background, prominent at high pile-up, becomes a standard systematic to check in Run 3 trigger and offline selections.
- The linear-slope criterion over tau>4 ps provides a simple benchmark for future trigger decorrelation studies.
Reading between the lines
- If the simulation's pile-up model is representative, the same MoDe recipe should transfer to other LHCb inclusive triggers that use the same MLNN architecture, not only the two- and three-body lines.
- The method's reliance on true lifetime in training means it is best applied when signal simulation is trustworthy; an analogous data-driven version could use reconstructed decay-time sidebands to estimate the lifetime distribution.
- A natural next test is to verify that the efficiency flattening survives when the trigger is run in real Run 3 conditions, where track-association algorithms may behave differently under fluctuating pile-up.
- Since MoDe with l=2 explicitly enforces monotonicity at large lifetimes, it could be combined with the monotonic Lipschitz architecture to give a formal guarantee of no lifetime sculpting beyond a chosen scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses a lifetime-correlation problem in the LHCb topological b trigger, introduced by the higher pile-up conditions of Run 3 (mu = 5.3). The authors identify a background arising from tracks produced in different primary vertices being combined into a candidate, which creates a long-lifetime tail that can bias the trigger efficiency at large decay times. To mitigate this, they train monotonic Lipschitz neural networks for the 2-body and 3-body topological lines using two decorrelation penalties: distance correlation (DisCo) and moment decomposition (MoDe). They compare the resulting trigger efficiencies as a function of true b-hadron lifetime in simulated B+ -> J/psi K+ and B0 -> D- pi+ samples, fit linear slopes above 4 ps for the B0 sample, and conclude that MoDe is better suited for lifetime decorrelation in the topological b trigger.
Significance. If the central claim is established, the paper is practically significant: it identifies a concrete Run 3 trigger limitation and demonstrates two workable mitigation strategies, with potential to preserve large-lifetime acceptance for time-dependent LHCb analyses. The paper is empirically grounded: the authors train and evaluate actual MLNN models in full Pythia/EvtGen/Geant4 simulation, include a no-decorrelation baseline that makes the problem visible, and provide slope fits with uncertainties. The identification of the PV-misassociation background and the demonstration that both penalties reduce the negative large-lifetime slope are useful contributions. However, the headline conclusion that MoDe is superior is currently supported only by point estimates with overlapping uncertainties and an unmatched operating-point comparison; that conclusion needs stronger statistical and methodological support before it is suitable for publication.
major comments (3)
- [Sec. 5, Table 1] The claim that MoDe is "better suited" is not statistically established. For the 2-body models, the fitted slopes are -11.23 +/- 1.6 ns^-1 (MoDe) and -12.9 +/- 4.7 ns^-1 (DisCo); for 3-body, -6.6 +/- 5.5 and -8.8 +/- 1.2 ns^-1. In both cases the uncertainties overlap substantially, and for 3-body the MoDe slope is consistent with both zero and the DisCo value. No test of the slope difference is provided, and the fits are independent rather than joint, so the stated ranking rests on point estimates. Please add a quantitative comparison of the slopes (e.g., a difference with propagated uncertainty, a chi-square/confidence-interval statement, or a test that accounts for the correlation between models) before claiming superiority.
- [Sec. 5, Figs. 3 and 4] Slopes are reported only for B0 -> D- pi+ (Fig. 4 and Table 1), while the conclusion is stated generally for the topological b trigger. The B+ -> J/psi K+ efficiencies in Fig. 3 are shown only as curves, with no fitted slopes or uncertainties, so the reader cannot judge whether the MoDe-vs-DisCo ranking holds for that decay mode. Either fit and report slopes for both decay modes, or explicitly limit the quantitative claim to B0 -> D- pi+ and provide qualitative evidence for the other sample.
- [Sec. 5 and Appendix B] The comparison uses lambda = 0.8 for DisCo and lambda = 0.2 for MoDe, selected from Fig. 5 based on precision/recall/ROC AUC rather than on decorrelation strength. Since the penalty strength controls the degree of decorrelation, differences in slopes could reflect the unmatched lambda values rather than an intrinsic difference between the methods. To support the ranking, the authors should either match the methods by an achieved decorrelation level (e.g., equal dCorr or equal large-lifetime slope) or show that the conclusion is robust to reasonable choices of lambda on both sides. As written, a different lambda for either method could plausibly change the ordering, especially for the 3-body case.
minor comments (4)
- [General] The phrase "Appendix Appendix A" in Sec. 2 should read "Appendix A". Similarly, "Table. 1" in Sec. 5 should be "Table 1". There are also minor grammatical issues, e.g., "the version of MoDe used not penalizing the turn-on curve" should be rephrased.
- [Figs. 3 and 4] The efficiency points are shown without uncertainty bands or error bars. Since Table 1 gives slope uncertainties, the figure would be clearer if the data points included statistical uncertainties, particularly at large lifetime where the event counts decrease.
- [Table 3] The structure of Table 3 is hard to read: the DisCo 3-body row appears to have no penalty strengths listed, and the MoDe 3-body row is not aligned with the 2-body row. Please reformat using separate sub-tables or explicit column entries.
- [Appendix B] The text says the 3-body MoDe model has a maximum precision at lambda = 0.1 and significant degradation beyond, but then declares lambda = 0.2 optimal. This is a defensible choice, but the reasoning would be clearer if the trade-off were quantified, e.g., by stating how much each metric changes between lambda = 0.1 and 0.2.
Circularity Check
No circular derivation: the lifetime-decorrelation comparison is an empirical simulation measurement, and the self-citations to MLNN/MoDe method papers are not load-bearing.
full rationale
The paper's central claim ('We thus deem the MoDe approach to be better suited to achieving lifetime-decorrelation in the topological b trigger') is supported by direct simulation measurements: the MLNNs are trained with two independent penalty terms (DisCo and MoDe), then their efficiencies are evaluated as functions of true b-hadron lifetime in B+→J/ψK+ and B0→D−π+ samples (Figs. 3-4, Table 1). No equation in the paper defines the measured slope in terms of the penalty term by construction; the slope is an empirical output of the trained model. The methods themselves (MLNNs, DisCo, MoDe) are taken as inputs from prior work, some of which has overlapping authors, but the paper does not rely on those citations to prove the decorrelation performance; it evaluates the models directly. There is no fitted parameter renamed as a prediction, no imported uniqueness theorem, no ansatz smuggled via citation that is load-bearing for the conclusion. The noted weaknesses—overlapping slope uncertainties for the MoDe-vs-DisCo comparison, unmatched λ working points, and the reliance on simulation—are statistical or external-validity concerns, not circularity. The derivation chain is therefore self-contained for the claims it makes.
Assumptions & free parameters
free parameters (5)
- DisCo penalty strength lambda (2- and 3-body) =
0.8 for both 2- and 3-body models
- MoDe penalty strength lambda (2- and 3-body) =
0.2 for both 2- and 3-body models
- MoDe Legendre order ell =
2
- Efficiency slope fit range =
tau > 4 ps
- Trigger threshold operating points =
Not specified; set to match output bandwidth
assumptions (4)
- domain assumption The LHCb simulation chain (Pythia, EvtGen, Photos, Geant4) faithfully models Run 3 pileup and the PV-misassociation background.
- domain assumption Decorrelation from the reconstructed/composite lifetime is sufficient to flatten efficiency versus true b-hadron lifetime.
- domain assumption The DisCo and MoDe penalties are correctly implemented according to Refs. [31,32].
- domain assumption The monotonic Lipschitz NN architecture (Refs. [5-7]) trains stably with the additional loss terms.
Cite this review
Pith. "Pith review of Decorrelation of neural networks from particle lifetimes in the LHCb topological $b$ trigger." pith.science (2026). https://pith.science/paper/4QHKXYEY
@misc{pith2026260722281,
author = {Pith},
title = {Pith review of: Decorrelation of neural networks from particle lifetimes in the LHCb topological $b$ trigger},
year = {2026},
howpublished = {\url{https://pith.science/paper/4QHKXYEY}},
note = {Machine review of arXiv:2607.22281}
}
abstract
The LHCb topological beauty trigger is the primary set of algorithms for selecting collision events containing $b$-hadrons in the fully software-based LHCb trigger. The algorithms apply monotonic Lipschitz neural networks (NNs) to select vertices of charged particles consistent with the distinct topology of a $b$ decay, i.e., those with large lifetimes and transverse momentum. Many analyses of the events recorded require that the selection must be unbiased with respect to the $b$-hadron lifetime at large lifetimes. Accurate reconstruction is challenging in busier detector environments, in which several visible proton-proton collisions occur simultaneously per bunch crossing, such that misassociation of decay products can result in vertices with artificially large measured lifetimes. This paper presents two approaches to mitigate correlations between NN scores and candidate lifetimes at large lifetime, and evaluates the performance of the resulting models.
Figures
Reference graph
Works this paper leans on
-
[1]
Alves Jr., et al., JINST3(LHCb-DP-2008-001), S08005 (2008)
A.A. Alves Jr., et al., JINST3(LHCb-DP-2008-001), S08005 (2008). DOI 10.1088/1748-0221/3/08/S08005
-
[2]
M. Williams, V.V. Gligorov, C. Thomas, H. Dijkstra, J. Nardulli, P. Spradlin, The HLT2 Topological Lines. Tech. rep., CERN, Geneva (2011). URLhttps://cds. cern.ch/record/1323557
arXiv 2011
-
[3]
V.V. Gligorov, C. Thomas, M. Williams, The HLT in- clusive B triggers. Tech. rep., CERN, Geneva (2011). URLhttps://cds.cern.ch/record/1384380. LHCb- INT-2011-030
arXiv 2011
-
[4]
Aaij, et al., JINST14(LHCb-DP-2019-001), P04013 (2019)
R. Aaij, et al., JINST14(LHCb-DP-2019-001), P04013 (2019). DOI 10.1088/1748-0221/14/04/P04013
-
[5]
Kitouni, N
O. Kitouni, N. Nolte, M. Williams, Machine Learning: Science and Technology4(3), 035020 (2023). DOI 10. 1088/2632-2153/aced80. URLhttp://dx.doi.org/10. 1088/2632-2153/aced80
2023
-
[6]
N. Schulte, B.R. Delaney, N. Nolte, G.M. Ciezarek, J. Al- brecht, M. Williams. Development of the topological trig- ger for lhcb run 3 (2023). URLhttps://arxiv.org/abs/ 2306.09873 8 Table 2: Input features of the two- and three-body MLNNs of the topologicalbtrigger. The MLNNs were required to increase monotonically in features marked with✓(∼) for all of R...
arXiv 2023
-
[7]
B. Delaney, N. Schulte, G. Ciezarek, N. Nolte, M. Williams, J. Albrecht. Applications of lipschitz neural networks to the run 3 lhcb trigger system (2024). DOI 10.1051/epjconf/202429509005. URLhttps://doi.org/ 10.1051/epjconf/202429509005
arXiv 2024
-
[8]
Aaij, et al., JINST19, P05065 (2024)
R. Aaij, et al., JINST19, P05065 (2024). DOI 10.1088/ 1748-0221/19/05/P05065
2024
Show all 33 references
-
[9]
Aaij, et al., JINST14, P04006 (2019)
R. Aaij, et al., JINST14, P04006 (2019). DOI 10.1088/ 1748-0221/14/04/P04006
2019
-
[10]
Aaij, et al., Comput
R. Aaij, et al., Comput. Softw. Big Sci.4(1), 7 (2020). DOI 10.1007/s41781-020-00039-7
2020 doi
-
[11]
Dujany, B
G. Dujany, B. Storaci, J. Phys. Conf. Ser.664, 082010 (2015). DOI 10.1088/1742-6596/664/8/082010
2015 doi
-
[12]
Mathad, M
A. Mathad, M. Ferrillo, S. Barr´ e, P. Koppenburg, P. Owen, G. Raven, E. Rodrigues, N. Serra, Com- put. Softw. Big Sci.8(1), 6 (2024). DOI 10.1007/ s41781-024-00116-1
2024
-
[13]
Abdelmotteleb, A
A. Abdelmotteleb, A. Bertolin, C. Burr, B. Couturier, E. Eckstein, D. Fazzini, N. Grieser, C. Haen, R. O’Neil, E. Rodrigues, N. Skidmore, M. Smith, A.R. Wiederhold, S. Zhang, Computing and Software for Big Science9(1) (2025). DOI 10.1007/s41781-025-00144-5. URLhttp: //dx.doi.o...
2025 doi
-
[14]
Abe, et al., Phys
K. Abe, et al., Phys. Rev. Lett.80, 660 (1998). DOI 10.1103/PhysRevLett.80.660
1998 doi
-
[15]
Gligorov, M
V.V. Gligorov, M. Williams, JINST8, P02013 (2013). DOI 10.1088/1748-0221/8/02/P02013
2013 doi
-
[16]
Likhomanenko, et al., J
T. Likhomanenko, et al., J. Phys. Conf. Ser.664, 082025 (2015). DOI 10.1088/1742-6596/664/8/082025
2015 doi
-
[17]
CERN-LHCC-2013-021, Geneva (2013)
LHCb collaboration, LHCb VELO Upgrade Technical Design Report. CERN-LHCC-2013-021, Geneva (2013)
2013
-
[18]
CERN-LHCC-2014-001, Geneva (2014)
LHCb collaboration, LHCb Tracker Upgrade Technical Design Report. CERN-LHCC-2014-001, Geneva (2014)
2014
-
[19]
Dziurda, et al., Eur
A. Dziurda, et al., Eur. Phys. J. C85(6), 609 (2025). DOI 10.1140/epjc/s10052-025-14225-7
2025 doi
-
[20]
Paszke, et al., in33rd Conference on Neural Informa- tion Processing Systems(2019)
A. Paszke, et al., in33rd Conference on Neural Informa- tion Processing Systems(2019)
2019
-
[21]
Kingma, J
D.P. Kingma, J. Ba, in3rd International Conference on Learning Representations(2014)
2014
-
[22]
Sj¨ ostrand, S
T. Sj¨ ostrand, S. Mrenna, P. Skands, Comput. Phys. Com- mun.178, 852 (2008). DOI 10.1016/j.cpc.2008.01.036
2008 doi
-
[23]
Sj¨ ostrand, S
T. Sj¨ ostrand, S. Mrenna, P. Skands, JHEP05, 026 (2006). DOI 10.1088/1126-6708/2006/05/026
2006 doi
-
[24]
Belyaev, et al., J
I. Belyaev, et al., J. Phys. Conf. Ser.331, 032047 (2011). DOI 10.1088/1742-6596/331/3/032047
2011 doi
-
[25]
Lange, Nucl
D.J. Lange, Nucl. Instrum. Meth.A462, 152 (2001). DOI 10.1016/S0168-9002(01)00089-4
2001 doi
-
[26]
Davidson, T
N. Davidson, T. Przedzinski, Z. Was, Comp. Phys. Comm.199, 86 (2016). DOI https://doi.org/10.1016/ j.cpc.2015.09.013
2016
-
[27]
Allison, K
J. Allison, K. Amako, J. Apostolakis, H. Araujo, P. Dubois, et al., IEEE Trans.Nucl.Sci.53, 270 (2006). DOI 10.1109/TNS.2006.869826
2006
-
[28]
Agostinelli, et al., Nucl
S. Agostinelli, et al., Nucl. Instrum. Meth.A506, 250 (2003). DOI 10.1016/S0168-9002(03)01368-8
2003 doi
-
[29]
Clemencic, et al., J
M. Clemencic, et al., J. Phys. Conf. Ser.331, 032023 (2011). DOI 10.1088/1742-6596/331/3/032023
2011 doi
-
[30]
Dunietz, J.L
I. Dunietz, J.L. Rosner, Phys. Rev. D34, 1404 (1986). DOI 10.1103/PhysRevD.34.1404
1986 doi
-
[31]
Kasieczka, D
G. Kasieczka, D. Shih, Physical Review Letters125(12) (2020). DOI 10.1103/physrevlett.125.122001. URLhttp: //dx.doi.org/10.1103/PhysRevLett.125.122001 9 0.0 0.2 0.4 0.6 0.8 1.0 Decorrelation strength, 0.80 0.85 0.90 0.95 1.00Precision LHCb Simulation DisCo MoDe 0.0 0.2 0.4 0.6...
2020 doi
-
[32]
Kitouni, B
O. Kitouni, B. Nachman, C. Weisser, M. Williams, Jour- nal of High Energy Physics2021(4) (2021). DOI 10.1007/jhep04(2021)070. URLhttp://dx.doi.org/10. 1007/JHEP04(2021)070
2021 doi
-
[33]
Sz´ ekely, M.L
G.J. Sz´ ekely, M.L. Rizzo, N.K. Bakirov, The An- nals of Statistics35(6), 2769 (2007). DOI 10.1214/ 009053607000000505. URLhttps://doi.org/10.1214/ 009053607000000505
2007
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.