REVIEW 2 major objections 1 minor 37 references
Hadronic screening masses on the lattice keep showing higher-order and non-perturbative effects all the way up to the electroweak scale.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 22:26 UTC pith:EKHHT5MV
load-bearing objection Abstract promises first non-perturbative screening masses (baryonic + non-static mesonic) up to the electroweak scale, but the supplied full text is an unrelated ML paper, so the lattice claims cannot be checked. the 2 major comments →
Hadronic screening masses in thermal QCD up to the electroweak scale
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Lattice results for hadronic screening masses, including baryonic modes and preliminary non-static mesonic modes, deviate from the leading perturbative predictions of the three-dimensional effective theory by amounts that stay visible up to the electroweak scale, revealing persistent higher-order effects of both perturbative and non-perturbative origin.
What carries the argument
Hadronic screening masses extracted from spatial correlators on the lattice; they quantify the medium’s correlation length and are compared with the systematic expansion of the dimensionally reduced three-dimensional effective theory valid at asymptotically high temperature.
Load-bearing premise
The new lattice strategies that reach electroweak temperatures produce continuum-extrapolated, finite-volume-controlled results whose residual systematics are smaller than the observed deviations from the three-dimensional effective theory.
What would settle it
A continuum extrapolation of the same screening masses at a still higher temperature (or with significantly smaller lattice spacing and larger volume) that collapses onto the next-to-leading-order three-dimensional effective-theory curve within the quoted errors would falsify the claim of persistent higher-order effects.
If this is right
- Perturbative descriptions of the quark-gluon plasma remain incomplete even at electroweak temperatures.
- Baryonic and non-static mesonic screening channels must be retained when modelling the early-universe plasma.
- Dimensionally reduced effective theories need systematic inclusion of non-perturbative matching coefficients up to the electroweak scale.
- Heavy-ion phenomenology at the highest LHC energies still requires non-perturbative input for correlation lengths.
Where Pith is reading between the lines
- The same lattice technology could be used to extract the Debye mass and other transport coefficients at electroweak temperatures, testing whether the same higher-order effects appear.
- If the residual non-perturbative pieces scale as expected from the magnetic sector, they may set a lower bound on the temperature at which a pure three-dimensional description becomes quantitative.
- The persistence of these effects suggests that holographic or other non-perturbative models calibrated at lower temperatures may still be useful near the electroweak scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submitted abstract claims novel lattice QCD results for hadronic screening masses (baryonic and preliminary non-static mesonic) at temperatures from the GeV scale to the electroweak scale, obtained via new theoretical and computational strategies, and compared against the perturbative expansion of the three-dimensional effective theory (EQCD). The comparison is said to reveal persistent higher-order and non-perturbative effects. However, the full manuscript text supplied under this arXiv identifier and title is an unrelated machine-learning paper on Off-Policy Learning with Limited Supply (OPLS) in contextual bandits for recommender systems; it contains no lattice ensembles, continuum extrapolations, screening-mass data, or EQCD comparisons.
Significance. If the abstract’s claims were supported by a matching lattice manuscript with controlled continuum and finite-volume results, the work would be significant for non-perturbative thermal QCD, testing the reach of dimensional reduction up to electroweak temperatures. As supplied, the manuscript has zero significance for that claim: the body is a completely different paper on constrained off-policy learning, so no scientific contribution to lattice QCD or screening masses can be evaluated or credited.
major comments (2)
- The full manuscript text does not match the title, abstract, primary category (hep-lat), or arXiv:2603.18700. It is instead the complete text of an unrelated WWW ’26 paper on Off-Policy Learning with Limited Supply (contextual bandits, coupon allocation, OPLS algorithm). No lattice QCD content, ensembles, continuum extrapolations, finite-volume studies, screening-mass values, or EQCD comparisons appear anywhere in the body. The central claim of the abstract is therefore unsupported by any verifiable material in the submission.
- Because the body contains none of the lattice results invoked in the abstract, it is impossible to assess the load-bearing assumption that residual systematics (continuum, volume, renormalization) are smaller than the reported deviations from the 3D effective theory. The comparison that constitutes the paper’s strongest claim cannot be inspected.
minor comments (1)
- Even the abstract alone is insufficient for a proceedings-style claim of ‘persistent higher-order effects \ldots up to the electroweak scale’ without any numerical tables, figures, or error budgets.
Circularity Check
No circularity detectable; abstract describes an external lattice-vs-perturbation comparison and the supplied full text is an unrelated manuscript.
full rationale
The claimed paper (hadronic screening masses) is presented only via its abstract, which frames a standard non-circular comparison: continuum lattice measurements of screening masses versus the independent high-T expansion of 3D EQCD. No equations, fitted parameters, or uniqueness theorems appear that would force the lattice numbers to equal any input by construction. The CACHEABLE PAPER SOURCE CONTEXT instead contains an entirely different manuscript (Off-Policy Learning with Limited Supply). That unrelated text likewise contains no self-definitional loops, fitted-input-as-prediction steps, or load-bearing self-citation chains that reduce its claims to tautologies; its method and empirical claims are ordinary. Consequently no circular steps can be exhibited by quotation, and the honest finding is score 0.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Continuum QCD at finite temperature is correctly regularized by the lattice ensembles employed, with residual cutoff and volume effects smaller than the reported deviations from the 3D effective theory.
- domain assumption The three-dimensional effective theory and its perturbative expansion become the correct asymptotic description of thermal QCD at large T.
- ad hoc to paper Novel theoretical and computational strategies exist that allow controlled lattice simulations up to the electroweak scale.
read the original abstract
Novel theoretical and computational strategies have opened the possibility of exploring thermal QCD at the non-perturbative level at unprecedented temperatures, reaching from the GeV scale up to the electroweak scale. A number of observable quantities are now being investigated in this regime. Key ones are the hadronic screening masses, which encode the correlation length of the medium and thus the extent to which strong interactions are screened in a thermal environment. In these proceedings we present recent lattice results for hadronic screening masses, including baryonic modes and preliminary non-static mesonic modes. These results can be compared with predictions from the perturbative expansion in the three-dimensional effective theory valid at asymptotically large temperatures. The comparison reveals persistent higher-order effects including those of non-perturbative origin, up to the electroweak scale, shedding new light on the microscopic structure of QCD at extreme temperatures.
Reference graph
Works this paper leans on
-
[1]
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins. 2018. Bandits with knapsacks.Journal of the ACM (JACM)65, 3 (2018), 1–55
2018
-
[2]
Ashwinkumar Badanidiyuru, John Langford, and Aleksandrs Slivkins. 2014. Re- sourceful contextual bandits. InConference on Learning Theory. PMLR, 1109–1134
2014
-
[3]
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson. 2013. Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising.Journal of Machine Learning Research14, 11 (2013)
2013
-
[4]
Konstantina Christakopoulou, Jaya Kawale, and Arindam Banerjee. 2017. Recom- mendation with capacity constraints. InProceedings of the 2017 ACM on Conference on Information and Knowledge Management. 1439–1448
2017
-
[5]
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li. 2014. Doubly Robust Policy Evaluation and Optimization.Statist. Sci.29, 4 (2014), 485–511
2014
-
[6]
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018. More Robust Doubly Robust Off-Policy Evaluation. InProceedings of the 35th Interna- tional Conference on Machine Learning, Vol. 80. PMLR, 1447–1456
2018
-
[7]
Nicolo Felicioni, Maurizio Ferrari Dacrema, Marcello Restelli, and Paolo Cre- monesi. 2022. Off-Policy Evaluation with Deficient Support Using Side Informa- tion.Advances in Neural Information Processing Systems35 (2022)
2022
-
[8]
Chongming Gao, Shijun Li, Wenqiang Lei, Jiawei Chen, Biao Li, Peng Jiang, Xiangnan He, Jiaxin Mao, and Tat-Seng Chua. 2022. KuaiRec: A fully-observed dataset and insights for evaluating recommender systems. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 540–550
2022
-
[9]
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018. Offline A/B Testing for Recommender Systems. InPro- ceedings of the 11th ACM International Conference on Web Search and Data Mining. 198–206
2018
-
[10]
Nathan Kallus, Yuta Saito, and Masatoshi Uehara. 2021. Optimal Off-Policy Evaluation from Multiple Logging Policies. InProceedings of the 38th International Conference on Machine Learning, Vol. 139. PMLR, 5247–5256
2021
-
[11]
Haruka Kiyohara, Ren Kishimoto, Kosuke Kawakami, Ken Kobayashi, Kazuhide Nakata, and Yuta Saito. 2024. Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation. InInternational Conference on Learning Repre- sentations
2024
-
[12]
Haruka Kiyohara, Masahiro Nomura, and Yuta Saito. 2024. Off-policy evaluation of slate bandit policies via optimizing abstraction. InProceedings of the ACM on Web Conference 2024. 3150–3161
2024
-
[13]
Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro, Yusuke Narita, Nobuyuki Shimizu, and Yasuo Yamamoto. 2022. Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model. InProceedings of the 15th International Conference on Web Search and Data Mining
2022
-
[14]
Haruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu, Yasuo Yamamoto, and Yuta Saito. 2023. Off-Policy Evaluation of Ranking Policies under Diverse User Behavior. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1154–1163
2023
-
[15]
Lihong Li, Shunbao Chen, Jim Kleban, and Ankur Gupta. 2015. Counterfactual estimation and optimization of click metrics in search engines: A case study. In Proceedings of the 24th International Conference on World Wide Web. 929–934
2015
-
[16]
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou. 2018. Breaking the curse of horizon: infinite-horizon off-policy estimation. InProceedings of the 32nd International Conference on Neural Information Processing Systems. 5361–5371
2018
-
[17]
Alberto Maria Metelli, Alessio Russo, and Marcello Restelli. 2021. Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning. Advances in Neural Information Processing Systems34 (2021)
2021
-
[18]
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011. Scikit-learn: Machine learning in Python.Journal of Machine Learning Re...
2011
-
[19]
Yuta Saito, Himan Abdollahpouri, Jesse Anderton, Ben Carterette, and Mounia Lalmas. 2024. Long-term Off-Policy Evaluation and Learning. InProceedings of the ACM on Web Conference 2024. 3432–3443
2024
-
[20]
Yuta Saito, Shunsuke Aihara, Megumi Matsutani, and Yusuke Narita. 2021. Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2021
-
[21]
Yuta Saito and Thorsten Joachims. 2021. Counterfactual Learning and Evaluation for Recommender Systems: Foundations, Implementations, and Recent Advances. InProceedings of the 15th ACM Conference on Recommender Systems. 828–830
2021
-
[22]
Yuta Saito and Thorsten Joachims. 2022. Off-Policy Evaluation for Large Action Spaces via Embeddings. InInternational Conference on Machine Learning. PMLR, 19089–19122
2022
-
[23]
Yuta Saito, Qingyang Ren, and Thorsten Joachims. 2023. Off-Policy Evaluation for Large Action Spaces via Conjunct Effect Modeling.arXiv preprint arXiv:2305.08062 (2023)
Pith/arXiv arXiv 2023
-
[24]
Yuta Saito, Takuma Udagawa, Haruka Kiyohara, Kazuki Mogi, Yusuke Narita, and Kei Tateno. 2021. Evaluating the Robustness of Off-Policy Evaluation. In Proceedings of the 15th ACM Conference on Recommender Systems. 114–123
2021
-
[25]
Yuta Saito, Jihan Yao, and Thorsten Joachims. 2024. POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition.arXiv preprint arXiv:2402.06151(2024)
Pith/arXiv arXiv 2024
-
[26]
Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, and Miroslav Dudík. 2020. Doubly Robust Off-Policy Evaluation with Shrinkage. InProceedings of the 37th International Conference on Machine Learning, Vol. 119. PMLR, 9167–9176
2020
-
[27]
Yi Su, Lequn Wang, Michele Santacatterina, and Thorsten Joachims. 2019. Cab: Continuous adaptive blending for policy evaluation and learning. InInternational Conference on Machine Learning, Vol. 84. 6005–6014
2019
-
[28]
Adith Swaminathan and Thorsten Joachims. 2015. Batch learning from logged bandit feedback through counterfactual risk minimization.The Journal of Machine Learning Research16, 1 (2015), 1731–1755
2015
-
[29]
Adith Swaminathan and Thorsten Joachims. 2015. Counterfactual risk mini- mization: Learning from logged bandit feedback. InInternational Conference on Machine Learning. PMLR, 814–823
2015
-
[30]
Adith Swaminathan and Thorsten Joachims. 2015. The Self-Normalized Estimator for Counterfactual Learning.Advances in Neural Information Processing Systems 28 (2015)
2015
-
[31]
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni. 2017. Off-Policy Evaluation for Slate Recommendation. InAdvances in Neural Information Processing Systems, Vol. 30. 3632–3642
2017
-
[32]
Masatoshi Uehara, Chengchun Shi, and Nathan Kallus. 2022. A review of off- policy evaluation in reinforcement learning.arXiv preprint arXiv:2212.06355 (2022)
Pith/arXiv arXiv 2022
-
[33]
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık. 2017. Optimal and adaptive off-policy evaluation in contextual bandits. InInternational Conference on Machine Learning. PMLR, 3589–3597
2017
-
[34]
Huasen Wu, Rayadurgam Srikant, Xin Liu, and Chong Jiang. 2015. Algorithms with logarithmic or sublinear regret for constrained contextual bandits.Advances in Neural Information Processing Systems28 (2015)
2015
-
[35]
Yafei Zhao and Long Yang. 2024. Constrained contextual bandit algorithm for limited-budget recommendation system.Engineering Applications of Artificial Intelligence128 (2024), 107558
2024
-
[36]
Wenliang Zhong, Rong Jin, Cheng Yang, Xiaowei Yan, Qi Zhang, and Qiang Li
-
[37]
InProceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Stock constrained recommendation in tmall. InProceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2287–2296
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.