REVIEW 3 major objections 3 minor 39 references
Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that four public vulnerability scoring systems rank the same real-world vulnerabilities substantially differently, and that some of them fail to track actual exploitation risk, based on a dataset of 600 Patch Tuesday vulne
desk verdict The abstract promises a useful four-way comparison of vulnerability scoring systems, but the manuscript body is an unrelated neural-network paper, so there is nothing to evaluate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The comparison is carried by an outcome-linked dataset: 600 real-world vulnerabilities from four months of Microsoft Patch Tuesday disclosures, each connected to whether it was actually exploited. The four scoring systems—the Common Vulnerability Scoring System (CVSS), the Stakeholder-Specific Vulnerability Categorization (SSVC), the Exploit Prediction Scoring System (EPSS), and the Exploitability Index—are the objects whose rankings and risk assessments are compared against these exploitation outcomes. The outcome link is what allows the paper to test whether the scores capture real-world exploitation risk, not just internal severity.
What would settle it
Re-run the comparison on the same 600 vulnerabilities using independently verified exploitation labels (for example, from vendor incident reports rather than public exploit feeds). If the four scoring systems then agree closely in their rankings, the claim of significant disparity collapses; alternatively, if EPSS's predictive performance on independently verified labels is much worse than on the public-feed labels, the paper's assessment of EPSS is undercut.
Extended reading notes
Core claim
The central claim is that the four scoring systems diverge markedly when ranking the same vulnerabilities, and that at least some of them do not adequately reflect whether a vulnerability was actually exploited in the wild. The study builds an outcome-linked dataset of 600 vulnerabilities drawn from four months of Microsoft Patch Tuesday disclosures, linking each vulnerability to its real-world exploitation status. The authors report significant disparities in how the systems rank the same vulnerabilities, and they assess the systems' ability to capture real-world exploitation risk. If the paper is right, organizations using a single scoring system to triage vulnerabilities may be making sys
Load-bearing premise
The comparison rests on the premise that each of the 600 vulnerabilities' true exploitation status is known and correctly recorded, and that those labels are not drawn from the same public exploitation feeds that EPSS was trained on.
Editorial extensions
If this is right
- Security teams should not treat any single severity score as a reliable proxy for real-world exploitation risk.
- Organizations that triage vulnerabilities using one scoring system may be ordering their remediation priorities differently than they would under another system.
- The measured disparities motivate more transparent and consistent definitions of exploitability, risk, and severity.
- The findings suggest that relying on a single metric could lead to systematically misallocated patching effort, and that a consensus or blended approach may be worth exploring.
Reading between the lines
- The full text supplied with this record is a different paper (on generating neural-network weights from text descriptions), not the vulnerability-scoring study described in the abstract; this extraction is therefore based on the abstract and reader-pass notes, and the methods section behind the comparison is unavailable for verification.
- If EPSS was trained on the same public exploitation feeds used to label the 600 vulnerabilities as exploited or not, then the claimed test of EPSS is partly self-referential and would overstate its agreement with the outcome labels.
- Real-world exploitation data is incomplete and noisy: many exploited vulnerabilities are never publicly documented, so the outcome labels may be incomplete, which would make every scoring system look worse than it is.
- A testable extension is to run the same comparison with independently verified exploitation labels (e.g., from vendor incident reports) to see whether the measured disparities persist.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract of arXiv:2508.13644 claims a large-scale, outcome-linked empirical comparison of four vulnerability scoring systems (CVSS, SSVC, EPSS, and the Exploitability Index) on 600 Patch Tuesday vulnerabilities, reporting significant ranking disparities and assessing the systems' ability to capture real-world exploitation risk. However, the full text supplied with the submission is an entirely unrelated manuscript, 'Text2Weight: Bridging Natural Language and Neural Network Weight Spaces' (arXiv:2508.13633), which describes a diffusion transformer for generating neural network weights from text. No section of the provided document describes the vulnerability dataset, the scoring systems, the outcome-label definitions, or any analysis of the claimed study. The abstract is therefore the only in-scope content, and its central claims are unsupported by any methods, data, or results in the submission.
Significance. If the claimed vulnerability-scoring study existed and were methodologically sound, it would be a valuable contribution to vulnerability management, providing independent evidence about the agreement and predictive validity of widely used scoring systems. It could also inform organizations that rely on these scores for risk-based prioritization. However, as submitted, the paper contains none of the supporting content: no dataset, no methodology, no outcome-label definitions, no tables, no statistical analysis, and no reproducibility artifacts. There are no machine-checked proofs, parameter-free derivations, or falsifiable predictions that could be assessed. Because the central empirical claim is entirely unverifiable from the manuscript, the significance of the work cannot be evaluated in its current form.
major comments (3)
- [Full text (entire body)] The body of the submission is not the study described in the abstract. The full text is 'Text2Weight,' an unrelated machine-learning paper, and it contains no methods, dataset construction, outcome-label definitions, scoring-system implementations, tables, or figures relevant to vulnerability scoring. The claimed comparison of CVSS, SSVC, EPSS, and the Exploitability Index on 600 Patch Tuesday vulnerabilities is therefore unsupported. This is an internal inconsistency in the submission, not a matter of disagreement with consensus; the central claim cannot be checked at all.
- [Abstract] The abstract states that the study is 'outcome-linked' and evaluates the systems' 'ability to capture the real-world exploitation risk.' This premise is load-bearing because EPSS is itself a model trained on exploitation-in-the-wild data. If the outcome labels are drawn from the same public exploitation feeds used to train EPSS, the comparison would be partly self-referential; if the labels are incomplete or noisy, the performance of all systems would be underestimated. Since no methods or data are provided, the independence and accuracy of the outcome labels cannot be assessed. This concern cannot be resolved from the submitted manuscript.
- [Abstract] The abstract claims that this is 'the first large-scale, outcome-linked empirical comparison' of the four scoring systems. This novelty claim cannot be verified without a detailed description of the dataset, the vulnerability selection process, the outcome-label sources, and a comparison with prior work. None of these are present in the submission. Without this information, the claim of 'first' is unsubstantiated.
minor comments (3)
- [Metadata] The arXiv identifier and title in the abstract do not match the content of the full text, which is marked as arXiv:2508.13633. This appears to be a compilation or submission error that should be corrected.
- [Abstract] The abstract refers to 'the Exploitability Index' without defining which index is meant or specifying its version. In a proper methods section this would need to be clarified; here, the absence of any definition compounds the lack of content.
- [Full text (body)] Figures, equations, and references in the body all pertain to neural network weight generation, not vulnerability scoring. The document is internally inconsistent as a submission and cannot be reviewed as a coherent manuscript.
Circularity Check
No circularity can be established from the submitted material: the full text is an unrelated paper, so there is no derivation chain to reduce.
full rationale
The submission is internally inconsistent: the abstract describes an empirical comparison of vulnerability scoring systems (CVSS, SSVC, EPSS, Exploitability Index) on 600 Patch Tuesday vulnerabilities, but the full text is an entirely different paper, Text2Weight: Bridging Natural Language and Neural Network Weight Spaces. Because no methods, dataset construction, outcome-label definitions, or equations for the vulnerability study are present, there is no derivation chain that could be shown to be circular by construction. The only potential circularity concern raised by the reader's take is that EPSS is trained on public exploitation-in-the-wild data and the study's 'outcome-linked' ground truth might come from the same feeds. That concern is a data-independence validity hazard, not a demonstrated by-construction equivalence: the abstract does not state the source of the exploitation labels, and no supporting text exists in which to identify the specific overlap. Under the hard rule that circularity must be exhibited with quoted text and a specific reduction, no such exhibit is possible here. The absence of the study's methods is a serious completeness/integrity problem, but it is not a circularity finding. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (1)
- triage tier thresholds
assumptions (2)
- domain assumption Four months of Microsoft Patch Tuesday disclosures yields a representative sample of the real-world vulnerability population.
- domain assumption Real-world exploitation status for each of the 600 vulnerabilities is accurately known and independent of the scoring systems whose outputs are being tested.
Cite this review
Pith. "Pith review of Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems." pith.science (2026). https://pith.science/paper/XB2OUHNY
@misc{pith2026250813644,
author = {Pith},
title = {Pith review of: Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/XB2OUHNY}},
note = {Machine review of arXiv:2508.13644}
}
read the original abstract
Accurately assessing software vulnerabilities is essential for effective prioritization and remediation. While various scoring systems exist to support this task, their differing goals, methodologies and outputs often lead to inconsistent prioritization decisions. This work provides the first large-scale, outcome-linked empirical comparison of four publicly available vulnerability scoring systems: the Common Vulnerability Scoring System (CVSS), the Stakeholder-Specific Vulnerability Categorization (SSVC), the Exploit Prediction Scoring System (EPSS), and the Exploitability Index. We use a dataset of 600 real-world vulnerabilities derived from four months of Microsoft's Patch Tuesday disclosures to investigate the relationships between these scores, evaluate how they support vulnerability management task, how these scores categorize vulnerabilities across triage tiers, and assess their ability to capture the real-world exploitation risk. Our findings reveal significant disparities in how scoring systems rank the same vulnerabilities, with implications for organizations relying on these metrics to make data-driven, risk-based decisions. We provide insights into the alignment and divergence of these systems, highlighting the need for more transparent and consistent exploitability, risk, and severity assessments.
Reference graph
Works this paper leans on
-
[1]
Samuel K Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2022. Git re-basin: Merging models modulo permutation symmetries. arXiv preprint arXiv:2209.04836 (2022)
arXiv 2022
-
[2]
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al . 2023. Audiolm: a language modeling approach to audio generation. IEEE/ACM transactions on audio, speech, and language processing 31 (2023), 2523–2533
work page 2023
-
[3]
Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 249–256
work page 2010
-
[4]
Yifan Gong, Zheng Zhan, Yanyu Li, Yerlan Idelbayev, Andrey Zharkov, Kfir Aberman, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. 2024. Efficient Training with Denoised Neural Weights. In European Conference on Computer Vision . Springer, 18–34
work page 2024
-
[5]
Gregory Griffin, Alex Holub, and Pietro Perona. 2007. Caltech-256 object category dataset. (2007)
work page 2007
-
[6]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision . 1026–1034
2015
-
[7]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[8]
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089 (2022)
arXiv 2022
Show all 39 references
-
[9]
Xiaolong Jin, Kai Wang, Dongwen Tang, Wangbo Zhao, Yukun Zhou, Junshu Tang, and Yang You. 2024. Conditional lora parameter generation. arXiv preprint arXiv:2408.01415 (2024)
2024 arXiv
-
[10]
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2022. Audiogen: Textually guided audio generation. arXiv preprint arXiv:2209.15352 (2022)
2022 arXiv
-
[11]
Alex Krizhevsky, Geoffrey Hinton, et al. 2009. Learning multiple layers of features from tiny images.(2009)
2009
-
[12]
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018. Visualizing the loss landscape of neural nets. Advances in neural information processing systems 31 (2018)
2018
-
[13]
Zexi Li, Lingzhi Gao, and Chao Wu. 2024. Text-to-model: Text-conditioned neural network diffusion for train-once-for-all personalization. arXiv preprint arXiv:2405.14132 (2024)
2024 arXiv
-
[14]
Derek Lim, Haggai Maron, Marc T Law, Jonathan Lorraine, and James Lucas
-
[15]
mnmoustafa and Mohammed Ali. 2017. Tiny ImageNet. https://kaggle.com/ competitions/tiny-imagenet. Kaggle
2017
-
[16]
Elvis Nava, Seijin Kobayashi, Yifei Yin, Robert K Katzschmann, and Benjamin F Grewe. 2022. Meta-learning via classifier (-free) diffusion guidance.arXiv preprint arXiv:2210.08942 (2022)
2022 arXiv
-
[17]
Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron. 2023. Equivariant architectures for learning in deep weight spaces. In International Conference on Machine Learning . PMLR, 25790–25816
2023
-
[18]
William Peebles, Ilija Radosavovic, Tim Brooks, Alexei A Efros, and Jitendra Malik
-
[19]
William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. In Proceedings of the IEEE/CVF international conference on computer vision . 4195–4205
2023
-
[20]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[21]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
2022
-
[22]
Konstantin Schürholt, Boris Knyazev, Xavier Giró-i Nieto, and Damian Borth
-
[23]
Konstantin Schürholt, Dimche Kostadinov, and Damian Borth. 2021. Self- supervised representation learning on neural network weights for model charac- teristic prediction. Advances in Neural Information Processing Systems 34 (2021), 16481–16493
2021
-
[24]
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE interna- tional conference on computer vision . 618–626
2017
-
[25]
Advances in Neural Information Processing Systems 35 (2022), 27906–27920
Hyper-representations as generative models: Sampling unseen neural network weights. Advances in Neural Information Processing Systems 35 (2022), 27906–27920
2022
-
[26]
Bowen Tian, Songning Lai, Lujundong Li, Zhihao Shuai, Runwei Guan, Tian Wu, and Yutao Yue. 2025. Pepl: Precision-enhanced pseudo-labeling for fine- grained image classification in semi-supervised learning. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech ...
2025
-
[27]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[28]
Bedionita Soro, Bruno Andreis, Hayeon Lee, Wonyong Jeong, Song Chong, Frank Hutter, and Sung Ju Hwang. 2024. Diffusion-based neural network weights generation. arXiv preprint arXiv:2402.18153 (2024)
2024 arXiv
-
[29]
Yihang Wang, Bowen Tian, Yueyang Su, Yixing Fan, and Jiafeng Guo. 2025. MDPO: Customized Direct Preference Optimization with a Metric-based Sampler for Question and Answer Generation. In Proceedings of the 31st International Conference on Computational Linguistics . 10660–10671
2025
-
[30]
Yujia Wu, Yiming Shi, Jiwei Wei, Chengwei Sun, Yang Yang, and Heng Tao Shen. 2024. Difflora: Generating personalized low-rank adaptation weights with diffusion. arXiv preprint arXiv:2408.06740 (2024)
2024 arXiv
-
[31]
Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Efty- chios Protopapadakis. 2018. Deep learning for computer vision: A brief review. Computational intelligence and neuroscience 2018, 1 (2018), 7068349
2018
-
[32]
A photo of
Allan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace, Yiding Jiang, Samuel Sokota, J Zico Kolter, and Chelsea Finn. 2023. Permutation equivariant neural functionals. Advances in neural information processing systems 36 (2023), 24966– 24992. �� ���� ������� ������ ����� ������...
2023
-
[34]
Mixue Xie, Shuang Li, Binhui Xie, Chi Liu, Jian Liang, Zixun Sun, Ke Feng, and Chengwei Zhu. 2024. Weight Diffusion for Future: Learn to Generalize in Non- Stationary Environments. Advances in Neural Information Processing Systems 37 (2024), 6367–6392
2024
-
[36]
Layer Selection: Focus on the second residual block’s final convolutional layer (ResNet’s ����������������) to capture mid-level features
-
[37]
Gradient Preservation: Enable gradients only for target layer through context management: ��� ����� � ��� �� � ���������������������target� �������������������������������������� ������������������������������������ selective gradient flow (32)
-
[38]
Denormalization: Recover original RGB values using dataset statistics: �denorm � � � � � �� � � ������ 0�229 0�224 0�225 ������ � �� ������ 0�485 0�456 0�406 ������ (33)
-
[39]
Class Targeting: Compute gradients relative to ground-truth class through ���������������������� wrapper
-
[2022]
arXiv preprint arXiv:2209.12892 (2022)
Learning to learn with generative models of neural network checkpoints. arXiv preprint arXiv:2209.12892 (2022)
2022 arXiv
-
[2023]
arXiv preprint arXiv:2312.04501 (2023)
Graph metanetworks for processing diverse neural architectures. arXiv preprint arXiv:2312.04501 (2023)
2023 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.