Pith. sign in

REVIEW 3 major objections 3 minor 39 references

Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that four public vulnerability scoring systems rank the same real-world vulnerabilities substantially differently, and that some of them fail to track actual exploitation risk, based on a dataset of 600 Patch Tuesday vulne

desk verdict The abstract promises a useful four-way comparison of vulnerability scoring systems, but the manuscript body is an unrelated neural-network paper, so there is nothing to evaluate. read the letter →

arxiv 2508.13644 v1 pith:XB2OUHNY submitted 2025-08-19 cs.CR cs.SE

classification cs.CRcs.SE
keywords vulnerabilityscoringCVSSSSVCEPSSExploitabilityIndexPatchTuesdayprioritizationreal-worldexploitationrisk
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to provide the first large-scale, outcome-linked empirical comparison of four publicly available vulnerability scoring systems: CVSS, SSVC, EPSS, and the Exploitability Index. Using 600 real-world vulnerabilities from four months of Microsoft Patch Tuesday disclosures, it examines how the systems rank the same vulnerabilities, how they categorize them across triage tiers, and how well they capture real-world exploitation risk. The reported finding is significant disparity across systems, with implications for security teams that rely on these metrics for data-driven prioritization. The abstract stands alone here because the full text supplied with this record is a different paper; the summary therefore rests on the abstract and the stated discovery.

What carries the argument

The comparison is carried by an outcome-linked dataset: 600 real-world vulnerabilities from four months of Microsoft Patch Tuesday disclosures, each connected to whether it was actually exploited. The four scoring systems—the Common Vulnerability Scoring System (CVSS), the Stakeholder-Specific Vulnerability Categorization (SSVC), the Exploit Prediction Scoring System (EPSS), and the Exploitability Index—are the objects whose rankings and risk assessments are compared against these exploitation outcomes. The outcome link is what allows the paper to test whether the scores capture real-world exploitation risk, not just internal severity.

What would settle it

Re-run the comparison on the same 600 vulnerabilities using independently verified exploitation labels (for example, from vendor incident reports rather than public exploit feeds). If the four scoring systems then agree closely in their rankings, the claim of significant disparity collapses; alternatively, if EPSS's predictive performance on independently verified labels is much worse than on the public-feed labels, the paper's assessment of EPSS is undercut.

Watch

Extended reading notes

Core claim

The central claim is that the four scoring systems diverge markedly when ranking the same vulnerabilities, and that at least some of them do not adequately reflect whether a vulnerability was actually exploited in the wild. The study builds an outcome-linked dataset of 600 vulnerabilities drawn from four months of Microsoft Patch Tuesday disclosures, linking each vulnerability to its real-world exploitation status. The authors report significant disparities in how the systems rank the same vulnerabilities, and they assess the systems' ability to capture real-world exploitation risk. If the paper is right, organizations using a single scoring system to triage vulnerabilities may be making sys

Load-bearing premise

The comparison rests on the premise that each of the 600 vulnerabilities' true exploitation status is known and correctly recorded, and that those labels are not drawn from the same public exploitation feeds that EPSS was trained on.

Editorial extensions

If this is right

  • Security teams should not treat any single severity score as a reliable proxy for real-world exploitation risk.
  • Organizations that triage vulnerabilities using one scoring system may be ordering their remediation priorities differently than they would under another system.
  • The measured disparities motivate more transparent and consistent definitions of exploitability, risk, and severity.
  • The findings suggest that relying on a single metric could lead to systematically misallocated patching effort, and that a consensus or blended approach may be worth exploring.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The full text supplied with this record is a different paper (on generating neural-network weights from text descriptions), not the vulnerability-scoring study described in the abstract; this extraction is therefore based on the abstract and reader-pass notes, and the methods section behind the comparison is unavailable for verification.
  • If EPSS was trained on the same public exploitation feeds used to label the 600 vulnerabilities as exploited or not, then the claimed test of EPSS is partly self-referential and would overstate its agreement with the outcome labels.
  • Real-world exploitation data is incomplete and noisy: many exploited vulnerabilities are never publicly documented, so the outcome labels may be incomplete, which would make every scoring system look worse than it is.
  • A testable extension is to run the same comparison with independently verified exploitation labels (e.g., from vendor incident reports) to see whether the measured disparities persist.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The abstract of arXiv:2508.13644 claims a large-scale, outcome-linked empirical comparison of four vulnerability scoring systems (CVSS, SSVC, EPSS, and the Exploitability Index) on 600 Patch Tuesday vulnerabilities, reporting significant ranking disparities and assessing the systems' ability to capture real-world exploitation risk. However, the full text supplied with the submission is an entirely unrelated manuscript, 'Text2Weight: Bridging Natural Language and Neural Network Weight Spaces' (arXiv:2508.13633), which describes a diffusion transformer for generating neural network weights from text. No section of the provided document describes the vulnerability dataset, the scoring systems, the outcome-label definitions, or any analysis of the claimed study. The abstract is therefore the only in-scope content, and its central claims are unsupported by any methods, data, or results in the submission.

Significance. If the claimed vulnerability-scoring study existed and were methodologically sound, it would be a valuable contribution to vulnerability management, providing independent evidence about the agreement and predictive validity of widely used scoring systems. It could also inform organizations that rely on these scores for risk-based prioritization. However, as submitted, the paper contains none of the supporting content: no dataset, no methodology, no outcome-label definitions, no tables, no statistical analysis, and no reproducibility artifacts. There are no machine-checked proofs, parameter-free derivations, or falsifiable predictions that could be assessed. Because the central empirical claim is entirely unverifiable from the manuscript, the significance of the work cannot be evaluated in its current form.

major comments (3)
  1. [Full text (entire body)] The body of the submission is not the study described in the abstract. The full text is 'Text2Weight,' an unrelated machine-learning paper, and it contains no methods, dataset construction, outcome-label definitions, scoring-system implementations, tables, or figures relevant to vulnerability scoring. The claimed comparison of CVSS, SSVC, EPSS, and the Exploitability Index on 600 Patch Tuesday vulnerabilities is therefore unsupported. This is an internal inconsistency in the submission, not a matter of disagreement with consensus; the central claim cannot be checked at all.
  2. [Abstract] The abstract states that the study is 'outcome-linked' and evaluates the systems' 'ability to capture the real-world exploitation risk.' This premise is load-bearing because EPSS is itself a model trained on exploitation-in-the-wild data. If the outcome labels are drawn from the same public exploitation feeds used to train EPSS, the comparison would be partly self-referential; if the labels are incomplete or noisy, the performance of all systems would be underestimated. Since no methods or data are provided, the independence and accuracy of the outcome labels cannot be assessed. This concern cannot be resolved from the submitted manuscript.
  3. [Abstract] The abstract claims that this is 'the first large-scale, outcome-linked empirical comparison' of the four scoring systems. This novelty claim cannot be verified without a detailed description of the dataset, the vulnerability selection process, the outcome-label sources, and a comparison with prior work. None of these are present in the submission. Without this information, the claim of 'first' is unsubstantiated.
minor comments (3)
  1. [Metadata] The arXiv identifier and title in the abstract do not match the content of the full text, which is marked as arXiv:2508.13633. This appears to be a compilation or submission error that should be corrected.
  2. [Abstract] The abstract refers to 'the Exploitability Index' without defining which index is meant or specifying its version. In a proper methods section this would need to be clarified; here, the absence of any definition compounds the lack of content.
  3. [Full text (body)] Figures, equations, and references in the body all pertain to neural network weight generation, not vulnerability scoring. The document is internally inconsistent as a submission and cannot be reviewed as a coherent manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be established from the submitted material: the full text is an unrelated paper, so there is no derivation chain to reduce.

full rationale

The submission is internally inconsistent: the abstract describes an empirical comparison of vulnerability scoring systems (CVSS, SSVC, EPSS, Exploitability Index) on 600 Patch Tuesday vulnerabilities, but the full text is an entirely different paper, Text2Weight: Bridging Natural Language and Neural Network Weight Spaces. Because no methods, dataset construction, outcome-label definitions, or equations for the vulnerability study are present, there is no derivation chain that could be shown to be circular by construction. The only potential circularity concern raised by the reader's take is that EPSS is trained on public exploitation-in-the-wild data and the study's 'outcome-linked' ground truth might come from the same feeds. That concern is a data-independence validity hazard, not a demonstrated by-construction equivalence: the abstract does not state the source of the exploitation labels, and no supporting text exists in which to identify the specific overlap. Under the hard rule that circularity must be exhibited with quoted text and a specific reduction, no such exhibit is possible here. The absence of the study's methods is a serious completeness/integrity problem, but it is not a circularity finding. Accordingly, the circularity score is 0.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The manuscript body does not contain the described study, so the audit is performed at abstract level. No numeric free parameters are visible; triage-tier thresholds are flagged because the abstract describes categorization across triage tiers and threshold choices would be author-dependent. Two domain assumptions carry the empirical claim: the representativeness of the Patch Tuesday sample, and the accuracy and independence of the exploitation outcome labels. No invented entities are introduced; all four systems are pre-existing.

free parameters (1)
  • triage tier thresholds
    The abstract says the systems are compared across triage tiers. How tier boundaries are set, by each system's specification or by the authors' choice, is not stated; if chosen by the authors they are free parameters. Unverifiable without the methods section.
assumptions (2)
  • domain assumption Four months of Microsoft Patch Tuesday disclosures yields a representative sample of the real-world vulnerability population.
    Stated in the abstract. If Patch Tuesday is skewed by vendor, product mix, or disclosure timing, the comparison generalizes only to that subpopulation.
  • domain assumption Real-world exploitation status for each of the 600 vulnerabilities is accurately known and independent of the scoring systems whose outputs are being tested.
    Required for the 'capture the real-world exploitation risk' conclusion. Exploitation feeds are incomplete and noisy, and EPSS is trained on such feeds, so label quality and independence is the load-bearing premise. The methods that would verify this are absent from the provided text.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems." pith.science (2026). https://pith.science/paper/XB2OUHNY

@misc{pith2026250813644,
  author       = {Pith},
  title        = {Pith review of: Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XB2OUHNY}},
  note         = {Machine review of arXiv:2508.13644}
}
read the original abstract

Accurately assessing software vulnerabilities is essential for effective prioritization and remediation. While various scoring systems exist to support this task, their differing goals, methodologies and outputs often lead to inconsistent prioritization decisions. This work provides the first large-scale, outcome-linked empirical comparison of four publicly available vulnerability scoring systems: the Common Vulnerability Scoring System (CVSS), the Stakeholder-Specific Vulnerability Categorization (SSVC), the Exploit Prediction Scoring System (EPSS), and the Exploitability Index. We use a dataset of 600 real-world vulnerabilities derived from four months of Microsoft's Patch Tuesday disclosures to investigate the relationships between these scores, evaluate how they support vulnerability management task, how these scores categorize vulnerabilities across triage tiers, and assess their ability to capture the real-world exploitation risk. Our findings reveal significant disparities in how scoring systems rank the same vulnerabilities, with implications for organizations relying on these metrics to make data-driven, risk-based decisions. We provide insights into the alignment and divergence of these systems, highlighting the need for more transparent and consistent exploitability, risk, and severity assessments.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

39 extracted references · 24 canonical work pages

  1. [1]

    Samuel K Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2022. Git re-basin: Merging models modulo permutation symmetries. arXiv preprint arXiv:2209.04836 (2022)

  2. [2]

    Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al . 2023. Audiolm: a language modeling approach to audio generation. IEEE/ACM transactions on audio, speech, and language processing 31 (2023), 2523–2533

  3. [3]

    Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 249–256

  4. [4]

    Yifan Gong, Zheng Zhan, Yanyu Li, Yerlan Idelbayev, Andrey Zharkov, Kfir Aberman, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. 2024. Efficient Training with Denoised Neural Weights. In European Conference on Computer Vision . Springer, 18–34

  5. [5]

    Gregory Griffin, Alex Holub, and Pietro Perona. 2007. Caltech-256 object category dataset. (2007)

  6. [6]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision . 1026–1034

  7. [7]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778

  8. [8]

    Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089 (2022)

Show all 39 references
  1. [9]

    Xiaolong Jin, Kai Wang, Dongwen Tang, Wangbo Zhao, Yukun Zhou, Junshu Tang, and Yang You. 2024. Conditional lora parameter generation. arXiv preprint arXiv:2408.01415 (2024)

  2. [10]

    Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2022. Audiogen: Textually guided audio generation. arXiv preprint arXiv:2209.15352 (2022)

  3. [11]

    Alex Krizhevsky, Geoffrey Hinton, et al. 2009. Learning multiple layers of features from tiny images.(2009)

  4. [12]

    Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018. Visualizing the loss landscape of neural nets. Advances in neural information processing systems 31 (2018)

  5. [13]

    Zexi Li, Lingzhi Gao, and Chao Wu. 2024. Text-to-model: Text-conditioned neural network diffusion for train-once-for-all personalization. arXiv preprint arXiv:2405.14132 (2024)

  6. [14]

    Derek Lim, Haggai Maron, Marc T Law, Jonathan Lorraine, and James Lucas

  7. [15]

    mnmoustafa and Mohammed Ali. 2017. Tiny ImageNet. https://kaggle.com/ competitions/tiny-imagenet. Kaggle

  8. [16]

    Elvis Nava, Seijin Kobayashi, Yifei Yin, Robert K Katzschmann, and Benjamin F Grewe. 2022. Meta-learning via classifier (-free) diffusion guidance.arXiv preprint arXiv:2210.08942 (2022)

  9. [17]

    Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron. 2023. Equivariant architectures for learning in deep weight spaces. In International Conference on Machine Learning . PMLR, 25790–25816

  10. [18]

    William Peebles, Ilija Radosavovic, Tim Brooks, Alexei A Efros, and Jitendra Malik

  11. [19]

    William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. In Proceedings of the IEEE/CVF international conference on computer vision . 4195–4205

  12. [20]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...

  13. [21]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695

  14. [22]

    Konstantin Schürholt, Boris Knyazev, Xavier Giró-i Nieto, and Damian Borth

  15. [23]

    Konstantin Schürholt, Dimche Kostadinov, and Damian Borth. 2021. Self- supervised representation learning on neural network weights for model charac- teristic prediction. Advances in Neural Information Processing Systems 34 (2021), 16481–16493

  16. [24]

    Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE interna- tional conference on computer vision . 618–626

  17. [25]

    Advances in Neural Information Processing Systems 35 (2022), 27906–27920

    Hyper-representations as generative models: Sampling unseen neural network weights. Advances in Neural Information Processing Systems 35 (2022), 27906–27920

  18. [26]

    Bowen Tian, Songning Lai, Lujundong Li, Zhihao Shuai, Runwei Guan, Tian Wu, and Yutao Yue. 2025. Pepl: Precision-enhanced pseudo-labeling for fine- grained image classification in semi-supervised learning. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech ...

  19. [27]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)

  20. [28]

    Bedionita Soro, Bruno Andreis, Hayeon Lee, Wonyong Jeong, Song Chong, Frank Hutter, and Sung Ju Hwang. 2024. Diffusion-based neural network weights generation. arXiv preprint arXiv:2402.18153 (2024)

  21. [29]

    Yihang Wang, Bowen Tian, Yueyang Su, Yixing Fan, and Jiafeng Guo. 2025. MDPO: Customized Direct Preference Optimization with a Metric-based Sampler for Question and Answer Generation. In Proceedings of the 31st International Conference on Computational Linguistics . 10660–10671

  22. [30]

    Yujia Wu, Yiming Shi, Jiwei Wei, Chengwei Sun, Yang Yang, and Heng Tao Shen. 2024. Difflora: Generating personalized low-rank adaptation weights with diffusion. arXiv preprint arXiv:2408.06740 (2024)

  23. [31]

    Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Efty- chios Protopapadakis. 2018. Deep learning for computer vision: A brief review. Computational intelligence and neuroscience 2018, 1 (2018), 7068349

  24. [32]

    A photo of

    Allan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace, Yiding Jiang, Samuel Sokota, J Zico Kolter, and Chelsea Finn. 2023. Permutation equivariant neural functionals. Advances in neural information processing systems 36 (2023), 24966– 24992. �� ���� ������� ������ ����� ������...

  25. [34]

    Mixue Xie, Shuang Li, Binhui Xie, Chi Liu, Jian Liang, Zixun Sun, Ke Feng, and Chengwei Zhu. 2024. Weight Diffusion for Future: Learn to Generalize in Non- Stationary Environments. Advances in Neural Information Processing Systems 37 (2024), 6367–6392

  26. [36]

    Layer Selection: Focus on the second residual block’s final convolutional layer (ResNet’s ����������������) to capture mid-level features

  27. [37]

    Gradient Preservation: Enable gradients only for target layer through context management: ��� ����� � ��� �� � ���������������������target� �������������������������������������� ������������������������������������ selective gradient flow (32)

  28. [38]

    Denormalization: Recover original RGB values using dataset statistics: �denorm � � � � � �� � � ������ 0�229 0�224 0�225 ������ � �� ������ 0�485 0�456 0�406 ������ (33)

  29. [39]

    Class Targeting: Compute gradients relative to ground-truth class through ���������������������� wrapper

  30. [2022]

    arXiv preprint arXiv:2209.12892 (2022)

    Learning to learn with generative models of neural network checkpoints. arXiv preprint arXiv:2209.12892 (2022)

  31. [2023]

    arXiv preprint arXiv:2312.04501 (2023)

    Graph metanetworks for processing diverse neural architectures. arXiv preprint arXiv:2312.04501 (2023)

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.