Pith. sign in

REVIEW 1 cited by

Towards a better labeling process for network security datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.01337 v1 pith:LTJ4TMHN submitted 2023-05-02 cs.CR

classification cs.CR
keywords datasetslabelslabelnetworkassignmentlabelingontologyprocess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most network security datasets do not have comprehensive label assignment criteria, hindering the evaluation of the datasets, the training of models, the results obtained, the comparison with other methods, and the evaluation in real-life scenarios. There is no labeling ontology nor tools to help assign the labels, resulting in most analyzed datasets assigning labels in files or directory names. This paper addresses the problem of having a better labeling process by (i) reviewing the needs of stakeholders of the datasets, from creators to model users, (ii) presenting a new ontology of label assignment, (iii) presenting a new tool for assigning structured labels for Zeek network flows based on the ontology, and (iv) studying the differences between generating labels and consuming labels in real-life scenarios. We conclude that a process for structured label assignment is paramount for advancing research in network security and that the new ontology-based label assignation rules should be published as an artifact of every dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Labeling NIDS Rules with MITRE ATT&CK Techniques: Machine Learning vs. Large Language Models

    cs.CR 2024-12 conditional novelty 4.0 of 10

    On a dataset of 973 Snort rules, traditional ML models (SVM, F1 up to 0.87) outperformed ChatGPT, Claude, and Gemini (best F1 0.62) at labeling rules with MITRE ATT&CK techniques.

Pith tools