REVIEW 4 major objections 4 minor 112 references
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A deep learning sequence tagger fine-tuned on only 100 labeled log lines correctly identifies 99.5% of sensitive attributes in software logs, outperforming regular-expression pipelines.
desk verdict Useful benchmark and a credible regex-vs-NER comparison, but the RQ3 headline result likely overstates generalization because fine-tuning and test sets share templates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is sequence tagging in IOB format over whitespace-tokenized log lines, built on CodeBERT, a pre-trained transformer for code and natural language, with a single classification head. A hierarchical two-stage architecture carries the argument: the main SDLog model detects and labels eight coarse sensitive categories, and the second SDLog Net model fine-tunes the same backbone on the network-related subset to resolve IP addresses, ports, and host names individually. Treating network tokens as one coarse 'net' category lets the model handle compound tokens like IP:port that whitespace tokenization would otherwise split incorrectly.
What would settle it
Have an independent team re-annotate a random sample of the benchmark lines under the same protocol and compare their labels to the published ground truth; if agreement on the full token set drops well below the reported 0.87 kappa, or if SDLog's F1 on the independently labeled sample falls far below 98.4%, the central comparison against regex is not trustworthy. Alternatively, apply the 100-line fine-tuning protocol to an unseen log corpus with third-party expert labels and check whether the F1 remains above the low 90s.
Extended reading notes
Core claim
The central claim is that sensitive-attribute detection in logs should be treated as a token-level named entity recognition problem rather than a pattern-matching problem. SDLog fine-tunes a pre-trained code-aware language model with a softmax classification head over IOB labels (B-, I-, O), so each word in a log message is classified as non-sensitive, as the beginning of a sensitive attribute, or as a continuation of one. A two-level hierarchy handles ambiguity: a main model assigns coarse categories including a 'net' class that absorbs concatenated network tokens such as IP:port pairs, and a lightweight second model splits net spans into IP address, port, and host name. The paper reports that this design generalizes across 16 datasets and that the regex baseline's performance is unstable, affected by pattern choice and execution order, while SDLog's is not.
Load-bearing premise
The results stand or fall on the manual annotation of the 32,000 log lines: only 200 of the 1,363 templates were double-coded, reaching a kappa of 0.87, and a single coder labeled the rest, so if that one coder's notion of sensitivity drifted, every reported precision and recall number is measured against a skewed target.
Editorial extensions
If this is right
- Anonymization pipelines can use one fine-tuned model in place of a hand-tuned stack of regex rules, removing the need to tune pattern order for each dataset.
- Attributes that lack rigid formats, such as IDs, usernames, configuration details, and host names, become detectable, closing gaps that current regex tooling leaves open.
- A 100-line fine-tuning set, which the paper estimates takes about 2 to 3 days to label, is enough to adapt the detector to a new environment, with fine-tuning taking minutes on a consumer-grade GPU.
- Because the model is released as a pre-trained component, integration cost is comparable to calling a standard language model API.
- The two-stage coarse-to-fine design suggests that the same approach can be extended to other privacy-sensitive log attributes that are not represented in the current benchmark.
Reading between the lines
- If the benchmark's sensitivity definition becomes the de facto standard, the same coarse-to-fine tagging recipe could be applied to other privacy targets such as API keys, geolocation coordinates, or personal names in chat and issue-tracker logs.
- The reported regex order-sensitivity implies that many existing production anonymization pipelines may be silently under-performing, since reordering the same regex stack can change what gets masked; auditing that risk is a natural operational follow-up.
- The single-coder annotation of most templates is the point to probe: an independent re-annotation study would either confirm the 98.4% target or reveal label drift, and that would be more informative than another model comparison.
- A direct extension would be to test whether the 100-sample fine-tuning recipe transfers to logs from enterprise or government environments, which the paper explicitly leaves outside its evaluation scope.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SDLog, a CodeBERT-based sequence-labeling framework for detecting sensitive attributes in software logs, and evaluates it against regular-expression baselines. RQ1 systematically collects 41 regex patterns from literature and industry and measures their per-attribute performance on a 32,000-line annotated LogHub benchmark. RQ2 compares SDLog, trained with leave-one-dataset-out cross-validation, against the best regex pipeline and reports an overall F1 of 92.9%. RQ3 fine-tunes SDLog on 20, 50, or 100 labeled logs from each target dataset and reports that 100 samples suffice to reach 99.5% recall and 98.4% F1. The authors claim this is the first deep learning alternative to regex-based log anonymization and release models, annotations, and scripts.
Significance. If the results hold, the paper makes a useful contribution: it provides a sizable annotated log corpus for sensitivity detection, a systematic picture of regex variability, and a reproducible CodeBERT-based pipeline with a plausible deployment story. The release of models, data, and fine-tuning scripts is a clear strength. However, the two headline claims — that SDLog 'significantly outperforms' regex and that 100 fine-tuning samples give near-perfect detection — rest on evaluation choices that need to be fixed or substantially re-analyzed before the numbers can be taken at face value.
major comments (4)
- [§4.3.2, Table 8] The RQ3 fine-tuning protocol is not template-disjoint, so the headline '99.5% recall / 98.4% F1' likely measures template memorization rather than generalization to new log formats. Each LogHub subset contains 2,000 logs drawn from few templates (Table 2: Apache has 6 templates, HDFS 14, Proxifier 8), and Section 4.3.2 fine-tunes on 'the first 100 logs that contain at least one sensitive attribute' before testing on the remaining 1,900 logs. In low-template datasets those 100 logs can cover the entire template space, making the test set mostly repetitions of the same templates with different parameter values. Please re-run RQ3 with a template-disjoint split (e.g., hold out whole templates) and report results separately for seen and unseen templates; the current numbers cannot be read as evidence of deployment-ready generalization.
- [§4.2.3, Tables 6 and 4] The claim that 'SDLog not only outperforms regex across all categories' is contradicted by the URL row of Table 6, where SDLog has 0.0% precision, recall, and F1 on 128 URL instances, while the best regex pipeline in Table 4 reaches 95.1% F1 for URLs. The text's explanation about limited labeled examples is plausible, but the summary sentence and the RQ2 conclusion need to be revised to state that SDLog outperforms regex on most categories while remaining worse on URL detection, or the URL detector needs to be improved before the overall 'significantly outperforms' claim can stand.
- [§3.3 and §6] The ground-truth labels are the single target for every regex evaluation, SDLog training, and all reported F1 scores, but inter-rater reliability was measured only on 200 of the 1,363 templates (κ=0.87), with the remaining templates annotated by a single coder. As the threats-to-validity section acknowledges, this leaves room for subjective drift in what counts as sensitive. Please add a concrete validity check for the single-coder portion — for example, a second annotation pass on a random sample of the remaining templates, or a template-level agreement report — and discuss how label bias would affect the regex-vs-SDLog comparison.
- [§4.1.2.5] The regex evaluation counts a true positive when the regex-matched text 'overlaps' the labeled sensitive attribute, whereas SDLog is scored as a token-level NER system. If 'overlap' includes partial substring matches, the two methods are not scored under the same criterion and the RQ1/RQ2 comparison is not apples-to-apples. Please specify the exact matching rule (token-level exact match, span-level exact match, or partial overlap) and recompute the baseline with the same rule used for SDLog.
minor comments (4)
- [Headings] The headings 'Restuls' in Sections 4.1.3, 4.2.3, and 4.3.3 should be 'Results', and the Section 3.1 title 'Sensitivite Attributes' should be 'Sensitive Attributes'.
- [Tables 3 and 4] The port regex is printed as '(\d1,5)' in Table 3 and Table 4; it should be '(\d{1,5})' to be a valid regular expression.
- [Table 5] The HealthApp row reports a support of 1 with F1 0.0; the caption or surrounding text should clarify that this corresponds to one false positive, since the dataset contains no sensitive attributes.
- [Global] The tables consistently misspell 'Precision' as 'Percision'; please correct throughout.
Circularity Check
No significant circularity: the regex-vs-SDLog comparison is empirical and held-out; the sensitivity taxonomy is an adopted input, not a derived result.
full rationale
SDLog's central comparison is empirical rather than derivational. The regex baselines are collected from 185 papers and three industry partners, not derived from the SDLog model; SDLog is evaluated with leave-one-out cross-validation where each target dataset is held out during training, and the reported F1-score is computed from a unified confusion matrix over held-out predictions. The sensitivity taxonomy is adopted from the authors' own prior work (Aghili et al. [2]), and the ground-truth labels were produced by the same authors, but this is an input assumption and a measurement-validity concern, not a case where a predicted quantity reduces by construction to a fitted parameter or to the definition. RQ3's fine-tuning experiment trains on 20-100 target logs and tests on the remaining 1,900 logs; although the split is not template-disjoint, making the 99.5% recall figure vulnerable to template memorization, that is a generalization-threat issue, not a circularity of the kind where Eq. X equals Eq. Y by definition. The paper itself acknowledges this risk in Section 6: 'There is also a risk of overfitting, particularly when SDLog is fine-tuned on small datasets, potentially causing it to learn dataset-specific artifacts.' No load-bearing step reduces to a self-citation: the cited prior taxonomy is itself grounded in literature review, survey, and regulations, and no uniqueness theorem is invoked. Therefore no significant circularity is present.
Assumptions & free parameters
free parameters (4)
- fine-tuning sample size (headline result) =
100 logs
- learning rate =
5e-5
- training epochs =
2
- weight decay =
0.01
assumptions (5)
- domain assumption CodeBERT's pretrained representations transfer to software log tokens
- domain assumption Ground-truth labels on 32,000 log lines are correct
- domain assumption Whitespace tokenization is an adequate basis for sensitive span detection
- domain assumption LogHub datasets are representative of real-world logs
- ad hoc to paper The taxonomy of sensitive attributes from Aghili et al. [2] is accepted without independent validation
invented entities (1)
-
'net' parent label category
Cite this review
Pith. "Pith review of SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs." pith.science (2026). https://pith.science/paper/AVFIUODV
@misc{pith2026250514976,
author = {Pith},
title = {Pith review of: SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs},
year = {2026},
howpublished = {\url{https://pith.science/paper/AVFIUODV}},
note = {Machine review of arXiv:2505.14976}
}
read the original abstract
Software logs are messages recorded during the execution of a software system that provide crucial run-time information about events and activities. Although software logs have a critical role in software maintenance and operation tasks, publicly accessible log datasets remain limited, hindering advance in log analysis research and practices. The presence of sensitive information, particularly Personally Identifiable Information (PII) and quasi-identifiers, introduces serious privacy and re-identification risks, discouraging the publishing and sharing of real-world logs. In practice, log anonymization techniques primarily rely on regular expression patterns, which involve manually crafting rules to identify and replace sensitive information. However, these regex-based approaches suffer from significant limitations, such as extensive manual efforts and poor generalizability across diverse log formats and datasets. To mitigate these limitations, we introduce SDLog, a deep learning-based framework designed to identify sensitive information in software logs. Our results show that SDLog overcomes regex limitations and outperforms the best-performing regex patterns in identifying sensitive information. With only 100 fine-tuning samples from the target dataset, SDLog can correctly identify 99.5% of sensitive attributes and achieves an F1-score of 98.4%. To the best of our knowledge, this is the first deep learning alternative to regex-based methods in software log anonymization.
Figures
Reference graph
Works this paper leans on
-
[1]
Roozbeh Aghili, Heng Li, and Foutse Khomh. 2023. Studying the characteristics of AIOps projects on GitHub.Empirical Software Engineering28, 6 (2023), 143
2023
-
[2]
Roozbeh Aghili, Heng Li, and Foutse Khomh. 2025. Protecting Privacy in Software Logs: What Should Be Anonymized?Proceedings of the ACM on Software Engineering2, FSE (2025)
2025
-
[3]
Roozbeh Aghili, Qiaolin Qin, Heng Li, and Foutse Khomh. 2024. Understanding Web Application Workloads and Their Applications: Systematic Literature Review and Characterization. In2024 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 474–486
2024
-
[4]
Laith Alzubaidi, Jinglan Zhang, Amjad J Humaidi, Ayad Al-Dujaili, Ye Duan, Omran Al-Shamma, José Santamaría, Mohammed A Fadhel, Muthana Al-Amidie, and Laith Farhan. 2021. Review of deep learning: concepts, CNN architectures, challenges, applications, future directions.Journal of big Data8 (2021), 1–74
2021
-
[5]
Chinatsu Aone. 1999. A trainable summarizer with knowledge acquired from robust NLP techniques.Advances in automatic text summarization (1999), 71–80
1999
-
[6]
Ron Artstein and Massimo Poesio. 2008. Inter-coder agreement for computational linguistics.Computational linguistics34, 4 (2008), 555–596
2008
-
[7]
Prasasthy Balasubramanian, Justin Seby, and Panos Kostakos. 2023. Transformer-based llms in cybersecurity: An in-depth study on log anomaly detection and conversational defense mechanisms. In2023 IEEE International Conference on Big Data (BigData). IEEE, 3590–3599
2023
-
[8]
Titus Barik, Robert DeLine, Steven Drucker, and Danyel Fisher. 2016. The bones of the system: A case study of logging and telemetry at microsoft. InProceedings of the 38th International Conference on Software Engineering Companion. 92–101
2016
Show all 112 references
-
[9]
Victor R Basili, Richard W Selby, and David H Hutchens. 1986. Experimentation in software engineering.IEEE Transactions on software engineering 7 (1986), 733–743
1986
-
[10]
Mohamed Amine Batoun, Mohammed Sayagh, Roozbeh Aghili, Ali Ouni, and Heng Li. 2024. A literature review and existing challenges on software logging practices: From the creation to the analysis of software logs.Empirical Software Engineering29, 4 (2024), 103
2024
-
[11]
Daniel M Bikel, Scott Miller, Richard Schwartz, and Ralph Weischedel. 1998. Nymble: a high-performance learning name-finder.arXiv preprint cmp-lg/9803003(1998)
1998 arXiv
-
[12]
Jasmin Bogatinovski, Sasho Nedelkoski, Alexander Acker, Florian Schmidt, Thorsten Wittkopp, Soeren Becker, Jorge Cardoso, and Odej Kao. 2021. Artificial intelligence for it operations (aiops) workshop white paper.arXiv preprint arXiv:2101.06054(2021)
2021 arXiv
-
[13]
Andrew Borthwick, John Sterling, Eugene Agichtein, and Ralph Grishman. 1998. Description of the MENE Named Entity System as used in MUC-7. InProceedings of the Seventh Message Understanding Conference (MUC-7), Fairfax, Virginia, April 29-May 1, 1998
1998
-
[14]
Tønnes Brekne and André Årnes. 2005. Circumventing IP-address pseudonymization.. InCommunications and Computer Networks. 43–48
2005
-
[15]
Tønnes Brekne, André Årnes, and Arne Øslebø. 2005. Anonymization of ip traffic monitoring data: Attacks on two prefix-preserving anonymization schemes and some proposed remedies. InInternational Workshop on Privacy Enhancing Technologies. Springer, 179–196
2005
-
[16]
An Ran Chen, Tse-Hsun Chen, and Shaowei Wang. 2021. Demystifying the challenges and benefits of analyzing user-reported logs in bug reports. Empirical Software Engineering26 (2021), 1–30
2021
-
[17]
Marcello Cinque, Domenico Cotroneo, and Antonio Pecchia. 2012. Event logs for the analysis of software failures: A rule-based approach.IEEE Transactions on Software Engineering39, 6 (2012), 806–821
2012
-
[18]
Jacob Cohen. 1960. A coefficient of agreement for nominal scales.Educational and psychological measurement20, 1 (1960), 37–46
1960
-
[19]
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. (2011)
2011
-
[20]
Eli Cortez, Anand Bonde, Alexandre Muzio, Mark Russinovich, Marcus Fontoura, and Ricardo Bianchini. 2017. Resource central: Understanding and predicting workloads for improved resource management in large cloud platforms. InProceedings of the 26th Symposium on Operating System...
2017
-
[21]
Hetong Dai, Yiming Tang, Heng Li, and Weiyi Shang. 2023. PILAR: Studying and mitigating the influence of configurations on log parsing. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 818–829
2023
-
[22]
Biplob Debnath, Mohiuddin Solaimani, Muhammad Ali Gulzar Gulzar, Nipun Arora, Cristian Lumezanu, Jianwu Xu, Bo Zong, Hui Zhang, Guofei Jiang, and Latifur Khan. 2018. LogLens: A real-time log analysis system. In2018 IEEE 38th international conference on distributed computing sy...
2018
-
[23]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805(2018)
2018 arXiv
-
[24]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human ...
2019
-
[25]
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al. 2020. Codebert: A pre-trained model for programming and natural languages.arXiv preprint arXiv:2002.08155(2020)
2020 arXiv
-
[26]
Ulrich Flegel. 2002. Pseudonymizing Unix log files. InInternational Conference on Infrastructure Security. Springer, 162–179. Manuscript submitted to ACM 26 Roozbeh Aghili, Xingfang Wu, Foutse Khomh, and Heng Li
2002
-
[27]
Michalis Foukarakis, Demetres Antoniades, Spiros Antonatos, and Evangelos P Markatos. 2007. Flexible and high-performance anonymization of NetFlow records using anontool. In2007 Third International Conference on Security and Privacy in Communications Networks and the Workshops...
2007
-
[28]
Michael Foukarakis, Demetres Antoniades, and Michalis Polychronakis. 2009. Deep packet anonymization. InProceedings of the Second European Workshop on System Security. 16–21
2009
-
[29]
Tadayoshi Fushiki. 2011. Estimation of prediction error by using K-fold cross-validation.Statistics and Computing21 (2011), 137–146
2011
-
[30]
Th Gamer, Chr Mayer, and Marcus Schöller. 2008. Pktanon–a generic framework for profile-based traffic anonymization. (2008)
2008
-
[31]
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. 2016. LSTM: A search space odyssey.IEEE transactions on neural networks and learning systems28, 10 (2016), 2222–2232
2016
-
[32]
Xiaodan Gu and Kai Dong. 2023. PD-PAn: Prefix-and Distribution-Preserving Internet of Things Traffic Anonymization.Electronics12, 20 (2023), 4369
2023
-
[33]
Chunjing Han, Kunkun Sun, Haina Tang, Yulei Wu, and Xiaodan Zhang. 2020. AFT-Anon: A scaling method for online trace anonymization based on anonymous flow tables. In2020 IEEE Symposium on Computers and Communications (ISCC). IEEE, 1–7
2020
-
[34]
Pinjia He, Jieming Zhu, Shilin He, Jian Li, and Michael R Lyu. 2017. Towards automated log parsing for large-scale log data analysis.IEEE Transactions on Dependable and Secure Computing15, 6 (2017), 931–944
2017
-
[35]
Pinjia He, Jieming Zhu, Zibin Zheng, and Michael R Lyu. 2017. Drain: An online log parsing approach with fixed depth tree. In2017 IEEE international conference on web services (ICWS). IEEE, 33–40
2017
-
[36]
Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. 2021. Cosqa: 20,000+ web queries for code search and question answering.arXiv preprint arXiv:2105.13239(2021)
2021 arXiv
-
[37]
Shaohan Huang, Yi Liu, Carol Fung, Rong He, Yining Zhao, Hailong Yang, and Zhongzhi Luan. 2020. Paddy: An event log parsing approach using dynamic dictionary. InNOMS 2020-2020 IEEE/IFIP network operations and management symposium. IEEE, 1–8
2020
-
[38]
Zhiheng Huang, Wei Xu, and Kai Yu. 2015. Bidirectional LSTM-CRF models for sequence tagging.arXiv preprint arXiv:1508.01991(2015)
2015 arXiv
-
[39]
Hugging Face. 2025. Token Classification with Transformers. https://huggingface.co/docs/transformers/tasks/token_classification Accessed on May 5, 2025
2025
-
[40]
Alibaba Inc. 2025. Alibaba Traces. https://github.com/alibaba/clusterdata Accessed on May 5, 2025
2025
-
[41]
Google Inc. 2025. Google Traces. https://github.com/google/cluster-data Accessed on May 5, 2025
2025
-
[42]
Microsoft Inc. 2025. Azure Traces. https://github.com/Azure/AzurePublicDataset Accessed on May 5, 2025
2025
-
[43]
Zhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li, Junjie Huang, Yintong Huo, Pinjia He, Jiazhen Gu, and Michael R Lyu. 2024. Lilac: Log parsing using llms with adaptive parsing cache.Proceedings of the ACM on Software Engineering1, FSE (2024), 137–160
2024
-
[44]
Zhihan Jiang, Jinyang Liu, Junjie Huang, Yichen Li, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Jieming Zhu, and Michael R Lyu. 2024. A large-scale evaluation for log parsing techniques: How far are we?. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Te...
2024
-
[45]
Zhen Ming Jiang, Ahmed E Hassan, Parminder Flora, and Gilbert Hamann. 2008. Abstracting execution logs to execution events for enterprise applications (short paper). In2008 The Eighth International Conference on Quality Software. IEEE, 181–186
2008
-
[46]
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. InProceedings of the AAAI conference on artificial intelligence, Vol. 34-05. 8018–8025
2020
-
[47]
Łukasz Korzeniowski and Krzysztof Goczyła. 2022. Landscape of automated log analysis: A systematic literature review and mapping study.IEEE Access10 (2022), 21892–21913
2022
-
[48]
Dimitris Koukis, Spyros Antonatos, Demetres Antoniades, Evangelos P Markatos, and Panagiotis Trimintzios. 2006. A generic anonymization framework for network traffic. In2006 IEEE International Conference on Communications, Vol. 5. IEEE, 2302–2309
2006
-
[49]
Lawrence Berkeley National Laboratory. 1998. The Internet Traffic Archive. https://ita.ee.lbl.gov/html/traces.html Accessed on May 5, 2025
1998
-
[50]
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations.arXiv preprint arXiv:1909.11942(2019)
2019 arXiv
-
[51]
Max Landauer, Markus Wurzenberger, Florian Skopik, Giuseppe Settanni, and Peter Filzmoser. 2018. Dynamic log file analysis: An unsupervised cluster evolution approach for anomaly detection.computers & security79 (2018), 94–116
2018
-
[52]
Fengcun Li and Bo Hu. 2019. Deepjs: Job scheduling based on deep reinforcement learning in cloud data center. InProceedings of the 4th International Conference on Big Data and Computing. 48–53
2019
-
[53]
Heng Li, Weiyi Shang, Bram Adams, Mohammed Sayagh, and Ahmed E Hassan. 2020. A qualitative study of the benefits and costs of logging from developers’ perspectives.IEEE Transactions on Software Engineering47, 12 (2020), 2858–2873
2020
-
[54]
Jing Li, Aixin Sun, Jianglei Han, and Chenliang Li. 2020. A survey on deep learning for named entity recognition.IEEE transactions on knowledge and data engineering34, 1 (2020), 50–70
2020
-
[55]
Xiaoyun Li, Pengfei Chen, Linxiao Jing, Zilong He, and Guangba Yu. 2022. SwissLog: Robust anomaly detection and localization for interleaved unstructured logs.IEEE Transactions on Dependable and Secure Computing20, 4 (2022), 2762–2780
2022
-
[56]
Yufei Li, Yanchi Liu, Haoyu Wang, Zhengzhang Chen, Wei Cheng, Yuncong Chen, Wenchao Yu, Haifeng Chen, and Cong Liu. 2023. Glad: Content-aware dynamic graphs for log anomaly detection. In2023 IEEE International Conference on Knowledge Graph (ICKG). IEEE, 9–18. Manuscript submit...
2023
-
[57]
Zhijing Li, Qiuai Fu, Zhijun Huang, Jianbo Yu, Yiqian Li, Yuanhao Lai, and Yuchi Ma. 2024. Revisiting Log Parsing: The Present, the Future, and the Uncertainties.IEEE Transactions on Reliability73, 3 (2024), 1459–1472
2024
-
[58]
Zhenhao Li, Chuan Luo, Tse-Hsun Chen, Weiyi Shang, Shilin He, Qingwei Lin, and Dongmei Zhang. 2023. Did we miss something important? studying and exploring variable-aware log abstraction. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 830–842
2023
-
[59]
Ying-Dar Lin, Po-Ching Lin, Sheng-Hao Wang, I-Wei Chen, and Yuan-Cheng Lai. 2014. Pcaplib: A system of extracting, classifying, and anonymizing real packet traces.IEEE Systems Journal10, 2 (2014), 520–531
2014
-
[60]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692(2019)
2019 arXiv
-
[61]
Liam Daly Manocchio, Siamak Layeghy, David Gwynne, and Marius Portmann. 2024. A configurable anonymisation approach for network flow data: Balancing utility and privacy.Computers and Electrical Engineering118 (2024), 109465
2024
-
[62]
Ehsan Mashhadi and Hadi Hemmati. 2021. Applying codebert for automated program repair of java simple bugs. In2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR). IEEE, 505–509
2021
-
[63]
Andrew McCallum and Wei Li. 2003. Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons. InProceedings of the seventh conference on Natural language learning at HLT-NAACL 2003. 188–191
2003
-
[64]
Mary L McHugh. 2012. Interrater reliability: the kappa statistic.Biochemia medica22, 3 (2012), 276–282
2012
-
[65]
Frank McSherry and Ratul Mahajan. 2010. Differentially-private network trace analysis.ACM SIGCOMM Computer Communication Review40, 4 (2010), 123–134
2010
-
[66]
Salma Messaoudi, Annibale Panichella, Domenico Bianculli, Lionel Briand, and Raimondas Sasnauskas. 2018. A search-based approach for accurate identification of log message formats. InProceedings of the 26th Conference on Program Comprehension. 167–177
2018
-
[67]
Haibo Mi, Huaimin Wang, Yangfan Zhou, Michael Rung-Tsong Lyu, and Hua Cai. 2013. Toward fine-grained, unsupervised, scalable performance diagnosis for production cloud computing systems.IEEE Transactions on Parallel and Distributed Systems24, 6 (2013), 1245–1255
2013
-
[68]
Greg Minshall. 2005. TCPDPRIV. https://ita.ee.lbl.gov/html/contrib/tcpdpriv.html Accessed on May 5, 2025
2005
-
[69]
Diego Mollá, Menno Van Zaanen, and Steve Cassidy. 2007. Named entity recognition in question answering of speech data. InProceedings of the 2007 Australasian Language Technology Workshop. ALTA, 57–65
2007
-
[70]
David Nadeau and Satoshi Sekine. 2007. A survey of named entity recognition and classification.Lingvisticae Investigationes30, 1 (2007), 3–26
2007
-
[71]
Karthik Nagaraj, Charles Killian, and Jennifer Neville. 2012. Structured comparative analysis of systems logs to diagnose performance problems. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12). 353–366
2012
-
[72]
Radoslaw Naumiuk and Jarosław Legierski. 2014. Anonymization of data sets from Service Delivery Platforms. In2014 Federated Conference on Computer Science and Information Systems. IEEE, 955–960
2014
-
[73]
Weina Niu, Zimu Li, Zhaoxu He, Aduo Wang, Beibei Li, and Xiaosong Zhang. 2023. FSMFLog: Discovering Anomalous Logs Combining Full Semantic Information and Multifeature Fusion.IEEE Internet of Things Journal11, 3 (2023), 4442–4453
2023
-
[74]
National Congress of Brazil. 2018. General Data Protection Law. https://www.gov.br/anpd/pt-br/centrais-de-conteudo/outros-documentos-e- publicacoes-institucionais/lgpd-en-lei-no-13-709-capa.pdf Accessed on May 5, 2025
2018
-
[75]
State of California
U.S. State of California. 2018. California Consumer Privacy Act. https://oag.ca.gov/privacy/ccpa Accessed on May 5, 2025
2018
-
[76]
Privacy Commissioner of Canada. 2000. Personal Information Protection and Electronic Documents Act. https://www.priv.gc.ca/en/privacy- topics/privacy-laws-in-canada/the-personal-information-protection-and-electronic-documents-act-pipeda/ Accessed on May 5, 2025
2000
-
[77]
Department of Health and Human Services
U.S. Department of Health and Human Services. 1996. Health Insurance Portability and Accountability Act. https://www.hhs.gov/hipaa/index.html Accessed on May 5, 2025
1996
-
[78]
National People’s Congress of the People’s Republic of China. 2021. Data Security Law. https://digichina.stanford.edu/work/translation-data- security-law-of-the-peoples-republic-of-china/ Accessed on May 5, 2025
2021
-
[79]
National People’s Congress of the People’s Republic of China. 2021. Personal Information Protection Law. https://personalinformationprotectionlaw. com/ Accessed on May 5, 2025
2021
-
[80]
Momen Oqaily, Mohammad Ekramul Kabir, Suryadipta Majumdar, Yosr Jarraya, Mengyuan Zhang, Makan Pourzandi, Lingyu Wang, and Mourad Debbabi. 2023. iCAT+: An Interactive Customizable Anonymization Tool Using Automated Translation Through Deep Learning.IEEE Transactions on Dependa...
2023
-
[81]
A Paszke. 2019. Pytorch: An imperative style, high-performance deep learning library.arXiv preprint arXiv:1912.01703(2019)
2019 arXiv
-
[82]
Divya Pathak, Mudit Verma, Aishwariya Chakraborty, and Harshit Kumar. 2024. Self Adjusting Log Observability for Cloud Native Applications. In2024 IEEE 17th International Conference on Cloud Computing (CLOUD). IEEE, 482–493
2024
-
[83]
Dave Plonka. 2003. ip2anonip. https://pages.cs.wisc.edu/%7Eplonka/ip2anonip/ Accessed on May 5, 2025
2003
-
[84]
Qiaolin Qin, Roozbeh Aghili, Heng Li, and Ettore Merlo. 2024. Preprocessing is All You Need: Boosting the Performance of Log Parsers With a General Preprocessing Framework.arXiv preprint arXiv:2412.05254(2024)
2024 arXiv
-
[85]
Lance A Ramshaw and Mitchell P Marcus. 1999. Text chunking using transformation-based learning. InNatural language processing using very large corpora. Springer, 157–176
1999
-
[86]
Piotr Ryciak, Katarzyna Wasielewska, and Artur Janicki. 2022. Anomaly detection in log files using selected natural language processing methods. Applied Sciences12, 10 (2022), 5089. Manuscript submitted to ACM 28 Roozbeh Aghili, Xingfang Wu, Foutse Khomh, and Heng Li
2022
-
[87]
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.arXiv preprint arXiv:1910.01108(2019)
2019 arXiv
-
[88]
Shiwen Shan, Yintong Huo, Yuxin Su, Yichen Li, Dan Li, and Zibin Zheng. 2024. Face it yourselves: An llm-based two-stage strategy to localize configuration errors via logs. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 13–25
2024
-
[89]
Adam J Slagell, Kiran Lakkaraju, and Katherine Luo. 2006. FLAIM: A Multi-level Anonymization Framework for Computer and Network Logs.. In LISA, Vol. 6. 3–8
2006
-
[90]
Leszek Sliwko. 2024. Cluster Workload Allocation: A Predictive Approach Leveraging Machine Learning Efficiency.IEEE Access(2024)
2024
-
[91]
Cong Sun, Zhihao Yang, Lei Wang, Yin Zhang, Hongfei Lin, and Jian Wang. 2021. Biomedical named entity recognition using BERT in the machine reading comprehension framework.Journal of Biomedical Informatics118 (2021), 103799
2021
-
[92]
Latanya Sweeney. 2002. k-anonymity: A model for protecting privacy.International journal of uncertainty, fuzziness and knowledge-based systems 10, 05 (2002), 557–570
2002
-
[93]
Jeniya Tabassum, Mounica Maddela, Wei Xu, and Alan Ritter. 2020. Code and Named Entity Recognition in StackOverflow. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). https://www.aclweb.org/anthology/2020.acl-main.443/
2020
-
[94]
Adrian-Ioan Tuns and Adrian Spătaru. 2023. Cloud Service Failure Prediction on Google’s Borg Cluster Traces Using Traditional Machine Learning. In2023 25th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC). IEEE, 162–169
2023
-
[95]
European Union. 2022. General Data Protection Regulation. https://gdpr-info.eu/ Accessed on May 5, 2025
2022
-
[96]
Artur Varanda, Leonel Santos, Rogério Luís de C Costa, Adail Oliveira, and Carlos Rabadão. 2021. Log pseudonymization: Privacy maintenance in practice.Journal of Information Security and Applications63 (2021), 103021
2021
-
[97]
Abhishek Verma, Luis Pedrosa, Madhukar Korupolu, David Oppenheimer, Eric Tune, and John Wilkes. 2015. Large-scale cluster management at Google with Borg. InProceedings of the tenth european conference on computer systems. 1–17
2015
-
[98]
Paul Voigt and Axel Von dem Bussche. 2017. The eu general data protection regulation (gdpr).A practical guide, 1st ed., Cham: Springer International Publishing10, 3152676 (2017), 10–5555
2017
-
[99]
Zahin Wahab, Sadif Ahmed, Md Nafiu Rahman, Rifat Shahriyar, and Gias Uddin. 2024. Secret Breach Prevention in Software Issue Reports.arXiv preprint arXiv:2410.23657(2024)
2024 arXiv
-
[100]
Jinyuan Wang, Tong Li, Runzi Zhang, Zifang Tang, Di Wu, and Zhen Yang. 2024. VCRLog: Variable Contents Relationship Perception for Log-based Anomaly Detection. In2024 IEEE 35th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 156–167
2024
-
[101]
2012.Experimentation in software engineering
Claes Wohlin, Per Runeson, Martin Höst, Magnus C Ohlsson, Björn Regnell, Anders Wesslén, et al. 2012.Experimentation in software engineering. Vol. 236. Springer
2012
-
[102]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019. Huggingface’s transformers: State-of-the-art natural language processing.arXiv preprint arXiv:1910.03771(2019)
2019 arXiv
-
[103]
Tzu-Tsung Wong. 2015. Performance evaluation of classification algorithms by k-fold and leave-one-out cross validation.Pattern recognition48, 9 (2015), 2839–2846
2015
-
[104]
Antonios Xenakis, Sabrina Mamtaz Nourin, Zhiyuan Chen, George Karabatis, Ahmed Aleroud, and Jhancy Amarsingh. 2023. A self-adaptive and secure approach to share network trace data.Digital Threats: Research and Practice4, 4 (2023), 1–20
2023
-
[105]
Jun Xu, Jinliang Fan, Mostafa H Ammar, and Sue B Moon. 2002. Prefix-preserving ip address anonymization: Measurement-based security evaluation and a new cryptography-based scheme. In10th IEEE International Conference on Network Protocols, 2002. Proceedings.IEEE, 280–289
2002
-
[106]
Junjielong Xu, Qiuai Fu, Zhouruixing Zhu, Yutong Cheng, Zhijing Li, Yuchi Ma, and Pinjia He. 2023. Hue: A user-adaptive parser for hybrid logs. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineerin...
2023
-
[107]
Mukesh Yadav and Dhirendra S Mishra. 2023. Identification of network threats using live log stream analysis. In2023 2nd International Conference on Paradigm Shifts in Communications Embedded Systems, Machine Learning and Signal Processing (PCEMS). IEEE, 1–6
2023
-
[108]
Siyu Yu, Pinjia He, Ningjiang Chen, and Yifan Wu. 2023. Brain: Log parsing with bidirectional parallel tree.IEEE Transactions on Services Computing 16, 5 (2023), 3224–3237
2023
-
[109]
Siyu Yu, Yifan Wu, Ying Li, and Pinjia He. 2024. Unlocking the Power of Numbers: Log Compression via Numeric Token Parsing. InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 919–930
2024
-
[110]
Siyu Yu, Yifan Wu, Zhijing Li, Pinjia He, Ningjiang Chen, and Changjian Liu. 2023. Log parsing with generalization ability under new log types. In Proceedings of the 31st acm joint european software engineering conference and symposium on the foundations of software engineerin...
2023
-
[111]
Jun Zhang and Fenfen Wang. 2022. Research on the construction of log parsing system based on regular expression. In2022 International Conference on Intelligent Transportation, Big Data & Smart City (ICITBS). IEEE, 627–630
2022
-
[112]
Jieming Zhu, Shilin He, Pinjia He, Jinyang Liu, and Michael R Lyu. 2023. Loghub: A large collection of system log datasets for ai-driven log analytics. In2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 355–366. Manuscript submitted to ACM
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.