REVIEW 3 major objections 6 minor 177 references
Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey claims to be the first dedicated map of coverage-guided testing for deep learning, organizing 89 papers into a three-part taxonomy of coverage analysis, test generation, and test optimization.
desk verdict A solid, useful survey of coverage-guided DL testing whose main weakness is a citation-threshold selection criterion that undermines the 'comprehensive' claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the survey's three-part workflow model of CGT, in which coverage analysis sits at the center and feeds two loops: coverage-guided test input generation, an iterative fuzzing cycle with seed selection, seed scheduling, input mutation, execution, test result verification, and seed retention; and coverage-guided test optimization, which handles test suite minimization, prioritization, and selection of unlabeled inputs. The organizing instrument is the taxonomy built from the 89 surveyed papers, which classifies coverage criteria by access level and granularity, generator mutation strategies into four families, and optimization tasks into three categories. These taxonomies carry the argument by letting the survey read distributions directly off the map, for instance the 73.4% share of gray-box criteria and the concentration of evaluation on CNN image models, and by giving the field a vocabulary in which its open challenges can be stated.
What would settle it
Re-run the survey's literature search across the same repositories and time window with the 10-citation floor removed, and rebuild the three taxonomies from the enlarged sample: if the new sample adds categories the current taxonomy cannot absorb, or shifts the headline distributions by a wide margin, the claim that this map is comprehensive fails in a concrete, checkable way.
Extended reading notes
Core claim
The paper's central claim is that coverage-guided testing for deep learning models is a mature, recognizable paradigm whose core is a feedback loop: coverage analysis measures how thoroughly test inputs exercise a model's internal states, and that measurement guides both the generation of new inputs and the optimization of existing test suites. On top of this loop the survey organizes the literature into taxonomies, grouping coverage criteria by access level (white-box, gray-box, black-box) and by granularity (neuron-wise, layer-wise, path-wise, connection-wise) across feedforward, recurrent, transformer, and reinforcement-learning architectures; grouping generators by mutation strategy (gradient-based, metamorphic, search-based, and generative-model-based); and grouping optimization into test suite minimization, prioritization, and selection. Based on its sample of 89 papers, the survey further claims that the field is dominated by gray-box criteria (73.4% of the surveyed criteria) and by CNN image-classification benchmarks such as MNIST, CIFAR-10, and ImageNet, and it identifies the open problems that follow: weak correlation between structural coverage and testing objectives, limited generalizability across architectures and tasks, scalability overhead, and the absence of standardized evaluation protocols and tool support.
Load-bearing premise
The survey's map of the field is only as complete as its paper-selection filter, which keeps studies with at least 10 citations and could therefore leave out recent or niche work that would change the taxonomies and the reported trends.
Editorial extensions
If this is right
- If the taxonomy is correct, any existing or future CGT method can be located in the three-part structure, letting new work position itself against a named category instead of being described in isolation.
- The empirical evidence the survey assembles implies that purely structural coverage criteria are unreliable proxies for error detection, which pushes future criteria toward testing objectives such as robustness and fairness.
- Because evaluation practice is concentrated on CNNs and image benchmarks, findings about CGT effectiveness on RNNs, Transformers, and deep reinforcement learning agents remain largely unverified, a gap the survey's own distribution analysis exposes.
- The survey's call for standardized evaluation protocols implies that cross-study comparisons are not yet trustworthy enough to guide engineering practice until benchmarks, dataset treatments, and metrics converge.
Reading between the lines
- A replication of the survey's search without the 10-citation threshold would be the natural stress test of the taxonomy: if the categories or the distribution claims, such as gray-box dominance, shift materially, the map reflects the selection filter as much as the field.
- The survey's empirical summary points to a sharper experiment than any single surveyed study runs: ablating coverage feedback from a fuzzer while keeping its mutation operators fixed would settle whether coverage guidance itself, rather than the adversarial nature of generated inputs, drives error detection.
- The catalogue of a dozen dataset treatments could be converted directly into a benchmark battery; applying all of them to one fixed model set would let researchers rank coverage criteria by sensitivity, which the survey shows the current literature does not yet provide.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a systematic literature survey of coverage-guided testing (CGT) for deep learning models. It assembles a corpus of 89 papers through a Quasi-Gold Standard search process, proposes taxonomies for coverage criteria (white-box, gray-box, black-box), coverage-guided test input generation (fuzzing, falsification, combinatorial testing, concolic testing), and coverage-guided test optimization (minimization, prioritization, selection). It also analyzes evaluation practices in the surveyed literature—datasets, model architectures, dataset treatments, and evaluation aspects—and closes with open challenges and future directions. The paper claims to be the first comprehensive survey dedicated specifically to CGT for DL models.
Significance. If the corpus is representative, the survey is a useful reference: it systematically organizes a fragmented literature, provides comparative tables that practitioners will find helpful, and synthesizes empirical findings on the weak correlation between structural coverage and error detection. The authors give credit for the QGS methodology, the explicit search string, the inclusion/exclusion criteria, and the detailed taxonomy. The survey also usefully distinguishes evaluation aspects and dataset treatments, which supports reproducibility of future empirical comparisons. The central value, however, is bibliographic, and therefore depends on the completeness and representativeness of the 89-paper corpus.
major comments (3)
- [§3.2.1, Table 3, criterion ②] The 'comprehensive' claim is load-bearing and is not established under the current selection protocol. Criterion ② requires every included paper to have at least 10 citations. Because citation counts grow with time, this systematically excludes recent (2023–2025) relevant work, while the snowballing step in §3.2.2 cannot correct the bias: forward snowballing still must satisfy the same 10-citation criterion, and backward snowballing only recovers older references. Consequently, the publication-trend analysis (Figure 3), the taxonomy distributions (Figures 4, 8, Tables 5–9), and the 'observed gaps' in §8 (e.g., limited Transformer, GNN, and generative-model coverage) may reflect the citation filter rather than the actual state of the field. I request that the authors either (a) re-run the search without the citation threshold and report how the corpus, taxonomies, and trend analysis change, or (b) explicitly re-frame the survey as covering 'influential and well-cited work' and add a limitations paragraph stating that recent low-cited studies are underrepresented. The present wording 'comprehensive survey' is too strong for a corpus selected by citation count.
- [§3.2.2 and §7.5] The paper states that snowballing 'mitigate[s] the risk of overlooking relevant papers,' but the manual screening combined with the citation threshold is still highly selective: only 7 papers were added by snowballing out of 89. This is a small fraction, and the authors do not report the number of papers rejected at each stage (e.g., how many passed the full-text review but were excluded solely due to the 10-citation threshold). Without this information, the reader cannot gauge the severity of the selection bias. Please report the screening funnel with per-criterion rejection counts, and ideally provide a sensitivity analysis of the main descriptive statistics (e.g., dataset and architecture distributions in §7.1) against a corpus that relaxes criterion ②.
- [§1 and §9 (claims of 'first comprehensive survey')] The claim to be 'the first comprehensive survey dedicated to CGT' is presented without a systematic search for other dedicated surveys. The comparison in Table 1 covers general DL testing surveys and the 2024 test-optimization survey, but it does not demonstrate that no other CGT-focused survey exists, including non-English or workshop-level surveys. This claim is secondary to the survey's content, but it should be either verified with a targeted search or softened to 'to the best of our knowledge.'
minor comments (6)
- [§3.2.1, Table 3] The table header column is labeled 'Index' but the exclusion criteria are numbered ❺–❾; the text also refers to 'criterion ❽' for duplicate removal and 'criterion ❻' for scope exclusion. This is internally consistent but slightly confusing; consider labeling the columns 'Inclusion #' and 'Exclusion #'.
- [§4.1.1, DeepGauge description] Typo: 'DeepGuage' appears in the text and 'ℎ𝑖𝑔𝑛𝑛' appears where 'ℎ𝑖𝑔ℎ𝑛' is intended. Also, in §4.3 the text cites 'Neuron Boundary Coverage [82]' while Table 5 lists it under [80]; please reconcile the reference numbers.
- [Table 5, column header] The column header 'Training Inps' is a typo; it should read 'Training Inputs'.
- [§5.1.2, Table 8] In Table 8, the study labeled 'DEEPWALK [159]' is referred to as 'DeepWalk' in the text; please make the naming consistent. Also, the legend for symbols such as '†', '§', and '®' is given in different places; it would help to consolidate all footnotes into one legend.
- [§7.2, Table 10] The 'Criteria Setting' row lists '[6, 38, 47, 55, 72, 77, 79, 80, 96]' and '[126, 131, 140, 147, 148, 151, 167]', which together count 16 studies, but the row label '# studies' reads '16 /' with a stray slash. Please clean up the formatting so that counts are unambiguous.
- [§7.3, second paragraph] The sentence 'When multiple test suites cover the same number of classes, NLC introduced an additional measure using scaled entropy' would benefit from an explicit equation number or a reference to the original NLC paper, since the formula is introduced inline without context.
Circularity Check
No circular derivation: the survey's claims are descriptive summaries of its corpus; the 10-citation inclusion rule is a selection-bias concern, not circularity.
full rationale
This paper is a systematic literature survey, not a formal derivation or empirical prediction. Its central claims—that CGT approaches can be organized into taxonomies of coverage analysis, test input generation, and test optimization, and that evaluation practices are uneven—are descriptive summaries of the 89 selected papers, not outputs forced by fitted parameters or self-citation chains. The only methodological passage that could undermine the 'comprehensive' claim is inclusion criterion ② in §3.2.1 and Table 3, which requires at least 10 citations per paper; this systematically filters out recent or niche work, and the snowballing step in §3.2.2 can only recover papers connected to already-selected seeds. I weigh this as a corpus-completeness and selection-bias threat, not as circular reasoning, because the survey does not define its conclusions in terms of that same criterion. The authors do include several of their own prior works as surveyed primary studies (e.g., references [41], [42], and [46]), but those self-citations are not load-bearing: the taxonomy, evaluation analysis, and future-directions discussion do not depend on accepting any self-cited result as an external proof. No equation or classification step reduces by construction to its own input, so the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The Quasi-Gold Standard (QGS) method is an appropriate systematic review methodology for this field.
- ad hoc to paper The inclusion criterion requiring at least 10 citations selects influential and relevant papers.
- domain assumption The proposed taxonomy categories are comprehensive and non-overlapping.
Cite this review
Pith. "Pith review of Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/ZGJZCICR
@misc{pith2026250700496,
author = {Pith},
title = {Pith review of: Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZGJZCICR}},
note = {Machine review of arXiv:2507.00496}
}
read the original abstract
As Deep Learning (DL) models are increasingly applied in safety-critical domains, ensuring their quality has emerged as a pressing challenge in modern software engineering. Among emerging validation paradigms, coverage-guided testing (CGT) has gained prominence as a systematic framework for identifying erroneous or unexpected model behaviors. Despite growing research attention, existing CGT studies remain methodologically fragmented, limiting the understanding of current advances and emerging trends. This work addresses that gap through a comprehensive review of state-of-the-art CGT methods for DL models, including test coverage analysis, coverage-guided test input generation, and coverage-guided test input optimization. This work provides detailed taxonomies to organize these methods based on methodological characteristics and application scenarios. We also investigate evaluation practices adopted in existing studies, including the use of benchmark datasets, model architectures, and evaluation aspects. Finally, open challenges and future directions are highlighted in terms of the correlation between structural coverage and testing objectives, method generalizability across tasks and models, practical deployment concerns, and the need for standardized evaluation and tool support. This work aims to provide a roadmap for future academic research and engineering practice in DL model quality assurance.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Identifying relevant studies in software engineering.Information and Software Technology53, 6 (2011), 625–637
2011. Identifying relevant studies in software engineering.Information and Software Technology53, 6 (2011), 625–637. Special Section: Best papers from the APSEC. Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2025. Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey•111:45
2011
-
[2]
Stephanie Abrecht, Maram Akila, Sujan Sai Gannamaneni, Konrad Groh, Christian Heinzemann, Sebastian Houben, and Matthias Woehrle. 2020. Revisiting Neuron Coverage and Its Application to Test Generation. InProceedings of the Computer Safety, Reliability, and Security (SAFECOMP-Workshops, 2020, Vol. 12235). Springer, 289–301
2020
-
[3]
Zohreh Aghababaeyan, Manel Abdellatif, Mahboubeh Dadkhah, and Lionel C. Briand. 2024. DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 6 (2024), 158
2024
-
[4]
Raid Rafi Omar Al-Nima, Tingting Han, Saadoon A. M. Al-Sumaidaee, Taolue Chen, and Wai Lok Woo. 2021. Robustness and performance of Deep Reinforcement Learning.Applied Soft Computing (ASC)105 (2021), 107295
2021
-
[5]
Fahmy, Fabrizio Pastore, and Lionel C
Mohammed Walid Attaoui, Hazem M. Fahmy, Fabrizio Pastore, and Lionel C. Briand. 2023. Black-box Safety Analysis and Retraining of DNNs based on Feature Extraction and Clustering.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 3 (2023), 79:1–79:40
2023
-
[6]
Tongtong Bai, Song Huang, Yifan Huang, Xingya Wang, Chunyan Xia, Yubin Qu, and Zhen Yang. 2024. CriticalFuzz: A critical neuron coverage-guided fuzz testing framework for deep neural networks.Information and Software Technology (IST)172 (2024), 107476
2024
-
[7]
Agathe Balayn, Natasa Rikalo, Jie Yang, and Alessandro Bozzon. 2023. Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI, 2023). ACM, 11:1–11:20
2023
-
[8]
Houssem Ben Braiek and Foutse Khomh. 2019. DeepEvolution: A Search-Based Testing Approach for Deep Neural Networks. In Proceedings of the 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME, 2019). IEEE, 454–458
2019
Show all 177 references
-
[9]
Lucas, Peter I
Cameron Browne, Edward Jack Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez Liebana, Spyridon Samothrakis, and Simon Colton. 2012. A Survey of Monte Carlo Tree Search Methods.IEEE Transactions on Computational Inte...
2012
-
[10]
Taejoon Byun and Sanjai Rayadurgam. 2020. Manifold for machine learning assurance. InProceedings of the 42nd International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER, 2020). ACM, 97–100
2020
-
[11]
Taejoon Byun, Vaibhav Sharma, Abhishek Vijayakumar, Sanjai Rayadurgam, and Darren D. Cofer. 2019. Input Prioritization for Testing Neural Networks. InProceedings of the IEEE International Conference On Artificial Intelligence Testing (AITest, 2019). IEEE, 63–70
2019
-
[12]
Cristian Cadar, Daniel Dunbar, and Dawson R. Engler. 2008. KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs. InProceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI, 2008). USENIX Association, 209–224
2008
-
[13]
Sooyoung Cha, Seongjoon Hong, Junhee Lee, and Hakjoo Oh. 2018. Automatically generating search heuristics for concolic testing. In Proceedings of the 40th International Conference on Software Engineering (ICSE, 2018). ACM, 1244–1254
2018
-
[14]
Guillaume M. J-B. Chaslot, Mark H. M. Winands, H. Jaap van den Herik, Jos W. H. M. Uiterwijk, and Bruno Bouzy. 2008. Progressive Strategies for Monte-Carlo Tree Search.New Mathematics and Natural Computation04, 03 (2008), 343–357
2008
-
[15]
Jialuo Chen, Jingyi Wang, Xingjun Ma, Youcheng Sun, Jun Sun, Peixin Zhang, and Peng Cheng. 2023. QuoTe: Quality-oriented Testing for Deep Learning Systems.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 5 (2023), 125:1–125:33
2023
-
[16]
Junjie Chen, Zhuo Wu, Zan Wang, Hanmo You, Lingming Zhang, and Ming Yan. 2020. Practical Accuracy Estimation for Efficient Deep Neural Network Testing.ACM Transactions on Software Engineering and Methodology (TOSEM)29, 4 (2020), 30:1–30:35
2020
-
[17]
Tsong Yueh Chen, Shing-Chi Cheung, and Siu-Ming Yiu. 2020. Metamorphic Testing: A New Approach for Generating Next Test Cases. CoRRabs/2002.12543 (2020). arXiv:2002.12543 https://arxiv.org/abs/2002.12543
2020 arXiv
-
[18]
Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, Pak-Lok Poon, Dave Towey, T. H. Tse, and Zhi Quan Zhou. 2018. Metamorphic Testing: A Review of Challenges and Opportunities.ACM Computing Surveys (CSUR)51, 1 (2018), 4:1–4:27
2018
-
[19]
Pranav Singh Chib and Pravendra Singh. 2024. Recent Advancements in End-to-End Autonomous Driving Using Deep Learning: A Survey.IEEE Transactions on Intelligent Vehicles (TIV)9, 1 (2024), 103–118
2024
-
[20]
Hepeng Dai, Chang-Ai Sun, and Huai Liu. 2022. DeepController: Feedback-Directed Fuzzing for Deep Learning Systems. InProceedings of the 34th International Conference on Software Engineering and Knowledge Engineering (SEKE, 2022). KSI Research Inc., 531–536
2022
-
[21]
Hepeng Dai, Chang-Ai Sun, Huai Liu, and Xiangyu Zhang. 2024. DFuzzer: Diversity-Driven Seed Queue Construction of Fuzzing for Deep Learning Models.IEEE Transactions on Reliability (TR)73, 2 (2024), 1075–1089
2024
-
[22]
Bissyandé, and Yves Le Traon
Xueqi Dang, Yinghua Li, Mike Papadakis, Jacques Klein, Tegawendé F. Bissyandé, and Yves Le Traon. 2024. GraphPrior: Mutation-based Test Input Prioritization for Graph Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 1 (2024), 22:1–22:40
2024
-
[23]
Samet Demir, Hasan Ferit Eniser, and Alper Sen. 2020. DeepSmartFuzzer: Reward Guided Test Generation For Deep Learning. In Proceedings of the Workshop on Artificial Intelligence Safety co-located with the 29th International Joint Conference on Artificial Intelligence and the 1...
2020
-
[24]
Dwyer, and Mary Lou Soffa
Swaroopa Dola, Matthew B. Dwyer, and Mary Lou Soffa. 2021. Distribution-Aware Testing of Neural Networks Using Generative Models. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE 2021). IEEE, 226–237
2021
-
[25]
Dwyer, and Mary Lou Soffa
Swaroopa Dola, Matthew B. Dwyer, and Mary Lou Soffa. 2023. Input Distribution Coverage: Measuring Feature Interaction Adequacy in Neural Network Testing.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 3 (2023), 81:1–81:48. Proc. ACM Meas. Anal. Comput. Syst...
2023
-
[26]
Dwyer, and Mary Lou Soffa
Swaroopa Dola, Rory McDaniel, Matthew B. Dwyer, and Mary Lou Soffa. 2024. CIT4DNN: Generating Diverse and Rare Inputs for Neural Networks Using Latent Space Combinatorial Testing. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE 2024). ...
2024
-
[27]
Yizhen Dong, Peixin Zhang, Jingyi Wang, Shuang Liu, Jun Sun, Jianye Hao, Xinyu Wang, Li Wang, Jin Song Dong, and Ting Dai
-
[28]
Alexandre Donzé and Oded Maler. 2010. Robust Satisfaction of Temporal Logic over Real-Valued Signals. InProceedings of the 8th international conference on Formal modeling and analysis of timed systems (FORMATS, 2010, Vol. 6246). Springer, 92–106
2010
-
[29]
Chengwen Du and Tao Chen. 2024. Contexts Matter: An Empirical Study on Contextual Influence in Fairness Testing for Deep Learning Systems. InProceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM, 2024). ACM, 107–118
2024
-
[30]
Xiaoning Du, Xiaofei Xie, Yi Li, Lei Ma, Yang Liu, and Jianjun Zhao. 2019. DeepStellar: model-based quantitative analysis of stateful deep learning systems. InProceedings of the 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations...
2019
-
[31]
Xiaoning Du, Xiaofei Xie, Yi Li, Lei Ma, Jianjun Zhao, and Yang Liu. 2018. DeepCruiser: Automated Guided Testing for Stateful Deep Learning Systems.CoRRabs/1812.05339 (2018). arXiv:1812.05339 http://arxiv.org/abs/1812.05339
2018 arXiv
-
[32]
Fahmy, Fabrizio Pastore, Lionel C
Hazem M. Fahmy, Fabrizio Pastore, Lionel C. Briand, and Thomas Stifter. 2023. Simulator-based Explanation and Debugging of Hazard-triggering Events in DNN-based Safety-critical Systems.ACM Transactions on Software Engineering and Methodology (TOSEM) 32, 4 (2023), 104:1–104:47
2023
-
[33]
Fainekos and George J
Georgios E. Fainekos and George J. Pappas. 2009. Robustness of temporal logic specifications for continuous-time signals.Theoretical Computer Science (TCS)410, 42 (2009), 4262–4291
2009
-
[34]
Yang Feng, Qingkai Shi, Xinyu Gao, Jun Wan, Chunrong Fang, and Zhenyu Chen. 2020. DeepGini: prioritizing massive tests to enhance the robustness of deep neural networks. InProceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2020...
2020
-
[35]
Xinyu Gao, Yang Feng, Yining Yin, Zixi Liu, Zhenyu Chen, and Baowen Xu. 2022. Adaptive Test Selection for Deep Neural Networks. InProceedings of the IEEE/ACM 44th International Conference on Software Engineering (ICSE, 2022). ACM, 73–85
2022
-
[36]
Saha, Mukul R
Xiang Gao, Ripon K. Saha, Mukul R. Prasad, and Abhik Roychoudhury. 2020. Fuzz testing based data augmentation to improve robustness of deep neural networks. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering (ICSE, 2020). ACM, 1147–1158
2020
-
[37]
Xuanqi Gao, Juan Zhai, Shiqing Ma, Chao Shen, Yufei Chen, and Qian Wang. 2022. Fairneuron: Improving Deep Neural Network Fairness with Adversary Games on Selective Neurons. InProceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE, 2022). A...
2022
-
[38]
Simos Gerasimou, Hasan Ferit Eniser, Alper Sen, and Alper Cakan. 2020. Importance-driven deep learning system testing. InProceedings of the 42nd International Conference on Software Engineering (ICSE, 2020). ACM, 702–713
2020
-
[39]
Goodfellow, Jonathon Shlens, and Christian Szegedy
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. InProceedings of the 3rd International Conference on Learning Representations (ICLR, 2015)
2015
-
[40]
Matthew Groh, Omar Badri, Roxana Daneshjou, Arash Koochek, Caleb Harris, Luis R Soenksen, P Murali Doraiswamy, and Rosalind Picard. 2024. Deep learning-aided decision support for diagnosis of skin disease across skin tones.Nature Medicine30, 2 (2024), 573–583
2024
-
[41]
Hongjing Guo, Chuanqi Tao, and Zhiqiu Huang. 2023. Multi-Objective White-Box Test Input Selection for Deep Neural Network Model Enhancement. InProceedings of the 34th IEEE International Symposium on Software Reliability Engineering (ISSRE, 2023). IEEE, 521–532
2023
-
[42]
Hongjing Guo, Chuanqi Tao, and Zhiqiu Huang. 2024. Neuron importance-aware coverage analysis for deep neural network testing. Empirical Software Engineering (EMSE)29, 5 (2024), 118
2024
-
[43]
Jianmin Guo, Yu Jiang, Yue Zhao, Quan Chen, and Jiaguang Sun. 2018. DLFuzz: differential fuzzing testing of deep learning systems. In Proceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering ...
2018
-
[44]
Jianmin Guo, Quan Zhang, Yue Zhao, Heyuan Shi, Yu Jiang, and Jia-Guang Sun. 2022. RNN-Test: Towards Adversarial Testing for Recurrent Neural Network Systems.IEEE Transactions on Software Engineering (TSE)48, 10 (2022), 4167–4180
2022
-
[45]
Ge Han, Zheng Li, Peng Tang, Chengyu Hu, and Shanqing Guo. 2022. FuzzGAN: A Generation-Based Fuzzing Framework for Testing Deep Neural Networks. InProceedings of the 24th IEEE Int Conf on High Performance Computing & Communications (HPCC, 2022). IEEE, 1601–1608
2022
-
[46]
Yao Hao, Zhiqiu Huang, Hongjing Guo, and Guohua Shen. 2023. Test Input Selection for Deep Neural Network Enhancement Based on Multiple-Objective Optimization. InProceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2023). IE...
2023
-
[47]
Fabrice Harel-Canada, Lingxiao Wang, Muhammad Ali Gulzar, Quanquan Gu, and Miryung Kim. 2020. Is neuron coverage a meaningful measure for testing deep neural networks?. InProceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundation...
2020
-
[48]
Adrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish, Mathias Payer, and Antony L. Hosking. 2021. Seed selection for successful fuzzing. InProceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2021). ACM, 230–243
2021
-
[49]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. InProceedings of the 31st International Conference on Neural Information Processing Systems (NIPS...
2017
-
[50]
Qiang Hu, Yuejun Guo, Maxime Cordy, Xiaofei Xie, Lei Ma, Mike Papadakis, and Yves Le Traon. 2022. An Empirical Study on Data Distribution-Aware Test Selection for Deep Learning Enhancement.ACM Transactions on Software Engineering and Methodology (TOSEM)31, 4 (2022), 78:1–78:30
2022
-
[51]
Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Lei Ma, Mike Papadakis, and Yves Le Traon. 2024. Test Optimization in DNN Testing: A Survey.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 4 (2024), 111:1–111:42
2024
-
[52]
Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Lei Ma, and Yves Le Traon. 2023. Aries: Efficient Testing of Deep Neural Networks via Labeling-Free Accuracy Estimation. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, ...
2023
-
[53]
Qianchao Hu, Feng Wang, Binglin Liu, and Haitian Liu. 2023. Research on Deep Neural Network Testing Techniques. InProceedings of the 4th International Conference on Machine Learning and Computer Application (ICMLCA, 2023). ACM, 113–119
2023
-
[54]
Shengyou Hu, Huayao Wu, Peng Wang, Jing Chang, Yongjun Tu, Xiu Jiang, Xintao Niu, and Changhai Nie. 2023. ATOM: Automated Black-Box Testing of Multi-Label Image Classification Systems. InProceedings of the 38th IEEE/ACM International Conference on Automated Software Engineerin...
2023
-
[55]
Dong Huang, Tsz On Li, Xiaofei Xie, and Heming Cui. 2024. Themis: Automatic and Efficient Deep Learning System Testing with Strong Fault Detection Capability. InProceedings of the 35th IEEE International Symposium on Software Reliability Engineering (ISSRE, 2024). IEEE, 451–462
2024
-
[56]
Wei Huang, Youcheng Sun, Xingyu Zhao, James Sharp, Wenjie Ruan, Jie Meng, and Xiaowei Huang. 2022. Coverage-Guided Testing for Recurrent Neural Networks.IEEE Transactions on Reliability (TR)71, 3 (2022), 1191–1206
2022
-
[57]
Xiaowei Huang, Daniel Kroening, Wenjie Ruan, James Sharp, Youcheng Sun, Emese Thamo, Min Wu, and Xinping Yi. 2020. A survey of safety and trustworthiness of deep neural networks: Verification, testing, adversarial attack and defence, and interpretability.Computer Science Revie...
2020
-
[58]
Belongie, and Jan Kautz
Xun Huang, Ming-Yu Liu, Serge J. Belongie, and Jan Kautz. 2018. Multimodal Unsupervised Image-to-Image Translation. InProceedings of the 15th European Conference on Computer Vision (ECCV, 2018, Vol. 11207). Springer, 179–196
2018
-
[59]
Gunel Jahangirova and Paolo Tonella. 2020. An Empirical Evaluation of Mutation Operators for Deep Learning Systems. InProceedings of the 13th IEEE International Conference on Software Testing, Validation and Verification (ICST, 2020). IEEE, 74–84
2020
-
[60]
Zhenlan Ji, Pingchuan Ma, Yuanyuan Yuan, and Shuai Wang. 2023. CC: Causality-Aware Coverage Criterion for Deep Neural Networks. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 1788–1800
2023
-
[61]
Yue Jia and Mark Harman. 2011. An Analysis and Survey of the Development of Mutation Testing.IEEE Transactions on Software Engineering, TSE37, 5 (2011), 649–678
2011
-
[62]
Zhonghao Jiang, Meng Yan, Li Huang, Weifeng Sun, Chao Liu, Song Sun, and David Lo. 2025. DeepVec: State-Vector Aware Test Case Selection for Enhancing Recurrent Neural Network.IEEE Transactions on Software Engineering (TSE)51, 6 (2025), 1702–1723
2025
-
[63]
Jinhan Kim, Robert Feldt, and Shin Yoo. 2019. Guiding deep learning system testing using surprise adequacy. InProceedings of the 41st International Conference on Software Engineering (ICSE, 2019). IEEE / ACM, 1039–1049
2019
-
[64]
Jinhan Kim, Robert Feldt, and Shin Yoo. 2023. Evaluating Surprise Adequacy for Deep Learning System Testing.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 2 (2023), 42:1–42:29
2023
-
[65]
Jinhan Kim, Jeongil Ju, Robert Feldt, and Shin Yoo. 2020. Reducing DNN labelling cost using surprise adequacy: an industrial case study for autonomous driving. InProceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundati...
2020
-
[66]
Seah Kim and Shin Yoo. 2020. Evaluating Surprise Adequacy for Question Answering. InProceedings of the 42nd International Conference on Software Engineering Workshops (ICSE Workshops, 2020). ACM, 197–202
2020
-
[67]
George Klees, Andrew Ruef, Benji Cooper, Shiyi Wei, and Michael Hicks. 2018. Evaluating Fuzz Testing. InProceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security (CCS, 2018). ACM, 2123–2138
2018
-
[68]
Richard Kuhn, Itzel Dominguez Mendoza, Raghu Kacker, and Yu Lei
D. Richard Kuhn, Itzel Dominguez Mendoza, Raghu Kacker, and Yu Lei. 2013. Combinatorial Coverage Measurement Concepts and Applications. InProceedings of the IEEE 6th International Conference on Software Testing, Verification and Validation Workshops (ICST, 2013). IEEE Computer...
2013
-
[69]
Goodfellow, and Samy Bengio
Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. 2017. Adversarial examples in the physical world. InProceedings of the 5th International Conference on Learning Representations (ICLR, 2017). OpenReview.net
2017
-
[70]
Seokhyun Lee, Sooyoung Cha, Dain Lee, and Hakjoo Oh. 2020. Effective white-box testing of deep neural networks with adaptive neuron-selection strategy. InProceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2020). ACM, 165–176
2020
-
[71]
Bingdong Li, Jinlong Li, Ke Tang, and Xin Yao. 2015. Many-Objective Evolutionary Algorithms: A Survey.ACM Computing Surveys (CSUR)48, 1 (2015), 13:1–13:35
2015
-
[72]
Zenan Li, Xiaoxing Ma, Chang Xu, and Chun Cao. 2019. Structural coverage criteria for neural networks could be misleading. In Proceedings of the 41st International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER, 2019). IEEE / ACM, 89–92
2019
-
[73]
Zenan Li, Xiaoxing Ma, Chang Xu, Chun Cao, Jingwei Xu, and Jian Lü. 2019. Boosting operational DNN testing efficiency through conditioning. InProceedings of the ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineeri...
2019
-
[74]
Zhong Li, Minxue Pan, Tian Zhang, and Xuandong Li. 2021. Testing DNN-based Autonomous Driving Systems under Critical Environmental Conditions. InProceedings of the 38th International Conference on Machine Learning - Volume 139 (ICML, 2021). PMLR, 6471–6482
2021
-
[75]
Zhong Li, Zhengfeng Xu, Ruihua Ji, Minxue Pan, Tian Zhang, Linzhang Wang, and Xuandong Li. 2024. Distance-Aware Test Input Selection for Deep Neural Networks. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2024). ACM, 248–260
2024
-
[76]
Weiguang Liu, Senlin Luo, Limin Pan, and Zhao Zhang. 2025. DeepCNP: An efficient white-box testing of deep neural networks by aligning critical neuron paths.Information and Software Technology179 (2025), 107640
2025
-
[77]
Yibing Liu, Chris Xing Tian, Haoliang Li, Lei Ma, and Shiqi Wang. 2024. Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11,
2024
-
[78]
Zixi Liu, Yang Feng, Yining Yin, and Zhenyu Chen. 2022. DeepState: Selecting Test Suites to Enhance the Robustness of Recurrent Neural Networks. InProceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE, 2022). ACM, 598–609
2022
-
[79]
Lei Ma, Felix Juefei-Xu, Minhui Xue, Bo Li, Li Li, Yang Liu, and Jianjun Zhao. 2019. DeepCT: Tomographic Combinatorial Testing for Deep Learning Systems. InProceedings of the 26th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2019). IE...
2019
-
[80]
Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chunyang Chen, Ting Su, Li Li, Yang Liu, Jianjun Zhao, and Yadong Wang. 2018. DeepGauge: multi-granularity testing criteria for deep learning systems. InProceedings of the 33rd ACM/IEEE International Confere...
2018
-
[81]
Pingchuan Ma, Shuai Wang, and Jin Liu. 2020. Metamorphic Testing and Certified Mitigation of Fairness Violations in NLP Models. In Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI, 2020). ijcai.org, 458–465
2020
-
[82]
Shiqing Ma, Yingqi Liu, Wen-Chuan Lee, Xiangyu Zhang, and Ananth Grama. 2018. MODE: automated neural network model debugging via state differential analysis and input selection. InProceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposi...
2018
-
[83]
Wei Ma, Mike Papadakis, Anestis Tsakmalis, Maxime Cordy, and Yves Le Traon. 2021. Test Selection for Deep Learning Systems.ACM Transactions on Software Engineering and Methodology (TOSEM)30, 2 (2021), 13:1–13:22
2021
-
[84]
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards Deep Learning Models Resistant to Adversarial Attacks. InInternational Conference on Learning Representations. https://openreview.net/forum?id=rJzIBfZAb
2018
-
[85]
Sanoop Mallissery and Yu-Sung Wu. 2024. Demystify the Fuzzing Methods: A Comprehensive Survey.ACM Computing Surveys (CSUR) 56, 3 (2024), 71:1–71:38
2024
-
[86]
Senthil Mani, Anush Sankaran, Srikanth Tamilselvam, and Akshay Sethi. 2019. Coverage Testing of Deep Learning Models using Dataset Characterization.CoRRabs/1911.07309 (2019). arXiv:1911.07309 http://arxiv.org/abs/1911.07309
2019 arXiv
-
[87]
Sondess Missaoui, Simos Gerasimou, and Nicholas Matragkas. 2023. Semantic Data Augmentation for Deep Learning Testing Using Generative AI. InProceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2023). IEEE, 1694–1698
2023
-
[88]
Vasilii Mosin, Miroslaw Staron, Darko Durisic, Francisco Gomes de Oliveira Neto, Sushant Kumar Pandey, and Ashok Chaitanya Koppisetty. 2022. Comparing Input Prioritization Techniques for Testing Deep Learning Algorithms. InProceedings of the 48th Euromicro Conference on Softwa...
2022
-
[89]
de Albuquerque
Khan Muhammad, Amin Ullah, Jaime Lloret, Javier Del Ser, and Victor Hugo C. de Albuquerque. 2021. Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions.IEEE Transactions on Intelligent Transportation Systems, (TITS)22, 7 (2021), 4316–4336. Proc. A...
2021
-
[90]
Briand, and Yvan Labiche
Daniel Di Nardo, Nadia Alshahwan, Lionel C. Briand, and Yvan Labiche. 2015. Coverage-based regression test case selection, minimization and prioritization: a case study on an industrial system.Software Testing, Verification and Reliability (STVR)25, 4 (2015), 371–396
2015
-
[91]
Neelofar and Aldeida Aleti. 2024. Towards Reliable AI: Adequacy Metrics for Ensuring the Quality of System-level Testing of Autonomous Vehicles. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE, 2024). ACM, 68:1–68:12
2024
-
[92]
Mahdi Nejadgholi and Jinqiu Yang. 2019. A Study of Oracle Approximations in Testing Deep Learning Libraries. InProceedings of the 34th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2019). IEEE, 785–796
2019
-
[93]
Andersen, and Ian J
Augustus Odena, Catherine Olsson, David G. Andersen, and Ian J. Goodfellow. 2019. TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing. InProceedings of the 36th International Conference on Machine Learning - Volume 97 (ICML, 2019). PMLR, 4901–4911
2019
-
[94]
Mallikarjuna Paramesha, Nitin Liladhar Rane, and Jayesh Rane. 2024. Artificial Intelligence, Machine Learning, Deep Learning, and Blockchain in Financial and Banking Services: A Comprehensive Review.Partners Universal Multidisciplinary Research Journal1, 2 (Jul. 2024), 51–67
2024
-
[95]
Leo Hyun Park, Soochang Chung, Jaeuk Kim, and Taekyoung Kwon. 2023. GradFuzz: Fuzzing deep neural networks with gradient vector coverage for adversarial examples.Neurocomputing522 (2023), 165–180
2023
-
[96]
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017. DeepXplore: Automated Whitebox Testing of Deep Learning Systems. In Proceedings of the 26th Symposium on Operating Systems Principles (SOSP, 2017). ACM, 1–18
2017
-
[97]
Goran Petrovic, Marko Ivankovic, Gordon Fraser, and René Just. 2021. Does mutation testing improve testing practices?. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE, 2021). IEEE, 910–921
2021
-
[98]
Hung Viet Pham, Thibaud Lutellier, Weizhen Qi, and Lin Tan. 2019. CRADLE: cross-backend validation to detect and localize bugs in deep learning libraries. InProceedings of the 41st International Conference on Software Engineering (ICSE, 2019). IEEE / ACM, 1027–1038
2019
-
[99]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings ...
2021
-
[100]
Alexandre Rebert, Sang Kil Cha, Thanassis Avgerinos, Jonathan Foote, David Warren, Gustavo Grieco, and David Brumley. 2014. Optimizing Seed Selection for Fuzzing. InProceedings of the 23rd USENIX Security Symposium (USENIX Security, 2014). USENIX Association, 861–875
2014
-
[101]
Vincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, and Paolo Tonella. 2021. DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation Score. InProceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2021). IEEE, 355–367
2021
-
[102]
Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella. 2020. Testing machine learning based systems: a systematic mapping.Empirical Software Engineering (EMSE)25, 6 (2020), 5193–5254
2020
-
[103]
Vincenzo Riccio and Paolo Tonella. 2023. When and Why Test Generators for Deep Learning Produce Invalid Inputs: an Empirical Study. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 1161–1173
2023
-
[104]
Buttazzo
Giulio Rossolini, Alessandro Biondi, and Giorgio C. Buttazzo. 2023. Increasing the Confidence of Deep Neural Networks by Coverage Analysis.IEEE Transactions on Software Engineering (TSE)49, 2 (2023), 802–815
2023
-
[105]
Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen
Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved Techniques for Training GANs. InProceedings of the 30th International Conference on Neural Information Processing Systems (NIPS, 2016). 2226–2234
2016
-
[106]
Sergio Segura, Gordon Fraser, Ana Belén Sánchez, and Antonio Ruiz Cortés. 2016. A Survey on Metamorphic Testing.IEEE Transactions on Software Engineering (TSE)42, 9 (2016), 805–824
2016
-
[107]
Dwyer, and Yanjun Qi
Arshdeep Sekhon, Yangfeng Ji, Matthew B. Dwyer, and Yanjun Qi. 2022. White-box Testing of NLP models with Mask Neuron Coverage. InProceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL, 2022). Association for Computa...
2022
-
[108]
Koushik Sen, Darko Marinov, and Gul Agha. 2005. CUTE: a concolic unit testing engine for C. InProceedings of the 10th Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE, 2005). ACM, 263–272
2005
-
[109]
Weijun Shen, Yanhui Li, Lin Chen, Yuanlei Han, Yuming Zhou, and Baowen Xu. 2020. Multiple-Boundary Clustering and Prioritization to Promote Neural Network Retraining. InProceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2020). IEE...
2020
-
[110]
Ying Shi, Beibei Yin, and Jing-Ao Shi. 2025. Markov model based coverage testing of deep learning software systems.Information and Software Technology (IST)179 (2025), 107628
2025
-
[111]
Ying Shi, Beibei Yin, and Zheng Zheng. 2024. Multi-granularity coverage criteria for deep reinforcement learning systems.Journal of Systems and Software (JSS)212 (2024), 112016
2024
-
[112]
Ying Shi, Beibei Yin, Zheng Zheng, and Tiancheng Li. 2021. An Empirical Study on Test Case Prioritization Metrics for Deep Neural Networks. InProceedings of the 21st IEEE International Conference on Software Quality, Reliability and Security (QRS, 2021). IEEE, Proc. ACM Meas. ...
2021
-
[113]
Jiaze Sun, Juan Li, and Sulei Wen. 2023. DeepMC: DNN test sample optimization method jointly guided by misclassification and coverage.Appl. Intell.53, 12 (2023), 15787–15801
2023
-
[114]
Weidi Sun, Yuteng Lu, Xiaokun Luan, and Meng Sun. 2023. HeatC: A Variable-Grained Coverage Criterion for Deep Learning Systems. InProceedings of the 9th International Symposium on Dependable Software Engineering. Theories, Tools, and Applications (SETTA, 2023, Vol. 14464). Spr...
2023
-
[115]
Weidi Sun, Yuteng Lu, and Meng Sun. 2021. Are Coverage Criteria Meaningful Metrics for DNNs?. InProceedings of the International Joint Conference on Neural Networks (IJCNN, 2021). IEEE, 1–8
2021
-
[116]
Weidi Sun, Xiaoyong Xue, Yuteng Lu, Jia Zhao, and Meng Sun. 2023. HashC: Making deep learning coverage testing finer and faster. Journal of Systems Architecture (JSA)144 (2023), 102999
2023
-
[117]
Weifeng Sun, Meng Yan, Zhongxin Liu, and David Lo. 2023. Robust Test Selection for Deep Neural Networks.IEEE Transactions on Software Engineering (TSE)49, 12 (2023), 5250–5278
2023
-
[118]
Youcheng Sun, Xiaowei Huang, Daniel Kroening, James Sharp, Matthew Hill, and Rob Ashmore. 2019. Structural Test Coverage Criteria for Deep Neural Networks.ACM Transactions on Embedded Computing Systems (TECS)18, 5s (2019), 94:1–94:23
2019
-
[119]
Youcheng Sun, Min Wu, Wenjie Ruan, Xiaowei Huang, Marta Kwiatkowska, and Daniel Kroening. 2018. Concolic testing for deep neural networks. InProceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (ASE, 2018). ACM, 109–119
2018
-
[120]
Chuanqi Tao, Yali Tao, Hongjing Guo, Zhiqiu Huang, and Xiaobing Sun. 2023. DLRegion: Coverage-guided fuzz testing of deep neural networks with region-based neuron selection strategies.Information and Software Technology (IST)162 (2023), 107266
2023
-
[121]
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. 2018. DeepTest: automated testing of deep-neural-network-driven autonomous cars. InProceedings of the 40th International Conference on Software Engineering (ICSE, 2018). ACM, 303–314
2018
-
[122]
Yongqiang Tian, Wuqi Zhang, Ming Wen, Shing-Chi Cheung, Chengnian Sun, Shiqing Ma, and Yu Jiang. 2023. Finding Deviated Behaviors of the Compressed DNN Models for Image Classifications.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 5 (2023), 128:1–128:32
2023
-
[123]
Miller Trujillo, Mario Linares-Vásquez, Camilo Escobar-Velásquez, Ivana Dusparic, and Nicolás Cardozo. 2020. Does Neuron Coverage Matter for Deep Reinforcement Learning?: A Preliminary Study. InProceedings of the IEEE/ACM 42nd International Conference on Software Engineering W...
2020
-
[124]
Pasareanu
Muhammad Usman, Youcheng Sun, Divya Gopinath, Rishi Dange, Luca Manolache, and Corina S. Pasareanu. 2023. An overview of structural coverage metrics for testing neural networks.International Journal on Software Tools for Technology Transfer (STTT)25, 3 (2023), 393–405
2023
-
[125]
Xiaohui Wan, Tiancheng Li, Weibin Lin, Yi Cai, and Zheng Zheng. 2024. Coverage-guided fuzzing for deep reinforcement learning systems.Journal of Systems and Software (JSS)210 (2024), 111963
2024
-
[126]
Dong Wang, Ziyuan Wang, Chunrong Fang, Yanshan Chen, and Zhenyu Chen. 2019. DeepPath: Path-Driven Testing Criteria for Deep Neural Networks. InProceedings of the IEEE International Conference On Artificial Intelligence Testing (AITest, 2019). IEEE, 119–120
2019
-
[127]
Jingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma, Dongxia Wang, Jun Sun, and Peng Cheng. 2021. RobOT: Robustness-Oriented Testing for Deep Learning Systems. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE, 2021). IEEE, 300–311
2021
-
[128]
Peng Wang, Shengyou Hu, Huayao Wu, Xintao Niu, Changhai Nie, and Lin Chen. 2024. A Combinatorial Interaction Testing Method for Multi-Label Image Classifier. InProceedings of the 35th IEEE International Symposium on Software Reliability Engineering (ISSRE 2024). IEEE, 463–474
2024
-
[129]
Eric Wong
Shengrong Wang, Dongcheng Li, Hui Li, Man Zhao, and W. Eric Wong. 2024. A Survey on Test Input Selection and Prioritization for Deep Neural Networks. InProceedings of the 10th International Symposium on System Security, Safety, and Reliability (ISSSR, 2024). 232–243
2024
-
[130]
Shuai Wang and Zhendong Su. 2020. Metamorphic Object Insertion for Testing Object Detection Systems. InProceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2020). IEEE, 1053–1065
2020
-
[131]
Zhiyu Wang, Sihan Xu, Lingling Fan, Xiangrui Cai, Linyu Li, and Zheli Liu. 2024. Can Coverage Criteria Guide Failure Discovery for Image Classifiers? An Empirical Study.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 7, Article 190 (Sept. 2024), 28 pages
2024
-
[132]
Zan Wang, Hanmo You, Junjie Chen, Yingyi Zhang, Xuyuan Dong, and Wenbin Zhang. 2021. Prioritizing Test Inputs for Deep Neural Networks via Mutation Analysis. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE, 2021). IEEE, 397–409
2021
-
[133]
Moshi Wei, Yuchao Huang, Jinqiu Yang, Junjie Wang, and Song Wang. 2023. CoCoFuzzing: Testing Neural Code Models With Coverage-Guided Fuzzing.IEEE Transactions on Reliability (TR)72, 3 (2023), 1276–1289
2023
-
[134]
Zhengyuan Wei and W. K. Chan. 2021. Fuzzing Deep Learning Models against Natural Robustness with Filter Coverage. InProceedings of the 21st IEEE International Conference on Software Quality, Reliability and Security (QRS, 2021). IEEE, 608–619. Proc. ACM Meas. Anal. Comput. Sys...
2021
-
[135]
Michael Weiss and Paolo Tonella. 2022. Simple techniques work surprisingly well for neural network test prioritization and active learning (replicability study). InProceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2022). ACM, 139–150
2022
-
[136]
Tsui-Wei Weng, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, Dong Su, Yupeng Gao, Cho-Jui Hsieh, and Luca Daniel. 2018. Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach. InInternational Conference on Learning Representations. https://openreview.net/forum?i...
2018
-
[137]
Xiaoxue Wu, Jinjin Shen, Wei Zheng, Lidan Lin, Yulei Sui, and Abubakar Omari Abdallah Semasaba. 2023. RNNtcs: A test case selection method for Recurrent Neural Networks.Knowledge-Based Systems (KBS)279 (2023), 110955
2023
-
[138]
Dongwei Xiao, Zhibo Liu, Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2022. Metamorphic Testing of Deep Learning Compilers. Proceedings of the ACM on Measurement and Analysis of Computing Systems (POMACS)6, 1 (2022), 15:1–15:28
2022
-
[139]
Tao Xie, Darko Marinov, Wolfram Schulte, and David Notkin. 2005. Symstra: A Framework for Generating Object-Oriented Unit Tests Using Symbolic Execution. InProceedings of the 11th International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TA...
2005
-
[140]
Xiaofei Xie, Tianlin Li, Jian Wang, Lei Ma, Qing Guo, Felix Juefei-Xu, and Yang Liu. 2022. NPC: Neuron Path Coverage via Characterizing Decision Logic of Deep Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)31, 3 (2022), 47:1–47:27
2022
-
[141]
Xiaofei Xie, Lei Ma, Felix Juefei-Xu, Minhui Xue, Hongxu Chen, Yang Liu, Jianjun Zhao, Bo Li, Jianxiong Yin, and Simon See. 2019. DeepHunter: a coverage-guided fuzz testing framework for deep neural networks. InProceedings of the 28th ACM SIGSOFT International Symposium on Sof...
2019
-
[142]
Xiaofei Xie, Lei Ma, Haijun Wang, Yuekang Li, Yang Liu, and Xiaohong Li. 2019. DiffChaser: Detecting Disagreements for Deep Neural Networks. InProceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI, 2019). ijcai.org, 5772–5778
2019
-
[143]
Yutao Xie, Jiayi Lin, Hande Dong, Lei Zhang, and Zhonghai Wu. 2024. Survey of Code Search Based on Deep Learning.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 2 (2024), 54:1–54:42
2024
-
[144]
Huan Xu and Shie Mannor. 2012. Robustness and generalization.Mach. Learn.86, 3 (2012), 391–423
2012
-
[145]
Ahmed Haj Yahmed, Houssem Ben Braiek, Foutse Khomh, Sonia Bouzidi, and Rania Zaatour. 2022. DiverGet: a Search-Based Software Testing approach for Deep Neural Network Quantization assessment.Empirical Software Engineering (EMSE)27, 7 (2022), 193
2022
-
[146]
Yoriyuki Yamagata, Shuang Liu, Takumi Akazaki, Yihai Duan, and Jianye Hao. 2021. Falsification of Cyber-Physical Systems Using Deep Reinforcement Learning.IEEE Transactions on Software Engineering (TSE)47, 12 (2021), 2823–2840
2021
-
[147]
Ming Yan, Junjie Chen, Xuejie Cao, Zhuo Wu, Yuning Kang, and Zan Wang. 2024. Revisiting deep neural network test coverage from the test effectiveness perspective.Journal of Software: Evolution and Process (JSEP)36, 4 (2024)
2024
-
[148]
Shenao Yan, Guanhong Tao, Xuwei Liu, Juan Zhai, Shiqing Ma, Lei Xu, and Xiangyu Zhang. 2020. Correlations between deep neural network model coverage criteria and model quality. InProceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposiu...
2020
-
[149]
Minghao Yang, Shunkun Yang, and Wenda Wu. 2022. DeepRTest: A Vulnerability-Guided Robustness Testing and Enhancement Framework for Deep Neural Networks. InProceedings of the 22nd IEEE International Conference on Software Quality, Reliability and Security (QRS, 2022). IEEE, 754–762
2022
-
[150]
Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2022. Revisiting Neuron Coverage Metrics and Quality of Deep Neural Networks. InProceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2022). IEEE, 408–419
2022
-
[151]
Aoshuang Ye, Lina Wang, Lei Zhao, and Jianpeng Ke. 2022. DANCe: Dynamic Adaptive Neuron Coverage for Fuzzing Deep Neural Networks. InProceedings of the International Joint Conference on Neural Networks (IJCNN, 2022). IEEE, 1–8
2022
-
[152]
Shin Yoo and Mark Harman. 2007. Pareto efficient multi-objective test case selection. InProceedings of the ACM/SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2007). ACM, 140–150
2007
-
[153]
Shin Yoo and Mark Harman. 2012. Regression testing minimization, selection and prioritization: a survey.Software Testing, Verification and Reliability (STVR)22, 2 (2012), 67–120
2012
-
[154]
Hanmo You, Zan Wang, Junjie Chen, Shuang Liu, and Shuochuan Li. 2023. Regression Fuzzing for Deep Learning Systems. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 82–94
2023
-
[155]
Guangba Yu, Gou Tan, Haojia Huang, Zhenyu Zhang, Pengfei Chen, Roberto Natella, and Zibin Zheng. 2024. A Survey on Failure Analysis and Fault Injection in AI Systems. arXiv:2407.00125 [cs.SE] https://arxiv.org/abs/2407.00125
2024
-
[156]
Jing Yu, Shukai Duan, and Xiaojun Ye. 2023. A White-Box Testing for Deep Neural Networks Based on Neuron Coverage.IEEE Transactions on Neural Networks and Learning Systems (TNNLS)34, 11 (2023), 9185–9197
2023
-
[157]
Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2022. Unveiling Hidden DNN Defects with Decision-Based Metamorphic Testing. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2022). ACM, 113:1–113:13
2022
-
[158]
Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2023. Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 1200–1212. Proc. ACM Meas. Anal. Com...
2023
-
[159]
Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2024. Provably Valid and Diverse Mutations of Real-World Media Data for DNN Testing. IEEE Transactions on Software Engineering (TSE)50, 5 (2024), 1040–1064
2024
-
[160]
Yuanyuan Yuan, Shuai Wang, and Zhendong Su. 2024. See the Forest, not Trees: Unveiling and Escaping the Pitfalls of Error-Triggering Inputs in Neural Network Testing. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2024). ...
2024
-
[161]
Chenhan Zhang, Shui Yu, Zhiyi Tian, and James J. Q. Yu. 2024. Generative Adversarial Networks: A Survey on Attack and Defense Perspective.ACM Computing Surveys (CSUR)56, 4 (2024), 91:1–91:35
2024
-
[162]
Zhang, Mark Harman, Lei Ma, and Yang Liu
Jie M. Zhang, Mark Harman, Lei Ma, and Yang Liu. 2022. Machine Learning Testing: Survey, Landscapes and Horizons.IEEE Transactions on Software Engineering (TSE)48, 2 (2022), 1–36
2022
-
[163]
Pengcheng Zhang, Bin Ren, Hai Dong, and Qiyin Dai. 2022. CAGFuzz: Coverage-Guided Adversarial Generative Fuzzing Testing for Image-Based Deep Learning Systems.IEEE Transactions on Software Engineering (TSE)48, 11 (2022), 4630–4646
2022
-
[164]
Peixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong, Xinyu Wang, Xingen Wang, Jin Song Dong, and Ting Dai. 2020. White-box fairness testing through adversarial sampling. InProceedings of the 42nd International Conference on Software Engineering (ICSE, 2020). ACM, 949–960
2020
-
[165]
Xiaoyu Zhang, Weipeng Jiang, Chao Shen, Qi Li, Qian Wang, Chenhao Lin, and Xiaohong Guan. 2025. Deep Learning Library Testing: Definition, Methods and Challenges.ACM Computing Surveys (CSUR)57, 7 (2025), 187:1–187:37
2025
-
[166]
Zhenya Zhang, Gidon Ernst, Sean Sedwards, Paolo Arcaini, and Ichiro Hasuo. 2018. Two-Layered Falsification of Hybrid Systems Guided by Monte Carlo Tree Search.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)37, 11 (2018), 2894–2905
2018
-
[167]
Zhenya Zhang, Deyun Lyu, Paolo Arcaini, Lei Ma, Ichiro Hasuo, and Jianjun Zhao. 2023. FalsifAI: Falsification of AI-Enabled Hybrid Control Systems Guided by Time-Aware Coverage Criteria.IEEE Transactions on Software Engineering (TSE)49, 4 (2023), 1842–1859
2023
-
[168]
Zhuangyu Zhang, Zhiyi Zhang, Ziyuan Wang, Fang Chen, and Zhiqiu Huang. 2023. BTM: Black-Box Testing for DNN Based on Meta-Learning. InProceedings of the 23rd IEEE International Conference on Software Quality, Reliability, and Security (QRS, 2023). IEEE, 581–592
2023
-
[169]
Chunyu Zhao, Yanzhou Mu, Xiang Chen, Jingke Zhao, Xiaolin Ju, and Gan Wang. 2022. Can test input selection methods for deep neural network guarantee test diversity? A large-scale empirical study.Information and Software Technology (IST)150 (2022), 106982
2022
-
[170]
Wei Zheng, Lidan Lin, Xiaoxue Wu, and Xiang Chen. 2024. An Empirical Study on Correlations Between Deep Neural Network Fairness and Neuron Coverage Criteria.IEEE Transactions on Software Engineering (TSE)50, 3 (2024), 391–412
2024
-
[171]
Yuhan Zhi, Xiaofei Xie, Chao Shen, Jun Sun, Xiaoyu Zhang, and Xiaohong Guan. 2024. Seed Selection for Testing Deep Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 1 (2024), 23:1–23:33
2024
-
[172]
Jianyi Zhou, Feng Li, Jinhao Dong, Hongyu Zhang, and Dan Hao. 2020. Cost-Effective Testing of a Deep Learning Model through Input Reduction. InProceedings of the 31st IEEE International Symposium on Software Reliability Engineering (ISSRE, 2020). IEEE, 289–300
2020
-
[173]
Zhiyang Zhou, Wensheng Dou, Jie Liu, Chenxin Zhang, Jun Wei, and Dan Ye. 2021. DeepCon: Contribution Coverage Testing for Deep Learning Systems. InProceedings of the 28th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2021). IEEE, 189–200
2021
-
[174]
Zenghui Zhou, Pak-Lok Poon, Tsong Yueh Chen, Kun Qiu, Qinghua Zhao, and Zheng Zheng. 2025. Evaluating the effectiveness of neuron coverage metrics: a metamorphic-testing approach.Software Quality Journal (SQJ)33, 2 (2025), 19
2025
-
[175]
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. 2017. Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. InProceedings of the IEEE International Conference on Computer Vision (ICCV 2017). IEEE Computer Society, 2242–2251
2017
-
[176]
Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, and Paolo Tonella. 2021. DeepHyperion: exploring the feature space of deep learning-based systems through illumination search. InProceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis (IS...
2021
-
[2020]
InProceedings of the 25th International Conference on Engineering of Complex Computer Systems (ICECCS, 2020)
An Empirical Study on Correlation between Coverage and Robustness for Deep Neural Networks. InProceedings of the 25th International Conference on Engineering of Complex Computer Systems (ICECCS, 2020). IEEE, 73–82
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.