Pith. sign in

REVIEW 3 major objections 6 minor 177 references

Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This survey claims to be the first dedicated map of coverage-guided testing for deep learning, organizing 89 papers into a three-part taxonomy of coverage analysis, test generation, and test optimization.

desk verdict A solid, useful survey of coverage-guided DL testing whose main weakness is a citation-threshold selection criterion that undermines the 'comprehensive' claim. read the letter →

arxiv 2507.00496 v1 pith:ZGJZCICR submitted 2025-07-01 cs.SE

classification cs.SE
keywords coverage-guidedtestingdeeplearningcoveragecriteriatestinputgenerationoptimizationneuronsystematicsurveytaxonomy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims that coverage-guided testing (CGT) has become a distinct paradigm for deep learning quality assurance and that the field's scattered literature can be organized into one coherent map. It positions itself as the first survey devoted specifically to CGT rather than to general DL testing, and it builds the map from 89 selected papers sorted into three pillars: coverage criteria, which measure how much of a model's behavior space a test suite exercises; coverage-guided test input generation, which uses those signals to synthesize error-revealing inputs; and coverage-guided test optimization, which prioritizes, selects, or minimizes unlabeled test inputs to save labeling cost. The value the survey offers, if its claim holds, is a shared vocabulary and structure for future work, along with a documented picture of how CGT methods are evaluated today. It also records the field's emerging doubts, noting that empirical studies increasingly find weak correlation between structural coverage and error detection.

What carries the argument

The load-bearing object is the survey's three-part workflow model of CGT, in which coverage analysis sits at the center and feeds two loops: coverage-guided test input generation, an iterative fuzzing cycle with seed selection, seed scheduling, input mutation, execution, test result verification, and seed retention; and coverage-guided test optimization, which handles test suite minimization, prioritization, and selection of unlabeled inputs. The organizing instrument is the taxonomy built from the 89 surveyed papers, which classifies coverage criteria by access level and granularity, generator mutation strategies into four families, and optimization tasks into three categories. These taxonomies carry the argument by letting the survey read distributions directly off the map, for instance the 73.4% share of gray-box criteria and the concentration of evaluation on CNN image models, and by giving the field a vocabulary in which its open challenges can be stated.

What would settle it

Re-run the survey's literature search across the same repositories and time window with the 10-citation floor removed, and rebuild the three taxonomies from the enlarged sample: if the new sample adds categories the current taxonomy cannot absorb, or shifts the headline distributions by a wide margin, the claim that this map is comprehensive fails in a concrete, checkable way.

Watch

Extended reading notes

Core claim

The paper's central claim is that coverage-guided testing for deep learning models is a mature, recognizable paradigm whose core is a feedback loop: coverage analysis measures how thoroughly test inputs exercise a model's internal states, and that measurement guides both the generation of new inputs and the optimization of existing test suites. On top of this loop the survey organizes the literature into taxonomies, grouping coverage criteria by access level (white-box, gray-box, black-box) and by granularity (neuron-wise, layer-wise, path-wise, connection-wise) across feedforward, recurrent, transformer, and reinforcement-learning architectures; grouping generators by mutation strategy (gradient-based, metamorphic, search-based, and generative-model-based); and grouping optimization into test suite minimization, prioritization, and selection. Based on its sample of 89 papers, the survey further claims that the field is dominated by gray-box criteria (73.4% of the surveyed criteria) and by CNN image-classification benchmarks such as MNIST, CIFAR-10, and ImageNet, and it identifies the open problems that follow: weak correlation between structural coverage and testing objectives, limited generalizability across architectures and tasks, scalability overhead, and the absence of standardized evaluation protocols and tool support.

Load-bearing premise

The survey's map of the field is only as complete as its paper-selection filter, which keeps studies with at least 10 citations and could therefore leave out recent or niche work that would change the taxonomies and the reported trends.

Editorial extensions

If this is right

  • If the taxonomy is correct, any existing or future CGT method can be located in the three-part structure, letting new work position itself against a named category instead of being described in isolation.
  • The empirical evidence the survey assembles implies that purely structural coverage criteria are unreliable proxies for error detection, which pushes future criteria toward testing objectives such as robustness and fairness.
  • Because evaluation practice is concentrated on CNNs and image benchmarks, findings about CGT effectiveness on RNNs, Transformers, and deep reinforcement learning agents remain largely unverified, a gap the survey's own distribution analysis exposes.
  • The survey's call for standardized evaluation protocols implies that cross-study comparisons are not yet trustworthy enough to guide engineering practice until benchmarks, dataset treatments, and metrics converge.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A replication of the survey's search without the 10-citation threshold would be the natural stress test of the taxonomy: if the categories or the distribution claims, such as gray-box dominance, shift materially, the map reflects the selection filter as much as the field.
  • The survey's empirical summary points to a sharper experiment than any single surveyed study runs: ablating coverage feedback from a fuzzer while keeping its mutation operators fixed would settle whether coverage guidance itself, rather than the adversarial nature of generated inputs, drives error detection.
  • The catalogue of a dozen dataset treatments could be converted directly into a benchmark battery; applying all of them to one fixed model set would let researchers rank coverage criteria by sensitivity, which the survey shows the current literature does not yet provide.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript presents a systematic literature survey of coverage-guided testing (CGT) for deep learning models. It assembles a corpus of 89 papers through a Quasi-Gold Standard search process, proposes taxonomies for coverage criteria (white-box, gray-box, black-box), coverage-guided test input generation (fuzzing, falsification, combinatorial testing, concolic testing), and coverage-guided test optimization (minimization, prioritization, selection). It also analyzes evaluation practices in the surveyed literature—datasets, model architectures, dataset treatments, and evaluation aspects—and closes with open challenges and future directions. The paper claims to be the first comprehensive survey dedicated specifically to CGT for DL models.

Significance. If the corpus is representative, the survey is a useful reference: it systematically organizes a fragmented literature, provides comparative tables that practitioners will find helpful, and synthesizes empirical findings on the weak correlation between structural coverage and error detection. The authors give credit for the QGS methodology, the explicit search string, the inclusion/exclusion criteria, and the detailed taxonomy. The survey also usefully distinguishes evaluation aspects and dataset treatments, which supports reproducibility of future empirical comparisons. The central value, however, is bibliographic, and therefore depends on the completeness and representativeness of the 89-paper corpus.

major comments (3)
  1. [§3.2.1, Table 3, criterion ②] The 'comprehensive' claim is load-bearing and is not established under the current selection protocol. Criterion ② requires every included paper to have at least 10 citations. Because citation counts grow with time, this systematically excludes recent (2023–2025) relevant work, while the snowballing step in §3.2.2 cannot correct the bias: forward snowballing still must satisfy the same 10-citation criterion, and backward snowballing only recovers older references. Consequently, the publication-trend analysis (Figure 3), the taxonomy distributions (Figures 4, 8, Tables 5–9), and the 'observed gaps' in §8 (e.g., limited Transformer, GNN, and generative-model coverage) may reflect the citation filter rather than the actual state of the field. I request that the authors either (a) re-run the search without the citation threshold and report how the corpus, taxonomies, and trend analysis change, or (b) explicitly re-frame the survey as covering 'influential and well-cited work' and add a limitations paragraph stating that recent low-cited studies are underrepresented. The present wording 'comprehensive survey' is too strong for a corpus selected by citation count.
  2. [§3.2.2 and §7.5] The paper states that snowballing 'mitigate[s] the risk of overlooking relevant papers,' but the manual screening combined with the citation threshold is still highly selective: only 7 papers were added by snowballing out of 89. This is a small fraction, and the authors do not report the number of papers rejected at each stage (e.g., how many passed the full-text review but were excluded solely due to the 10-citation threshold). Without this information, the reader cannot gauge the severity of the selection bias. Please report the screening funnel with per-criterion rejection counts, and ideally provide a sensitivity analysis of the main descriptive statistics (e.g., dataset and architecture distributions in §7.1) against a corpus that relaxes criterion ②.
  3. [§1 and §9 (claims of 'first comprehensive survey')] The claim to be 'the first comprehensive survey dedicated to CGT' is presented without a systematic search for other dedicated surveys. The comparison in Table 1 covers general DL testing surveys and the 2024 test-optimization survey, but it does not demonstrate that no other CGT-focused survey exists, including non-English or workshop-level surveys. This claim is secondary to the survey's content, but it should be either verified with a targeted search or softened to 'to the best of our knowledge.'
minor comments (6)
  1. [§3.2.1, Table 3] The table header column is labeled 'Index' but the exclusion criteria are numbered ❺–❾; the text also refers to 'criterion ❽' for duplicate removal and 'criterion ❻' for scope exclusion. This is internally consistent but slightly confusing; consider labeling the columns 'Inclusion #' and 'Exclusion #'.
  2. [§4.1.1, DeepGauge description] Typo: 'DeepGuage' appears in the text and 'ℎ𝑖𝑔𝑛𝑛' appears where 'ℎ𝑖𝑔ℎ𝑛' is intended. Also, in §4.3 the text cites 'Neuron Boundary Coverage [82]' while Table 5 lists it under [80]; please reconcile the reference numbers.
  3. [Table 5, column header] The column header 'Training Inps' is a typo; it should read 'Training Inputs'.
  4. [§5.1.2, Table 8] In Table 8, the study labeled 'DEEPWALK [159]' is referred to as 'DeepWalk' in the text; please make the naming consistent. Also, the legend for symbols such as '†', '§', and '®' is given in different places; it would help to consolidate all footnotes into one legend.
  5. [§7.2, Table 10] The 'Criteria Setting' row lists '[6, 38, 47, 55, 72, 77, 79, 80, 96]' and '[126, 131, 140, 147, 148, 151, 167]', which together count 16 studies, but the row label '# studies' reads '16 /' with a stray slash. Please clean up the formatting so that counts are unambiguous.
  6. [§7.3, second paragraph] The sentence 'When multiple test suites cover the same number of classes, NLC introduced an additional measure using scaled entropy' would benefit from an explicit equation number or a reference to the original NLC paper, since the formula is introduced inline without context.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the survey's claims are descriptive summaries of its corpus; the 10-citation inclusion rule is a selection-bias concern, not circularity.

full rationale

This paper is a systematic literature survey, not a formal derivation or empirical prediction. Its central claims—that CGT approaches can be organized into taxonomies of coverage analysis, test input generation, and test optimization, and that evaluation practices are uneven—are descriptive summaries of the 89 selected papers, not outputs forced by fitted parameters or self-citation chains. The only methodological passage that could undermine the 'comprehensive' claim is inclusion criterion ② in §3.2.1 and Table 3, which requires at least 10 citations per paper; this systematically filters out recent or niche work, and the snowballing step in §3.2.2 can only recover papers connected to already-selected seeds. I weigh this as a corpus-completeness and selection-bias threat, not as circular reasoning, because the survey does not define its conclusions in terms of that same criterion. The authors do include several of their own prior works as surveyed primary studies (e.g., references [41], [42], and [46]), but those self-citations are not load-bearing: the taxonomy, evaluation analysis, and future-directions discussion do not depend on accepting any self-cited result as an external proof. No equation or classification step reduces by construction to its own input, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey's central claim rests on the completeness and representativeness of the literature search, the fairness of the citation threshold, and the validity of the taxonomy. There are no free parameters fitted to data and no new postulated entities.

assumptions (3)
  • domain assumption The Quasi-Gold Standard (QGS) method is an appropriate systematic review methodology for this field.
    The paper relies on QGS to collect and screen papers, but this is one of several possible review methodologies and could affect which papers are included.
  • ad hoc to paper The inclusion criterion requiring at least 10 citations selects influential and relevant papers.
    Section 3.2.1 introduces this threshold, which systematically excludes recent or less-cited work and may bias the survey's coverage and conclusions.
  • domain assumption The proposed taxonomy categories are comprehensive and non-overlapping.
    The categorization of coverage criteria and generation methods is an interpretive framework; other equally valid classifications may exist, and the boundaries between categories are not formally proven.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/ZGJZCICR

@misc{pith2026250700496,
  author       = {Pith},
  title        = {Pith review of: Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZGJZCICR}},
  note         = {Machine review of arXiv:2507.00496}
}
read the original abstract

As Deep Learning (DL) models are increasingly applied in safety-critical domains, ensuring their quality has emerged as a pressing challenge in modern software engineering. Among emerging validation paradigms, coverage-guided testing (CGT) has gained prominence as a systematic framework for identifying erroneous or unexpected model behaviors. Despite growing research attention, existing CGT studies remain methodologically fragmented, limiting the understanding of current advances and emerging trends. This work addresses that gap through a comprehensive review of state-of-the-art CGT methods for DL models, including test coverage analysis, coverage-guided test input generation, and coverage-guided test input optimization. This work provides detailed taxonomies to organize these methods based on methodological characteristics and application scenarios. We also investigate evaluation practices adopted in existing studies, including the use of benchmark datasets, model architectures, and evaluation aspects. Finally, open challenges and future directions are highlighted in terms of the correlation between structural coverage and testing objectives, method generalizability across tasks and models, practical deployment concerns, and the need for standardized evaluation and tool support. This work aims to provide a roadmap for future academic research and engineering practice in DL model quality assurance.

Figures

Figures reproduced from arXiv: 2507.00496 by the authors.

Figure 1
Figure 1. Workflow of Coverage-Guided Testing for DL Models [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Study Collection and Selection Process [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Publication Trends of Coverage-Guided Testing Studies [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Distribution of Coverage Criteria for DL Models (b) Layer-Wise (c) Path-Wise [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: High-Level Illustrations of White-box and Gray-box Coverage Criteria for FNNs [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Workflow of Coverage-Guided Fuzzing for DL Models [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Illustrations of Test Optimization Tasks for DL Models [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 8
Figure 8. Figure 8: Overview of Model Architectures and Datasets in Coverage-Guided Testing Studies [PITH_FULL_IMAGE:figures/full_fig_p030_8.png]
Figure 9
Figure 9. Figure 9: Illustrations of Dataset Treatments for Coverage Criteria Evaluation [PITH_FULL_IMAGE:figures/full_fig_p031_9.png]
Figure 10
Figure 10. Figure 10: Publication Trends and Research Focuses of Empirical Studies on Coverage-Guided Testing [PITH_FULL_IMAGE:figures/full_fig_p039_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

177 extracted references · 78 canonical work pages

  1. [1]

    Identifying relevant studies in software engineering.Information and Software Technology53, 6 (2011), 625–637

    2011. Identifying relevant studies in software engineering.Information and Software Technology53, 6 (2011), 625–637. Special Section: Best papers from the APSEC. Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2025. Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey•111:45

  2. [2]

    Stephanie Abrecht, Maram Akila, Sujan Sai Gannamaneni, Konrad Groh, Christian Heinzemann, Sebastian Houben, and Matthias Woehrle. 2020. Revisiting Neuron Coverage and Its Application to Test Generation. InProceedings of the Computer Safety, Reliability, and Security (SAFECOMP-Workshops, 2020, Vol. 12235). Springer, 289–301

  3. [3]

    Zohreh Aghababaeyan, Manel Abdellatif, Mahboubeh Dadkhah, and Lionel C. Briand. 2024. DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 6 (2024), 158

  4. [4]

    Raid Rafi Omar Al-Nima, Tingting Han, Saadoon A. M. Al-Sumaidaee, Taolue Chen, and Wai Lok Woo. 2021. Robustness and performance of Deep Reinforcement Learning.Applied Soft Computing (ASC)105 (2021), 107295

  5. [5]

    Fahmy, Fabrizio Pastore, and Lionel C

    Mohammed Walid Attaoui, Hazem M. Fahmy, Fabrizio Pastore, and Lionel C. Briand. 2023. Black-box Safety Analysis and Retraining of DNNs based on Feature Extraction and Clustering.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 3 (2023), 79:1–79:40

  6. [6]

    Tongtong Bai, Song Huang, Yifan Huang, Xingya Wang, Chunyan Xia, Yubin Qu, and Zhen Yang. 2024. CriticalFuzz: A critical neuron coverage-guided fuzz testing framework for deep neural networks.Information and Software Technology (IST)172 (2024), 107476

  7. [7]

    Agathe Balayn, Natasa Rikalo, Jie Yang, and Alessandro Bozzon. 2023. Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI, 2023). ACM, 11:1–11:20

  8. [8]

    Houssem Ben Braiek and Foutse Khomh. 2019. DeepEvolution: A Search-Based Testing Approach for Deep Neural Networks. In Proceedings of the 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME, 2019). IEEE, 454–458

Show all 177 references
  1. [9]

    Lucas, Peter I

    Cameron Browne, Edward Jack Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez Liebana, Spyridon Samothrakis, and Simon Colton. 2012. A Survey of Monte Carlo Tree Search Methods.IEEE Transactions on Computational Inte...

  2. [10]

    Taejoon Byun and Sanjai Rayadurgam. 2020. Manifold for machine learning assurance. InProceedings of the 42nd International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER, 2020). ACM, 97–100

  3. [11]

    Taejoon Byun, Vaibhav Sharma, Abhishek Vijayakumar, Sanjai Rayadurgam, and Darren D. Cofer. 2019. Input Prioritization for Testing Neural Networks. InProceedings of the IEEE International Conference On Artificial Intelligence Testing (AITest, 2019). IEEE, 63–70

  4. [12]

    Cristian Cadar, Daniel Dunbar, and Dawson R. Engler. 2008. KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs. InProceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI, 2008). USENIX Association, 209–224

  5. [13]

    Sooyoung Cha, Seongjoon Hong, Junhee Lee, and Hakjoo Oh. 2018. Automatically generating search heuristics for concolic testing. In Proceedings of the 40th International Conference on Software Engineering (ICSE, 2018). ACM, 1244–1254

  6. [14]

    Guillaume M. J-B. Chaslot, Mark H. M. Winands, H. Jaap van den Herik, Jos W. H. M. Uiterwijk, and Bruno Bouzy. 2008. Progressive Strategies for Monte-Carlo Tree Search.New Mathematics and Natural Computation04, 03 (2008), 343–357

  7. [15]

    Jialuo Chen, Jingyi Wang, Xingjun Ma, Youcheng Sun, Jun Sun, Peixin Zhang, and Peng Cheng. 2023. QuoTe: Quality-oriented Testing for Deep Learning Systems.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 5 (2023), 125:1–125:33

  8. [16]

    Junjie Chen, Zhuo Wu, Zan Wang, Hanmo You, Lingming Zhang, and Ming Yan. 2020. Practical Accuracy Estimation for Efficient Deep Neural Network Testing.ACM Transactions on Software Engineering and Methodology (TOSEM)29, 4 (2020), 30:1–30:35

  9. [17]

    Tsong Yueh Chen, Shing-Chi Cheung, and Siu-Ming Yiu. 2020. Metamorphic Testing: A New Approach for Generating Next Test Cases. CoRRabs/2002.12543 (2020). arXiv:2002.12543 https://arxiv.org/abs/2002.12543

  10. [18]

    Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, Pak-Lok Poon, Dave Towey, T. H. Tse, and Zhi Quan Zhou. 2018. Metamorphic Testing: A Review of Challenges and Opportunities.ACM Computing Surveys (CSUR)51, 1 (2018), 4:1–4:27

  11. [19]

    Pranav Singh Chib and Pravendra Singh. 2024. Recent Advancements in End-to-End Autonomous Driving Using Deep Learning: A Survey.IEEE Transactions on Intelligent Vehicles (TIV)9, 1 (2024), 103–118

  12. [20]

    Hepeng Dai, Chang-Ai Sun, and Huai Liu. 2022. DeepController: Feedback-Directed Fuzzing for Deep Learning Systems. InProceedings of the 34th International Conference on Software Engineering and Knowledge Engineering (SEKE, 2022). KSI Research Inc., 531–536

  13. [21]

    Hepeng Dai, Chang-Ai Sun, Huai Liu, and Xiangyu Zhang. 2024. DFuzzer: Diversity-Driven Seed Queue Construction of Fuzzing for Deep Learning Models.IEEE Transactions on Reliability (TR)73, 2 (2024), 1075–1089

  14. [22]

    Bissyandé, and Yves Le Traon

    Xueqi Dang, Yinghua Li, Mike Papadakis, Jacques Klein, Tegawendé F. Bissyandé, and Yves Le Traon. 2024. GraphPrior: Mutation-based Test Input Prioritization for Graph Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 1 (2024), 22:1–22:40

  15. [23]

    Samet Demir, Hasan Ferit Eniser, and Alper Sen. 2020. DeepSmartFuzzer: Reward Guided Test Generation For Deep Learning. In Proceedings of the Workshop on Artificial Intelligence Safety co-located with the 29th International Joint Conference on Artificial Intelligence and the 1...

  16. [24]

    Dwyer, and Mary Lou Soffa

    Swaroopa Dola, Matthew B. Dwyer, and Mary Lou Soffa. 2021. Distribution-Aware Testing of Neural Networks Using Generative Models. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE 2021). IEEE, 226–237

  17. [25]

    Dwyer, and Mary Lou Soffa

    Swaroopa Dola, Matthew B. Dwyer, and Mary Lou Soffa. 2023. Input Distribution Coverage: Measuring Feature Interaction Adequacy in Neural Network Testing.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 3 (2023), 81:1–81:48. Proc. ACM Meas. Anal. Comput. Syst...

  18. [26]

    Dwyer, and Mary Lou Soffa

    Swaroopa Dola, Rory McDaniel, Matthew B. Dwyer, and Mary Lou Soffa. 2024. CIT4DNN: Generating Diverse and Rare Inputs for Neural Networks Using Latent Space Combinatorial Testing. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE 2024). ...

  19. [27]

    Yizhen Dong, Peixin Zhang, Jingyi Wang, Shuang Liu, Jun Sun, Jianye Hao, Xinyu Wang, Li Wang, Jin Song Dong, and Ting Dai

  20. [28]

    Alexandre Donzé and Oded Maler. 2010. Robust Satisfaction of Temporal Logic over Real-Valued Signals. InProceedings of the 8th international conference on Formal modeling and analysis of timed systems (FORMATS, 2010, Vol. 6246). Springer, 92–106

  21. [29]

    Chengwen Du and Tao Chen. 2024. Contexts Matter: An Empirical Study on Contextual Influence in Fairness Testing for Deep Learning Systems. InProceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM, 2024). ACM, 107–118

  22. [30]

    Xiaoning Du, Xiaofei Xie, Yi Li, Lei Ma, Yang Liu, and Jianjun Zhao. 2019. DeepStellar: model-based quantitative analysis of stateful deep learning systems. InProceedings of the 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations...

  23. [31]

    Xiaoning Du, Xiaofei Xie, Yi Li, Lei Ma, Jianjun Zhao, and Yang Liu. 2018. DeepCruiser: Automated Guided Testing for Stateful Deep Learning Systems.CoRRabs/1812.05339 (2018). arXiv:1812.05339 http://arxiv.org/abs/1812.05339

  24. [32]

    Fahmy, Fabrizio Pastore, Lionel C

    Hazem M. Fahmy, Fabrizio Pastore, Lionel C. Briand, and Thomas Stifter. 2023. Simulator-based Explanation and Debugging of Hazard-triggering Events in DNN-based Safety-critical Systems.ACM Transactions on Software Engineering and Methodology (TOSEM) 32, 4 (2023), 104:1–104:47

  25. [33]

    Fainekos and George J

    Georgios E. Fainekos and George J. Pappas. 2009. Robustness of temporal logic specifications for continuous-time signals.Theoretical Computer Science (TCS)410, 42 (2009), 4262–4291

  26. [34]

    Yang Feng, Qingkai Shi, Xinyu Gao, Jun Wan, Chunrong Fang, and Zhenyu Chen. 2020. DeepGini: prioritizing massive tests to enhance the robustness of deep neural networks. InProceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2020...

  27. [35]

    Xinyu Gao, Yang Feng, Yining Yin, Zixi Liu, Zhenyu Chen, and Baowen Xu. 2022. Adaptive Test Selection for Deep Neural Networks. InProceedings of the IEEE/ACM 44th International Conference on Software Engineering (ICSE, 2022). ACM, 73–85

  28. [36]

    Saha, Mukul R

    Xiang Gao, Ripon K. Saha, Mukul R. Prasad, and Abhik Roychoudhury. 2020. Fuzz testing based data augmentation to improve robustness of deep neural networks. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering (ICSE, 2020). ACM, 1147–1158

  29. [37]

    Xuanqi Gao, Juan Zhai, Shiqing Ma, Chao Shen, Yufei Chen, and Qian Wang. 2022. Fairneuron: Improving Deep Neural Network Fairness with Adversary Games on Selective Neurons. InProceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE, 2022). A...

  30. [38]

    Simos Gerasimou, Hasan Ferit Eniser, Alper Sen, and Alper Cakan. 2020. Importance-driven deep learning system testing. InProceedings of the 42nd International Conference on Software Engineering (ICSE, 2020). ACM, 702–713

  31. [39]

    Goodfellow, Jonathon Shlens, and Christian Szegedy

    Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. InProceedings of the 3rd International Conference on Learning Representations (ICLR, 2015)

  32. [40]

    Matthew Groh, Omar Badri, Roxana Daneshjou, Arash Koochek, Caleb Harris, Luis R Soenksen, P Murali Doraiswamy, and Rosalind Picard. 2024. Deep learning-aided decision support for diagnosis of skin disease across skin tones.Nature Medicine30, 2 (2024), 573–583

  33. [41]

    Hongjing Guo, Chuanqi Tao, and Zhiqiu Huang. 2023. Multi-Objective White-Box Test Input Selection for Deep Neural Network Model Enhancement. InProceedings of the 34th IEEE International Symposium on Software Reliability Engineering (ISSRE, 2023). IEEE, 521–532

  34. [42]

    Hongjing Guo, Chuanqi Tao, and Zhiqiu Huang. 2024. Neuron importance-aware coverage analysis for deep neural network testing. Empirical Software Engineering (EMSE)29, 5 (2024), 118

  35. [43]

    Jianmin Guo, Yu Jiang, Yue Zhao, Quan Chen, and Jiaguang Sun. 2018. DLFuzz: differential fuzzing testing of deep learning systems. In Proceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering ...

  36. [44]

    Jianmin Guo, Quan Zhang, Yue Zhao, Heyuan Shi, Yu Jiang, and Jia-Guang Sun. 2022. RNN-Test: Towards Adversarial Testing for Recurrent Neural Network Systems.IEEE Transactions on Software Engineering (TSE)48, 10 (2022), 4167–4180

  37. [45]

    Ge Han, Zheng Li, Peng Tang, Chengyu Hu, and Shanqing Guo. 2022. FuzzGAN: A Generation-Based Fuzzing Framework for Testing Deep Neural Networks. InProceedings of the 24th IEEE Int Conf on High Performance Computing & Communications (HPCC, 2022). IEEE, 1601–1608

  38. [46]

    Yao Hao, Zhiqiu Huang, Hongjing Guo, and Guohua Shen. 2023. Test Input Selection for Deep Neural Network Enhancement Based on Multiple-Objective Optimization. InProceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2023). IE...

  39. [47]

    Fabrice Harel-Canada, Lingxiao Wang, Muhammad Ali Gulzar, Quanquan Gu, and Miryung Kim. 2020. Is neuron coverage a meaningful measure for testing deep neural networks?. InProceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundation...

  40. [48]

    Adrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish, Mathias Payer, and Antony L. Hosking. 2021. Seed selection for successful fuzzing. InProceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2021). ACM, 230–243

  41. [49]

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. InProceedings of the 31st International Conference on Neural Information Processing Systems (NIPS...

  42. [50]

    Qiang Hu, Yuejun Guo, Maxime Cordy, Xiaofei Xie, Lei Ma, Mike Papadakis, and Yves Le Traon. 2022. An Empirical Study on Data Distribution-Aware Test Selection for Deep Learning Enhancement.ACM Transactions on Software Engineering and Methodology (TOSEM)31, 4 (2022), 78:1–78:30

  43. [51]

    Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Lei Ma, Mike Papadakis, and Yves Le Traon. 2024. Test Optimization in DNN Testing: A Survey.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 4 (2024), 111:1–111:42

  44. [52]

    Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Lei Ma, and Yves Le Traon. 2023. Aries: Efficient Testing of Deep Neural Networks via Labeling-Free Accuracy Estimation. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, ...

  45. [53]

    Qianchao Hu, Feng Wang, Binglin Liu, and Haitian Liu. 2023. Research on Deep Neural Network Testing Techniques. InProceedings of the 4th International Conference on Machine Learning and Computer Application (ICMLCA, 2023). ACM, 113–119

  46. [54]

    Shengyou Hu, Huayao Wu, Peng Wang, Jing Chang, Yongjun Tu, Xiu Jiang, Xintao Niu, and Changhai Nie. 2023. ATOM: Automated Black-Box Testing of Multi-Label Image Classification Systems. InProceedings of the 38th IEEE/ACM International Conference on Automated Software Engineerin...

  47. [55]

    Dong Huang, Tsz On Li, Xiaofei Xie, and Heming Cui. 2024. Themis: Automatic and Efficient Deep Learning System Testing with Strong Fault Detection Capability. InProceedings of the 35th IEEE International Symposium on Software Reliability Engineering (ISSRE, 2024). IEEE, 451–462

  48. [56]

    Wei Huang, Youcheng Sun, Xingyu Zhao, James Sharp, Wenjie Ruan, Jie Meng, and Xiaowei Huang. 2022. Coverage-Guided Testing for Recurrent Neural Networks.IEEE Transactions on Reliability (TR)71, 3 (2022), 1191–1206

  49. [57]

    Xiaowei Huang, Daniel Kroening, Wenjie Ruan, James Sharp, Youcheng Sun, Emese Thamo, Min Wu, and Xinping Yi. 2020. A survey of safety and trustworthiness of deep neural networks: Verification, testing, adversarial attack and defence, and interpretability.Computer Science Revie...

  50. [58]

    Belongie, and Jan Kautz

    Xun Huang, Ming-Yu Liu, Serge J. Belongie, and Jan Kautz. 2018. Multimodal Unsupervised Image-to-Image Translation. InProceedings of the 15th European Conference on Computer Vision (ECCV, 2018, Vol. 11207). Springer, 179–196

  51. [59]

    Gunel Jahangirova and Paolo Tonella. 2020. An Empirical Evaluation of Mutation Operators for Deep Learning Systems. InProceedings of the 13th IEEE International Conference on Software Testing, Validation and Verification (ICST, 2020). IEEE, 74–84

  52. [60]

    Zhenlan Ji, Pingchuan Ma, Yuanyuan Yuan, and Shuai Wang. 2023. CC: Causality-Aware Coverage Criterion for Deep Neural Networks. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 1788–1800

  53. [61]

    Yue Jia and Mark Harman. 2011. An Analysis and Survey of the Development of Mutation Testing.IEEE Transactions on Software Engineering, TSE37, 5 (2011), 649–678

  54. [62]

    Zhonghao Jiang, Meng Yan, Li Huang, Weifeng Sun, Chao Liu, Song Sun, and David Lo. 2025. DeepVec: State-Vector Aware Test Case Selection for Enhancing Recurrent Neural Network.IEEE Transactions on Software Engineering (TSE)51, 6 (2025), 1702–1723

  55. [63]

    Jinhan Kim, Robert Feldt, and Shin Yoo. 2019. Guiding deep learning system testing using surprise adequacy. InProceedings of the 41st International Conference on Software Engineering (ICSE, 2019). IEEE / ACM, 1039–1049

  56. [64]

    Jinhan Kim, Robert Feldt, and Shin Yoo. 2023. Evaluating Surprise Adequacy for Deep Learning System Testing.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 2 (2023), 42:1–42:29

  57. [65]

    Jinhan Kim, Jeongil Ju, Robert Feldt, and Shin Yoo. 2020. Reducing DNN labelling cost using surprise adequacy: an industrial case study for autonomous driving. InProceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundati...

  58. [66]

    Seah Kim and Shin Yoo. 2020. Evaluating Surprise Adequacy for Question Answering. InProceedings of the 42nd International Conference on Software Engineering Workshops (ICSE Workshops, 2020). ACM, 197–202

  59. [67]

    George Klees, Andrew Ruef, Benji Cooper, Shiyi Wei, and Michael Hicks. 2018. Evaluating Fuzz Testing. InProceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security (CCS, 2018). ACM, 2123–2138

  60. [68]

    Richard Kuhn, Itzel Dominguez Mendoza, Raghu Kacker, and Yu Lei

    D. Richard Kuhn, Itzel Dominguez Mendoza, Raghu Kacker, and Yu Lei. 2013. Combinatorial Coverage Measurement Concepts and Applications. InProceedings of the IEEE 6th International Conference on Software Testing, Verification and Validation Workshops (ICST, 2013). IEEE Computer...

  61. [69]

    Goodfellow, and Samy Bengio

    Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. 2017. Adversarial examples in the physical world. InProceedings of the 5th International Conference on Learning Representations (ICLR, 2017). OpenReview.net

  62. [70]

    Seokhyun Lee, Sooyoung Cha, Dain Lee, and Hakjoo Oh. 2020. Effective white-box testing of deep neural networks with adaptive neuron-selection strategy. InProceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2020). ACM, 165–176

  63. [71]

    Bingdong Li, Jinlong Li, Ke Tang, and Xin Yao. 2015. Many-Objective Evolutionary Algorithms: A Survey.ACM Computing Surveys (CSUR)48, 1 (2015), 13:1–13:35

  64. [72]

    Zenan Li, Xiaoxing Ma, Chang Xu, and Chun Cao. 2019. Structural coverage criteria for neural networks could be misleading. In Proceedings of the 41st International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER, 2019). IEEE / ACM, 89–92

  65. [73]

    Zenan Li, Xiaoxing Ma, Chang Xu, Chun Cao, Jingwei Xu, and Jian Lü. 2019. Boosting operational DNN testing efficiency through conditioning. InProceedings of the ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineeri...

  66. [74]

    Zhong Li, Minxue Pan, Tian Zhang, and Xuandong Li. 2021. Testing DNN-based Autonomous Driving Systems under Critical Environmental Conditions. InProceedings of the 38th International Conference on Machine Learning - Volume 139 (ICML, 2021). PMLR, 6471–6482

  67. [75]

    Zhong Li, Zhengfeng Xu, Ruihua Ji, Minxue Pan, Tian Zhang, Linzhang Wang, and Xuandong Li. 2024. Distance-Aware Test Input Selection for Deep Neural Networks. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2024). ACM, 248–260

  68. [76]

    Weiguang Liu, Senlin Luo, Limin Pan, and Zhao Zhang. 2025. DeepCNP: An efficient white-box testing of deep neural networks by aligning critical neuron paths.Information and Software Technology179 (2025), 107640

  69. [77]

    Yibing Liu, Chris Xing Tian, Haoliang Li, Lei Ma, and Shiqi Wang. 2024. Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11,

  70. [78]

    Zixi Liu, Yang Feng, Yining Yin, and Zhenyu Chen. 2022. DeepState: Selecting Test Suites to Enhance the Robustness of Recurrent Neural Networks. InProceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE, 2022). ACM, 598–609

  71. [79]

    Lei Ma, Felix Juefei-Xu, Minhui Xue, Bo Li, Li Li, Yang Liu, and Jianjun Zhao. 2019. DeepCT: Tomographic Combinatorial Testing for Deep Learning Systems. InProceedings of the 26th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2019). IE...

  72. [80]

    Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chunyang Chen, Ting Su, Li Li, Yang Liu, Jianjun Zhao, and Yadong Wang. 2018. DeepGauge: multi-granularity testing criteria for deep learning systems. InProceedings of the 33rd ACM/IEEE International Confere...

  73. [81]

    Pingchuan Ma, Shuai Wang, and Jin Liu. 2020. Metamorphic Testing and Certified Mitigation of Fairness Violations in NLP Models. In Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI, 2020). ijcai.org, 458–465

  74. [82]

    Shiqing Ma, Yingqi Liu, Wen-Chuan Lee, Xiangyu Zhang, and Ananth Grama. 2018. MODE: automated neural network model debugging via state differential analysis and input selection. InProceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposi...

  75. [83]

    Wei Ma, Mike Papadakis, Anestis Tsakmalis, Maxime Cordy, and Yves Le Traon. 2021. Test Selection for Deep Learning Systems.ACM Transactions on Software Engineering and Methodology (TOSEM)30, 2 (2021), 13:1–13:22

  76. [84]

    Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards Deep Learning Models Resistant to Adversarial Attacks. InInternational Conference on Learning Representations. https://openreview.net/forum?id=rJzIBfZAb

  77. [85]

    Sanoop Mallissery and Yu-Sung Wu. 2024. Demystify the Fuzzing Methods: A Comprehensive Survey.ACM Computing Surveys (CSUR) 56, 3 (2024), 71:1–71:38

  78. [86]

    Senthil Mani, Anush Sankaran, Srikanth Tamilselvam, and Akshay Sethi. 2019. Coverage Testing of Deep Learning Models using Dataset Characterization.CoRRabs/1911.07309 (2019). arXiv:1911.07309 http://arxiv.org/abs/1911.07309

  79. [87]

    Sondess Missaoui, Simos Gerasimou, and Nicholas Matragkas. 2023. Semantic Data Augmentation for Deep Learning Testing Using Generative AI. InProceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2023). IEEE, 1694–1698

  80. [88]

    Vasilii Mosin, Miroslaw Staron, Darko Durisic, Francisco Gomes de Oliveira Neto, Sushant Kumar Pandey, and Ashok Chaitanya Koppisetty. 2022. Comparing Input Prioritization Techniques for Testing Deep Learning Algorithms. InProceedings of the 48th Euromicro Conference on Softwa...

  81. [89]

    de Albuquerque

    Khan Muhammad, Amin Ullah, Jaime Lloret, Javier Del Ser, and Victor Hugo C. de Albuquerque. 2021. Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions.IEEE Transactions on Intelligent Transportation Systems, (TITS)22, 7 (2021), 4316–4336. Proc. A...

  82. [90]

    Briand, and Yvan Labiche

    Daniel Di Nardo, Nadia Alshahwan, Lionel C. Briand, and Yvan Labiche. 2015. Coverage-based regression test case selection, minimization and prioritization: a case study on an industrial system.Software Testing, Verification and Reliability (STVR)25, 4 (2015), 371–396

  83. [91]

    Neelofar and Aldeida Aleti. 2024. Towards Reliable AI: Adequacy Metrics for Ensuring the Quality of System-level Testing of Autonomous Vehicles. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE, 2024). ACM, 68:1–68:12

  84. [92]

    Mahdi Nejadgholi and Jinqiu Yang. 2019. A Study of Oracle Approximations in Testing Deep Learning Libraries. InProceedings of the 34th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2019). IEEE, 785–796

  85. [93]

    Andersen, and Ian J

    Augustus Odena, Catherine Olsson, David G. Andersen, and Ian J. Goodfellow. 2019. TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing. InProceedings of the 36th International Conference on Machine Learning - Volume 97 (ICML, 2019). PMLR, 4901–4911

  86. [94]

    Mallikarjuna Paramesha, Nitin Liladhar Rane, and Jayesh Rane. 2024. Artificial Intelligence, Machine Learning, Deep Learning, and Blockchain in Financial and Banking Services: A Comprehensive Review.Partners Universal Multidisciplinary Research Journal1, 2 (Jul. 2024), 51–67

  87. [95]

    Leo Hyun Park, Soochang Chung, Jaeuk Kim, and Taekyoung Kwon. 2023. GradFuzz: Fuzzing deep neural networks with gradient vector coverage for adversarial examples.Neurocomputing522 (2023), 165–180

  88. [96]

    Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017. DeepXplore: Automated Whitebox Testing of Deep Learning Systems. In Proceedings of the 26th Symposium on Operating Systems Principles (SOSP, 2017). ACM, 1–18

  89. [97]

    Goran Petrovic, Marko Ivankovic, Gordon Fraser, and René Just. 2021. Does mutation testing improve testing practices?. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE, 2021). IEEE, 910–921

  90. [98]

    Hung Viet Pham, Thibaud Lutellier, Weizhen Qi, and Lin Tan. 2019. CRADLE: cross-backend validation to detect and localize bugs in deep learning libraries. InProceedings of the 41st International Conference on Software Engineering (ICSE, 2019). IEEE / ACM, 1027–1038

  91. [99]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings ...

  92. [100]

    Alexandre Rebert, Sang Kil Cha, Thanassis Avgerinos, Jonathan Foote, David Warren, Gustavo Grieco, and David Brumley. 2014. Optimizing Seed Selection for Fuzzing. InProceedings of the 23rd USENIX Security Symposium (USENIX Security, 2014). USENIX Association, 861–875

  93. [101]

    Vincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, and Paolo Tonella. 2021. DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation Score. InProceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2021). IEEE, 355–367

  94. [102]

    Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella. 2020. Testing machine learning based systems: a systematic mapping.Empirical Software Engineering (EMSE)25, 6 (2020), 5193–5254

  95. [103]

    Vincenzo Riccio and Paolo Tonella. 2023. When and Why Test Generators for Deep Learning Produce Invalid Inputs: an Empirical Study. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 1161–1173

  96. [104]

    Buttazzo

    Giulio Rossolini, Alessandro Biondi, and Giorgio C. Buttazzo. 2023. Increasing the Confidence of Deep Neural Networks by Coverage Analysis.IEEE Transactions on Software Engineering (TSE)49, 2 (2023), 802–815

  97. [105]

    Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen

    Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved Techniques for Training GANs. InProceedings of the 30th International Conference on Neural Information Processing Systems (NIPS, 2016). 2226–2234

  98. [106]

    Sergio Segura, Gordon Fraser, Ana Belén Sánchez, and Antonio Ruiz Cortés. 2016. A Survey on Metamorphic Testing.IEEE Transactions on Software Engineering (TSE)42, 9 (2016), 805–824

  99. [107]

    Dwyer, and Yanjun Qi

    Arshdeep Sekhon, Yangfeng Ji, Matthew B. Dwyer, and Yanjun Qi. 2022. White-box Testing of NLP models with Mask Neuron Coverage. InProceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL, 2022). Association for Computa...

  100. [108]

    Koushik Sen, Darko Marinov, and Gul Agha. 2005. CUTE: a concolic unit testing engine for C. InProceedings of the 10th Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE, 2005). ACM, 263–272

  101. [109]

    Weijun Shen, Yanhui Li, Lin Chen, Yuanlei Han, Yuming Zhou, and Baowen Xu. 2020. Multiple-Boundary Clustering and Prioritization to Promote Neural Network Retraining. InProceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2020). IEE...

  102. [110]

    Ying Shi, Beibei Yin, and Jing-Ao Shi. 2025. Markov model based coverage testing of deep learning software systems.Information and Software Technology (IST)179 (2025), 107628

  103. [111]

    Ying Shi, Beibei Yin, and Zheng Zheng. 2024. Multi-granularity coverage criteria for deep reinforcement learning systems.Journal of Systems and Software (JSS)212 (2024), 112016

  104. [112]

    Ying Shi, Beibei Yin, Zheng Zheng, and Tiancheng Li. 2021. An Empirical Study on Test Case Prioritization Metrics for Deep Neural Networks. InProceedings of the 21st IEEE International Conference on Software Quality, Reliability and Security (QRS, 2021). IEEE, Proc. ACM Meas. ...

  105. [113]

    Jiaze Sun, Juan Li, and Sulei Wen. 2023. DeepMC: DNN test sample optimization method jointly guided by misclassification and coverage.Appl. Intell.53, 12 (2023), 15787–15801

  106. [114]

    Weidi Sun, Yuteng Lu, Xiaokun Luan, and Meng Sun. 2023. HeatC: A Variable-Grained Coverage Criterion for Deep Learning Systems. InProceedings of the 9th International Symposium on Dependable Software Engineering. Theories, Tools, and Applications (SETTA, 2023, Vol. 14464). Spr...

  107. [115]

    Weidi Sun, Yuteng Lu, and Meng Sun. 2021. Are Coverage Criteria Meaningful Metrics for DNNs?. InProceedings of the International Joint Conference on Neural Networks (IJCNN, 2021). IEEE, 1–8

  108. [116]

    Weidi Sun, Xiaoyong Xue, Yuteng Lu, Jia Zhao, and Meng Sun. 2023. HashC: Making deep learning coverage testing finer and faster. Journal of Systems Architecture (JSA)144 (2023), 102999

  109. [117]

    Weifeng Sun, Meng Yan, Zhongxin Liu, and David Lo. 2023. Robust Test Selection for Deep Neural Networks.IEEE Transactions on Software Engineering (TSE)49, 12 (2023), 5250–5278

  110. [118]

    Youcheng Sun, Xiaowei Huang, Daniel Kroening, James Sharp, Matthew Hill, and Rob Ashmore. 2019. Structural Test Coverage Criteria for Deep Neural Networks.ACM Transactions on Embedded Computing Systems (TECS)18, 5s (2019), 94:1–94:23

  111. [119]

    Youcheng Sun, Min Wu, Wenjie Ruan, Xiaowei Huang, Marta Kwiatkowska, and Daniel Kroening. 2018. Concolic testing for deep neural networks. InProceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (ASE, 2018). ACM, 109–119

  112. [120]

    Chuanqi Tao, Yali Tao, Hongjing Guo, Zhiqiu Huang, and Xiaobing Sun. 2023. DLRegion: Coverage-guided fuzz testing of deep neural networks with region-based neuron selection strategies.Information and Software Technology (IST)162 (2023), 107266

  113. [121]

    Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. 2018. DeepTest: automated testing of deep-neural-network-driven autonomous cars. InProceedings of the 40th International Conference on Software Engineering (ICSE, 2018). ACM, 303–314

  114. [122]

    Yongqiang Tian, Wuqi Zhang, Ming Wen, Shing-Chi Cheung, Chengnian Sun, Shiqing Ma, and Yu Jiang. 2023. Finding Deviated Behaviors of the Compressed DNN Models for Image Classifications.ACM Transactions on Software Engineering and Methodology (TOSEM)32, 5 (2023), 128:1–128:32

  115. [123]

    Miller Trujillo, Mario Linares-Vásquez, Camilo Escobar-Velásquez, Ivana Dusparic, and Nicolás Cardozo. 2020. Does Neuron Coverage Matter for Deep Reinforcement Learning?: A Preliminary Study. InProceedings of the IEEE/ACM 42nd International Conference on Software Engineering W...

  116. [124]

    Pasareanu

    Muhammad Usman, Youcheng Sun, Divya Gopinath, Rishi Dange, Luca Manolache, and Corina S. Pasareanu. 2023. An overview of structural coverage metrics for testing neural networks.International Journal on Software Tools for Technology Transfer (STTT)25, 3 (2023), 393–405

  117. [125]

    Xiaohui Wan, Tiancheng Li, Weibin Lin, Yi Cai, and Zheng Zheng. 2024. Coverage-guided fuzzing for deep reinforcement learning systems.Journal of Systems and Software (JSS)210 (2024), 111963

  118. [126]

    Dong Wang, Ziyuan Wang, Chunrong Fang, Yanshan Chen, and Zhenyu Chen. 2019. DeepPath: Path-Driven Testing Criteria for Deep Neural Networks. InProceedings of the IEEE International Conference On Artificial Intelligence Testing (AITest, 2019). IEEE, 119–120

  119. [127]

    Jingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma, Dongxia Wang, Jun Sun, and Peng Cheng. 2021. RobOT: Robustness-Oriented Testing for Deep Learning Systems. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE, 2021). IEEE, 300–311

  120. [128]

    Peng Wang, Shengyou Hu, Huayao Wu, Xintao Niu, Changhai Nie, and Lin Chen. 2024. A Combinatorial Interaction Testing Method for Multi-Label Image Classifier. InProceedings of the 35th IEEE International Symposium on Software Reliability Engineering (ISSRE 2024). IEEE, 463–474

  121. [129]

    Eric Wong

    Shengrong Wang, Dongcheng Li, Hui Li, Man Zhao, and W. Eric Wong. 2024. A Survey on Test Input Selection and Prioritization for Deep Neural Networks. InProceedings of the 10th International Symposium on System Security, Safety, and Reliability (ISSSR, 2024). 232–243

  122. [130]

    Shuai Wang and Zhendong Su. 2020. Metamorphic Object Insertion for Testing Object Detection Systems. InProceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2020). IEEE, 1053–1065

  123. [131]

    Zhiyu Wang, Sihan Xu, Lingling Fan, Xiangrui Cai, Linyu Li, and Zheli Liu. 2024. Can Coverage Criteria Guide Failure Discovery for Image Classifiers? An Empirical Study.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 7, Article 190 (Sept. 2024), 28 pages

  124. [132]

    Zan Wang, Hanmo You, Junjie Chen, Yingyi Zhang, Xuyuan Dong, and Wenbin Zhang. 2021. Prioritizing Test Inputs for Deep Neural Networks via Mutation Analysis. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE, 2021). IEEE, 397–409

  125. [133]

    Moshi Wei, Yuchao Huang, Jinqiu Yang, Junjie Wang, and Song Wang. 2023. CoCoFuzzing: Testing Neural Code Models With Coverage-Guided Fuzzing.IEEE Transactions on Reliability (TR)72, 3 (2023), 1276–1289

  126. [134]

    Zhengyuan Wei and W. K. Chan. 2021. Fuzzing Deep Learning Models against Natural Robustness with Filter Coverage. InProceedings of the 21st IEEE International Conference on Software Quality, Reliability and Security (QRS, 2021). IEEE, 608–619. Proc. ACM Meas. Anal. Comput. Sys...

  127. [135]

    Michael Weiss and Paolo Tonella. 2022. Simple techniques work surprisingly well for neural network test prioritization and active learning (replicability study). InProceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2022). ACM, 139–150

  128. [136]

    Tsui-Wei Weng, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, Dong Su, Yupeng Gao, Cho-Jui Hsieh, and Luca Daniel. 2018. Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach. InInternational Conference on Learning Representations. https://openreview.net/forum?i...

  129. [137]

    Xiaoxue Wu, Jinjin Shen, Wei Zheng, Lidan Lin, Yulei Sui, and Abubakar Omari Abdallah Semasaba. 2023. RNNtcs: A test case selection method for Recurrent Neural Networks.Knowledge-Based Systems (KBS)279 (2023), 110955

  130. [138]

    Dongwei Xiao, Zhibo Liu, Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2022. Metamorphic Testing of Deep Learning Compilers. Proceedings of the ACM on Measurement and Analysis of Computing Systems (POMACS)6, 1 (2022), 15:1–15:28

  131. [139]

    Tao Xie, Darko Marinov, Wolfram Schulte, and David Notkin. 2005. Symstra: A Framework for Generating Object-Oriented Unit Tests Using Symbolic Execution. InProceedings of the 11th International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TA...

  132. [140]

    Xiaofei Xie, Tianlin Li, Jian Wang, Lei Ma, Qing Guo, Felix Juefei-Xu, and Yang Liu. 2022. NPC: Neuron Path Coverage via Characterizing Decision Logic of Deep Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)31, 3 (2022), 47:1–47:27

  133. [141]

    Xiaofei Xie, Lei Ma, Felix Juefei-Xu, Minhui Xue, Hongxu Chen, Yang Liu, Jianjun Zhao, Bo Li, Jianxiong Yin, and Simon See. 2019. DeepHunter: a coverage-guided fuzz testing framework for deep neural networks. InProceedings of the 28th ACM SIGSOFT International Symposium on Sof...

  134. [142]

    Xiaofei Xie, Lei Ma, Haijun Wang, Yuekang Li, Yang Liu, and Xiaohong Li. 2019. DiffChaser: Detecting Disagreements for Deep Neural Networks. InProceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI, 2019). ijcai.org, 5772–5778

  135. [143]

    Yutao Xie, Jiayi Lin, Hande Dong, Lei Zhang, and Zhonghai Wu. 2024. Survey of Code Search Based on Deep Learning.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 2 (2024), 54:1–54:42

  136. [144]

    Huan Xu and Shie Mannor. 2012. Robustness and generalization.Mach. Learn.86, 3 (2012), 391–423

  137. [145]

    Ahmed Haj Yahmed, Houssem Ben Braiek, Foutse Khomh, Sonia Bouzidi, and Rania Zaatour. 2022. DiverGet: a Search-Based Software Testing approach for Deep Neural Network Quantization assessment.Empirical Software Engineering (EMSE)27, 7 (2022), 193

  138. [146]

    Yoriyuki Yamagata, Shuang Liu, Takumi Akazaki, Yihai Duan, and Jianye Hao. 2021. Falsification of Cyber-Physical Systems Using Deep Reinforcement Learning.IEEE Transactions on Software Engineering (TSE)47, 12 (2021), 2823–2840

  139. [147]

    Ming Yan, Junjie Chen, Xuejie Cao, Zhuo Wu, Yuning Kang, and Zan Wang. 2024. Revisiting deep neural network test coverage from the test effectiveness perspective.Journal of Software: Evolution and Process (JSEP)36, 4 (2024)

  140. [148]

    Shenao Yan, Guanhong Tao, Xuwei Liu, Juan Zhai, Shiqing Ma, Lei Xu, and Xiangyu Zhang. 2020. Correlations between deep neural network model coverage criteria and model quality. InProceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposiu...

  141. [149]

    Minghao Yang, Shunkun Yang, and Wenda Wu. 2022. DeepRTest: A Vulnerability-Guided Robustness Testing and Enhancement Framework for Deep Neural Networks. InProceedings of the 22nd IEEE International Conference on Software Quality, Reliability and Security (QRS, 2022). IEEE, 754–762

  142. [150]

    Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2022. Revisiting Neuron Coverage Metrics and Quality of Deep Neural Networks. InProceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2022). IEEE, 408–419

  143. [151]

    Aoshuang Ye, Lina Wang, Lei Zhao, and Jianpeng Ke. 2022. DANCe: Dynamic Adaptive Neuron Coverage for Fuzzing Deep Neural Networks. InProceedings of the International Joint Conference on Neural Networks (IJCNN, 2022). IEEE, 1–8

  144. [152]

    Shin Yoo and Mark Harman. 2007. Pareto efficient multi-objective test case selection. InProceedings of the ACM/SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2007). ACM, 140–150

  145. [153]

    Shin Yoo and Mark Harman. 2012. Regression testing minimization, selection and prioritization: a survey.Software Testing, Verification and Reliability (STVR)22, 2 (2012), 67–120

  146. [154]

    Hanmo You, Zan Wang, Junjie Chen, Shuang Liu, and Shuochuan Li. 2023. Regression Fuzzing for Deep Learning Systems. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 82–94

  147. [155]

    Guangba Yu, Gou Tan, Haojia Huang, Zhenyu Zhang, Pengfei Chen, Roberto Natella, and Zibin Zheng. 2024. A Survey on Failure Analysis and Fault Injection in AI Systems. arXiv:2407.00125 [cs.SE] https://arxiv.org/abs/2407.00125

  148. [156]

    Jing Yu, Shukai Duan, and Xiaojun Ye. 2023. A White-Box Testing for Deep Neural Networks Based on Neuron Coverage.IEEE Transactions on Neural Networks and Learning Systems (TNNLS)34, 11 (2023), 9185–9197

  149. [157]

    Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2022. Unveiling Hidden DNN Defects with Decision-Based Metamorphic Testing. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (ASE, 2022). ACM, 113:1–113:13

  150. [158]

    Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2023. Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE, 2023). IEEE, 1200–1212. Proc. ACM Meas. Anal. Com...

  151. [159]

    Yuanyuan Yuan, Qi Pang, and Shuai Wang. 2024. Provably Valid and Diverse Mutations of Real-World Media Data for DNN Testing. IEEE Transactions on Software Engineering (TSE)50, 5 (2024), 1040–1064

  152. [160]

    Yuanyuan Yuan, Shuai Wang, and Zhendong Su. 2024. See the Forest, not Trees: Unveiling and Escaping the Pitfalls of Error-Triggering Inputs in Neural Network Testing. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA, 2024). ...

  153. [161]

    Chenhan Zhang, Shui Yu, Zhiyi Tian, and James J. Q. Yu. 2024. Generative Adversarial Networks: A Survey on Attack and Defense Perspective.ACM Computing Surveys (CSUR)56, 4 (2024), 91:1–91:35

  154. [162]

    Zhang, Mark Harman, Lei Ma, and Yang Liu

    Jie M. Zhang, Mark Harman, Lei Ma, and Yang Liu. 2022. Machine Learning Testing: Survey, Landscapes and Horizons.IEEE Transactions on Software Engineering (TSE)48, 2 (2022), 1–36

  155. [163]

    Pengcheng Zhang, Bin Ren, Hai Dong, and Qiyin Dai. 2022. CAGFuzz: Coverage-Guided Adversarial Generative Fuzzing Testing for Image-Based Deep Learning Systems.IEEE Transactions on Software Engineering (TSE)48, 11 (2022), 4630–4646

  156. [164]

    Peixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong, Xinyu Wang, Xingen Wang, Jin Song Dong, and Ting Dai. 2020. White-box fairness testing through adversarial sampling. InProceedings of the 42nd International Conference on Software Engineering (ICSE, 2020). ACM, 949–960

  157. [165]

    Xiaoyu Zhang, Weipeng Jiang, Chao Shen, Qi Li, Qian Wang, Chenhao Lin, and Xiaohong Guan. 2025. Deep Learning Library Testing: Definition, Methods and Challenges.ACM Computing Surveys (CSUR)57, 7 (2025), 187:1–187:37

  158. [166]

    Zhenya Zhang, Gidon Ernst, Sean Sedwards, Paolo Arcaini, and Ichiro Hasuo. 2018. Two-Layered Falsification of Hybrid Systems Guided by Monte Carlo Tree Search.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)37, 11 (2018), 2894–2905

  159. [167]

    Zhenya Zhang, Deyun Lyu, Paolo Arcaini, Lei Ma, Ichiro Hasuo, and Jianjun Zhao. 2023. FalsifAI: Falsification of AI-Enabled Hybrid Control Systems Guided by Time-Aware Coverage Criteria.IEEE Transactions on Software Engineering (TSE)49, 4 (2023), 1842–1859

  160. [168]

    Zhuangyu Zhang, Zhiyi Zhang, Ziyuan Wang, Fang Chen, and Zhiqiu Huang. 2023. BTM: Black-Box Testing for DNN Based on Meta-Learning. InProceedings of the 23rd IEEE International Conference on Software Quality, Reliability, and Security (QRS, 2023). IEEE, 581–592

  161. [169]

    Chunyu Zhao, Yanzhou Mu, Xiang Chen, Jingke Zhao, Xiaolin Ju, and Gan Wang. 2022. Can test input selection methods for deep neural network guarantee test diversity? A large-scale empirical study.Information and Software Technology (IST)150 (2022), 106982

  162. [170]

    Wei Zheng, Lidan Lin, Xiaoxue Wu, and Xiang Chen. 2024. An Empirical Study on Correlations Between Deep Neural Network Fairness and Neuron Coverage Criteria.IEEE Transactions on Software Engineering (TSE)50, 3 (2024), 391–412

  163. [171]

    Yuhan Zhi, Xiaofei Xie, Chao Shen, Jun Sun, Xiaoyu Zhang, and Xiaohong Guan. 2024. Seed Selection for Testing Deep Neural Networks.ACM Transactions on Software Engineering and Methodology (TOSEM)33, 1 (2024), 23:1–23:33

  164. [172]

    Jianyi Zhou, Feng Li, Jinhao Dong, Hongyu Zhang, and Dan Hao. 2020. Cost-Effective Testing of a Deep Learning Model through Input Reduction. InProceedings of the 31st IEEE International Symposium on Software Reliability Engineering (ISSRE, 2020). IEEE, 289–300

  165. [173]

    Zhiyang Zhou, Wensheng Dou, Jie Liu, Chenxin Zhang, Jun Wei, and Dan Ye. 2021. DeepCon: Contribution Coverage Testing for Deep Learning Systems. InProceedings of the 28th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER, 2021). IEEE, 189–200

  166. [174]

    Zenghui Zhou, Pak-Lok Poon, Tsong Yueh Chen, Kun Qiu, Qinghua Zhao, and Zheng Zheng. 2025. Evaluating the effectiveness of neuron coverage metrics: a metamorphic-testing approach.Software Quality Journal (SQJ)33, 2 (2025), 19

  167. [175]

    Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. 2017. Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. InProceedings of the IEEE International Conference on Computer Vision (ICCV 2017). IEEE Computer Society, 2242–2251

  168. [176]

    Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, and Paolo Tonella. 2021. DeepHyperion: exploring the feature space of deep learning-based systems through illumination search. InProceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis (IS...

  169. [2020]

    InProceedings of the 25th International Conference on Engineering of Complex Computer Systems (ICECCS, 2020)

    An Empirical Study on Correlation between Coverage and Robustness for Deep Neural Networks. InProceedings of the 25th International Conference on Engineering of Complex Computer Systems (ICECCS, 2020). IEEE, 73–82

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.