Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Fault Localization in Deep Learning-based Software: A System-level Approach

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read FL4Deep claims that system-level fault localization for deep learning software—covering data, model construction, and deployment—beats model-centric tools, with best accuracy on data, library-mismatch, and loss-function faults in a…

desk verdict A genuinely novel system-level fault localization idea, but the headline numbers rest on a training/validation overlap the paper never rules out. read the letter →

arxiv 2411.08172 v1 pith:F5D5HEHM submitted 2024-11-12 cs.SE cs.LG

classification cs.SEcs.LG
keywords faultlocalizationdeeplearningsoftwareknowledgegraphsystem-leveldebuggingtestingdatapipelinefaultsdeploymentmismatchstaticanddynamicanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

FL4Deep's central claim is that fault localization for deep learning software should treat the whole development pipeline—data preparation, model construction and training, and deployment—as the search space, not just the neural network model. The paper argues that model-centric tools miss whole fault classes, and that a knowledge graph built from static facts plus dynamic training logs can localize a wider range of faults. On 100 faulty DL scripts, the approach reports the best accuracy for data faults (84%), training/deployment library mismatches (100%), and loss-function faults (69%), and the most balanced precision and recall across five of six fault categories. The authors' intended contribution is the first system-level fault localizer for DL software, with static information carrying most of the performance.

What carries the argument

The load-bearing object is the system-level knowledge graph: a labeled directed graph that encodes facts about the dataset, model structure and hyperparameters, training environment, deployment environment, and per-epoch training logs as RDF triples. Faults are represented as inference rules in Notation3, and a reasoning engine derives fault-related facts and links them to the system parts where the root cause lives. Dynamic training traces are compressed with eight statistical operators and fed to Random Forest, Decision Tree, and KNN classifiers whose majority vote predicts training-phase faults; NodePiece, an anchor-based inductive link predictor, completes missing KG relationships; and the final ranked list orders root causes by how often each fault type appears in prior DL-bug studies.

What would settle it

Check every script in the 100-script validation set for exact or near-duplicate overlap with the 75-script training set, especially the 20 SO posts reused from prior studies and any mutants derived from them; if overlap exists, recompute precision and recall on a strictly disjoint hold-out set and see whether the reported data-fault accuracy stays near 84%.

Watch

Extended reading notes

Core claim

On the authors' own terms, the discovery is that a knowledge graph spanning the entire DL pipeline can serve as the backbone of fault localization: static and dynamic information is extracted from the system, encoded as RDF triples, enriched by rule-based reasoning and inductive link prediction, and converted into a ranked list of root causes. This lets FL4Deep identify faults that originate outside the model—such as a train/deploy library version mismatch—which prior techniques that analyze only model training cannot see. In the reported evaluation, FL4Deep wins three of six fault categories on accuracy, and on precision/recall it is the best-balanced method for data, library mismatch, loss function, insufficient iteration, and activation-function faults, with precision/recall of 1.0/0.84, 1.0/1.0, 0.85/0.69, 0.89/0.62, and 0.89/0.92 respectively.

Load-bearing premise

The 100 validation scripts must be independent of the 75 training scripts; if the same Stack Overflow posts or mutated versions of them appear in both sets, the measured precision and recall for dynamic faults are inflated.

Editorial extensions

If this is right

  • Faults that live outside the model—such as mismatched library versions between training and deployment—are localizable at 100% accuracy, a fault class the four compared approaches miss entirely.
  • A single execution of the buggy script suffices for the method, avoiding the ten-run training strategy that DeepFD uses, while still matching or beating it on most fault categories.
  • Static information is the highest-leverage component: removing it drops performance by 67% with a p-value of 0.04, so future DL fault localizers should invest in pipeline-wide static facts.
  • The knowledge-graph design lets the method rank multiple simultaneous root causes by prior fault frequency, giving developers an ordered debugging checklist rather than a single suspect line.
  • Because the approach targets the whole pipeline, it can be extended to DL frameworks beyond Keras and TensorFlow and to deployment faults that only appear after model export.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the validation scripts are not strictly disjoint from the 75 training scripts, the reported dynamic-fault numbers are an upper bound; the paper does not state disjointness, so a leakage check is the first thing an independent evaluator should run.
  • The ranking step uses global fault frequencies from prior studies; an adaptive ranker that uses knowledge-graph confidence scores could improve recall for rare faults like optimizer issues, where FL4Deep lags DeepFD.
  • The same knowledge-graph-plus-rules recipe should transfer to classical ML pipelines and to data-centric debugging, since the data and deployment rules are not specific to neural architectures.
  • A cross-framework test on PyTorch would reveal how much of the reported accuracy comes from the static extraction rules versus the learned components, since only the former should port without retraining.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes FL4Deep, a knowledge-graph-based fault localization technique for deep learning software. FL4Deep extracts static information about the dataset, model, training environment, and deployment environment, plus dynamic information from training logs, and uses this to build a knowledge graph. Faults are identified by a set of KG rules, by random forest/decision tree/KNN classifiers trained on statistics of dynamic traces, and by a NodePiece link-prediction model that completes missing KG relationships; the final output is a ranked list of root causes. The approach is evaluated on a validation set of 100 buggy DL scripts (with six fault-type counts summing to 106) and compared with DeepFD, DeepLocalize, AutoTrainer, and UMLAUT. The authors report best or near-best performance on data, library mismatch, loss function, insufficient iteration, and activation function faults, and an ablation study showing static information is the most important component.

Significance. If the evaluation is sound, the paper makes a useful contribution to fault localization for DL-based systems. Its strengths include the system-level scope beyond the trained model, the combination of static rules with learned dynamic classifiers and graph link prediction, the use of a real-world bug dataset with a publicly available replication package, and the comparison with four existing tools using their own replication packages. The ablation study and the use of Fisher's exact test are also positive features. The central caveat is that the headline comparison depends on the independence of the 100 validation scripts from the 75 training scripts used to fit the learned components; the manuscript does not currently establish this, and there are also internal numerical inconsistencies in the reported results that must be resolved before the claims can be taken at face value.

major comments (3)
  1. [Sections 4.1, 5.2, 6.1] Section 4.1 states that the 75-script training set includes 17 buggy codes from Cao et al. [23] and 30 from Defect4ML [64], with 12 reported by both sources, while the 100-script validation set was built from 20 SO posts gathered from prior studies [23, 107] plus mutation operators applied to those posts. Because [23] is a common source, and Section 6.1 confirms that validation scripts are "buggy DL scripts previously employed by other studies" and then mutated, the paper never establishes that the validation scripts (or their mutated descendants) are disjoint from the training scripts. The dynamic fault classifiers (RF/DT/KNN, Section 4.2) and the NodePiece link predictor (Section 4.4) are fit on the 75 training samples and then scored on the 100 validation samples; any overlap or near-duplicate mutation between the two sets will inflate the reported accuracy for loss, activation, insufficient iteration, and optimization faults, and will make inductive link prediction artificially easy because validation KGs would be nearly isomorphic to training KGs. This independence is load-bearing for the central claim of outperforming DeepFD, DeepLocalize, AutoTrainer, and UMLAUT. The authors should provide explicit evidence of disjointness (e.g., a full list of training and validation sample IDs and mutation lineages, or a deduplication analysis) or re-run the evaluation on a properly separated hold-out set.
  2. [Section 5.2/5.4, Tables 5 and 7] The six issue-type sample counts in Table 5 and Table 7 sum to 106 (19+20+16+15+10+26), not the 100 scripts stated in the abstract and in Section 5.2. If scripts may contain multiple simultaneous faults, this should be stated explicitly and the per-fault-type evaluation protocol defined. Independently, specific rows are internally inconsistent: for UMLAUT on Data, Table 5 reports 8 detected faults while Table 7 reports FP=27, FN=10 with 19 samples, implying TP=9 and yielding precision 0.25 and recall 0.47, neither matching the printed PR=0.23/RC=0.44; for FL4Deep on Insufficient iteration, Table 5 reports 8 detected faults while Table 7 reports FP=1, FN=5, RC=0.62, which with 15 positive samples implies either TP=8/FN=7 (recall 0.53) or TP=10/FN=5 (recall 0.67). The paper must reconcile these numbers and specify how a predicted fault is counted as "identified" (e.g., top-1, top-k, or any rank in the output list); without that, the headline accuracy and precision/recall comparisons are not verifiable.
  3. [Section 5.3] The sensitivity analysis opens with "Using the same 20 samples we used to compare approaches," but Section 5.2 reports the comparison on 100 scripts. If the ablation was actually run on only 20 scripts, the conclusions of RQ2 (e.g., static information removal causes a 67% drop with p=0.04) do not apply to the 100-script evaluation and should be redone; if it is a typo, it must be corrected. In addition, the Fisher's exact test should report the contingency tables for each ablation comparison and should account for the paired nature of the 100-sample evaluation and the multiple comparisons across the six fault types.
minor comments (5)
  1. [Section 5.4, Finding 4] Finding 4 says UMLAUT's precision on activation function faults is "0.26%" but the correct value is 26% (0.26 as a proportion); please fix.
  2. [Section 6.1] The phrase "unseen samples" in Section 6.1 is ambiguous: clarify whether "unseen" refers to the baselines' prior evaluations or to FL4Deep's training set, and if the latter, explain how disjointness was ensured.
  3. [Table 4] Table 4 is hard to read because each tool's output is a variable-length list (e.g., UMLAUT's warnings and FL4Deep's ranked root causes) with inconsistent formatting; a clearer layout or legend mapping each tool's output to the six fault categories would help readers interpret the comparison.
  4. [Abstract and Section 1] The claim "For the first time" is strong given that UMLAUT and DeepDiagnosis already analyze multiple stages of the DL pipeline; consider softening the novelty claim to "system-level" or "full-pipeline" framing.
  5. [Section 4.5] The ranking of root causes uses priors from Humbatova et al. without a sensitivity analysis; a short discussion of how the ranking would change under alternative priors would strengthen the paper.

Circularity Check

1 steps flagged · score 5.0 of 10

Validation scripts are mutated descendants of the same SO posts used to build the training set, so learned-component predictions are partly fitted-input evaluation; static-rule results are independent.

  1. fitted input called prediction [Section 4.1 (Dataset Preparation); applied in Sections 4.2 and 4.4; evaluation in Section 5.2]
    "The training dataset consists of 75 real-world buggy DL programs, extracted from SO and GitHub repositories. ... Of 75 samples. 17 buggy codes were sourced from the research conducted by Cao et al.[23] ... The validation dataset includes 100 buggy DL scripts ... To create this dataset, we utilized 20 SO posts gathered from prior studies [23, 107]. Additionally, we applied various mutation operators specifically designed for DL software systems (such as modifying the loss function and changing the activation function [26, 47]) to these 20 samples to generate new buggy samples."

    The validation posts are taken from [23] and [107], while the training set already contains 17 scripts from [23] and 12 additional samples reported by both [23] and Defect4ML. The paper never asserts that the validation set is disjoint from training, and the described construction makes overlap likely. Each mutated validation script inherits the seed post, code, and training dynamics of a training sample. The RF/DT/KNN classifiers (Section 4.2) and the NodePiece link predictor (Section 4.4) are fitted on those training samples and then scored on these mutated descendants, so the reported dynamic-fault predictions for loss function, insufficient iteration, optimization, and activation faults partly reproduce near-duplicate training traces rather than independent generalization.

full rationale

The paper's static-rule results (data 84%, library mismatch 100%) are derived from explicit rules and external fault literature, so those claims are not circular. The central concern is the learned components: dynamic classifiers and NodePiece are trained on 75 samples and evaluated on 100 scripts constructed from 20 SO posts from the same cited corpora [23,107] with mutation operators. Because the training set also draws from [23] (17 samples, plus 12 also in Defect4ML), the validation set is not established as independent; mutated validation scripts can be near-duplicates of training scripts, making dynamic predictions and KG link predictions artificially easy. The ablation study shows removing dynamic information (p=0.06) and link prediction (p=0.10) is not statistically significant, while removing static information is (p=0.04), consistent with the learned components being the fragile, partly circular part of the evaluation. No uniqueness theorem or ansatz is smuggled via self-citation; the self-citations to Defect4ML and earlier rule-based tools supply data and fault definitions but are not the load-bearing reduction. Overall, the headline three-out-of-six accuracy claim is only partly affected because the two strongest categories (data, library mismatch) are static-rule based; the precision/recall wins on dynamic fault categories are the part that reduces by construction to training-set overlap.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central claim rests on hand-written rules, fitted classifiers, and an assumption that the validation set is representative and disjoint. No new physical entities are introduced.

free parameters (3)
  • KG fault rules in Table 2
    Hand-crafted rules encoding when a fault is present (e.g., 'missing activation function', 'suboptimal train/test ratio'). The conditions are chosen by the authors from literature and are not fitted to data, but they fully determine the static fault localization output.
  • Ranking priors from Humbatova et al.
    The ranking of candidate faults is based on relative frequencies reported in prior studies; these are fixed constants not measured in this paper.
  • Trained classifier parameters (RF, DT, KNN, NodePiece)
    The dynamic fault predictors are fitted on 75 samples; hyperparameters are not given, so the fitted model parameters are free parameters of the evaluation.
assumptions (3)
  • domain assumption Faults in DL software can be mapped to KG facts and inferred from static/dynamic features.
    Section 4.3: 'Fault-related facts refer to the faults identified by FL4Deep. ... a relationship between a fault-related fact and a basic fact ... indicates the location of the fault's root cause.'
  • domain assumption The fault taxonomy in Table 2 is complete and correct for Keras/TensorFlow systems.
    The rules are gathered from previous studies and Q&A forums (Section 4.3); their correctness is assumed.
  • domain assumption The 75-sample training set has accurate labels.
    Section 4.1: labels are assigned by the first two authors based on accepted answers or PRs; no inter-rater agreement is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fault Localization in Deep Learning-based Software: A System-level Approach." pith.science (2026). https://pith.science/paper/F5D5HEHM

@misc{pith2026241108172,
  author       = {Pith},
  title        = {Pith review of: Fault Localization in Deep Learning-based Software: A System-level Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F5D5HEHM}},
  note         = {Machine review of arXiv:2411.08172}
}
read the original abstract

Over the past decade, Deep Learning (DL) has become an integral part of our daily lives. This surge in DL usage has heightened the need for developing reliable DL software systems. Given that fault localization is a critical task in reliability assessment, researchers have proposed several fault localization techniques for DL-based software, primarily focusing on faults within the DL model. While the DL model is central to DL components, there are other elements that significantly impact the performance of DL components. As a result, fault localization methods that concentrate solely on the DL model overlook a large portion of the system. To address this, we introduce FL4Deep, a system-level fault localization approach considering the entire DL development pipeline to effectively localize faults across the DL-based systems. In an evaluation using 100 faulty DL scripts, FL4Deep outperformed four previous approaches in terms of accuracy for three out of six DL-related faults, including issues related to data (84%), mismatched libraries between training and deployment (100%), and loss function (69%). Additionally, FL4Deep demonstrated superior precision and recall in fault localization for five categories of faults including three mentioned fault types in terms of accuracy, plus insufficient training iteration and activation function.

Figures

Figures reproduced from arXiv: 2411.08172 by the authors.

Figure 1
Figure 1. High-level view of a DL-based tax software system [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Sample KG of a DL-based system software systems requires expertise not only in ML but also in related fields like data manage￾ment [35]. DL frameworks (such as Keras, TensorFlow, and PyTorch), which are designed to simplify the development of DL-based systems, play a significant role in modern ML development [117]. However, creating test cases based on these frameworks can be particularly challenging due to internal… view at source ↗
Figure 3
Figure 3. Link prediction in KG based on the type of inference (a) transductive link prediction, (b) inductive link [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: High-level view of FL4Deep methodology ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article . Publication date: November 2024 [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Pipeline of DL-based system development process divided into three parts (adopted from [10]) [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Process of creating training dataset and training Random Forest (RF), Decision Tree (DT), and K-nearest [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Training a model for KG link prediction ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article . Publication date: November 2024 [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs

    cs.SE 2026-08 reject novelty 6.0 of 10

    An empirical GitHub mining study finds vLLM is the most adopted LLM serving framework, parallel and memory optimizations dominate, and multi-framework use is rare.

  2. Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy

    cs.SE 2026-03 conditional novelty 6.0 of 10

    MCP server faults form five empirical categories—server setting, server/tool configuration, server/host configuration, documentation, and general programming—confirmed by a 41-practitioner survey.

Reference graph

Works this paper leans on

124 extracted references · 51 canonical work pages · cited by 2 Pith papers

  1. [23]

    Jialun Cao, Meiziniu Li, Xiao Chen, Ming Wen, Yongqiang Tian, Bo Wu, and Shing-Chi Cheung. 2022. DeepFD: Automated Fault Diagnosis and Localization for Deep Learning Programs. arXiv preprint arXiv:2205.01938 (2022)

  2. [64]

    Mohammad Mehdi Morovati, Amin Nikanjam, Foutse Khomh, and Zhen Ming Jiang. 2023. Bugs in machine learning- based systems: a faultload benchmark. Empirical Software Engineering 28, 3 (2023), 62

  3. [1]

    [n. d.]. RDFLib: a Python library for working with RDF. https://github.com/RDFLib/rdflib. Accessed: 2024-08

  4. [2]

    Cwm: a general-purpose data processor for the semantic web

    2002. Cwm: a general-purpose data processor for the semantic web. http://www.w3.org/2000/10/swap/doc/cwm.html

  5. [3]

    Apache Jena: a Java framework for writing Semantic Web applications

    2009. Apache Jena: a Java framework for writing Semantic Web applications. https://jena.apache.org/

  6. [4]

    RDF 1.1 Primer

    2014. RDF 1.1 Primer. https://www.w3.org/TR/rdf11-primer/

  7. [5]

    ISO/IEC/IEEE International Standard - Systems and software engineering–Vocabulary

    2017. ISO/IEC/IEEE International Standard - Systems and software engineering–Vocabulary. ISO/IEC/IEEE 24765:2017(E) (2017), 1–541. https://doi.org/10.1109/IEEESTD.2017.8016712

  8. [6]

    Notation3 Language

    2024. Notation3 Language. https://w3c.github.io/N3/spec/

Show all 124 references
  1. [7]

    ISO/IEC 29881: 2010. 2010. Information technology, Systems and software engineering, FiSMA 1.1 functional size measurement method

  2. [8]

    Mehdi Ali, Max Berrendorf, Charles Tapley Hoyt, Laurent Vermue, Sahand Sharifzadeh, Volker Tresp, and Jens Lehmann. 2021. PyKEEN 1.0: A Python Library for Training and Evaluating Knowledge Graph Embeddings. Journal of Machine Learning Research 22, 82 (2021), 1–6. http://jmlr.o...

  3. [9]

    Mohammad Amin Alipour. 2012. Automated fault localization techniques: a survey. Oregon State University 54, 3 (2012)

  4. [10]

    Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software engineering for machine learning: A case study. In 2019 IEEE/ACM 41st International Conference on Software Engineerin...

  5. [11]

    Paul Ammann and Jeff Offutt. 2016. Introduction to software testing . Cambridge University Press

  6. [12]

    Andrea Arcuri and Lionel Briand. 2014. A hitchhiker’s guide to statistical tests for assessing randomized algorithms in software engineering. Software Testing, Verification and Reliability 24, 3 (2014), 219–250

  7. [13]

    Aitor Arrieta, Sergio Segura, Urtzi Markiegi, Goiuria Sagardui, and Leire Etxeberria. 2018. Spectrum-based fault localization in software product lines. Information and Software Technology 100 (2018), 18–31. https://doi.org/10. 1016/j.infsof.2018.03.008

  8. [14]

    Jinheon Baek, Dong Bok Lee, and Sung Ju Hwang. 2020. Learning to extrapolate knowledge: Transductive few-shot out-of-graph link prediction. Advances in neural information processing systems 33 (2020), 546–560

  9. [15]

    Barrasa, J

    J. Barrasa, J. Webber, and J. Webber. 2023. Building Knowledge Graphs: A Practitioner’s Guide . O’Reilly. https: //books.google.ca/books?id=Ztb5zgEACAAJ

  10. [16]

    Jason Bell. 2020. Machine learning: hands-on for developers and technical professionals . John Wiley & Sons

  11. [17]

    Houssem Ben Braiek and Foutse Khomh. 2023. Testing feedforward neural networks training programs. ACM Transactions on Software Engineering and Methodology 32, 4 (2023), 1–61

  12. [18]

    Tim Berners-Lee and Dan Connolly. 2008. Notation3 (N3): A readable RDF syntax. https://www.w3.org/ TeamSubmission/2008/SUBM-n3-20080114/

  13. [19]

    Sweta Bhattacharya, Praveen Kumar Reddy Maddikunta, Quoc-Viet Pham, Thippa Reddy Gadekallu, Chiranji Lal Chowdhary, Mamoun Alazab, Md Jalil Piran, et al. 2021. Deep learning and medical image processing for coronavirus (COVID-19) pandemic: A survey. Sustainable cities and soci...

  14. [20]

    Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data . 1247–1250

  15. [21]

    Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in neural information processing systems 26 (2013)

  16. [22]

    Léon Bottou. 2014. From machine learning to machine reasoning: An essay. Machine learning 94 (2014), 133–149

  17. [24]

    Bahzad Charbuty and Adnan Abdulazeez. 2021. Classification based on decision tree algorithm for machine learning. Journal of Applied Science and Technology Trends 2, 01 (2021), 20–28

  18. [25]

    Xiaojun Chen, Shengbin Jia, and Yang Xiang. 2020. A review: Knowledge reasoning over knowledge graph. Expert Systems with Applications 141 (2020), 112948

  19. [26]

    Zhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang, Tao Xie, and Xuanzhe Liu. 2020. A Comprehensive Study on Challenges in Deploying Deep Learning Based Software. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Fo...

  20. [27]

    Patrick Cousot and Radhia Cousot. 1977. Abstract interpretation: a unified lattice model for static analysis of programs by construction or approximation of fixpoints. InProceedings of the 4th ACM SIGACT-SIGPLAN symposium on Principles of programming languages. 238–252

  21. [28]

    Ben De Meester, Dörthe Arndt, Pieter Bonte, Jabran Bhatti, Wim Dereuddre, Ruben Verborgh, Femke Ongenae, Filip De Turck, Erik Mannens, and Rik Van de Walle. 2015. Event-driven rule-based reasoning using EYE. In ISWC2015. CEUR

  22. [29]

    Nan Duan, Duyu Tang, and Ming Zhou. 2020. Machine reasoning: Technology, dilemma and future. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts . 1–6

  23. [30]

    Michael D Ernst. 2003. Static and dynamic analysis: Synergy and duality. In WODA 2003: ICSE Workshop on Dynamic Analysis. 24–27

  24. [31]

    Fensel, U

    D. Fensel, U. Şimşek, K. Angele, E. Huaman, E. Kärle, O. Panasiuk, I. Toma, J. Umbrich, and A. Wahler. 2020.Knowledge Graphs: Methodology, Tools and Selected Use Cases . Springer International Publishing. https://books.google.ca/books? id=1qnNDwAAQBAJ

  25. [32]

    Daniel Galin. 2004. Software quality assurance: from theory to implementation . Pearson education

  26. [33]

    Mikhail Galkin, Max Berrendorf, and Charles Tapley Hoyt. 2022. An open challenge for inductive link prediction on knowledge graphs. arXiv preprint arXiv:2203.01520 (2022)

  27. [34]

    Mikhail Galkin, Etienne Denis, Jiapeng Wu, and William L Hamilton. 2021. Nodepiece: Compositional and parameter- efficient representations of large knowledge graphs. arXiv preprint arXiv:2106.12144 (2021)

  28. [35]

    Jerry Gao, Chuanqi Tao, Dou Jie, and Shengqiang Lu. 2019. What is AI software testing? and why. In 2019 IEEE International Conference on Service-Oriented System Engineering (SOSE) . IEEE, 27–2709

  29. [36]

    Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 249–256

  30. [37]

    Liang Gong, Hongyu Zhang, Hyunmin Seo, and Sunghun Kim. 2014. Locating crashing faults based on crash stack traces. arXiv preprint arXiv:1404.4100 (2014)

  31. [38]

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. MIT press

  32. [39]

    Cyril Goutte and Eric Gaussier. 2005. A probabilistic interpretation of precision, recall and F-score, with implication for evaluation. In European conference on information retrieval . Springer, 345–359

  33. [40]

    Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing . Ieee, IEEE, Vancouver, BC, Canada, 6645–6649

  34. [41]

    Sakshi Gupta. 2021. What Is the Best Language for Machine Learning? https://www.springboard.com/blog/data- science/best-language-for-machine-learning. Accessed: 2021-10-06

  35. [42]

    Allan Hackshaw and Amy Kirkwood. 2011. Interpreting and reporting clinical trials with results of borderline significance. Bmj 343 (2011)

  36. [43]

    Mark Harman and Robert Hierons. 2001. An overview of program slicing. software focus 2, 3 (2001), 85–92

  37. [44]

    Stefanus A Haryono, Ferdian Thung, David Lo, Julia Lawall, and Lingxiao Jiang. 2021. Characterization and automatic updates of deprecated machine-learning api usages. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, IEEE, Luxembourg, 137–147

  38. [45]

    Bruno Miranda Henrique, Vinicius Amorim Sobreiro, and Herbert Kimura. 2019. Literature review: Machine learning techniques applied to financial market prediction. Expert Systems with Applications 124 (2019), 226–251

  39. [46]

    Nargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio, Andrea Stocco, and Paolo Tonella. 2020. Taxonomy of real faults in deep learning systems. In Proceedings of the ACM/IEEE 42nd international conference on software engineering. Association for Computing Mach...

  40. [47]

    Nargiz Humbatova, Gunel Jahangirova, and Paolo Tonella. 2021. Deepcrime: mutation testing of deep learning systems based on real faults. In Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis. 67–78

  41. [48]

    IEEE. 2010. ISO/IEC/IEEE International Standard - Systems and software engineering – Vocabulary . IEEE, 3 Park Avenue, New York NY 10016-5997, USA. 1–418 pages. https://doi.org/10.1109/IEEESTD.2010.5733835

  42. [49]

    Md Johirul Islam, Giang Nguyen, Rangeet Pan, and Hridesh Rajan. 2019. A comprehensive study on deep learning bug characteristics. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineer...

  43. [50]

    Md Johirul Islam, Rangeet Pan, Giang Nguyen, and Hridesh Rajan. 2020. Repairing deep neural networks: Fix patterns and challenges. In 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE) . IEEE, Association for Computing Machinery, New York, NY, USA, 1135...

  44. [51]

    ISO/PAS. 2019. ISO/PAS 21448:2019 Road vehicles — Safety of the intended functionality. ISO/PAS 21448:2019 (2019)

  45. [52]

    V Roshan Joseph. 2022. Optimal ratio for data splitting. Statistical Analysis and Data Mining: The ASA Data Science Journal 15, 4 (2022), 531–538

  46. [53]

    Andrej Karpathy. 2019. A Recipe for Training Neural Networks . https://karpathy.github.io/2019/04/25/recipe/

  47. [54]

    Leo Katz. 1953. A new status index derived from sociometric analysis. Psychometrika 18, 1 (1953), 39–43

  48. [55]

    Diksha Khurana, Aditya Koli, Kiran Khatter, and Sukhdev Singh. 2023. Natural language processing: State of the art, current trends and challenges. Multimedia tools and applications 82, 3 (2023), 3713–3744

  49. [56]

    John P Klein and Melvin L Moeschberger. 2006. Survival analysis: techniques for censored and truncated data . Springer Science & Business Media

  50. [57]

    Jitender Kumar Chhabra and Varun Gupta. 2010. A survey of dynamic software metrics. Journal of computer science and technology 25 (2010), 1016–1029

  51. [58]

    Dong Kyu Lee, Junyong In, and Sangseok Lee. 2015. Standard deviation and standard error of the mean. Korean journal of anesthesiology 68, 3 (2015), 220–223

  52. [59]

    Michael R Lyu. 2007. Software reliability engineering: A roadmap. In Future of Software Engineering (FOSE’07) . IEEE, 153–170

  53. [60]

    Christopher Manning, Prabhakar Raghavan, and Hinrich Schütze. 2010. Introduction to information retrieval.Natural Language Engineering 16, 1 (2010), 100–103

  54. [61]

    John McCarthy. 2007. What is artificial intelligence? (2007)

  55. [62]

    Richard Meyes, Melanie Lu, Constantin Waubert de Puiseau, and Tobias Meisen. 2019. Ablation studies in artificial neural networks. arXiv preprint arXiv:1901.08644 (2019)

  56. [63]

    Mohammad Mehdi Morovati, Amin Nikanjam, and Foutse Khomh. 2024. Paper replication package. https://github. com/mohmehmo/fl4deep. Accessed: 2024-07

  57. [65]

    Mohammad Mehdi Morovati, Amin Nikanjam, Florian Tambon, Foutse Khomh, and Zhen Ming Jiang. 2024. Bug characterization in machine learning-based systems. Empirical Software Engineering 29, 1 (2024), 14

  58. [66]

    Glenford J Myers, Corey Sandler, and Tom Badgett. 2011. The art of software testing . John Wiley & Sons

  59. [67]

    Lee Naish, Hua Jie Lee, and Kotagiri Ramamohanarao. 2011. A Model for Spectra-Based Software Diagnosis. ACM Trans. Softw. Eng. Methodol. 20, 3, Article 11 (aug 2011), 32 pages. https://doi.org/10.1145/2000791.2000795

  60. [68]

    Maximilian Nickel, Kevin Murphy, Volker Tresp, and Evgeniy Gabrilovich. 2015. A review of relational machine learning for knowledge graphs. Proc. IEEE 104, 1 (2015), 11–33

  61. [69]

    Amin Nikanjam, Houssem Ben Braiek, Mohammad Mehdi Morovati, and Foutse Khomh. 2021. Automatic Fault Detection for Deep Learning Programs Using Graph Transformations. ACM Trans. Softw. Eng. Methodol. 31, 1, Article 14 (sep 2021), 27 pages. https://doi.org/10.1145/3470006

  62. [70]

    Amin Nikanjam, Mohammad Mehdi Morovati, Foutse Khomh, and Houssem Ben Braiek. 2022. Faults in deep reinforcement learning programs: a taxonomy and a detection approach. Automated Software Engineering 29, 1 (2022), 1–32

  63. [71]

    Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall. 2018. Activation functions: Com- parison of trends in practice and research for deep learning. arXiv preprint arXiv:1811.03378 (2018)

  64. [72]

    Lawrence Page, Sergey Brin, Rajeev Motwani, Terry Winograd, et al. 1999. The pagerank citation ranking: Bringing order to the web. (1999), 1–17

  65. [73]

    Jeff Z Pan. 2009. Resource description framework. InHandbook on ontologies. Springer, 71–90. https://jena.apache.org/

  66. [74]

    Annibale Panichella and Cynthia CS Liem. 2021. What are we really testing in mutation testing for machine learning? a critical reflection. In 2021 IEEE/ACM 43rd International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER). IEEE, 66–70

  67. [75]

    Mike Papadakis and Yves Le Traon. 2015. Metallaxis-FL: Mutation-Based Fault Localization. Softw. Test. Verif. Reliab. 25, 5–7 (aug 2015), 605–628. https://doi.org/10.1002/stvr.1509

  68. [76]

    Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013. On the difficulty of training Recurrent Neural Networks. arXiv:1211.5063 [cs.LG] https://arxiv.org/abs/1211.5063

  69. [77]

    Ernst, Deric Pang, and Benjamin Keller

    Spencer Pearson, José Campos, René Just, Gordon Fraser, Rui Abreu, Michael D. Ernst, Deric Pang, and Benjamin Keller. 2017. Evaluating and Improving Fault Localization. In2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE). 609–620. https://doi.org/10.11...

  70. [78]

    Goran Petrović, Marko Ivanković, Gordon Fraser, and René Just. 2021. Does mutation testing improve testing practices?. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 910–921

  71. [79]

    Danijel Radjenović, Marjan Heričko, Richard Torkar, and Aleš Živkovič. 2013. Software fault prediction metrics: A systematic literature review. Information and Software Technology 55, 8 (2013), 1397–1418. https://doi.org/10.1016/j. infsof.2013.02.009 ACM Trans. Softw. Eng. Met...

  72. [80]

    Foyzur Rahman, Daryl Posnett, Abram Hindle, Earl Barr, and Premkumar Devanbu. 2011. BugCache for inspections: hit or miss?. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering. 322–331

  73. [81]

    Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella. 2020. Testing machine learning based systems: a systematic mapping.Empirical Software Engineering 25, 6 (2020), 5193–5254

  74. [82]

    Steven J Rigatti. 2017. Random forest. Journal of Insurance Medicine 47, 1 (2017), 31–39

  75. [83]

    Emilio Rivera-Landos, Foutse Khomh, and Amin Nikanjam. 2021. The challenge of reproducible ML: an empirical study on the impact of bugs

  76. [84]

    Andrea Rossi, Denilson Barbosa, Donatella Firmani, Antonio Matinata, and Paolo Merialdo. 2021. Knowledge graph embedding for link prediction: A comparative analysis. ACM Transactions on Knowledge Discovery from Data (TKDD) 15, 2 (2021), 1–49

  77. [85]

    Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. 2018. Modeling relational data with graph convolutional networks. InThe semantic web: 15th international conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, procee...

  78. [86]

    Eldon Schoop, Forrest Huang, and Bjoern Hartmann. 2021. Umlaut: Debugging deep learning programs using program structure and model behavior. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–16

  79. [87]

    David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. 2015. Hidden technical debt in machine learning systems. Advances in neural information processing systems 28 (2015)

  80. [88]

    Diogo Seca. 2021. A review on oracle issues in machine learning. arXiv preprint arXiv:2105.01407 (2021)

  81. [89]

    Mehil B Shah, Mohammad Masudur Rahman, and Foutse Khomh. 2024. Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study. arXiv preprint arXiv:2401.03069 (2024)

  82. [90]

    Jonathan Richard Shewchuk. 2022. Concise Machine Learning

  83. [91]

    Nasim Shirvani-Mahdavi, Farahnaz Akrami, Mohammed Samiul Saeef, Xiao Shi, and Chengkai Li. 2023. Comprehen- sive analysis of freebase and dataset creation for robust evaluation of knowledge graph link prediction models. In International Semantic Web Conference. Springer, 113–133

  84. [92]

    Amit Singhal. 2012. Introducing the Knowledge Graph: things, not strings. https://blog.google/products/search/ introducing-knowledge-graph-things-not/

  85. [93]

    Ezekiel Soremekun, Lukas Kirschner, Marcel Böhme, and Andreas Zeller. 2021. Locating faults with program slicing: an empirical analysis. Empirical Software Engineering 26 (2021), 1–45

  86. [94]

    Kishore Sugali. 2021. Software testing: Issues and challenges of artificial intelligence & machine learning.International Journal of Artificial Intelligence & Applications 12 (2021)

  87. [95]

    Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision . 843–852

  88. [96]

    Florian Tambon, Amin Nikanjam, Le An, Foutse Khomh, and Giuliano Antoniol. 2021. Silent Bugs in Deep Learning Frameworks: An Empirical Study of Keras and TensorFlow. arXiv preprint arXiv:2112.13314 (2021)

  89. [97]

    Jimin Tan, Jianan Yang, Sai Wu, Gang Chen, and Jake Zhao. 2021. A critical look at the current train/test split in machine learning. arXiv preprint arXiv:2106.04525 (2021)

  90. [98]

    Muhammad Usman, Youcheng Sun, Divya Gopinath, Rishi Dange, Luca Manolache, and Corina S Păsăreanu. 2023. An overview of structural coverage metrics for testing neural networks. International Journal on Software Tools for Technology Transfer 25, 3 (2023), 393–405

  91. [99]

    Shikhar Vashishth, Soumya Sanyal, Vikram Nitin, and Partha Talukdar. 2019. Composition-based multi-relational graph convolutional networks. arXiv preprint arXiv:1911.03082 (2019)

  92. [100]

    Kavuri, and Kewen Yin

    Venkat Venkatasubramanian, Raghunathan Rengaswamy, Surya N. Kavuri, and Kewen Yin. 2003. A review of process fault detection and diagnosis: Part III: Process history based methods. Computers & Chemical Engineering 27, 3 (2003), 327–346. https://doi.org/10.1016/S0098-1354(02)00162-X

  93. [101]

    Celine Vens, Jan Struyf, Leander Schietgat, Sašo Džeroski, and Hendrik Blockeel. 2008. Decision trees for hierarchical multi-label classification. Machine learning 73 (2008), 185–214

  94. [102]

    Ruben Verborgh and Jos De Roo. 2015. Drawing conclusions from linked data on the web: The EYE reasoner. IEEE Software 32, 3 (2015), 23–27

  95. [103]

    Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J

    Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, E...

  96. [104]

    Christina Voskoglou. 2017. What is the best programming language for Machine Learning . Towards Data Science. https://towardsdatascience.com/what-is-the-best-programming-language-for-machine-learning-a745c156d6b7

  97. [105]

    Mohammad Wardat, Breno Dantas Cruz, Wei Le, and Hridesh Rajan. 2022. Deepdiagnosis: Automatically diagnosing faults and recommending actionable fixes in deep learning programs. InProceedings of the 44th International Conference on Software Engineering. 561–572

  98. [106]

    Mohammad Wardat, Breno Dantas Cruz, Wei Le, and Hridesh Rajan. 2023. An Effective Data-Driven Approach for Localizing Deep Learning Faults. arXiv preprint arXiv:2307.08947 (2023)

  99. [107]

    Mohammad Wardat, Wei Le, and Hridesh Rajan. 2021. Deeplocalize: Fault localization for deep neural networks. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 251–262

  100. [108]

    Robert West, Evgeniy Gabrilovich, Kevin Murphy, Shaohua Sun, Rahul Gupta, and Dekang Lin. 2014. Knowledge base completion via search-based question answering. In Proceedings of the 23rd international conference on World wide web. 515–526

  101. [109]

    W Eric Wong, Ruizhi Gao, Yihao Li, Rui Abreu, and Franz Wotawa. 2016. A survey on software fault localization. IEEE Transactions on Software Engineering 42, 8 (2016), 707–740

  102. [110]

    W Eric Wong and TH Tse. 2023. Handbook of software fault localization: foundations and advances . John Wiley & Sons

  103. [111]

    Qingyao Wu, Mingkui Tan, Hengjie Song, Jian Chen, and Michael K Ng. 2016. ML-FOREST: A multi-label tree ensemble method for multi-label classification. IEEE transactions on knowledge and data engineering 28, 10 (2016), 2665–2680

  104. [112]

    Baowen Xu, Ju Qian, Xiaofang Zhang, Zhongqiang Wu, and Lin Chen. 2005. A brief survey of program slicing. ACM SIGSOFT Software Engineering Notes 30, 2 (2005), 1–36

  105. [113]

    Orhan G. Yalçın. 2021. Top 5 Deep Learning Frameworks to Watch in 2021 and Why Tensor- Flow. https://towardsdatascience.com/top-5-deep-learning-frameworks-to-watch-in-2021-and-why-tensorflow- 98d8d6667351 Accessed: 2022-12-29

  106. [114]

    Yilin Yang, Tianxing He, Zhilong Xia, and Yang Feng. 2022. A comprehensive empirical study on bug characteristics of deep learning frameworks. Information and Software Technology 151 (2022), 107004

  107. [115]

    Xiao Yu, Kwabena Ebo Bennin, Jin Liu, Jacky Wai Keung, Xiaofei Yin, and Zhou Xu. 2019. An empirical study of learning to rank techniques for effort-aware defect prediction. In 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER) . I...

  108. [116]

    Ekim Yurtsever, Jacob Lambert, Alexander Carballo, and Kazuya Takeda. 2020. A survey of autonomous driving: Common practices and emerging technologies. IEEE access 8 (2020), 58443–58469

  109. [117]

    Jie M Zhang, Mark Harman, Lei Ma, and Yang Liu. 2020. Machine learning testing: Survey, landscapes and horizons. IEEE Transactions on Software Engineering (2020)

  110. [118]

    Min-Ling Zhang and Zhi-Hua Zhou. 2005. A k-nearest neighbor based algorithm for multi-label classification. In 2005 IEEE international conference on granular computing , Vol. 2. IEEE, 718–721

  111. [119]

    Shichao Zhang. 2021. Challenges in KNN classification. IEEE Transactions on Knowledge and Data Engineering 34, 10 (2021), 4663–4675

  112. [120]

    Xiangyu Zhang, Neelam Gupta, and Rajiv Gupta. 2006. Locating faults through automated predicate switching. In Proceedings of the 28th international conference on Software engineering . 272–281

  113. [121]

    Xiangyu Zhang, Neelam Gupta, and Rajiv Gupta. 2007. A study of effectiveness of dynamic slicing in locating real faults. Empirical Software Engineering 12, 2 (2007), 143–160

  114. [122]

    Xiaoyu Zhang, Juan Zhai, Shiqing Ma, and Chao Shen. 2021. AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair System. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 359–371

  115. [123]

    Yuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong, Shing-Chi Cheung, and Tao Xie. 2020. Detecting numerical bugs in neural network architectures. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Softw...

  116. [124]

    Daming Zou, Jingjing Liang, Yingfei Xiong, Michael D Ernst, and Lu Zhang. 2019. An empirical study of fault localization families and their combinations. IEEE Transactions on Software Engineering 47, 2 (2019), 332–347. ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article ....

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.