REVIEW 3 major objections 5 minor 2 cited by
Fault Localization in Deep Learning-based Software: A System-level Approach
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read FL4Deep claims that system-level fault localization for deep learning software—covering data, model construction, and deployment—beats model-centric tools, with best accuracy on data, library-mismatch, and loss-function faults in a…
desk verdict A genuinely novel system-level fault localization idea, but the headline numbers rest on a training/validation overlap the paper never rules out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the system-level knowledge graph: a labeled directed graph that encodes facts about the dataset, model structure and hyperparameters, training environment, deployment environment, and per-epoch training logs as RDF triples. Faults are represented as inference rules in Notation3, and a reasoning engine derives fault-related facts and links them to the system parts where the root cause lives. Dynamic training traces are compressed with eight statistical operators and fed to Random Forest, Decision Tree, and KNN classifiers whose majority vote predicts training-phase faults; NodePiece, an anchor-based inductive link predictor, completes missing KG relationships; and the final ranked list orders root causes by how often each fault type appears in prior DL-bug studies.
What would settle it
Check every script in the 100-script validation set for exact or near-duplicate overlap with the 75-script training set, especially the 20 SO posts reused from prior studies and any mutants derived from them; if overlap exists, recompute precision and recall on a strictly disjoint hold-out set and see whether the reported data-fault accuracy stays near 84%.
Extended reading notes
Core claim
On the authors' own terms, the discovery is that a knowledge graph spanning the entire DL pipeline can serve as the backbone of fault localization: static and dynamic information is extracted from the system, encoded as RDF triples, enriched by rule-based reasoning and inductive link prediction, and converted into a ranked list of root causes. This lets FL4Deep identify faults that originate outside the model—such as a train/deploy library version mismatch—which prior techniques that analyze only model training cannot see. In the reported evaluation, FL4Deep wins three of six fault categories on accuracy, and on precision/recall it is the best-balanced method for data, library mismatch, loss function, insufficient iteration, and activation-function faults, with precision/recall of 1.0/0.84, 1.0/1.0, 0.85/0.69, 0.89/0.62, and 0.89/0.92 respectively.
Load-bearing premise
The 100 validation scripts must be independent of the 75 training scripts; if the same Stack Overflow posts or mutated versions of them appear in both sets, the measured precision and recall for dynamic faults are inflated.
Editorial extensions
If this is right
- Faults that live outside the model—such as mismatched library versions between training and deployment—are localizable at 100% accuracy, a fault class the four compared approaches miss entirely.
- A single execution of the buggy script suffices for the method, avoiding the ten-run training strategy that DeepFD uses, while still matching or beating it on most fault categories.
- Static information is the highest-leverage component: removing it drops performance by 67% with a p-value of 0.04, so future DL fault localizers should invest in pipeline-wide static facts.
- The knowledge-graph design lets the method rank multiple simultaneous root causes by prior fault frequency, giving developers an ordered debugging checklist rather than a single suspect line.
- Because the approach targets the whole pipeline, it can be extended to DL frameworks beyond Keras and TensorFlow and to deployment faults that only appear after model export.
Reading between the lines
- If the validation scripts are not strictly disjoint from the 75 training scripts, the reported dynamic-fault numbers are an upper bound; the paper does not state disjointness, so a leakage check is the first thing an independent evaluator should run.
- The ranking step uses global fault frequencies from prior studies; an adaptive ranker that uses knowledge-graph confidence scores could improve recall for rare faults like optimizer issues, where FL4Deep lags DeepFD.
- The same knowledge-graph-plus-rules recipe should transfer to classical ML pipelines and to data-centric debugging, since the data and deployment rules are not specific to neural architectures.
- A cross-framework test on PyTorch would reveal how much of the reported accuracy comes from the static extraction rules versus the learned components, since only the former should port without retraining.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FL4Deep, a knowledge-graph-based fault localization technique for deep learning software. FL4Deep extracts static information about the dataset, model, training environment, and deployment environment, plus dynamic information from training logs, and uses this to build a knowledge graph. Faults are identified by a set of KG rules, by random forest/decision tree/KNN classifiers trained on statistics of dynamic traces, and by a NodePiece link-prediction model that completes missing KG relationships; the final output is a ranked list of root causes. The approach is evaluated on a validation set of 100 buggy DL scripts (with six fault-type counts summing to 106) and compared with DeepFD, DeepLocalize, AutoTrainer, and UMLAUT. The authors report best or near-best performance on data, library mismatch, loss function, insufficient iteration, and activation function faults, and an ablation study showing static information is the most important component.
Significance. If the evaluation is sound, the paper makes a useful contribution to fault localization for DL-based systems. Its strengths include the system-level scope beyond the trained model, the combination of static rules with learned dynamic classifiers and graph link prediction, the use of a real-world bug dataset with a publicly available replication package, and the comparison with four existing tools using their own replication packages. The ablation study and the use of Fisher's exact test are also positive features. The central caveat is that the headline comparison depends on the independence of the 100 validation scripts from the 75 training scripts used to fit the learned components; the manuscript does not currently establish this, and there are also internal numerical inconsistencies in the reported results that must be resolved before the claims can be taken at face value.
major comments (3)
- [Sections 4.1, 5.2, 6.1] Section 4.1 states that the 75-script training set includes 17 buggy codes from Cao et al. [23] and 30 from Defect4ML [64], with 12 reported by both sources, while the 100-script validation set was built from 20 SO posts gathered from prior studies [23, 107] plus mutation operators applied to those posts. Because [23] is a common source, and Section 6.1 confirms that validation scripts are "buggy DL scripts previously employed by other studies" and then mutated, the paper never establishes that the validation scripts (or their mutated descendants) are disjoint from the training scripts. The dynamic fault classifiers (RF/DT/KNN, Section 4.2) and the NodePiece link predictor (Section 4.4) are fit on the 75 training samples and then scored on the 100 validation samples; any overlap or near-duplicate mutation between the two sets will inflate the reported accuracy for loss, activation, insufficient iteration, and optimization faults, and will make inductive link prediction artificially easy because validation KGs would be nearly isomorphic to training KGs. This independence is load-bearing for the central claim of outperforming DeepFD, DeepLocalize, AutoTrainer, and UMLAUT. The authors should provide explicit evidence of disjointness (e.g., a full list of training and validation sample IDs and mutation lineages, or a deduplication analysis) or re-run the evaluation on a properly separated hold-out set.
- [Section 5.2/5.4, Tables 5 and 7] The six issue-type sample counts in Table 5 and Table 7 sum to 106 (19+20+16+15+10+26), not the 100 scripts stated in the abstract and in Section 5.2. If scripts may contain multiple simultaneous faults, this should be stated explicitly and the per-fault-type evaluation protocol defined. Independently, specific rows are internally inconsistent: for UMLAUT on Data, Table 5 reports 8 detected faults while Table 7 reports FP=27, FN=10 with 19 samples, implying TP=9 and yielding precision 0.25 and recall 0.47, neither matching the printed PR=0.23/RC=0.44; for FL4Deep on Insufficient iteration, Table 5 reports 8 detected faults while Table 7 reports FP=1, FN=5, RC=0.62, which with 15 positive samples implies either TP=8/FN=7 (recall 0.53) or TP=10/FN=5 (recall 0.67). The paper must reconcile these numbers and specify how a predicted fault is counted as "identified" (e.g., top-1, top-k, or any rank in the output list); without that, the headline accuracy and precision/recall comparisons are not verifiable.
- [Section 5.3] The sensitivity analysis opens with "Using the same 20 samples we used to compare approaches," but Section 5.2 reports the comparison on 100 scripts. If the ablation was actually run on only 20 scripts, the conclusions of RQ2 (e.g., static information removal causes a 67% drop with p=0.04) do not apply to the 100-script evaluation and should be redone; if it is a typo, it must be corrected. In addition, the Fisher's exact test should report the contingency tables for each ablation comparison and should account for the paired nature of the 100-sample evaluation and the multiple comparisons across the six fault types.
minor comments (5)
- [Section 5.4, Finding 4] Finding 4 says UMLAUT's precision on activation function faults is "0.26%" but the correct value is 26% (0.26 as a proportion); please fix.
- [Section 6.1] The phrase "unseen samples" in Section 6.1 is ambiguous: clarify whether "unseen" refers to the baselines' prior evaluations or to FL4Deep's training set, and if the latter, explain how disjointness was ensured.
- [Table 4] Table 4 is hard to read because each tool's output is a variable-length list (e.g., UMLAUT's warnings and FL4Deep's ranked root causes) with inconsistent formatting; a clearer layout or legend mapping each tool's output to the six fault categories would help readers interpret the comparison.
- [Abstract and Section 1] The claim "For the first time" is strong given that UMLAUT and DeepDiagnosis already analyze multiple stages of the DL pipeline; consider softening the novelty claim to "system-level" or "full-pipeline" framing.
- [Section 4.5] The ranking of root causes uses priors from Humbatova et al. without a sensitivity analysis; a short discussion of how the ranking would change under alternative priors would strengthen the paper.
Circularity Check
Validation scripts are mutated descendants of the same SO posts used to build the training set, so learned-component predictions are partly fitted-input evaluation; static-rule results are independent.
-
fitted input called prediction
[Section 4.1 (Dataset Preparation); applied in Sections 4.2 and 4.4; evaluation in Section 5.2]
"The training dataset consists of 75 real-world buggy DL programs, extracted from SO and GitHub repositories. ... Of 75 samples. 17 buggy codes were sourced from the research conducted by Cao et al.[23] ... The validation dataset includes 100 buggy DL scripts ... To create this dataset, we utilized 20 SO posts gathered from prior studies [23, 107]. Additionally, we applied various mutation operators specifically designed for DL software systems (such as modifying the loss function and changing the activation function [26, 47]) to these 20 samples to generate new buggy samples."
The validation posts are taken from [23] and [107], while the training set already contains 17 scripts from [23] and 12 additional samples reported by both [23] and Defect4ML. The paper never asserts that the validation set is disjoint from training, and the described construction makes overlap likely. Each mutated validation script inherits the seed post, code, and training dynamics of a training sample. The RF/DT/KNN classifiers (Section 4.2) and the NodePiece link predictor (Section 4.4) are fitted on those training samples and then scored on these mutated descendants, so the reported dynamic-fault predictions for loss function, insufficient iteration, optimization, and activation faults partly reproduce near-duplicate training traces rather than independent generalization.
full rationale
The paper's static-rule results (data 84%, library mismatch 100%) are derived from explicit rules and external fault literature, so those claims are not circular. The central concern is the learned components: dynamic classifiers and NodePiece are trained on 75 samples and evaluated on 100 scripts constructed from 20 SO posts from the same cited corpora [23,107] with mutation operators. Because the training set also draws from [23] (17 samples, plus 12 also in Defect4ML), the validation set is not established as independent; mutated validation scripts can be near-duplicates of training scripts, making dynamic predictions and KG link predictions artificially easy. The ablation study shows removing dynamic information (p=0.06) and link prediction (p=0.10) is not statistically significant, while removing static information is (p=0.04), consistent with the learned components being the fragile, partly circular part of the evaluation. No uniqueness theorem or ansatz is smuggled via self-citation; the self-citations to Defect4ML and earlier rule-based tools supply data and fault definitions but are not the load-bearing reduction. Overall, the headline three-out-of-six accuracy claim is only partly affected because the two strongest categories (data, library mismatch) are static-rule based; the precision/recall wins on dynamic fault categories are the part that reduces by construction to training-set overlap.
Assumptions & free parameters
free parameters (3)
- KG fault rules in Table 2
- Ranking priors from Humbatova et al.
- Trained classifier parameters (RF, DT, KNN, NodePiece)
assumptions (3)
- domain assumption Faults in DL software can be mapped to KG facts and inferred from static/dynamic features.
- domain assumption The fault taxonomy in Table 2 is complete and correct for Keras/TensorFlow systems.
- domain assumption The 75-sample training set has accurate labels.
Cite this review
Pith. "Pith review of Fault Localization in Deep Learning-based Software: A System-level Approach." pith.science (2026). https://pith.science/paper/F5D5HEHM
@misc{pith2026241108172,
author = {Pith},
title = {Pith review of: Fault Localization in Deep Learning-based Software: A System-level Approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/F5D5HEHM}},
note = {Machine review of arXiv:2411.08172}
}
read the original abstract
Over the past decade, Deep Learning (DL) has become an integral part of our daily lives. This surge in DL usage has heightened the need for developing reliable DL software systems. Given that fault localization is a critical task in reliability assessment, researchers have proposed several fault localization techniques for DL-based software, primarily focusing on faults within the DL model. While the DL model is central to DL components, there are other elements that significantly impact the performance of DL components. As a result, fault localization methods that concentrate solely on the DL model overlook a large portion of the system. To address this, we introduce FL4Deep, a system-level fault localization approach considering the entire DL development pipeline to effectively localize faults across the DL-based systems. In an evaluation using 100 faulty DL scripts, FL4Deep outperformed four previous approaches in terms of accuracy for three out of six DL-related faults, including issues related to data (84%), mismatched libraries between training and deployment (100%), and loss function (69%). Additionally, FL4Deep demonstrated superior precision and recall in fault localization for five categories of faults including three mentioned fault types in terms of accuracy, plus insufficient training iteration and activation function.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 2 Pith papers
-
LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs
An empirical GitHub mining study finds vLLM is the most adopted LLM serving framework, parallel and memory optimizations dominate, and multi-framework use is rare.
-
Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy
MCP server faults form five empirical categories—server setting, server/tool configuration, server/host configuration, documentation, and general programming—confirmed by a 41-practitioner survey.
Reference graph
Works this paper leans on
-
[23]
Jialun Cao, Meiziniu Li, Xiao Chen, Ming Wen, Yongqiang Tian, Bo Wu, and Shing-Chi Cheung. 2022. DeepFD: Automated Fault Diagnosis and Localization for Deep Learning Programs. arXiv preprint arXiv:2205.01938 (2022)
work page Pith review arXiv 2022
-
[64]
Mohammad Mehdi Morovati, Amin Nikanjam, Foutse Khomh, and Zhen Ming Jiang. 2023. Bugs in machine learning- based systems: a faultload benchmark. Empirical Software Engineering 28, 3 (2023), 62
work page 2023
-
[1]
[n. d.]. RDFLib: a Python library for working with RDF. https://github.com/RDFLib/rdflib. Accessed: 2024-08
2024
-
[2]
Cwm: a general-purpose data processor for the semantic web
2002. Cwm: a general-purpose data processor for the semantic web. http://www.w3.org/2000/10/swap/doc/cwm.html
2002
-
[3]
Apache Jena: a Java framework for writing Semantic Web applications
2009. Apache Jena: a Java framework for writing Semantic Web applications. https://jena.apache.org/
2009
-
[4]
RDF 1.1 Primer
2014. RDF 1.1 Primer. https://www.w3.org/TR/rdf11-primer/
2014
-
[5]
ISO/IEC/IEEE International Standard - Systems and software engineering–Vocabulary
2017. ISO/IEC/IEEE International Standard - Systems and software engineering–Vocabulary. ISO/IEC/IEEE 24765:2017(E) (2017), 1–541. https://doi.org/10.1109/IEEESTD.2017.8016712
arXiv 2017
-
[6]
Notation3 Language
2024. Notation3 Language. https://w3c.github.io/N3/spec/
2024
Show all 124 references
-
[7]
ISO/IEC 29881: 2010. 2010. Information technology, Systems and software engineering, FiSMA 1.1 functional size measurement method
2010
-
[8]
Mehdi Ali, Max Berrendorf, Charles Tapley Hoyt, Laurent Vermue, Sahand Sharifzadeh, Volker Tresp, and Jens Lehmann. 2021. PyKEEN 1.0: A Python Library for Training and Evaluating Knowledge Graph Embeddings. Journal of Machine Learning Research 22, 82 (2021), 1–6. http://jmlr.o...
2021
-
[9]
Mohammad Amin Alipour. 2012. Automated fault localization techniques: a survey. Oregon State University 54, 3 (2012)
2012
-
[10]
Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software engineering for machine learning: A case study. In 2019 IEEE/ACM 41st International Conference on Software Engineerin...
2019
-
[11]
Paul Ammann and Jeff Offutt. 2016. Introduction to software testing . Cambridge University Press
2016
-
[12]
Andrea Arcuri and Lionel Briand. 2014. A hitchhiker’s guide to statistical tests for assessing randomized algorithms in software engineering. Software Testing, Verification and Reliability 24, 3 (2014), 219–250
2014
-
[13]
Aitor Arrieta, Sergio Segura, Urtzi Markiegi, Goiuria Sagardui, and Leire Etxeberria. 2018. Spectrum-based fault localization in software product lines. Information and Software Technology 100 (2018), 18–31. https://doi.org/10. 1016/j.infsof.2018.03.008
2018
-
[14]
Jinheon Baek, Dong Bok Lee, and Sung Ju Hwang. 2020. Learning to extrapolate knowledge: Transductive few-shot out-of-graph link prediction. Advances in neural information processing systems 33 (2020), 546–560
2020
-
[15]
Barrasa, J
J. Barrasa, J. Webber, and J. Webber. 2023. Building Knowledge Graphs: A Practitioner’s Guide . O’Reilly. https: //books.google.ca/books?id=Ztb5zgEACAAJ
2023
-
[16]
Jason Bell. 2020. Machine learning: hands-on for developers and technical professionals . John Wiley & Sons
2020
-
[17]
Houssem Ben Braiek and Foutse Khomh. 2023. Testing feedforward neural networks training programs. ACM Transactions on Software Engineering and Methodology 32, 4 (2023), 1–61
2023
-
[18]
Tim Berners-Lee and Dan Connolly. 2008. Notation3 (N3): A readable RDF syntax. https://www.w3.org/ TeamSubmission/2008/SUBM-n3-20080114/
2008
-
[19]
Sweta Bhattacharya, Praveen Kumar Reddy Maddikunta, Quoc-Viet Pham, Thippa Reddy Gadekallu, Chiranji Lal Chowdhary, Mamoun Alazab, Md Jalil Piran, et al. 2021. Deep learning and medical image processing for coronavirus (COVID-19) pandemic: A survey. Sustainable cities and soci...
2021
-
[20]
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data . 1247–1250
2008
-
[21]
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in neural information processing systems 26 (2013)
2013
-
[22]
Léon Bottou. 2014. From machine learning to machine reasoning: An essay. Machine learning 94 (2014), 133–149
2014
-
[24]
Bahzad Charbuty and Adnan Abdulazeez. 2021. Classification based on decision tree algorithm for machine learning. Journal of Applied Science and Technology Trends 2, 01 (2021), 20–28
2021
-
[25]
Xiaojun Chen, Shengbin Jia, and Yang Xiang. 2020. A review: Knowledge reasoning over knowledge graph. Expert Systems with Applications 141 (2020), 112948
2020
-
[26]
Zhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang, Tao Xie, and Xuanzhe Liu. 2020. A Comprehensive Study on Challenges in Deploying Deep Learning Based Software. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Fo...
2020
-
[27]
Patrick Cousot and Radhia Cousot. 1977. Abstract interpretation: a unified lattice model for static analysis of programs by construction or approximation of fixpoints. InProceedings of the 4th ACM SIGACT-SIGPLAN symposium on Principles of programming languages. 238–252
1977
-
[28]
Ben De Meester, Dörthe Arndt, Pieter Bonte, Jabran Bhatti, Wim Dereuddre, Ruben Verborgh, Femke Ongenae, Filip De Turck, Erik Mannens, and Rik Van de Walle. 2015. Event-driven rule-based reasoning using EYE. In ISWC2015. CEUR
2015
-
[29]
Nan Duan, Duyu Tang, and Ming Zhou. 2020. Machine reasoning: Technology, dilemma and future. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts . 1–6
2020
-
[30]
Michael D Ernst. 2003. Static and dynamic analysis: Synergy and duality. In WODA 2003: ICSE Workshop on Dynamic Analysis. 24–27
2003
-
[31]
Fensel, U
D. Fensel, U. Şimşek, K. Angele, E. Huaman, E. Kärle, O. Panasiuk, I. Toma, J. Umbrich, and A. Wahler. 2020.Knowledge Graphs: Methodology, Tools and Selected Use Cases . Springer International Publishing. https://books.google.ca/books? id=1qnNDwAAQBAJ
2020
-
[32]
Daniel Galin. 2004. Software quality assurance: from theory to implementation . Pearson education
2004
-
[33]
Mikhail Galkin, Max Berrendorf, and Charles Tapley Hoyt. 2022. An open challenge for inductive link prediction on knowledge graphs. arXiv preprint arXiv:2203.01520 (2022)
2022 arXiv
-
[34]
Mikhail Galkin, Etienne Denis, Jiapeng Wu, and William L Hamilton. 2021. Nodepiece: Compositional and parameter- efficient representations of large knowledge graphs. arXiv preprint arXiv:2106.12144 (2021)
2021 arXiv
-
[35]
Jerry Gao, Chuanqi Tao, Dou Jie, and Shengqiang Lu. 2019. What is AI software testing? and why. In 2019 IEEE International Conference on Service-Oriented System Engineering (SOSE) . IEEE, 27–2709
2019
-
[36]
Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 249–256
2010
-
[37]
Liang Gong, Hongyu Zhang, Hyunmin Seo, and Sunghun Kim. 2014. Locating crashing faults based on crash stack traces. arXiv preprint arXiv:1404.4100 (2014)
2014 arXiv
-
[38]
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. MIT press
2016
-
[39]
Cyril Goutte and Eric Gaussier. 2005. A probabilistic interpretation of precision, recall and F-score, with implication for evaluation. In European conference on information retrieval . Springer, 345–359
2005
-
[40]
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing . Ieee, IEEE, Vancouver, BC, Canada, 6645–6649
2013
-
[41]
Sakshi Gupta. 2021. What Is the Best Language for Machine Learning? https://www.springboard.com/blog/data- science/best-language-for-machine-learning. Accessed: 2021-10-06
2021
-
[42]
Allan Hackshaw and Amy Kirkwood. 2011. Interpreting and reporting clinical trials with results of borderline significance. Bmj 343 (2011)
2011
-
[43]
Mark Harman and Robert Hierons. 2001. An overview of program slicing. software focus 2, 3 (2001), 85–92
2001
-
[44]
Stefanus A Haryono, Ferdian Thung, David Lo, Julia Lawall, and Lingxiao Jiang. 2021. Characterization and automatic updates of deprecated machine-learning api usages. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, IEEE, Luxembourg, 137–147
2021
-
[45]
Bruno Miranda Henrique, Vinicius Amorim Sobreiro, and Herbert Kimura. 2019. Literature review: Machine learning techniques applied to financial market prediction. Expert Systems with Applications 124 (2019), 226–251
2019
-
[46]
Nargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio, Andrea Stocco, and Paolo Tonella. 2020. Taxonomy of real faults in deep learning systems. In Proceedings of the ACM/IEEE 42nd international conference on software engineering. Association for Computing Mach...
2020
-
[47]
Nargiz Humbatova, Gunel Jahangirova, and Paolo Tonella. 2021. Deepcrime: mutation testing of deep learning systems based on real faults. In Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis. 67–78
2021
-
[48]
IEEE. 2010. ISO/IEC/IEEE International Standard - Systems and software engineering – Vocabulary . IEEE, 3 Park Avenue, New York NY 10016-5997, USA. 1–418 pages. https://doi.org/10.1109/IEEESTD.2010.5733835
2010
-
[49]
Md Johirul Islam, Giang Nguyen, Rangeet Pan, and Hridesh Rajan. 2019. A comprehensive study on deep learning bug characteristics. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineer...
2019
-
[50]
Md Johirul Islam, Rangeet Pan, Giang Nguyen, and Hridesh Rajan. 2020. Repairing deep neural networks: Fix patterns and challenges. In 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE) . IEEE, Association for Computing Machinery, New York, NY, USA, 1135...
2020
-
[51]
ISO/PAS. 2019. ISO/PAS 21448:2019 Road vehicles — Safety of the intended functionality. ISO/PAS 21448:2019 (2019)
2019
-
[52]
V Roshan Joseph. 2022. Optimal ratio for data splitting. Statistical Analysis and Data Mining: The ASA Data Science Journal 15, 4 (2022), 531–538
2022
-
[53]
Andrej Karpathy. 2019. A Recipe for Training Neural Networks . https://karpathy.github.io/2019/04/25/recipe/
2019
-
[54]
Leo Katz. 1953. A new status index derived from sociometric analysis. Psychometrika 18, 1 (1953), 39–43
1953
-
[55]
Diksha Khurana, Aditya Koli, Kiran Khatter, and Sukhdev Singh. 2023. Natural language processing: State of the art, current trends and challenges. Multimedia tools and applications 82, 3 (2023), 3713–3744
2023
-
[56]
John P Klein and Melvin L Moeschberger. 2006. Survival analysis: techniques for censored and truncated data . Springer Science & Business Media
2006
-
[57]
Jitender Kumar Chhabra and Varun Gupta. 2010. A survey of dynamic software metrics. Journal of computer science and technology 25 (2010), 1016–1029
2010
-
[58]
Dong Kyu Lee, Junyong In, and Sangseok Lee. 2015. Standard deviation and standard error of the mean. Korean journal of anesthesiology 68, 3 (2015), 220–223
2015
-
[59]
Michael R Lyu. 2007. Software reliability engineering: A roadmap. In Future of Software Engineering (FOSE’07) . IEEE, 153–170
2007
-
[60]
Christopher Manning, Prabhakar Raghavan, and Hinrich Schütze. 2010. Introduction to information retrieval.Natural Language Engineering 16, 1 (2010), 100–103
2010
-
[61]
John McCarthy. 2007. What is artificial intelligence? (2007)
2007
-
[62]
Richard Meyes, Melanie Lu, Constantin Waubert de Puiseau, and Tobias Meisen. 2019. Ablation studies in artificial neural networks. arXiv preprint arXiv:1901.08644 (2019)
2019 arXiv
-
[63]
Mohammad Mehdi Morovati, Amin Nikanjam, and Foutse Khomh. 2024. Paper replication package. https://github. com/mohmehmo/fl4deep. Accessed: 2024-07
2024
-
[65]
Mohammad Mehdi Morovati, Amin Nikanjam, Florian Tambon, Foutse Khomh, and Zhen Ming Jiang. 2024. Bug characterization in machine learning-based systems. Empirical Software Engineering 29, 1 (2024), 14
2024
-
[66]
Glenford J Myers, Corey Sandler, and Tom Badgett. 2011. The art of software testing . John Wiley & Sons
2011
-
[67]
Lee Naish, Hua Jie Lee, and Kotagiri Ramamohanarao. 2011. A Model for Spectra-Based Software Diagnosis. ACM Trans. Softw. Eng. Methodol. 20, 3, Article 11 (aug 2011), 32 pages. https://doi.org/10.1145/2000791.2000795
2011
-
[68]
Maximilian Nickel, Kevin Murphy, Volker Tresp, and Evgeniy Gabrilovich. 2015. A review of relational machine learning for knowledge graphs. Proc. IEEE 104, 1 (2015), 11–33
2015
-
[69]
Amin Nikanjam, Houssem Ben Braiek, Mohammad Mehdi Morovati, and Foutse Khomh. 2021. Automatic Fault Detection for Deep Learning Programs Using Graph Transformations. ACM Trans. Softw. Eng. Methodol. 31, 1, Article 14 (sep 2021), 27 pages. https://doi.org/10.1145/3470006
2021 doi
-
[70]
Amin Nikanjam, Mohammad Mehdi Morovati, Foutse Khomh, and Houssem Ben Braiek. 2022. Faults in deep reinforcement learning programs: a taxonomy and a detection approach. Automated Software Engineering 29, 1 (2022), 1–32
2022
-
[71]
Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall. 2018. Activation functions: Com- parison of trends in practice and research for deep learning. arXiv preprint arXiv:1811.03378 (2018)
2018 arXiv
-
[72]
Lawrence Page, Sergey Brin, Rajeev Motwani, Terry Winograd, et al. 1999. The pagerank citation ranking: Bringing order to the web. (1999), 1–17
1999
-
[73]
Jeff Z Pan. 2009. Resource description framework. InHandbook on ontologies. Springer, 71–90. https://jena.apache.org/
2009
-
[74]
Annibale Panichella and Cynthia CS Liem. 2021. What are we really testing in mutation testing for machine learning? a critical reflection. In 2021 IEEE/ACM 43rd International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER). IEEE, 66–70
2021
-
[75]
Mike Papadakis and Yves Le Traon. 2015. Metallaxis-FL: Mutation-Based Fault Localization. Softw. Test. Verif. Reliab. 25, 5–7 (aug 2015), 605–628. https://doi.org/10.1002/stvr.1509
2015 doi
-
[76]
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013. On the difficulty of training Recurrent Neural Networks. arXiv:1211.5063 [cs.LG] https://arxiv.org/abs/1211.5063
2013 arXiv
-
[77]
Ernst, Deric Pang, and Benjamin Keller
Spencer Pearson, José Campos, René Just, Gordon Fraser, Rui Abreu, Michael D. Ernst, Deric Pang, and Benjamin Keller. 2017. Evaluating and Improving Fault Localization. In2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE). 609–620. https://doi.org/10.11...
2017 doi
-
[78]
Goran Petrović, Marko Ivanković, Gordon Fraser, and René Just. 2021. Does mutation testing improve testing practices?. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 910–921
2021
-
[79]
Danijel Radjenović, Marjan Heričko, Richard Torkar, and Aleš Živkovič. 2013. Software fault prediction metrics: A systematic literature review. Information and Software Technology 55, 8 (2013), 1397–1418. https://doi.org/10.1016/j. infsof.2013.02.009 ACM Trans. Softw. Eng. Met...
2013 doi
-
[80]
Foyzur Rahman, Daryl Posnett, Abram Hindle, Earl Barr, and Premkumar Devanbu. 2011. BugCache for inspections: hit or miss?. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering. 322–331
2011
-
[81]
Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella. 2020. Testing machine learning based systems: a systematic mapping.Empirical Software Engineering 25, 6 (2020), 5193–5254
2020
-
[82]
Steven J Rigatti. 2017. Random forest. Journal of Insurance Medicine 47, 1 (2017), 31–39
2017
-
[83]
Emilio Rivera-Landos, Foutse Khomh, and Amin Nikanjam. 2021. The challenge of reproducible ML: an empirical study on the impact of bugs
2021
-
[84]
Andrea Rossi, Denilson Barbosa, Donatella Firmani, Antonio Matinata, and Paolo Merialdo. 2021. Knowledge graph embedding for link prediction: A comparative analysis. ACM Transactions on Knowledge Discovery from Data (TKDD) 15, 2 (2021), 1–49
2021
-
[85]
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. 2018. Modeling relational data with graph convolutional networks. InThe semantic web: 15th international conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, procee...
2018
-
[86]
Eldon Schoop, Forrest Huang, and Bjoern Hartmann. 2021. Umlaut: Debugging deep learning programs using program structure and model behavior. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–16
2021
-
[87]
David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. 2015. Hidden technical debt in machine learning systems. Advances in neural information processing systems 28 (2015)
2015
-
[88]
Diogo Seca. 2021. A review on oracle issues in machine learning. arXiv preprint arXiv:2105.01407 (2021)
2021 arXiv
-
[89]
Mehil B Shah, Mohammad Masudur Rahman, and Foutse Khomh. 2024. Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study. arXiv preprint arXiv:2401.03069 (2024)
2024 arXiv
-
[90]
Jonathan Richard Shewchuk. 2022. Concise Machine Learning
2022
-
[91]
Nasim Shirvani-Mahdavi, Farahnaz Akrami, Mohammed Samiul Saeef, Xiao Shi, and Chengkai Li. 2023. Comprehen- sive analysis of freebase and dataset creation for robust evaluation of knowledge graph link prediction models. In International Semantic Web Conference. Springer, 113–133
2023
-
[92]
Amit Singhal. 2012. Introducing the Knowledge Graph: things, not strings. https://blog.google/products/search/ introducing-knowledge-graph-things-not/
2012
-
[93]
Ezekiel Soremekun, Lukas Kirschner, Marcel Böhme, and Andreas Zeller. 2021. Locating faults with program slicing: an empirical analysis. Empirical Software Engineering 26 (2021), 1–45
2021
-
[94]
Kishore Sugali. 2021. Software testing: Issues and challenges of artificial intelligence & machine learning.International Journal of Artificial Intelligence & Applications 12 (2021)
2021
-
[95]
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision . 843–852
2017
-
[96]
Florian Tambon, Amin Nikanjam, Le An, Foutse Khomh, and Giuliano Antoniol. 2021. Silent Bugs in Deep Learning Frameworks: An Empirical Study of Keras and TensorFlow. arXiv preprint arXiv:2112.13314 (2021)
2021 arXiv
-
[97]
Jimin Tan, Jianan Yang, Sai Wu, Gang Chen, and Jake Zhao. 2021. A critical look at the current train/test split in machine learning. arXiv preprint arXiv:2106.04525 (2021)
2021 arXiv
-
[98]
Muhammad Usman, Youcheng Sun, Divya Gopinath, Rishi Dange, Luca Manolache, and Corina S Păsăreanu. 2023. An overview of structural coverage metrics for testing neural networks. International Journal on Software Tools for Technology Transfer 25, 3 (2023), 393–405
2023
-
[99]
Shikhar Vashishth, Soumya Sanyal, Vikram Nitin, and Partha Talukdar. 2019. Composition-based multi-relational graph convolutional networks. arXiv preprint arXiv:1911.03082 (2019)
2019 arXiv
-
[100]
Kavuri, and Kewen Yin
Venkat Venkatasubramanian, Raghunathan Rengaswamy, Surya N. Kavuri, and Kewen Yin. 2003. A review of process fault detection and diagnosis: Part III: Process history based methods. Computers & Chemical Engineering 27, 3 (2003), 327–346. https://doi.org/10.1016/S0098-1354(02)00162-X
2003 doi
-
[101]
Celine Vens, Jan Struyf, Leander Schietgat, Sašo Džeroski, and Hendrik Blockeel. 2008. Decision trees for hierarchical multi-label classification. Machine learning 73 (2008), 185–214
2008
-
[102]
Ruben Verborgh and Jos De Roo. 2015. Drawing conclusions from linked data on the web: The EYE reasoner. IEEE Software 32, 3 (2015), 23–27
2015
-
[103]
Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, E...
2020
-
[104]
Christina Voskoglou. 2017. What is the best programming language for Machine Learning . Towards Data Science. https://towardsdatascience.com/what-is-the-best-programming-language-for-machine-learning-a745c156d6b7
2017
-
[105]
Mohammad Wardat, Breno Dantas Cruz, Wei Le, and Hridesh Rajan. 2022. Deepdiagnosis: Automatically diagnosing faults and recommending actionable fixes in deep learning programs. InProceedings of the 44th International Conference on Software Engineering. 561–572
2022
-
[106]
Mohammad Wardat, Breno Dantas Cruz, Wei Le, and Hridesh Rajan. 2023. An Effective Data-Driven Approach for Localizing Deep Learning Faults. arXiv preprint arXiv:2307.08947 (2023)
2023 arXiv
-
[107]
Mohammad Wardat, Wei Le, and Hridesh Rajan. 2021. Deeplocalize: Fault localization for deep neural networks. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 251–262
2021
-
[108]
Robert West, Evgeniy Gabrilovich, Kevin Murphy, Shaohua Sun, Rahul Gupta, and Dekang Lin. 2014. Knowledge base completion via search-based question answering. In Proceedings of the 23rd international conference on World wide web. 515–526
2014
-
[109]
W Eric Wong, Ruizhi Gao, Yihao Li, Rui Abreu, and Franz Wotawa. 2016. A survey on software fault localization. IEEE Transactions on Software Engineering 42, 8 (2016), 707–740
2016
-
[110]
W Eric Wong and TH Tse. 2023. Handbook of software fault localization: foundations and advances . John Wiley & Sons
2023
-
[111]
Qingyao Wu, Mingkui Tan, Hengjie Song, Jian Chen, and Michael K Ng. 2016. ML-FOREST: A multi-label tree ensemble method for multi-label classification. IEEE transactions on knowledge and data engineering 28, 10 (2016), 2665–2680
2016
-
[112]
Baowen Xu, Ju Qian, Xiaofang Zhang, Zhongqiang Wu, and Lin Chen. 2005. A brief survey of program slicing. ACM SIGSOFT Software Engineering Notes 30, 2 (2005), 1–36
2005
-
[113]
Orhan G. Yalçın. 2021. Top 5 Deep Learning Frameworks to Watch in 2021 and Why Tensor- Flow. https://towardsdatascience.com/top-5-deep-learning-frameworks-to-watch-in-2021-and-why-tensorflow- 98d8d6667351 Accessed: 2022-12-29
2021
-
[114]
Yilin Yang, Tianxing He, Zhilong Xia, and Yang Feng. 2022. A comprehensive empirical study on bug characteristics of deep learning frameworks. Information and Software Technology 151 (2022), 107004
2022
-
[115]
Xiao Yu, Kwabena Ebo Bennin, Jin Liu, Jacky Wai Keung, Xiaofei Yin, and Zhou Xu. 2019. An empirical study of learning to rank techniques for effort-aware defect prediction. In 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER) . I...
2019
-
[116]
Ekim Yurtsever, Jacob Lambert, Alexander Carballo, and Kazuya Takeda. 2020. A survey of autonomous driving: Common practices and emerging technologies. IEEE access 8 (2020), 58443–58469
2020
-
[117]
Jie M Zhang, Mark Harman, Lei Ma, and Yang Liu. 2020. Machine learning testing: Survey, landscapes and horizons. IEEE Transactions on Software Engineering (2020)
2020
-
[118]
Min-Ling Zhang and Zhi-Hua Zhou. 2005. A k-nearest neighbor based algorithm for multi-label classification. In 2005 IEEE international conference on granular computing , Vol. 2. IEEE, 718–721
2005
-
[119]
Shichao Zhang. 2021. Challenges in KNN classification. IEEE Transactions on Knowledge and Data Engineering 34, 10 (2021), 4663–4675
2021
-
[120]
Xiangyu Zhang, Neelam Gupta, and Rajiv Gupta. 2006. Locating faults through automated predicate switching. In Proceedings of the 28th international conference on Software engineering . 272–281
2006
-
[121]
Xiangyu Zhang, Neelam Gupta, and Rajiv Gupta. 2007. A study of effectiveness of dynamic slicing in locating real faults. Empirical Software Engineering 12, 2 (2007), 143–160
2007
-
[122]
Xiaoyu Zhang, Juan Zhai, Shiqing Ma, and Chao Shen. 2021. AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair System. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 359–371
2021
-
[123]
Yuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong, Shing-Chi Cheung, and Tao Xie. 2020. Detecting numerical bugs in neural network architectures. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Softw...
2020
-
[124]
Daming Zou, Jingjing Liang, Yingfei Xiong, Michael D Ernst, and Lu Zhang. 2019. An empirical study of fault localization families and their combinations. IEEE Transactions on Software Engineering 47, 2 (2019), 332–347. ACM Trans. Softw. Eng. Methodol., Vol. 1, No. 1, Article ....
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.