Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Evidence is All We Need: Do Self-Admitted Technical Debts Impact Method-Level Maintenance?

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Self-admitted technical debt comments mark methods that become larger, more complex, and more prone to bugs and changes.

desk verdict Method-level SATD study with real associations, but missing age/size controls make the causal claims premature. read the letter →

arxiv 2411.13777 v2 pith:VLLOB6Z2 submitted 2024-11-21 cs.SE

classification cs.SE
keywords self-admittedtechnicaldebtmethod-levelanalysiscodemetricsbug-pronenesschange-pronenesssoftwaremaintenanceempiricalengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether self-admitted technical debt (SATD)—comments in which developers acknowledge substandard code—genuinely harms software maintenance, and it answers at the granularity of individual methods rather than whole files or classes. On 774,051 Java methods from 49 open-source projects, methods that begin life with a SATD comment end up significantly larger, more complex, less readable, and less maintainable than methods that never carry such comments, across all 14 code metrics studied. The same methods undergo more revisions and are more likely to be touched by bug-fixing commits, and the debt itself is rarely paid off: more than 60% of SATD comments are never removed, and 20% of those that are removed take over 1,000 days. If the claim holds, it supplies the missing empirical evidence that SATD is not just a label but a genuine maintenance liability, justifying early detection and remediation.

What carries the argument

The load-bearing mechanism is the method-level comparison design built on two tools. CodeShovel reconstructs each method's full history across renames, moves, and signature changes, so a method can be followed from birth to its current version; the SATD detector, a text-mining classifier, labels comments as SATD. A method counts as SATD if its initial version contains a flagged comment, and as NOT-SATD only if no version ever contains one, which lets the authors attribute later maintenance differences to the debt present at the method's origin. Code metrics, revision counts, edit distances, and bug-fix associations are then compared between the two groups with CDF plots, Wilcoxon rank-sum tests, and Cliff's delta effect sizes, with a 2-year age normalization applied to the change-proneness analysis.

What would settle it

Take methods introduced in the same project and year with similar initial size and cyclomatic complexity, split them by whether the first version carries a SATD comment, and compare future revision counts and bug-fix rates; if the gap shrinks to nothing, the claim that SATD itself drives maintenance problems is falsified.

Watch

Extended reading notes

Core claim

The central discovery is that SATD's widely assumed harm becomes visible once analysis moves from the file or class level to the method level. Methods whose first version contains a comment flagged as SATD by an automated detector are, at method level, statistically distinct from never-SATD methods: larger size, higher cyclomatic complexity, lower readability scores, worse maintainability index, more revisions and larger edit distances, and higher bug ratios under three different bug-definition datasets. The paper also finds the debt is persistent: with method history tracked across renames and moves, 61% of SATDs are never resolved, and the resolution-time distribution shows 60% of resolved debts taking at least 100 days and 20% taking more than 1,000 days. The authors present these results as the concrete evidence that earlier file- and class-level studies failed to find.

Load-bearing premise

The comparison assumes methods with and without SATD comments are otherwise alike, so if SATD methods tend to be older or larger from the start, the observed differences could come from age or size rather than from the debt itself.

Editorial extensions

If this is right

  • A single comment in a method's first version is a cheap early warning that the method will consume disproportionate maintenance effort.
  • Because the differences show up in every one of 14 code metrics, SATD deserves a slot as a feature in bug- and change-prediction models, which the paper notes currently underuse it.
  • Long-lived debts, with 20% taking over 1,000 days to resolve, mean deferring SATD cleanup is likely to compound maintenance cost rather than save it.
  • The method-level granularity resolves the earlier null results from file- and class-level studies, so future SATD impact studies should operate at method level.
  • Practitioner-facing tools can flag methods with SATD comments as candidates for prioritized refactoring or testing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same data could support a stronger test: match SATD and NOT-SATD methods on initial size and complexity at introduction, then compare their growth rates, which would isolate the debt's marginal effect from method size.
  • Because the paper pools all SATD types, a natural extension is to split by debt category (design, defect, documentation, requirement, test); effect sizes likely vary by type and would tell developers which debts to repay first.
  • The age normalization was applied only to change-proneness; repeating the readability and bug analyses on 2-year-normalized cohorts would show whether the quality gaps are age artifacts.
  • If developers write SATD comments under time pressure, the comment may be a proxy for rushed work; controlling for commit timing relative to releases would separate the debt signal from the haste signal.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies 774,051 Java methods from 49 open-source projects, comparing methods that begin with self-admitted technical debt (SATD) comments to methods that never contain SATD comments. Four research questions assess code quality metrics (RQ1), change-proneness (RQ2), bug-proneness (RQ3), and SATD resolution time (RQ4). The authors report that SATD methods are larger, more complex, less readable, more change-prone, and more bug-prone, and that more than 61% of SATD comments are never removed. They conclude that SATD has a significant negative impact on method-level maintainability and should be proactively managed.

Significance. The contribution is potentially significant: it is one of the first large-scale method-level studies to link SATD to maintenance indicators, and it uses a substantially larger dataset than prior class-level studies, with method histories tracked by CodeShovel, a replication package on Zenodo, and multiple robustness checks (Benjamini-Yekutieli correction, three bug-labeling datasets, project-level analyses). If the observed associations are robust to confounding, the results would overturn prior null findings and support investment in SATD detection and remediation. However, the paper's central causal claim is not yet supported because the main comparisons do not control for method age or size, as detailed below.

major comments (3)
  1. [§IV-C, RQ3] The bug-proneness comparison in RQ3 compares raw bug ratios between SATD and NOT-SATD methods without controlling for method age or size. The paper itself notes in §IV-B that age is strongly correlated with change- and bug-proneness and applies a 2-year age normalization there, but no equivalent adjustment is made for the bug ratios in Table VI and Figure 3. The aggregated gap (0.396 vs 0.213 for HighRecall) is exactly what one would expect if SATD methods are older or larger on average, since older methods have longer exposure to bug-fix commits. The authors should use matching (e.g., on age, size, and introduction-time complexity) or a multivariate model with grouped project effects before claiming that SATD causes bug-proneness. As written, the abstract's claim of a 'higher tendency for bugs' is only an unadjusted association.
  2. [§IV-A, RQ1] For RQ1, code metrics are measured 'from the most recent versions of the methods' (Section IV-A.1), and no baseline at method introduction is reported. Because the labeling approach (Section III-D) classifies methods by whether they start with SATD, the observed cross-sectional differences in Table II could reflect pre-existing size, complexity, or age differences at the time of introduction rather than evolution caused by SATD. For example, a method may receive a SATD comment precisely because it is already large and complex. The authors should either report the same metric comparisons at the introduction version, or match methods on introduction-time size, complexity, and age. Without this, RQ1 does not support the summary statement that SATD 'degrad[es]' code quality over time.
  3. [§V, Threats to Validity] The threats-to-validity section does not mention the comparability of the SATD and NOT-SATD groups, which is the central threat to RQ1-RQ3. Internal validity is discussed only in terms of statistical tests and SATD comment identification, while the age/size confounding that the authors acknowledge and address in RQ2 is omitted for RQ1 and RQ3. The paper should include this as an explicit threat and provide the adjusted or matched results (or, if adjustment is infeasible with the current data, temper the causal wording throughout the abstract, RQ summaries, and conclusion).
minor comments (4)
  1. [§IV-A, Table III] The columns N, S, M, L appear to be computed only among projects with statistically significant differences, while P>0.05 is computed over all projects; please state this in the table caption or legend, since the row sums currently exceed 100% (e.g., Readability sums to 106.52).
  2. [§IV-B] The 2-year age normalization is described in one sentence but its effect is not quantified; report the results before and after normalization (or at least a summary in the text) so readers can evaluate the robustness claim.
  3. [§IV-C, Table VI] Report the counts behind the bug ratios (number of bug-prone methods and total methods per dataset); the HighPrecision ratios are very small and counts would aid interpretation of the small effect size.
  4. [Figures 1–3] The text and figures use both 'NOT-SATD' and 'NOT_SATD'; make the labeling consistent.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the study's central comparisons are observational and do not reduce to their inputs by construction.

full rationale

This is an observational empirical study rather than a derivation. The paper labels a method as SATD if it contains at least one SATD comment in its initial version and as NOT-SATD if it never contains such a comment; it then compares externally measured quantities across these groups. None of the compared metrics is defined in terms of the SATD label: size, readability, complexity, maintainability index, revision counts, diff sizes, edit distances, and bug-ratios are all measured independently of the grouping, so RQ1-RQ3 do not reduce by definition. RQ4's 'resolution' is a measurement convention (comment disappearance), and the paper explicitly acknowledges the associated construct-validity threat (comments may remain after code is fixed), making it a limitation rather than a circular step. The paper cites the authors' prior work (CodeShovel [30], the age-normalization recommendation [9], the HighPrecision bug-labeling approach [4], and the non-critical-change exclusion [52]), but these citations supply tools, definitions, or robustness checks rather than the conclusion itself. The central SATD versus NOT-SATD comparisons would stand on the observed data even if those citations were removed. The strongest concern raised in the skeptical reading is that RQ1 and RQ3 compare groups without age or size matching, so older or larger methods may drive the observed differences. That is a confounding and construct-validity concern, not circularity: it questions whether the comparison supports the causal interpretation, but it does not show that any claimed result is equivalent to its own input by construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The empirical comparisons rest on measurement assumptions about SATD detection, method history, and bug labeling, plus a comparability assumption that is only partially tested. The paper introduces no free parameters or invented entities.

assumptions (4)
  • domain assumption A comment flagged by the SATD detector (Liu et al.) validly indicates self-admitted technical debt.
    Used in Section III-C; no independent accuracy check is reported on the studied projects, and the paper notes LLM detectors may be more accurate.
  • domain assumption CodeShovel reconstructs complete method histories, including renames and moves, without systematic bias.
    Relied on throughout Section III-B and RQ4; CodeTracker is cited as more accurate on a researcher-built oracle, so the chosen tool has known error risk.
  • domain assumption Bug-fix keyword heuristics identify bug-prone methods well enough for comparison.
    Used in Section IV-C; three datasets with different recall and precision are built on heuristics that have known false positives and false negatives.
  • domain assumption Method age and size are not confounding factors in RQ1 and RQ3 comparisons.
    No matching or multivariate control is applied in RQ1 and RQ3; the paper applies age normalization only in RQ2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evidence is All We Need: Do Self-Admitted Technical Debts Impact Method-Level Maintenance?." pith.science (2026). https://pith.science/paper/VLLOB6Z2

@misc{pith2026241113777,
  author       = {Pith},
  title        = {Pith review of: Evidence is All We Need: Do Self-Admitted Technical Debts Impact Method-Level Maintenance?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VLLOB6Z2}},
  note         = {Machine review of arXiv:2411.13777}
}
read the original abstract

Self-Admitted Technical Debt (SATD) refers to the phenomenon where developers explicitly acknowledge technical debt through comments in the source code. While considerable research has focused on detecting and addressing SATD, its true impact on software maintenance remains underexplored. The few studies that have examined this critical aspect have not provided concrete evidence linking SATD to negative effects on software maintenance. These studies, however, focused only on file- or class-level code granularity. This paper aims to empirically investigate the influence of SATD on various facets of software maintenance at the method level. We assess SATD's effects on code quality, bug susceptibility, change frequency, and the time practitioners typically take to resolve SATD. By analyzing a dataset of 774,051 methods from 49 open-source projects, we discovered that methods containing SATD are not only larger and more complex but also exhibit lower readability and a higher tendency for bugs and changes. We also found that SATD often remains unresolved for extended periods, adversely affecting code quality and maintainability. Our results provide empirical evidence highlighting the necessity of early identification, resource allocation, and proactive management of SATD to mitigate its long-term impacts on software quality and maintenance costs.

Figures

Figures reproduced from arXiv: 2411.13777 by the authors.

Figure 1
Figure 1. Differences in distribution for Size, McCabe, totalFanOut, Readability, SimpleReadability, and MaintainabilityIndex in [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Differences in distribution for #Revisions, DiffSizes, and CriticalEditDistances. Clearly, SATD methods tend to have [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Differences in the distribution of bug ratios in three datasets: SATD methods tend to have more bugs than NOT-SATD [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Figures (a) and (b) show the distribution of the percent of unresolved SATD methods and the distribution of SATD [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding the Effectiveness of LLMs in Automated Self-Admitted Technical Debt Repayment

    cs.SE 2025-01 conditional novelty 6.0 of 10

    Large new Python/Java benchmarks, diff-based metrics (BLEU-diff, CrystalBLEU-diff, LEMOD), and an LLM evaluation showing ~10% exact-match SATD repayment and larger models winning on fine-grained metrics.

Reference graph

Works this paper leans on

89 extracted references · 74 canonical work pages · cited by 1 Pith paper

  1. [1]

    The use of software complexity metrics in software maintenance,

    D. Kafura and G. Reddy, “The use of software complexity metrics in software maintenance,” IEEE Transactions on Software Engineering , vol. SE-13, no. 3, pp. 335–343, 1987

  2. [2]

    The role of method chains and comments in software readability and comprehension – an experiment,

    J. B ¨orstler and B. Paech, “The role of method chains and comments in software readability and comprehension – an experiment,” IEEE Transactions on Software Engineering , vol. 42, pp. 1–1, 09 2016

  3. [3]

    A study of reasoning processes in software maintenance management,

    M. Carr and C. Wagner, “A study of reasoning processes in software maintenance management,” Information Technology and Management , vol. 3, pp. 181–203, 2002

  4. [4]

    Method- level bug prediction: Problems and promises,

    S. Chowdhury, G. Uddin, H. Hemmati, and R. Holmes, “Method- level bug prediction: Problems and promises,” ACM Trans. Softw. Eng. Methodol., vol. 33, no. 4, apr 2024

  5. [5]

    On the performance of method-level bug prediction: A negative result,

    L. Pascarella, F. Palomba, and A. Bacchelli, “On the performance of method-level bug prediction: A negative result,” Journal of Systems and Software, vol. 161, p. 110487, 2020

  6. [6]

    Using source code metrics to predict change-prone java interfaces,

    D. Romano and M. Pinzger, “Using source code metrics to predict change-prone java interfaces,” in 2011 27th IEEE International Con- ference on Software Maintenance , 2011, pp. 303–312

  7. [7]

    On the relation of test smells to software code quality,

    D. Spadini, F. Palomba, A. Zaidman, M. Bruntink, and A. Bacchelli, “On the relation of test smells to software code quality,” in 2018 IEEE International Conference on Software Maintenance and Evolution, 2018

  8. [8]

    On the ability of complexity metrics to predict fault-prone classes in object-oriented systems,

    Y . Zhou, B. Xu, and H. Leung, “On the ability of complexity metrics to predict fault-prone classes in object-oriented systems,” Journal of Systems and Software , vol. 83, no. 4, pp. 660–674, 2010

Show all 89 references
  1. [9]

    Revisiting the debate: Are code metrics useful for measuring maintenance effort?

    S. Chowdhury, R. Holmes, A. Zaidman, and R. Kazman, “Revisiting the debate: Are code metrics useful for measuring maintenance effort?” Empirical Software Engineering (EMSE) , vol. 27, no. 6, p. 31, 2022

  2. [10]

    Tracy: A business-driven technical debt prioritization framework,

    R. R. De Almeida, C. Treude, and U. Kulesza, “Tracy: A business-driven technical debt prioritization framework,” in 2019 IEEE International Conference on Software Maintenance and Evolution , 2019

  3. [11]

    An exploratory study on self-admitted technical debt,

    A. Potdar and E. Shihab, “An exploratory study on self-admitted technical debt,” in Software Maintenance and Evolution (ICSME), 2014 IEEE International Conference on . IEEE, 2014, pp. 91–100

  4. [12]

    An exploratory study of the relationship between satd and other software development activities,

    S. Esfandiari and A. Sami, “An exploratory study of the relationship between satd and other software development activities,” in 2023 13th International Conference on Computer and Knowledge Engineering (ICCKE). IEEE, 2023, pp. 096–101

  5. [13]

    A systematic mapping study on technical debt and its management,

    Z. Li, P. Avgeriou, and P. Liang, “A systematic mapping study on technical debt and its management,” Journal of Systems and Software , vol. 101, pp. 193–220, 2015

  6. [14]

    Satd detector: a text-mining-based self-admitted technical debt detection tool,

    Z. Liu, Q. Huang, X. Xia, E. Shihab, D. Lo, and S. Li, “Satd detector: a text-mining-based self-admitted technical debt detection tool,” in Pro- ceedings of the 40th International Conference on Software Engineering: Companion Proceeedings, 2018, p. 9–12

  7. [15]

    Neural network-based detection of self-admitted technical debt: From perfor- mance to explainability,

    X. Ren, Z. Xing, X. Xia, D. Lo, X. Wang, and J. Grundy, “Neural network-based detection of self-admitted technical debt: From perfor- mance to explainability,” ACM transactions on software engineering and methodology (TOSEM), vol. 28, no. 3, pp. 1–45, 2019

  8. [16]

    Detecting and quantifying different types of self-admitted technical debt,

    E. d. S. Maldonado and E. Shihab, “Detecting and quantifying different types of self-admitted technical debt,” in 2015 IEEE 7Th international workshop on managing technical debt (MTD) , 2015, pp. 9–15

  9. [17]

    Automatic identification of self- admitted technical debt from four different sources,

    Y . Li, M. Soliman, and P. Avgeriou, “Automatic identification of self- admitted technical debt from four different sources,” Empirical Software Engineering, vol. 28, no. 3, p. 65, 2023

  10. [18]

    How far have we progressed in identifying self-admitted technical debts? a com- prehensive empirical study,

    Z. Guo, S. Liu, J. Liu, Y . Li, L. Chen, H. Lu, and Y . Zhou, “How far have we progressed in identifying self-admitted technical debts? a com- prehensive empirical study,”ACM Transactions on Software Engineering and Methodology (TOSEM) , vol. 30, no. 4, pp. 1–56, 2021

  11. [19]

    Self- admitted technical debt in r: detection and causes,

    R. Sharma, R. Shahbazi, F. H. Fard, Z. Codabux, and M. Vidoni, “Self- admitted technical debt in r: detection and causes,” Automated Software Engineering, vol. 29, no. 2, p. 53, 2022

  12. [20]

    Automatically learning patterns for self-admitted technical debt removal,

    F. Zampetti, A. Serebrenik, and M. Di Penta, “Automatically learning patterns for self-admitted technical debt removal,” in 2020 IEEE 27th International conference on software analysis, evolution and reengineer- ing (SANER), 2020, pp. 355–366

  13. [21]

    Was self-admitted technical debt removal a real removal? an in- depth perspective,

    ——, “Was self-admitted technical debt removal a real removal? an in- depth perspective,” in Proceedings of the 15th international conference on mining software repositories , 2018, pp. 526–536

  14. [22]

    An empirical study on the removal of self-admitted technical debt,

    E. D. S. Maldonado, R. Abdalkareem, E. Shihab, and A. Serebrenik, “An empirical study on the removal of self-admitted technical debt,” in 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2017, pp. 238–248

  15. [23]

    An empirical study on self-admitted technical debt in modern code review,

    Y . Kashiwa, R. Nishikawa, Y . Kamei, M. Kondo, E. Shihab, R. Sato, and N. Ubayashi, “An empirical study on self-admitted technical debt in modern code review,” Information and Software Technology , vol. 146, p. 106855, 2022

  16. [24]

    Presti: Predicting repayment effort of self-admitted technical debt using textual information,

    Y . Li, M. Soliman, and P. Avgeriou, “Presti: Predicting repayment effort of self-admitted technical debt using textual information,” Available at SSRN 4724886

  17. [25]

    Examining the impact of self- admitted technical debt on software quality,

    S. Wehaibi, E. Shihab, and L. Guerrouj, “Examining the impact of self- admitted technical debt on software quality,” in 2016 IEEE 23rd inter- national conference on software analysis, evolution, and reengineering (SANER), vol. 1, 2016, pp. 179–188

  18. [26]

    A large-scale empirical study on self-admitted technical debt,

    G. Bavota and B. Russo, “A large-scale empirical study on self-admitted technical debt,” in Proceedings of the 13th international conference on mining software repositories , 2016, pp. 315–326

  19. [27]

    A survey of self-admitted technical debt,

    G. Sierra, E. Shihab, and Y . Kamei, “A survey of self-admitted technical debt,” Journal of Systems and Software , vol. 152, pp. 70–82, 2019

  20. [28]

    Evidence-based soft- ware engineering,

    B. A. Kitchenham, T. Dyba, and M. Jorgensen, “Evidence-based soft- ware engineering,” in Proceedings. 26th International Conference on Software Engineering, 2004, pp. 273–281

  21. [29]

    Coverage is not strongly correlated with test suite effectiveness,

    L. Inozemtseva and R. Holmes, “Coverage is not strongly correlated with test suite effectiveness,” in Proceedings of the 36th International Conference on Software Engineering , 2014, p. 435–445

  22. [30]

    Codeshovel: Constructing method-level source code histories,

    F. Grund, S. Chowdhury, N. C. Bradley, B. Hall, and R. Holmes, “Codeshovel: Constructing method-level source code histories,” in Pro- ceedings of the International Conference on Software Engineering (ICSE), 2021, pp. 1510–1522

  23. [31]

    On the correlation between size and metric validity,

    Y . Gil and G. Lalouche, “On the correlation between size and metric validity,” Emp. Soft. Eng. , vol. 22, no. 5, pp. 2585–2611, 2017

  24. [32]

    The confounding effect of class size on the validity of object-oriented metrics,

    K. El Emam, S. Benlarbi, N. Goel, and S. N. Rai, “The confounding effect of class size on the validity of object-oriented metrics,” IEEE Transactions on Software Engineering, vol. 27, no. 7, pp. 630–650, 2001

  25. [33]

    Quantifying the effect of code smells on maintenance effort,

    D. I. K. Sjøberg, A. Yamashita, B. C. D. Anda, A. Mockus, and T. Dyb ˚a, “Quantifying the effect of code smells on maintenance effort,” IEEE Transactions on Software Engineering , vol. 39, no. 8, 2013

  26. [34]

    Empirical analysis of the relationship between cc and sloc in a large corpus of java methods,

    D. Landman, A. Serebrenik, and J. Vinju, “Empirical analysis of the relationship between cc and sloc in a large corpus of java methods,” in IEEE International Conference on Software Maintenance and Evolution, 2014, pp. 221–230

  27. [35]

    Accurate method and variable tracking in commit history,

    M. Jodavi and N. Tsantalis, “Accurate method and variable tracking in commit history,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2022, pp. 183–195

  28. [36]

    Using natural language processing to automatically detect self-admitted technical debt,

    E. Maldonado, E. Shihab, and N. Tsantalis, “Using natural language processing to automatically detect self-admitted technical debt,” IEEE Transactions on Software Engineering , 2017

  29. [37]

    An empirical study on the co-occurrence between refactoring actions and self-admitted technical debt removal,

    M. Iammarino, F. Zampetti, L. Aversano, and M. Di Penta, “An empirical study on the co-occurrence between refactoring actions and self-admitted technical debt removal,” Journal of Systems and Software , vol. 178, p. 110976, 2021

  30. [38]

    An exploratory study on the introduction and removal of different types of technical debt in deep learning frameworks,

    J. Liu, Q. Huang, X. Xia, E. Shihab, D. Lo, and S. Li, “An exploratory study on the introduction and removal of different types of technical debt in deep learning frameworks,” Empirical Software Engineering, vol. 26, pp. 1–36, 2021

  31. [39]

    Towards automatically addressing self-admitted technical debt: How far are we?

    A. Mastropaolo, M. Di Penta, and G. Bavota, “Towards automatically addressing self-admitted technical debt: How far are we?” in 2023 38th IEEE/ACM International Conference on Automated Software Engineer- ing (ASE), 2023, pp. 585–597

  32. [40]

    Recommending when design technical debt should be self-admitted,

    F. Zampetti, C. Noiseux, G. Antoniol, F. Khomh, and M. Di Penta, “Recommending when design technical debt should be self-admitted,” in 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2017, pp. 216–226

  33. [41]

    Evaluating complexity, code churn, and developer activity metrics as indicators of software vulnerabilities,

    Y . Shin, A. Meneely, L. Williams, and J. A. Osborne, “Evaluating complexity, code churn, and developer activity metrics as indicators of software vulnerabilities,” IEEE Transactions on Software Engineering , vol. 37, no. 6, pp. 772–787, 2011

  34. [42]

    A complexity measure,

    T. J. McCabe, “A complexity measure,” IEEE Transactions on Software Engineering, vol. SE-2, no. 4, pp. 308–320, 1976

  35. [43]

    An exploratory study of bug prediction at the method level,

    R. Mo, S. Wei, Q. Feng, and Z. Li, “An exploratory study of bug prediction at the method level,” Information and Software Technology , vol. 144, p. 110616, 2022

  36. [44]

    So you need more method level datasets for your software defect prediction? voil `a!

    T. Shippey, T. Hall, S. Counsell, and D. Bowes, “So you need more method level datasets for your software defect prediction? voil `a!” in ESEM ’16, 2016

  37. [45]

    Analysis of the reliability of a subset of change metrics for defect prediction,

    R. Moser, W. Pedrycz, and G. Succi, “Analysis of the reliability of a subset of change metrics for defect prediction,” in Proceedings of the Second ACM-IEEE International Symposium on Empirical Software Engineering and Measurement , 2008, p. 309–311

  38. [46]

    An exploratory study of the impact of antipatterns on class change- and fault-proneness,

    F. Khomh, M. D. Penta, Y .-G. Gu ´eh´eneuc, and G. Antoniol, “An exploratory study of the impact of antipatterns on class change- and fault-proneness,” Empirical software engineering : an international journal, vol. 17, no. 3, pp. 243–275, 2012

  39. [47]

    Toward a smell-aware bug prediction model,

    F. Palomba, M. Zanoni, F. A. Fontana, A. D. Lucia, and R. Oliveto, “Toward a smell-aware bug prediction model,” in IEEE Transactions on Software Engineering, vol. 45, no. 2, 2017, pp. 194–218

  40. [48]

    Toward effective deployment of design patterns for software extension: A case study,

    T. H. Ng, S.-C. Cheung, W. K. Chan, and Y .-T. Yu, “Toward effective deployment of design patterns for software extension: A case study,” in Proceedings of the 2006 international workshop on Software quality , 2006, pp. 51–56

  41. [49]

    Construct validity in software engineering research and software metrics,

    P. Ralph and E. Tempero, “Construct validity in software engineering research and software metrics,” in Proceedings of the 22nd International Conference on Evaluation and Assessment in Software Engineering 2018 (Christchurch, New Zealand), 2018, pp. 13–23

  42. [50]

    On the ”naturalness

    B. Ray, V . Hellendoorn, S. Godhane, Z. Tu, A. Bacchelli, and P. De- vanbu, “On the ”naturalness” of buggy code,” in Proceedings of the 38th International Conference on Software Engineering (ICSE) , 2016

  43. [51]

    On tracking java methods with git mechanisms,

    Y . Higo, S. Hayashi, and S. Kusumoto, “On tracking java methods with git mechanisms,” JSS, vol. 165, p. 110571, 2020

  44. [52]

    Impact of methodological choices on the analysis of code metrics and maintenance,

    S. I. Ahmad, S. Chowdhury, and R. Holmes, “Impact of methodological choices on the analysis of code metrics and maintenance,” Journal of Systems and Software , 2024

  45. [53]

    An empirical study of self-admitted technical debt in machine learning software,

    A. Bhatia, F. Khomh, B. Adams, and A. E. Hassan, “An empirical study of self-admitted technical debt in machine learning software,” arXiv preprint arXiv:2311.12019, 2023

  46. [54]

    Power comparisons of shapiro-wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests,

    N. M. Razali, Y . B. Wah et al. , “Power comparisons of shapiro-wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests,” Journal of statistical modeling and analytics , vol. 2, no. 1, pp. 21–33, 2011

  47. [55]

    An empirical study on software defect prediction with a simplified metric set,

    P. He, B. Li, X. Liu, J. Chen, and Y . Ma, “An empirical study on software defect prediction with a simplified metric set,” Information and Software Technology, vol. 59, pp. 170–190, 2015

  48. [56]

    Savior: Towards bug-driven hybrid testing,

    Y . Chen, P. Li, J. Xu, S. Guo, R. Zhou, Y . Zhang, T. Wei, and L. Lu, “Savior: Towards bug-driven hybrid testing,” in 2020 IEEE Symposium on Security and Privacy (SP) , 2020, pp. 1580–1596

  49. [57]

    Testing of mobile applications in the wild: A large-scale empirical study on android apps,

    F. Pecorelli, G. Catolino, F. Ferrucci, A. De Lucia, and F. Palomba, “Testing of mobile applications in the wild: A large-scale empirical study on android apps,” in Proceedings of the 28th international conference on program comprehension, pp. 296–307

  50. [58]

    Greenscaler: training software energy models with automatic test generation,

    S. Chowdhury, S. Borle, S. Romansky, and A. Hindle, “Greenscaler: training software energy models with automatic test generation,” Em- pirical software engineering , vol. 24, no. 4, pp. 1649–1692, 2019

  51. [59]

    On the time-based conclusion stability of cross-project defect prediction models,

    A. A. Bangash, H. Sahar, A. Hindle, and K. Ali, “On the time-based conclusion stability of cross-project defect prediction models,”Empirical Softw. Engg., vol. 25, no. 6, 2020

  52. [60]

    An empirical study on maintainable method size in java,

    S. A. Chowdhury, G. Uddin, and R. Holmes, “An empirical study on maintainable method size in java,” in 19th International Conference on Mining Software Repositories , 2022, p. 252–264

  53. [61]

    Improving code readability models with textual features,

    S. Scalabrino, M. Linares-V ´asquez, D. Poshyvanyk, and R. Oliveto, “Improving code readability models with textual features,” in IEEE 24th International Conference on Program Comprehension , 2016, pp. 1–10

  54. [62]

    Learning a metric for code readabil- ity,

    R. P. L. Buse and W. R. Weimer, “Learning a metric for code readabil- ity,” IEEE Trans. Softw. Eng. , vol. 36, no. 4, pp. 546–558, 2010

  55. [63]

    A simpler model of software readability,

    D. Posnett, A. Hindle, and P. Devanbu, “A simpler model of software readability,” in Proceedings of the 8th Working Conference on Mining Software Repositories (MSR) , 2011, pp. 73–82

  56. [64]

    The role of method chains and comments in software readability and comprehension—an experiment,

    J. B ¨orstler and B. Paech, “The role of method chains and comments in software readability and comprehension—an experiment,” IEEE Trans- actions on Software Engineering , vol. 42, no. 9, pp. 886–898, 2016

  57. [65]

    An empirical study assessing source code readability in comprehension,

    J. Johnson, S. Lubo, N. Yedla, J. Aponte, and B. Sharif, “An empirical study assessing source code readability in comprehension,” in2019 IEEE International Conference on Software Maintenance and Evolution, 2019, pp. 513–523

  58. [66]

    A model for program complexity analysis,

    C. L. McClure, “A model for program complexity analysis,” in Pro- ceedings of the 3rd International Conference on Software Engineering , 1978, p. 149–157

  59. [67]

    Reading beside the lines: Indentation as a proxy for complexity metric,

    A. Hindle, M. W. Godfrey, and R. C. Holt, “Reading beside the lines: Indentation as a proxy for complexity metric,” in 16th IEEE International Conference on Program Comprehension , 2008, pp. 133– 142

  60. [68]

    The control of the false discovery rate in multiple testing under dependency,

    Y . Benjamini and D. Yekutieli, “The control of the false discovery rate in multiple testing under dependency,” Annals of statistics, pp. 1165–1188, 2001

  61. [69]

    Measuring the psychological complexity of software maintenance tasks with the halstead and mccabe metrics,

    B. Curtis, S. B. Sheppard, P. Milliman, M. A. Borst, and T. Love, “Measuring the psychological complexity of software maintenance tasks with the halstead and mccabe metrics,” IEEE Transactions on Software Engineering, vol. SE-5, no. 2, pp. 96–104, 1979

  62. [70]

    Cyclomatic complexity metric for component based software,

    U. Tiwari and S. Kumar, “Cyclomatic complexity metric for component based software,” SIGSOFT Softw. Eng. Notes , vol. 39, no. 1, p. 1–6, Feb. 2014

  63. [71]

    Evaluation of halstead and cyclomatic complexity metrics in measuring defect density,

    M. Alfadel, A. Kobilica, and J. Hassine, “Evaluation of halstead and cyclomatic complexity metrics in measuring defect density,” in 2017 9th IEEE-GCC Conference and Exhibition , 2017, pp. 1–9

  64. [72]

    Evaluating software complexity measures,

    E. J. Weyuker, “Evaluating software complexity measures,” IEEE Trans- actions on Software Engineering , vol. 14, no. 9, pp. 1357–1365, 1988

  65. [73]

    A critique of cyclomatic complexity as a software metric,

    M. Shepperd, “A critique of cyclomatic complexity as a software metric,” Software Engineering Journal , vol. 3, no. 2, pp. 30–36, 1988

  66. [74]

    A method for assessing class change proneness,

    E.-M. Arvanitou, A. Ampatzoglou, A. Chatzigeorgiou, and P. Avgeriou, “A method for assessing class change proneness,” in Proceedings of the 21st International Conference on Evaluation and Assessment in Software Engineering, 2017, pp. 186–195

  67. [75]

    An exploratory study of the impact of code smells on software change-proneness,

    F. Khomh, M. Di Penta, and Y .-G. Gueheneuc, “An exploratory study of the impact of code smells on software change-proneness,” in 2009 16th Working Conference on Reverse Engineering , 2009, pp. 75–84

  68. [76]

    Enhancing change prediction models using developer-related factors,

    G. Catolino, F. Palomba, A. De Lucia, F. Ferrucci, and A. Zaidman, “Enhancing change prediction models using developer-related factors,” Journal of Systems and Software , vol. 143, pp. 14–28, 2018

  69. [77]

    Improving change prediction models with code smell- related information,

    G. Catolino, F. Palomba, F. A. Fontana, A. De Lucia, A. Zaidman, and F. Ferrucci, “Improving change prediction models with code smell- related information,” Empirical Software Engineering , vol. 25, pp. 49– 95, 2020

  70. [78]

    Identifying risky areas of software code in agile/lean software development: An industrial experience report,

    V . Antinyan, M. Staron, W. Meding, P. ¨Osterstr¨om, E. Wikstrom, J. Wranker, A. Henriksson, and J. Hansson, “Identifying risky areas of software code in agile/lean software development: An industrial experience report,” in IEEE Conference on Software Maintenance, Reengineerin...

  71. [79]

    Soft- ware quality analysis by code clones in industrial legacy software,

    A. Monden, D. Nakae, T. Kamiya, S. Sato, and K. Matsumoto, “Soft- ware quality analysis by code clones in industrial legacy software,” in Proceedings IEEE Symposium on Software Metrics , 2002, pp. 87–94

  72. [80]

    Identifying complex functions: By investigating various aspects of code complexity,

    V . Antinyan, M. Staron, J. Derehag, M. Runsten, E. Wikstr ¨om, W. Med- ing, A. Henriksson, and J. Hansson, “Identifying complex functions: By investigating various aspects of code complexity,” in 2015 Science and Information Conference (SAI) , 2015, pp. 879–888

  73. [81]

    From aristotle to ringel- mann: a large-scale analysis of team productivity and coordination in open source software projects,

    I. Scholtes, P. Mavrodiev, and F. Schweitzer, “From aristotle to ringel- mann: a large-scale analysis of team productivity and coordination in open source software projects,” Empirical Software Engineering (EMSE), vol. 21, no. 2, pp. 642–683, 2016

  74. [82]

    Binary codes capable of correcting deletions, inser- tions, and reversals,

    V . I. Levenshtein, “Binary codes capable of correcting deletions, inser- tions, and reversals,” Soviet physics doklady , vol. 10, 1966

  75. [83]

    Big bangs and small pops: On critical cyclomatic complexity and developer integration behavior,

    D. St ˚ahl, A. Martini, and T. M ˚artensson, “Big bangs and small pops: On critical cyclomatic complexity and developer integration behavior,” in 2019 IEEE/ACM 41st International Conference on Software Engi- neering: (ICSE-SEIP), 2019, pp. 81–90

  76. [84]

    Automatically assessing code understandability: How far are we?

    S. Scalabrino, G. Bavota, C. Vendome, M. Linares-V ´asquez, D. Poshy- vanyk, and R. Oliveto, “Automatically assessing code understandability: How far are we?” in 32nd IEEE/ACM International Conference on Automated Software Engineering , 2017, pp. 417–427

  77. [85]

    Hey! are you committing tangled changes?

    H. Kirinuki, Y . Higo, K. Hotta, and S. Kusumoto, “Hey! are you committing tangled changes?” in Proceedings of the 22nd International Conference on Program Comprehension , 2014, pp. 262–265

  78. [86]

    The impact of tangled code changes on defect prediction models,

    K. Herzig, S. Just, and A. Zeller, “The impact of tangled code changes on defect prediction models,” Empirical Software Engineering , vol. 21, no. 2, pp. 303–336, 2016

  79. [87]

    Identifying reasons for software changes using historic databases,

    A. Mocku and L. G. V otta, “Identifying reasons for software changes using historic databases,” in Proceedings 2000 International Conference on Software Maintenance , 2000, pp. 120–130

  80. [88]

    Evaluating szz implementations through a developer- informed oracle,

    G. Rosa, L. Pascarella, S. Scalabrino, R. Tufano, G. Bavota, M. Lanza, and R. Oliveto, “Evaluating szz implementations through a developer- informed oracle,” in Proceedings of the 43rd International Conference on Software Engineering , 2021, p. 436–447

  81. [89]

    An empirical study on the effectiveness of large language models for satd identification and classification,

    M. S. Sheikhaei, Y . Tian, S. Wang, and B. Xu, “An empirical study on the effectiveness of large language models for satd identification and classification,” Empirical Software Engineering , vol. 29, no. 6, p. 159, 2024

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.