REVIEW 3 major objections 6 minor 15 references
Exploring Dependence, Overreliance, and Addiction Related Behaviors Associated with Large Language Model Use Among Software Engineers
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Survey of 119 developers finds LLM use is habitual and functional, not compulsive.
desk verdict Useful exploratory data on an understudied population, but the headline claim that addiction-related behavior is 'marginal' rests on an unstated weighting of discordant structured and open-ended responses. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a conceptual distinction among three constructs—dependence, overreliance, and addiction-related behaviors—operationalized through a qualitative survey. Frequency items (Q1–Q18) measure each construct, while open-ended questions ask participants to describe a concrete LLM use and to imagine a workday without these tools. Descriptive statistics summarize the structured responses and reflexive thematic analysis codes the open-ended ones; the interpretive pivot is that participants' narrative accounts of life without LLMs carry more weight than the frequency scales in deciding whether use is functional or compulsive.
What would settle it
Open the paper's Figure 6 and count the structured addiction responses at face value: for most items Q9–Q18, over half of participants selected 'Often' or 'Very Often,' which contradicts the claim that addiction-related behavior is marginal if those items are treated as the primary evidence. A field test would give developers a diary or telemetry study over several weeks to see whether self-reported loss of control—failed attempts to cut back, restlessness when blocked—matches actual usage and work disruption.
Extended reading notes
Core claim
The central claim is that among practicing software engineers, LLM use is habitual and preferential rather than compulsive. In the paper's account, functional dependence shows up as reduced efficiency when LLMs are unavailable—slower tasks, more manual searching—rather than an inability to work. Overreliance shows up not as blind acceptance of generated code (most respondents reject deploying LLM code without review) but as LLMs becoming the default first stop for information, pushing colleagues and documentation into a fallback role. Addiction-related behaviors are the least supported by participants' own descriptions: structured items captured frequent time overruns and difficulty reducing use, but open-ended accounts rarely described impaired control or significant work disruption, so the authors characterize those behaviors as marginal.
Load-bearing premise
The strongest conclusion—that addiction-related behaviors are marginal—depends on treating the open-ended comments as more diagnostic than the structured survey answers, where many addiction items drew 'Often' or 'Very Often' from majorities; the paper never states or justifies that weighting rule.
Editorial extensions
If this is right
- Organizations should treat frequent LLM use as functional dependence rather than a sign of addiction, and support developers in maintaining skills for independent work.
- Policy and training should target trust calibration: helping engineers decide when LLMs are the right source and when documentation, prior experience, or colleagues should lead.
- Because overreliance appears as LLMs displacing peers and documentation as the first source of support, teams should actively preserve peer consultation and documentation habits.
- Support should be tailored to career stage, since early- and mid-career developers report more LLM-centered information seeking than senior developers.
- Intensive professional use should not be labeled addiction without evidence of impaired control or disruption to professional work.
Reading between the lines
- Editorial inference: if the structured addiction items are weighted as heavily as the open-ended comments, the paper's 'marginal addiction' conclusion would likely reverse, since most Q9–Q18 items show majorities reporting frequent behavior.
- Editorial inference: the data suggest a practical metric for overreliance—capturing whether an engineer's first move is an LLM rather than a colleague or documentation—could be more diagnostic than the standard 'do you verify outputs' question.
- Editorial inference: a longitudinal study would test whether overreliance grows as LLMs become more reliable, since developers may verify less when errors become rarer, a trend the cross-sectional snapshot cannot capture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an exploratory survey of 119 software practitioners recruited through Prolific and mailing lists, examining behavioral patterns associated with dependence, overreliance, and addiction-related behaviors in professional LLM use. The instrument combines structured Likert items (Q1–Q18) and two open-ended questions (QC1, QC2), analyzed with descriptive statistics and reflexive thematic analysis. The authors report that LLM use is habitual and goal-oriented: dependence manifests as functional integration into routine work, overreliance as a preference for LLMs over documentation and peers, and addiction-related behaviors are described as marginal. The central claim, stated in Section 4.3, is that engineers' relationship with LLMs is 'habitual and preferential rather than compulsive.'
Significance. If the central claim is supported, the study makes a useful contribution by separating functional dependence from problematic technology use in a professional software engineering population, extending prior work that focused on students or general users. The paper includes several strengths: an internationally diverse sample, explicit data-quality screening, a positionality statement, and a discussion of trust calibration as a complement to verification. However, the headline conclusion that addiction-related behaviors are marginal is not directly supported by the structured data. Figure 6 shows majorities reporting 'Often' or 'Very Often' on ten addiction-related items, and the authors' countervailing evidence consists mainly of responses to QC2, a hypothetical question about a day without LLM access that was not designed to elicit impaired control or withdrawal. Because the weighting of the two discordant data sources is not stated or justified, the significance of the paper's main finding is conditional on a revision that either provides a defensible decision rule or qualifies the claim.
major comments (3)
- [§4.3, §4.2.3, Figure 6] The claim that addiction-related behaviors are 'present only marginally' is not warranted by the structured responses. Figure 6 shows that 56–71% of respondents answered 'Often' or 'Very Often' on Q9, Q14, Q15, Q16, Q17, and Q18, which cover unplanned overuse, unsuccessful reduction attempts, restlessness, irritability, and negative effects on job performance. Section 4.2.3 acknowledges these items 'captured behaviors commonly discussed in the technology addiction literature' but discounts them because open-ended comments 'rarely indicated loss of control.' Neither Section 3.5 nor Section 3.8 states a decision rule for weighting structured versus open-ended data, nor is the frequency of addiction-related codes in the open-ended data reported. Without such a rule, the marginal-status conclusion is an interpretive choice rather than a finding.
- [§4.2.3, Table 2, §3.5] The use of QC2 as the main counter-evidence is problematic. QC2 asks participants to imagine a day without LLM access and describe how their daily tasks would be affected; it was designed as a contextual workflow probe, not as a diagnostic elicitation of craving, withdrawal, or impaired control. The fact that participants mention switching to search engines or slower completion therefore does not indicate the absence of loss-of-control experiences. The paper should either report codes from open-ended questions that directly probe self-regulation and control, or explicitly acknowledge that QC2 cannot speak to the addiction construct.
- [§3.1, §3.2, §2.1] The operationalization of the constructs is partially circular. The structured items were derived 'directly from the behavioral characteristics identified in the literature' using the same definitions that later structure the interpretation (Section 2.1, Table 1), and Section 3.2 explicitly disclaims formal psychometric validation. Consequently, the mapping from item responses to constructs is uncertain, and the pattern in Figure 6 is partly by construction. The open-ended responses are the only independent grounding, but they are used selectively to override the structured items. The paper should either treat both sources as complementary and report convergence and divergence explicitly, or clearly label the structured items as non-validated exploratory indicators and adjust the strength of the claims accordingly.
minor comments (6)
- [§4.2.2, Figure 4] The text states that in Q7 'roughly half' reported 'Often' or 'Very Often' and that in Q8 a 'clear majority' did so, but Figure 4 shows 65% for Q7 and 53% for Q8; the descriptions should be corrected to match the figure.
- [§4.2.1] 'Forty five participants' should be hyphenated as 'Forty-five.'
- [§4.2.3, Figure 7] Figure 7 is referenced in the text but is not present in the provided manuscript; the experience-level results for addiction-related items cannot be verified without the figure.
- [§3.5, Figure 1] Figure 1 (Thematic Analysis) is referenced in Section 3.5 but is not included in the provided text; the four-stage process cannot be inspected.
- [§4.2.3] In the sentence describing emotional reactions, the quotation marks around 'tedious,' 'annoying,' and 'draining' are inconsistently rendered; please use matching quotation marks.
- [§3.2] The description of the licensed psychologist's review would benefit from a sentence summarizing the specific feedback that led to questionnaire changes, rather than only stating that feedback was incorporated.
Circularity Check
No significant circularity: the survey's central characterization is data-driven, and the under-justified weighting of structured versus open-ended responses is a validity limitation, not a circular reduction.
full rationale
The paper is an exploratory qualitative survey, not a formal derivation chain. The central claim—that dependence and overreliance dominate while addiction-related behaviors are marginal—is presented as an interpretation of participant responses, not as a consequence of the conceptual definitions in Section 2. Section 3.1 states that questionnaire items were 'derived directly from the behavioral characteristics identified in the literature,' but the findings in Sections 4.2 and 4.3 are supported by reported frequencies and thematic codes. The structured items Q9–Q18 could have supported an addiction-heavy conclusion given the high 'Often/Very Often' rates, yet the authors chose to weight the open-ended QC2 accounts more heavily. That weighting rule is not stated in Sections 3.5 or 3.8, and this is a genuine methodological limitation, but it is not circular: the open-ended responses are independent of the construct definitions, and the conclusion is not forced by the definitions themselves. The self-citations (e.g., Santos et al., 2025; De Sousa et al., 2025; Coutinho et al., 2024) are supporting references for LLM use in software engineering and positionality, and they are not load-bearing for the main claim. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no ansatz is smuggled in via citation. Therefore, no significant circularity is identified.
Assumptions & free parameters
assumptions (4)
- domain assumption Self-reported survey responses are valid indicators of actual dependence, overreliance, and addiction-related behaviors.
- ad hoc to paper The self-designed questionnaire items measure the intended constructs without formal psychometric validation.
- ad hoc to paper Open-ended qualitative responses are more diagnostic than structured Likert responses when the two conflict.
- domain assumption A sample recruited via Prolific and mailing lists represents professional software engineering experiences broadly enough for exploratory pattern identification.
Cite this review
Pith. "Pith review of Exploring Dependence, Overreliance, and Addiction Related Behaviors Associated with Large Language Model Use Among Software Engineers." pith.science (2026). https://pith.science/paper/YKHNAMZW
@misc{pith2026260805561,
author = {Pith},
title = {Pith review of: Exploring Dependence, Overreliance, and Addiction Related Behaviors Associated with Large Language Model Use Among Software Engineers},
year = {2026},
howpublished = {\url{https://pith.science/paper/YKHNAMZW}},
note = {Machine review of arXiv:2608.05561}
}
read the original abstract
The widespread adoption of Large Language Models (LLMs) has changed how software engineers perform everyday development activities. While these systems provide substantial support for tasks such as code generation, debugging, and documentation, their increasing integration into professional workflows has also raised questions regarding developers' reliance on these tools and the emergence of dependence, overreliance, and addiction-related behaviors. This study investigates how software engineers experience the use of LLMs during professional software development, with attention to behavioral patterns associated with dependence, overreliance, and addiction-related behaviors. An exploratory survey was conducted with 119 software practitioners. The data were analyzed using descriptive statistics and qualitative thematic analysis of participants' open-ended responses. Participants primarily described functional dependence, with LLMs becoming integrated into routine software engineering activities because of the productivity and efficiency they provide. Responses also suggested patterns consistent with overreliance, particularly through prioritizing LLMs over documentation or peer consultation while continuing to verify generated outputs. Reports associated with addiction-related behaviors were less common and primarily reflected difficulty moderating use or emotional attachment to the technology rather than impaired control. The findings suggest that LLMs are becoming a habitual component of professional software engineering practice. While most reported use appears functional, the results indicate the importance of promoting appropriate reliance by supporting trust calibration, professional judgment, and verification throughout software development.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[2]
Theonline survey as a qualitative research tool
Braun,V.,Clarke,V.,Boulton,E.,Davey,L.,McEvoy,C.,2021. Theonline survey as a qualitative research tool. International journal of social research methodology 24, 641–654. Chamberlain, S.R., Lochner, C., Stein, D.J., Goudriaan, A.E., van Holst, R.J., Zohar, J., Grant, J.E.,
work page 2021
-
[7]
ACM Transactions on Software Engineering and Methodology 34, 1–30
The current challenges of software engineering in the era of large language models. ACM Transactions on Software Engineering and Methodology 34, 1–30. George,D.,Mallery,P.,2018.Descriptivestatistics,in:IBMSPSSStatistics 25 Step by Step. Routledge, pp. 126–134. Haman,M.,Školník,M.,2023. Behindthechatgpthype:areitssuggestions contributing to addiction? Anna...
work page 2018
-
[8]
Fostering appropriate reliance on large language models: The roleofexplanations,sources,andinconsistencies,in:Proceedingsofthe 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–19. Kitchenham,B.A.,Pfleeger,S.L.,2008.Personalopinionsurveys,in:Guide to advanced empirical software engineering. Springer, pp. 63–92. Klingbeil, A., Grützner, C., ...
work page 2025
-
[9]
arXiv preprint arXiv:2506.08872
Your brain on chatgpt: Accumula- tion of cognitive debt when using an ai assistant for essay writing task. arXiv preprint arXiv:2506.08872 . Linaker,J.,Sulaman,S.M.,Höst,M.,deMello,R.M.,2015. Guidelinesfor conducting surveys in software engineering v. 1.1. Lund University 50, 1–64. Lo,D.,2023. Trustworthyandsynergisticartificialintelligenceforsoftware eng...
arXiv 2015
-
[12]
arXiv preprint arXiv:2203.14695
Recruiting software engineers on prolific. arXiv preprint arXiv:2203.14695 . Sallou, J., Durieux, T., Panichella, A.,
-
[28]
International Journal of Educational Technology in Higher Education 21,
Zhang,S.,Zhao,X.,Zhou,T.,Kim,J.H.,2024.Doyouhaveaidependency? the roles of academic self-efficacy, academic stress, and performance expectations on problematic ai usage behavior. International Journal of Educational Technology in Higher Education 21,
work page 2024
-
[85]
URL: http://www.sciencedirect.com/science/article/pii/S0749597896926746, doi:https://doi.org/10.1006/obhd.1996.2674. Will,R.P.,1991. Trueandfalsedependenceontechnology:Evaluationwith an expert system. Computers in human behavior 7, 171–183. Yankouskaya, A., Liebherr, M., Ali, R.,
-
[94]
Understandinghumanover-relianceontechnol- ogy
Baxter,N.,Kabi,F.,2017. Understandinghumanover-relianceontechnol- ogy. Long-Term Care 5,
work page 2017
Show all 15 references
-
[2016]
aiholic”: Questioning the “chatgpt addiction
Behavioural addiction—a rising tide? European neuropsychopharmacology 26, 841–855. Chen,P.,Alias,S.B.,2024. Opportunitiesandchallengesinthecultivation of software development professionals in the context of large language models, in: Proceedings of the 2024 International Sympo...
2024
-
[2020]
arXiv preprint arXiv:2010.03525
Empirical standards for software engineering research. arXiv preprint arXiv:2010.03525 . Rasnayaka, S., Wang, G., Shariffdeen, R., Iyer, G.N.,
2010
-
[2021]
Do you really code? designing and evaluating screening questions for online surveys with programmers, in: 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE), IEEE. pp. 537–548. De Sousa, B.F., de Souza Santos, R., Gama, K.,
2021
-
[2022]
arXiv preprint arXiv:2201.05348
Software engi- neering user study recruitment on prolific: An experience report. arXiv preprint arXiv:2201.05348 . Russo, D.,
-
[2023]
Large language models for software engineering: Survey and open problems, in: 2023 IEEE/ACM International Confer- ence on Software Engineering: Future of Software Engineering (ICSE- FoSE), IEEE. pp. 31–53. Fan, L., Liu, X., Wang, B., Wang, L.,
2023
-
[2024]
Breaking the silence: the threatsofusingllmsinsoftwareengineering,in:Proceedingsofthe2024 ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, pp. 102–106. Santos,I.,Magalhaes,C.,Santos,R.D.S.,2025.Model-assistedandhuman- guided: Perc...
2025
-
[2025]
Integrating po- sitionality statements in empirical software engineering research, in: 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE), IEEE. pp. 28–35. Fan, A., Gokkaya, B., Harman, M., Lyubarskiy, M., Sengu...
2025
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.