Pith. sign in

REVIEW 3 major objections 5 minor 70 references

The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A systematic review of 39 studies finds LLM coding assistants speed up developers and also create over-reliance, with mixed effects on code quality.

desk verdict A careful, reproducible first synthesis of LLM-assistant productivity evidence; the mandatory 'Productivity' search term puts a small but real question mark on the exact percentages. read the letter →

arxiv 2507.03156 v3 pith:2HESWYOD submitted 2025-07-03 cs.SE cs.AIcs.HC

classification cs.SEcs.AIcs.HC
keywords systematicreviewmappingstudyLLMassistantsdeveloperproductivitySPACEframeworkcodequalityGitHubCopilotChatGPT
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks what the peer-reviewed evidence actually says about how LLM-based coding assistants change software developer productivity, and it offers the first systematic synthesis of that evidence. The authors gathered 39 primary studies published between 2014 and December 2024 and found that most report real gains—faster task completion, less time spent searching for code, and automation of repetitive work—while a substantial minority report risks such as over-reliance, disrupted flow, and weakened team collaboration. The review also maps each study onto the SPACE framework's five productivity dimensions, showing that 90% of studies look at least two dimensions but only 15% look at four or more. A key unresolved point is code quality: studies report it improving in some contexts and degrading in others, with no settled explanation of when each happens. If the review is right, the field now has a reference map of what is known, what is measured, and where the gaps are.

What carries the argument

The SPACE framework (Satisfaction and well-being, Performance, Activity, Communication and collaboration, Efficiency and flow) is the organizing lens the review uses to classify every primary study's productivity measures into dimensions and sub-dimensions. That mapping carries the central quantitative claims of the paper: 90% of studies cover at least two SPACE dimensions, only 15% cover four or more, and Satisfaction, Performance, and Efficiency dominate while Communication and Activity lag. The review also uses the McLuhan Tetrad as an interpretive device to discuss enhancement, obsolescence, retrieval, and reversal of developer practices, but the SPACE mapping is what produces the paper's measurable findings.

What would settle it

Re-run the search with broadened queries that drop the strict proximity-to-'productivity' requirement and include grey literature and short papers, then check whether the 90%-at-least-two-SPACE-dimensions and 15%-at-least-four figures, and the conclusion that code-quality findings are contradictory, still hold; if the broadened set substantially changes these proportions or resolves the code-quality conflict, the review's synthesis would need revision.

Watch

Extended reading notes

Core claim

The paper establishes that the literature on LLM-assistants and developer productivity is young, fast-growing, and uneven: 90% of the 39 included studies examine at least two SPACE dimensions, yet only 15% extend beyond three, and the most studied dimensions are Satisfaction, Performance, and Efficiency while Communication and Activity remain under-explored. Across the corpus, commonly reported benefits are accelerated development, minimized code search, automation of trivial and repetitive tasks, and support for learning and code-adjacent work; commonly reported risks are failing to meet requirements, over-reliance and cognitive offloading, flow disruption, and reduced human-human collaboration. The single most contested outcome is code quality, which appears as both a benefit and a risk depending on task, context, and evaluation metrics. Methodologically, 59% of the studies are exploratory, 38% are laboratory experiments, and 90% rely on self-report data, with little longitudinal or team-based evaluation. The authors conclude that productivity is a multidimensional construct, that current evaluation tools are partial, and that future work should adopt shared frameworks, validated instruments, and team- and field-based designs.

Load-bearing premise

The synthesis and its headline percentages assume that the six-database search query with title, abstract, and keyword terms, plus snowballing, retrieved all or a representative sample of the relevant peer-reviewed studies; if a substantial share of studies used different terminology or were published in venues the query missed, the reported themes and SPACE percentages could shift.

Editorial extensions

If this is right

  • If the mapping is correct, future studies of LLM-assistants should routinely measure more than perceived speed and satisfaction, because the corpus shows that communication, activity, and well-being are rarely captured.
  • The unresolved code-quality signal implies that organizations should not assume throughput gains from LLM-assistants translate into better software, and that evaluation should track quality metrics separately from time savings.
  • The heavy reliance on self-reported data and short laboratory experiments means existing evidence says little about long-term skill erosion, cumulative technical debt, or team dynamics, so longitudinal and field studies would be the most informative next step.
  • The concentration of 77% of included studies in 2024 suggests the evidence base is extremely recent and may shift quickly as model capabilities and developer practices evolve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the SPACE percentages are representative, a practical testable extension is to run the same mapping on studies published after December 2024 and check whether the Communication and Activity shares have grown; the paper's own logic predicts they are the next dimensions to fill.
  • The contradictory code-quality findings hint that the relevant variable is not the assistant itself but the validation workflow around it; one could test this by comparing code-quality outcomes in studies that mandate review versus those that do not.
  • The near-total absence of well-being measures suggests an implicit blind spot: productivity gains that come with higher stress or burnout would not be visible in the current corpus, so a review of mental-health outcomes in LLM-assisted development would complement this synthesis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript reports a systematic review and mapping study of 39 peer-reviewed studies, published between January 2014 and December 2024, that examine the impact of LLM-based assistants on software developer productivity. The authors follow a Kitchenham-and-Charters-based protocol, report a PRISMA flow diagram, apply a quality assessment, and provide a public replication package. The review addresses four research questions: publication and venue characteristics, methodological strategies and instruments, reported benefits and risks, and mapping of the evidence onto the SPACE framework. The headline findings are that most primary studies (90%) adopt a multidimensional view of productivity, only 15% address four or more SPACE dimensions, Satisfaction/Performance/Efficiency are the most studied dimensions, Communication and Activity are under-explored, and code quality outcomes are contradictory across contexts. The discussion extends the synthesis with McLuhan's Tetrad and derives recommendations for practitioners and researchers.

Significance. If the corpus-level claims are reliable, this would be a useful reference synthesis for a fast-moving subfield, and the paper would provide a structured basis for future empirical work. The manuscript has clear strengths: the protocol is grounded in established SLR guidelines, the selection process is transparently reported, the quality assessment is described, and the replication package makes the study list and exclusion decisions publicly available. The SPACE-based mapping is a sensible organizing device, and the distinction between benefits and risks is clearly presented. However, the value of the synthesis depends heavily on whether the search strategy actually retrieved the relevant literature, and the mandatory literal term "Productivity" in all database queries creates a potentially serious recall limitation that is not adequately quantified or mitigated.

major comments (3)
  1. [Section 3.1.2 and Table 1] The mandatory literal term "Productivity" in every database query (ACM, ScienceDirect, and Springer explicitly; IEEE, Web of Science, and Scopus via NEAR/5 Productivity) is a recall risk for the corpus-level claims. The 17-paper control set demonstrates only that the final strings retrieve a known set of papers; it does not measure recall across studies that report outcomes such as "efficiency," "throughput," "velocity," "developer experience," "cognitive load," or "code quality" without using the word "productivity" in title, abstract, or keywords. Because the 90% and 15% SPACE percentages and the thematic frequencies are computed over the 39 included studies, one missed study shifts the headline percentages by about 2.5 points, and the under-explored-dimension conclusions are even more sensitive. The snowballing described in Section 3.2 starts from the included seed set and cannot recover a class of studies that the seed set systematically excludes. Please add a sensitivity analysis using broader outcome terms (e.g., efficiency, throughput, developer experience, flow, cognitive load) or a manual audit of recent LLM-assistant evaluation papers that do not use "productivity," and report how the RQ2 and RQ3 distributions would change.
  2. [Table 1, ScienceDirect row] The ScienceDirect query is under-specified: the formal query is shown only as ((Language Model OR "LM" OR "LMs" OR "LLM" OR "LLMs" OR "Artificial Intelligence" OR "AI") AND (Productivity)), and the developer-term restriction is presented as a separate "Advanced Search" line without an explicit Boolean conjunction. It is therefore unclear whether the software-developer terms were applied as an additional conjunct, a separate search, or a filter. Because ScienceDirect contributed 3,734 of the 9,756 raw records, this ambiguity has a material impact on reproducibility. Please provide the exact field-level Boolean expression that was executed, or at minimum confirm that the precise query is stored in the replication package with enough detail for another researcher to reproduce the 3,734 raw results.
  3. [Section 3.2] The initial title/abstract screening of 8,953 records was performed by the first author alone, with the second and last authors validating only excluded records. This leaves the positive inclusion decisions subject to single-screener bias that is not addressed by the reported validation process. Since the inclusion set determines all downstream corpus-level percentages, please report inter-rater agreement on a screened sample, describe a second screening pass, or explain why single-author screening is unlikely to bias the final set. If no re-screening is possible, the threat should be discussed more substantively in Section 9.1 rather than only in terms of conservative inclusion of uncertain records.
minor comments (5)
  1. [Figure 2] The year labels in Figure 2 are garbled in the manuscript text (e.g., "00 00 00 00 00 00 00 11 33 33 3030 22"); please re-render the figure so that the counts per year are legible and consistent with the reported 77% figure for 2024.
  2. [Table 5] The percentages in Table 5 sum to 99% due to rounding; please adjust the individual values or add a rounding note.
  3. [Section 3.2 / Figure 1] The text says that snowballing "expanding the set to 44 primary studies," while the PRISMA figure labels the final included set as 39; please clarify in the text or figure that five studies were subsequently excluded during quality assessment, so the reader is not left with an apparent inconsistency.
  4. [Section 7 and Section 8.3] The statement that "well-being is not examined by any of the empirical studies" is a strong negative claim; consider softening it to "not directly measured" or "not operationalized as a distinct outcome," since some satisfaction-related instruments may touch on well-being indirectly.
  5. [Section 4.2] The summary says "seven authors published two or more papers," which is consistent with the preceding text (six authors with two papers and one with three), but the phrasing could be made more explicit to avoid a perceived numerical mismatch.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the review's claims are aggregations of 39 external primary studies; no fitted parameters or load-bearing self-citation chain is present.

full rationale

This paper is a systematic review and mapping study, not a derivation. Its central claims—the 39-study corpus, the 90% multi-dimensional and 15% four-or-more SPACE percentages, and the thematic benefit/risk frequencies—are counts and thematic syntheses of independent primary studies [PS1]–[PS39]. The organizing frameworks (SPACE [19] and McLuhan's Tetrad [68]) are external conceptual tools, and the SPACE sub-dimension taxonomy is adapted from prior related work [65] with data-derived additions, not from the authors' own prior results. The only self-references are the replication package [20] and the authors' earlier LLM-Cure paper [11], which is cited only as an example of LLM use for debugging/maintenance in the introduction and is not load-bearing for any synthesized finding. The search-string validation against 17 control papers is a recall check, not a fitted input that forces the reported percentages; the paper explicitly acknowledges in Section 9.1 that relevant studies using different terminology could be missed, which is a threat to validity rather than circularity. No equation or definition reduces a predicted or derived quantity to an input of the same argument. The review is self-contained with respect to its own claims, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The review introduces no fitted parameters and no invented constructs. Its claims rest on selection assumptions: that the search protocol captures the relevant literature, that formal peer review guarantees quality, that primary-study self-reports measure productivity, and that the SPACE framework is an appropriate organizing lens. These are standard domain assumptions for a systematic review.

assumptions (4)
  • domain assumption Search recall assumption: the six-database, title/abstract/keyword queries in Table 1, validated against 17 control papers, retrieve all or a representative sample of relevant peer-reviewed studies.
    Section 3.1.2 and Table 1. If the queries miss a substantial slice of relevant studies, the reported themes and percentages lose validity.
  • domain assumption Peer-review and accessibility filters (EC3, EC5) separate quality evidence from grey literature.
    Section 3.1.1. The review treats formal peer review as a proxy for reliability and excludes short, inaccessible, or workshop work, which may omit relevant results.
  • domain assumption Primary-study self-reports and performance metrics are treated as evidence of productivity impact.
    Sections 5 and 6. The synthesis aggregates reported effects without independently verifying the primary studies' measurements.
  • domain assumption The SPACE framework is an appropriate lens for organizing productivity in LLM-assisted development, including the authors' adapted sub-dimensions.
    Section 7. The framework was designed for general developer productivity; applying it to human-AI collaboration requires interpretive decisions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study." pith.science (2026). https://pith.science/paper/2HESWYOD

@misc{pith2026250703156,
  author       = {Pith},
  title        = {Pith review of: The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2HESWYOD}},
  note         = {Machine review of arXiv:2507.03156}
}
read the original abstract

Large language model assistants (LLM-assistants) present new opportunities to transform software development. Developers are increasingly adopting these tools across tasks, including coding, testing, debugging, documentation, and design. Yet, despite growing interest, there is no synthesis of how LLM-assistants affect software developer productivity. In this paper, we present a systematic review and mapping of 39 peer-reviewed studies published between January 2014 and December 2024 that examine this impact. Our analysis reveals that the majority of studies report considerable benefits from LLM-assistants, though a notable subset identifies critical risks. Commonly reported gains include accelerated development, minimized code search, and the automation of trivial and repetitive tasks. However, studies also highlight concerns around cognitive offloading and reduced team collaboration. Our study reveals that whether LLM-based assistants improve or degrade code quality remains unresolved, as existing studies report contradictory outcomes contingent on context and evaluation criteria. While the majority of studies (90%) adopt a multi-dimensional perspective by examining at least two SPACE dimensions, reflecting increased awareness of the complexity of developer productivity, only 15% extend beyond three dimensions, indicating substantial room for more integrated evaluations. Satisfaction, Performance, and Efficiency are the most frequently investigated dimensions, whereas Communication and Activity remain underexplored. Most studies are exploratory (59%) and methodologically diverse, but lack longitudinal and team-based evaluations. This review surfaces key research gaps and provides recommendations for future research and practice. All artifacts associated with this study are publicly available at https://zenodo.org/records/18489222

Figures

Figures reproduced from arXiv: 2507.03156 by the authors.

Figure 1
Figure 1. Overview of the selection process for primary studies included in the review using PRISMA flow chart [ [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Publication frequency per year. studies. The majority of the studies were rated above 3. Detailed results of the quality assessment are available in the replication package material [20]. 3.4 Data Extraction and Synthesis In the final stage of the review process, the synthesis of findings for RQs across these studies were conducted over a period of three months. we followed a qualitative synthesis approach, consiste… view at source ↗
Figure 3
Figure 3. The distribution of empirical procedures across methodological strategies. [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Classification of methodological procedures classification and their overlap. Each bar on the left represents the number of [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Mapping between empirical strategies and instruments in primary studies. [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Radar plots summarizing how frequently each theme appears as benefit (left) and risk (right) across primary studies on [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Distribution of sub-dimensions across each dimension of the [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]
Figure 8
Figure 8. Figure 8: The distribution of investigated SPACE dimensions and their overlap. [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]
Figure 9
Figure 9. Figure 9: McLuhan’s Tetrad diagram illustrates the implications of LLM assistants on the productivity of software developers. The [PITH_FULL_IMAGE:figures/full_fig_p028_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 60 canonical work pages

  1. [1]

    Large language models for software engineering: A systematic literature review

    Xinyi Hou et al. “Large language models for software engineering: A systematic literature review”. In:ACM Transactions on Software Engineering and Methodology33.8 (2024), pp. 1–79

  2. [2]

    Accessed: 2025-06-12

    OpenAI.GPT-4 Technical Report. Accessed: 2025-06-12. 2023.url: https://openai.com/research/gpt-4

  3. [3]

    https://github.com/features/copilot

    GitHub.GitHub Copilot. https://github.com/features/copilot. Accessed: 2025-06-12. 2021

  4. [4]

    A comparative study of code generation using chatgpt 3.5 across 10 programming languages

    Alessio Buscemi. “A comparative study of code generation using chatgpt 3.5 across 10 programming languages”. In:arXiv preprint arXiv:2308.04477(2023)

  5. [5]

    On the effectiveness of large language models in domain-specific code generation

    Xiaodong Gu et al. “On the effectiveness of large language models in domain-specific code generation”. In: ACM Transactions on Software Engineering and Methodology34.3 (2025), pp. 1–22

  6. [6]

    Code prediction by feeding trees to transformers

    Seohyun Kim et al. “Code prediction by feeding trees to transformers”. In:2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE. 2021, pp. 150–162

  7. [7]

    PATCH: Empowering Large Language Model with Programmer-Intent Guidance and Collaborative-Behavior Simulation for Automatic Bug Fixing

    Yuwei Zhang et al. “PATCH: Empowering Large Language Model with Programmer-Intent Guidance and Collaborative-Behavior Simulation for Automatic Bug Fixing”. In:ACM Transactions on Software Engineering and Methodology(2025)

  8. [8]

    Automatic bug fixing via deliberate problem solving with large language models

    Guoyang Weng and Artur Andrzejak. “Automatic bug fixing via deliberate problem solving with large language models”. In:2023 IEEE 34th International Symposium on Software Reliability Engineering Workshops (ISSREW). IEEE. 2023, pp. 34–36

Show all 70 references
  1. [9]

    Debugbench: Evaluating debugging capability of large language models

    Runchu Tian et al. “Debugbench: Evaluating debugging capability of large language models”. In:arXiv preprint arXiv:2401.04621(2024)

  2. [10]

    A comparative analysis of large language models for code documentation generation

    Shubhang Shekhar Dvivedi et al. “A comparative analysis of large language models for code documentation generation”. In:Proceedings of the 1st ACM International Conference on AI-Powered Software. 2024, pp. 65–73

  3. [11]

    LLM-Cure: LLM-based Competitor User Review Analysis for Feature Enhancement

    Maram Assi, Safwat Hassan, and Ying Zou. “LLM-Cure: LLM-based Competitor User Review Analysis for Feature Enhancement”. In:ACM Trans. Softw. Eng. Methodol.(June 2025).issn: 1049-331X.doi: 10.1145/3744644. url: https://doi.org/10.1145/3744644

  4. [12]

    Software testing with large language models: Survey, landscape, and vision

    Junjie Wang et al. “Software testing with large language models: Survey, landscape, and vision”. In:IEEE Transactions on Software Engineering(2024)

  5. [13]

    Large language models for software engineering: Survey and open problems

    Angela Fan et al. “Large language models for software engineering: Survey and open problems”. In:2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE). IEEE. 2023, pp. 31–53

  6. [14]

    Large language models based automatic synthesis of software specifications

    Shantanu Mandal et al. “Large language models based automatic synthesis of software specifications”. In: arXiv preprint arXiv:2304.09181(2023)

  7. [15]

    Chatgpt prompt patterns for improving code quality, refactoring, requirements elicitation, and software design

    Jules White et al. “Chatgpt prompt patterns for improving code quality, refactoring, requirements elicitation, and software design”. In:Generative ai for effective software development. Springer, 2024, pp. 71–108

  8. [16]

    Taking Flight with Copilot: Early insights and opportunities of AI-powered pair- programming tools

    Christian Bird et al. “Taking Flight with Copilot: Early insights and opportunities of AI-powered pair- programming tools”. In:Queue20.6 (2022), pp. 35–57

  9. [17]

    What predicts software developers’ productivity?

    Emerson Murphy-Hill et al. “What predicts software developers’ productivity?” In:IEEE Transactions on Software Engineering47.3 (2019), pp. 582–594

  10. [18]

    Mind the gap: on the relationship between automatically measured and self-reported productivity

    Moritz Beller et al. “Mind the gap: on the relationship between automatically measured and self-reported productivity”. In:IEEE Software38.5 (2020), pp. 24–31. Manuscript submitted to ACM 38 Amr Mohamed, Maram Assi, and Mariam Guizani

  11. [19]

    acmqueue 19 (1), 20–48

    N Forsgren et al.The SPACE of developer productivity: There’s more to it than you think. acmqueue 19 (1), 20–48. 2021

  12. [20]

    Supplemental Material

    Amr Mohamed, Maram Assi, and Mariam Guizani.The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study. Supplemental Material. 2025.doi: 10.5281/zenodo. 18489222.url: https://zenodo.org/records/18489222

  13. [21]

    Pearson Education, 1995

    Frederick P Brooks Jr.The mythical man-month: essays on software engineering. Pearson Education, 1995

  14. [22]

    Addison-Wesley, 2013

    Tom DeMarco and Tim Lister.Peopleware: productive projects and teams. Addison-Wesley, 2013

  15. [23]

    The geography of coordination: Dealing with distance in R&D work

    Rebecca E Grinter, James D Herbsleb, and Dewayne E Perry. “The geography of coordination: Dealing with distance in R&D work”. In:Proceedings of the 1999 ACM International Conference on Supporting Group Work. 1999, pp. 306–315

  16. [24]

    Two case studies of open source software development: Apache and Mozilla

    Audris Mockus, Roy T Fielding, and James D Herbsleb. “Two case studies of open source software development: Apache and Mozilla”. In:ACM Transactions on Software Engineering and Methodology (TOSEM)11.3 (2002), pp. 309–346

  17. [25]

    A systematic review of productivity factors in software development

    Stefan Wagner and Melanie Ruhe. “A systematic review of productivity factors in software development”. In: arXiv preprint arXiv:1801.06475(2018)

  18. [26]

    What happens when software developers are (un) happy

    Daniel Graziotin et al. “What happens when software developers are (un) happy”. In:Journal of Systems and Software140 (2018), pp. 32–47

  19. [27]

    Measuring and predicting software productivity: A systematic map and review

    Kai Petersen. “Measuring and predicting software productivity: A systematic map and review”. In:Information and Software Technology53.4 (2011), pp. 317–343

  20. [28]

    Continuous deployment at Facebook and OANDA

    Tony Savor et al. “Continuous deployment at Facebook and OANDA”. In:Proceedings of the 38th International Conference on software engineering companion. 2016, pp. 21–30

  21. [29]

    Developer fluency: Achieving true mastery in software projects

    Minghui Zhou and Audris Mockus. “Developer fluency: Achieving true mastery in software projects”. In: Proceedings of the eighteenth ACM SIGSOFT international symposium on Foundations of software engineering. 2010, pp. 137–146

  22. [30]

    IT Revolution, 2018

    Nicole Forsgren, Jez Humble, and Gene Kim.Accelerate: The science of lean software and devops: Building and scaling high performing technology organizations. IT Revolution, 2018

  23. [31]

    DevEX: What actually drives productivity?

    Abi Noda et al. “DevEX: What actually drives productivity?” In:Communications of the ACM66.11 (2023), pp. 44–49

  24. [32]

    Developer experience: Concept and definition

    Fabian Fagerholm and Jürgen Münch. “Developer experience: Concept and definition”. In:2012 International Conference on Software and System Process (ICSSP). 2012, pp. 73–77.doi: 10.1109/ICSSP.2012.6225984

  25. [33]

    Changes in perceived productivity of software engineers during COVID-19 pandemic: The voice of evidence

    Darja Smite et al. “Changes in perceived productivity of software engineers during COVID-19 pandemic: The voice of evidence”. In:Journal of Systems and Software186 (2022), p. 111197

  26. [34]

    What improves developer productivity at google? code quality

    Lan Cheng et al. “What improves developer productivity at google? code quality”. In:Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 2022, pp. 1302–1313

  27. [35]

    An exploratory study of productivity perceptions in software teams

    Anastasia Ruvimova et al. “An exploratory study of productivity perceptions in software teams”. In:Proceedings of the 44th International Conference on Software Engineering. 2022, pp. 99–111

  28. [36]

    Improving productivity through corporate hackathons: A multiple case study of two large-scale agile organizations

    Nils Brede Moe et al. “Improving productivity through corporate hackathons: A multiple case study of two large-scale agile organizations”. In:arXiv preprint arXiv:2112.05528(2021). Manuscript submitted to ACM The Impact of LLM-Assistants on Software Developer Productivity: A S...

  29. [37]

    How developers and managers define and trade productivity for quality

    Margaret-Anne Storey, Brian Houck, and Thomas Zimmermann. “How developers and managers define and trade productivity for quality”. In:Proceedings of the 15th International Conference on Cooperative and Human Aspects of Software Engineering. 2022, pp. 26–35

  30. [38]

    Impostor phenomenon in software engineers

    Paloma Guenes et al. “Impostor phenomenon in software engineers”. In:Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Society. 2024, pp. 96–106

  31. [39]

    From forced Working-From-Home to voluntary working-from-anywhere: Two revolutions in telework

    Darja Šmite et al. “From forced Working-From-Home to voluntary working-from-anywhere: Two revolutions in telework”. In:Journal of Systems and Software195 (2023), p. 111509

  32. [40]

    Staffs Keele et al.Guidelines for performing systematic literature reviews in software engineering. Tech. rep. Technical report, ver. 2.3 ebse technical report. ebse, 2007

  33. [41]

    Search strategy formulation for systematic reviews: Issues, challenges and opportunities

    Andrew MacFarlane, Tony Russell-Rose, and Farhad Shokraneh. “Search strategy formulation for systematic reviews: Issues, challenges and opportunities”. In:Intelligent Systems with Applications15 (2022), p. 200091

  34. [42]

    Identifying relevant studies in software engineering

    He Zhang, Muhammad Ali Babar, and Paolo Tell. “Identifying relevant studies in software engineering”. In: Information and Software Technology53.6 (2011), pp. 625–637

  35. [43]

    Systematic literature review of commercial participation in open source software

    Xuetao Li et al. “Systematic literature review of commercial participation in open source software”. In:ACM Transactions on Software Engineering and Methodology34.2 (2025), pp. 1–31

  36. [44]

    A Systematic Literature Review of Multi-Label Learning in Software Engineering

    Joonas Hämäläinen, Teerath Das, and Tommi Mikkonen. “A Systematic Literature Review of Multi-Label Learning in Software Engineering”. In:ACM Transactions on Software Engineering and Methodology(2024)

  37. [45]

    A systematic literature review on federated machine learning: From a software engineering perspective

    Sin Kit Lo et al. “A systematic literature review on federated machine learning: From a software engineering perspective”. In:ACM Computing Surveys (CSUR)54.5 (2021), pp. 1–39

  38. [46]

    A Systematic Literature Review on the Influence of Enhanced Developer Experience on Developers’ Productivity: Factors, Practices, and Recommendations

    Abdul Razzaq et al. “A Systematic Literature Review on the Influence of Enhanced Developer Experience on Developers’ Productivity: Factors, Practices, and Recommendations”. In:ACM Computing Surveys57.1 (2024), pp. 1–46

  39. [47]

    The PRISMA 2020 statement: an updated guideline for reporting systematic reviews

    Matthew J Page et al. “The PRISMA 2020 statement: an updated guideline for reporting systematic reviews”. In:BMJ372 (2021).doi: 10.1136/bmj.n71. eprint: https://www.bmj.com/content/372/bmj.n71.full.pdf.url: https://www.bmj.com/content/372/bmj.n71

  40. [48]

    A systematic literature review on technical debt prioritization: Strategies, processes, factors, and tools

    Valentina Lenarduzzi et al. “A systematic literature review on technical debt prioritization: Strategies, processes, factors, and tools”. In:Journal of Systems and Software171 (2021), p. 110827

  41. [49]

    A systematic literature review of literature reviews in software testing

    Vahid Garousi and Mika V Mäntylä. “A systematic literature review of literature reviews in software testing”. In:Information and Software Technology80 (2016), pp. 195–216

  42. [50]

    The ABC of software engineering research

    Klaas-Jan Stol and Brian Fitzgerald. “The ABC of software engineering research”. In:ACM Transactions on Software Engineering and Methodology (TOSEM)27.3 (2018), pp. 1–51

  43. [51]

    Research in software engineering: an analysis of the literature

    Robert L. Glass, Iris Vessey, and Venkataraman Ramesh. “Research in software engineering: an analysis of the literature”. In:Information and Software technology44.8 (2002), pp. 491–506

  44. [52]

    Criteria for evaluating usability evaluation methods

    H Rex Hartson, Terence S Andre, and Robert C Williges. “Criteria for evaluating usability evaluation methods”. In:International journal of human-computer interaction13.4 (2001), pp. 373–410

  45. [53]

    Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research

    Sandra G Hart and Lowell E Staveland. “Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research”. In:Advances in psychology. Vol. 52. Elsevier, 1988, pp. 139–183

  46. [54]

    Davis’ technology acceptance model (TAM)(1989)

    Patrícia Silva. “Davis’ technology acceptance model (TAM)(1989)”. In:Information seeking behavior and technology adoption: Theories and trends(2015), pp. 205–219

  47. [55]

    Computer self-efficacy: Development of a measure and initial test

    Deborah R Compeau and Christopher A Higgins. “Computer self-efficacy: Development of a measure and initial test”. In:MIS quarterly(1995), pp. 189–211. Manuscript submitted to ACM 40 Amr Mohamed, Maram Assi, and Mariam Guizani

  48. [56]

    Overcoming open source project entry barriers with a portal for newcomers

    Igor Steinmacher et al. “Overcoming open source project entry barriers with a portal for newcomers”. In: Proceedings of the 38th International Conference on Software Engineering. 2016, pp. 273–284

  49. [57]

    A circumplex model of affect

    James A Russell. “A circumplex model of affect.” In:Journal of personality and social psychology39.6 (1980), p. 1161

  50. [58]

    Measuring the cognitive load of software developers: An extended Systematic Mapping Study

    Lucian José Gonçales, Kleinner Farias, and Bruno C da Silva. “Measuring the cognitive load of software developers: An extended Systematic Mapping Study”. In:Information and Software Technology136 (2021), p. 106563

  51. [59]

    The effect of work environments on productivity and satisfaction of software engineers

    Brittany Johnson, Thomas Zimmermann, and Christian Bird. “The effect of work environments on productivity and satisfaction of software engineers”. In:IEEE Transactions on Software Engineering47.4 (2019), pp. 736–757

  52. [60]

    Software metrics: good, bad and missing

    Capers Jones. “Software metrics: good, bad and missing”. In:Computer27.9 (1994), pp. 98–100

  53. [61]

    Analytical and empirical evaluation of software reuse metrics

    Prem Devanbu et al. “Analytical and empirical evaluation of software reuse metrics”. In:Proceedings of IEEE 18th International Conference on Software Engineering. IEEE. 1996, pp. 189–199

  54. [62]

    Grounded copilot: How programmers interact with code-generating models

    Shraddha Barke, Michael B James, and Nadia Polikarpova. “Grounded copilot: How programmers interact with code-generating models”. In:Proceedings of the ACM on Programming Languages7.OOPSLA1 (2023), pp. 85–111

  55. [63]

    Complacency and bias in human use of automation: An attentional integration

    Raja Parasuraman and Dietrich H Manzey. “Complacency and bias in human use of automation: An attentional integration”. In:Human factors52.3 (2010), pp. 381–410

  56. [64]

    https://services.google.com/fh/files/ misc/state-of-devops-2021.pdf

    Google.DevOps Research and Assessment: 2022 State of DevOps Report. https://services.google.com/fh/files/ misc/state-of-devops-2021.pdf. [Online; accessed 12-Oct-2025]. Dec. 2022

  57. [65]

    How much SPACE do metrics have in GenAI assisted software development?

    Samarth Sikand et al. “How much SPACE do metrics have in GenAI assisted software development?” In: Proceedings of the 17th Innovations in Software Engineering Conference. 2024, pp. 1–5

  58. [66]

    From today’s code to tomorrow’s symphony: The AI transformation of developer’s routine by 2030

    Ketai Qiu et al. “From today’s code to tomorrow’s symphony: The AI transformation of developer’s routine by 2030”. In:ACM Transactions on Software Engineering and Methodology34.5 (2025), pp. 1–17

  59. [67]

    Software Engineering by and for Humans in an AI Era

    Silvia Abrahão et al. “Software Engineering by and for Humans in an AI Era”. In:ACM Transactions on Software Engineering and Methodology34.5 (2025), pp. 1–46

  60. [68]

    Laws of the Media

    Marshall McLuhan. “Laws of the Media”. In:ETC: A Review of General Semantics34.2 (1977). Retrieved from JSTOR, pp. 173–179.url: http://www.jstor.org/stable/42575246

  61. [69]

    Why do developers struggle with documentation while excelling at programming

    Daniel Fontanet Losquiño and Tomas Urdell. “Why do developers struggle with documentation while excelling at programming”. B.S. thesis. Universitat Politècnica de Catalunya, 2014

  62. [70]

    Ironies of generative AI: understanding and mitigating productivity loss in Human-AI interaction

    Auste Simkute et al. “Ironies of generative AI: understanding and mitigating productivity loss in Human-AI interaction”. In:International Journal of Human–Computer Interaction41.5 (2025), pp. 2898–2919. Primary Studies [PS1] Frank F Xu, Bogdan Vasilescu, and Graham Neubig. “In...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.