REVIEW 1 major objections 5 minor 39 references
Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research Scientists
T0 review · 1 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that research scientists mostly learn to program on their own, lack formal training in readable code, and rely on comments and AI assistants while struggling with documentation and identifier names.
desk verdict Useful exploratory survey with honest limitations; the LLM 'trend' claim overreaches the cross-sectional data but the artifact and documentation paradox make it worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The instrument is a 20-question survey administered through the Qualtrics platform, combining single-choice, multiple-choice, Likert-scale, and free-text items. Recruitment used convenience and snowball sampling from one university's departments, yielding 57 fully completed responses; three authors independently coded the free-text answers and resolved disagreements through consensus. That survey plus its thematic coding is the entire load-bearing mechanism, since every percentage and ranking in the results comes directly from it.
What would settle it
A probability-sampled survey of several hundred scientists across multiple institutions that found most had formal instruction in maintainable code, or that documentation was not among the top comprehension barriers, would undercut the paper's generalizations.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a profile: research scientists who write code rely on self-study and on-the-job learning, not formal software-engineering education; 33 of 57 respondents (57.9%) never received education or training in writing readable, maintainable code. When they try to understand others' code, the top reported obstacles are lack of comments (44 reports), missing project documentation (33), poor method/variable names (31), poor project structure (31), and unexplained hardcoded values (24). Nearly half (49.12%) say they never use automated code-quality tools, and among those who do, 41.38% use AI/LLM tools, mainly ChatGPT and Claude. All respondents rate code readability as at least slightly important for reproducibility, yet 54.55% say they only sometimes understand others' scientific code, and cryptic or short identifier names are the most common naming problem (40 reports).
Load-bearing premise
The 57 volunteers from one university, recruited by convenience and snowball sampling, represent research scientists broadly enough that the reported percentages reflect the wider community's practices.
Editorial extensions
If this is right
- Universities and research institutions should extend programming education for scientists beyond syntax to cover maintainability, naming, and documentation practices.
- Code comprehension research needs to shift its focus from Java-centric studies to Python and R, the languages scientists actually use.
- New tools for scientific code should target documentation generation and identifier-quality checks, the two most-reported pain points.
- Because many scientists adopt LLMs for code quality without formal software-engineering training, guidance for critically evaluating AI-generated code is needed.
- The gap between documentation's perceived importance and its inadequate delivery, the paper's 'documentation paradox,' warrants direct study of actual documentation behavior.
Reading between the lines
- If the self-selection bias runs in the opposite direction, with scientists who care about readability overrepresented among volunteers, the true level of formal training among all research scientists could be even lower than the reported 57.9%.
- The heavy reliance on LLMs among scientists who do use quality tools suggests a natural experiment: comparing code quality and reproducibility of LLM-assisted versus traditional scientific code could quantify the risk of unchecked AI-generated code.
- The prominence of cryptic and generic names in scientific code may reflect domain-specific shorthand, meaning generic linting tools may not catch names that are meaningful only within a lab and context-aware naming tools may be needed.
- The documentation paradox could be an artifact of self-report: scientists may believe they document more than they actually do, so observational studies of real repositories would test whether the perceived gap is real.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This exploratory survey paper reports on 57 research scientists at the University of Hawai'i at Manoa, recruited through convenience and snowball sampling, to understand their programming backgrounds, code comprehension practices, challenges, and tool usage. The main descriptive findings are that most respondents are self-taught or learned on the job, 57.9% report no formal training in writing readable code, Python and R dominate, comments and documentation are the most common readability practices, inadequate documentation and poor identifier names are the top challenges, traditional code quality tool adoption is low, and LLM-based tools (notably ChatGPT) are the most frequently mentioned tool category among those who use automated tools. The paper includes a Zenodo artifact with the dataset and thematic coding, and it explicitly acknowledges threats to validity from its single-institution convenience sample.
Significance. If taken as a preliminary descriptive study, this paper addresses a genuine gap: most program-comprehension research targets traditional software, not scientific programming. Its value lies in surfacing concrete candidate phenomena—self-taught programmers, reliance on comments, documentation deficits, cryptic names, and low traditional tool adoption—that can motivate follow-up work with larger, more representative samples. The authors provide the survey data and thematic coding in an artifact, which supports reproducibility and independent checking. The main advertised 'trend' claim about LLM adoption is not supported by the cross-sectional design and must be reframed, but the remaining descriptive findings are internally consistent and useful as an early-stage foundation.
major comments (1)
- [Abstract and Section I-B; Section III RQ2; Section IV-A] The abstract and the contribution list state, respectively, that the findings show 'a trend towards utilizing large language models' and 'a trend towards the adoption of AI-based code generation tools.' Section IV-A similarly refers to 'the shift toward LLMs.' However, the survey is cross-sectional: Section III RQ2 only reports current tool use, namely that among the 50.88% of respondents who use automated code quality tools, 41.38% of responses mention LLMs such as ChatGPT or Claude. No survey question asks about past behavior, changes in tool use, or adoption over time, so the data cannot support a temporal trend or shift. This is a load-bearing claim because it appears as one of the paper's three advertised contributions. The 'trend' and 'shift' language should be removed or explicitly reframed as 'current use' or 'current popularity of LLM-based tools among users of automated tools'; the Discussion's softer word 'popular' is appropriate, but the abstract and contributions must be made consistent with what the data actually show.
minor comments (5)
- [Section III, RQ2 (Likert frequency question)] The reported percentages 54.55%, 38.18%, 5.45%, and 1.82% sum to 100% but correspond to a denominator of 55 respondents, not the 57 valid participants stated in Section II-B. Please report the exact counts (30/55, 21/55, 3/55, 1/55) and explain the two missing responses, so readers can verify the percentages.
- [Section III, RQ1 and RQ2 (percentage reporting)] Several passages mix participant counts with percentage bases. For example, '11 participants (12.50%) working in Economics' uses a denominator of total domain responses (88), not the 57 participants, and similar issues appear in the programming-language and environment statistics. Please consistently state whether a percentage is of participants or of total selections, and add this clarification to the table captions for Tables I and II.
- [Tables I and II] The percentage columns in both tables are proportions of total challenge selections, not of participants. Since the text does not state this explicitly, readers may misread the percentages as prevalence among the 57 respondents. Add a note such as 'Percentages are computed over the total number of challenge selections across all participants.'
- [Section II-A] The survey instrument is said to be omitted for space and is available in the Zenodo artifact. Given that the artifact is a central part of the reproducibility story, consider including the full questionnaire as an appendix if the venue permits, or at least quoting the exact wording of the Likert and multiple-choice questions in the paper.
- [Section III, RQ2 (tool-use wording)] The phrase 'a fair number never use automated code quality tools' is vague; the preceding sentence gives 49.12%, so please use the percentage directly to avoid ambiguity.
Circularity Check
No significant circularity: the paper is a self-contained empirical survey with no derivation chain, fitted parameters, or predictions that reduce to the inputs.
full rationale
This paper reports descriptive survey results from 57 research scientists. There is no analytic derivation, no fitted model, no prediction generated from fitted parameters, and no uniqueness theorem or formal result imported from prior work. The central claims (e.g., 57.9% lacking formal instruction, reliance on comments and documentation, top challenges around documentation and identifier naming) are direct summaries of the survey responses reported in Section III. The only self-citations are to prior papers by some of the authors (references [10] and [13]), and these are used for context on identifier structure and readability research, not as load-bearing justification for the present findings. The 'documentation paradox' in Section IV-A is an interpretive observation about two survey results, not a circular argument: the paper does not define 'documentation practice' in terms of the reported documentation challenge, nor does it claim to derive one from the other. The skeptic's concern that the abstract's 'trend towards utilizing large language models' is a temporal inference from a cross-sectional question is a construct-validity or overclaim concern, not a circularity concern, because no survey result is being used as both input and output of a derivation. Under the stated criteria, the paper is fully self-contained against its own empirical data, so no circular step can be exhibited and the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Survey respondents' self-reported programming practices and challenges accurately reflect their actual behavior.
- domain assumption The convenience and snowball sample from a single university is adequate for identifying preliminary trends in scientific code comprehension.
- domain assumption The survey instrument and thematic coding process yield valid and consistent data.
Cite this review
Pith. "Pith review of Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research Scientists." pith.science (2026). https://pith.science/paper/YB3IQRSL
@misc{pith2026250110037,
author = {Pith},
title = {Pith review of: Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research Scientists},
year = {2026},
howpublished = {\url{https://pith.science/paper/YB3IQRSL}},
note = {Machine review of arXiv:2501.10037}
}
read the original abstract
Scientific software-defined as computer programs, scripts, or code used in scientific research, data analysis, modeling, or simulation-has become central to modern research. However, there is limited research on the readability and understandability of scientific code, both of which are vital for effective collaboration and reproducibility in scientific research. This study surveys 57 research scientists from various disciplines to explore their programming backgrounds, practices, and the challenges they face regarding code readability. Our findings reveal that most participants learn programming through self-study or on the-job training, with 57.9% lacking formal instruction in writing readable code. Scientists mainly use Python and R, relying on comments and documentation for readability. While most consider code readability essential for scientific reproducibility, they often face issues with inadequate documentation and poor naming conventions, with challenges including cryptic names and inconsistent conventions. Our findings also show low adoption of code quality tools and a trend towards utilizing large language models to improve code quality. These findings offer practical insights into enhancing coding practices and supporting sustainable development in scientific software.
Reference graph
Works this paper leans on
-
[1]
Software engineering practices for scientific software de velopment: A systematic mapping study,
E.-M. Arvanitou, A. Ampatzoglou, A. Chatzigeorgiou, an d J. C. Carver, “Software engineering practices for scientific software de velopment: A systematic mapping study,” Journal of Systems and Software , vol. 172, p. 110848, 2021. 1
work page 2021
-
[2]
L. V aughan, Python Tools for Scientists: An Introduction to Using Anaconda, JupyterLab, and Python’s Scientific Libraries . No Starch Press, 2023. 1
work page 2023
-
[3]
How do scientists develop and use scientific soft ware?,
J. E. Hannay, C. MacLeod, J. Singer, H. P . Langtangen, D. P fahl, and G. Wilson, “How do scientists develop and use scientific soft ware?,” in 2009 ICSE W orkshop on Software Engineering for Computation al Science and Engineering , pp. 1–8, 2009. 1, 4
work page 2009
-
[4]
A survey of the practice of computational s cience,
P . Prabhu, T. B. Jablin, A. Raman, Y . Zhang, J. Huang, H. Ki m, N. P . Johnson, F. Liu, S. Ghosh, S. Beard, T. Oh, M. Zoufaly, D. Walk er, and D. I. August, “A survey of the practice of computational s cience,” in State of the Practice Reports , SC ’11, (New Y ork, NY , USA), Association for Computing Machinery, 2011. 1, 4
work page 2011
-
[5]
Program comprehension du ring software maintenance and evolution,
A. V on Mayrhauser and A. V ans, “Program comprehension du ring software maintenance and evolution,” Computer, vol. 28, no. 8, 1995. 1
work page 1995
-
[6]
Meas uring Pro- gram Comprehension: A Large-Scale Field Study with Profess ionals,
X. Xia, L. Bao, D. Lo, Z. Xing, A. E. Hassan, and S. Li, “Meas uring Pro- gram Comprehension: A Large-Scale Field Study with Profess ionals,” IEEE Transactions on Software Engineering , vol. 44, no. 10, pp. 951– 976, 2018. 1
work page 2018
-
[7]
Reassessing Java Code Readability Models w ith a Human-Centered Approach,
A. Sergeyuk, O. Lvova, S. Titov, A. Serova, F. Bagirov, E. Kirillova, and T. Bryksin, “Reassessing Java Code Readability Models w ith a Human-Centered Approach,” in Proceedings of the 32nd IEEE/ACM International Conference on Program Comprehension , ICPC ’24, (New Y ork, NY , USA), p. 225–235, Association for Computing Machi nery,
-
[8]
Linguistic antipatterns: what they are and how developers perceive them,
V . Arnaoudova, M. Di Penta, and G. Antoniol, “Linguistic antipatterns: what they are and how developers perceive them,” Empirical Software Engineering, vol. 21, pp. 104–158, Feb 2016. 1, 4
work page 2016
Show all 39 references
-
[9]
A Survey on Renamings of S oftware Entities,
G. Li, H. Liu, and A. S. Nyamawe, “A Survey on Renamings of S oftware Entities,” ACM Comput. Surv. , vol. 53, Apr. 2020. 1, 4
2020
-
[10]
On the Generation, Structure, an d Semantics of Grammar Patterns in Source Code Identifiers,
C. D. Newman, R. S. Alsuhaibani, M. J. Decker, A. Peruma, D. Kaushik, M. W. Mkaouer, and E. Hill, “On the Generation, Structure, an d Semantics of Grammar Patterns in Source Code Identifiers,” CoRR, vol. abs/2007.08033, 2020. 1
2007 arXiv
-
[11]
What’s in a Name? A Study of Identifiers,
D. Lawrie, C. Morrell, H. Feild, and D. Binkley, “What’s in a Name? A Study of Identifiers,” in 14th IEEE International Conference on Program Comprehension (ICPC’06) , pp. 3–12, 2006. 1
2006
-
[12]
A Systematic Literature Review on Empirical A nalysis of the Relationship Between Code Smells and Software Quality Attr ibutes,
A. Kaur, “A Systematic Literature Review on Empirical A nalysis of the Relationship Between Code Smells and Software Quality Attr ibutes,” Archives of Computational Methods in Engineering , vol. 27, pp. 1267– 1296, Sep 2020. 1
2020
-
[13]
An eye tracking study assessing source code r eadability rules for program comprehension,
K. Park, J. Johnson, C. S. Peterson, N. Y edla, I. Baysing er, J. Aponte, and B. Sharif, “An eye tracking study assessing source code r eadability rules for program comprehension,” Empir . Softw. Eng., vol. 29, no. 6, p. 160, 2024. 1
2024
-
[14]
40 Y ears of Designi ng Code Comprehension Experiments: A Systematic Mapping Study,
M. Wyrich, J. Bogner, and S. Wagner, “40 Y ears of Designi ng Code Comprehension Experiments: A Systematic Mapping Study,” ACM Comput. Surv., vol. 56, Nov. 2023. 1, 4
2023
-
[15]
Exploration and Ex planation in Computational Notebooks,
A. Rule, A. Tabard, and J. D. Hollan, “Exploration and Ex planation in Computational Notebooks,” in Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems , CHI ’18, (New Y ork, NY , USA), p. 1–12, Association for Computing Machinery, 2018. 1 , 4
2018
-
[16]
Scientists and software engineers: A tale of two cultures,
J. Segal, “Scientists and software engineers: A tale of two cultures,” in PPIG 2008: Proceedings of the 20th Annual Meeting of the Psyc hology of Programming Interest Group , Lancaster University, 2008. 1
2008
-
[17]
Manag- ing Messes in Computational Notebooks,
A. Head, F. Hohman, T. Barik, S. M. Drucker, and R. DeLine , “Manag- ing Messes in Computational Notebooks,” in Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , CHI ’19, (New Y ork, NY , USA), p. 1–12, Association for Computing Mach inery,
2019
-
[18]
A Comparison of Mac hine Learning Code Quality in Python Scripts and Jupyter Noteboo ks,
K. Adams, A. Vilkomir, and M. Hills, “A Comparison of Mac hine Learning Code Quality in Python Scripts and Jupyter Noteboo ks,” J. Comput. Sci. Coll. , vol. 39, p. 96–108, Nov. 2023. 1
2023
-
[19]
Understanding and improving the quality and reproducibility of Jupyter no tebooks,
J. F. Pimentel, L. Murta, V . Braganholo, and J. Freire, “ Understanding and improving the quality and reproducibility of Jupyter no tebooks,” Empirical Software Engineering , vol. 26, p. 65, May 2021. 1, 4
2021
-
[20]
What’s Wrong with Computational Notebooks? Pain Points, N eeds, and Design Opportunities,
S. Chattopadhyay, I. Prasad, A. Z. Henley, A. Sarma, and T. Barik, “What’s Wrong with Computational Notebooks? Pain Points, N eeds, and Design Opportunities,” in Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , CHI ’20, (New Y ork, NY , USA), p. 1...
2020
-
[21]
Better code, better shari ng: on the need of analyzing jupyter notebooks,
J. Wang, L. Li, and A. Zeller, “Better code, better shari ng: on the need of analyzing jupyter notebooks,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering: Ne w Ideas and Emerging Results , ICSE-NIER ’20, (New Y ork, NY , USA), p. 53–56, As...
2020
-
[22]
A large- scale comparison of Python code in Jupyter notebooks and scr ipts,
K. Grotov, S. Titov, V . Sotnikov, Y . Golubev, and T. Bryk sin, “A large- scale comparison of Python code in Jupyter notebooks and scr ipts,” in Proceedings of the 19th International Conference on Mining Software Repositories, MSR ’22, (New Y ork, NY , USA), p. 353–364, Assoc...
2022
-
[23]
A Large-Scale Study About Quality and Reproducibility of Jupyter Noteboo ks,
J. F. Pimentel, L. Murta, V . Braganholo, and J. Freire, “ A Large-Scale Study About Quality and Reproducibility of Jupyter Noteboo ks,” in 2019 IEEE/ACM 16th International Conference on Mining Soft ware Repositories (MSR) , pp. 507–517, 2019. 1
2019
-
[24]
Qualtrics XM: The Leading Experience Management Soft ware
“Qualtrics XM: The Leading Experience Management Soft ware.” https://www.qualtrics.com/. 2
-
[25]
Guidelines for conducting surveys in software engineering,
J. Lin˚ aker, S. M. Sulaman, R. M. de Mello, and M. H¨ ost, “ Guidelines for conducting surveys in software engineering,” 2015. 2
2015
- [26]
-
[27]
Sampling in software engineeri ng research: a critical review and guidelines,
S. Baltes and P . Ralph, “Sampling in software engineeri ng research: a critical review and guidelines,” Empirical Software Engineering, vol. 27, p. 94, Apr 2022. 2
2022
-
[28]
Wagner, D
S. Wagner, D. Mendez, M. Felderer, D. Graziotin, and M. K alinowski, Challenges in Survey Research , pp. 93–125. Cham: Springer Interna- tional Publishing, 2020. 2
2020
-
[29]
Descriptive compound identifier names improve so urce code comprehension,
A. Schankin, A. Berger, D. V . Holt, J. C. Hofmeister, T. R iedel, and M. Beigl, “Descriptive compound identifier names improve so urce code comprehension,” in Proceedings of the 26th Conference on Program Comprehension, ICPC ’18, (New Y ork, NY , USA), p. 31–40, Association fo...
2018
-
[30]
On t he Role of Computer Languages in Scientific Computing,
D. Leroy, J. Sallou, J. Bourcier, and B. Combemale, “On t he Role of Computer Languages in Scientific Computing,” Computing in Science & Engineering , vol. 24, no. 4, pp. 55–59, 2022. 4
2022
-
[31]
Bug Analysis in Jupyter Notebook Projects: An Empirical St udy,
T. L. De Santana, P . A. D. M. S. Neto, E. S. De Almeida, and I . Ahmed, “Bug Analysis in Jupyter Notebook Projects: An Empirical St udy,” ACM Trans. Softw. Eng. Methodol. , vol. 33, Apr. 2024. 4
2024
-
[32]
Do Code Smells Impact the Effort of Different Maintenance Programm ing Ac- tivities?,
Z. Soh, A. Y amashita, F. Khomh, and Y .-G. Gu´ eh´ eneuc, “ Do Code Smells Impact the Effort of Different Maintenance Programm ing Ac- tivities?,” in 2016 IEEE 23rd International Conference on Software Analysis, Evolution, and Reengineering (SANER) , 2016. 4
2016
-
[33]
The e ffect of poor source code lexicon and readability on developers’ cogniti ve load,
S. Fakhoury, Y . Ma, V . Arnaoudova, and O. Adesope, “The e ffect of poor source code lexicon and readability on developers’ cogniti ve load,” in Proceedings of the 26th Conference on Program Comprehensio n, ICPC ’18, (New Y ork, NY , USA), p. 286–296, Association for Comput i...
2018
-
[34]
Ten simple rules for writing and sharing computational ana lyses in Jupyter Notebooks,
A. Rule, A. Birmingham, C. Zuniga, I. Altintas, S.-C. Hu ang, R. Knight, N. Moshiri, M. H. Nguyen, S. B. Rosenthal, F. P´ erez, and P . W. Rose, “Ten simple rules for writing and sharing computational ana lyses in Jupyter Notebooks,” PLOS Computational Biology , 07 2019. 4
2019
-
[35]
Future of software development with generative AI,
J. Sauvola, S. Tarkoma, M. Klemettinen, J. Riekki, and D . Doermann, “Future of software development with generative AI,” Automated Soft- ware Engineering, vol. 31, Mar 2024. 4
2024
-
[36]
A Large-Scale Surv ey on the Usability of AI Programming Assistants: Successes and Chal lenges,
J. T. Liang, C. Y ang, and B. A. Myers, “A Large-Scale Surv ey on the Usability of AI Programming Assistants: Successes and Chal lenges,” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , ICSE ’24, (New Y ork, NY , USA), Association for Com...
2024
-
[37]
Assessing the Readabi lity of ChatGPT Code Snippet Recommendations: A Comparative Study ,
C. Dantas, A. Rocha, and M. Maia, “Assessing the Readabi lity of ChatGPT Code Snippet Recommendations: A Comparative Study ,” in Proceedings of the XXXVII Brazilian Symposium on Software E ngineer- ing, SBES ’23, (New Y ork, NY , USA), p. 283–292, Association for Computing Mac...
2023
-
[38]
Q uality Assessment of ChatGPT Generated Code and their Use by Develo pers,
M. L. Siddiq, L. Roney, J. Zhang, and J. C. D. S. Santos, “Q uality Assessment of ChatGPT Generated Code and their Use by Develo pers,” in Proceedings of the 21st International Conference on Mining Software Repositories, MSR ’24, (New Y ork, NY , USA), p. 152–156, Association ...
2024
-
[39]
Refining ChatGPT-Generated Code: Characteri zing and Mitigating Code Quality Issues,
Y . Liu, T. Le-Cong, R. Widyasari, C. Tantithamthavorn, L. Li, X.-B. D. Le, and D. Lo, “Refining ChatGPT-Generated Code: Characteri zing and Mitigating Code Quality Issues,” ACM Trans. Softw. Eng. Methodol. , vol. 33, June 2024. 4
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.