REVIEW 2 major objections 6 minor 116 references
Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software
T0 review · 2 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Scientific open-source software projects are abandoned at an 8% lower rate than matched non-scientific open-source projects, according to survival models on a curated set of 18,247 repositories.
desk verdict The RQ3 matching concern that sank the reader's verdict is explicitly answered in Section 4.1; this is a solid empirical paper that deserves serious peer review, not a desk reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a two-stage pipeline that classifies repositories from README content, combined with survival models on the resulting labeled set. An initial large-language-model pass separates scientific from non-scientific software and detects mentions of publications or funding; a second pass assigns each scientific repository to one of three software-stack layers — scientific infrastructure, domain-specific code, or publication-specific code — and one of 13 STEM fields. Longevity is measured as time from earliest to latest commit, with abandonment defined as six months without commits, and the paper uses Kaplan-Meier curves and Cox proportional-hazards regression to estimate which factors raise or lower abandonment risk. The decisive comparison for the headline result is the binary 'is scientific software' indicator estimated on scientific repositories matched to non-scientific repositories by number of commits, number of authors, and earliest commit year.
What would settle it
Re-run the RQ3 comparison after passing the non-scientific sample through the exact full eligibility filter used for the scientific catalog (at least 10 files, more than 300 commits, at least 3 authors, more than 6 active months, and a last commit after November 2018) before matching; if the 0.92 hazard ratio for 'is scientific software' moves to or above 1, the headline result is an artifact of how the non-scientific projects were sampled.
Extended reading notes
Core claim
The paper's discovery is that scientific open-source software, far from being especially fragile, tends to be more durable than otherwise comparable non-scientific open-source software. Using a matched sample of 36,494 non-scientific repositories selected to resemble the 18,247 scientific ones in commit count, author count, and era, the authors fit a Cox proportional-hazards model in which the indicator 'is scientific software' has a hazard ratio of 0.92 (p < 0.001), corresponding to an 8% lower risk of abandonment. The same model shows scientific infrastructure outlasting domain-specific code, which in turn outlasts publication-specific code; projects with more downstream dependents, explicit mentions of publications or funding, and government participants also survive longer, whereas newer projects and projects with academic participants face higher abandonment risk. The authors interpret this as evidence that science-specific incentives such as publication credit and funding support, while short-term, give scientific projects a measurable durability advantage over generic open-source software.
Load-bearing premise
The comparison assumes that the non-scientific sample was selected under the same rules as the scientific one — in particular the requirement of a recent commit after November 2018 and the same minimum size and activity filters — so that the survival gap reflects genuine longevity rather than different selection windows.
Editorial extensions
If this is right
- The common assumption that scientific software is unusually prone to abandonment should be qualified: for large, collaboratively developed projects, scientific software appears to be the sturdier category.
- Funding agencies concerned with sustainability can allocate attention by layer, since scientific infrastructure shows the lowest abandonment hazard and publication-specific code the highest.
- Measurable project attributes — downstream dependents, documentation that mentions publications or funding, and government participation — are associated with longer lifespans and can serve as early indicators of durability.
- The catalog of over 18,000 labeled repositories gives future studies a population from which to sample projects for qualitative or longitudinal work on sustainability.
Reading between the lines
- If the longevity advantage is driven by external anchors such as grants and publication credit, then non-scientific OSS could borrow the practice of tying development to documented outcomes or institutional sponsorship — a transfer the paper mentions only as a possibility.
- Because the classifier depends on READMEs, scientific software with sparse or outdated documentation may be underrepresented; applying the same method to code content or paper-mention mining could shift the estimated 8% advantage.
- The matched comparison holds commit count, author count, and era constant but not popularity directly; a replication that also matches on stars or forks would test whether the scientific advantage survives once visibility is equalized.
- The average scientific advantage hides real variation — computer science and data science projects show shorter lifespans than astronomy or mathematics — so policy responses aimed at the whole class may miss the projects most at risk.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a two-part empirical study of scientific open-source software longevity. The authors build SciCat, a dataset of 18,247 repositories, by filtering World of Code with size/activity heuristics and then using a two-stage LLM pipeline (GPT-3.5 prescreening, GPT-4 classification) on READMEs to assign each repository a STEM field and a layer in Hinsen's software stack; the classification is validated with stratified human annotation and cross-checked against five external collections. Using SciCat, the paper estimates Kaplan-Meier survival curves and Cox proportional-hazards models to identify correlates of longevity (RQ2), and compares SciCat to a matched sample of 36,494 non-scientific repositories (RQ3). The headline findings are that infrastructure-layer position, downstream dependents, publication/funding mentions, and government participation are associated with lower abandonment hazard, that academic participation and recent start dates are associated with higher hazard, and that scientific software has an 8% lower abandonment hazard than matched non-scientific software.
Significance. The paper's dataset and classification methodology are solid and useful: the authors provide a replication package, report inter-rater reliability and per-class precision/recall, and transparently discuss classification errors. The RQ2 survival models are standard, with multiple operationalization checks, and the direction of effects is consistent with prior OSS ecosystem work. If the RQ3 comparison were valid, the finding that scientific software outlives comparable non-scientific software would be a notable and policy-relevant result. However, the current evidence for RQ3 does not support that conclusion because the non-scientific comparison sample is not described as being subject to the same eligibility filters as SciCat; the central claim therefore requires a substantial re-analysis. The dataset and RQ2 contributions remain valuable independently.
major comments (2)
- [Section 6, RQ3 Methods; Section 4.1 sampling] The paper's headline RQ3 result (HR = 0.92 for Is Scientific Software, Figure 7) is not supported by the sampling procedure as described. SciCat was restricted to repositories with at least 10 files, more than 300 commits, at least 3 authors, more than 6 consecutive active months, and a last commit after November 2018. The non-scientific sample is described only as matched on bins of commit count, author count, and earliest commit year, with the first commit bin being '<750 commits' and the first author bin being '<10 authors'. As written, the non-scientific sample can therefore include projects with very few commits or authors and, crucially, projects whose last commit predates the observation window. Because SciCat is selected on surviving to late 2018 while the non-scientific sample is not, the survival comparison confounds longevity with sample-selection criteria. The authors' assertion in Section 4.1 that 'we use the same filtering criteria' for the matched sample is not reflected in the RQ3 methods, and Section 6's own requirement that sampling be 'noninformative or ignorable' is not met. The star-count robustness check in Section 7 does not address this temporal-selection concern. The authors should either apply the identical eligibility filters, including the recent-commit filter, to the non-scientific sample and re-run the analysis, or adopt a left-truncation/delayed-entry Cox model and explain why the current comparison is valid.
- [Section 5.1, Figure 5; Section 5.2] The Kaplan-Meier curves and restricted mean survival times (e.g., 6.44 years over 15 years) are estimated from time of earliest commit for a sample in which every project had a last commit after November 2018. This is length-biased sampling: repositories that died before November 2018 are excluded by construction, so the curves and RMST describe only projects that survived to that date, and they overstate typical longevity. The same conditioning affects the Cox models in RQ2, since they use the same sample without accounting for left truncation. Please report estimates with delayed entry (entry at the later of project start and November 2018) or explicitly reframe all RQ2 descriptive results as conditional on survival to the sample-entry date.
minor comments (6)
- [Section 6, Results] The sentence 'which is has negative HR = 0.92' is ungrammatical; it should read 'which has an HR below 1'.
- [Section 5, Methods] The censoring scheme should be described more precisely: a project with no commit in the last six months is treated as an event at its last commit, while an active project is censored; this is interval-censoring rather than pure right-censoring, and the exact censoring time used in the Cox model should be stated.
- [Section 4.2] What is called 'recall' in the cross-validation against external datasets is better described as overlap after filtering; consider renaming to avoid implying a well-defined ground-truth recall.
- [Figure 7] The forest plot does not show confidence intervals; include them in the figure or a companion table to support the reported significance levels.
- [Section 4.1] The November 2018 last-commit filter is motivated as avoiding 'less relevant repositories,' but this is precisely the condition that induces length-biased sampling; please discuss this explicitly in Threats to Validity.
- [Section 8] The word 'quantiative' should be 'quantitative'.
Circularity Check
No circularity found: labels and outcomes are independent, and the RQ3 comparison uses matched samples with the same filtering criteria.
full rationale
The paper's central quantities are measured independently. The outcome (survival time from earliest to latest commit, with a six-month inactivity window defining abandonment) comes from World of Code commit histories, while the scientific/non-scientific label comes from LLM classification of README contents, validated against manual raters and external datasets. The Cox models estimate hazard ratios from these variables; no coefficient is defined in terms of the target result, and no fitted parameter is renamed as a prediction. The RQ3 comparison between scientific and non-scientific projects is not circular: Section 4.1 explicitly states that the same filtering criteria were used to identify the matching non-scientific repositories, and the stratified matching on commit count, author count, and earliest commit year is described in Section 6. Although the comparison may be vulnerable to statistical threats to validity (e.g., left-truncation or selection-window effects), these are methodological concerns, not circularity. The authors' self-citations (e.g., World of Code infrastructure, author identity resolution, prior survival analyses of OSS) provide data infrastructure or methodological precedent, and they are not invoked as uniqueness theorems or as substitute evidence for the paper's longevity findings. External benchmarks (JOSS, Papers with Code, NSF Soft-Search, Research Software Directory) provide independent validation of the curated dataset. Thus, no load-bearing step reduces by construction to the paper's own inputs.
Assumptions & free parameters
free parameters (2)
- Abandonment window =
6 months
- RQ3 matching bin boundaries =
commits: 750/1800/5000; authors: 10/25/60; earliest year: 2016/2019
assumptions (5)
- domain assumption World of Code commit and dependency data are sufficiently accurate for measuring project activity and lifespan.
- domain assumption LLM labels of field, layer, and paper/funding mentions are accurate enough for use as covariates.
- domain assumption Hinsen's scientific software stack layers map onto meaningful longevity classes.
- ad hoc to paper The RQ3 matched sample is ignorably selected given commits, authors, and earliest commit year.
- domain assumption Abandonment can be operationalized as a six-month gap without commits.
Cite this review
Pith. "Pith review of Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software." pith.science (2026). https://pith.science/paper/MHOED423
@misc{pith2026250418971,
author = {Pith},
title = {Pith review of: Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software},
year = {2026},
howpublished = {\url{https://pith.science/paper/MHOED423}},
note = {Machine review of arXiv:2504.18971}
}
read the original abstract
Scientific software is essential to scientific innovation and in many ways it is distinct from other types of software. Abandoned (or unmaintained), buggy, and hard to use software, a perception often associated with scientific software can hinder scientific progress, yet, in contrast to other types of software, its longevity is poorly understood. Existing data curation efforts are fragmented by science domain and/or are small in scale and lack key attributes. We use large language models to classify public software repositories in World of Code into distinct scientific domains and layers of the software stack, curating a large and diverse collection of over 18,000 scientific software projects. Using this data, we estimate survival models to understand how the domain, infrastructural layer, and other attributes of scientific software affect its longevity. We further obtain a matched sample of non-scientific software repositories and investigate the differences. We find that infrastructural layers, downstream dependencies, mentions of publications, and participants from government are associated with a longer lifespan, while newer projects with participants from academia had shorter lifespan. Against common expectations, scientific projects have a longer lifetime than matched non-scientific open-source software projects. We expect our curated attribute-rich collection to support future research on scientific software and provide insights that may help extend longevity of both scientific and other projects.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Carl S Adorf, Vyas Ramasubramani, Joshua A Anderson, and Sharon C Glotzer. 2018. How to professionally develop reusable scientific software—and when not to. Comput. Sci. Eng. 21, 2 (2018), 66–79
2018
-
[2]
Adem Ait, Javier Luis Cánovas Izquierdo, and Jordi Cabot. 2022. An Empirical Study on the Survival Rate of GitHub Projects. In Int’l Conf. MSR (MSR) . 365–375
2022
-
[3]
Carina Alves, Joyce Oliveira, and Slinger Jansen. 2018. Understanding governance mechanisms and health in software ecosystems: A systematic literature review. In Int’l Conf. Enterprise Inf. Systems (ICEIS) . Springer, 517–542
2018
-
[4]
Taylor Arnold and Lauren Tilton. 2024. Humanities Data in R: Exploring Networks, Geospatial Data, Images, and Text (2nd ed.). Springer
2024
-
[5]
SoS: BIO: Evaluating the Impact of Biomedical Tools and Methods
Albert-Laszlo Barabasi. 2024. NSF Award “SoS: BIO: Evaluating the Impact of Biomedical Tools and Methods”. https://reporter.nih.gov/project-details/11120817
-
[6]
Michelle Barker and Daniel Katz. 2022. Overview of research software funding landscape. Zenodo (2022)
2022
-
[7]
Rob Baxter, N Chue Hong, Dirk Gorissen, James Hetherington, and Ilian Todorov. 2012. The research software engineer. In Digital Research Conf., Oxford . 1–3
2012
-
[8]
Christoph Becker, Ruzanna Chitchyan, Leticia Duboc, Steve Easterbrook, Birgit Penzenstadler, Norbert Seyff, and Colin C Venters. 2015. Sustainability design and software: The Karlskrona manifesto. In Int’l Conf. Software Eng. , Vol. 2. IEEE, 467–476
2015
Show all 116 references
-
[9]
Chris Bogart, Christian Kästner, James Herbsleb, and Ferdian Thung. 2021. When and how to make breaking changes: Policies and practices in 18 open source software ecosystems. ACM Trans. Softw. Eng. Methodol. 30, 4 (2021), 1–56
2021
-
[10]
Massimiliano Bonomi. 2019. Promoting transparency and reproducibility in enhanced molecular simulations. Nature Methods 16, 8 (2019), 670–673
2019
-
[11]
Hudson Borges and Marco Tulio Valente. 2018. What’s in a GitHub star? Understanding repository starring practices in a social coding platform. J. Syst. Softw. 146 (2018), 112–129
2018
-
[12]
Barry Bozeman. 2000. Technology transfer and public policy: a review of research and theory. Research Policy 29, 4 (2000), 627–655
2000
-
[13]
Alys Brett, Michael Croucher, Robert Haines, Simon Hettrick, James Hetherington, Mark Stillwell, and Claire Wyatt
-
[14]
Eva Maxfield Brown, Lindsey Schwartz, Richard Lewei Huang, and Nicholas Weber. 2023. Soft-Search: Two Datasets to Study the Identification and Production of Research Software. InJoint Conf. Digital Libraries (JCDL). IEEE, 228–231
2023
-
[15]
Fabio Calefato, Marco Aurelio Gerosa, Giuseppe Iaffaldano, Filippo Lanubile, and Igor Steinmacher. 2022. Will you come back to contribute? Investigating the inactivity of OSS core developers in GitHub.Empir. Softw. Eng. 27, 3 (2022), 76
2022
-
[16]
Juan Andrés Carruthers, Jorge Andrés Diaz-Pace, and Emanuel Agustín Irrazábal. 2022. How are software datasets constructed in Empirical Software Engineering studies? A systematic mapping study. In Euromicro Conf. Software Eng. and Advanced Applications (SEAA) . IEEE, 442–450
2022
-
[17]
Vinton G Cerf. 2024. The Boundary Hunters. Commun. ACM 67, 8 (2024), 5–5
2024
-
[18]
Kaylea Champion and Benjamin Mako Hill. 2021. Underproduction: An approach for measuring risk in open source software. In Int’l Conf. Software Analysis, Evolution and Reengineering (SANER) . IEEE, 388–399
2021
-
[19]
Adam S Charles, Benjamin Falk, Nicholas Turner, Talmo D Pereira, Daniel Tward, Benjamin D Pedigo, Jaewon Chung, Randal Burns, Satrajit S Ghosh, Justus M Kebschull, et al. 2020. Toward Community-Driven big Open brain science: open big data and tools for structure, function, and...
2020
-
[20]
InduShobha Chengalur-Smith, Anna Sidorova, and Sherae L Daniel. 2010. Sustainability of free/libre open source projects: A longitudinal study. J. Assoc. Inf. Syst. 11, 11 (2010), 5
2010
-
[21]
Jailton Coelho and Marco Tulio Valente. 2017. Why modern open source projects fail. In Int’l Conf. Foundations of Software Eng. (FSE). ACM, 186–196
2017
-
[22]
Filipe R Cogo, Gustavo A Oliva, and Ahmed E Hassan. 2021. Deprecation of packages and releases in software ecosystems: A case study on npm. IEEE Trans. Softw. Eng. 48, 7 (2021), 2208–2223
2021
-
[23]
Jeremy Cohen, Daniel S Katz, Michelle Barker, Neil Chue Hong, Robert Haines, and Caroline Jay. 2020. The four pillars of research software engineering. IEEE Software 38, 1 (2020), 97–105
2020
-
[24]
Colazo and Yulin Fang
Jorge A. Colazo and Yulin Fang. 2010. Following the sun: Temporal dispersion and performance in open-source software project teams. J. Assoc. Inf. Syst. 11, 11 (2010), 684–707
2010
-
[25]
D. R. Cox. 1972. Regression models and life-tables. J. R. Stat. Soc. 34, 2 (1972), 187–220
1972
-
[26]
Alexandre Decan, Tom Mens, and Philippe Grosjean. 2019. An empirical comparison of dependency network evolution in seven software packaging ecosystems. Empir. Softw. Eng. 24, 1 (2019), 381–416
2019
-
[27]
Tapajit Dey, Yuxing Ma, and Audris Mockus. 2019. Are Software Dependency Supply Chain Metrics Useful in Predicting Change of Popularity of NPM Packages?. In Int’l Conf. Predictive Models and Data Analytics in Software Eng. ACM, 1–10. , Vol. 1, No. 1, Article . Publication date...
2019
-
[28]
Anshu Dubey, Jared O’Neal, Klaus Weide, and Saurabh Chawdhary. 2020. Distillation of best practices from refactoring flash for exascale. SN Computer Science 1, 4 (2020), 1–9
2020
-
[29]
George Dyson. 2012. Turing’s cathedral: the origins of the digital universe . Vintage
2012
-
[30]
Claudia Eitzen. 2020. Research Software-Publication and Sustainability . Ph. D. Dissertation. Kiel University
2020
-
[31]
This is damn slick!
Hongbo Fang, Hemank Lamba, James Herbsleb, and Bogdan Vasilescu. 2022. “This is damn slick!” Estimating the impact of tweets on open source project popularity and new contributors. In Int’l Conf. Software Engineering (ICSE) . 2116–2129
2022
-
[32]
Stuart Faulk, Eugene Loh, Michael L Van De Vanter, Susan Squires, and Lawrence G Votta. 2009. Scientific computing’s productivity gridlock: How software engineering can help. Comput. Sci. Eng. 11, 6 (2009), 30–39
2009
-
[33]
Daniela Feitosa, Christina von Flach, and Joenio Costa. 2023. Understanding practices and challenges of developing sustainable research software: A pilot interview. In OpenScience Worksh. Universidade Federal da Bahia
2023
-
[34]
Brian Fitzgerald. 2006. The transformation of open-source software. MIS Quarterly 30, 3 (2006), 587–598
2006
-
[35]
Joseph L Fleiss, Bruce Levin, Myunghee Cho Paik, et al. 1981. The measurement of interrater agreement. Statistical methods for rates and proportions 2, 212-236 (1981), 22–23
1981
-
[36]
Anne Fouilloux, Jean Iaquinta, Alok Kumar Gupta, Hamish Struthers, Oskar Landgren, Prashanth Dwarakanath, Tommi Bergman, and Yanchun He. 2023. Building on Communities to Further Software Sustainability. Comput. Sci. Eng. 25, 3 (2023), 84–88
2023
-
[37]
Tanner Fry, Tapajit Dey, Andrey Karnauch, and Audris Mockus. 2020. A Dataset and an Approach for Identity Resolution of 38 Million Author IDs extracted from 2B Git Commits. In Int’l Conf. MSR (MSR)
2020
-
[38]
Jonas Gamalielsson and Björn Lundell. 2014. Sustainability of Open Source software communities beyond a fork: How and why has the LibreOffice project evolved? J. Syst. Softw. 89 (2014), 128–145
2014
-
[39]
Daniel Garijo, Miguel Arroyo, Esteban Gonzalez, Christoph Treude, and Nicola Tarocco. 2024. Bidirectional paper- repository tracing in software engineering. In Int’l Conf. MSR (MSR) . IEEE, 642–646
2024
-
[40]
Rishab Aiyer Ghosh. 2005. Understanding free software developers: Findings from the FLOSS study. In Perspectives on Free and Open Source Software . MIT Press, 23–46
2005
-
[41]
Michael Haider, Michael Riesch, and Christian Jirauschek. 2021. Realization of best practices in software engineering and scientific writing through ready-to-use project skeletons. Optical and Quantum Electronics 53, 10 (2021), 1–17
2021
-
[42]
Wilhelm Hasselbring, Leslie Carr, Simon Hettrick, Heather Packer, and Thanassis Tiropanis. 2020. Open source research software. IEEE Computer (2020)
2020
-
[43]
Runzhi He, Hengzhi Ye, and Minghui Zhou. 2024. Revealing the value of Repository Centrality in lifespan prediction of Open Source Software Projects. arXiv:2405.07508 [cs.SE] https://arxiv.org/abs/2405.07508
2024 arXiv
-
[44]
Konrad Hinsen. 2019. Dealing With Software Collapse. Comput. Sci. Eng. 21, 3 (2019), 104–108
2019
-
[45]
James Howison and Julia Bullard. 2016. Software in the scientific literature: Problems with seeing, finding, and using software mentioned in the biology literature. J. Assoc. Inf. Sci. Technol. 67, 9 (2016), 2137–2155
2016
-
[46]
James Howison, Ewa Deelman, Michael J McLennan, Rafael Ferreira da Silva, and James D Herbsleb. 2015. Under- standing the scientific software ecosystem and its impact: Current and future measures. Research Evaluation 24, 4 (2015), 454–470
2015
-
[47]
James Howison and James D Herbsleb. 2011. Scientific software production: incentives and collaboration. In ACM Conf. Computer Supported Cooperative Work (CSCW) . 513–522
2011
-
[48]
Yu Huang, Denae Ford, and Thomas Zimmermann. 2021. Leaving my fingerprints: Motivations and challenges of contributing to OSS for social good. In Int’l Conf. Software Engineering (ICSE) . IEEE, 1020–1032
2021
-
[49]
Ana-Maria Istrate, Donghui Li, Dario Taraborelli, Michaela Torkar, Boris Veytsman, and Ivana Williams. 2022. A large dataset of software mentions in the biomedical literature. arXiv preprint arXiv:2209.00693 (2022)
2022 arXiv
-
[50]
Slinger Jansen, Elena Baninemeh, Siamak Farshidi, et al. 2022. FAIRSECO: An infrastructure for measuring impact of research software. In CEUR Worksh. Proc., Vol. 3245. CEUR WS
2022
-
[51]
Mitchell Joblin, Sven Apel, Claus Hunsen, and Wolfgang Mauerer. 2017. Classifying developers into core and peripheral: An empirical study on count and network metrics. In Int’l Conf. Software Eng. (ICSE) . IEEE, 164–174
2017
-
[52]
Arne Johanson and Wilhelm Hasselbring. 2018. Software engineering for computational science: Past, present, future. Comput. Sci. Eng. 20, 2 (2018), 90–109
2018
-
[53]
Niklas Jørgensen. 2021. Assessing Open Source Software as a Scholarly Contribution. Commun. ACM 64, 6 (2021), 62–67
2021
-
[54]
Eirini Kalliamvakou, Georgios Gousios, Kelly Blincoe, Leif Singer, Daniel M German, and Daniela Damian. 2016. An in-depth study of the promises and perils of mining GitHub. Empir. Softw. Eng. 21 (2016), 2035–2071
2016
-
[55]
Upulee Kanewala and James M. Bieman. 2014. Testing scientific software: A systematic literature review. Inf. Softw. Technol. 56, 10 (2014), 1219–1232
2014
-
[56]
Daniel Katz, Kyle Niemeyer, and Arfon Smith. 2018. Publish your software: Introducing the Journal of Open Source Software (JOSS). Comput. Sci. Eng. 20, 3 (2018), 84–88. , Vol. 1, No. 1, Article . Publication date: April 2025. Scientific Open-Source Software Is Less Likely to B...
2018
-
[57]
Aidan Kelley and Daniel Garijo. 2021. A framework for creating knowledge graphs of scientific software metadata. Quantitative Science Studies 2, 4 (2021), 1423–1446
2021
-
[58]
Kleinbaum and Mitchel Klein
David G. Kleinbaum and Mitchel Klein. 2012. The Cox Proportional Hazards Model and Its Characteristics . Springer, 97–159
2012
-
[59]
Julia Koehler Leman, Brian D Weitzner, Pamela D Renfrew, et al . 2020. Better together: Elements of successful scientific software development in a distributed collaborative community. PLOS Comput. Biol. 16, 5 (2020)
2020
-
[60]
Lakhani and Eric von Hippel
Karim R. Lakhani and Eric von Hippel. 2003. How Open Source Software Works: “Free” User-to-User Assistance. Research Policy 32, 6 (2003), 923–943
2003
-
[61]
Anna-Lena Lamprecht, Leyla Garcia, Mateusz Kuzak, Carlos Martinez, Ricardo Arcila, Eva Martin Del Pico, Victoria Dominguez Del Angel, Stephanie Van De Sandt, Jon Ison, Paula Andrea Martinez, et al. 2020. Towards FAIR principles for research software. Data Science 3, 1 (2020), 37–59
2020
-
[62]
J Richard Landis and Gary G Koch. 1977. An Application of Hierarchical Kappa-type Statistics in the Assessment of Majority Agreement among Multiple Observers. Biometrics 33, 2 (1977), 363–374
1977
-
[63]
Benjamin D Lee. 2018. Ten simple rules for documenting scientific software. PLOS Comput. Biol. 14, 12 (2018), e1006561
2018
-
[64]
Zhifang Liao, Benhong Zhao, Shengzong Liu, Haozhi Jin, Dayu He, Liu Yang, Yan Zhang, and Jinsong Wu. 2019. A prediction model of the project life-span in open source software ecosystem. Mobile Networks and Applications 24 (2019), 1382–1391
2019
-
[65]
R. J. A. Little and D. B. Rubin. 1987. Statistical Analysis with Missing Data . John Willey & Sons
1987
-
[66]
Yuxing Ma, Tapajit Dey, Chris Bogart, Sadika Amreen, Marat Valiev, Adam Tutko, David Kennard, Russell Zaretzki, and Audris Mockus. 2021. World of code: enabling a research workflow for mining and analyzing the universe of open source VCS data. Empir. Softw. Eng. 26 (2021), 1–42
2021
-
[67]
Konstantinos Manikas and Klaus Marius Hansen. 2013. Software ecosystems–A systematic literature review. J. Syst. Softw. 86, 5 (2013), 1294–1306
2013
-
[68]
Allen Mao, Daniel Garijo, and Shobeir Fakhraei. 2019. SoMEF: A framework for capturing scientific software metadata from its documentation. In 2019 IEEE Int’l Conf. on Big Data (Big Data) . IEEE, 3032–3037
2019
-
[69]
Reed Milewicz, Gustavo Pinto, and Paige Rodeghero. 2019. Characterizing the roles of contributors in open-source scientific software projects. In Int’l Conf. MSR (MSR) . IEEE, 421–432
2019
-
[70]
Courtney Miller, Mahmoud Jahanshahi, Audris Mockus, Bogdan Vasilescu, and Christian Kästner. 2025. Understanding the Response to Open-Source Dependency Abandonment in the npm Ecosystem. In Int’l Conf. Software Eng. (ICSE)
2025
-
[71]
We Feel Like We’re Winging It:
Courtney Miller, Christian Kästner, and Bogdan Vasilescu. 2023. “We Feel Like We’re Winging It:” A Study on Navigating Open-Source Dependency Abandonment. In Int’l Conf. Foundations of Software Eng. (FSE) . 1281–1293
2023
-
[72]
Fielding, and James Herbsleb
Audris Mockus, Roy T. Fielding, and James Herbsleb. 2002. Two Case Studies of Open Source Software Development: Apache and Mozilla. ACM Trans. Softw. Eng. Methodol. 11, 3 (2002), 309–346
2002
-
[73]
Audris Mockus, Diomidis Spinellis, Zoe Kotti, and Gabriel John Dusing. 2020. A Complete Set of Related Git Repositories Identified via Community Detection Approaches Based on Shared Commits. In Int’l Conf. MSR (MSR)
2020
-
[74]
Dirk F Moore et al. 2016. Applied survival analysis using R . Vol. 473. Springer
2016
-
[75]
Lee Morris. 2021. Understanding Software Sustainability in the field of Research Software Engineering. Ph. D. Dissertation. University of Huddersfield
2021
-
[76]
Nuthan Munaiah, Steven Kroh, Craig Cabrey, and Meiyappan Nagappan. 2017. Curating GitHub for engineered software projects. Empir. Softw. Eng. 22 (2017), 3219–3253
2017
-
[77]
Justin Murphy, Elias T Brady, Shazibul Islam Shamim, and Akond Rahman. 2020. A curated dataset of security defects in scientific software projects. In Proc. of the 7th Symp. on Hot Topics in the Science of Security . 1–2
2020
-
[78]
Udit Nangia and Daniel Katz. 2017. Track 1 Paper: Surveying the U.S. National Postdoctoral Assoc. Regarding Software Use and Training in Research. (2017)
2017
-
[79]
Richard R. Nelson. 1993. National Innovation Systems: A Comparative Analysis . Oxford University Press, New York, NY
1993
-
[80]
Andrew Nesbitt, Boris Veytsman, Daniel Mietchen, Eva Maxfield Brown, James Howison, João Felipe Pimentel, Laurent Hébert-Dufresne, and Stephan Druskat. 2024. Biomedical open source software: Crucial packages and hidden heroes. arXiv preprint arXiv:2404.06672 (2024)
2024 arXiv
-
[81]
Christopher Newfield. 2025. Humanities Decline in Darkness: How Humanities Research Funding Works. Public Humanities 1 (2025), e31
2025
-
[82]
Ileana Ober and Iulian Ober. 2017. On patterns of multi-domain interaction for scientific software development focused on separation of concerns. Procedia computer science 108 (2017), 2298–2302
2017
-
[83]
Gede Artha Azriadi Prana, Christoph Treude, Ferdian Thung, Thushari Atapattu, and David Lo. 2019. Categorizing the content of GitHub README files. Empir. Softw. Eng. 24 (2019), 1296–1327. , Vol. 1, No. 1, Article . Publication date: April 2025. 24 A. Malviya Thakur, R. Milewic...
2019
-
[84]
F Queiroz, R Silva, J Miller, S Brockhauser, and H Fangohr. 2017. Good Usability Practices in Scientific Software Development. In Worksh. on Sustainable Software for Science (WSSSPE5) . WSSSPE
2017
-
[85]
Karthik Ram, Carl Boettiger, Scott Chamberlain, Noam Ross, Maëlle Salmon, and Stefanie Butland. 2018. A community of practice around peer review for long-term research software sustainability. Comput. Sci. Eng. 21, 2 (2018), 59–65
2018
-
[86]
Rahul Ramachandran, Kaylin Bugbee, and Kevin Murphy. 2021. From Open Data to Open Science. Earth and Space Science 8, 5 (2021), e2020EA001562
2021
-
[87]
Klaus Rechert, Jurek Oberhauser, and Rafael Gieschke. 2021. How long can we build it? ensuring usability of a scientific code base. Int’l Journal of Digital Curation 16, 1 (2021), 11–11
2021
-
[88]
Bernard Reddy and David Evans. 2002. Government Preferences for Promoting Open-Source Software: A Solution in Search of a Problem. Michigan Telecommunications and Technology Law Review 9 (05 2002)
2002
-
[89]
Paulo Ribeiro, E
J. Paulo Ribeiro, E. Lima, and G. Rossi. 2020. Governments adoption of open-source software: A systematic review of barriers and enablers. Government Inf. Quarterly 37, 3 (2020), 101447
2020
-
[90]
Patrick Royston and Mahesh KB Parmar. 2013. Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome. BMC medical research methodology 13 (2013), 1–15
2013
-
[91]
Rebecca Sanders, Diane Kelly, and Terry Shepard. 2008. The development and use of scientific software. (2008)
2008
-
[92]
Samuel D Schwartz, Stephen F Fickas, Boyana Norris, and Anshu Dubey. 2024. A survey of open source software repositories in the us department of energy’s national laboratories. Comput. Sci. Eng. (2024)
2024
-
[93]
Schweik and Robert C
Charles M. Schweik and Robert C. English. 2012. Internet success: A study of open-source software commons . MIT Press
2012
-
[94]
Philip Sedgwick. 2012. Multiple significance tests: the Bonferroni correction. BMJ 344 (2012)
2012
-
[95]
Shubh Sharma and Paul Sturges. 2007. Making use of open-source software in government: The practicalities of policy. Inf. Development 23, 2 (2007), 115–125
2007
-
[96]
Vanessa Sochat, Nicholas May, Ian Cosden, Carlos Martinez-Ortiz, and Sadie Bartholomew. 2022. The Research Software Encyclopedia: a community framework to define research software. J. Open Research Software (2022)
2022
-
[97]
David AW Soergel. 2014. Rampant software errors may undermine scientific results. F1000Research 3 (2014)
2014
-
[98]
Jurriaan H Spaaks, J Maassen, T Klaver, S Verhoeven, W Van Hage, L Ridder, L Kulik, T Bakker, V van Hees, L Bogaardt, et al. 2018. Research Software Directory. Zenodo (2018)
2018
-
[99]
Stoltz and Marshall A
Dustin S. Stoltz and Marshall A. Taylor. 2024. Mapping Texts: Computational Text Analysis for the Social Sciences . Oxford University Press
2024
-
[100]
Carly Strasser, Kate Hertweck, Josh Greenberg, Dario Taraborelli, and Elizabeth Vu. 2022. Ten simple rules for funding scientific open source software. PLOS Comput. Biol. 18, 11 (2022), e1010627
2022
-
[101]
Elizabeth A Stuart, Stephen R Cole, Catherine P Bradshaw, and Philip J Leaf. 2011. The use of propensity scores to assess the generalizability of results from randomized trials. J. R. Stat. Soc. (2011)
2011
-
[102]
Jiayi Sun, Aarya Patil, Youhai Li, Jin LC Guo, and Shurui Zhou. 2024. How to Sustain a Scientific Open-Source Software Ecosystem: Learning from the Astropy Project. arXiv preprint arXiv:2402.15081 (2024)
2024 arXiv
-
[103]
Elizabeth Tipton. 2013. Improving generalizations from experiments using propensity score subclassification: Assumptions, properties, and contexts. Journal of Educational and Behavioral Statistics 38, 3 (2013), 239–266
2013
-
[104]
Erik H Trainer, Chalalai Chaihirunkarn, Arun Kalyanasundaram, and James D Herbsleb. 2014. Community code engagements: summer of code & hackathons for community building in scientific software. In Int’l Conf. Supporting Group Work (GROUP). 111–121
2014
-
[105]
Joanna Tucker. 2022. Facing the Challenge of Digital Sustainability as Humanities Researchers. Journal of the British Academy 10 (2022), 93–120
2022
-
[106]
Hajime Uno, Brian Claggett, Lu Tian, Eisuke Inoue, Paul Gallo, Toshio Miyata, Deborah Schrag, Masahiro Takeuchi, Yoshiaki Uyama, Lihui Zhao, et al. 2014. Moving beyond the hazard ratio in quantifying the between-group difference in survival analysis. Journal of clinical Oncolo...
2014
-
[107]
Marat Valiev, Bogdan Vasilescu, and James Herbsleb. 2018. Ecosystem-level determinants of sustained activity in open-source projects: A case study of the PyPI ecosystem. In Int’l Conf. Foundations of Software Eng. (FSE) . 644–655
2018
-
[108]
Supatsara Wattanakriengkrai, Bodin Chinthanet, Hideaki Hata, Raula Gaikovina Kula, Christoph Treude, Jin Guo, and Kenichi Matsumoto. 2022. GitHub repositories with links to academic papers: Public access, traceability, and evolution. J. Syst. Softw. (2022)
2022
-
[109]
Michael Weiss. 2005. The Promise of Government Incentives for Open-Source Projects. IEEE Software 22, 5 (2005), 74–77
2005
-
[110]
James Willenbring and Gursimran Singh Walia. 2021. Evaluating the Sustainability of Computational Science and Engineering Software: Empirical Observations. Technical Report. Sandia National Lab.(SNL-NM), Albuquerque, NM (United States). , Vol. 1, No. 1, Article . Publication d...
2021
-
[111]
Bruno Lopes Xavier, Rodrigo Pereira dos Santos, and Davi Viana dos Santos. 2020. Software Ecosystems and Digital Games: Understanding the Financial Sustainability Aspect.. In ICEIS (2). 450–457
2020
-
[112]
Minghui Zhou and Audris Mockus. 2011. Does the initial environment impact the future of developers?. In Int’l Conf. Software Eng. (ICSE). 271–280
2011
-
[113]
Minghui Zhou and Audris Mockus. 2012. What make long term contributors: Willingness and opportunity in OSS community. In Int’l Conf. Software Eng. (ICSE) . IEEE, 518–528
2012
-
[114]
Minghui Zhou, Audris Mockus, Xiujuan Ma, Lu Zhang, and Hong Mei. 2016. Inflow and retention in oss communities with commercial involvement: A case study of three hybrid projects. ACM Trans. Softw. Eng. Methodol. 25, 2 (2016), 1–29
2016
-
[115]
Shurui Zhou, Bogdan Vasilescu, and Christian Kästner. 2019. What the fork: a study of inefficient and efficient forking practices in social coding. In Int’l Conf. Foundations of Software Eng. (FSE) . 350–361. , Vol. 1, No. 1, Article . Publication date: April 2025
2019
-
[2017]
Research software engineers: State of the nation report 2017. (2017)
2017
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.