Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Maintenance and Support in Community-Driven Scientific Pipeline Ecosystems: A Cross-Platform Empirical Study of nf-core

T0 review · 3 major / 6 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Sustaining scientific pipeline ecosystems requires actionable support, review-ready contributions, and stronger links between user forums and repository work—not just engines and templates.

desk verdict Solid, large-scale map of how nf-core actually maintains pipelines and supports users across GitHub and the Seqera forum; useful and carefully scoped, not a theory paper. read the letter →

arxiv 2607.10839 v1 pith:ZGEYSEIC submitted 2026-07-12 cs.SE

classification cs.SE
keywords nf-corescientificworkflowsystemspipelinemaintenancecommunitysupportcross-platformminingempiricalsoftwareengineeringGitHubissuesandpullrequeststraceability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that community-driven scientific pipeline ecosystems stay usable only when communities can maintain pipelines, integrate contributions, and support users across messy real-world environments. Mining tens of thousands of GitHub issues and pull requests plus nearly nine hundred forum threads from nf-core, it shows that the three spaces do different jobs: issues report and coordinate problems, pull requests implement and review fixes, and forums capture execution failures around containers, cloud, HPC, reporting, and Nextflow usage. Resolution is associated with actionability, coordination, and diagnostic evidence—assignees, comments, milestones, bug labels, checklists, linked issues, code blocks, and sustained interaction—while cloud, HPC, and workflow-semantics questions stay harder to settle. Explicit links are dense inside GitHub but almost absent between forums and repositories, so user-facing knowledge often never becomes durable documentation or code change. The practical claim is that sustainability needs structured templates, review-aware automation, infrastructure-specific guidance, and lightweight cross-platform traceability.

What carries the argument

Cross-platform empirical design treating GitHub issues, pull requests, and Seqera Community Forum discussions as complementary units of analysis, combining BERTopic topic modeling, statistical outcome models, direct-link mining, semantic similarity, technical-signal overlap, and qualitative inspection.

What would settle it

Find that explicit forum-to-GitHub links are common once sampling covers more years or platforms, or that after controlling for unobserved maintainer attention the measured associations of assignees, checklists, code blocks, and diagnostic mentions with closure, merge, and accepted answers disappear.

Watch

Extended reading notes

Core claim

In nf-core, maintenance and support are distributed across complementary spaces whose resolution depends on actionability, coordination, and diagnostic evidence, yet problem–solution flow is strongly traceable inside GitHub and weakly linked to community forum support, leaving many user-facing execution problems only implicitly connected to repository maintenance.

Load-bearing premise

The study assumes that the filtered, de-duplicated GitHub and forum records, keyword technical signals, and manually labeled topic clusters are a complete enough picture of real maintenance and support work.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a large-scale cross-platform empirical study of maintenance and support in the nf-core ecosystem, analyzing 15,760 GitHub issues, 35,411 pull requests, and 895 Seqera Community Forum discussions. Using BERTopic, statistical outcome models, direct-link mining, semantic similarity, technical-signal overlap, and qualitative inspection, the authors address four RQs: what maintenance/support topics arise; how they differ across artifact types; which factors associate with resolution (issue closure, PR merge/closed-without-merge, forum accepted answers); and how problems and solutions flow between repository-centered and community-centered spaces. The central claims are that issues, PRs, and forums play complementary roles; resolution co-occurs with actionability, coordination, and diagnostic evidence; and GitHub has dense issue–PR traceability while explicit forum–GitHub links are rare.

Significance. If the results hold, the paper makes a useful contribution to empirical software engineering and scientific workflow research by treating community-driven pipeline ecosystems as socio-technical systems rather than only as workflow-engine or template problems. Strengths include carefully filtered multi-artifact corpora, separate topic models per artifact type, multivariable models with FDR correction and effect sizes, Kaplan–Meier lifecycle curves, manual validation of 100+ linkage candidates, open acknowledgment of observational limits, a replication package, and concrete recommendations for templates, automation, infrastructure guidance, and cross-platform traceability. The work extends prior nf-core and SWS mining studies by jointly analyzing issues, PRs, and forum support and by quantifying weak forum–repository linkage—an actionable gap for similar ecosystems.

major comments (3)
  1. [Section 3.3 / 4.1] Section 3.3 and 4.1: Topic labels that underpin RQ1–RQ2 rest on primary labeling by the first author (top terms + ≥25 artifacts per topic) with co-author consensus, but no quantitative agreement metric (e.g., Cohen’s κ or percent agreement on a held-out sample) is reported. Because the paper’s claim that issues, PRs, and forums expose distinct layers of work depends on these labels, please report inter-rater reliability on a stratified sample of topics/artifacts, or at minimum document disagreement rates and how borderline cases were resolved, so readers can assess label stability.
  2. [Section 3.2 / Table 21] Section 3.2 and Table 21 (RQ4 technical-signal overlap): Technical signals (cloud, HPC, container, MultiQC, etc.) are keyword-based free parameters. Overlap percentages are used to argue that similar concerns recur across spaces even without explicit links. A brief sensitivity check (alternate keyword lists, stemming variants, or precision/recall on a labeled sample) would show whether the cross-platform overlap pattern is robust or keyword-list-dependent. Without this, the aggregate-signal layer of RQ4 is harder to trust at the reported precision.
  3. [Section 3.5 / 4.3.4] Section 3.5 and 4.3.4: Forum resolution uses accepted-answer status, with activity span as a lifecycle proxy because accepted-answer timestamps are unavailable. The paper correctly flags this, but several claims (e.g., code blocks associated with faster movement toward accepted answers; cloud/HPC harder to resolve) mix outcome rates with exploratory KM-style curves over activity span. Please more sharply separate (i) accepted-answer incidence associations from (ii) time-to-resolution claims for forums, and avoid language that implies exact time-to-answer when only activity span is observed.
minor comments (6)
  1. [Section 3.1 / Table 1] Table 1 vs. abstract/body counts: Final topic-modeling corpora (15,760 / 35,411 / 895) are clear, but intermediate counts (15,763 issues; 36,521 then 35,430 PRs) appear in multiple places. A single canonical data-flow sentence in Section 3.1 would reduce reader friction.
  2. [Figure 3 / Table 2] Figure 3 topic names sometimes differ slightly from Table 2 (e.g., “Linting and Template-Check Failures” vs. “Automated Linting and Template Compliance”). Align labels across tables and figures.
  3. [Table 17] Section 4.3.3, Table 17: The large positive OR for posts count and negative OR for replies count is explained as overlapping engagement measures, but a short note on multicollinearity diagnostics (or a model without both) would help readers interpret the regularized forum model.
  4. [Section 2.5 / 3.3] Section 2.5: BERTopic is well motivated; a one-sentence comparison of final topic counts/coherence against a simple LDA baseline (even if only in the replication package) would strengthen the modeling choice for readers who still expect LDA in SE mining papers.
  5. [Tables 2–4] Minor typos/consistency: e.g., “igenomes” / “iGenomes”, “nfcoretool” spacing, and occasional topic-name truncation in tables. A pass for notation consistency (nf-core vs nfcore) would help.
  6. [Section 5.6] Discussion 5.6 generalizes to Galaxy, Snakemake, and Pegasus. Soften or qualify this as a hypothesis for future comparative work, since all empirical results are nf-core-specific (as Section 6.2 already notes).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: observational mining study with associations, not predictions forced by construction.

full rationale

This is a cross-platform empirical mining study of nf-core (issues, PRs, forum posts). Topics come from BERTopic plus manual labeling of the collected corpora; resolution analyses report lifecycle associations (logistic/Cox/regularized models, KM curves) and are explicitly framed as non-causal; problem–solution flow is measured via direct links, TF–IDF similarity candidates, and keyword technical-signal overlap. No quantity is fitted to data and then re-presented as an independent prediction of a closely related target. Self-citations to Alam & Roy and Alam et al. supply prior nf-core/SWS context and methods precedent; they do not uniquely force or define the new cross-platform findings. The derivation chain is data → descriptive topics/associations/link counts, which is self-contained against the stated observational scope. Score 0 with empty steps is appropriate.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

Empirical mining study; almost all content is data-driven. The few modeling choices (BERTopic hyperparameters, keyword lists for technical signals, one-month observation buffer, exclusion of archived repos and linked PRs from the PR corpus) are free parameters or domain assumptions rather than invented physical entities.

free parameters (3)
  • BERTopic UMAP/HDBSCAN hyperparameters (n_neighbors, n_components, min_cluster_size)
    Tuned per corpus for coherence; different settings would alter topic granularity and therefore the topic-level resolution rates reported in RQ1–RQ3.
  • Keyword lists for technical signals (cloud, HPC, container, MultiQC, etc.)
    Hand-crafted extractors used for both outcome models and cross-platform overlap; false positives/negatives affect association strengths.
  • One-month observation buffer (cutoff 8 April 2026)
    Chosen to reduce right-censoring; shorter or longer windows change open/closed counts.
assumptions (3)
  • domain assumption GitHub issues, pull requests, and Seqera forum posts are sufficiently complete and representative records of nf-core maintenance and support activity.
    Stated in Section 3 and Threats; private chat, local debugging, and unlogged work are acknowledged but treated as secondary.
  • domain assumption Closure, merge, and accepted-answer status are valid proxies for resolution.
    Used throughout RQ3; authors note that a closed issue or accepted answer may not fully solve the underlying problem.
  • ad hoc to paper BERTopic clusters plus manual inspection of ≥25 documents per topic yield interpretable, non-overlapping maintenance themes.
    Core of RQ1 labeling procedure; inter-rater discussion is described but no quantitative agreement metric is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Maintenance and Support in Community-Driven Scientific Pipeline Ecosystems: A Cross-Platform Empirical Study of nf-core." pith.science (2026). https://pith.science/paper/ZGEYSEIC

@misc{pith2026260710839,
  author       = {Pith},
  title        = {Pith review of: Maintenance and Support in Community-Driven Scientific Pipeline Ecosystems: A Cross-Platform Empirical Study of nf-core},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZGEYSEIC}},
  note         = {Machine review of arXiv:2607.10839}
}
read the original abstract

Community-driven scientific pipeline ecosystems are increasingly important for reproducible data-intensive research, but their sustainability depends on more than workflow engines, templates, and testing infrastructure. It also depends on how communities maintain pipelines, integrate contributions, and support users across heterogeneous execution environments. This paper presents a cross-platform empirical study of maintenance and support in nf-core, a large ecosystem of standardized Nextflow pipelines. We analyze 15,760 GitHub issues, 35,411 GitHub pull requests, and 895 Seqera Community Forum discussions to examine what maintenance and support concerns arise, how they differ across artifact types, which factors are associated with resolution outcomes, and how problems and solutions flow between repository-centered and community-centered spaces. We find that issues primarily capture repository-level problem reporting and maintenance coordination; pull requests capture implementation, review, testing, dependency, and template-update work; and forum discussions capture user-facing support around execution failures, containers, cloud and HPC environments, MultiQC reporting, and Nextflow usage. Resolution outcomes are associated with actionability, coordination, and diagnostic evidence. Issue closure is linked to assignees, comments, milestones, bug labels, error mentions, and version information. Pull request integration varies by author role, automation type, draft status, checklists, linked issues, and review routing. Forum accepted answers are more likely when discussions include code blocks, sustained interaction, and concrete technical evidence, while cloud, HPC, and workflow-semantics questions are harder to resolve. Cross-platform analysis reveals strong repository-internal traceability within GitHub, but limited explicit linkage between forum discussions and repository artifacts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Federated-learning developers' biggest public pain points are environment setup, API/version breakage, non-IID training instability, and evaluation/privacy integration, with Stack Overflow skewing How and GitHub skewing Why.

Reference graph

Works this paper leans on

79 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Nature biotechnology 38(3), 276–278 (2020)

    Ewels, P.A., Peltzer, A., Fillinger, S., Patel, H., Alneberg, J., Wilm, A., Garcia, M.U., Di Tommaso, P., Nahnsen, S.: The nf-core framework for community-curated bioinformatics pipelines. Nature biotechnology 38(3), 276–278 (2020)

  2. [2]

    Journal of Grid Computing13, 457–493 (2015)

    Liu, J., Pacitti, E., Valduriez, P., Mattoso, M.: A survey of data-intensive scientific workflow management. Journal of Grid Computing13, 457–493 (2015)

  3. [3]

    arXiv preprint arXiv:2601.09612 (2026)

    Alam, K., Roy, B.: Analyzing github issues and pull requests in nf-core pipelines: Insights into nf-core pipeline repositories. arXiv preprint arXiv:2601.09612 (2026)

  4. [4]

    Nature biotechnology35(4), 316–319 (2017)

    Di Tommaso, P., Chatzou, M.,et al.: Nextflow enables reproducible computational workflows. Nature biotechnology35(4), 316–319 (2017)

  5. [5]

    Nature biotechnology37(4), 358–367 (2019)

    Sansone, S.-A., McQuilton, P., Rocca-Serra, P., Gonzalez-Beltran, A., Izzo, M., Lister, A.L., Thurston, M., Community, F.: Fairsharing as a community approach to standards, repositories and policies. Nature biotechnology37(4), 358–367 (2019)

  6. [6]

    Scientific data3(1), 1–9 (2016)

    Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., Silva Santos, L.B., Bourne, P.E.,et al.: The fair guiding principles for scientific data management and stewardship. Scientific data3(1), 1–9 (2016)

  7. [7]

    FGCS75, 284–298 (2017)

    Cohen-Boulakia, S., Belhajjame, K.,et al.: Scientific workflows for computational reproducibility in the life sciences: Status, challenges and opportunities. FGCS75, 284–298 (2017)

  8. [8]

    Nature methods18(10), 1161–1168 (2021)

    Wratten, L., Wilm, A., G¨ oke, J.: Reproducible, scalable, and shareable analysis pipelines with bioinformatics workflow managers. Nature methods18(10), 1161–1168 (2021)

Show all 79 references
  1. [9]

    Scientific data9(1), 622 (2022)

    Barker, M., Chue Hong, N.P., Katz, D.S., Lamprecht, A.-L., Martinez-Ortiz, C., Psomopoulos, F., Harrow, J., Castro, L.J., Gruenpeter, M., Martinez, P.A.,et al.: Introducing the fair principles for research software. Scientific data9(1), 622 (2022)

  2. [10]

    Empirical Software Engineering30(5), 151 (2025)

    Alam, K., Roy, B., Roy, C.K., Mittal, K.: An empirical investigation on the challenges in scientific workflow systems development. Empirical Software Engineering30(5), 151 (2025)

  3. [11]

    Bioinformatics28(19), 2520– 2522 (2012)

    K¨ oster, J., Rahmann, S.: Snakemake—a scalable bioinformatics workflow engine. Bioinformatics28(19), 2520– 2522 (2012)

  4. [12]

    Genome biology11, 1–13 (2010)

    Goecks, J., Nekrutenko, A.,et al.: Galaxy: a comprehensive approach for supporting accessible, reproducible, and transparent computational research in the life sciences. Genome biology11, 1–13 (2010)

  5. [13]

    FGCS46, 17–35 (2015)

    Deelman, E., Vahi, K.,et al.: Pegasus, a workflow management system for science automation. FGCS46, 17–35 (2015)

  6. [14]

    Bioinformatics20(17), 3045–3054 (2004)

    Oinn, T., Addis, M., Ferris, J., Marvin, D., Senger, M., Greenwood, M., Carver, T., Glover, K., Pocock, M.R., Wipat, A.,et al.: Taverna: a tool for the composition and enactment of bioinformatics workflows. Bioinformatics20(17), 3045–3054 (2004)

  7. [15]

    Future Generation Computer Systems174, 107974 (2026)

    Suter, F., Coleman, T., Altinta¸ s,˙I., Badia, R.M., Balis, B., Chard, K., Colonnelli, I., Deelman, E., Di Tom- maso, P., Fahringer, T.,et al.: A terminology for scientific workflow systems. Future Generation Computer Systems174, 107974 (2026)

  8. [16]

    Briefings in bioinformatics18(3), 530–536 (2017)

    Leipzig, J.: A review of bioinformatic pipeline frameworks. Briefings in bioinformatics18(3), 530–536 (2017)

  9. [17]

    Future Generation Computer Systems75, 228–238 38 (2017)

    Da Silva, R.F., Filgueira, R., Pietri, I., Jiang, M., Sakellariou, R., Deelman, E.: A characterization of work- flow management systems for extreme-scale applications. Future Generation Computer Systems75, 228–238 38 (2017)

  10. [18]

    Genome Biology26(1), 228 (2025)

    Langer, B.E., Amaral, A., Baudement, M.-O., Bonath, F., Charles, M., Chitneedi, P.K., Clark, E.L., Di Tom- maso, P., Djebali, S., Ewels, P.A.,et al.: Empowering bioinformatics communities with nextflow and nf-core. Genome Biology26(1), 228 (2025)

  11. [19]

    In: 2009 31st International Conference on Software Engineering-Companion Volume, pp

    Jansen, S., Finkelstein, A., Brinkkemper, S.: A sense of community: A research agenda for software ecosystems. In: 2009 31st International Conference on Software Engineering-Companion Volume, pp. 187–190 (2009). IEEE

  12. [20]

    Journal of Systems and Software117, 84–103 (2016)

    Manikas, K.: Revisiting software ecosystems research: A longitudinal literature study. Journal of Systems and Software117, 84–103 (2016)

  13. [21]

    IEEE Transactions on Software Engineering43(2), 185–204 (2016)

    Storey, M.-A., Zagalsky, A., Figueira Filho, F., Singer, L., German, D.M.: How social and communication channels shape and challenge a participatory culture in software development. IEEE Transactions on Software Engineering43(2), 185–204 (2016)

  14. [22]

    Updated: January 4, 2019; Accessed: 2026-04-09 (2009)

    Preston-Werner, T.: GitHub Issue Tracker! GitHub Blog. Updated: January 4, 2019; Accessed: 2026-04-09 (2009). https://github.blog/news-insights/github-issue-tracker/

  15. [23]

    https://github.blog/news-insights/pull-requests-2-0/

    Tomayko, R.: Pull Requests 2.0. https://github.blog/news-insights/pull-requests-2-0/. Updated: 2019-12-16; Accessed: 2026-05-30 (2010)

  16. [24]

    https://community.seqera.io/

    Seqera: Seqera Community. https://community.seqera.io/. Accessed: 2026-04-21 (2026)

  17. [25]

    In: Proceedings of the 36th International Conference on Software Engineering, pp

    Tsay, J., Dabbish, L., Herbsleb, J.: Influence of social and technical factors for evaluating contribution in github. In: Proceedings of the 36th International Conference on Software Engineering, pp. 356–366 (2014)

  18. [26]

    In: Proceedings of the 36th International Conference on Software Engineering, pp

    Gousios, G., Pinzger, M., Deursen, A.v.: An exploratory study of the pull-based software development model. In: Proceedings of the 36th International Conference on Software Engineering, pp. 345–355 (2014)

  19. [27]

    Empirical Software Engineering21(5), 2035–2071 (2016)

    Kalliamvakou, E., Gousios, G., Blincoe, K., Singer, L., German, D.M., Damian, D.: An in-depth study of the promises and perils of mining github. Empirical Software Engineering21(5), 2035–2071 (2016)

  20. [28]

    Empirical Software Engineering27(1), 3 (2022)

    Hata, H., Novielli, N., Baltes, S., Kula, R.G., Treude, C.: Github discussions: An exploratory study of early adoption. Empirical Software Engineering27(1), 3 (2022)

  21. [29]

    In: Proceedings of the 15th International Conference on Cooperative and Human Aspects of Software Engineering, pp

    Hellman, J., Chen, J., Uddin, M.S., Cheng, J., Guo, J.L.: Characterizing user behaviors in open-source software user forums: an empirical study. In: Proceedings of the 15th International Conference on Cooperative and Human Aspects of Software Engineering, pp. 46–55 (2022)

  22. [30]

    In: 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR), pp

    Rahman, M.S., Codabux, Z., Roy, C.K.: Investigating the understandability of review comments on code change requests. In: 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR), pp. 539–551 (2025). IEEE

  23. [31]

    In: Proceedings of the 11th Working Conference on Mining Software Repositories, pp

    Kalliamvakou, E., Gousios, G., Blincoe, K., Singer, L., German, D.M., Damian, D.: The promises and perils of mining github. In: Proceedings of the 11th Working Conference on Mining Software Repositories, pp. 92–101 (2014)

  24. [32]

    In: Proceedings of the 11th Working Conference on Mining Software Repositories, pp

    Rahman, M.M., Roy, C.K.: An insight into the pull requests of github. In: Proceedings of the 11th Working Conference on Mining Software Repositories, pp. 364–367 (2014)

  25. [33]

    Computational and Structural Biotechnology Journal21, 2075–2085 (2023)

    Djaffardjy, M., Marchment, G., Sebe, C., Blanchet, R., Belhajjame, K., Gaignard, A., Lemoine, F., Cohen- Boulakia, S.: Developing and reusing bioinformatics data analysis pipelines using scientific workflow systems. Computational and Structural Biotechnology Journal21, 2075–20...

  26. [34]

    In: 2023 30th APSEC, pp

    Alam, K., Roy, B., Serebrenik, A.: Reusability challenges of scientific workflows: A case study for galaxy. In: 2023 30th APSEC, pp. 289–298 (2023). IEEE

  27. [35]

    Data Intelligence2(1-2), 108–121 (2020) 39

    Goble, C., Cohen-Boulakia, S., Soiland-Reyes, S., Garijo, D., Gil, Y., Crusoe, M.R., Peters, K., Schober, D.: Fair computational workflows. Data Intelligence2(1-2), 108–121 (2020) 39

  28. [36]

    In: 2022 IEEE/ACM Workshop on Workflows in Support of Large-Scale Science (WORKS), pp

    Alam, K., Roy, B.: Challenges of provenance in scientific workflow management systems. In: 2022 IEEE/ACM Workshop on Workflows in Support of Large-Scale Science (WORKS), pp. 10–18 (2022). IEEE

  29. [37]

    Methods in Ecology and Evolution14(6), 1364–1380 (2023)

    Braga, P.H.P., H´ ebert, K., Hudgins, E.J., Scott, E.R., Edwards, B.P., S´ anchez Reyes, L.L., Grainger, M.J., Foroughirad, V., Hillemann, F., Binley, A.D.,et al.: Not just for programmers: How github can accelerate collaborative and reproducible research in ecology and evolut...

  30. [38]

    Journal of Software: Evolution and Process36(4), 2554 (2024)

    Tan, S.H., Li, Z., Yan, L.: Crossfix: Resolution of github issues via similar bugs recommendation. Journal of Software: Evolution and Process36(4), 2554 (2024)

  31. [39]

    In: Proceedings of the 30th Annual ACM Symposium on Applied Computing, pp

    Soares, D.M., Lima J´ unior, M.L., Murta, L., Plastino, A.: Acceptance factors of pull requests in open-source projects. In: Proceedings of the 30th Annual ACM Symposium on Applied Computing, pp. 1541–1546 (2015)

  32. [40]

    Journal of Systems and Software171, 110806 (2021)

    Lenarduzzi, V., Nikkola, V., Saarim¨ aki, N., Taibi, D.: Does code quality affect pull request acceptance? an empirical study. Journal of Systems and Software171, 110806 (2021)

  33. [41]

    Empirical Software Engineering25, 2694–2747 (2020)

    Han, J., Shihab, E., Wan, Z., Deng, S., Xia, X.: What do programmers discuss about deep learning frameworks. Empirical Software Engineering25, 2694–2747 (2020)

  34. [42]

    In: ICSME, pp

    Li, H., Khomh, F., Openja, M.,et al.: Understanding quantum software engineering challenges an empirical study on stack exchange forums and github issues. In: ICSME, pp. 343–354. IEEE, ??? (2021). IEEE

  35. [43]

    In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pp

    Yang, Z., Wang, C., Shi, J., Hoang, T., Kochhar, P., Lu, Q., Xing, Z., Lo, D.: What do users ask in open- source ai repositories? an empirical study of github issues. In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pp. 79–91 (2023). IEEE

  36. [44]

    Multimedia Tools and Applications78, 15169–15211 (2019)

    Jelodar, H., Wang, Y., Yuan, C., Feng, X., Jiang, X., Li, Y., Zhao, L.: Latent dirichlet allocation (lda) and topic modeling: models, applications, a survey. Multimedia Tools and Applications78, 15169–15211 (2019)

  37. [45]

    In: Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering-Volume 1, pp

    Asuncion, H.U., Asuncion, A.U., Taylor, R.N.: Software traceability with topic modeling. In: Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering-Volume 1, pp. 95–104 (2010)

  38. [46]

    Journal of Information Science40(5), 621–636 (2014)

    Bagheri, A., Saraee, M., De Jong, F.: Adm-lda: An aspect detection model based on topic modelling using the structure of review sentences. Journal of Information Science40(5), 621–636 (2014)

  39. [47]

    In: 2010 IEEE International Conference on Software Maintenance, pp

    Gethers, M., Poshyvanyk, D.: Using relational topic models to capture coupling among classes in object- oriented software systems. In: 2010 IEEE International Conference on Software Maintenance, pp. 1–10 (2010). IEEE

  40. [48]

    935–946 (2016)

    Nadi, S., Kr¨ uger, S., Mezini, M., Bodden, E.: Jumping through hoops: Why do java developers struggle with cryptography apis? In: Proceedings of the 38th International Conference on Software Engineering, pp. 935–946 (2016)

  41. [49]

    In: 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR), pp

    Scoccia, G.L., Migliarini, P., Autili, M.: Challenges in developing desktop web apps: a study of stack overflow and github. In: 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR), pp. 271–282 (2021). IEEE

  42. [50]

    In: Proceedings of the ACM/IEEE 19th International Conference on Model Driven Engineering Languages and Systems, pp

    Kahani, N., Bagherzadeh, M., Dingel, J., Cordy, J.R.: The problems with eclipse modeling tools: a topic analysis of eclipse forums. In: Proceedings of the ACM/IEEE 19th International Conference on Model Driven Engineering Languages and Systems, pp. 227–237 (2016)

  43. [51]

    Online Learning28(1), 175–195 (2024)

    Barker, H.A., Lee, H.S., Kellogg, S., Anderson, R.: The viability of topic modeling to identify participant motivations for enrolling in online professional development. Online Learning28(1), 175–195 (2024)

  44. [52]

    In: Proceedings of the 13th Innovations in Software Engineering Conference on Formerly Known as India Software Engineering Conference, pp

    Dhasade, A.B., Venigalla, A.S.M., Chimalakonda, S.: Towards prioritizing github issues. In: Proceedings of the 13th Innovations in Software Engineering Conference on Formerly Known as India Software Engineering Conference, pp. 1–5 (2020)

  45. [53]

    ” (2021) 40

    Jokhio, M.: Mining github issues for bugs, feature requests and questions. ” (2021) 40

  46. [54]

    94–97 (2019)

    Wang, X., Lee, M., Pinchbeck, A., Fard, F.: Where does lda sit for github? In: 2019 34th IEEE/ACM International Conference on Automated Software Engineering Workshop (ASEW), pp. 94–97 (2019). IEEE

  47. [55]

    Journal of machine Learning research3(Jan), 993–1022 (2003)

    Blei, D.M., Ng, A.Y., Jordan, M.I.: Latent dirichlet allocation. Journal of machine Learning research3(Jan), 993–1022 (2003)

  48. [56]

    arXiv preprint arXiv:2203.05794 (2022)

    Grootendorst, M.: Bertopic: Neural topic modeling with a class-based tf-idf procedure. arXiv preprint arXiv:2203.05794 (2022)

  49. [57]

    bioRxiv, 2024–05 (2024)

    Langer, B.E., Amaral, A., et al.: Empowering bioinformatics communities with nextflow and nf-core. bioRxiv, 2024–05 (2024)

  50. [58]

    Packt Publishing Ltd, ??? (2016)

    Hardeniya, N., Perkins, J., Chopra, D., Joshi, N., Mathur, I.: Natural Language Processing: Python and NLTK. Packt Publishing Ltd, ??? (2016)

  51. [59]

    No Starch Press, San Francisco, CA 94103, USA (2020)

    Vasiliev, Y.: Natural Language Processing with Python and spaCy: A Practical Introduction. No Starch Press, San Francisco, CA 94103, USA (2020)

  52. [60]

    Reimers, N., Gurevych, I.: Sentence-bert: Sentence embeddings using siamese bert-networks. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3...

  53. [61]

    https://huggingface.co/ sentence-transformers

    Sentence Transformers community: Sentence Transformers — Hugging Face Hub. https://huggingface.co/ sentence-transformers. Accessed: 2025-10-07 (2025)

  54. [62]

    https://huggingface.co/models

    Community, H.F.: Hugging Face Models. https://huggingface.co/models. Accessed: 2025-10-07 (2025)

  55. [63]

    arXiv preprint arXiv:1802.03426 (2018)

    McInnes, L., Healy, J., Melville, J.: Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018)

  56. [64]

    McInnes, L., Healy, J., Astels, S.,et al.: hdbscan: Hierarchical density based clustering. J. Open Source Softw. 2(11), 205 (2017)

  57. [65]

    In: Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, pp

    R¨ oder, M., Both, A., Hinneburg, A.: Exploring the space of topic coherence measures. In: Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, pp. 399–408 (2015)

  58. [66]

    In: MSR, pp

    Abdellatif, A., Costa, D., Badran, K., Abdalkareem, R., Shihab, E.: Challenges in chatbot development: A study of stack overflow posts. In: MSR, pp. 174–185 (2020)

  59. [67]

    John Wiley & Sons, Hoboken, New Jersey

    Agresti, A.: Categorical Data Analysis. John Wiley & Sons, Hoboken, New Jersey. (2013)

  60. [68]

    Chapman and Hall/CRC, Boca Raton, Florida (2017)

    Fagerland, M., Lydersen, S., Laake, P.: Statistical Analysis of Contingency Tables. Chapman and Hall/CRC, Boca Raton, Florida (2017)

  61. [69]

    Annals of statistics, 1165–1188 (2001)

    Benjamini, Y., Yekutieli, D.: The control of the false discovery rate in multiple testing under dependency. Annals of statistics, 1165–1188 (2001)

  62. [70]

    John Wiley & Sons, Hoboken, New Jersey

    Hosmer Jr, D.W., Lemeshow, S., Sturdivant, R.X.: Applied Logistic Regression. John Wiley & Sons, Hoboken, New Jersey. (2013)

  63. [71]

    Annual Review of Statistics and Its Application 10(1), 1–23 (2023)

    Kalbfleisch, J.D., Schaubel, D.E.: Fifty years of the cox model. Annual Review of Statistics and Its Application 10(1), 1–23 (2023)

  64. [72]

    Journal of statistical software33, 1–22 (2010)

    Friedman, J.H., Hastie, T., Tibshirani, R.: Regularization paths for generalized linear models via coordinate descent. Journal of statistical software33, 1–22 (2010)

  65. [73]

    Oxidative medicine and cellular longevity2021(1), 2290120 (2021)

    D’Arrigo, G., Leonardis, D., Abd ElHafeez, S., Fusaro, M., Tripepi, G., Roumeliotis, S.: Methods to analyse time-to-event data: The kaplan-meier survival curve. Oxidative medicine and cellular longevity2021(1), 2290120 (2021)

  66. [74]

    In: 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), pp

    Openja, M., Adams, B., Khomh, F.: Analysis of modern release engineering topics:–a large-scale study using 41 stackoverflow–. In: 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), pp. 104–114 (2020). IEEE

  67. [75]

    JCST31, 910–924 (2016)

    Yang, X.-L., Lo, D., Xia, X., Wan, Z.-Y., Sun, J.-L.: What security questions do developers ask? a large-scale study of stack overflow posts. JCST31, 910–924 (2016)

  68. [76]

    In: Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp

    Bagherzadeh, M., Khatchadourian, R.: Going big: a large-scale study on what big data developers ask. In: Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp. 432–442 (2019)

  69. [77]

    Organizational Research Methods10(2), 393 (2007)

    Bean, C.J.: Qualitative research design: An interactive approach. Organizational Research Methods10(2), 393 (2007)

  70. [78]

    Springer Science & Business Media (2012)

    Wohlin, C., Runeson, P., H¨ ost, M., Ohlsson, M.C., Regnell, B., Wessl´ en, A.: Experimentation in software engineering. Springer Science & Business Media (2012)

  71. [79]

    Maintenance and Support in Community-Driven Scientific Pipeline Ecosystems: A Cross-Platform Empirical Study of Nf-core

    Anonymous: Artifact of the Paper “Maintenance and Support in Community-Driven Scientific Pipeline Ecosystems: A Cross-Platform Empirical Study of Nf-core”. https://doi.org/10.5281/zenodo.21324489 . https://doi.org/10.5281/zenodo.21324489 42

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.