Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that six million suspected fake stars accumulated on GitHub between 2019 and 2024, that fake-star campaigns surged in 2024, and that most served to promote short-lived malware or hype repositories.

desk verdict First solid global measurement of fake GitHub stars, but the headline counts are operating-point estimates until someone runs a sensitivity analysis. read the letter →

arxiv 2412.13459 v2 pith:YSD7NU7I submitted 2024-12-18 cs.CR cs.SE

classification cs.CRcs.SE
keywords GitHubfakestarsstarfraudsupplychainsecuritylongitudinalmeasurementmalwarepopularitysignalsGHArchive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that fake GitHub stars are not a fringe phenomenon but a systemic, growing one visible in the platform's public event stream: an estimated six million suspected fake stars were cast between July 2019 and December 2024, and the coordinated campaigns behind them surged sharply in 2024. The authors build StarScout, a detector that runs as SQL queries over the GHArchive event dataset and flags two anomalous starring patterns: accounts that star once and then go silent, and lockstep clusters of accounts starring the same repositories inside short time windows. They argue that if the counts are right, star counts are a partially corrupted popularity signal, a security hazard because a large share of fake-star repositories turn out to be phishing or malware vehicles, and a weak growth-hacking strategy because fake stars attract real attention for under two months and then become a liability. The paper's contribution is the first systematic global longitudinal measurement of fake-star campaigns on GitHub, with downstream stakes for how practitioners, platforms, and supply-chain security tools treat star counts.

What carries the argument

The carrying mechanism is StarScout, which detects two signatures of anomalous starring in the stargazer bipartite graph of accounts and repositories. The low activity signature flags accounts with a single WatchEvent plus at most one other event, i.e., accounts that star one repository and then go stale. The lockstep signature flags groups of at least 50 accounts that repeatedly star at least 10 repositories such that each repository receives at least 25 stars from the group within a 30-day window, found by running the CopyCatch algorithm expressed as SQL on the GHArchive data. A postprocessing step then keeps only repositories with an anomalous monthly spike of fake stars (more than 50 fake stars and over 50% fake in one month, plus over 10% fake all-time), converting noisy per-star signals into a repository-level claim about coordinated fake star campaigns.

What would settle it

Re-run StarScout over the same GHArchive window while sweeping the thresholds (e.g., lockstep n between 20 and 100, rho between 0.3 and 0.7, the 50-star minimum between 20 and 200, and the postprocessing cutoffs around their set values): if the 2024 surge and the 18,617-campaign count do not persist across most of the parameter space, the headline numbers are artifacts of the chosen cutoffs. Independently, take a random sample of accounts flagged as low-activity and check GitHub's deletion, suspension, or IP-level telemetry records: if a large share of flagged accounts remain active for years with organic behavior, the low-activity signature overcounts real users.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that fake-star activity is large, growing, and increasingly malicious: StarScout identifies 6.0 million suspected fake stars across 26,254 repositories before postprocessing, and 3.81 million fake stars within 18,617 repositories judged to be running fake star campaigns, involving 301,000 accounts. The lockstep signature alone accounts for 4.93 million of those stars. The paper further reports that campaigns surged in 2024; that most participating accounts and repositories have highly trivial activity patterns; that the largest category of still-accessible fake-star repositories is spam or phishing, while deleted repositories carry names suggesting pirated software, cryptocurrency bots, and game cheats; and that panel autoregression models show fake stars have a short-lived promotional effect roughly five times weaker than real stars, with a negative long-run association between accumulated fake stars and later real-star gains.

Load-bearing premise

The load-bearing premise is that StarScout's hand-set thresholds — the 50-star minimum, the one-star-plus-one-event low-activity rule, the lockstep parameters (50 accounts, 10 repositories, half the accounts co-starring within 30 days), and the postprocessing cutoffs (monthly fake stars >50, fake ratio >50%, all-time fake ratio >10%) — cleanly separate fake from authentic starring, and the paper itself notes these parameters were chosen by judgment without thorough sensitivity tests.

Editorial extensions

If this is right

  • If the counts are right, raw star counts are a materially distorted popularity and trust signal on GitHub, since in July 2024 about one in six popular repositories (3,499 of those with 50+ monthly stars) had a fake star campaign.
  • Fake stars and malware are linked: 90.42% of campaign repositories were deleted by January 2025, and open coding of the survivors puts spam or phishing as the largest category at roughly 30%, so star inflation should be treated as a supply-chain security indicator.
  • Buying stars is counterproductive for growth hacking: a 1% increase in fake stars is associated with only about 0.07% more real stars one month later (versus 0.36% for real stars), and the accumulated fake-star stock is associated with lower future real-star gains.
  • Current ranking mechanisms already filter much of the fraud: only 78 fake-star campaign repositories (0.42%) ever appeared in GitHub Trending during the study window.
  • Supply-chain exposure is real but narrow: 229 fake-star campaign repositories were linked to 738 packages in registries, most with no dependent packages or repositories, so the main current risk is the malware subset rather than widespread downstream adoption.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The six-million figure is a lower bound by construction, because StarScout only sees public GHArchive events and ignores repositories with fewer than 50 fake stars; if merchants shift to smaller batches or higher-activity accounts, the measured volume would understate the true market.
  • The same lockstep signature could be ported to other popularity surfaces such as npm download counts, PyPI mirrors, model-download counts, or social-media view counts, turning the paper's method into a generic coordinated-burst fraud detector rather than a GitHub-specific study.
  • A testable extension of the promotional-effect result is to follow repositories after GitHub strips their fake stars: if traffic and attention drop sharply, the boost was cosmetic, whereas steady traffic would suggest campaigns do convert into real adoption.
  • The parameter set matches merchant batch sizes observed in 2024; re-running the detector on later data would reveal whether the arms race pushes merchants toward slower, smaller, or more realistic patterns, which is exactly the sensitivity the paper could not run at 20 TB scale.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents StarScout, a scalable BigQuery-based detector for fake GitHub stars, and applies it to all GHArchive event data from July 2019 to December 2024. StarScout combines a low-activity signature (accounts with exactly one WatchEvent and at most one other event) and a lockstep signature (clusters of at least n=50 accounts repeatedly starring at least m=10 repositories within Δt=30 days, with ρ=0.5 density), followed by postprocessing thresholds (monthly fake stars >50, monthly fake-star ratio >50%, and all-time fake-star ratio >10%). The detector identifies 6.0 million suspected fake stars across 26,254 repositories before postprocessing, and 18,617 repositories, 301k accounts, and 3.81 million fake stars after postprocessing. Evaluation against the Stargazer Ghost Network yields 81% repository recall and 76% account recall, and detected repositories/accounts show deletion ratios up to 90%, about 16x higher than random baselines. The measurement study reports a 2024 surge in fake-star campaigns, characterizes the involved accounts and repositories as having trivial activity patterns and skew toward spam/phishing and hyped domains, and uses panel autoregression to argue that fake stars provide only a short-term promotional effect and a long-term negative effect on real star gain.

Significance. If the headline numbers are robust, this is the first systematic, global, longitudinal measurement of fake GitHub stars, and it is a valuable contribution to software supply-chain security and platform moderation research. The paper has several genuine strengths: the detector is implemented at scale on public GHArchive data; recall is anchored to an external, independently documented malware campaign; the deletion-ratio evidence is suggestive and consistently higher than baselines; the artifact, data, and scripts are publicly released; and the responsible-disclosure process is clearly described. The RQ4 result that fake stars have only a short-term promotional effect, if credible, would directly inform practitioner guidance. However, the central counts are point estimates produced by hand-set thresholds that the authors themselves describe as ad hoc, and precision is only assessed indirectly through deletion rates. These issues do not invalidate the study, but they make the absolute prevalence numbers and the 2024-surge conclusion less anchored than the paper's framing suggests.

major comments (3)
  1. [Sections 3.2, 3.3, 3.5]
  2. [Section 3.4]
  3. [Section 4.4, Table 6]
minor comments (5)
  1. [Section 3.4]
  2. [Section 3.4]
  3. [Section 4.2, Table 3]
  4. [Section 4.4, Table 6]
  5. [Appendix C]

Circularity Check

1 steps flagged · score 2.0 of 10

One self-definitional account-activity finding in RQ2; the core detection counts and effect analyses are not circular.

  1. self definitional [Section 4.2 (RQ2: Activity Patterns), 'Duration of Activity' paragraph and Finding 5]
    "Since accounts with the low-activity signature (31.99% of accounts in fake star campaigns) have one day of activity by construction, we focus on the activity duration of accounts with the lockstep signature. ... Finding 5: Most repositories and accounts with fake star campaigns have trivial GitHub activity patterns."

    Accounts qualify as being in fake star campaigns only through the two StarScout signatures: the low-activity signature requires exactly one WatchEvent and at most one additional event, and the lockstep signature requires membership in clusters that repeatedly star repositories in short time windows. Therefore the account-level observation that activity is 'trivial' is entailed by the inclusion criteria: 31.99% of these accounts have one day of activity by construction, and the remaining lockstep accounts are selected because their starring is coordinated and repetitive. Reporting this as an empirical finding is a restatement of the detector's definition rather than an independent discovery.

full rationale

I walked the derivation chain from the two detection signatures to the headline numbers. The 6.0M/18,617/301k/3.81M figures are operational outputs of StarScout at explicitly stated, hand-set thresholds; they are not derived from the conclusions they are used to support. Section 3.4 anchors recall to the external Stargazer Ghost Network ground truth and reports criterion validity through deletion rates; neither is a fitted parameter renamed as a prediction. Section 3.5 openly discloses the ad-hoc heuristics and the absence of thorough sensitivity tests, which is a robustness and validity limitation, not circularity. No load-bearing self-citation or imported uniqueness theorem is used; the only self-citation is the arXiv availability pointer. The one exhibitable circular step is in RQ2: claiming that fake-star-campaign accounts have trivial activity patterns is guaranteed for low-activity accounts and largely built into the lockstep definition. The paper itself acknowledges the low-activity duration is 'by construction.' Because this is a characterization result rather than the central prevalence claim or the RQ4 effect estimate, the overall circularity score is low.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on observable event data and on the assumption that fake-star merchants leave detectable low-activity or lockstep footprints; several thresholds are hand-chosen, which adds free parameters. No new physical or theoretical entities are introduced.

free parameters (6)
  • Minimum fake stars per repository (50-star cutoff) = 50
    Hand-set in Section 3.2 from a gray-literature observation that merchants sell a minimum of 50 stars; used in low-activity, lockstep, and postprocessing steps, so it shapes the 6.0M and 18,617 counts.
  • Low activity signature definition = 1 WatchEvent, at most 1 additional event on the same repository/day
    Section 3.2 defines this heuristic; the paper says it was simplified and tuned to avoid costly API calls, not derived from data or sensitivity analysis.
  • Lockstep parameters (n, m, Δt, ρ) = n=50, m=10, Δt=30 days, ρ=0.5
    Section 3.3 sets these values to find clusters of at least 50 accounts and 10 repositories with at least 25 accounts starring each repo within 30 days; no sensitivity tests reported (Section 3.5).
  • Postprocessing campaign thresholds = monthly fake stars >50, monthly fake ratio >50%, all-time fake ratio >10%
    Section 3.2 defines which repositories are called 'fake star campaigns'; changes here directly change the 18,617 campaign count and all RQ1-RQ4 results.
  • Panel autoregression specification (AR order, log transforms, controls) = AR(2), log-transformed variables, fixed/random effects
    Section 4.4; the short-term positive and long-term negative effects are fitted coefficients from this model family. The authors report consistency across AR(1) to AR(6), which reduces, but does not remove, the dependency on this choice.
  • Number of clusters in k-means for account activity profiles = k=3
    Section 4.2 chooses k=3 by highest Silhouette among k=2..8; affects only the RQ2 cluster description, not the headline fake-star counts.
assumptions (6)
  • domain assumption GHArchive WatchEvents are a complete and correct representation of GitHub starring events.
    All detection runs on GHArchive event data; if WatchEvents are missed, delayed, or include non-star events, every count is affected. Invoked in Sections 3.2-3.3.
  • domain assumption Accounts controlled by fake-star merchants must be either newly created low-activity throwaways or act in synchrony within short windows.
    Stated in Section 3.1 as the core detectability premise; if merchants create realistic active accounts that star slowly and non-coordinately, StarScout will miss them.
  • domain assumption Lockstep star patterns are unlikely to form naturally and are highly correlated with fraud.
    Section 3.2 invokes prior social-network research (CopyCatch, FRAUDAR); this assumption justifies treating all detected lockstep groups as suspicious.
  • domain assumption A high GitHub deletion ratio is a valid criterion for validating fake-star detection.
    Section 3.4 uses deletion status as criterion validity; this assumes GitHub deletes fraudulent or malicious accounts and repositories more often than typical ones.
  • ad hoc to paper Fake-star merchants sell a minimum of 50 stars, so a 50-star threshold captures fake campaigns.
    Section 3.2 and Appendix A rely on gray-market listings; the threshold is not derived from measured data and could exclude smaller campaigns.
  • domain assumption Panel regression controls and log transformations sufficiently account for unobserved heterogeneity.
    Section 4.4 states fixed/random effects handle unobserved heterogeneity but also concedes real causality may hide in exogenous variables; the long-term negative effect is only as credible as this assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware." pith.science (2026). https://pith.science/paper/YSD7NU7I

@misc{pith2026241213459,
  author       = {Pith},
  title        = {Pith review of: Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YSD7NU7I}},
  note         = {Machine review of arXiv:2412.13459}
}
read the original abstract

GitHub, the de facto platform for open-source software development, provides a set of social-media-like features to signal high-quality repositories. Among them, the star count is the most widely used popularity signal, but it is also at risk of being artificially inflated (i.e., faked), decreasing its value as a decision-making signal and posing a security risk to all GitHub users. In this paper, we present a systematic, global, and longitudinal measurement study of fake stars in GitHub. To this end, we build StarScout, a scalable tool able to detect anomalous starring behaviors across all GitHub metadata between 2019 and 2024. Analyzing the data collected using StarScout, we find that: (1) fake-star-related activities have rapidly surged in 2024; (2) the accounts and repositories in fake star campaigns have highly trivial activity patterns; (3) the majority of fake stars are used to promote short-lived phishing malware repositories; the remaining ones are mostly used to promote AI/LLM, blockchain, tool/application, and tutorial/demo repositories; (4) while repositories may have acquired fake stars for growth hacking, fake stars only have a promotion effect in the short term (i.e., less than two months) and become a liability in the long term. Our study has implications for platform moderators, open-source practitioners, and supply chain security researchers.

Figures

Figures reproduced from arXiv: 2412.13459 by the authors.

Figure 1
Figure 1. An example malware repository detected by our [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. A high-level overview of StarScout. a b c x a x Account Repository Legend Starred At <a, x>: 2024-01-08 <b, x>: 2024-01-11 <c, x>: 2024-01-05 Parameters 𝑛 = 3 𝑚 = 3 𝜌 = 2/3 Δ𝑡 = 30 𝑑𝑎𝑦𝑠 a b z y c x a b z y c x Starred At <a, y>: 2024-03-25 <c, y>: 2024-04-02 Starred At <b, z>: 2024-09-13 <c, z>: 2024-09-14 z y [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Under the parameter setting shown above, accounts [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: The percentage (%) of GitHub stars, accounts, and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: The number (#) of GitHub repositories with fake star [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 7
Figure 7. Figure 7: Comparing the distribution of GitHub events be [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: An example phishing repository claiming to be [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: An example phishing repository claiming to be a [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: In repository Uttampatel1/temp1 (197 fake stars as detected by StarScout, 175 fake stars remaining at the time of writing), its README and repository files initially indicated a machine learning repository. The phishing malware link was only added in a later commit si…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Study of Cursorrules Files in GitHub Open Source Projects

    cs.SE 2026-08 conditional novelty 6.0 of 10

    A study of 12,110 .cursorrules files on GitHub shows these AI prompt files appear mostly in small personal projects, focus on code quality, and rarely mention security.

Reference graph

Works this paper leans on

105 extracted references · 74 canonical work pages · cited by 1 Pith paper

  1. [1]

    Retrieved Sept 19, 2024 from https://www.gharchive.org/

    2011.GHArchive. Retrieved Sept 19, 2024 from https://www.gharchive.org/

  2. [2]

    2019.Can we trust GitHub stars?Retrieved July 9, 2024 from https://traefik.io/ blog/can-we-trust-github-stars-e8aa8b6b0baa/

  3. [3]

    Retrieved July 9, 2024 from https://mp.weixin.qq.com/s/q62sSpd6y7_ JGt7mKG0LbQ

    2021.Controversy arises as Ali OceanBase offers gifts for GitHub stars, CTO apologizes. Retrieved July 9, 2024 from https://mp.weixin.qq.com/s/q62sSpd6y7_ JGt7mKG0LbQ

  4. [4]

    Retrieved Sept 24, 2024 from https://juejin.cn/post/7115303814483148807

    2022.GitHub ’Patriotic Package, ’ Actually Just Selling Stars. Retrieved Sept 24, 2024 from https://juejin.cn/post/7115303814483148807

  5. [5]

    Retrieved Nov 14, 2024 from https://www.darkreading.com/cloud-security/ phishing-campaign-targets-pypi-users-to-distribute-malicious-code

    2022.Phishing Campaign Targets PyPI Users to Distribute Malicious Code. Retrieved Nov 14, 2024 from https://www.darkreading.com/cloud-security/ phishing-campaign-targets-pypi-users-to-distribute-malicious-code

  6. [6]

    Retrieved Sep 23, 2024 from https://www.bleepingcomputer.com/news/ security/malicious-microsoft-vscode-extensions-steal-passwords-open- remote-shells/

    2023.Malicious Microsoft VSCode extensions steal passwords, open remote shells. Retrieved Sep 23, 2024 from https://www.bleepingcomputer.com/news/ security/malicious-microsoft-vscode-extensions-steal-passwords-open- remote-shells/

  7. [7]

    Retrieved Apr 28, 2024 from https://checkmarx.com/blog/npm-account-takeover-results- in-crypto-supply-chain-attack/

    2023.NPM Account Takeover Results in Crypto Supply Chain Attack. Retrieved Apr 28, 2024 from https://checkmarx.com/blog/npm-account-takeover-results- in-crypto-supply-chain-attack/

  8. [8]

    Retrieved Mar 11, 2025 from https://stackoverflow.blog/2023/11/08/the-product-approach-to- open-source-communities/

    2023.The product approach to open-source communities. Retrieved Mar 11, 2025 from https://stackoverflow.blog/2023/11/08/the-product-approach-to- open-source-communities/

Show all 105 references
  1. [9]

    2023.Should startups worry about GitHub stars?Retrieved July 9, 2024 from https://technical.ly/software-development/should-startups-worry-about- github-stars/

  2. [10]

    Tracking the Fake GitHub Star Black Market with Dagster, dbt and BigQuery

    2023. Tracking the Fake GitHub Star Black Market with Dagster, dbt and BigQuery. Retrieved July 9, 2024 from https://dagster.io/blog/fake-stars

  3. [11]

    Retrieved Sept 18, 2024 from https://www.wired.com/story/xz-backdoor-everything-you-need- to-know/

    2024.Everything you need to know about the xz util backdoor. Retrieved Sept 18, 2024 from https://www.wired.com/story/xz-backdoor-everything-you-need- to-know/

  4. [12]

    Retrieved Sep 24, 2024 from https: //docs.github.com/en/site-policy/acceptable-use-policies

    2024.GitHub Acceptable Use Policies. Retrieved Sep 24, 2024 from https: //docs.github.com/en/site-policy/acceptable-use-policies

  5. [13]

    Retrieved July 9, 2024 from https://www.wired.com/story/github-stars-black- market-coders-cheat/

    2024.The GitHub Black Market That Helps Coders Cheat the Popularity Contest. Retrieved July 9, 2024 from https://www.wired.com/story/github-stars-black- market-coders-cheat/

  6. [14]

    Retrieved Sept 24, 2024 from https://gitstar.com.cn/

    2024.GitStar. Retrieved Sept 24, 2024 from https://gitstar.com.cn/

  7. [15]

    Retrieved Nov 11, 2024 from https: //blog.phylum.io/the-great-npm-garbage-patch/

    2024.The Great npm Garbage Patch. Retrieved Nov 11, 2024 from https: //blog.phylum.io/the-great-npm-garbage-patch/

  8. [16]

    Retrieved Nov 14, 2024 from https://www.infosecurity-magazine.com/news/malicious-containers- found-docker/

    2024.Millions of Malicious Containers Found on Docker Hub. Retrieved Nov 14, 2024 from https://www.infosecurity-magazine.com/news/malicious-containers- found-docker/

  9. [17]

    Retrieved Sept 18, 2024 from https://www.sonatype.com/state-of-the-software-supply-chain/ open-source-supply-and-demand

    2024.Open Source Supply, Demand, and Security. Retrieved Sept 18, 2024 from https://www.sonatype.com/state-of-the-software-supply-chain/ open-source-supply-and-demand

  10. [18]

    Retrieved Sept 23, 2024 from https://socket.dev/blog/openssf-warns- of-reputation-farming-using-closed-github-issues-and-prs

    2024.OpenSSF Warns of Reputation Farming Leveraging Closed GitHub Issues and PRs. Retrieved Sept 23, 2024 from https://socket.dev/blog/openssf-warns- of-reputation-farming-using-closed-github-issues-and-prs

  11. [19]

    Retrieved Apr 28, 2024 from https: //cran.r-project.org/web/packages/plm

    2024.plm: Linear Models for Panel Data. Retrieved Apr 28, 2024 from https: //cran.r-project.org/web/packages/plm

  12. [20]

    Retrieved Sept 18, 2024 from https://docs

    2024.Rate limits for the REST API. Retrieved Sept 18, 2024 from https://docs. github.com/en/rest/using-the-rest-api/rate-limits-for-the-rest-api

  13. [21]

    Retrieved Sept 17, 2024 from https://research

    2024.Stargazer Ghost Network. Retrieved Sept 17, 2024 from https://research. checkpoint.com/2024/stargazers-ghost-network/

  14. [22]

    Retrieved Oct 28, 2024 from https://github.com/trending

    2024.Trending repositories on GitHub today. Retrieved Oct 28, 2024 from https://github.com/trending

  15. [23]

    Retrieved Nov 8, 2024 from https://www.virustotal.com/

    2024.VirusTotal. Retrieved Nov 8, 2024 from https://www.virustotal.com/

  16. [24]

    Retrieved Feb 25, 2025 from https://ecosyste.ms/

    2025.ecosyste.ms. Retrieved Feb 25, 2025 from https://ecosyste.ms/

  17. [25]

    Retrieved Feb 25, 2025 from https://github.com/ larsbijl/trending_archive

    2025.larsbijl/trending_archive. Retrieved Feb 25, 2025 from https://github.com/ larsbijl/trending_archive

  18. [26]

    Malak Saleh Aljabri, Rachid Zagrouba, Afrah Shaahid, Fatima Alnasser, Asalah Saleh, and Dorieh M. Alomari. 2023. Machine learning-based social media bot detection: A comprehensive literature review.Soc. Netw. Anal. Min.13, 1 (2023), 20

  19. [27]

    Shahram Amini, Michael S Delgado, Daniel J Henderson, and Christopher F Parmeter. 2012. Fixed vs random: The Hausman test four decades later. In Essays in Honor of Jerry Hausman. Vol. 29. Emerald Group Publishing Limited, 479–513

  20. [28]

    David Arthur and Sergei Vassilvitskii. 2007. k-means++: The advantages of careful seeding. InProceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA 2007. SIAM, 1027–1035

  21. [29]

    Andrew Begel, Jan Bosch, and Margaret-Anne D. Storey. 2013. Social Network- ing Meets Software Development: Perspectives from GitHub, MSDN, Stack Exchange, and TopCoder.IEEE Softw.30, 1 (2013), 52–66

  22. [30]

    Alex Beutel, Wanhong Xu, Venkatesan Guruswami, Christopher Palow, and Christos Faloutsos. 2013. CopyCatch: Stopping group attacks by spotting lock- step behavior in social networks. In22nd International World Wide Web Confer- ence, WWW ’13, Rio de Janeiro, Brazil, May 13-17, 2...

  23. [31]

    Enrico Blanzieri and Anton Bryl. 2008. A survey of learning-based techniques of email spam filtering.Artif. Intell. Rev.29, 1 (2008), 63–92

  24. [32]

    René Bohnsack and Meike Malena Liesner. 2019. What the hack? A growth hacking taxonomy and practical applications for firms.Business Horizons62, 6 (2019), 799–818

  25. [33]

    Hudson Borges and Marco Túlio Valente. 2018. What’s in a GitHub Star? Understanding Repository Starring Practices in a Social Coding Platform.J. Syst. Softw.146 (2018), 112–129

  26. [34]

    Donald T Campbell. 1979. Assessing the impact of planned social change. Evaluation and Program Planning2, 1 (1979), 67–90

  27. [35]

    Alan Cao and Brendan Dolan-Gavitt. 2022. What the fork? Finding and analyzing malware in GitHub forks. InProc. of NDSS, Vol. 22

  28. [36]

    Qiang Cao, Xiaowei Yang, Jieqi Yu, and Christopher Palow. 2014. Uncovering Large Groups of Active Malicious Accounts in Online Social Networks. InPro- ceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, Scottsdale, AZ, USA, November 3-7, 2014. A...

  29. [37]

    David Nevado Catalán, Sergio Pastrana, Narseo Vallina-Rodriguez, and Juan Tapiador. 2023. An analysis of fake social media engagement services.Comput. Secur.124 (2023), 103013

  30. [38]

    1996.Psychological Testing and Assessment: An Introduction to Tests and Measurement

    Ronald Jay Cohen, Mark E Swerdlik, and Suzanne M Phillips. 1996.Psychological Testing and Assessment: An Introduction to Tests and Measurement. Mayfield Publishing Co

  31. [39]

    Stefano Cresci, Roberto Di Pietro, Marinella Petrocchi, Angelo Spognardi, and Maurizio Tesconi. 2015. Fame for sale: Efficient detection of fake Twitter fol- lowers.Decis. Support Syst.80 (2015), 56–71

  32. [40]

    Stefano Cresci, Roberto Di Pietro, Marinella Petrocchi, Angelo Spognardi, and Maurizio Tesconi. 2017. The Paradigm-Shift of Social Spambots: Evidence, Theories, and Tools for the Arms Race. InProceedings of the 26th International Conference on World Wide Web Companion. ACM, 963–972

  33. [41]

    Emiliano De Cristofaro, Arik Friedman, Guillaume Jourjon, Mohamed Ali Kâafar, and Muhammad Zubair Shafiq. 2014. Paying for Likes?: Understanding Facebook Like Fraud Using Honeypots. InProceedings of the 2014 Internet Measurement Conference, IMC 2014. ACM, 129–136

  34. [42]

    Alexandre Decan, Tom Mens, and Philippe Grosjean. 2019. An empirical compar- ison of dependency network evolution in seven software packaging ecosystems. Empir. Softw. Eng.24, 1 (2019), 381–416

  35. [43]

    Chrysanthos Dellarocas. 2010. Online reputation systems: How to design one that does what you need.MIT Sloan Management Review51, 3 (2010), 33

  36. [44]

    Kun Du, Hao Yang, Yubao Zhang, Haixin Duan, Haining Wang, Shuang Hao, Zhou Li, and Min Yang. 2020. Understanding Promotion-as-a-Service on GitHub. InACSAC ’20: Annual Computer Security Applications Conference, Virtual Event / Austin, TX, USA, 7-11 December, 2020. ACM, 597–610

  37. [45]

    Elena-Ivona Dumitrescu and Christophe Hurlin. 2012. Testing for Granger non- causality in heterogeneous panels.Economic Modelling29, 4 (2012), 1450–1460

  38. [46]

    2016.Roads and Bridges: The Unseen Labor Behind Our Digital Infrastructure

    Nadia Eghbal. 2016.Roads and Bridges: The Unseen Labor Behind Our Digital Infrastructure. Ford Foundation

  39. [47]

    Robert F Engle, David F Hendry, and Jean-Francois Richard. 1983. Exogeneity. Econometrica: Journal of the Econometric Society(1983), 277–304

  40. [48]

    Herbsleb, and Bogdan Vasilescu

    Hongbo Fang, Daniel Klug, Hemank Lamba, James D. Herbsleb, and Bogdan Vasilescu. 2020. Need for Tweet: How Open Source Developers Talk About Their GitHub Work on Twitter. InMSR ’20: 17th International Conference on Mining Software Repositories. ACM, 322–326

  41. [49]

    This Is Damn Slick!

    Hongbo Fang, Hemank Lamba, James D. Herbsleb, and Bogdan Vasilescu. 2022. "This Is Damn Slick!" Estimating the Impact of Tweets on Open Source Project Popularity and New Contributors. In44th IEEE/ACM 44th International Confer- ence on Software Engineering, ICSE 2022. ACM, 2116–2129

  42. [50]

    Robert Faris, Hal Roberts, Bruce Etling, Nikki Bourassa, Ethan Zuckerman, and Yochai Benkler. 2017. Partisanship, propaganda, and disinformation: Online media and the 2016 US presidential election.Berkman Klein Center Research Publication6 (2017)

  43. [51]

    Chris Grier, Kurt Thomas, Vern Paxson, and Chao Michael Zhang. 2010. @spam: The underground on 140 characters or less. InProceedings of the 17th ACM Conference on Computer and Communications Security, CCS 2010. ACM, 27–37

  44. [52]

    Nick Hagar and Aaron Shaw. 2022. Concentration without cumulative advan- tage: The distribution of news source attention in online communities.Journal of Communication72, 6 (2022), 675–686

  45. [53]

    Christoph Hanck, Martin Arnold, Alexander Gerber, and Martin Schmelzer. 2021. Regression with Panel Data. InIntroduction to Econometrics with R. Universität Duisburg-Essen. https://www.econometrics-with-r.org/10-rwpd.html

  46. [54]

    Hao He, Haoqin Yang, Philipp Burckhardt, Alexandros Kapravelos, Bogdan Vasilescu, and Christian Kästner. 2024. Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware.CoRR abs/2412.13459 (2024). arXiv:2412.13459

  47. [55]

    Li He, Xianzhi Wang, Hongxu Chen, and Guandong Xu. 2022. Online Spam Review Detection: A Survey of Literature.Hum. Centric Intell. Syst.2, 1-2 (2022), 14–30

  48. [56]

    Runzhi He, Hengzhi Ye, and Minghui Zhou. 2024. Revealing the Value of Repository Centrality in Lifespan Prediction of Open Source Software Projects. 12 Six Million (Suspected) Fake Stars on GitHub: A Growing Spiral of Popularity Contests, Spam, and Malware ICSE ’26, April 12–1...

  49. [57]

    Bryan Hooi, Hyun Ah Song, Alex Beutel, Neil Shah, Kijung Shin, and Christos Faloutsos. 2016. FRAUDAR: Bounding Graph Fraud in the Face of Camouflage. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 895–904

  50. [58]

    Zubair Shafiq

    Muhammad Ikram, Lucky Onwuzurike, Shehroze Farooqi, Emiliano De Cristo- faro, Arik Friedman, Guillaume Jourjon, Mohamed Ali Kâafar, and M. Zubair Shafiq. 2017. Measuring, Characterizing, and Detecting Facebook Like Farms. ACM Trans. Priv. Secur.20, 4 (2017), 13:1–13:28

  51. [59]

    Shahedul Huq Khandkar. 2009. Open coding.University of Calgary(2009)

  52. [60]

    2006.Community Building on the Web: Secret Strategies for Successful Online Communities

    Amy Jo Kim. 2006.Community Building on the Web: Secret Strategies for Successful Online Communities. Peachpit Press

  53. [61]

    Simon Koch, David Klein, and Martin Johns. 2024. The Fault in Our Stars: An Analysis of GitHub Stars as an Importance Metric for Web Source Code. In MADWeb: Workshop on Measurements; Attacks and Defenses for the Web

  54. [62]

    Tadayoshi Kohno, Yasemin Acar, and Wulf Loh. 2023. Ethical Frameworks and Computer Security Trolley Problems: Foundations for Conversations. In32nd USENIX Security Symposium, USENIX Security 2023, Anaheim, CA, USA, August 9-11, 2023. USENIX Association, 5145–5162

  55. [63]

    Jouni Kuha. 2004. AIC and BIC: Comparisons of assumptions and performance. Sociological Methods & Research33, 2 (2004), 188–229

  56. [64]

    Piergiorgio Ladisa, Henrik Plate, Matias Martinez, and Olivier Barais. 2023. SoK: Taxonomy of Attacks on Open-Source Software Supply Chains. In44th IEEE Symposium on Security and Privacy, SP 2023. IEEE, 1509–1526

  57. [65]

    J Richard Landis and Gary G Koch. 1977. The measurement of observer agree- ment for categorical data.Biometrics(1977), 159–174

  58. [66]

    Majd Latah. 2020. Detection of malicious social bots: A survey and a refined taxonomy.Expert Syst. Appl.151 (2020), 113383

  59. [67]

    Yuxing Ma, Tapajit Dey, Chris Bogart, Sadika Amreen, Marat Valiev, Adam Tutko, David Kennard, Russell Zaretzki, and Audris Mockus. 2021. World of Code: Enabling a research workflow for mining and analyzing the universe of open source VCS data.Empir. Softw. Eng.26, 2 (2021), 22

  60. [68]

    Dana Movshovitz-Attias, Yair Movshovitz-Attias, Peter Steenkiste, and Christos Faloutsos. 2013. Analysis of the reputation system and user contributions on a question answering website: StackOverflow. InAdvances in Social Networks Analysis and Mining 2013, ASONAM ’13. ACM, 886–893

  61. [69]

    Suhaib Mujahid, Rabe Abdalkareem, and Emad Shihab. 2023. What are the characteristics of highly-selected packages? A case study on the npm ecosystem. J. Syst. Softw.198 (2023), 111588

  62. [70]

    Arjun Mukherjee, Vivek Venkataraman, Bing Liu, and Natalie Glance. 2013. What Yelp fake review filter might be doing?. InProceedings of the International AAAI Conference on Web and Social Media, Vol. 7. 409–418

  63. [71]

    Nuthan Munaiah, Steven Kroh, Craig Cabrey, and Meiyappan Nagappan. 2017. Curating GitHub for engineered software projects.Empir. Softw. Eng.22, 6 (2017), 3219–3253

  64. [72]

    Marc Ohm, Henrik Plate, Arnold Sykosch, and Michael Meier. 2020. Back- stabber’s Knife Collection: A Review of Open Source Software Supply Chain Attacks. InDetection of Intrusions and Malware, and Vulnerability Assessment - 17th International Conference, DIMV A 2020, Lisbon, P...

  65. [73]

    Shashank Pandit, Duen Horng Chau, Samuel Wang, and Christos Faloutsos

  66. [74]

    René Peeters. 2003. The maximum edge biclique problem is NP-complete.Discret. Appl. Math.131, 3 (2003), 651–654

  67. [75]

    Francesco Pierri, Alessandro Artoni, and Stefano Ceri. 2020. Investigating Italian disinformation spreading on Twitter in the context of 2019 European elections. PloS One15, 1 (2020), e0227821

  68. [76]

    Huilian Sophie Qiu, Yucen Lily Li, Hema Susmita Padala, Anita Sarma, and Bogdan Vasilescu. 2019. The Signals that Potential Contributors Look for When Choosing Open-source Projects.Proc. ACM Hum. Comput. Interact.3, CSCW (2019), 122:1–122:29

  69. [77]

    Sanjeev Rao, Anil Kumar Verma, and Tarunpreet Bhatia. 2021. A review on social spam detection: Challenges, open issues, and future directions.Expert Syst. Appl.186 (2021), 115742

  70. [78]

    Marten Risius and Kevin Marc Blasiak. 2024. Shadowbanning: An Opaque Form of Content Moderation.Business & Information Systems Engineering(2024), 1–13

  71. [79]

    Peter J Rousseeuw. 1987. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis.J. Comput. Appl. Math.20 (1987), 53–65

  72. [80]

    William Schueller and Johannes Wachs. 2024. Modeling interconnected social and technical risks in open source software ecosystems.Collective Intelligence 3, 1 (2024), 26339137241231912

  73. [81]

    Indira Sen, Anupama Aggarwal, Shiven Mian, Siddharth Singh, Ponnurangam Kumaraguru, and Anwitaman Datta. 2018. Worth its Weight in Likes: Towards Detecting Fake Likes on Instagram. InProceedings of the 10th ACM Conference on Web Science, WebSci 2018. ACM, 205–209

  74. [82]

    Chengcheng Shao, Giovanni Luca Ciampaglia, Onur Varol, Kai-Cheng Yang, Alessandro Flammini, and Filippo Menczer. 2018. The spread of low-credibility content by social bots.Nature Communications9, 1 (2018), 1–9

  75. [83]

    Jonghyuk Song, Sangho Lee, and Jong Kim. 2015. CrowdTarget: Target-based Detection of Crowdturfing in Online Social Networks. InProceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, Denver, CO, USA, October 12-16, 2015. ACM, 793–804

  76. [84]

    Gianluca Stringhini, Gang Wang, Manuel Egele, Christopher Kruegel, Giovanni Vigna, Haitao Zheng, and Ben Y. Zhao. 2013. Follow the green: Growth and dynamics in Twitter follower markets. InProceedings of the 2013 Internet Mea- surement Conference, IMC 2013. ACM, 163–176

  77. [85]

    Nishat Ara Tania, Md Rayhanul Masud, Md Omar Faruk Rokon, Qian Zhang, and Michalis Faloutsos. 2024. Who is Creating Malware Repositories on GitHub and Why?. InCompanion Proceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, Singapore, May 13-17, 2024. ACM, 955–958

  78. [86]

    Kurt Thomas, Chris Grier, Dawn Song, and Vern Paxson. 2011. Suspended accounts in retrospect: An analysis of Twitter spam. InProceedings of the 11th ACM SIGCOMM Internet Measurement Conference, IMC ’11, Berlin, Germany, November 2-, 2011. ACM, 243–258

  79. [87]

    Asher Trockman, Shurui Zhou, Christian Kästner, and Bogdan Vasilescu. 2018. Adding sparkle to social coding: An empirical study of repository badges in the npmecosystem. InProceedings of the 40th International Conference on Software Engineering, ICSE 2018. ACM, 511–522

  80. [88]

    Herbsleb

    Jason Tsay, Laura Dabbish, and James D. Herbsleb. 2014. Influence of social and technical factors for evaluating contribution in GitHub. In36th International Conference on Software Engineering, ICSE ’14. ACM, 356–366

  81. [89]

    Enrique Larios Vargas, Maurício Finavaro Aniche, Christoph Treude, Magiel Bruntink, and Georgios Gousios. 2020. Selecting third-party libraries: The practitioners’ perspective. InESEC/FSE ’20: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundat...

  82. [90]

    Gummadi, Balachander Krishnamurthy, and Alan Mislove

    Bimal Viswanath, Muhammad Ahmad Bashir, Mark Crovella, Saikat Guha, Krishna P. Gummadi, Balachander Krishnamurthy, and Alan Mislove. 2014. Towards Detecting Anomalous User Behavior in Online Social Networks. In Proceedings of the 23rd USENIX Security Symposium, San Diego, CA, ...

  83. [91]

    Duc-Ly Vu, Ivan Pashchenko, Fabio Massacci, Henrik Plate, and Antonino Sabetta. 2020. Typosquatting and Combosquatting Attacks on the Python Ecosystem. InIEEE European Symposium on Security and Privacy Workshops, EuroS&P Workshops 2020, Genoa, Italy, September 7-11, 2020. IEEE...

  84. [92]

    Gang Wang, Tianyi Wang, Haitao Zheng, and Ben Y. Zhao. 2014. Man vs. Machine: Practical Adversarial Detection of Malicious Crowdsourcing Workers. InProceedings of the 23rd USENIX Security Symposium, San Diego, CA, USA, August 20-22, 2014. USENIX Association, 239–254

  85. [93]

    Gang Wang, Christo Wilson, Xiaohan Zhao, Yibo Zhu, Manish Mohanlal, Haitao Zheng, and Ben Y. Zhao. 2012. Serf and turf: Crowdturfing for fun and profit. InProceedings of the 21st World Wide Web Conference 2012, WWW 2012, Lyon, France, April 16-20, 2012. ACM, 679–688

  86. [94]

    Weippl, Sebastian Schrittwieser, and Sylvi Rennert

    Edgar R. Weippl, Sebastian Schrittwieser, and Sylvi Rennert. 2016. Empirical Research and Research Ethics in Information Security. InInformation Systems Security and Privacy - Second International Conference, ICISSP 2016, Rome, Italy, February 19-21, 2016, Revised Selected Pap...

  87. [95]

    Klemmer, Marcel Fourné, Yasemin Acar, and Sascha Fahl

    Dominik Wermke, Noah Wöhler, Jan H. Klemmer, Marcel Fourné, Yasemin Acar, and Sascha Fahl. 2022. Committed to Trust: A Qualitative Study on Security & Trust in Open Source Software Projects. In43rd IEEE Symposium on Security and Privacy, SP 2022, San Francisco, CA, USA, May 22...

  88. [96]

    Evan D Wolff, KatE M Growley, Maida O Lerner, Matthew B Welling, Michael G Gruden, and Jacob Canter. 2021. Navigating the SolarWinds supply chain attack. Procurement Law.56 (2021), 3

  89. [97]

    Mengying Wu, Geng Hong, Wuyuao Mai, Xinyi Wu, Lei Zhang annd Yingyuan Pu, Huajun Chai, Lingyun Ying, Haixin Duan, and Min Yang. 2025. Exposing the Hidden Layer: Software Repositories in the Service of SEO Manip- ulation. InICSE 2025: The 47th International Confenrece on Softwa...

  90. [98]

    Cao Xiao, David Mandell Freeman, and Theodore Hwa. 2015. Detecting Clusters of Fake Accounts in Online Social Networks. InProceedings of the 8th ACM Workshop on Artificial Intelligence and Security, AISec 2015. ACM, 91–101

  91. [99]

    Kai-Cheng Yang, Onur Varol, Clayton A Davis, Emilio Ferrara, Alessandro Flam- mini, and Filippo Menczer. 2019. Arming the public with artificial intelligence to counter social bots.Human Behavior and Emerging Technologies1, 1 (2019), 48–61

  92. [100]

    Zhao, and Yafei Dai

    Zhi Yang, Christo Wilson, Xiao Wang, Tingting Gao, Ben Y. Zhao, and Yafei Dai

  93. [101]

    Williams

    Nusrat Zahan, Parth Kanakiya, Brian Hambleton, Shohanuzzaman Shohan, and Laurie A. Williams. 2023. OpenSSF Scorecard: On the Path Toward Ecosystem- Wide Automated Security Metrics.IEEE Secur. Priv.21, 6 (2023), 76–88. 13 ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Hao ...

  94. [102]

    Linhong Zhu and Kristina Lerman. 2016. Attention Inequality in Social Media. CoRRabs/1601.07200 (2016). arXiv:1601.07200

  95. [103]

    Markus Zimmermann, Cristian-Alexandru Staicu, Cam Tenny, and Michael Pradel. 2019. Small World with High Risks: A Study of Security Threats in the npm Ecosystem. In28th USENIX Security Symposium, USENIX Security 2019, Santa Clara, CA, USA, August 14-16, 2019. USENIX Associatio...

  96. [2007]

    InProceedings of the 16th International Conference on World Wide Web, WWW 2007, Banff, Alberta, Canada, May 8-12, 2007

    Netprobe: A fast and scalable system for fraud detection in online auction networks. InProceedings of the 16th International Conference on World Wide Web, WWW 2007, Banff, Alberta, Canada, May 8-12, 2007. ACM, 201–210

  97. [2014]

    Uncovering social network Sybils in the wild.ACM Trans. Knowl. Discov. Data8, 1 (2014), 2:1–2:29

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.