REVIEW 3 major objections 4 minor 42 references
MurkySky: Analyzing News Reliability on Bluesky
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read On Bluesky in summer 2024, only about 2% of rated news links came from unreliable sources.
desk verdict Useful first Bluesky reliability snapshot with a valuable public tool, but the 2% figure is overgeneralized beyond the NewsGuard-rated denominator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is MurkySky itself, a collection and rating pipeline that taps the Bluesky Firehose for posts, likes, and reposts, extracts embedded URLs, and evaluates the linked domains using NewsGuard's source-level reliability ratings, with a score of 60 as the reliable/unreliable threshold. Around this, the paper builds hashtag co-occurrence networks weighted by reliability (Eq. 1), k-core and community detection on the like/repost network for audience segmentation, and log-odds ratio analysis (Eq. 2) to characterize each audience's vocabulary. The source-level rating is the mechanism that turns raw sharing data into the prevalence estimate.
What would settle it
Perform a content-level credibility check on a random sample of the news links collected in the same period and compare the percentage of articles judged unreliable to the 2% source-level figure; a large gap would show the source-level proxy misstates the true prevalence.
Extended reading notes
Core claim
On Bluesky during summer 2024, only about 2% of news links carrying a NewsGuard rating pointed to unreliable sources, making unreliable-source content a small fraction of news sharing. Reliable-source content dominated, and more than 80% of reliable news posts and reposts came from left or left-leaning sources; unreliable-source content also skewed left but was more partisan overall. Hashtag co-occurrence analysis shows that unreliable links concentrated on topics such as Covid-19, the Israel–Hamas conflict, environmental issues, and specific local controversies, rather than spreading uniformly. The authors read this as evidence that Bluesky, at this stage, hosted largely reliable and predominantly left-leaning news discourse.
Load-bearing premise
The entire prevalence estimate rests on accepting NewsGuard's source-level rating, with a cut-off score of 60, as a faithful proxy for whether a linked article is actually unreliable, and on the June–August 2024 firehose capture being complete enough to avoid bias from disconnection lag.
Editorial extensions
If this is right
- If the 2% figure holds, Bluesky in summer 2024 was a relatively low-misinformation environment, supporting claims that it offers an alternative to Twitter/X for reliable news.
- The concentration of unreliable content around specific topics suggests that platform moderation or fact-checking resources can be targeted at those issue clusters.
- The open MurkySky API allows researchers to extend the snapshot into continuous longitudinal monitoring and to test whether growth changes the picture.
- The left-leaning skew of both reliable and unreliable content implies that Bluesky's information ecosystem may be ideologically homogeneous, which could shape future migration and discourse.
- Comparing the 2% with the Iffy Quotient for Twitter/Facebook would put Bluesky's reliability in context across platforms.
Reading between the lines
- We infer that if the source-level rating overestimates unreliability, as the paper itself suspects, then the true share of misleading content may be even lower than 2%; a content-level study of the same posts would settle this.
- The 2% snapshot predates the post-election influx of users; a natural extension is to run MurkySky continuously and test whether the unreliable share rises with platform growth.
- The clustering of unreliable links around German anti-AfD protests and the Israel–Hamas conflict suggests that geopolitical events can create temporary misinformation spikes, and the tool's temporal view could be used to study such spikes.
- The heavy left lean of reliable sources may mean that Bluesky's reputation for reliability is partly a function of its user base's ideology, and a right-leaning influx might change both the reliability mix and the topic clusters.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents MurkySky, a tool that ingests Bluesky's firehose, extracts posts containing URLs, and labels the domains with NewsGuard reliability ratings, alongside descriptive analyses of hashtag co-occurrence, audience segmentation via k-core and word clouds, political-orientation distributions, and rank-frequency plots. Using data from June 14 to August 29, 2024, the authors report that about 2% of NewsGuard-rated news links come from unreliable sources, that both reliable and unreliable source content on the platform skew left-leaning, and that unreliable-source content clusters around specific topics. The paper claims to be the first comprehensive analysis of news reliability on Bluesky and deploys the tool as a public API.
Significance. The study is a timely, purely observational contribution to an under-studied platform. Its strengths are the directness of the measurement against an external rating system (no fitted parameters, no circular benchmark), the public deployment of the collection tool as an API, and a candid limitations section that acknowledges the coarse nature of source-level ratings and the snapshot character of the data. If the headline figure is correctly scoped to NewsGuard-rated links, the paper provides a useful baseline and a reusable monitoring instrument. The main risk is that the unconditional phrasing of the central claim outruns the measured denominator, and the absence of a coverage analysis for unrated domains weakens the generalizability of the 2% figure.
major comments (3)
- [Results, 'Relative Values chart'; Abstract; Conclusion] The 2% figure is explicitly defined as 'the proportion of unreliable links compared to the total number of links with a NewsGuard rating.' The Abstract nevertheless states that unreliable-source content accounts for 'a small fraction of all news-linking posts,' and the Conclusion states that 'unreliable information constitutes only about 2% of the total news content on the platform.' These are unconditional statements, but the denominator excludes every post whose domain lacks a NewsGuard rating, and the political-orientation analysis is further restricted to English-language sources. The Limitations section concedes that NewsGuard 'cannot identify each and every news source, especially those in the long tail of the popularity distribution.' Because no coverage or sensitivity analysis is provided for unrated domains, the unconditional claims are unsupported. Please either rephrase the headline claims to be explicitly conditional on NewsGuard-rated (and, where relevant, English-language) sources, or provide a sensitivity analysis of the 2% estimate under different assumptions about the reliability distribution of unrated domains.
- [Methodology, 'Source Reliability Assessment'; Limitations] The paper asserts that source-level ratings are 'a likely upper bound on the true prevalence of unreliable content' because unreliable sources may republish reliable articles. This reasoning applies to the distinction between source-level and article-level classification of the same link, but it does not bound the unconditional prevalence over all news-linking posts, because the denominator in the reported ratio excludes unrated domains. The upper-bound claim therefore requires an additional assumption about the unrated share. Please scope the upper-bound statement to the NewsGuard-rated subset or combine it with a discussion of how unrated domains are likely to affect the ratio.
- [Methodology, 'The MurkySky Tool'; Limitations] The restricted observation window (June-August 2024) is justified by the claim that the collection script 'ran without interruptions,' but no evidence is provided for the completeness of the firehose capture. A systematic gap in capturing events (e.g., a daily pattern of disconnects or a slowdown during high-volume periods) could bias the numerator and denominator of the prevalence ratio in opposite directions. Please add a completeness validation, for example by comparing captured event counts with platform-reported totals or by reporting the fraction of time the connection was active and the backlog at resumption.
minor comments (4)
- [Eq. (1)] The description of the edge weight says it is 'the difference between the frequency of reliable-source link-posts to that of unreliable-source ones,' but the formula (WUT − WT)/(WUT + WT) subtracts reliable counts from unreliable counts. The sign should be reconciled with the description.
- [Eq. (2)] The log-odds formula should be checked against Monroe et al. (2008); the denominator terms as written (ni + a0 − yiw + aw) are not the standard form, and the prior mass terms appear to require a0 − aw rather than a0 − yiw + aw. If the formula is correct as implemented, a reference to the specific variant would help.
- [Table 1] The 'Unreliable Domains' column includes msnbc.com and dailykos.com; a brief note explaining how NewsGuard's criteria place these outlets below the reliability threshold would prevent confusion for readers.
- [Figures 3 and 4] The qualitative word clouds and orientation-prevalence charts are not accompanied by underlying counts (e.g., number of users per modularity class, number of domains or links per orientation bin). Including a small table would make the descriptive claims auditable.
Circularity Check
No circularity: all load-bearing claims are direct measurements against external ratings, not reductions to the paper's own inputs.
full rationale
This paper makes no derivation claim that reduces to its inputs. The 2% prevalence figure is a descriptive statistic: the Results section explicitly defines it as 'the proportion of unreliable links compared to the total number of links with a NewsGuard rating,' and the underlying counts come from the Bluesky firehose labeled by the external NewsGuard rubric (threshold 60), not from a model fitted to the data. The hashtag edge weight (Eq. 1), log-odds ratio (Eq. 2), k-core segmentation, and rank-frequency plots are all standard descriptive transformations of the collected posts; none of them is used to produce the headline prevalence, nor do they smuggle the outcome into a fitted parameter. The only in-paper citations to the authors' own prior work (Shao et al. 2016, 2018a, 2018b) are background references to misinformation measurement and spread; the central measurement does not rest on a self-citation chain or an imported uniqueness theorem. The Limitations section concedes that NewsGuard coverage is incomplete and that source-level ratings are coarse, and the abstract's wording ('small fraction of all news-linking posts') does go beyond the rated denominator; that is a scope or generalizability concern, not a circularity, because the conditional 2% estimate is honestly reported and externally anchored. No fitted input is relabeled as a prediction, and no known result is renamed as organization. Accordingly, the paper scores 0 on circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption NewsGuard source-level ratings with threshold 60 correspond to the reliability of linked articles.
- domain assumption The Bluesky firehose stream was uninterrupted and representative during the June-August 2024 window.
- domain assumption Restricting political-orientation analysis to English-language sources does not materially bias the ideological distribution.
- domain assumption The Media Bias Fact Check over AllSides over NewsGuard hierarchy yields valid political labels when ratings conflict.
- domain assumption A single left-right political orientation axis captures the relevant ideological dimension of news sources.
Cite this review
Pith. "Pith review of MurkySky: Analyzing News Reliability on Bluesky." pith.science (2026). https://pith.science/paper/KUCU2HJL
@misc{pith2026250110557,
author = {Pith},
title = {Pith review of: MurkySky: Analyzing News Reliability on Bluesky},
year = {2026},
howpublished = {\url{https://pith.science/paper/KUCU2HJL}},
note = {Machine review of arXiv:2501.10557}
}
read the original abstract
Bluesky has recently emerged as a lively competitor to Twitter/X for a platform for public discourse and news sharing. Most of the research on Bluesky so far has focused on characterizing its adoption due to migration. There has been less interest on characterizing the properties of Bluesky as a platform for news sharing and discussion, and in particular the prevalence of unreliable information on it. To fill this gap, this research provides the first comprehensive analysis of news reliability on Bluesky. We introduce MurkySky, a public tool to track the prevalence of content from unreliable news sources on Bluesky. Using firehose data from the summer of 2024, we find that on Bluesky reliable-source news content is prevalent, and largely originating from left-leaning sources. Content from unreliable news sources, while accounting for a small fraction of all news-linking posts, tends to originate from more partisan sources, but largely reflects the left-leaning skew of the platform. Analysis of the language and hashtags used in news-linking posts shows that unreliable-source content concentrates on specific topics of discussion.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Acemoglu, D.; Ozdaglar, A.; and ParandehGheibi, A. 2010. Spread of (mis)information in social networks. Games and Economic Behavior, 70(2): 194--227
work page 2010
-
[4]
Anderson, A.; Huttenlocher, D.; Kleinberg, J.; Leskovec, J.; and Tiwari, M. 2015. Global Diffusion via Cascading Invitations: Structure, Growth, and Homophily. In Proceedings of the 24th International Conference on World Wide Web, WWW '15, 66–76. Republic and Canton of Geneva, CHE: International World Wide Web Conferences Steering Committee. ISBN 9781450334693
work page 2015
-
[5]
Baumgartner, J.; Zannettou, S.; Keegan, B.; Squire, M.; and Blackburn, J. 2020. The Pushshift Reddit Dataset. In Proceedings of the International AAAI Conference on Web and Social Media, volume 14, 830--839
work page 2020
-
[6]
Bianchi, J.; Pratelli, M.; Petrocchi, M.; and Pinelli, F. 2024. Evaluating Trustworthiness of Online News Publishers via Article Classification. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing, SAC '24, 671--678. New York, NY, USA: Association for Computing Machinery. ISBN 9798400702433
work page 2024
-
[7]
D.; Guillaume, J.-L.; Lambiotte, R.; and Lefebvre, E
Blondel, V. D.; Guillaume, J.-L.; Lambiotte, R.; and Lefebvre, E. 2008. Fast unfolding of communities in large networks. Journal of Statistical Mechanics: Theory and Experiment, 2008(10): P10008
work page 2008
-
[8]
Cohen, R.; Moffatt, K.; Ghenai, A.; Yang, A.; Corwin, M.; Lin, G.; Zhao, R.; Ji, Y.; Parmentier, A.; P’ng, J.; Tan, W.; and Gray, L. 2020. Addressing misinformation in online social networks: Diverse platforms and the potential of multiagent trust modeling. Information (Switzerland), 11(11): 1--40
work page 2020
Show all 42 references
-
[9]
Cristelli, M.; Batty, M.; and Pietronero, L. 2012. There is More than a Power Law in Zipf. Scientific Reports, 2: 812
2012
-
[10]
Datta, A.; et al. 2010. Decentralized online social networks. In Handbook of social network technologies and applications, 349--378. Springer
2010
-
[11]
A.; Birnholtz, J.; and Hancock, J
DeVito, M. A.; Birnholtz, J.; and Hancock, J. T. 2017. Platforms, People, and Perception: Using Affordances to Understand Self-Presentation on Social Media. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing, CSCW '17, 740--75...
2017
-
[12]
Doctorow, C. 2023. Social Quitting. Locus, 90(1): 29. Commentary
2023
-
[13]
FORCE11 . 2020. The FAIR Data principles. https://force11.org/info/the-fair-data-principles/
2020
-
[14]
W.; Wallach, H.; Iii, H
Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J. W.; Wallach, H.; Iii, H. D.; and Crawford, K. 2021. Datasheets for datasets. Communications of the ACM, 64(12): 86--92
2021
-
[15]
Goel, S.; Broder, A.; Gabrilovich, E.; and Pang, B. 2010. Anatomy of the Long Tail: Ordinary People with Extraordinary Tastes. In Proceedings of the Third ACM International Conference on Web Search and Data Mining, WSDM ’10, 201–210. New York, NY, USA: Association for Computin...
2010
-
[16]
Green, J.; McCabe, S.; Shugars, S.; Chwe, H.; Horgan, L.; Cao, S.; and Lazer, D. 2024. Curation Bubbles. American Political Science Review. Forthcoming
2024
-
[17]
B.; Castro, I.; Raman, A.; Sastry, N.; and Tyson, G
He, J.; Zia, H. B.; Castro, I.; Raman, A.; Sastry, N.; and Tyson, G. 2023. Flocking to Mastodon: Tracking the Great Twitter Migration. In Proceedings of the 2023 ACM on Internet Measurement Conference, IMC '23, 111–123. New York, NY, USA: Association for Computing Machinery. I...
2023
-
[18]
Hogg, L.; DiResta, R.; Fukuyama, F.; Reisman, R.; Keller, D.; Ovadya, A.; Thorburn, L.; Stray, J.; and Mathur, S. 2024. Shaping the Future of Social Media with Middleware . Technical report, CoRR. ArXiv:2412.10283 [cs]
2024 arXiv
-
[19]
Hou, A.; and Shiau, W.-L. 2020. Understanding Facebook to Instagram migration: a push-pull migration model perspective. Information Technology & People, 33(1): 272--295
2020
-
[20]
R.; and Liu, H
Jeong, U.; Sheth, P.; Tahir, A.; Alatawi, F.; Bernard, H. R.; and Liu, H. 2024. Exploring Platform Migration Patterns between Twitter and Mastodon: A User Behavior Study. Proceedings of the International AAAI Conference on Web and Social Media, 18(1): 738--750
2024
-
[21]
Kates, S.; Tucker, J.; Nagler, J.; and Bonneau, R. 2021. The Times They Are Rarely A-Changin’: Circadian Regularities in Social Media Use. Journal of Quantitative Description: Digital Media, 1
2021
-
[22]
Kleppmann, M.; Frazee, P.; Gold, J.; Graber, J.; Holmgren, D.; Ivy, D.; Johnson, J.; Newbold, B.; and Volpert, J. 2024. Bluesky and the AT Protocol: Usable Decentralized Social Media. In Proceedings of the ACM Conext-2024 Workshop on the Decentralization of the Internet, DIN '...
2024
-
[23]
M.; and Tagarelli, A
La Cava, L.; Aiello, L. M.; and Tagarelli, A. 2023. Drivers of social influence in the Twitter migration to Mastodon . Scientific Reports, 13(1): 21626
2023
-
[24]
La Cava, L.; Greco, S.; and Tagarelli, A. 2021. Understanding the growth of the Fediverse through the lens of Mastodon. Applied Network Science, 6(1): 1--15
2021
-
[25]
Lazer, D. M. J.; Baum, M. A.; Benkler, Y.; Berinsky, A. J.; Greenhill, K. M.; Menczer, F.; Metzger, M. J.; Nyhan, B.; Pennycook, G.; Rothschild, D.; Schudson, M.; Sloman, S. A.; Sunstein, C. R.; Thorson, E. A.; Watts, D. J.; and Zittrain, J. L. 2018. The science of fake news. ...
2018
-
[26]
Lemmer-Webber, C.; Tallon, J.; Shepherd, E.; Guy, A.; and Prodromou, E. 2018. ActivityPub . W3C recommendation, World Wide Web Consortium, Wakefield, MA 01880, USA
2018
-
[27]
G.; and Pennycook, G
Lin, H.; Lasser, J.; Lewandowsky, S.; Cole, R.; Gully, A.; Rand, D. G.; and Pennycook, G. 2023. High level of correspondence across different news domain quality rating sets. PNAS Nexus, 2(9)
2023
-
[28]
Mendelsohn, J.; Le Bras, R.; Choi, Y.; and Sap, M. 2023. From Dogwhistles to Bullhorns : Unveiling Coded Rhetoric with Language Models . In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Computational Linguisti...
2023
-
[29]
Monroe, B.; Colaresi, M.; and Quinn, K. 2008. Fightin’ Words: Lexical Feature Selection and Evaluation for Identifying the Content of Political Conflict. Political Analysis, 16(4): 372--403
2008
-
[30]
Murtfeldt, R.; Alterman, N.; Kahveci, I.; and West, J. D. 2024. RIP Twitter API : A eulogy to its vast research contributions. Technical report, CoRR. ArXiv:2404.07340 [cs]
2024
-
[31]
NewsGuard . 2024. NewsGuard Rating Process and Criteria. Accessed: January 10, 2024
2024
-
[32]
Quelle, D.; and Bovet, A. 2024. Bluesky: Network Topology, Polarisation, and Algorithmic Curation. Technical report, CoRR. ArXiv:2405.17571 [cs]
2024 arXiv
-
[33]
Raman, A.; Joglekar, S.; de Cristofaro, E.; Sastry, N.; and Tyson, G. 2019. Challenges in the decentralised web: The Mastodon case. In Proceedings of the ACM SIGCOMM Internet Measurement Conference (IMC), 217--229. ACM
2019
-
[34]
Resnick, P.; Ovadya, A.; and Gilchrist, G. 2023. Iffy Quotient: A Platform Health Metric for Misinformation. White Paper v3, University of Michigan Center for Social Media Responsibility, Ann Arbor, MI, USA
2023
-
[35]
Reuben, M.; Friedland, L.; Puzis, R.; and Grinberg, N. 2024. Leveraging Exposure Networks for Detecting Fake News Sources . In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , KDD '24, 5635--5646. New York, NY, USA: Association for Computi...
2024
-
[36]
Ribeiro, B.; and Faloutsos, C. 2015. Modeling Website Popularity Competition in the Attention - Activity Marketplace . In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining , WSDM '15, 389--398. New York, NY, USA: Association for Computing Mac...
2015
-
[37]
Schneider, N. 2019. Decentralization: an incomplete ambition. Journal of Cultural Economy, 12(4): 265--285
2019
-
[38]
L.; Flammini, A.; and Menczer, F
Shao, C.; Ciampaglia, G. L.; Flammini, A.; and Menczer, F. 2016. Hoaxy: A Platform for Tracking Online Misinformation. In Proceedings of the 25 th International Conference Companion on World Wide Web , WWW '16 Companion, 745--750. Republic and Canton of Geneva, Switzerland: In...
2016
-
[39]
L.; Varol, O.; Yang, K.; Flammini, A.; and Menczer, F
Shao, C.; Ciampaglia, G. L.; Varol, O.; Yang, K.; Flammini, A.; and Menczer, F. 2018 a . The spread of low-credibility content by social bots. Nature Communications, 9(1): 4787
2018
-
[40]
Shao, C.; Hui, P.-M.; Wang, L.; Jiang, X.; Flammini, A.; Menczer, F.; and Ciampaglia, G. L. 2018 b . Anatomy of an online misinformation network. PLOS ONE, 13(4): 1--23
2018
-
[41]
Don ’t research us
Wähner, M.; Deubel, A.; Breuer, J.; and Weller, K. 2024. “ Don ’t research us”— How Mastodon instance rules connect to research ethics. Publizistik, 69(3): 357--380
2024
-
[42]
Zhang, Z.; Zhao, J.; Wang, G.; Johnston, S.-K.; Chalhoub, G.; Ross, T.; Liu, D.; Tinsman, C.; Zhao, R.; Van Kleek, M.; and Shadbolt, N. 2024. Trouble in Paradise? Understanding Mastodon Admin's Motivations, Experiences, and Challenges Running Decentralised Social Media. Proc. ...
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.