{"id":"172f1116-04e6-4e2c-a053-f676c68cfe29","arxiv_id":"2507.18840","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A large-scale descriptive study finds that scholarly posts on Bluesky surged from late 2024 and show higher engagement and more original text than previously reported for X.","lead":"This paper maps how science is discussed on Bluesky by analyzing 2.6 million posts that link to over 530,000 scholarly articles from 2023 to mid-2025. It finds a sharp rise in scholarly activity after November 2024 and reports that Bluesky posts get more likes, reposts, replies, and quotes, and are less likely to just repeat a paper's title, compared with earlier accounts of X.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline gap over X rests on unmatched, methodologically different prior baselines; the comparison is not established by the paper's own data.","rationale":"The paper's descriptive core—temporal surge, disciplinary distribution, language use, and engagement levels on Bluesky—is credible and partially supported by an independent API check in the Appendix. The central advertised contribution, however, is the comparison with X, and that comparison is built on prior-study baselines that differ in time period, data source, platform mechanics, and metric definitions. The reader's weakest_assumption identifies exactly this issue, and it is indeed load-bearing: the abstract's 'substantially higher interaction' and 'greater textual originality' statements depend on it. I did not find a more fundamental flaw in the data pipeline or the internal logic of the analyses; the missing matched baseline is the key soft spot. A matched, same-window, same-metric reanalysis of X data would settle the question, but without it the headline comparative claim should be treated as conditional rather than established. The reader's CONDITIONAL verdict is therefore appropriate, and I would not change it.","tokens_in":15253,"tokens_out":3879,"duration_ms":45115,"concrete_test":"Construct a matched X comparison: for the same 532k OpenAlex DOIs, collect all X posts containing DOI or publisher links from 1 November 2024 to 31 July 2025, retrieve their like, repost, reply, and quote counts, and run the identical cleaning and TF-IDF cosine-similarity pipeline used in Sections 2.3 and 3.4. Then recompute the Section 4.1 percentages (≥10 likes, ≥10 reposts, ≥10 quotes, ≥10 replies, and title-similarity ≥0.9) on this matched sample. If the X values approach or exceed the Bluesky values, the claimed 'substantially higher' engagement and originality collapse; if the values remain an order of magnitude apart under the matched design, the comparison survives this test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing comparative claim in Section 4.1 and the abstract—that Bluesky posts receive 'substantially higher' engagement and show 'greater textual originality' than X—is not supported by the paper's own measurements, because the X side of the comparison is not measured with the same instrument. The Fang et al. (2022) percentages at the ≥10 threshold refer to an earlier period, likely to a different tweet sample (Crossref Event Data versus Altmetric), and to X's then-current mechanics. Bluesky's small, early-adopter, migration-driven user base and chronological feed can inflate per-post engagement rates for reasons unrelated to platform quality. The originality comparison is even more fragile: the paper's 6.3% 'near-verbatim' figure is a TF-IDF cosine-similarity score ≥0.9 computed after stripping @usernames, hashtags, and URLs, whereas the cited X 'title replication rates' (11.8–92.4%) come from older exact or near-duplicate matching studies with heterogeneous definitions. This is not an internal inconsistency, but it is an external-validity gap. Unless the X baseline is recomputed under the same collection, cleaning, and metric definitions and in a contemporaneous window, the conclusion that Bluesky is more interactive and interpretive than X is unproven. The descriptive temporal-surge result is independently corroborated by the Appendix API check and is not the problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents the first large-scale empirical study of how scholarly articles are discussed on Bluesky, analyzing over 2.6 million posts referencing more than 530,000 articles from January 2023 to July 2025. The authors collect post metadata via the Bluesky API (with post URLs from Altmetric) and article metadata from OpenAlex, then describe temporal trends, disciplinary coverage, language use, textual similarity between posts and article titles, and engagement metrics. They report a sharp increase in scholarly activity between November 2024 and January 2025, corroborated by an independent API-based check in the Appendix. The paper's headline comparative claims are that Bluesky posts show 'substantially higher levels of interaction' and 'greater textual originality' than previously reported for X, suggesting that Bluesky supports more interactive and interpretive scholarly communication.","tokens_in":15536,"tokens_out":3922,"duration_ms":39496,"significance":"If the descriptive conclusions hold, this paper is a valuable early map of scholarly communication on a rapidly growing platform. Its strengths include a detailed data-collection pipeline, a large openly described dataset, and an independent corroboration of the temporal surge using a separate retrieval strategy. The paper is likely to be widely cited as a reference point for Bluesky's role in post-Twitter science communication and altmetrics. However, the central comparative claims about engagement and originality rest on previously published X statistics that are not methodologically matched to the Bluesky measurements. The descriptive findings are sound, but the comparative conclusions require substantial additional evidence or a rescoped interpretation.","major_comments":[{"comment":"The claim that Bluesky posts receive 'substantially higher' engagement than X is based on a comparison with previously published X statistics from Fang et al. (2022), but those statistics were collected under different conditions: a different time period, a different data source (Crossref Event Data versus Altmetric), a different sampling approach, and X's then-current platform mechanics. For example, the paper contrasts 48.2% of Bluesky posts receiving at least ten likes with 3.9-7.5% for X, yet without a matched contemporaneous X sample processed through the same pipeline, the gap could reflect differences in user base, post universe, or data-collection coverage rather than a genuine platform difference. The current comparison is not sufficient to support the paper's central 'substantially higher interaction' claim.","section":"Section 4.1 (and Abstract)"},{"comment":"The originality comparison is similarly unmatched. The paper computes a TF-IDF cosine-similarity threshold of 0.9 on cleaned post text and article titles, finding 6.3% near-verbatim posts, and compares this to 'title replication rates' of 11.8-92.4% from prior studies (Didegah et al., 2018; Kumar et al., 2019; Na, 2015; Sergiadis, 2018; Thelwall, Tsou, et al., 2013). Those prior studies use heterogeneous definitions of replication, including exact title matching and near-duplicate detection with different cleaning rules. Because the metrics and definitions are not aligned, the conclusion that Bluesky posts have 'greater textual originality' than X is not established by the data presented in this paper.","section":"Section 4.1 (title replication rates)"}],"minor_comments":[{"comment":"There is an inconsistency in the reported start date of Altmetric's Bluesky tracking: Section 2.1 says Bluesky was added in December 2024 (Kidambi, 2024a), while Section 3.1 says 'Altmetric began collecting Bluesky posts on 24 October 2024' (Altmetric team, 2025). Please reconcile these dates.","section":"Section 2.1 and Section 3.1"},{"comment":"The treatment of 40,593 post-article mentions (1.5%) that precede the official publication date as occurring 'within one month after publication' may inflate the reported 50.7% within-first-week and 65.7% within-first-month figures; consider reporting these cases separately or as negative time lags.","section":"Figure 3 note"},{"comment":"The limitations section does not mention the lack of a matched X baseline for the engagement and originality comparisons, even though this is the most important caveat for the paper's headline claims.","section":"Section 4.2 (Limitations)"},{"comment":"The sample posts in Table 2 are paraphrased and anonymized to protect privacy; this is reasonable, but the paper should clarify whether the paraphrasing was performed before or after computing the cosine similarity scores, since the displayed text may not correspond exactly to the analyzed text.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The descriptive core of the paper is solid and the temporal surge is well corroborated. The main issue is the unmatched X baseline in Section 4.1, which undermines the headline comparative claims. Since the authors include researchers who have published extensive X-based comparisons (e.g., Fang et al., 2022), they are well positioned to either recompute the X baseline under identical definitions or substantially soften the comparative conclusion. I would recommend major revision rather than rejection, as the descriptive findings and dataset are valuable and the comparative claims are potentially fixable with additional analysis or careful rescoping."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what this paper actually gives you: the first large-scale picture of how scholarly articles circulate on Bluesky — 2.6 million posts, 530,000 articles, with a transparent pipeline from Altmetric and OpenAlex, an independent API check in the appendix that corroborates the November 2024–January 2025 surge, and topic-level data on Figshare. For anyone tracking where altmetrics is heading, the descriptive core is genuinely useful and the temporal story is convincing. The finding that most posts appear within a week of publication, and the discipline/language distributions, are exactly the kind of baseline the field needs right now.\n\nThe soft spot is exactly where the reader's stress test lands: the comparative claims in Section 4.1 and the abstract. The “substantially higher engagement” line compares Bluesky's own numbers to X statistics from Fang et al. (2022), collected in a different period, from a different aggregator, under different platform mechanics, and the title-originality comparison is even messier — the 6.3% near-verbatim figure is TF-IDF cosine ≥0.9 after stripping usernames/hashtags/URLs, while the cited X rates come from studies using exact or near-duplicate matching with heterogeneous definitions. That is not an internal inconsistency, but it is an external-validity gap. The paper's own limitations section is honest about data omissions and normalization issues, but it never flags this mismatch. The fix is straightforward: either re-compute X baselines with the same collection and metric definitions in a contemporaneous window, or drop the direct comparison and frame the engagement/originality numbers as descriptive. As written, the headline claim is unproven.\n\nMinor points: no code or full data released (only topic-level data), and the language analysis relies on Bluesky's own `langs` field, which users can override — a small uncertainty but not a distorting one for the main conclusions.\n\nThis paper is for altmetrics researchers, science-of-science people, and anyone thinking about replacing X as a source of social media metrics. It deserves a serious referee: the descriptive contribution is new, the data work is careful, and the comparative weakness is repairable. I would send it to peer review expecting major revision on the X comparison, not rejection.","headline":"A solid descriptive baseline of scholarly posting on Bluesky, but the headline comparison with X rests on unmatched prior-study numbers and should be either recomputed or softened before publication.","tokens_in":16031,"tokens_out":1392,"would_cite":true,"duration_ms":16369,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Bluesky science posts beat X on engagement and original commentary","keywords":["Bluesky","altmetrics","science communication","scholarly communication","social media metrics","OpenAlex","textual similarity","user engagement"],"falsifier":"A matched-baseline study would settle the matter: collect Bluesky and X posts linking to the same set of scholarly articles over the same calendar months, compute the same engagement thresholds and cosine-similarity-to-title scores for both platforms, and compare. If the engagement and originality gaps shrink to near zero under matching, the paper's headline claim that Bluesky discourse is more interactive and more interpretive would be unsupported.","tokens_in":15077,"feed_emoji":"🦋","tokens_out":2718,"duration_ms":28856,"temperature":0.7,"pith_summary":"This paper asks whether Bluesky, the decentralized social platform that many academics migrated to after leaving X, can serve as a credible venue for science communication and as a new source of altmetrics. Analyzing more than 2.6 million Bluesky posts referencing over 530,000 scholarly articles, it finds that scholarly activity on Bluesky surged sharply between November 2024 and January 2025 as researchers left X. The paper's central claim is that Bluesky posts attract substantially more likes, reposts, replies, and quotes than previously reported for X, and that Bluesky posts are more textually original, less often copying article titles verbatim. If true, this means Bluesky is not merely a replacement for X but a place where scientific discourse is more participatory and more interpretive.","feed_headline":"Bluesky science posts beat X on engagement and original commentary","feed_subtitle":"A 2.6-million-post study finds researchers who left X talk about papers more and copy titles less.","key_machinery":"The machinery is a three-stage data-collection and measurement pipeline: first, scholarly post URLs and DOIs were pulled from the Altmetric Explorer; second, post metadata and engagement counts were retrieved through the official Bluesky API; third, article metadata and disciplinary classifications were matched through the OpenAlex database. The two central measurement instruments are complementary cumulative distribution functions of likes, reposts, replies, and quotes, and cosine similarity scores between cleaned post text and the referenced article's title, computed from TF-IDF vectors. The cosine similarity score supplies the operational definition of textual originality, distinguishing posts that merely repeat a title from those that summarize, comment on, or reinterpret the research.","core_discovery":"The paper establishes that scholarly discourse on Bluesky has become both more interactive and more interpretive than comparable discourse previously reported for X. Specifically, 48.2% of scholarly Bluesky posts received at least ten likes and 34.4% at least ten reposts, versus prior X findings of 3.9-7.5% and 1.4-4.4% respectively; replies and quotes on Bluesky likewise occur at rates up to two orders of magnitude higher than earlier X reports. Only 6.3% of Bluesky posts nearly replicate an article title, compared with earlier X title-replication rates ranging from 11.8% to 92.4%. The paper also documents that 50.7% of scholarly Bluesky posts appear within one week of publication, that health, social, and environmental sciences dominate the discourse, and that 91% of posts are in English, leading the authors to position Bluesky as a promising altmetrics source and a stable post-X venue for science communication.","pith_inferences":["If the engagement gap between Bluesky and X were re-tested with matched samples (same articles, same days, same types of users), the gap could shrink; the paper's comparison rests on numbers from different studies, time periods, and sampling designs.","The higher textual originality may partly reflect Bluesky's early-adopter user base of active academics rather than a permanent feature of the platform; as users grow more diverse, title-copying behavior could rise.","A testable extension would be to compare Bluesky engagement rates before and after November 2024 to see whether the migration itself, rather than platform culture, drives the interactive style the paper documents.","Because the dataset excludes posts that discuss articles without including a link, the originality and engagement numbers are estimates for link-bearing discourse only; full-text conversations on Bluesky remain unmeasured."],"forward_implications":["Bluesky can serve as a bona fide altmetrics data source, capturing early post-publication attention in a way that complements traditional citations.","Science communication metrics built on Bluesky may reflect more substantive engagement than metrics built on X, because the platform's posts skew toward commentary and interpretation rather than title repetition.","The sharp post-migration surge indicates that altmetric aggregators tracking Bluesky from late 2024 will record a genuinely growing stream of scholarly attention, not an artifact of indexing alone.","Disciplinary and linguistic biases observed on X, such as English dominance and the prominence of health, social, and environmental sciences, carry over to Bluesky, meaning new-platform altmetrics inherit the same coverage limitations.","Because most posts appear within a week of publication, Bluesky can be used for near-real-time monitoring of research visibility."],"supporting_citations":[{"why":"Documents when Altmetric began collecting Bluesky posts and the enhancement of historical coverage, which the paper uses to rule out an indexing artifact as the main driver of the surge.","marker":"Altmetric team (2025)"},{"why":"Provides an independent Bluesky API-based journal case study corroborating the November 2024-to-January 2025 increase in scholarly activity.","marker":"Arroyo-Machado et al. (2025)"},{"why":"Supplies the prior X benchmark for engagement thresholds (likes, reposts, quotes, replies) that the Bluesky findings are compared against.","marker":"Fang et al. (2022)"},{"why":"One of the earlier X studies used to frame interaction quality and to contextualize the originality comparison.","marker":"Didegah et al. (2018)"},{"why":"Provides the tweet-to-article-title similarity baseline from X that underpins the claim that Bluesky posts are more textually original.","marker":"Thelwall, Tsou, et al. (2013)"},{"why":"Documents the mass migration of researchers and scientific institutions to Bluesky, which the paper links to the observed activity surge.","marker":"Kupferschmidt (2024)"},{"why":"Gives the commercial altmetric observation that Bluesky hosted more posts linked to new research than X on most days in March 2025, supporting the platform's relevance claim.","marker":"Taylor & Areia (2025)"},{"why":"Provides part of the prior X evidence on user motivations and title-repetition rates used in the originality comparison.","marker":"Na (2015)"}],"fun_headline_variants":["Bluesky science chatter: 2.6M posts show higher engagement than X","Exodus to Bluesky: science posts see 10x more likes and replies than X","Researchers on Bluesky: more talk, less copy-paste of titles than on X","Bluesky outshines X in science talk: 2.6M posts analyzed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central comparison assumes that engagement and title-repetition statistics from earlier studies of X are directly comparable to the new Bluesky numbers, even though the two sets of numbers come from different time periods, sampling methods, user populations, and platform mechanics.","fun_headline_variants_meta":{"raw":{"variants":["Bluesky science chatter: 2.6M posts show higher engagement than X","Exodus to Bluesky: science posts see 10x more likes and replies than X","Researchers on Bluesky: more talk, less copy-paste of titles than on X","Bluesky outshines X in science talk: 2.6M posts analyzed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000795,"raw_usage":{"total_tokens":3514,"prompt_tokens":974,"completion_tokens":2540,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":2447}},"tokens_in":590,"tokens_out":2540,"duration_ms":19351,"temperature":1.0,"reasoning_tokens":2447,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:06:42.721390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched-baseline study would settle the matter: collect Bluesky and X posts linking to the same set of scholarly articles over the same calendar months, compute the same engagement thresholds and cosine-similarity-to-title scores for both platforms, and compare. If the engagement and originality gaps shrink to near zero under matching, the paper's headline claim that Bluesky discourse is more interactive and more interpretive would be unsupported.","supporting_citations":[],"review_version":1}