{"id":"26bf481b-640f-4ecf-a52c-657d0da5c15c","arxiv_id":"2505.10266","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"AI-generated misinformation on X, identified via Community Notes and an LLM, differs from conventional misinformation in content, source accounts, and virality, but the evidence is weakened by selection and measurement issues.","lead":"This study compares 91,452 misinformation posts flagged on X's Community Notes, using an LLM to label which ones are AI-generated. It reports that AI-generated misinformation is more entertaining, more viral, and more often posted by smaller accounts, which matters for platform moderation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The AI-generation label is read off the Community Note text, and the validation reuses that same note text; the virality coefficients in Eq. 1 may therefore measure note wording rather than AI-generated content.","rationale":"The reader's weakest assumption correctly identifies the circular validation: the AI label is derived from Community Note text and validated with the same note text visible to raters. My independent reading of the methods confirms this, with the additional detail that the LLM prompt explicitly instructs classification 'based on the community note provided,' and the human validation protocol states raters reviewed 'the corresponding textual Community Note.' This is not merely a measurement-quality nit; it is the construct validity of the independent variable. The virality regression (Eq. 1) controls for media type and account characteristics, but if the treatment indicator is a proxy for note wording, the coefficients are biased in an unknown direction. For instance, notes that say 'AI-generated' may be written for posts that are already attention-grabbing, or note-writers may label only certain types of AI content; selection and measurement are entangled. The reported validation statistics do not resolve this: 75% agreement on AI posts and 22% on non-AI posts show only that some AI-related signal is present, but because raters saw the note, the signal could be the note's wording rather than the media content. The low Fleiss kappa (0.322) also indicates substantial rater disagreement. Additionally, I noticed an internal inconsistency between the reported prevalence of 5.06% AI posts and Table S1's mean of 0.12 for the AI-generated indicator, which suggests the analysis may have used different labeling criteria or data versions. This reinforces that the central claim is not robustly supported. I therefore see no reason to alter the reader's REJECT verdict; the paper may still be a useful descriptive dataset, but the headline virality claim is not established by the present design.","tokens_in":14803,"tokens_out":4215,"duration_ms":47805,"concrete_test":"Re-run the validation study on the same 400 posts with the Community Note text withheld: show raters only the original post and attached media. If raters' AI/non-AI discrimination (e.g., difference in mean ratings or AUC) collapses toward chance, the AI label is an artifact of note text and the Eq. 1 virality coefficients cannot be attributed to AI generation. As a secondary check, compute a keyword-based label from whether the note contains terms like 'AI-generated', 'deepfake', or 'synthetic'; if the LLM label and keyword label are nearly identical, this confirms the label is driven by note wording.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that AI-generated misinformation is more viral depends on the validity of the AI label. In 'Identification of AI-generated misinformation', the LLM is instructed to identify 'whether you expect the original post to contain AI-generated content based on the community note provided' — the classifier sees only the note, not the post or its media. The validation then asks human raters to judge the same posts while showing them the same Community Note. High agreement between the LLM and raters is thus largely agreement that notes mentioning AI are recognized as mentioning AI; it does not establish that the underlying post contains AI-generated content. The reported low interrater agreement (Fleiss kappa = 0.322) and the 22% rate at which raters labeled non-AI posts as AI reinforce the concern. If the 'AI-generated' indicator is essentially a note-text feature, then the 10.81% retweet, 34.16% like, and 10.32% impression advantages in Eq. 1 may reflect properties of notes that mention AI (e.g., novelty, controversy, or visibility of deepfakes) rather than properties of AI-generated content itself. This is load-bearing because every downstream comparison — virality, topics, sentiment, believability, account characteristics — uses the same label. A further inconsistency compounds the problem: the text reports 4,577 AI posts (5.06%) while Table S1 reports an AI-generated mean of 0.12 (12%) for the same N=91,452, suggesting ambiguity in how the label was computed or applied.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a large-scale observational study of misinformation posts flagged on X's Community Notes platform between January 2023 and January 2025. The authors use GPT-4-turbo to classify posts as AI-generated based on the text of the associated Community Note, and then compare AI-generated and non-AI-generated misleading posts across four research questions: content characteristics (sentiment, topic, media type), author account characteristics, virality (retweets, likes, impressions), and perceived believability/harmfulness. The central claim is that AI-generated misinformation is more viral than other misinformation: the negative binomial regressions in Eq. (1) report 10.81% more retweets, 34.16% more likes, and 10.32% more impressions for AI-generated posts after controlling for media type, account characteristics, and month-year fixed effects. Secondary findings are that AI-generated misinformation tends to be more positive in sentiment, more entertainment-focused, more likely to come from smaller accounts, and slightly less believable and harmful. The paper also provides an LLM-based annotation of a 3,000-post subsample for sentiment, topic, believability, and harmfulness.","tokens_in":15092,"tokens_out":5138,"duration_ms":50205,"significance":"If the identification strategy were valid, this would be one of the first large-scale empirical accounts of real-world AI-generated misinformation, with direct implications for platform moderation and misinformation research. The paper contributes a sizable public dataset, transparent prompts, and regression specifications with fixed effects and standard robustness checks. The main empirical claims, however, rest entirely on the validity of the AI-generated label, and that label is currently derived from Community Note text in a way that is not independently validated against the posts' actual content. The internal inconsistencies in the reported numbers further reduce confidence. The significance of the findings is therefore conditional on resolving these measurement and reporting issues.","major_comments":[{"comment":"The AI-generated indicator is constructed by an LLM that sees only the Community Note text, not the post's media or original content ('Identify whether you expect the original post to contain AI-generated content based on the community note provided'). The validation study then presents human raters with the same Community Note plus the post, so agreement between the LLM and raters largely demonstrates that notes mentioning AI are recognized as mentioning AI, not that the underlying posts contain AI-generated media. The reported Fleiss kappa of 0.322 (fair) and the 22% rate at which non-AI posts were rated as AI further indicate label noise. This issue is load-bearing for every downstream comparison, including the virality coefficients in Eq. (1). The authors should re-validate the label with raters who judge the post's media and text without seeing the Community Note, and ideally cross-check a sample against external provenance information.","section":"Identification of AI-generated misinformation / Validation"},{"comment":"The sample consists only of posts that received a helpful Community Note, which requires users to write a note and other users to rate it as helpful. Visible and controversial posts are more likely to enter the sample, and the probability of being flagged may differ systematically between AI-generated and non-AI-generated posts. If AI-generated posts are disproportionately flagged when they are already viral, the regression coefficient on AIGenerated in Eq. (1) is biased upward. The Limitations section acknowledges the selection problem qualitatively but does not quantify it. A bounding exercise, a selection model, or a sensitivity analysis based on note-writing rates should be added before the virality claim can be accepted.","section":"Data source / Virality (RQ3, Eq. 1)"},{"comment":"The reported number of AI-generated posts is internally inconsistent. The text in Section 'Empirical Analysis' states that the dataset includes 4,577 AI-generated posts (5.06%), which is consistent with N=91,452 only if the AI variable is coded as 0.050. Yet Table S1 reports the mean of AI-generated as 0.12 (12%) for the same N=91,452. Additionally, the abstract given at the start of the manuscript says 82,076 misleading posts, whereas the full text and Table S1 use 91,452. These discrepancies must be resolved because the prevalence of AI-generated misinformation is a central descriptive finding and the AI/non-AI split enters every subsequent analysis.","section":"Empirical Analysis (RQ1) / Table S1"},{"comment":"The sentence in Section 'Author characteristics' reads 'AI-generated misleading posts tend to originate from accounts with significantly more followers (950,660 vs. 585,671),' but Table S1 reports the opposite: the mean follower count is 585,671 for AI-generated posts and 950,660 for non-AI-generated posts. The Discussion and abstract explicitly rely on the 'smaller accounts' direction for RQ2, so this reversal is not a harmless typo; it directly contradicts the paper's own summary statistics and must be corrected.","section":"Author characteristics (RQ2)"}],"minor_comments":[{"comment":"The model checks list is numbered '(1) ... (3) ... (iii)', which appears to be a typographical inconsistency; the items should be numbered consistently.","section":"Model checks"},{"comment":"The final sentence of the Implications section ends with an incomplete phrase 'types of misleading information.' that appears to be a leftover from an earlier draft and should be removed.","section":"Implications"},{"comment":"The Introduction contains the phrase 'addresses this gap with by characterizing'; the word 'with' should be deleted.","section":"Introduction"},{"comment":"The means for Believability (0.34) and Harmfulness (0.41) in Table S1 appear to be binary-coded conversions, but the main text reports three-level distributions (low/medium/high) for the same LLM-annotated sub-sample (N=3,000). The coding scheme should be clarified so that the table and the text are reconcilable.","section":"Table S1"},{"comment":"The footnote about the Community Notes data download is referenced as '1Available via ...' with no spacing after the footnote marker; also, the URL is not rendered as a clickable link in the text.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The reader's report recommends rejection, but I see the core problem as a measurement validity issue that is, in principle, fixable within a major revision. The paper needs (1) an independent validation of the AI-generated label using raters who do not see the Community Note, (2) sensitivity analyses addressing selection into the Community Notes sample, and (3) correction of the internal inconsistencies in the prevalence and follower-count numbers. If the authors cannot provide a valid label after such a revision, the central virality claim and the descriptive comparisons would not be supported. I therefore recommend major_revision rather than immediate rejection, but the revision must be thorough and the re-validation results must be reported transparently."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this about the paper: it is the first large-scale attempt to characterize AI-generated misinformation circulating on X via Community Notes, and the dataset is real (91,452 posts, two years). The descriptive parts—topics, sentiment, believability—are plausible starting points. But the central claim, that AI-generated misinformation is more viral, is not supported by the design as currently written, and there are a couple of internal errors that will need fixing.\n\nWhat is genuinely new: using Community Notes plus an LLM to label AI-generated posts at scale. The prompts are included, the validation is described, and the authors acknowledge selection bias in the limitations. That alone is worth a referee's time.\n\nWhere it breaks: the LLM is asked to infer AI generation 'based on the community note provided'—it never sees the post's media. The human validation then shows the raters the same note. High LLM-human agreement mostly tells you the note mentions AI, not that the post contains AI media. The low Fleiss kappa (0.322) and 22% false-positive rate among non-AI posts reinforce that. So the virality coefficients (10.8% more retweets, 34.2% more likes, 10.3% more impressions) may be measuring properties of notes that mention AI, or the selection of posts that receive such notes, rather than AI generation itself.\n\nThere are also concrete inconsistencies. The prose in 'Author characteristics' says AI posts come from accounts with more followers, but the abstract and Table S1 say the opposite (AI accounts are smaller). The text reports 4,577 AI posts (5.06%), while Table S1 shows a mean of 0.12 (12%) for AI-generated. And the misinformation exposure means in Table S1 look impossible given the reported t-test. These are fixable, but they matter.\n\nIs the paper a waste? No. The dataset and the descriptive findings on entertainment/sentiment are a useful resource, and the 'smaller accounts' implication for moderation is practically relevant if it holds up. But the headline virality claim needs either a better label (e.g., analyzing the media, not the note) or a much more cautious interpretation.\n\nMy take: it deserves serious peer review, but not acceptance in this form. I'd send it to referees with instructions to focus on label validity. For my own reading group, I'd maybe include it to discuss the label circularity.","headline":"A genuinely new dataset for studying AI misinformation in the wild, but the AI label is read off the Community Note rather than the post, and internal inconsistencies undercut the headline virality claim.","tokens_in":15588,"tokens_out":4166,"would_cite":false,"duration_ms":37639,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-generated misinformation on X is measurably more viral than other misleading posts, gaining about 11% more retweets, 34% more likes, and 10% more views after controls for account size and content.","keywords":["AI-generated misinformation","deepfakes","Community Notes","virality","misinformation characteristics","X social media","negative binomial regression","crowdsourced fact-checking"],"falsifier":"Inspect the original images or videos of a random sample of posts the model labels AI-generated and check provenance or run forensic deepfake detectors; if most contain no detectable AI-created media, the virality gap is an artifact of note wording rather than of AI generation.","tokens_in":14619,"feed_emoji":"🤖","tokens_out":9147,"duration_ms":81238,"temperature":0.7,"pith_summary":"This paper tries to establish that AI-generated misinformation on social media is not just a hypothetical risk but a measurable class of content with its own spread patterns. Using crowd-flagged misleading posts on X, it compares more than 90,000 posts, about 5% of them identified as AI-generated, and finds that the AI-generated ones are reliably more viral: roughly 11% more retweets, 34% more likes, and 10% more impressions after controlling for media type and account characteristics. It also reports that AI-generated misinformation skews toward entertainment and positive sentiment, tends to come from smaller but older and more conservative accounts, and is rated only slightly less believable and harmful than ordinary misinformation. A sympathetic reader would care because these patterns imply that the AI-generated nature of a post is an independent driver of engagement, not a proxy for account reach or content negativity.","feed_headline":"AI-made fake posts outpace ordinary misinformation on X","feed_subtitle":"In a 91,452-post X sample, AI fakes got 10.8% more retweets and 34.2% more likes than ordinary fakes.","key_machinery":"The central object is the binary indicator AIGenerated_i, obtained by giving a large language model each Community Note and asking whether the note indicates AI-generated content; a balanced subset of 3,000 posts is further annotated for sentiment, topic, harmfulness, and believability. That indicator is the key regressor in three count-regression models (negative binomial) explaining retweets, likes, and impressions, with media type and account attributes as controls and month-year fixed effects. The size and significance of its coefficient is what supports the virality claim.","core_discovery":"Across 91,452 misleading posts flagged on X's Community Notes from January 2023 to January 2025, the paper identifies a subset as AI-generated by having a language model read each Community Note and decide whether it refers to AI-created content. Its central claim is that this subset behaves differently enough to be treated as its own category. AI-generated misleading posts receive 10.81% more retweets, 34.16% more likes, and 10.32% more impressions than non-AI misleading posts in negative binomial regressions that control for media type, follower and followee counts, account age, verification status, and month-year fixed effects. They are also 1.33 times more likely to carry media, more often categorized as entertainment with positive sentiment, more likely to originate from smaller yet older and more conservative accounts, and slightly less believable and harmful than conventional misinformation.","pith_inferences":["Because AI-generation is inferred from Community Notes, the virality estimates may partly reflect selection in who writes notes and which posts get them; a post needs a note mentioning AI to enter the AI group, and such notes may be more likely on striking or widely seen posts.","The larger effect on likes (34%) than retweets (11%) hints that AI-generated content succeeds by triggering an emotional or entertainment response that rewards approval more than sharing; a follow-up experiment could test whether labeling a post as AI-generated changes users' engagement.","Extending the same design to a different platform would clarify whether the virality premium is a property of AI content or of X's recommendation algorithm; persistence across platforms would suggest content-intrinsic appeal."],"forward_implications":["AI-generation status should become a standard explanatory variable in social-media misinformation research; the paper finds its engagement effect survives controls for content, sentiment, and account size.","Platform moderation should not rely on account prominence as the primary risk signal, because AI-generated misinformation spreads further despite coming from smaller accounts.","Fact-checking and detection systems geared toward negative or overtly political content will under-serve the part of the AI-generated misinformation ecosystem that is entertaining and positive in tone.","The finding that AI-generated posts are slightly less believable and harmful than conventional fakes implies that their viral advantage is not explained by higher perceived credibility.","Text-only sentiment analysis misses part of the positive tone carried by attached media, so multimodal annotation is needed to characterize AI-generated content."],"supporting_citations":[{"why":"Provides the Community Notes platform and data download that the entire dataset is built on.","marker":"X 2021"},{"why":"Establishes community-based fact-checking on X's Birdwatch/Community Notes as a source for identifying misleading posts.","marker":"Pröllochs 2022"},{"why":"Supports the wisdom-of-crowds premise that aggregated community ratings can identify misinformation accurately.","marker":"Allen et al. 2021"},{"why":"Prior study of how community fact-checked misinformation diffuses; supplies the diffusion controls and comparison baseline.","marker":"Drolsbach and Pröllochs 2023a"},{"why":"Provides the method for estimating author partisanship and misinformation exposure scores used in RQ2.","marker":"Mosleh and Rand 2022"},{"why":"Supplies the tweet sentiment model used as a text-only validation of the positive-sentiment finding.","marker":"Barbieri et al. 2020"}],"fun_headline_variants":["AI fakes out-viral human misinformation on X","AI misinformation more likely to go viral on X","AI misinformation more entertaining, positive, and viral on X","AI fake posts get 34% more likes and 11% more retweets on X"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a post actually contains AI-generated media when a language model reads its Community Note and judges that the note refers to AI; the human validation shows raters agree the note mentions AI, not that the underlying media was truly AI-made.","fun_headline_variants_meta":{"raw":{"variants":["AI fakes out-viral human misinformation on X","AI misinformation more likely to go viral on X","AI misinformation more entertaining, positive, and viral on X","AI fake posts get 34% more likes and 11% more retweets on X"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001165,"raw_usage":{"total_tokens":4811,"prompt_tokens":921,"completion_tokens":3890,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":3819}},"tokens_in":537,"tokens_out":3890,"duration_ms":24881,"temperature":1.0,"reasoning_tokens":3819,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:12:45.947810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the original images or videos of a random sample of posts the model labels AI-generated and check provenance or run forensic deepfake detectors; if most contain no detectable AI-created media, the virality gap is an artifact of note wording rather than of AI generation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the method for estimating author partisanship and misinformation exposure scores used in RQ2."}],"review_version":1}