{"id":"6c16e0ef-5561-4737-ae34-700e6393e437","arxiv_id":"2412.13280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"During the 2022 U.S. midterms, fewer than half of prominent misinformation narratives received fact-checks, the median fact-check came four days late, and fact-check posts were about 1.2% of narrative-related conversation.","lead":"This study measured how often, how quickly, and how widely institutional fact-checks reached audiences during the 2022 U.S. midterm elections on X/Twitter. It found that fewer than half of major misinformation narratives were fact-checked, the median fact-check arrived four days after a rumor first appeared, and fact-check posts were about 1.2% of the conversation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fact-check identification via Google search and MBFC-filtered sources may undercount corrections, making coverage and speed magnitudes upper/lower bounds; the qualitative claim survives but headline values need a sensitivity check.","rationale":"The reader's weakest assumption is the composition of the 135-narrative corpus. My strongest concern overlaps but is slightly different: fact-check identification is the load-bearing step for the speed and reach headlines, while corpus completeness primarily drives the coverage headline. The paper's own Google-search design, its restriction to MBFC-rated unbiased sources, and its requirement that a correction explicitly use the term 'fact-check' all bias the fact-check set toward certain organizations and away from others. This is a measurement problem on the numerator, not just the denominator, and it affects all three claimed barriers. That said, the concern is not disqualifying. The paper is transparent about its identification procedure, and it frames the work as an observational snapshot with acknowledged limitations. The qualitative claims—fact-checks are a minority of the conversation, arrive only after a substantial fraction of misinformation posts, and spread within rather than across partisan communities—would likely survive even a substantial expansion of the fact-check set, because the conversation-volume and partisan-siloing results are structurally robust to adding a few thousand more fact-check posts to a conversation of over a million. The honest verdict remains CONDITIONAL: the paper's magnitudes should be treated with caution until data and code are released and a sensitivity analysis against a comprehensive fact-check archive is performed, but no internal inconsistency or fatal flaw is present. I would keep the reader's CONDITIONAL verdict rather than escalating, because all identified issues are addressable with released data and reanalysis.","tokens_in":20101,"tokens_out":1753,"duration_ms":15867,"concrete_test":"Re-run the coverage and speed analyses using a comprehensive fact-check archive (e.g., Duke Reporters Lab or ClaimReview tagged articles from all IFCN signatories) instead of the Google-search procedure, releasing a full list of the 135 narratives and the 164 identified fact-checks. For every narrative currently coded as not fact-checked, verify against the archive whether at least one attributable fact-check exists. If the coverage rate moves beyond a few percentage points, or if the median delay changes by more than a day, the paper should report both estimates and recast its headline values as bounds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claims about coverage (47% of narratives) and speed (four-day median delay) depend on the completeness of the ElectionRumors2022 denominator. That dataset was constructed in real time by manual curation, and the paper itself says very low-spread cases were excluded and duplication was removed. If the curation missed narratives, coverage could be higher or lower in an unknown direction. However, the more decisive internal check concerns fact-check identification. The authors only count fact-checks found via Google search, restricted to sources rated high credibility/unbiased by Media Bias Fact Check, and only those that use the term 'fact-check'. Fact-checks published on platforms that do not surface in the first ten Google results, checks by organizations that MBFC rates as biased, and corrections that do not use the label 'fact-check' (e.g., news debunks or platform labels) would be missed. A missing fact-check for a narrative classified as 'not fact-checked' would inflate the 47% coverage gap and the speed measure (a missed earlier fact-check would lengthen the four-day median). The speed number is additionally sensitive to the manual validation of the earliest post per narrative; the authors state they validated the earliest post is related, but any error in that timestamp shifts the delay directly. The reach analysis further depends on counting fact-check shares only when a post contains a link to a URL in an identified fact-check; non-link correcting posts and corrections from the missed sources are invisible, making the 1.2% conversation share a lower bound rather than an unbiased estimate. None of these issues is necessarily fatal—the qualitative conclusion that fact-checking was partial, slow, and partisan-bounded is plausible—but the headline magnitudes require a sensitivity analysis against a comprehensive fact-check archive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates real-world constraints on institutional fact-checking during the 2022 U.S. midterm elections by combining three data sources: ElectionTweets2022 (446 million election-related posts), ElectionRumors2022 (135 manually curated false/misleading narratives), and a new corpus of 164 fact-checks identified by Google search with Media Bias Fact Check (MBFC) source filtering. The central empirical claims are that only 47% of prominent narratives were fact-checked, that the median first fact-check appeared four days after a narrative's first post and after 79% of its misinformation posts, and that fact-check-linked posts constitute only 1.2% of the narrative conversation and circulate almost exclusively within partisan communities. The paper also argues that partisan asymmetry in coverage attenuates once narrative type is controlled, using logistic and OLS regressions with robustness checks on virality and influencer thresholds.","tokens_in":20273,"tokens_out":3106,"duration_ms":33559,"significance":"If the headline estimates hold, the paper makes an important contribution: it measures fact-checking effectiveness in a naturalistic setting using a rumor corpus that is not built from fact-check archives, thereby avoiding the selection/circularity problem that affects much of this literature. The study is also transparent in reporting manual validation of narrative timelines, robustness checks on variable cut-offs, and a stated intention to release code and data, all of which strengthen confidence in its descriptive findings. The coverage, speed, and reach magnitudes, should they survive sensitivity analysis, would be directly relevant to debates about whether institutional fact-checking can serve as a primary correction mechanism.","major_comments":[{"comment":"The headline coverage figure of 47% is sensitive to the completeness of fact-check identification, which relies on the first ten Google results, a requirement that pages use the term \"fact-check,\" and MBFC ratings of unbiased/high credibility. A debunk that does not surface in that search, is published by an organization MBFC rates as biased, or is formatted as a news correction or platform label would be missed, and if any such debunk corresponds to a narrative coded as \"not fact-checked,\" the 47% figure is inflated and the four-day median delay is lengthened. The manuscript needs a sensitivity analysis that varies these inclusion rules, for example by (a) matching narratives against IFCN/International Fact-Checking Network signatory databases or the Duke Reporter's Lab list, (b) relaxing the MBFC filter to include all fact-checking organizations regardless of MBFC bias rating, and (c) reporting how many \"not fact-checked\" narratives have at least one candidate correction from these broader sources. This is load-bearing for the paper's central claim.","section":"Data & Methods, Fact-check Identification (p. 9)"},{"comment":"The 47% coverage rate and the four-day median delay are conditional on the completeness of the ElectionRumors2022 denominator, which was constructed in real time by manual observation and which the authors state eliminated \"very low-spread\" cases. If the curation missed prominent narratives or systematically underrepresented particular communities or types of claims, the coverage, speed, and partisan-comparison results all shift, and the direction of the bias is not obvious a priori. The paper should state this limitation explicitly in the main text and provide an external validation check, such as comparing the 135-narrative corpus against an independently constructed list of prominent election-claims (e.g., from the Election Integrity Partnership or media casebooks), including how many narratives overlap and whether the coverage rate changes on the union of the two lists.","section":"Data & Methods, False and Misleading Narratives (p. 8)"},{"comment":"The logistic regression in Table 3 shows signs of quasi-complete separation: the coefficients for \"Partisanship: Neutral\" (-16.036, SE 1.239) and \"Classification: Improbable\" (16.484, SE 0.832) are implausibly large with small standard errors, which is a classic separation artifact rather than a meaningful estimate. Consequently, the claim in the Results that partisanship has \"limited evidence\" of influencing coverage rests on unstable model estimates, especially given only 3 neutral narratives and 9 improbable narratives. The authors should re-estimate with Firth's penalized likelihood or a Bayesian model, or present the OLS specification in Table 6 as the primary model, and should note in the text that the \"Improbable\" and \"Neutral\" coefficients are not interpretable as finite odds ratios.","section":"Results, Coverage; SI Table 3 (logistic regression)"},{"comment":"The reach analysis defines a fact-check post only as a post containing a link to a URL of an identified fact-check. This measure excludes quotes, replies, and text-based corrections that name or describe the fact-check without a URL, and the Discussion later acknowledges that \"textual corrections which do not link to an external fact-check\" are outside the analysis. As a result, the 1.2% conversation share and the 4.15%/0.30%/0.38% sharing rates are lower bounds, not point estimates. The main text should state this directional bias explicitly, and ideally the authors should hand-code a sample of posts that discuss fact-checks without linking to quantify how much the reach estimate would change. The Discussion's statement that \"fewer than 2% of users ... also shared a fact-check\" should also be reconciled with the Table 2 percentages, which use different denominators.","section":"Results, Reach (pp. 16-18)"}],"minor_comments":[{"comment":"The citation \"election period Schafer et al., 2024\" is missing an opening parenthesis; it should read \"election period (Schafer et al., 2024).\"","section":"p. 3, Introduction"},{"comment":"The text says \"at least 100,00 followers\" and should read \"100,000.\"","section":"p. 12, Results, Coverage"},{"comment":"The sentence beginning \"While this outcome likely comes of no surprise to fact-checkers, who prioritize ... based on their as their human and technological resources\" contains a duplicated \"as their\" and should be rephrased.","section":"p. 15, Results, Speed"},{"comment":"The caption refers to the \"modal fact-check (red)\" in one place and to the \"aggregated (median) fact-check response time (vertical line)\" in another, which is confusing about whether the vertical line marks the mean, median, or mode; the text and legend should use one consistent statistic.","section":"Figure 2 caption"},{"comment":"The abstract's phrase \"most comprehensive assessment to date\" is an overclaim for a single-platform, single-election study with a manually curated corpus; I suggest tempering it to something like \"a comprehensive assessment\" or \"one of the first campaign-wide assessments.\"","section":"Abstract and Discussion"},{"comment":"Several author names appear with a spacing artifact (e.g., \"V ogels\" for Vogels, \"V osoughi\" for Vosoughi, \"V .\" for V.). These should be corrected for a publication version.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid descriptive core and the non-circular design is a real strength. The main risk is that the coverage and speed headlines depend on potentially incomplete fact-check identification and on the manually curated denominator; both are addressable with sensitivity analyses and external validation rather than requiring new data collection at a different scale. I would not recommend rejection, but the load-bearing sensitivity issue in the fact-check identification should be resolved before the paper is accepted. There is also a fit question for the editor: the paper is methodologically straightforward, and its contribution lies more in the empirical magnitudes and the design than in new methods; that is appropriate for a data-oriented journal but should be weighed against the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper measures fact-check coverage, speed, and reach starting from an independent rumor corpus rather than from fact-check archives. That design choice is the real contribution, and it makes the headline numbers worth engaging with.\n\nThe good parts are real. ElectionRumors2022 gives the authors 135 narratives and 1.8M posts collected in real time, and they manually validated the earliest post for each narrative so the timing analysis has a solid anchor. The results — 47% of narratives ever fact-checked, a four-day median delay between first post and first fact-check, 1.2% fact-check share of conversation — are plausible and, as far as I know, new as a package. The partisan finding is more nuanced than the raw percentages suggest: right-leaning narratives are fact-checked more often, but once you control for claim type, the gap mostly disappears. That is a useful corrective to the 'moderation is biased against the right' narrative in public debate.\n\nThe soft spots are exactly where your reader flags them. Fact-check identification relies on Google search, restricted to MBFC high-credibility sources and the literal term 'fact-check'. That will miss corrections, which means the 47% is a lower bound on coverage and the four-day delay is an upper bound on speed. The 1.2% reach figure only counts posts containing a link to an identified fact-check, so it too is a lower bound. The denominator — the 135 narratives — was manually curated in real time, and the paper admits very low-spread cases were excluded; the abstract says 'prominent', but the precision of the percentages should not be oversold. The logistic regression also shows signs of perfect separation (the neutral and improbable-results coefficients are around ±16 with standard errors under 1), so the OLS re-estimation is a sensible check and should perhaps be the primary specification for those predictors. Finally, the arXiv version contains a data-availability promise but no code or data. All of these are addressable rather than fatal.\n\nWho is this for: misinformation researchers, fact-checking practitioners, and platform policy staff. The qualitative conclusion — institutional fact-checking is partial, slow, and partisan-bounded — survives the identification worries. The exact magnitudes need a sensitivity check against a comprehensive fact-check archive and a public code/data deposit.\n\nRecommendation: send it to review. It deserves a serious referee. I would ask for the sensitivity analysis, confidence intervals, and data/code before accepting.","headline":"A genuinely new measurement of fact-checking limits built on an independent rumor corpus; the headline numbers are plausible but the fact-check identification needs a sensitivity check before they are treated as exact.","tokens_in":20961,"tokens_out":3041,"would_cite":true,"duration_ms":25953,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Institutional fact-checking during the 2022 U.S. midterms covered fewer than half of prominent election rumors, lagged four days behind them, and reached barely one percent of the conversation.","keywords":["fact-checking","misinformation","election rumors","coverage","speed","reach","partisanship","X/Twitter"],"falsifier":"A fully independent census of 2022 midterm election misinformation—assembled without real-time observer judgment, for example by mining a complete archive of election-related posts with a different narrative-clustering method or by using post-election retrospective media reviews—that found more than half of prominent narratives were fact-checked, or that the median lag was under 24 hours, would contradict the core coverage and speed claims. A narrower test: locate a widely shared voter-suppression narrative from the period that was quickly fact-checked by multiple organizations; the paper's finding that suppression narratives are checked at 19% and largely ignored would need revision.","tokens_in":19812,"feed_emoji":"📉","tokens_out":6946,"duration_ms":60662,"temperature":0.7,"pith_summary":"This paper tries to establish that institutional fact-checking, as practiced during the 2022 U.S. midterm elections on X/Twitter, was too selective, too slow, and too partisan-bounded to function as the primary defense against political misinformation. Drawing on a real-time dataset of 135 prominent false or misleading election narratives, the authors find that fewer than half were ever fact-checked; for those that were, the median fact-check appeared four days after the narrative first surfaced and after 79% of the narrative's misinformation posts had already been published; and fact-check posts made up under 1.2% of narrative-related conversation, mostly circulating within a single partisan community. Once claim type and timing are held constant, the data show little evidence of right-leaning bias in coverage; the apparent partisan asymmetry tracks the kinds of claims each side produced. If these findings hold, fact-checking as currently organized cannot carry the burden of correcting election misinformation, and faster, platform-embedded or better-coordinated corrections are needed.","feed_headline":"Fact-checks missed most 2022 election rumors","feed_subtitle":"Half of rumors went unchecked; the median check arrived four days late and reached 1.2% of the conversation.","key_machinery":"The load-bearing comparison is between two corpora built independently: ElectionRumors2022 (a real-time, manually curated dataset of 135 false or misleading election narratives on X/Twitter during the 2022 midterms) and a manually collected set of 164 fact-checks found by Google searches for each narrative. To that is added a partisanship assignment for users and narratives built from co-engagement networks of reposts, plus a logistic regression of which narratives get fact-checked. Together these allow the paper to measure coverage (share of narratives with at least one fact-check), speed (delay from a narrative's first post to its first fact-check), and reach (share of posts containing fact-check links and the partisanship of users sharing them).","core_discovery":"The paper's central discovery is that the three practical constraints—coverage, speed, and reach—compound to make institutional fact-checking a marginal force in a live U.S. election. Only 63 of 135 prominent election rumors (47%) received any fact-check; the median first fact-check came four days after the rumor's first appearance and three days after peak activity, with 79% of related posts already published; and among posts about even fact-checked narratives, only 1.2% contained or linked to a fact-check. Sharing of fact-checks was heavily partisan: left-leaning users posted 85% of fact-checks for left-leaning narratives and 80% for right-leaning narratives, while right-leaning users posted about 10%; only 0.30% of right-leaning users and 4.15% of left-leaning users shared any fact-check. The paper also shows that selection into fact-checking tracks claim type and timing more than virality or influencer involvement: all improbable-result narratives were checked, while only 19% of voter-suppression narratives were, and post-election narratives were checked more often. It reads the near-zero partisan coefficient after controls as evidence against the claim that fact-checking is systematically biased against the political right.","pith_inferences":["Editorial inference: the same coverage-speed-reach triad is a ready-made template for evaluating platform-embedded corrections, such as community notes, in the same 2022 data or in later elections.","Editorial inference: because suppression narratives are the least-covered type, a concrete reform—dedicating a share of fact-checking capacity to first-person, hard-to-verify claims—would directly target the largest coverage gap.","Editorial inference: the reach figure was measured on X/Twitter only; corrections attached to posts rather than circulated as links might cross partisan lines differently, so transferring the 1.2% number to other platforms is an extrapolation, not a finding.","Editorial inference: the archive-bias result implies that prior misinformation studies sampling from fact-check databases may have overestimated effect sizes conditional on claim type; re-running such studies on fact-check-independent corpora would show whether their conclusions survive."],"forward_implications":["If the 47% coverage figure is correct, reliance on institutional fact-checks alone leaves more than half of prominent election rumors publicly unrebutted in the critical window.","A median four-day delay with 79% of posts already out implies the modal fact-check functions as a record of a rumor rather than a timely correction.","Fact-check posts being 1.2% of the conversation and mostly within one partisan community means the audiences most exposed to a rumor rarely encounter the correction.","Using fact-check archives as a sampling frame for misinformation research inherits a bias toward easily debunkable, post-election narratives and overstates partisan asymmetry.","Coverage decisions driven by claim type rather than virality or influencer engagement suggest that important rumors, such as voter-suppression claims, are systematically under-served."],"supporting_citations":[{"why":"Provides ElectionRumors2022, the 135-narrative corpus and timelines that form the denominator for the coverage and speed analyses.","marker":"Schafer et al., 2024"},{"why":"Supplies the definition of fact-checking and the meta-analytic evidence that fact-checks work in experimental settings, the benchmark this field study extends.","marker":"Walter et al., 2020"},{"why":"Documents prior real-world evidence that few users who encounter misinformation read corresponding fact-checks, motivating the reach analysis.","marker":"Guess et al., 2020"},{"why":"Provides the estimate that a median X/Twitter post loses information value in about 80 minutes, the speed baseline against which the four-day delay is judged.","marker":"Pfeffer et al., 2023"},{"why":"Documents partisan disparities in fact-check output during 2022 that the paper's controls qualify.","marker":"Stencel et al., 2023"},{"why":"Supplies the claim that corrections are most effective early, the normative basis for the speed analysis.","marker":"Vraga & Bode, 2020"},{"why":"Used as the credibility and partisanship filter for deciding which fact-check sources to include.","marker":"MBFC, 2023"},{"why":"Frames demand for fact-checks as a missing factor, motivating the reach analysis.","marker":"Graham & Porter, 2024"}],"fun_headline_variants":["Fact-checks reach 1.2% of election rumor conversations","Election fact-checks: half missed, four days late, 1.2% reach","Fact-checking falls short on coverage, speed, and reach","Study: fact-checks too sparse, slow, and partisan","Fact-checks miss half of rumors, lag 4 days, reach 1.2%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results depend on the 135 narratives in ElectionRumors2022 being a complete and unbiased record of prominent election misinformation; if the real-time observation missed major narratives or systematically excluded certain communities, every headline rate—47% coverage, four-day delay, and the partisan comparisons—would shift.","fun_headline_variants_meta":{"raw":{"variants":["Fact-checks reach 1.2% of election rumor conversations","Election fact-checks: half missed, four days late, 1.2% reach","Fact-checking falls short on coverage, speed, and reach","Study: fact-checks too sparse, slow, and partisan","Fact-checks miss half of rumors, lag 4 days, reach 1.2%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1388,"prompt_tokens":1092,"completion_tokens":296,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":195}},"tokens_in":708,"tokens_out":296,"duration_ms":3332,"temperature":1.0,"reasoning_tokens":195,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:16:27.791414+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A fully independent census of 2022 midterm election misinformation—assembled without real-time observer judgment, for example by mining a complete archive of election-related posts with a different narrative-clustering method or by using post-election retrospective media reviews—that found more than half of prominent narratives were fact-checked, or that the median lag was under 24 hours, would contradict the core coverage and speed claims. A narrower test: locate a widely shared voter-suppression narrative from the period that was quickly fact-checked by multiple organizations; the paper's finding that suppression narratives are checked at 19% and largely ignored would need revision.","supporting_citations":[],"review_version":1}