{"id":"ebeee3e4-0ad0-47d0-8e9c-85927d4af3d0","arxiv_id":"2505.09526","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review classifying misinformation warning evaluation metrics into behavioral, trust, usability, and cognitive/psychological categories, and highlighting standardization challenges.","lead":"This paper reviews how researchers measure whether misinformation warning labels, pop-ups, and browser extensions actually change user behavior and beliefs. It groups existing metrics into four categories and lists challenges such as inconsistency and lack of standard measures, but it offers no new experimental data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The four-category taxonomy is undermined by internal inconsistencies: CTR is defined oppositely in §III.A vs Table VI, and Table VI lists metrics with no category in §III; the review's central classification claim is unsupported until every metric is mapped and definitions reconciled.","rationale":"The reader's weakest assumption was that the PRISMA search and the four-category taxonomy are complete and non-overlapping. I agree the taxonomy is the load-bearing point, but the more specific and more decisive failure is internal: the paper's own tables contradict the narrative and introduce metrics that are never categorized. This is more damaging than the missing PRISMA flow diagram because even a perfect search cannot yield a comprehensive review if the resulting tables cannot be mapped onto the claimed categories. The CTR contradiction is particularly telling because a single metric is assigned opposite interpretations, and both appear under the paper's own analysis. A revised table that maps every metric to one category and one operational definition would settle whether the taxonomy is exhaustive and consistent. The uniqueness claim about no previous metric-focused review is not directly tested by this check; it would require a separate comparison with prior frameworks such as Hartwig et al. and Smith et al. Because this is a review with no strong empirical claim, rejection is not warranted; the issues are fixable with a mapping table and definition reconciliation. The reader's CONDITIONAL verdict therefore remains appropriate, and I would not change it.","tokens_in":13427,"tokens_out":5330,"duration_ms":56094,"concrete_test":"Produce a complete metric-to-category mapping from the final included study set: for every metric row in Tables II–VI, specify the Section III category (or an explicit 'uncategorized' bucket) and the operational definition taken from the cited source. Then check consistency: (1) no metric row is unmapped; if one is, the taxonomy is not exhaustive. (2) No metric is mapped to two different categories without justification; if one is, the categories overlap. (3) For metrics appearing in both narrative and tables (CTR, cognitive load, attitude shift), the definitions agree verbatim or are explicitly reconciled. The CTR contradiction in §III.A versus Table VI can be settled by returning to Kaiser et al. [10] to determine whether higher CTR means attention to the warning or disregard of it, then correcting one of the two passages.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a comprehensive, four-category taxonomy of metrics for misinformation warning interventions (Abstract, §III). That claim requires every included metric to be assignable to exactly one category and each category's definitions to be consistent. The manuscript does not currently satisfy this. Section III.A defines CTR as 'how frequently users interact with a warning intervention by clicking on it' and says a high CTR may indicate effectiveness in capturing attention, but Table VI defines CTR as 'the proportion of users who proceed beyond a warning to access potential misinformation' and says a high CTR indicates frequent disregard for warnings while a low CTR reflects user compliance. These are opposite interpretations of the same metric; if the meaning is unresolved, the categorization of behavioral metrics is unreliable. Similarly, Table VI contains metrics never defined in Section III: Awareness of Retraction, Durability of Warning, Sharing Discernment, Memory Bias Awareness, Perceived Risk of Misinformation, Misperception, and Skepticism. The four categories therefore do not cover all metrics the review itself identifies. Table IV places Cognitive Load under usability, while Section III.D places it under cognitive and psychological metrics, showing category overlap. Because the review's usefulness rests on the classification, these gaps and contradictions, not only the missing PRISMA details, directly undermine the completeness claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reviews metrics used to evaluate the effectiveness of misinformation warning interventions. It claims to be the first systematic review explicitly focused on such metrics (Section I) and proposes a four-category taxonomy: behavioral, trust and credibility, usability, and cognitive and psychological metrics (Section III, Figure 1). The authors adopt the PRISMA framework (Section II) and summarize metrics and their pros and cons in Tables II–VI. They identify challenges including variation in behavioral metrics, lack of standardization, context dependence, lack of diversity, limited dimensional focus, and ambiguity in measurement (Section V.A), and propose future directions such as multidimensional, inclusive, adaptive, and standardized metrics (Section V.B).","tokens_in":13653,"tokens_out":3572,"duration_ms":34625,"significance":"If the taxonomy and the inventory of metrics were reliable, this review would be a useful reference for researchers designing evaluations of misinformation warnings, and it would fill a genuine gap in a literature that is largely intervention-outcome oriented rather than metric-centric. The paper shows a broad engagement with the relevant literature (54 references), and the explicit pros/cons tables are a helpful start. However, the central classification claim is currently undermined by internal inconsistencies between the prose definitions in Section III and the metric table in Section VI, by category overlap, and by the presence of unmapped metrics in Table VI. Because the review's value rests on the completeness and consistency of the taxonomy, these issues are load-bearing. The PRISMA process is also underreported, which weakens the completeness claim.","major_comments":[{"comment":"Click-through rate (CTR) is defined in opposite ways. Section III.A states that CTR 'is employed to assess how frequently users interact with a warning intervention by clicking on it' and treats a high CTR as possibly indicating effectiveness in capturing attention. Table VI defines CTR as 'the proportion of users who proceed beyond a warning to access potential misinformation' and states that a high CTR 'indicates frequent disregard for warnings, while a low CTR reflects user compliance.' These are contradictory interpretations of the same metric. Because CTR is the first behavioral metric presented and is used as an example in Section V.A, the behavioral category is unreliable until this definition is reconciled.","section":"Section III.A vs. Table VI"},{"comment":"Table VI lists several metrics that are not defined or categorized in Section III: Awareness of Retraction, Durability of the Warning, Sharing Discernment, Memory Bias Awareness, Perceived Risk of Misinformation, Misperception, and Skepticism. If these are part of the review's inventory, the four-category taxonomy must either place them in the appropriate category or explicitly state that they are outside the taxonomy. As written, the taxonomy does not cover all metrics the review itself identifies, contradicting the claim of a comprehensive classification in Section III.","section":"Table VI"},{"comment":"The category boundary between usability and cognitive/psychological metrics is inconsistent. Cognitive Load is listed as a usability metric in Table IV, but Section III.D describes cognitive load as a cognitive and psychological metric. In addition, Table IV includes Task Completion Rate, User Satisfaction, Aesthetic Appeal, and Time on Task, none of which are defined in Section III.C's list of usability metrics. The taxonomy therefore has both overlaps and omissions, which undermines its claim of being exhaustive and mutually exclusive.","section":"Section III.C vs. Section III.D and Table IV"},{"comment":"The PRISMA process is underreported. The paper does not provide a PRISMA flow diagram, does not state the search dates (only 'studies up to 2024'), and does not report the number of records identified, screened, excluded, or included. Without these details, the completeness of the literature base for the review cannot be assessed, and the 'comprehensive' claim in the abstract and Section I is not verifiable.","section":"Section II"},{"comment":"The claim that 'perceived accuracy emerges as the most frequently used metric, followed by sharing intention or likelihood of sharing' is stated without the counting methodology or results to support it. Table VI does list many citations for perceived accuracy, but there is no systematic frequency analysis described, and some metrics in the table have overlapping definitions (e.g., Sharing intention vs. Sharing likelihood, Perceived accuracy vs. Misperception). This claim needs either a quantitative summary from the screening process or a clear qualitative basis.","section":"Section IV"}],"minor_comments":[{"comment":"The uniqueness claim 'To our knowledge, no previous review has explicitly focused on reviewing metrics for misinformation warning interventions' is unsubstantiated and should be qualified, especially since Hartwig et al. [18] is described as a systematic literature review of user-centered misinformation interventions and Smith et al. [20] explicitly discusses standardized outcome measures. A short comparison with these prior reviews would make the incremental contribution clearer.","section":"Section I"},{"comment":"Metric terminology is not consistent across sections and tables. For example, 'Believe change' in Section III.D appears as 'Belief change' in Table V, 'Memory retention' in Section III.D appears as 'Memory recall' in Table V, 'Sharing likelihood' in Section III.A appears as 'Sharing intention/Likelihood to Share' in Table VI, and 'Attitude shift' in Table V corresponds to a different definition than 'Attitude shift' in Section III.D. Please standardize names and definitions.","section":"Throughout"},{"comment":"The table title contains a typo: 'COGNITIVE AND PSYCOLOGICAL' should be 'COGNITIVE AND PSYCHOLOGICAL'. Also, the rows 'Belief change', 'Memory recall', and 'Attitude change' are not defined in Section III.D; please align the table with the prose.","section":"Table V"},{"comment":"The entry 'Engagement (likes, shares, and reactions)' contains a grammatical error: 'This metrics are use to measures' should be 'These metrics are used to measure'. Additionally, the table would be more useful if it included a column indicating which of the four proposed categories each metric belongs to, reflecting the paper's taxonomy.","section":"Table VI"},{"comment":"The inclusion criteria in Table I require peer-reviewed journal articles and conference proceedings, but the reference list includes sources such as arXiv preprints (e.g., [31], [33], [47], [53], [54]). Please clarify whether these sources met the inclusion criteria or whether the criteria were applied more loosely than stated.","section":"Section II"},{"comment":"Figure 1 presents the proposed classification but is not described in the text beyond a generic reference in Section III. A brief description of the figure and its categories would help readers follow the taxonomy, and the categories should be named consistently with the section headings.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is an early draft, as indicated by the 'Author's draft for soliciting feedback' note, and it shows promise as a reference review. However, the central taxonomy is not yet internally consistent, and the PRISMA reporting does not meet the standard expected for a systematic review. These issues are fixable within the paper's scope, so I recommend major revision. I would also encourage the editor to ask the authors to verify the uniqueness claim against the nearby literature, since Hartwig et al. [18] is closely related."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a useful review of evaluation metrics for misinformation warning interventions, but it is not ready as-is. The four-category taxonomy (behavioral, trust/credibility, usability, cognitive/psychological) is a reasonable organizing scheme, and the paper does a genuine service by assembling a broad set of metrics and mapping them to the studies that use them. The challenges section makes sensible points about standardization, context, and diversity. If you work on warning interventions, this is a useful starting bibliography.\n\nThe soft spots are real and they hit the paper's central claim. First, CTR is defined opposite ways in the text and in Table VI: the text says high CTR means the warning captures attention, the table says high CTR means users disregard the warning. That is not a minor inconsistency; it calls into question the reliability of the taxonomy. Second, Table VI lists metrics like Awareness of Retraction, Sharing Discernment, Memory Bias Awareness, Perceived Risk, Misperception, and Skepticism that are never discussed in Section III where the four categories are defined. So the taxonomy does not actually cover all the metrics the review identifies. Third, Cognitive Load appears in the usability table but under cognitive/psychological in the text, so the categories are not mutually exclusive. These issues mean the paper's main contribution, the clean classification, is not currently supported.\n\nThe PRISMA reporting is also thin: no flow diagram, no search dates, no screening counts. That may be a style problem with the author's draft, but for a systematic review it matters.\n\nThe claim of uniqueness ('no previous review has explicitly focused on metrics') is not substantiated by a comparative analysis of prior reviews. It might be true, but the paper doesn't show the work.\n\nWho is this for? Researchers designing studies on misinformation warnings who want a quick inventory of metrics and known pitfalls. It is not a definitive theoretical contribution. A serious referee could help the authors fix the inconsistencies and report the search properly; I would send it to review, not desk reject, but I would expect major revision. I would not cite it in its current form.","headline":"A useful but under-polished systematic review of metrics for misinformation warnings; the taxonomy is a good start, but internal inconsistencies undercut the completeness claim until fixed.","tokens_in":14193,"tokens_out":2484,"would_cite":false,"duration_ms":22435,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that no prior review focused on metrics for misinformation warning interventions, and that existing measures split into four families—behavioral, trust/credibility, usability, cognitive—whose fragmentation hinders…","keywords":["misinformation warnings","evaluation metrics","behavioral metrics","trust and credibility","usability metrics","cognitive and psychological metrics","misinformation interventions","systematic review"],"falsifier":"A re-analysis of a larger search corpus that includes non-English studies and platform industry reports would falsify the taxonomy's exhaustiveness if it turned up a metric family fitting none of the four categories—for example, cost-per-correction or temporal decay of warning effects. Likewise, a coding exercise showing that perceived accuracy and perceived believability are treated as interchangeable in the underlying studies would weaken the claim that the categories are distinct.","tokens_in":13222,"feed_emoji":"⚠️","tokens_out":8056,"duration_ms":72262,"temperature":0.7,"pith_summary":"This paper is a systematic review with a narrow target: the metrics researchers use to judge whether misinformation warning interventions work. It tries to establish that no earlier review has focused specifically on these metrics, that the metrics can be sorted into four families—behavioral, trust and credibility, usability, and cognitive and psychological—and that this fragmentation is a root cause of contradictory findings about warning effectiveness. The paper documents which metrics are common (perceived accuracy and sharing intention lead), which are missing (affective and emotional impact, accessibility), and what would need to change for evaluations to be comparable. The stakes are practical: platforms deploy warning labels widely, and without agreed metrics we cannot tell which label designs reduce belief in, and spread of, misinformation.","feed_headline":"Review: metrics for misinformation warnings fall into four families","feed_subtitle":"Perceived accuracy and sharing intention lead, but no standard makes studies hard to compare.","key_machinery":"The central organizing device is a four-category taxonomy of metrics, with each category named and defined in the paper: behavioral metrics (click-through rates, sharing rates, sharing likelihood, engagement duration, behavioral change), trust and credibility metrics (perceived accuracy, perceived objectivity, perceived credibility), usability metrics (perceived usefulness, perceived disruption, accessibility), and cognitive and psychological metrics (belief change, cognitive load, memory retention, attitude shift, cognitive resistance). The taxonomy does the argumentative work: mapping each published metric to these categories and to tables of pros and cons lets the review turn scattered findings into a list of challenges and future directions.","core_discovery":"The paper's central claim is that the field lacks a dedicated, systematic account of evaluation metrics for misinformation warning interventions, and that the metrics in use form four recognizable families. On this view, studies reporting positive or limited effects of warnings often are not measuring the same thing: some track clicks, shares, or alternative-source visits; others measure perceived accuracy, credibility, or trust; still others assess perceived usefulness, accessibility, cognitive load, or belief change. The paper assembles these into a taxonomy, identifies perceived accuracy and sharing intention as the most frequent measures, and argues that the lack of standardization, context dependence, missing accessibility considerations, and single-dimensional focus are why warning effectiveness remains contested.","pith_inferences":["If the taxonomy is right, a natural next step is a benchmark set of standardized tasks and scales—one the review calls for but does not build; a consortium could validate it by running several warning designs through the same battery.","The review's context-dependence challenge implies that the same warning may need platform-specific metric suites; that is testable by deploying identical labels on different platforms and measuring whether relative effectiveness rankings shift.","The absence of affective metrics suggests biometric or physiological measures, such as skin conductance or facial expression, could capture emotional responses that self-report misses; this is an extension the paper mentions as a gap but does not test.","A cost-side dimension—for example, false-positive warnings eroding trust, or the effort users spend verifying flagged content—is absent from the taxonomy; adding it would make the framework more useful to platform operators weighing intervention costs."],"forward_implications":["Researchers comparing warning studies will need to state which metric family they are sampling; results from a sharing-rate study and a perceived-accuracy study are not interchangeable.","A standardized metric battery would let platform designers test the same warning design across different platforms and see whether effectiveness rankings change.","Evaluations that ignore usability and accessibility may overstate effectiveness for the general population while missing failures for visually impaired users.","Future intervention studies should pair an observable behavior metric with a cognitive or trust metric, since single-dimensional measures can miss belief change without behavior change."],"supporting_citations":[{"why":"Prior proposal of a comprehensive evaluation framework for misinformation interventions; the paper positions itself against this by focusing on metrics rather than broader implications.","marker":"[17]"},{"why":"A systematic literature review of user-centered misinformation interventions that the paper draws on for review design and for the claim that no prior taxonomy centered on metrics.","marker":"[18]"},{"why":"Earlier review of fact-checking and warning interventions focused on short-term outcomes and cultural applicability, not on evaluation metrics.","marker":"[19]"},{"why":"Identifies widely used outcome measures such as perceived accuracy and willingness to share, supplying the baseline the paper extends by reviewing metrics in depth.","marker":"[20]"},{"why":"Primary source for sharing rate and sharing-likelihood metrics in the behavioral category.","marker":"[8]"},{"why":"Source for click-through rate, alternative visit rate, and behavioral-change metrics.","marker":"[10]"},{"why":"Source for perceived accuracy and engagement metrics, and for the finding that trust shapes perceived usefulness.","marker":"[25]"},{"why":"Supplies the accessibility gap argument by studying misinformation warnings with low-vision and blind users.","marker":"[30]"}],"fun_headline_variants":["Warning metrics: four families, no shared standard","Misinfo warnings: metrics fall into 4 groups, review finds","Four metric families for warning interventions, but no consensus","Misinformation warning metrics: four groups, but no shared standard","Warning metrics review: four families, no standardization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's completeness rests on the literature search finding all relevant metric families and on the four-category taxonomy being exhaustive and non-overlapping; if a family is missing or categories overlap, the list of challenges and the claim of comprehensive coverage weaken.","fun_headline_variants_meta":{"raw":{"variants":["Warning metrics: four families, no shared standard","Misinfo warnings: metrics fall into 4 groups, review finds","Four metric families for warning interventions, but no consensus","Misinformation warning metrics: four groups, but no shared standard","Warning metrics review: four families, no standardization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3487,"prompt_tokens":827,"completion_tokens":2660,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":2581}},"tokens_in":443,"tokens_out":2660,"duration_ms":19238,"temperature":1.0,"reasoning_tokens":2581,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:28:36.167739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-analysis of a larger search corpus that includes non-English studies and platform industry reports would falsify the taxonomy's exhaustiveness if it turned up a metric family fitting none of the four categories—for example, cost-per-correction or temporal decay of warning effects. Likewise, a coding exercise showing that perceived accuracy and perceived believability are treated as interchangeable in the underlying studies would weaken the claim that the categories are distinct.","supporting_citations":[{"cited_title":"A focus shift in the evaluation of misinformation interventions,","cited_arxiv_id":null,"evidence_quote":"Prior proposal of a comprehensive evaluation framework for misinformation interventions; the paper positions itself against this by focusing on metrics rather than broader implications."},{"cited_title":"The landscape of user-centered misinformation interventions-a systematic literature review,","cited_arxiv_id":null,"evidence_quote":"A systematic literature review of user-centered misinformation interventions that the paper draws on for review design and for the claim that no prior taxonomy centered on metrics."},{"cited_title":"Interventions to mitigate covid-19 misinformation: a systematic review and meta-analysis,","cited_arxiv_id":null,"evidence_quote":"Earlier review of fact-checking and warning interventions focused on short-term outcomes and cultural applicability, not on evaluation metrics."},{"cited_title":"A systematic review of covid-19 misinformation interventions: lessons learned: study examines covid-19 misinformation interventions and lessons learned,","cited_arxiv_id":null,"evidence_quote":"Identifies widely used outcome measures such as perceived accuracy and willingness to share, supplying the baseline the paper extends by reviewing metrics in depth."},{"cited_title":"Exploring lightweight interventions at posting time to reduce the sharing of misinformation on social media,","cited_arxiv_id":null,"evidence_quote":"Primary source for sharing rate and sharing-likelihood metrics in the behavioral category."},{"cited_title":"Adapting security warnings to counter online disinformation,","cited_arxiv_id":null,"evidence_quote":"Source for click-through rate, alternative visit rate, and behavioral-change metrics."},{"cited_title":"Real solutions for fake news? measuring the effectiveness of general warnings and fact- check tags in reducing belief in false stories on social media,","cited_arxiv_id":null,"evidence_quote":"Source for perceived accuracy and engagement metrics, and for the finding that trust shapes perceived usefulness."},{"cited_title":"Designing and conducting usability research on social media misinformation with low vision or blind users,","cited_arxiv_id":null,"evidence_quote":"Supplies the accessibility gap argument by studying misinformation warnings with low-vision and blind users."}],"review_version":1}