{"id":"5616f8f4-57e1-4ea1-b333-32e1e2a44371","arxiv_id":"2503.05711","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Warning labels for AI-generated content raise users' belief that images are AI-made and shift trust in the label design, but leave like, comment, and share intentions largely unchanged.","lead":"This study tested ten warning label designs for AI-generated social media content with 911 participants. It finds labels increase belief that content is AI-generated, but do not significantly change likes, comments, or shares.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ3/F3 is not supported by the reported analysis: Table 3 tests within-design variation across images, not between-design differences in trust, so the headline claim that trust varies by label design lacks a demonstrated statistical basis.","rationale":"This pass treats the paper as an exploratory evaluation of ten warning-label designs, with headline empirical findings that labels raise belief that content is AI-generated, trust differs by label design, engagement is not significantly changed, and label trust correlates with platform trust. The OSF link, posted code, and explicit limitations are real strengths, and the reader's repeated-measures concern is valid. However, the most load-bearing gap is in RQ3: the reported Table 3 analysis is within-condition rather than between-condition. It cannot establish that trust varied across label designs, and the only between-group result is relegated to a supplementary mention. Since Finding F3 and the abstract's 'trust in the label significantly varied based on the label design' rest on this missing comparison, the central claim is at risk. A participant-level between-design ANOVA or mixed model would settle the issue directly, and the existing conditional verdict remains appropriate if the condition is expanded to require a transparent between-design trust comparison. Agreement is partial because the reader's weakest assumption was repeated-measures independence, whereas the trust-by-design comparison is a more direct threat to a headline claim: even a fully independent ANOVA of the kind presented would not answer RQ3.","tokens_in":30487,"tokens_out":6962,"duration_ms":68946,"concrete_test":"Open the OSF repository (osf.io/m8sg2) and locate the Q8 (Trust-In-Label) analysis. Check whether there is a between-group ANOVA or mixed model with Trust-In-Label as the dependent variable and the ten label designs (or control) as the factor, with one participant-level mean per design or random intercepts for participant and image. If present, report the test statistic, degrees of freedom, p-value, and the Tukey HSD cell for TRT10 versus each other design. If absent, re-run it from the posted data; if the overall between-design effect is non-significant or TRT10 is not significantly above the other designs, Finding F3 and the associated 'trust varied by design' statement should be removed or explicitly downgraded to a descriptive mean comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3.1 purports to answer RQ3, but the analysis in Table 3 does not compare label designs. Each row is a separate ANOVA inside one treatment group, using that group's eight image ratings; a significant F for TRT1 means trust in that label differed across the eight images, not that Design 1 was trusted differently from Design 10. The text's claim that 'trust varies based on the label design' and Finding F3 ('Design Sample 10 ... elicited the highest trust level') depends on a between-group ANOVA that is only mentioned parenthetically, with no F statistic, degrees of freedom, p-value, or effect size reported in the main text. If that between-group analysis is absent from the OSF supplement, the headline trust-by-design claim is unverified. This is independent of the repeated-measures concern: even if all image-level ratings were independent, Table 3 still would not answer RQ3. The repeated-measures issue compounds the problem, since the reported tests use eight correlated ratings per participant and can inflate significance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a four-dimensional design space for warning labels on AI-generated social media content (label sentiment, icon/color, position, level of detail) and evaluates ten prototype labels derived from it in a randomized between-subjects experiment with 911 Prolific participants. Participants rated eight images (four political, four entertainment; half AI-generated/deepfake, half real) either without a label (control) or with one of ten label designs. The authors report that labels significantly increased belief that content was AI-generated (RQ2), that trust in the label varied by design (RQ3), that label trust correlated with platform trust (RQ4), and that labels did not significantly change like/comment/share intentions (RQ5), while content type did affect engagement. The paper contributes a design-space framework and empirical evidence on label efficacy, with data and code shared on OSF.","tokens_in":30646,"tokens_out":5144,"duration_ms":45847,"significance":"This is a timely and well-scoped empirical contribution to the emerging literature on labeling AI-generated content. The randomized design, the use of realistic social-media-style stimuli, and the open sharing of materials and code are clear strengths. The design space itself is a useful synthesis of platform practice and prior work, and the null finding for engagement (F5) is a valuable counterpoint to the misinformation-labeling literature, suggesting that AI-content labels may not reduce sharing and liking in the way traditional fact-check labels do. If the trust-by-design finding were properly substantiated, the results would give platform designers concrete guidance. However, the statistical support for the headline trust-by-design claim and the trust-platform correlation is currently incomplete, and the repeated-measures structure of the data is not appropriately modeled in most analyses. These issues are fixable within the scope of the manuscript, but they need to be addressed before the central claims are reliable.","major_comments":[{"comment":"The reported analysis does not support the claim that trust in the label significantly varied based on the label design. Table 3 reports one-way ANOVAs run separately within each treatment group across the eight images; a significant F for a given group means trust in that label differed across the eight images, not that Design 1 was trusted differently from Design 10. The between-group ANOVA is mentioned only parenthetically ('To confirm the results in trust with the design samples, we performed a one-way ANOVA between the groups, it showed that TRT10 with warning label design 10 had the highest'), with no F statistic, degrees of freedom, p-value, or effect size reported in the main text, and the pairwise comparisons are deferred to a supplementary file. The abstract's statement that 'their trust in the label significantly varied based on the label design' and Finding F3 ('Design Sample 10 ... elicited the highest trust level') therefore lack a demonstrated statistical basis in the manuscript. Please report the full between-group ANOVA with pairwise comparisons in the main text, or temper the claim to what the within-group analysis actually shows.","section":"§4.3.1, Table 3, Finding F3"},{"comment":"The ANOVAs treat the roughly 7,000 image-level ratings as independent observations, even though each of the 911 participants rated eight images within a single condition. This nesting within participant and within image can inflate test statistics and shrink p-values. The authors acknowledge this in Section 5.4 ('Our repeated statistical analysis may affect the significance test') and state an intention to use linear mixed models, but the results presented in the main text are all based on the independence assumption. Until mixed models or some appropriate correction are reported, the significance of the belief effect (F1) and the between-treatment comparisons in Section 4.2.2 should be treated as provisional. At minimum, the paper should clarify the unit of analysis, justify the independent-observation assumption, or replace the affected ANOVAs with models that account for the repeated-measures structure.","section":"§4.2.1, §4.2.2, §4.5.1–§4.5.4, §5.4"},{"comment":"The correlation between trust in the label and trust in the platform is reported as r(85) = .73, p < .001, but the degrees of freedom do not match any obvious unit of analysis. The study has 10 treatment groups and 911 participants; a group-level correlation would yield df = 8, while a participant-level correlation would yield df ≈ 900. The reported df = 85 is inconsistent with both. Please clarify the unit of analysis on which the correlation was computed and report the corresponding sample size. If the correlation was computed on aggregated group means, the claim that individual users who trust the label also trust the platform is much weaker than presented.","section":"§4.4.2, Finding F4"}],"minor_comments":[{"comment":"The between-group ANOVAs in these two subsections report nearly identical F(9, 6468) values and p-values, yet Section 4.2.1 describes a comparison of control vs. treatment groups while Section 4.2.2 describes a comparison among the 10 treatment groups. It is unclear whether these are the same test presented twice or two different tests with coincidentally identical statistics; please clarify the actual models and report the N used in each analysis.","section":"§4.2.1–§4.2.2"},{"comment":"The text refers to 'control group (n = 648)' and 'treatment group (n = 6,593)' for the engagement analyses, but these are image-level observations, not numbers of participants. This is confusing for readers and should be explicitly labeled as image-level observations (e.g., '648 image ratings from 81 control participants').","section":"§4.5.1–§4.5.3"},{"comment":"The caption contains a typo: 'ward-cloud' should be 'word cloud'.","section":"Figure 9 caption"},{"comment":"Several references are incomplete, e.g., [12] lists no author names and no publication year, and a few other entries lack venue or retrieval details. Please complete the reference list.","section":"References"},{"comment":"The effect size for the treatment effect on belief is reported as eta-squared = 0.006. This is a very small effect, and while the paper acknowledges small effect sizes in the Discussion, the abstract's wording ('a significant effect') may overstate practical importance; a brief note in the abstract or results about the small magnitude would help calibrate reader expectations.","section":"§4.2.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a CHI 2025 submission, and the OSF supplement is referenced but not included in the arXiv version. The missing between-group ANOVA for RQ3 is a key point; if it is present in the OSF supplement, the authors should integrate the statistics into the main text. The repeated-measures issue is acknowledged by the authors, but the current analysis does not adequately address it; a mixed-model reanalysis is needed. The paper's scope and topic are well suited to the venue, and the empirical contribution is potentially valuable once the statistical reporting is fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful paper with a clear contribution, but one of its headline findings isn't supported by the analysis as reported. The four-dimension design space and the ten prototypes are the real new thing, and the null result on engagement is credible. The trust-by-design claim (F3) is not backed by the statistics in the main text: Table 3 runs separate within-group ANOVAs across the eight images for each treatment, so a significant F for TRT1 means trust varied across the images for that label, not that Design 1 was trusted more than Design 10. The paper mentions a between-group ANOVA and says pairwise comparisons are in the supplement, but no F, df, p, or effect size appears in the main text. That needs to be fixed before this finding can be assessed.\n\nThe repeated-measures issue compounds this: with 911 participants rating eight images each, the image-level observations aren't independent. The authors acknowledge this in Section 5.4 and plan mixed models, but the reported p-values are likely too small. The belief effect and the engagement null are probably robust to this, but the exact numbers should be re-estimated.\n\nCredit where due: the design space is a genuine mapping exercise grounded in platform practice and prior literature, the experiment has a reasonable sample, attention checks, randomization checks, and the data/code are on OSF. The qualitative analysis is thoughtful and generates useful hypotheses about label wording and placement.\n\nThe small effect sizes are worth noting but not disqualifying; the authors are open about them. The main fix is statistical transparency: report the between-group ANOVA for trust, and run mixed models for the repeated-measures structure. If the supplement already contains the between-group analysis, the authors need to bring it into the main text. If it doesn't, F3 should be presented as a descriptive observation, not a finding.\n\nWho benefits: HCI researchers, platform designers, and people implementing AI-content labeling under the EU AI Act. It deserves serious peer review, but the revisions on the trust analysis are mandatory, not cosmetic.","headline":"A useful design-space paper with a credible null on engagement, but the headline trust-by-design claim rests on an unreported between-group ANOVA.","tokens_in":31156,"tokens_out":2042,"would_cite":true,"duration_ms":18858,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Warning labels make users believe social media images are AI-made, but do not change how they engage with them.","keywords":["warning label design","AI-generated content","deepfake","user perception","social media engagement","trust in labels","human-computer interaction","content labeling"],"falsifier":"Re-analyze the posted dataset with linear mixed models that include random intercepts for participant and image; if the belief effect (reported as F(9,6468)=4.05, p<.001) becomes non-significant or the trust differences between label designs disappear, then the headline findings are artifacts of the independence assumption. A field experiment with real like, comment, and share data would also settle whether the null engagement finding holds outside survey intentions.","tokens_in":30273,"feed_emoji":"🏷️","tokens_out":4486,"duration_ms":40349,"temperature":0.7,"pith_summary":"The paper sets out to define a design space for warning labels on AI-generated social media content and to test whether such labels work. It derives ten label prototypes from four dimensions—sentiment, iconography and color, position, and level of detail—and runs a 911-participant experiment with a control group. The central result is that any label made users more likely to believe an image was AI-generated or edited, while trust in the label itself depended on its design. At the same time, labels did not significantly change likes, comments, or shares, although political content drew more engagement than entertainment content. The authors conclude that labels inform belief and build platform trust but are not a standalone fix for engagement with synthetic media.","feed_headline":"AI warning labels change beliefs, not engagement","feed_subtitle":"Study of 911 users and 10 label designs finds belief shifts, trust varies, while likes, comments, and shares stay flat.","key_machinery":"The carrying instrument is a four-dimensional design space for AI-content warning labels—label sentiment (hazardous, neutral without AI, neutral with AI), iconography and color (warning icon in red, neutral icon, no icon), position (on content obscuring, above content, on content non-obscuring), and level of detail (simple text versus detailed provenance). From this space the authors built ten prototype labels and embedded them into eight realistic social-media image posts, four AI-generated or edited and four real, split evenly between political and entertainment content. The experiment randomly assigned 911 participants to the control or one of the ten treatments and measured belief that content was AI-generated, trust in the label, trust in the platform, and like, comment, and share intentions using five-point Likert items. This design lets the paper attribute differences in belief and trust to label design while using the control group to isolate the mere presence of a label.","core_discovery":"The paper's core claim is that warning labels for AI-generated content are effective as transparency signals but weak as behavior-change tools. Across ten label designs tested against a no-label control, the presence of a label significantly raised users' belief that content was AI-generated, deepfake, or AI-edited regardless of their familiarity with the content, with the strongest effects for labels reading 'Made with AI' with a neutral diamond icon. Trust in labels varied by design: the Content Credentials-style detailed label earned the highest trust, while red 'Deepfake' warnings scored lower and were read as ambiguous. Label trust and platform trust were strongly correlated, yet engagement—liking, commenting, sharing—did not differ significantly between labeled and unlabeled content; content category, not labeling, drove engagement differences. The authors present this as evidence that label design choices matter for belief and trust, and that labels should be part of a broader intervention strategy rather than a standalone solution.","pith_inferences":["If the engagement null holds in real-world settings, regulators should not expect transparency labels alone to slow the spread of AI-generated misinformation; complementary measures such as friction or accuracy prompts would be needed.","The finding that labels made users believe even real images were AI-generated suggests a false-positive cost: over-labeling could cultivate blanket skepticism toward authentic content, an effect the paper does not test.","A testable extension is to vary label prevalence across a feed to see whether trust erodes or the belief effect weakens under warning fatigue, and to measure actual clicks and shares rather than Likert intentions.","The design space could be combined with provenance metadata to test whether a two-layer label—simple 'Made with AI' text plus click-through details—captures both belief and trust, since the most trusted design was not the one that best communicated that content was AI-generated."],"forward_implications":["Platforms that adopt any of the ten tested label designs can expect users to be more likely to believe an image is AI-generated or edited, even when the image is real.","Label wording matters: designs using 'Made with AI' with a neutral icon produced the strongest belief effect, while designs using the word 'Deepfake' were less clear to users and scored lower on trust.","Trust in the label transfers to trust in the platform: label trust and platform trust correlated at r(85)=.73, p<.001.","Labels alone will not reduce engagement: no significant difference in like, comment, or share appeared between labeled and unlabeled conditions, but political content drew more reactions and comments than entertainment content.","Familiarity with the content does not change whether users believe the label's assertion that content is AI-generated."],"supporting_citations":[{"why":"Supplies the call for empirical work on labeling AI-generated content and the framing of promises and perils that this study answers.","marker":"[51]"},{"why":"Provides the established evidence that misinformation warning labels reduce belief and sharing, the baseline this study compares against.","marker":"[31]"},{"why":"Supplies prior evidence on which label terms users associate with AI-generated content, informing the wording dimension.","marker":"[12]"},{"why":"Precedent showing that warnings make users roughly twice as likely to detect deepfake videos, setting expectations for label effects.","marker":"[29]"},{"why":"Evidence that detailed provenance can backfire, used to interpret why detailed labels may fail to communicate their message.","marker":"[16]"},{"why":"Provides the implied-truth-effect framing and the direct Likert measurement approach the study adopts for engagement intentions.","marker":"[45]"}],"fun_headline_variants":["Labels boost AI detection, fail to curb engagement","Warning labels change minds, not clicks","Trust in AI labels depends on design","AI warnings: belief up, sharing unchanged","Label design matters for AI trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The statistical conclusion that labels change belief rests on treating each of the eight image ratings by the same participant as independent observations, even though those ratings are correlated within participant and within image; the authors acknowledge in Section 5.4 that repeated-measures analysis can affect significance and plan linear mixed models instead.","fun_headline_variants_meta":{"raw":{"variants":["Labels boost AI detection, fail to curb engagement","Warning labels change minds, not clicks","Trust in AI labels depends on design","AI warnings: belief up, sharing unchanged","Label design matters for AI trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1352,"prompt_tokens":948,"completion_tokens":404,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":341}},"tokens_in":564,"tokens_out":404,"duration_ms":4605,"temperature":1.0,"reasoning_tokens":341,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T19:32:14.737070+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the posted dataset with linear mixed models that include random intercepts for participant and image; if the belief effect (reported as F(9,6468)=4.05, p<.001) becomes non-significant or the trust differences between label designs disappear, then the headline findings are artifacts of the independence assumption. A field experiment with real like, comment, and share data would also settle whether the null engagement finding holds outside survey intentions.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the call for empirical work on labeling AI-generated content and the framing of promises and perils that this study answers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the established evidence that misinformation warning labels reduce belief and sharing, the baseline this study compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies prior evidence on which label terms users associate with AI-generated content, informing the wording dimension."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent showing that warnings make users roughly twice as likely to detect deepfake videos, setting expectations for label effects."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Evidence that detailed provenance can backfire, used to interpret why detailed labels may fail to communicate their message."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the implied-truth-effect framing and the direct Likert measurement approach the study adopts for engagement intentions."}],"review_version":1}