{"id":"023ccaa7-c4fa-4df2-889b-12eae2c51ec9","arxiv_id":"2506.16622","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new dataset and AI model for public perception of science news shows that perceived importance, surprise, and fun predict how much engagement science posts receive on Reddit.","lead":"Researchers created a large dataset of how people perceive science news on twelve qualities, from importance to fun, and trained an AI model to score new stories. They found that these predicted scores connect with real engagement: science posts perceived as more important, surprising, or fun get more upvotes and comments on Reddit.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The natural-experiment claim depends on within-URL validity of the perception model, which is untested: the n=50 Reddit validation gives Controversy r=0.16 (CI spans zero), and no paired same-URL validation exists.","rationale":"The reader's weakest assumption was domain transfer of the perception model to Reddit, and I agree that is the crux, but the sharper issue is the within-URL contrast. The paper builds a substantial resource (10,489 annotations) and the predictive correlations on news are non-trivial; the large-scale Reddit associations are plausible. However, the causal language is justified only by the natural experiment, which controls for the underlying science. That control raises the measurement bar: the model must rank different framings of the same article correctly. The reported validation cannot establish this, both because n=50 is small and because one of the five final-regression dimensions (Controversy) has a correlation CI spanning zero on Reddit. In addition, the model's input distribution (news articles) differs from Reddit post text, and no paired-frame validation is reported. A concrete paired-frame validation would settle whether the model's within-URL predictions are meaningful. This does not warrant rejection; the paper's predictive findings and dataset are valuable. It does warrant a conditional decision requiring this validation (or explicit softening of causal language) and release of data and code.","tokens_in":17709,"tokens_out":7588,"duration_ms":86642,"concrete_test":"Obtain human perception ratings for paired Reddit posts sharing the same URL (e.g., 100 distinct URLs, each with at least two posts), using the same 12-dimension instrument with 8 annotators per post. For each dimension, compute the correlation between model-predicted within-URL differences (post A minus post B) and human-rated within-URL differences. If the dimensions used in the regression (Importance, Surprisingness, Fun, Controversy, Expertise) do not show significantly positive within-URL agreement (e.g., Spearman rho > 0.3 with a CI excluding 0), the natural-experiment coefficients cannot be interpreted as evidence about public perception. Also report per-dimension confidence intervals for the existing n=50 validation to establish whether Controversy r=0.16 is distinguishable from zero.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central causal claim is the natural experiment in Figs. 5–6: for the same science (URL), posts with higher predicted Importance, Surprisingness, Fun, Controversy, and lower Expertise receive more comments and upvotes. This inference requires that the model's scores measure public perception on Reddit, not merely text style. The only domain-transfer evidence is the 50-post Reddit set (Appendix C, Table 3). For the five dimensions retained in the regression, Pearson r with human ratings is 0.62, 0.56, 0.76, 0.16, 0.40, respectively; at n=50, the 95% CI for r=0.16 is approximately [-0.12, 0.44], so Controversy predictions are statistically indistinguishable from zero on Reddit. Yet Figures 5–6 report a significant positive Controversy coefficient for comments. Moreover, the natural experiment requires more than marginal model validity: it requires that within-URL differences in predicted scores track human-perceived framing differences. The model was trained on full news article title+body pairs, and no validation is reported on paired posts sharing the same URL. Without such validation, the reported coefficients may be artifacts of lexical or engagement-bait features correlated with model output, rather than evidence that public perception drives engagement. Thus the strongest claim is not yet supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a computational framework for modeling public perceptions of science news along twelve dimensions, a crowdsourced dataset of 10,489 annotations from 2,101 US and UK participants across 1,506 science news stories, and fine-tuned RoBERTa models that predict the perception scores. The authors use the estimated perceptions in two analyses: first, mixed-effect regressions of individual background and content factors on perception scores; second, regressions of Reddit engagement (post scores and comment counts) on model-predicted perception scores, including a within-URL 'natural experiment' comparing different framings of the same science news story. The central claim is that posts estimated to be more important, surprising, fun, and controversial attract more engagement, and that this reflects a direct connection between public perception and engagement.","tokens_in":17972,"tokens_out":3570,"duration_ms":41673,"significance":"If the results hold, the paper would make a useful contribution: the annotation framework and dataset are substantial, the modeling pipeline is sensible, and the Reddit engagement analysis addresses an important question in science communication. The authors also provide a credible discussion of low inter-annotator agreement and attempt to validate their model on external domains, which is a strength. The main value would be in showing that automatically estimated perception dimensions can predict engagement beyond simple surface features. However, the central causal claim currently depends on model predictions whose validity for the specific within-URL Reddit setting is not established, so the significance is contingent on additional validation rather than being established by the present evidence.","major_comments":[{"comment":"The natural-experiment claim that framing the same science differently changes engagement via perceived dimensions rests entirely on model-predicted perception scores for Reddit posts. The only domain-transfer validation (Appendix C, Table 3) uses n=50 posts; for CONTROVERSY, Pearson r=0.16 with a 95% confidence interval spanning zero, and the five retained dimensions have correlations of 0.62, 0.56, 0.76, 0.16, and 0.40. No validation is reported on paired posts sharing the same URL, which is the exact setting of the within-URL regressions. Because the model was trained on title+body news text, the significant CONTROVERSY coefficient in Figure 6, and the other coefficients in Figures 5–6, could reflect lexical or engagement-bait features rather than perceived public perception. The causal wording in the Introduction ('strong causal relationship') and Discussion is not supported by the current evidence; I recommend adding same-URL human-rated validation or substantially softening the causal claims.","section":"§3.2, Figures 5–6"},{"comment":"The average Krippendorff's α across all statements is 0.11, and the dimensions that appear in the final Reddit regressions are among the least agreed-upon (SURPRISINGNESS α=0.096, CONTROVERSY α=0.147). The argument that low-IAA data can still train reliable models relies on one prior example and does not directly establish that the aggregated mean scores for these particular dimensions are stable enough to support the reported effect sizes. Please report the reliability of the averaged scores (e.g., variance components or split-half reliability) and, if possible, rerun the engagement regressions on the human-rated 50-post Reddit subset to show that the pattern is not an artifact of low annotation reliability.","section":"Appendix B.3, Table 2"},{"comment":"The stepwise VIF-based variable removal is not described in enough detail to know which of the twelve perception dimensions were removed and whether the final five (IMPORTANCE, SURPRISINGNESS, FUN, CONTROVERSY, EXPERTISE) are robust to the ordering of removal. Since the perception dimensions are intercorrelated, stepwise selection can produce unstable coefficients and optimistic significance levels. Please report the full coefficient table before removal, the VIF values for all dimensions, and a robustness check with alternative dimension subsets or a regularized regression.","section":"§5, Regression"}],"minor_comments":[{"comment":"There is an unresolved placeholder '(author?)' in the sentence discussing the newsworthiness annotation task; this reference needs to be completed.","section":"Appendix B.3"},{"comment":"The heading 'SURPRINGNESS' is a typo for 'SURPRISINGNESS'; please fix throughout if the misspelling appears elsewhere.","section":"Section 2"},{"comment":"The word 'Predicing' in the caption should be 'Predicting'.","section":"Figure 5 caption"},{"comment":"The name 'Krippendorrf' should be 'Krippendorff'.","section":"Appendix B.3"},{"comment":"The sentence 'In this section, I describe the creation process of this dataset' uses a first-person style inconsistent with the rest of the paper; please make it impersonal.","section":"Appendix B"},{"comment":"The paper does not state whether the dataset, trained models, or code will be publicly released; for a dataset-centered contribution, an availability statement would be valuable.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The core concern is the mismatch between the strength of the causal claim and the evidence for the model's validity in the within-URL Reddit setting. The n=50 domain-transfer validation and the low IAA for key dimensions make the central result vulnerable; if the authors can provide same-URL validation or a human-rated subset analysis, the paper could become a solid contribution. The paper is otherwise well within scope for a computational social science / NLP venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset and the 12-dimension perception framework are the real contributions here. Ten thousand annotations on 1,500 science news stories from representative US/UK samples is a meaningful resource, and the RoBERTa models achieve respectable correlation on several dimensions. The finding that science news consumption frequency and trust in science matter more than demographics is clean and useful, and the idea of comparing posts that share the same URL as a natural experiment is a clever way to hold the underlying science fixed.\n\nThe soft spot is the natural-experiment inference itself. The paper claims that predicted perception scores drive Reddit engagement, but those scores are validated on only 50 Reddit posts, and per-dimension correlations are weak for some of the exact dimensions used in the regression: Controversy at 0.16 and Expertise at 0.40. More importantly, the within-URL analysis requires that the model captures framing differences for the same science, and no validation exists on paired posts sharing a URL. Without that, the coefficients in Figures 5–6 could reflect lexical or engagement-bait cues rather than public perception. The phrase \"strong causal relationship\" overstates what is an observational design with URL fixed effects. The paper would be much stronger if the authors tempered that claim and either added a paired validation or explicitly acknowledged the untested assumption.\n\nThe low inter-annotator agreement (average alpha = 0.11) is a real limitation, but the authors handle it honestly, noting similar results in prior work and showing the averaged scores retain relative ordering (correlation 0.8 with ranking). I also did not see a data or code availability statement in the version I read; for a dataset-centric paper, that should be required.\n\nWho should read this: science communication researchers will find the dataset and the perception determinants analysis useful. Those interested in computational social science should read it as a cautionary example of how model-based predictions can drift when applied to a new domain without strong validation.\n\nI would send this to a serious referee. The core resource is valuable, the descriptive findings are worth publishing, and the engagement analysis is worth reporting with clearer limits. Expect major revision.","headline":"Solid dataset and framework, but the causal engagement claim leans on a perception model that is weakly validated for Reddit and never validated within the same-URL comparisons.","tokens_in":18470,"tokens_out":2974,"would_cite":true,"duration_ms":35149,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that public perception of science news can be measured on twelve dimensions, and that these measured perceptions predict engagement with science on Reddit — even when the same underlying research is framed differently.","keywords":["science communication","public perception","news values","perception prediction","Reddit engagement","natural experiment","framing","computational social science"],"falsifier":"Collect human perception ratings on a random sample of several hundred Reddit science posts and compare them with the model's predictions; if correlations for importance, surprisingness, or fun are near zero, the engagement regressions are artifacts of domain transfer. A second check would take one well-known study, post two framings with matched content but different predicted perception scores, and measure whether the higher-scoring framing reliably wins on upvotes and comments.","tokens_in":17525,"feed_emoji":"📈","tokens_out":4666,"duration_ms":50975,"temperature":0.7,"pith_summary":"This paper argues that how the public perceives a piece of science news can be measured along twelve dimensions, and that those measured perceptions predict how much people will engage with the science online. To support this, the authors build a dataset of over ten thousand ratings of fifteen hundred science news stories, train a language model to estimate the twelve perception scores from text, and then apply the model to roughly ninety-five thousand Reddit posts about science. They find that posts estimated higher on importance, surprisingness, and fun receive more upvotes and comments, while posts requiring specialized knowledge receive less engagement. The pattern holds even when the same underlying research is described in different framings, which the authors read as evidence that perception drives engagement rather than merely reflecting it.","feed_headline":"Perception scores predict Reddit upvotes and comments on science","feed_subtitle":"Same science framed as important, fun, or surprising draws more engagement, pointing to a causal link.","key_machinery":"The load-bearing machinery is a twelve-dimension perception framework paired with a supervised text model. Each dimension is defined by one or more Likert-scaled statements; human raters from representative US and UK samples scored the statements, and a fine-tuned RoBERTa-Large multi-task regression model was trained to reproduce the average scores from a news article's title and body. The same model is then applied to Reddit posts, and engagement regressions control for the shared URL, subreddit, domain, and first-sharing status, with a stepwise variance-inflation-factor procedure to handle correlated perception dimensions. The natural experiment component uses the fact that the same underlying science is often posted in multiple framings, allowing the authors to compare perception and engagement while holding the science itself fixed.","core_discovery":"On the paper's own terms, the central discovery is that public perception of science information is both measurable and predictive: a twelve-dimension framework (newsworthiness, understandability, expertise, importance, fun, surprisingness, controversy, exaggeration, interestingness, benefit, sharing willingness, reading willingness) can be scored by human raters, approximated by a text model, and used to forecast public engagement. The key empirical result is that estimated perception scores correlate with Reddit engagement across a large corpus, and that this correlation survives a natural experiment comparing different posts about the same science. The authors conclude that more positive perceptions cause more engagement, not merely that popular topics are perceived positively.","pith_inferences":["If the causal reading holds, the framework could be used for A/B-style message testing at scale: render alternative framings of the same finding, estimate each framing's perception scores, and select the version predicted to engage broader audiences before any human sees it.","The same perception scores could be inverted as a diagnostic for science communication inequity: posts that score low on importance or high on expertise may systematically exclude readers with less science background, and the model could audit which scientific fields receive such treatment.","The model's weak domain transfer to Twitter suggests the relationship between perception and engagement may differ across platforms; testing whether the same engagement pattern appears on shorter-form or image-first platforms would sharpen or limit the causal claim.","Because engagement metrics count interaction rather than understanding, using perception-driven engagement as a success measure risks optimizing for amusing or surprising content; a fuller account would pair perception scores with comprehension or trust outcomes."],"forward_implications":["Science communicators could estimate a draft's likely reception before publishing, flagging posts that read as overly specialized or low in perceived importance.","Framing changes engagement: a communicator can raise expected upvotes and comments by making importance, surprisingness, or fun salient without altering the underlying finding.","Since content domain and outlet type shape perception more than demographics, targeting messages by scientific field may matter more than tailoring by age, gender, or education.","The trained perception model offers a reusable measurement instrument for studying science communication at scale across news, Reddit, and potentially other platforms."],"supporting_citations":[{"why":"Supplies the news-values framework underlying most of the twelve perception dimensions.","marker":"[21]"},{"why":"Grounds the importance and benefit dimensions in science-journalism-specific news values.","marker":"[22]"},{"why":"Provides the crowd-rating approach and the Arxiv newsworthiness dataset used to validate the model's transfer.","marker":"[23]"},{"why":"Supplies the altmetric-based data processing pipeline for building the raw news/article corpus.","marker":"[51]"},{"why":"Provides additional preprocessing steps for the news corpus.","marker":"[52]"},{"why":"The annotation tool used to collect the crowdsourced perception ratings.","marker":"[53]"},{"why":"The pretrained language model fine-tuned into the perception predictor.","marker":"[54]"}],"fun_headline_variants":["Perception scores drive Reddit engagement, study shows","Twelve perception dimensions predict Reddit interaction","Model links perception scores to upvotes and comments","Perception scores causally affect Reddit comments and upvotes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The perception model trained on news articles generalizes to Reddit posts, so that the predicted scores used in the engagement regressions approximate the scores human raters would give those posts.","fun_headline_variants_meta":{"raw":{"variants":["Perception scores drive Reddit engagement, study shows","Twelve perception dimensions predict Reddit interaction","Model links perception scores to upvotes and comments","Perception scores causally affect Reddit comments and upvotes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2787,"prompt_tokens":949,"completion_tokens":1838,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":1776}},"tokens_in":565,"tokens_out":1838,"duration_ms":17239,"temperature":1.0,"reasoning_tokens":1776,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:37:31.752122+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect human perception ratings on a random sample of several hundred Reddit science posts and compare them with the model's predictions; if correlations for importance, surprisingness, or fun are near zero, the engagement regressions are artifacts of domain transfer. A second check would take one well-known study, post two framings with matched content but different predicted perception scores, and measure whether the higher-scoring framing reliably wins on upvotes and comments.","supporting_citations":[{"cited_title":"What is news? news values revisited (again)","cited_arxiv_id":null,"evidence_quote":"Supplies the news-values framework underlying most of the twelve perception dimensions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the importance and benefit dimensions in science-journalism-specific news values."},{"cited_title":"From crowd ratings to predictive models of news- worthiness to support science journalism","cited_arxiv_id":null,"evidence_quote":"Provides the crowd-rating approach and the Arxiv newsworthiness dataset used to validate the model's transfer."},{"cited_title":"Modeling Information Change in Science Communication with Semantically Matched Paraphrases","cited_arxiv_id":"2210.13001","evidence_quote":"Supplies the altmetric-based data processing pipeline for building the raw news/article corpus."},{"cited_title":"Potato: The portable text annotation tool","cited_arxiv_id":null,"evidence_quote":"The annotation tool used to collect the crowdsourced perception ratings."}],"review_version":1}