Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News Coverage

T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read An LLM pipeline labels every news article for lean, tone, and topic, then rolls the labels up into a dynamic, topic-level view of each publisher's bias.

desk verdict A genuinely useful LLM media-bias dashboard whose usability evidence is solid, but the political-lean dimension is under-validated to the point of being uninformative within topics; the authors' relative-value defense needs an external test. read the letter →

arxiv 2502.06009 v2 pith:5OZIYEZQ submitted 2025-02-09 cs.HC cs.CY

classification cs.HCcs.CY
keywords mediabiasnewsanalysislargelanguagemodelsLLM-driventoolsselectionframinghuman-in-the-loopdashboard
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the two forms of media bias that actually mislead readers—selection bias (which stories, facts, and people get covered) and framing bias (the tone and political lean of that coverage)—can be measured at the level of individual articles, every day, across major publishers, using a large language model with human oversight. To show this, the authors built the Media Bias Detector, a live dashboard that labels each of the top stories from ten major U.S. news outlets by category, topic, subtopic, article type, five-point political lean, and five-point tone, then aggregates those labels to reveal how coverage differs across publishers, topics, and time. The central benefit claimed is that readers can explore bias for themselves instead of accepting a static left-or-right rating of an entire outlet, and that both expert and general users found the tool usable, useful, and more trustworthy once they learned humans review the labels.

What carries the argument

The load-bearing mechanism is the article-level annotation pipeline: a GPT-4o model prompted to classify every article into category, topic, subtopic, article type, one of five political-lean levels, and one of five tone levels, with facts and event memberships extracted at sentence level. Humans stay in the loop at two points: they curate the topic hierarchy as news themes emerge, and they weekly validate the model's labels for coherence. The aggregation layer then rolls these per-article labels up to publisher, topic, and date-range views (the Coverage dashboard) and clusters same-day articles on the same incident into events whose top facts can be compared across publishers (the Events dashboard). The design rationale is that aggregation of article-level labels preserves within-publisher variation—for example, a publisher can be neutral overall but lean Republican on immigration and Democratic on the environment—which static single-score ratings cannot express.

What would settle it

Take a random sample of about 200 articles spanning the ten publishers and the period covered by the dashboard, have a panel of independent human coders label each article for lean and tone using the same five-point scales as the prompt, and measure agreement with GPT-4o's labels; if agreement for political lean on contested topics such as immigration, crime, or the election is near chance, the publisher-level lean distributions the tool displays are not supported. A weaker but still decisive version is to hold the articles fixed and swap the labeling model, checking whether a publisher's lean profile reverses with the model.

Watch

Extended reading notes

Core claim

The paper's central claim is that media bias is best exposed not by a single left-right publisher score but by article-level measurements of what gets covered and how it gets framed, aggregated over time. The Media Bias Detector operationalizes selection bias as the differential attention paid to news categories, topics, and subtopics, and framing bias as two dimensions computed per article: political lean (Democrat to Republican, with neutral levels) and tone (very negative to very positive). Each day, GPT-4o processes the top stories from ten prominent publishers and extracts topic, subtopic, article type, lean, tone, facts, and event clusters; human annotators review the model's output weekly for coherence and soundness rather than re-labeling from scratch. The authors report that in a study with 13 journalism, communications, and political science experts, along with a survey of 150 news consumers, users valued exploring bias from multiple angles, updated their beliefs about specific outlets when the data contradicted their priors, and trusted the tool more once told that humans validate the labels.

Load-bearing premise

Everything the dashboard shows rests on the assumption that the language model's automatic labels for political lean, tone, topic, and facts are accurate and stable enough to aggregate into publisher-level comparisons; the paper reports no agreement statistics against independent human labels, and its own limitations section concedes the model leans left and ties certain topics to parties regardless of article content.

Editorial extensions

If this is right

  • Static left-right publisher ratings become replaceable by distributions, so a publisher's lean can be reported per topic and over time rather than as one fixed label.
  • Users can compare how the same breaking event is framed across outlets at the level of individual facts, seeing which facts are emphasized, omitted, or worded differently.
  • Tone becomes a trackable dimension of framing bias, letting researchers measure the negativity or positivity of coverage per topic and publisher, a dimension most existing bias tools ignore.
  • The pipeline's cost scales linearly with the number of publishers and stories, so the same design can expand beyond ten outlets and beyond the top twenty daily stories as budgets allow.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the approach is right, the natural next study is longitudinal: tracking lean and tone distributions through an election cycle would test whether the tool's dynamic view detects real shifts in framing that static ratings would hide, something the paper's snapshot user studies do not yet show.
  • Because the paper concedes that GPT models lean left and tie topics like climate change to Democrats and immigration to Republicans regardless of content, a testable extension would be to re-label a fixed set of articles with several different models or with topic-blind prompts and check whether publisher-level rankings survive the swap.
  • The fact-extraction layer could be validated independently by checking whether the top facts attributed to an event are verifiable against a trusted record, which would separate framing differences from outright distortion.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents the Media Bias Detector, a web-based dashboard that uses GPT-4o to annotate individual news articles from ten major U.S. publishers in near real time for topic, subtopic, article type, tone, and political lean, then aggregates these labels to support publisher-level and topic-level comparison of selection and framing bias. The authors describe the system architecture and interface, report design considerations (broad exploration, deep exploration, comparison), and evaluate the tool through semi-structured interviews with 13 media experts and a follow-up survey of 150 news consumers. The main claims are that the tool makes granular, dynamic media bias exploration feasible for researchers, journalists, and the public, and that both expert and consumer users find it useful and valuable for understanding media bias.

Significance. If the central claims hold, the paper makes a useful contribution to HCI and computational media studies: it demonstrates an open, transparent, human-in-the-loop LLM pipeline for article-level bias annotation, and it provides qualitative evidence about the needs and trust concerns of expert and lay users of media bias tools. The authors are candid about several limitations, and the transparency of their methodology (prompts, topic hierarchy, and human validation process are made available) is a strength. However, the scientific value of the tool's central output, especially the political-lean labels, rests on validity evidence that the paper does not provide; the user studies measure perceptions, not the accuracy of the underlying measurements. The tool and its evaluation are therefore promising but not yet fully demonstrated.

major comments (4)
  1. [§9.3, §4.1, §2.3] The construct validity of the political-lean dimension is load-bearing and unsupported. Section 9.3 concedes that GPT-4o 'inherently associat[es] certain topics with Democratic leanings ... and others with Republican leanings ... regardless of the specific arguments presented in the article.' If this is true, then within-topic comparisons of lean across publishers—the central framing-bias use case described in Sections 3 and 4.1—would be confounded by topic composition, because two articles on the same topic with opposite frames could both receive the same label. The paper's assertion that 'their relative values ... are informative' is not backed by any evidence. To support the central claim, the authors should provide direct verification, for example by measuring label variation on paired articles about the same event with opposing frames, or by reporting human-LLM agreement or agreement against an independent benchmark on a held-out sample. The current validation described in Section 2.3, where humans read GPT outputs for coherence, cannot detect systematic topic-driven bias.
  2. [§7.2, Figure 8] The follow-up survey treats the tool's unvalidated labels as ground truth. For example, the text states that 101 of 150 respondents 'correctly stated' that the Wall Street Journal's economic coverage was neutral, and 107 of 150 gave the 'correct answer' on the Fox News vs. New York Times Biden-age question. These judgments rely entirely on the tool's outputs, whose accuracy is not established. Consequently, the observed post-tool shifts measure alignment with the model, not learning of an external truth. The authors should either validate the specific survey answers against independent human-coded or otherwise benchmarked labels, or reframe the findings as belief updating toward the tool's outputs rather than toward ground truth.
  3. [§2.3] The paper explicitly avoids independent labeling: 'we engage them in a validation task where they read GPT-generated responses to assess their coherence and soundness.' This design choice is understandable given subjectivity, but it means the paper provides no quantitative evidence about label accuracy, inter-annotator agreement, or human-model agreement for any of the four annotation dimensions. Given that the tool's entire value proposition depends on the trustworthiness of these labels, at minimum the authors should report a small-scale agreement study (e.g., two or three independent coders labeling a random sample of 50–100 articles per dimension) and disclose the resulting agreement statistics, even if the statistics are imperfect.
  4. [§5.5, §6, Figure 6] The expert evaluation's quantitative results show no significant improvement on most measures, and the authors explain this as expected due to short-term use. However, the qualitative findings are then used to support fairly strong claims about the tool's value. The manuscript would be strengthened by more explicit reporting of which claims rest on qualitative themes versus quantitative evidence, and by avoiding language that implies the quantitative study supports a broad effectiveness claim. This is a presentation and interpretation issue that should be addressed in the revision.
minor comments (6)
  1. [§4.2] The Events dashboard is described as showing coverage from 'the past three days' in the text and Figure 4, but Section 4.2 earlier states the page is 'sorted in descending order of the number of articles written about them in the past day.' Please clarify which time window applies.
  2. [Figure 8 caption] The caption contains a typo: 'Biden s age' should be 'Biden's age.' Please also fix similar apostrophe issues elsewhere (e.g., 'it s' in Figure 7).
  3. [§7.2] The shifts in survey responses displayed in Figure 8 are described as 'significant' and 'impactful,' but no statistical tests, confidence intervals, or effect sizes are reported for pre-post comparisons. Please add appropriate statistical reporting.
  4. [§2.3] The term 'validation' is used for a process that checks coherence and reasonableness, not label accuracy. Consider renaming this to 'coherence review' or 'qualitative verification' to avoid conflating it with accuracy validation, which the paper does not currently report.
  5. [§5.1] Table 2 lists participants P10 and P12, but their evaluation results were dropped from the quantitative analysis due to time constraints. The text notes this, but it is also worth stating in the table caption or the analysis section that the reported Likert and NASA-TLX results are based on N=11 rather than N=13.
  6. [§9.3] The statement that 'individual political lean labels could be subject to LLM biases, their relative values ... are informative' is a key assumption. Please either provide a citation to empirical evidence for this claim or mark it explicitly as a hypothesis to be tested.

Circularity Check

1 steps flagged · score 6.0 of 10

The follow-up survey defines 'correct answer' as the Media Bias Detector's own LLM labels, so measured post-tool learning reduces to alignment with the system; expert interviews and tool construction are not circular.

  1. self definitional [Section 7.2 Analysis, Figure 8a (same pattern in Figure 8b)]
    "The training phase revealed that the Media Bias Detector classified a majority of the Wall Street Journal’s economic coverage as neutral, and most respondents arrived at the same conclusion after using the tool. Furthermore, this had a significant impact on their post-training responses, with over two thirds of respondents (101/150) correctly stating that the publisher’s coverage was indeed neutral."

    The 'correct' answer is the tool's own GPT-4o classification, which participants were shown during the training/exploration stage before being re-asked the question. The survey therefore measures how faithfully respondents reproduce the tool's labels, not whether the labels are accurate or whether users learned an independently established fact. The same reduction appears in Figure 8b, where 'Fox News, the correct answer' is correct only because it matches the tool's article counts. Consequently, the paper's claims that the tool 'correct[s] priors' and 'introduces new insights' are, for these quantitative measures, true by construction rather than evidence of the tool's validity.

full rationale

The tool itself is a system, not a derivation, and most of the paper's evidence is non-circular: the expert interviews (Sections 5-6) report usable qualitative findings, and the design rationale in Sections 3-4 does not depend on treating the tool's output as ground truth. No load-bearing self-citation chain, imported uniqueness theorem, or ansatz-by-citation is present. The circularity is concentrated in the follow-up survey's learning-outcome measure: Section 7.2 equates 'correct' with the Media Bias Detector's own LLM-generated classifications, shows those classifications to participants during training, and then counts post-training convergence as evidence of impact (Figure 8a-c). That is a self-definitional evaluation step, not an independent validation of the lean, tone, or coverage labels. Section 9.3 explicitly concedes that GPT-4o associates certain topics with Democratic or Republican leanings 'regardless of the specific arguments presented in the article,' and asserts without evidence that 'relative values' remain informative. I do not count this as a circularity step, because it is a correctness/validity caveat rather than a definitional reduction, but it strengthens the need for the independent label validation that the paper does not provide. Overall score 6 reflects partial circularity: one central quantitative evaluation claim reduces by construction, while the tool implementation and qualitative expert findings remain independent.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The tool's central capability assumes that GPT-4o zero-shot labels for lean, tone, topics, and facts are usable without per-task training, and that the hand-built topic list and event window capture the news landscape. The paper makes no independent accuracy measurement, and Section 9.3 concedes the LLM shows left-leaning tendencies. These are the main axioms and free choices on which the reported insights rest.

free parameters (5)
  • One-day event detection window = 1 day (initially 3 days)
    Events are detected by clustering similar articles within a one-day window, and the paper states the window was adjusted from three days to one day. This hand-tuned choice determines which events appear and how coverage is grouped.
  • Top-20 articles per publisher per interval = 20
    Data collection captures only the top 20 articles per publisher in each interval, a resource-constrained cutoff that directly shapes the selection bias measurements and excludes less prominent coverage.
  • Five-level political lean scale = Democrat, Neutral Leaning Democrat, Neutral, Neutral Leaning Republican, Republican
    The discrete five-category lean scale is chosen by the authors; its granularity and thresholds affect all lean aggregations and user-facing comparisons.
  • Five-level tone scale = Very Negative, Negative, Neutral, Positive, Very Positive
    The tone classification uses a hand-picked five-level sentiment scale; no validation is reported for how these buckets map onto human judgments of news tone.
  • Hand-curated topic and subtopic hierarchy = Master list in Figure A3, updated by human annotators
    The set of categories, topics, and subtopics is defined and updated by research assistants when they notice major themes, so coverage analysis is limited to themes the team has already noticed.
assumptions (5)
  • domain assumption Political lean can be reliably operationalized as a five-level Democrat-to-Republican scale from article text.
    The entire lean analysis depends on this mapping. Section 4 says GPT-4o classifies lean, but Section 9.3 concedes the model's left-leaning tendencies and topic-lean associations.
  • domain assumption Tone can be classified into five discrete sentiment levels that capture framing bias.
    Section 4 defines the tone labels, but no human-LLM agreement or validation against human ratings is reported.
  • ad hoc to paper No ground truth for media bias exists, so observed patterns and differences between publishers serve as the standard for bias.
    Section 2.2 argues this explicitly. The survey then uses the tool's own outputs as the 'correct' answers, which is a circular use of the no-ground-truth stance.
  • domain assumption LLM zero-shot annotation is accurate enough for large-scale media analysis without task-specific training.
    Section 2.3 cites prior work on LLM annotation, but the paper provides no accuracy, precision, recall, or agreement statistics for its own labels.
  • domain assumption The hand-built topic list and one-day event clustering window capture the news landscape adequately.
    Section 9.2 notes that topic lists are driven by what annotators notice and that each article is assigned a single topic, which creates blind spots and undercounts multi-topic stories.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News Coverage." pith.science (2026). https://pith.science/paper/5OZIYEZQ

@misc{pith2026250206009,
  author       = {Pith},
  title        = {Pith review of: Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News Coverage},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5OZIYEZQ}},
  note         = {Machine review of arXiv:2502.06009}
}
read the original abstract

Mainstream media, through their decisions on what to cover and how to frame the stories they cover, can mislead readers without using outright falsehoods. Therefore, it is crucial to have tools that expose these editorial choices underlying media bias. In this paper, we introduce the Media Bias Detector, a tool for researchers, journalists, and news consumers. By integrating large language models, we provide near real-time granular insights into the topics, tone, political lean, and facts of news articles aggregated to the publisher level. We assessed the tool's impact by interviewing 13 experts from journalism, communications, and political science, revealing key insights into usability and functionality, practical applications, and AI's role in powering media bias tools. We explored this in more depth with a follow-up survey of 150 news consumers. This work highlights opportunities for AI-driven tools that empower users to critically engage with media content, particularly in politically charged environments.

Figures

Figures reproduced from arXiv: 2502.06009 by the authors.

Figure 1
Figure 1. The default view of the Coverage dashboard which allows [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. The grid view of the Coverage dashboard presents an alternative visualization to the stacked bar in Figure 1 by giving [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. To enable deep exploration of specific data (D2), we allow users to click through a hierarchy of news topics and subtopics to zoom into news of interest to them and be able to compare the volume, tone, and lean across publishers and date ranges. In this sequence of images, we show the user interacting with the dashboard to focus on the ‘Presidential Horse Race’ subtopic (D) after starting with the default all-catego… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Complementary to the Coverage view, the Events dashboard shown here presents an event-level view of the news that [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Overall results on the NASA-Task Load Index (NASA [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Average responses to Likert-scale measures on bias [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: In response to questions about the Media Bias De￾tector’s customization features, the vast majority of partic￾ipants stated that they found these options to be ‘extremely’ or ‘very’ useful. However, when asked whether this level of customization would be appropriate fo…
Figure 8
Figure 8. Figure 8: Participants showed a significant shift in their re [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: In contrast to our expert participants, the majority [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 11
Figure 11. Figure 11: A majority of participants said the Media Bias Detector had the potential to change their views on media bias over time. They also stated that they would find edu￾cational materials explaining different kinds of media bias helpful for getting the most out of this tool…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Scent: Exploring Cross-Language Information Navigation

    cs.HC 2026-04 conditional novelty 6.5 of 10

    Language scent—the perceived value of switching search languages—guides multilingual strategy; a system that augments it yields more granular strategies and diverse information.

Reference graph

Works this paper leans on

117 extracted references · 48 canonical work pages · cited by 1 Pith paper

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Shaaron Ainsworth. 1999. The functions of multiple representations. Computers & education 33, 2-3 (1999), 131–152

  3. [3]

    Shaaron Ainsworth. 2008. The educational value of multiple-representations when learning complex scientific concepts. In Visualization: Theory and practice in science education. Springer, 191–208

  4. [4]

    Scott Alexander. 2022. The Media Very Rarely Lies. Astral Codex Ten. https: //www.astralcodexten.com/p/the-media-very-rarely-lies Accessed: 2024-08-17

  5. [5]

    Hunt Allcott and Matthew Gentzkow. 2017. Social media and fake news in the 2016 election. Journal of economic perspectives 31, 2 (2017), 211–236

  6. [6]

    Jennifer Allen, Baird Howland, Markus Mobius, David Rothschild, and Duncan J Watts. 2020. Evaluating the fake news problem at the scale of the information ecosystem. Science advances 6, 14 (2020), eaay3539

  7. [7]

    Jennifer Allen, Duncan J Watts, and David G Rand. 2024. Quantifying the impact of misinformation and vaccine-skeptical content on Facebook. Science 384, 6699 (2024), eadk3451

  8. [8]

    AllSides. 2024. https://www.allsides.com/media-bias/media-bias-rating- methods. Accessed: 2024-09-05

Show all 117 references
  1. [9]

    Michelle A Amazeen, Arunima Krishna, and Rob Eschmann. 2022. Cutting the bunk: Comparing the solo and aggregate effects of prebunking and debunking COVID-19 vaccine misinformation. Science Communication 44, 4 (2022), 387– 417

  2. [10]

    Maryam Amirizaniani, Jihan Yao, Adrian Lavergne, Elizabeth Snell Okada, Aman Chadha, Tanya Roosta, and Chirag Shah. 2024. Developing a framework for auditing large language models using human-in-the-loop. arXiv preprint arXiv:2402.09346 (2024)

  3. [11]

    Matthew A Baum and Tim Groeling. 2008. New media and the polarization of American political discourse. Political Communication 25, 4 (2008), 345–365

  4. [12]

    Md Momen Bhuiyan, Sang Won Lee, Nitesh Goyal, and Tanushree Mitra. 2023. Newscomp: Facilitating diverse news reading through comparative annotation. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17

  5. [13]

    Biasly. 2024. https://www.biasly.com/bias-meter-how-it-works/. Accessed: 2024-09-05

  6. [14]

    Dylan Bourgeois, Jérémie Rappaz, and Karl Aberer. 2018. Selection bias in news coverage: learning it, fighting it. In Companion Proceedings of the The Web Conference 2018. 535–543

  7. [15]

    Curtis Bram. 2024. Beyond partisan filters: Can underreported news reduce issue polarization? Plos one 19, 2 (2024), e0297808

  8. [16]

    Ivar Bråten and Helge I Strømsø. 2011. Measuring strategic processing when students read multiple texts. Metacognition and Learning 6 (2011), 111–130

  9. [17]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101

  10. [18]

    Brown and Rachel R

    Danielle K. Brown and Rachel R. Mourão. 2021. Protest Coverage Matters: How Media Framing and Visual Communication Affects Sup- port for Black Civil Rights Protests. Mass Communication and Soci- ety 24, 4 (2021), 576–596. https://doi.org/10.1080/15205436.2021.1884724 arXiv:htt...

  11. [19]

    Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)

  12. [20]

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al

  13. [21]

    Ceren Budak, Sharad Goel, and Justin M Rao. 2016. Fair and balanced? Quantify- ing media bias through crowdsourced content analysis.Public Opinion Quarterly 80, S1 (2016), 250–271. Media Bias Detector CHI ’25, April 26-May 1, 2025, Yokohama, Japan

  14. [22]

    sticky notes

    Heather Burgess, Kate Jongbloed, Anna Vorobyova, Sean Grieve, Sharyle Lyn- don, Tim Wesseling, Kate Salters, Robert S Hogg, Surita Parashar, and Margo E Pearce. 2021. The “sticky notes” method: Adapting interpretive description methodology for team-based qualitative analysis i...

  15. [23]

    Carroll and Caroline Carrithers

    John M. Carroll and Caroline Carrithers. 1984. Training wheels in a user interface. Commun. ACM 27, 8 (aug 1984), 800–806. https://doi.org/10.1145/ 358198.358218

  16. [24]

    Man-pui Sally Chan, Christopher R Jones, Kathleen Hall Jamieson, and Dolores Albarracín. 2017. Debunking: A meta-analysis of the psychological efficacy of messages countering misinformation. Psychological science 28, 11 (2017), 1531–1546

  17. [25]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  18. [26]

    Media Bias/Fact Check. 2024. https://mediabiasfactcheck.com/left-vs-right- bias-how-we-rate-the-bias-of-media-sources/. Accessed: 2024-09-05

  19. [27]

    Chun-Fang Chiang and Brian Knight. 2011. Media bias and influence: Evidence from newspaper endorsements. The Review of economic studies 78, 3 (2011), 795–820

  20. [28]

    Dennis Chong and James N Druckman. 2007. Framing theory. Annu. Rev. Polit. Sci. 10, 1 (2007), 103–126

  21. [29]

    Dave D’Alessio and Mike Allen. 2000. Media bias in presidential elections: A meta-analysis. Journal of communication 50, 4 (2000), 133–156

  22. [30]

    James W Dearing and Everett Rogers. 1996. Agenda-setting. Sage publications

  23. [31]

    Tilman Dingler, Benjamin Tag, Evangelos Karapanos, Koichi Kise, and Andreas Dengel. 2020. Workshop on Detection and Design for Cognitive Biases in People and Computing Systems. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI...

  24. [32]

    Jakob-Moritz Eberl, Hajo G Boomgaarden, and Markus Wagner. 2017. One bias fits all? Three types of media bias and their effects on party preferences. Communication Research 44, 8 (2017), 1125–1148

  25. [33]

    Robert M Entman. 1993. Framing: Toward clarification of a fractured paradigm. Journal of communication 43, 4 (1993), 51–58

  26. [34]

    Motahhare Eslami, Karrie Karahalios, Christian Sandvig, Kristen Vaccaro, Aimee Rickman, Kevin Hamilton, and Alex Kirlik. 2016. First I "like" it, then I hide it: Folk Theories of Social Feeds. Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (2016)....

  27. [35]

    Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. In Proceedings of the 61st Annual Meeting of the Association for Computation...

  28. [36]

    Kiran Garimella, Tim Smith, Rebecca Weiss, and Robert West. 2021. Political polarization in online news consumption. In Proceedings of the International AAAI Conference on Web and Social Media , Vol. 15. 152–162

  29. [37]

    Matthew Gentzkow and Jesse M Shapiro. 2006. Media bias and reputation. Journal of political Economy 114, 2 (2006), 280–316

  30. [38]

    Matthew Gentzkow and Jesse M Shapiro. 2010. What drives media slant? Evidence from US daily newspapers. Econometrica 78, 1 (2010), 35–71

  31. [39]

    Matthew Gentzkow, Jesse M Shapiro, and Daniel F Stone. 2015. Media bias in the marketplace: Theory. In Handbook of media economics . Vol. 1. Elsevier, 623–645

  32. [40]

    Angela A Geraci, Ardith Brunt, and Cindy Marihart. 2014. The work behind weight-loss surgery: a qualitative analysis of food intake after the first two years post-op. International Scholarly Research Notices 2014, 1 (2014), 427062

  33. [41]

    Akshay Goel, Almog Gueta, Omry Gilon, Chang Liu, Sofia Erell, Lan Huong Nguyen, Xiaohong Hao, Bolous Jaber, Shashir Reddy, Rupesh Kartha, et al

  34. [42]

    Felix Hamborg, Karsten Donnay, and Bela Gipp. 2019. Automated identification of media bias in news articles: an interdisciplinary literature review.International Journal on Digital Libraries 20, 4 (2019), 391–415

  35. [43]

    InMachine Learning for Health (ML4H)

    Llms accelerate annotation for medical information extraction. InMachine Learning for Health (ML4H) . PMLR, 82–100

  36. [44]

    Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting , Vol. 50. Sage publications Sage CA: Los Angeles, CA, 904–908

  37. [45]

    Tony Harcup and Deirdre O’neill. 2017. What is news? News values revisited (again). Journalism studies 18, 12 (2017), 1470–1488

  38. [46]

    Homa Hosseinmardi, Samuel Wolken, David M Rothschild, and Duncan J Watts

  39. [47]

    Jan Hartmann, Antonella De Angeli, and Alistair Sutcliffe. 2008. Framing the user experience: information biases on website quality judgement. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Florence, Italy) (CHI ’08). Association for Computing Ma...

  40. [48]

    Juan José Igartua and Lifen Cheng. 2009. Moderating Effect of Group Cue While Processing News on Immigration: Is the Framing Effect a Heuristic Process? Journal of Communication 59 (2009), 726–749. https://api.semanticscholar.org/ CorpusID:59474192

  41. [49]

    arXiv preprint arXiv:2310.18863 (2023)

    The diminishing state of shared reality on US television news. arXiv preprint arXiv:2310.18863 (2023)

  42. [50]

    Hong Huang, Hua Zhu, Wenshi Liu, Hua Gao, Hai Jin, and Bang Liu. 2024. Uncovering the essence of diverse media biases from the semantic embedding space. Humanities and Social Sciences Communications 11, 1 (2024), 1–12

  43. [51]

    Oliver Kiss, Markus Mobius, and Tanya Rosenblat. 2023. Assembling News Like Legos. (2023)

  44. [52]

    Shanto Iyengar and Kyu S Hahn. 2009. Red media, blue media: Evidence of ideological selectivity in media use. Journal of communication 59, 1 (2009), 19–39

  45. [53]

    Emily K. Johnson. 2022. Miro, Miro: Student perceptions of a visual discussion board. In Proceedings of the 40th ACM International Conference on Design of Com- munication (Boston, MA, USA) (SIGDOC ’22). Association for Computing Ma- chinery, New York, NY, USA, 96–101. https://...

  46. [54]

    George Lifchits, Ashton Anderson, Daniel G Goldstein, Jake M Hofman, and Duncan J Watts. 2021. Success stories cause false beliefs about success.Judgment and Decision Making 16, 6 (2021), 1439–1463

  47. [55]

    David M. J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Penny- cook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonat...

  48. [56]

    Robertson, Mary Czerwinski, and Desney S

    Bongshin Lee, Greg Smith, George G. Robertson, Mary Czerwinski, and Desney S. Tan. 2009. FacetLens: exposing trends and relationships to support sensemaking within faceted datasets. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, MA, USA)...

  49. [57]

    Yixin Liu, Kejian Shi, Katherine S He, Longtian Ye, Alexander R Fabbri, Pengfei Liu, Dragomir Radev, and Arman Cohan. 2023. On learning to summarize with large language models as references. arXiv preprint arXiv:2305.14239 (2023)

  50. [58]

    Sora Lim, Adam Jatowt, Michael Färber, and Masatoshi Yoshikawa. 2020. An- notating and analyzing biased sentences in news articles using crowdsourcing. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 1478–1484

  51. [59]

    Yixin Liu, Kejian Shi, Katherine He, Longtian Ye, Alexander Fabbri, Pengfei Liu, Dragomir Radev, and Arman Cohan. 2024. On Learning to Summarize with Large Language Models as References. InProceedings of the 2024 Conference of the North American Chapter of the Association for ...

  52. [60]

    Maxwell McCombs and DL Shaw. 2005. The agenda-setting function of the press. The Press. Oxford, England: Oxford University Press Inc (2005), 156–168

  53. [61]

    Jörg Matthes. 2009. What’s in a frame? A content analysis of media framing studies in the world’s leading communication journals, 1990-2005. Journalism & mass communication quarterly 86, 2 (2009), 349–367

  54. [62]

    Nolwenn Maudet. 2019. Dead Angles of Personalization: Integrating Curation Algorithms in the Fabric of Design. In Proceedings of the 2019 on Designing Interactive Systems Conference (San Diego, CA, USA) (DIS ’19). Association for Computing Machinery, New York, NY, USA, 1439–14...

  55. [63]

    Ad Fontes Media. n.d.. Interactive Media Bias Chart. Retrieved from https://adfontesmedia.com/interactive-media-bias-chart/

  56. [64]

    McCombs and Donald L

    Maxwell E. McCombs and Donald L. Shaw. 1972. The Agenda-Setting Function of Mass Media. The Public Opinion Quarterly 36, 2 (1972), 176–187. http: //www.jstor.org/stable/2747787

  57. [65]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 72 (nov 2019), 23 pages. https://doi.org/10.1145/3359174

  58. [66]

    Amy Mitchell, Jeffrey Gottfried, Galen Stocking, Katerina Eva Matsa, and Eliza- beth Grieco. 2017. Covering President Trump in a polarized media environment. Pew Research Center (2017)

  59. [67]

    Miro. 2024. https://miro.com/. Accessed: 2024-09-05

  60. [68]

    A Mitchell, J Gottfried, M Barthel, N Sumida, and A Mitchell. 2018. Distinguish- ing between factual and opinion statements in the news. Pew Research Center’s Journalism Project

  61. [69]

    Daniel Muise, Homa Hosseinmardi, Baird Howland, Markus Mobius, David Roth- schild, and Duncan J. Watts. 2022. Quantifying partisan news diets in Web and TV audiences. Science Advances 8, 28 (2022), eabn0083. https://doi.org/10.1126/ sciadv.abn0083 arXiv:https://www.science.org...

  62. [70]

    Corman, and Huan Liu

    Fred Morstatter, Liang Wu, Uraz Yavanoglu, Stephen R. Corman, and Huan Liu

  63. [71]

    Abu Naser

    Md. Abu Naser. 2020. Relevance and Challenges of the Agenda-Setting Theory in the Changed Media Landscape. https://api.semanticscholar.org/CorpusID: 245116933

  64. [72]

    Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2024. More human than human: measuring ChatGPT political bias. Public Choice 198, 1 (2024), 3–23

  65. [73]

    Janni Nielsen, Torkil Clemmensen, and Carsten Yssing. 2002. Getting access to what goes on in people’s heads? Reflections on the think-aloud technique. In Proceedings of the second Nordic conference on Human-computer interaction . 101–110

  66. [74]

    Lloyd H Nakatani and John A Rohrlich. 1983. Soft machines: A philosophy of user-computer interface design. In Proceedings of the SIGCHI conference on Human Factors in Computing Systems . 19–23

  67. [75]

    Deirdre O’neill and Tony Harcup. 2009. News values and selectivity. In The handbook of journalism studies . Routledge, 181–194

  68. [76]

    Ground News. 2024. https://ground.news/rating-system. Accessed: 2024-09-05

  69. [77]

    Sungkyu Park, Jaimie Yejean Park, Jeong-han Kang, and Meeyoung Cha. 2021. The presence of unexpected biases in online fact-checking.The Harvard Kennedy School Misinformation Review (2021)

  70. [78]

    Maria Ojala, Ashlee Cunsolo, Charles A Ogunbode, and Jacqueline Middleton

  71. [79]

    Gordon Pennycook and David G Rand. 2021. The psychology of fake news. Trends in cognitive sciences 25, 5 (2021), 388–402

  72. [80]

    Gordon Pennycook and David G Rand. 2022. Nudging social media toward accuracy. The Annals of the American Academy of Political and Social Science 700, 1 (2022), 152–164

  73. [81]

    Souneil Park, Seungwoo Kang, Sangyoung Chung, and Junehwa Song. 2009. NewsCube: delivering multiple aspects of news to mitigate media bias. In Pro- ceedings of the SIGCHI conference on human factors in computing systems . 443– 452

  74. [82]

    Riccardo Puglisi and James M Snyder Jr. 2015. Empirical studies of media bias. In Handbook of media economics . Vol. 1. Elsevier, 647–667

  75. [83]

    Alejandro Peña, Aythami Morales, Julian Fierrez, Ignacio Serna, Javier Ortega- Garcia, Iñigo Puente, Jorge Cordova, and Gonzalo Cordova. 2023. Leveraging large language models for topic classification in the domain of public affairs. In International Conference on Document Ana...

  76. [84]

    van den Haak, and Menno D

    Judith Ramey, Ted Boren, Elisabeth Cuddihy, Joe Dumas, Zhiwei Guan, Maaike J. van den Haak, and Menno D. T. De Jong. 2006. Does think aloud work? how do we know?. InCHI ’06 Extended Abstracts on Human Factors in Computing Systems (Montréal, Québec, Canada)(CHI EA ’06). Associa...

  77. [85]

    Mathieu Ravaut, Aixin Sun, Nancy Chen, and Shafiq Joty. 2024. On Con- text Utilization in Summarization with Large Language Models. In Proceed- ings of the 62nd Annual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martin...

  78. [86]

    Emily Pronin, Thomas Gilovich, and Lee Ross. 2004. Objectivity in the eye of the beholder: divergent perceptions of bias in self versus others. Psychological review 111, 3 (2004), 781

  79. [87]

    Claire E Robertson, Nicolas Pröllochs, Kaoru Schwarzenegger, Philip Pärnamets, Jay J Van Bavel, and Stefan Feuerriegel. 2023. Negativity drives online news consumption. Nature Human Behaviour 7, 5 (2023), 812–822

  80. [88]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  81. [89]

    Jérôme Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, and Markus Pauly

  82. [90]

    David Canfield Smith, Charles Irby, Ralph Kimball, and Eric Harslem. 1982. The Star user interface: An overview. In Proceedings of the June 7-10, 1982, national computer conference. 515–528

  83. [91]

    Aaron Springer and Steve Whittaker. 2019. Progressive disclosure: empirically motivated approaches to designing effective transparency. In Proceedings of the 24th international conference on intelligent user interfaces . 107–120

  84. [92]

    Mohi Reza, Nathan M Laundry, Ilya Musabirov, Peter Dushniku, Zhi Yuan “Michael” Yu, Kashish Mittal, Tovi Grossman, Michael Liut, Anastasia Kuzminykh, and Joseph Jay Williams. 2024. ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing...

  85. [93]

    Statista. 2023. Leading news and media sites in the U.S. 2023, by share of visits. https://www.statista.com/statistics/381569/leading-news-and-media- sites-usa-by-share-of-visits/ Accessed: December 5, 2024

  86. [94]

    Rothschild, Elliot Pickens, Gideon Heltzer, Jenny Wang, and Duncan J

    David M. Rothschild, Elliot Pickens, Gideon Heltzer, Jenny Wang, and Duncan J. Watts. 2023. Warped Front Pages. Columbia Journalism Review (2023)

  87. [95]

    Pál Susánszky, Ákos Kopper, and Frank T Zsigó. 2022. Media framing of political protests–reporting bias and the discrediting of political activism. Post-Soviet Affairs 38, 4 (2022), 312–328

  88. [96]

    The self-perception and political biases of ChatGPT. arXiv. preprint]. doi 10 (2023)

  89. [97]

    Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science 359, 6380 (2018), 1146–1151. https://doi.org/10.1126/science. aap9559 arXiv:https://www.science.org/doi/pdf/10.1126/science.aap9559

  90. [98]

    Xinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao

  91. [99]

    Steven A Stahl, Cynthia R Hynd, Bruce K Britton, Mary M McNish, and Dennis Bosquet. 1996. What happens when students read multiple source documents in history? Reading Research Quarterly 31, 4 (1996), 430–456

  92. [100]

    David H Weaver. 2007. Thoughts on agenda setting, framing, and priming. Journal of communication 57, 1 (2007), 142–147

  93. [101]

    Natalie Jomini Stroud. 2008. Media use and political predispositions: Revisiting the concept of selective exposure. Political behavior 30 (2008), 341–366

  94. [102]

    Derong Xu, Wei Chen, Wenjun Peng, Chao Zhang, Tong Xu, Xiangyu Zhao, Xian Wu, Yefeng Zheng, and Enhong Chen. 2023. Large language models for generative information extraction: A survey. arXiv preprint arXiv:2312.17617 (2023)

  95. [103]

    Adaku Uchendu, Jooyoung Lee, Hua Shen, Thai Le, Dongwon Lee, et al. 2023. Does human collaboration enhance the accuracy of identifying llm-generated deepfake texts?. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 11. 163–174

  96. [104]

    Kun Yu, Shlomo Berkovsky, Ronnie Taib, Dan Conway, Jianlong Zhou, and Fang Chen. 2017. User Trust Dynamics: An Investigation Driven by Differences in System Performance. In Proceedings of the 22nd International Conference on Intelligent User Interfaces (Limassol, Cyprus) (IUI ...

  97. [105]

    Liudmila Zavolokina, Kilian Sprenkamp, Zoya Katashinskaya, Daniel Gordon Jones, and Gerhard Schwabe. 2024. Think Fast, Think Slow, Think Critical: Designing an Automated Propaganda Detection Tool. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Hono...

  98. [106]

    Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al. 2023. Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792 (2023)

  99. [107]

    Watts, David M

    Duncan J. Watts, David M. Rothschild, and Markus Mobius. 2021. Measuring the news and its impact on democracy. Proceedings of the National Academy of Sciences 118, 15 (2021), e1912443118. https://doi.org/10.1073/pnas.1912443118 arXiv:https://www.pnas.org/doi/pdf/10.1073/pnas.1...

  100. [108]

    Vera Liao, and Rachel K

    Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy. 2020. Effect of confi- dence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona, Spain) (FAT* ’2...

  101. [109]

    Yi Xiang and Miklos Sarvary. 2007. News consumption and media bias. Market- ing Science 26, 5 (2007), 611–628

  102. [111]

    Ka-Ping Yee, Kirsten Swearingen, Kevin Li, and Marti Hearst. 2003. Faceted metadata for image search and browsing. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Ft. Lauderdale, Florida, USA) (CHI ’03). Association for Computing Machinery, New Yo...

  103. [115]

    Wenxuan Zhang, Yue Deng, Bing Liu, Sinno Jialin Pan, and Lidong Bing. 2023. Sentiment analysis in the era of large language models: A reality check. arXiv preprint arXiv:2305.15005 (2023)

  104. [117]

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can large language models transform computational social science? Computational Linguistics 50, 1 (2024), 237–291. Media Bias Detector CHI ’25, April 26-May 1, 2025, Yokohama, Japan A Append...

  105. [2018]

    Identifying Framing Bias in Online News. Trans. Soc. Comput. 1, 2, Article 5 (jun 2018), 18 pages. https://doi.org/10.1145/3204948 CHI ’25, April 26-May 1, 2025, Yokohama, Japan Wang et al

  106. [2021]

    Annual review of environment and resources 46, 1 (2021), 35–58

    Anxiety, worry, and grief in a time of environmental and climate crisis: A narrative review. Annual review of environment and resources 46, 1 (2021), 35–58

  107. [2023]

    arXiv preprint arXiv:2303.12712 (2023)

    Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712 (2023)

  108. [2024]

    In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24)

    Human-LLM Collaborative Annotation Through Effective Verification of LLM Labels. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 303, 21 pages. https://doi...

  109. [2781]

    https://aclanthology.org/2024.acl-long.153

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.