REVIEW 3 major objections 5 minor 45 references
News Source Citing Patterns in AI Search Systems
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper uses 24,000 real user conversations to show that AI search engines concentrate news citations among a few outlets, skew left, and face no user preference penalty for doing so.
desk verdict A solid observational audit of real AI search citations; the concentration and provider-difference findings are well supported, but the 'liberal bias' claim needs a baseline before it can carry the gatekeeping narrative. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the AI Search Arena dataset of real user-AI conversations with head-to-head ratings, from which citations are extracted and matched to news domains. Political leaning comes from audience-based scores tied to voter registration records, and quality comes from aggregated expert and fact-checker ratings. Concentration is measured with the Gini index on citation counts across domains, and user preference is modeled with the Bradley-Terry paired-comparison model, which estimates the probability that one response is preferred over another. This combination lets the paper observe citation behavior in naturalistic use and connect it to user choices.
What would settle it
A direct test: obtain product telemetry or a representative panel of general users from the same providers and measure the same citation concentration, political leaning, and user-preference associations. If the concentration and left-leaning bias shrink or user satisfaction moves with source leaning or quality in that sample, the paper's claim that these patterns characterize real-world AI search usage would be overturned.
Extended reading notes
Core claim
Using over 24,000 conversations and 366,087 citations from 12 AI search models by OpenAI, Perplexity, and Google, the paper finds that citation patterns are shared across providers: news citations concentrate heavily, with Gini coefficients of 0.83 for OpenAI, 0.77 for Perplexity, and 0.69 for Google, and the top 20 outlets account for 67.3% of OpenAI news citations. Left-leaning and center sources make up over 98% of news citations while right-leaning sources are under 1.2%, and low-credibility outlets are rarely cited, with high-quality sources dominating at 89.7% to 96.2% across model families. The paper also reports that in 1,534 head-to-head comparisons, the political leaning and quality of cited news sources show no significant association with user preference, while response length remains the main driver.
Load-bearing premise
The load-bearing premise is that the AI Search Arena dataset represents real-world AI search usage; it is a self-selected head-to-head evaluation platform, and if its users and queries differ systematically from the general public, the citation and preference findings describe the platform rather than all AI search systems.
Editorial extensions
If this is right
- AI search providers control news exposure more than traditional ranked search, since a few outlets receive most citations and users do not adjust their preferences accordingly.
- If users do not respond to the leaning or quality of cited sources, platforms face little competitive incentive to diversify or balance news citations.
- Choosing a provider materially changes the news diet: OpenAI users see fewer but higher-quality sources, while Google and Perplexity users see broader but somewhat lower-quality sets.
- Regulatory or design interventions aimed at source diversity would need to target the systems themselves rather than rely on user demand.
Reading between the lines
- Editorial inference: if the Arena population skews toward technically sophisticated users, the left-leaning bias in the general population could be stronger or weaker; the paper does not test this.
- Editorial inference: the null user-preference result suggests citation source characteristics may be nearly invisible in the interface; a follow-up experiment that makes source labels salient could reveal whether users care when they notice.
- Editorial inference: the same concentration methodology could be applied to non-news citations, which make up 91% of citations and may show different bias patterns.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes 366,087 citations extracted from 24,069 conversations in the AI Search Arena dataset, focusing on the 9% of citations that point to news domains. It identifies 1,795 news outlets, measures citation concentration with Gini coefficients and Lorenz curves, characterizes cited outlets with audience-based political-leaning scores (DomainDemo) and aggregated quality ratings (Lin et al.), and runs linear regressions to relate citation composition to question and response features. A Bradley-Terry analysis of 1,534 head-to-head user judgments examines whether news citation characteristics are associated with user preference. The reported findings are: same-provider models cite similar news domains; citation is heavily concentrated; cited sources skew left/center on the audience-based scale; high-quality sources dominate; and user preference is not significantly associated with the political leaning or quality of cited news sources. The paper frames these results as evidence that AI search systems act as unaccountable news gatekeepers.
Significance. If its descriptive claims hold, the paper is a valuable large-scale observational study of AI search citation behavior, with real user interactions, cross-provider coverage, and publicly available code and datasets. The concentration, provider-difference, and quality-preference findings are well supported by the data. The central 'liberal bias' claim, however, is not established because the paper lacks a baseline against which to call the observed composition biased. The external-validity claim that the Arena dataset represents real-world AI search usage is also asserted rather than demonstrated. The user-preference null result is used to draw strong policy conclusions despite being an absence-of-evidence finding. These issues are fixable, and the paper would be a solid contribution after they are addressed.
major comments (3)
- [Political Leaning and Quality of News Sources] The claim that 'all three model families exhibit a pronounced left-leaning bias in their news source citations' (also echoed in the Abstract) is not supported as a bias claim because no baseline is provided. Figure 3 and the Table 5 intercepts (18.5% left-leaning vs. 2.5% right-leaning citations) describe only the composition of the cited news domains. A composition is evidence of bias only relative to a benchmark, such as the distribution of political leaning across all news domains in the Yang (2025) list, the domains available in the retrieval index, the domains matching the user queries, or the citations returned by a traditional search engine for the same queries. Without such a baseline, the observed asymmetry is consistent with the alternative that the supply of news domains, or the set of high-quality news domains, is itself left/center-skewed on the audience-based DomainDemo scale. This issue is load-bearing because the liberal-bias finding is central to the gatekeeping narrative.
- [Data and Methods: AI Search Arena Dataset] The statement that 'the dataset captures real-world usage patterns' is not justified. The AI Search Arena is an opt-in head-to-head comparison platform in which users self-select and choose queries for side-by-side evaluation; these users and queries are likely not representative of general AI search product usage. The paper should provide evidence about representativeness, such as a comparison with product telemetry, user demographics, or query distributions, or it should consistently qualify the findings as descriptive of the Search Arena population. This concern affects the external validity of all provider-level citation rates, concentration, and preference results.
- [User Preference Analysis] The null result in Figure 5(c) is used to conclude that 'neither the political leaning nor the quality of cited news sources significantly influences user satisfaction' (Abstract) and later that 'the gatekeeping power of these systems remains largely unchallenged by end-users' (Implications). This is an absence-of-evidence claim based on 1,534 comparisons, a single feature specification, and conversation-level averaging of response variables. The paper should report the estimated coefficients, standard errors, and confidence intervals numerically, and it should either demonstrate adequate power to detect meaningful effects or soften the causal interpretation of the null.
minor comments (5)
- [Data and Methods: News Outlet Political Leaning] The DomainDemo measure is audience-based, not a content-based measure of outlet ideology; the paper sometimes uses 'political leaning' and 'bias' interchangeably. Clarify that Figure 3 shows audience-based leaning, which is an important interpretive distinction.
- [Data and Methods: News Outlet Political Leaning and Quality] The classification thresholds (-0.33, 0.33, and 0.5) are presented without justification. Report sensitivity analyses with alternative thresholds, or at least explain why these particular values are appropriate for the research questions.
- [Regression Analysis / Table 5] Table 5 reports R-squared values between 0.05 and 0.13, and the Discussion statement that 'the observed citation patterns stem primarily from the intrinsic characteristics of the AI systems themselves rather than from the nature of user questions' is stronger than the evidence. The model-family coefficients are small relative to the intercepts, but the low explained variance does not directly establish an intrinsic mechanism.
- [Appendix: Overrepresented News Domains] The log-odds ratio method uses the aggregated citation frequencies across all three model families as the background corpus. This choice may introduce circularity because the background is constructed from the same data being compared; state the interpretation of this choice or use an external background corpus.
- [References and formatting] The Appendix equation numbering skips from (2) to the second equation, and the GitHub URL in the Abstract contains a space ('ai search arena'), which may prevent the link from resolving correctly.
Circularity Check
No significant circularity: the empirical chain is descriptive and the author's prior datasets are externally validated.
full rationale
The paper's derivation is observational, not constructional: it labels citations with externally defined attributes and regresses those labels on covariates. The only author-self-citations that are load-bearing are the news-domain list (Yang 2025) and the DomainDemo political-leaning scores (Yang et al. 2025). Both are public artifacts assembled from independent external sources (NewsGuard, MBFC, ABYZ, and Media Cloud for the domain list; a voter-registration-matched Twitter panel and multiple validation sets for DomainDemo), so they qualify as independent support rather than circular inputs. The 'pronounced left-leaning bias' claim is a descriptive statement about the distribution of labeled citations; it may be criticized for lacking an explicit baseline of the available news supply or query distribution, but a missing baseline is a correctness/interpretation concern, not a circular reduction. The regression and Bradley-Terry analyses are explanatory/null-result models; no fitted parameter is renamed as a prediction, and no equation is defined in terms of its own output. No uniqueness theorem or ansatz is imported from the author's prior work to force the conclusion. Accordingly, the central findings have content independent of their inputs.
Assumptions & free parameters
free parameters (4)
- Political leaning classification thresholds =
-0.33 and 0.33
- Quality classification threshold =
0.5
- Number of embedding principal components =
10
- Number of BERTopic topics =
10
assumptions (5)
- domain assumption The AI Search Arena dataset captures real-world AI search usage patterns.
- domain assumption DomainDemo audience-based political leaning scores are valid measures of outlet ideology.
- domain assumption Lin et al. (2023) quality ratings are valid measures of news source quality.
- domain assumption The news domain list from Yang (2025) correctly identifies news outlets.
- standard math Standard statistical specifications (Gini, OLS, Bradley-Terry, log-odds ratios) are appropriate for these data.
Cite this review
Pith. "Pith review of News Source Citing Patterns in AI Search Systems." pith.science (2026). https://pith.science/paper/BMIWSHVR
@misc{pith2026250705301,
author = {Pith},
title = {Pith review of: News Source Citing Patterns in AI Search Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/BMIWSHVR}},
note = {Machine review of arXiv:2507.05301}
}
read the original abstract
AI-powered search systems are emerging as new information gatekeepers, fundamentally transforming how users access news and information. Despite their growing influence, the citation patterns of these systems remain poorly understood. We address this gap by analyzing data from the AI Search Arena, a head-to-head evaluation platform for AI search systems. The dataset comprises over 24,000 conversations and 65,000 responses from models across three major providers: OpenAI, Perplexity, and Google. Among the over 366,000 citations embedded in these responses, 9% reference news sources. We find that while models from different providers cite distinct news sources, they exhibit shared patterns in citation behavior. News citations concentrate heavily among a small number of outlets and display a pronounced liberal bias, though low-credibility sources are rarely cited. User preference analysis reveals that neither the political leaning nor the quality of cited news sources significantly influences user satisfaction. These findings reveal significant challenges in current AI search systems and have important implications for their design and governance.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Bakshy, E.; Messing, S.; and Adamic, L. A. 2015. Exposure to ideologically diverse news and opinion on Facebook . Science, 348(6239): 1130--1132
work page 2015
-
[4]
Barzilai-Nahon, K. 2008. Toward a theory of network gatekeeping: A framework for exploring information control. Journal of the American Society for Information Science and Technology, 59(9): 1493--1512
work page 2008
-
[5]
Bradley, R. A.; and Terry, M. E. 1952. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4): 324--345
work page 1952
-
[6]
Brantner, C.; Karlsson, M.; and Kuai, J. 2025. Sourcing behavior and the role of news media in AI-powered search engines in the digital media ecosystem: Comparing political news retrieval across five languages. Telecommunications Policy, 49(5): 102952
work page 2025
-
[7]
Buntain, C.; Bonneau, R.; Nagler, J.; and Tucker, J. A. 2023. Measuring the Ideology of Audiences for Web Links and Domains Using Differentially Private Engagement Data. Proceedings of the International AAAI Conference on Web and Social Media, 17(1): 72--83
work page 2023
-
[8]
N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M
Chiang, W.-L.; Zheng, L.; Sheng, Y.; Angelopoulos, A. N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M. I.; Gonzalez, J. E.; and Stoica, I. 2024. Chatbot arena: an open platform for evaluating LLMs by human preference. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org
work page 2024
Show all 45 references
-
[9]
Cronin, J.; Clemm von Hohenberg, B.; Gon c alves, J. F. F.; Menchen-Trevino, E.; and Wojcieszak, M. 2023. The (null) over-time effects of exposure to local news websites: Evidence from trace data. Journal of Information Technology & Politics, 1--15
2023
-
[10]
A.; and Nagler, J
Eady, G.; Bonneau, R.; Tucker, J. A.; and Nagler, J. 2025. News Sharing on Social Media: Mapping the Ideology of News Media, Politicians, and the Mass Public. Political Analysis, 33(2): 73--90
2025
-
[11]
Fischer, S.; Jaidka, K.; and Lelkes, Y. 2020. Auditing local news presence on Google News . Nature Human Behaviour, 4(12): 1236--1244
2020
-
[12]
Grootendorst, M. 2020. KeyBERT: Minimal keyword extraction with BERT
2020
-
[13]
Grootendorst, M. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794
2022 arXiv
-
[14]
Hagar, N.; and Diakopoulos, N. 2025. AI Overviews, Chatbots, and News Traffic: The Story So Far. https://generative-ai-newsroom.com/ai-overviews-chatbots-and-news-traffic-the-story-so-far-c010b3bf53cb (Accessed: 2025-07-01)
2025
-
[15]
D.; Gruppi, M.; Joseph, K.; Green, J.; Wihbey, J
Horne, B. D.; Gruppi, M.; Joseph, K.; Green, J.; Wihbey, J. P.; and Adal , S. 2022. NELA-Local: A Dataset of US Local News Articles for the Study of County-Level News Ecosystems. In Proceedings of the International AAAI Conference on Web and Social Media, volume 16, 1275--1284
2022
-
[16]
Jaidka, K.; and Furniturewala, S. 2025. The News with ChatGPT: An Audit and Survey Experiment on the Effects of GPT-Enabled News Search on User Attitudes
2025
-
[17]
V.; and Romano, S
Kuai, J.; Brantner, C.; Karlsson, M.; Couvering, E. V.; and Romano, S. 2025. AI chatbot accountability in the age of algorithmic gatekeeping: Comparing generative search engine political information retrieval across five languages. New Media & Society, 0(0): 14614448251321162
2025
-
[18]
T.; Simchon, A.; Carrella, F.; Garcia, D.; and Lewandowsky, S
Lasser, J.; Aroyehun, S. T.; Simchon, A.; Carrella, F.; Garcia, D.; and Lewandowsky, S. 2022. Social media sharing by political elites: An asymmetric American exceptionalism. arXiv:2207.06313
2022 arXiv
-
[19]
A.; Chiang, T.-W.; and Naaman, M
Le Qu \'e r \'e , M. A.; Chiang, T.-W.; and Naaman, M. 2022. Understanding Local News Social Coverage and Engagement at Scale during the COVID -19 Pandemic. In Proceedings of the International AAAI Conference on Web and Social Media, volume 16, 560--572
2022
-
[20]
Li, A.; and Sinnamon, L. 2024. Generative AI Search Engines as Arbiters of Public Knowledge: An Audit of Bias and Authority. Proceedings of the Association for Information Science and Technology, 61(1): 205--217
2024
-
[21]
Li, H.; and Aral, S. 2025. Human Trust in AI Search: A Large-Scale Experiment. arXiv:2504.06435
2025 arXiv
-
[22]
G.; and Pennycook, G
Lin, H.; Lasser, J.; Lewandowsky, S.; Cole, R.; Gully, A.; Rand, D. G.; and Pennycook, G. 2023. High level of correspondence across different news domain quality rating sets. PNAS nexus, 2(9): pgad286
2023
-
[23]
Lindemann, N. F. 2024. Chatbots, search engines, and the sealing of knowledges. AI & SOCIETY, 1--14
2024
-
[24]
A.; and West, J
Memon, S. A.; and West, J. D. 2024. Search Engines Post-ChatGPT: How Generative Artificial Intelligence Could Make Search Less Reliable. arXiv:2402.11707
2024 arXiv
-
[25]
N.; Darrell, T.; Norouzi, N.; and Gonzalez, J
Miroyan, M.; Wu, T.-H.; King, L.; Li, T.; Pan, J.; Hu, X.; Chiang, W.-L.; Angelopoulos, A. N.; Darrell, T.; Norouzi, N.; and Gonzalez, J. E. 2025. Search Arena: Analyzing Search-Augmented LLMs. arXiv:2506.05334
2025
-
[26]
L.; Colaresi, M
Monroe, B. L.; Colaresi, M. P.; and Quinn, K. M. 2008. Fightin'words: Lexical feature selection and evaluation for identifying the content of political conflict. Political Analysis, 16(4): 372--403
2008
-
[27]
Narayanan Venkit, P.; Laban, P.; Zhou, Y.; Mao, Y.; and Wu, C.-S. 2025. Search Engines in the AI Era: A Qualitative Understanding to the False Promise of Factual and Verifiable Source-Cited Responses in LLM-based Search. In Proceedings of the 2025 ACM Conference on Fairness, A...
2025
-
[28]
Pennycook, G.; and Rand, D. G. 2019. Fighting misinformation on social media using crowdsourced judgments of news source quality. Proceedings of the National Academy of Sciences, 116(7): 2521--2526
2019
-
[29]
Reimers, N.; and Gurevych, I. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084
2019 arXiv
-
[30]
E.; Green, J.; Ruck, D
Robertson, R. E.; Green, J.; Ruck, D. J.; Ognyanova, K.; Wilson, C.; and Lazer, D. 2023. Users choose to engage with more partisan news than they are exposed to on Google Search . Nature, 618(7964): 342--348
2023
-
[31]
E.; Jiang, S.; Joseph, K.; Friedland, L.; Lazer, D.; and Wilson, C
Robertson, R. E.; Jiang, S.; Joseph, K.; Friedland, L.; Lazer, D.; and Wilson, C. 2018. Auditing Partisan Audience Bias within Google Search . Proceedings of the ACM on Human-Computer Interaction, 2(CSCW)
2018
-
[32]
Shah, C.; and Bender, E. M. 2022. Situating Search. In Proceedings of the 2022 Conference on Human Information Interaction and Retrieval, CHIIR '22, 221--232. New York, NY, USA: Association for Computing Machinery. ISBN 9781450391863
2022
-
[33]
J.; and Vos, T
Shoemaker, P. J.; and Vos, T. 2009. Gatekeeping theory. Routledge
2009
-
[34]
Shu, K.; Sliva, A.; Wang, S.; Tang, J.; and Liu, H. 2017. Fake News Detection on Social Media: A Data Mining Perspective. SIGKDD Explor. Newsl., 19(1): 22–36
2017
-
[35]
E.; Rothschild, D
Spatharioti, S. E.; Rothschild, D. M.; Goldstein, D. G.; and Hofman, J. M. 2023. Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment. arXiv:2307.03744
2023 arXiv
-
[36]
W.; Andersen, R.; Buscher, G.; Manivannan, S.; Rangan, N.; and Yang, L
Suri, S.; Counts, S.; Wang, L.; Chen, C.; Wan, M.; Safavi, T.; Neville, J.; Shah, C.; White, R. W.; Andersen, R.; Buscher, G.; Manivannan, S.; Rangan, N.; and Yang, L. 2024. The Use of Generative Search Engines for Knowledge Work and Complex Tasks. arXiv:2404.04268
2024 arXiv
-
[37]
Tong, C. 2025. Unite or divide? Biased search queries and Google Search results in polarized politics. In Proceedings of the 17th ACM Web Science Conference 2025, Websci '25, 158–168. New York, NY, USA: Association for Computing Machinery. ISBN 9798400714832
2025
-
[38]
Trielli, D.; and Diakopoulos, N. 2019. Search as News Curator: The Role of Google in Shaping Attention to News Information. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI '19, 1–15. New York, NY, USA: Association for Computing Machinery. I...
2019
-
[39]
Trielli, D.; and Diakopoulos, N. 2022. Partisan search behavior and Google results in the 2018 U.S. midterm elections. Information, Communication & Society, 25(1): 145--161
2022
-
[40]
Van Dalen, A. 2023. Algorithmic gatekeeping for professional communicators: Power, trust, and legitimacy. Taylor & Francis
2023
-
[41]
B.; Croft, W
Wu, Z.; Sanderson, M.; Cambazoglu, B. B.; Croft, W. B.; and Scholer, F. 2020. Providing Direct Answers in Search Results: A Study of User Behavior. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, CIKM '20, 1635--1644. New York, NY...
2020
-
[42]
Xiong, H.; Bian, J.; Li, Y.; Li, X.; Du, M.; Wang, S.; Yin, D.; and Helal, S. 2024. When Search Engine Services Meet Large Language Models: Visions and Challenges. IEEE Transactions on Services Computing, 17(6): 4558--4577
2024
-
[43]
Yang, K.-C. 2025. A list of news domains. Zenodo
2025
-
[44]
D.; Grinberg, N.; Joseph, K.; and Lazer, D
Yang, K.-C.; Goel, P.; Quintana-Mathé, A.; Horgan, L.; McCabe, S. D.; Grinberg, N.; Joseph, K.; and Lazer, D. 2025. DomainDemo: a dataset of domain-sharing activities among different demographic groups on Twitter. arXiv:2501.09035
2025 arXiv
-
[45]
Yang, K.-C.; and Menczer, F. 2025. Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models. In Proceedings of the 17th ACM Web Science Conference, Websci '25, 127–137. New York, NY, USA: Association for Computing Machinery
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.