Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Search engines in polarized media environment: Auditing political information curation on Google and Bing prior to 2024 US elections

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Google and Bing skewed toward left-leaning news sources before the 2024 US election, an audit finds.

desk verdict Large, careful audit whose absolute 'left-leaning prioritization' claim rests on a debatable MBFC recoding, but the query-slant effect (Republican queries increase right-leaning sources) is real and consistent across engines. read the letter →

arxiv 2501.04763 v1 pith:FTTZVXMU submitted 2025-01-08 cs.CY cs.IRcs.SI

classification cs.CYcs.IRcs.SI
keywords searchengineauditpoliticalinformationcurationpartisangapideologicalslant2024USelectionsGoogleBingmediapolarization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that Google and Bing, audited in the six months before the 2024 US presidential election, systematically prioritized left-leaning journalistic media in top search results, and that the partisan slant of the query shifted the ideological composition of those results. Republican-focused queries consistently raised the share of right-leaning outlets, while Democratic-focused queries raised left-leaning outlets in most but not all conditions. If true, this means search engines are not neutral information gates but active mirrors of the polarized US media environment, with the potential to amplify perceived polarization among users.

What carries the argument

The study uses a virtual agent-based algorithm audit: cloud-hosted Firefox agents searched from fixed US IP addresses, cleared cookies between queries, and repeated the same queries monthly and daily in the run-up to the election. Media sources in the results were classified by type, then matched to MediaBiasFactCheck ratings which were recoded into binary left- and right-leaning dummy variables; generalized linear mixed-effects models with a logit link predicted the likelihood of encountering a left- or right-leaning source as a function of query type, controlling for date and location.

What would settle it

Re-run the audit using an independent measure of source ideology, such as expert panels or a different bias database, and check whether the overrepresentation of right-leaning media for Republican queries—odds ratios around 1.55 in daily organic results—persists. The claim would also be weakened if the left-right imbalance vanished once source quality, recency, or topical relevance were controlled for.

Watch

Extended reading notes

Core claim

The central claim is that information curation on Google and Bing during the 2024 US election cycle exhibited a partisan gap: roughly half of all journalistic media results were left-leaning across all query types, and right-leaning media were consistently overrepresented when queries named Republican candidates. Using odds ratios from mixed-effects models, right-leaning sources were 1.49 to 1.66 times more likely to appear in organic results for Republican queries in the monthly collection, and 1.55 to 1.57 times more likely in the daily collection, with even larger effects in Google's Newsblock (odds ratios of 2.79 to 3.13). Democratic queries increased left-leaning sources for Bing and Google organic results in the monthly data but not in the daily data, and never decreased right-leaning sources. The paper interprets this as search engines mirroring rather than correcting the partisan divides present in the US media environment.

Load-bearing premise

The paper assumes MediaBiasFactCheck's ideological labels, and the binary recoding of those labels into merely left-leaning or right-leaning, accurately measure the political leaning of news sources; if that coding is systematically off, the measured partisan gap could be an artifact of the labeler rather than of the search engines.

Editorial extensions

If this is right

  • Search engine users received an ideologically skewed media diet during the 2024 election, with left-leaning outlets dominating and right-leaning outlets surfacing mainly in response to Republican-focused queries.
  • The partisan gap in organic results was stable across US locations and over time, suggesting it is a structural feature of the engines' curation rather than a localized or transient artifact.
  • Google's Newsblock behaved differently and more erratically than organic results, flipping its partisan lean on adjacent days, so additional interface elements can amplify or counter the slant of organic results.
  • Source-level leaning does not directly measure whether a candidate is portrayed positively or negatively, so content-level analysis would be needed to assess tone.
  • The results give empirical support to public claims of search-engine skew, but they do not establish deliberate manipulation or a specific effect on voters.
  • Overrepresentation of right-leaning media for Republican queries was consistent across both engines, while Democratic queries never reduced right-leaning sources, suggesting an asymmetric response to query slant.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the MediaBiasFactCheck-based coding holds, the evidence suggests search engines are not diversifying political information but reproducing the existing imbalance of the US media ecosystem, which could strengthen perceptions of media bias among right-leaning users.
  • The asymmetric pattern—Republican queries polarize output more than Democratic queries—hints that the engines' response to query slant may be driven by the availability or authority of right-leaning sources, a hypothesis that could be tested in other elections and countries.
  • A natural extension is to analyze linked article content and tone rather than source-level ideology, to distinguish between visibility and sentiment in search results.
  • The audit used only cloud IP locations, missing most battleground states; expanding to purple-state locations is a concrete next test for the location-based claims.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper audits Google and Bing before the 2024 US presidential elections, using virtual agents to collect 305,901 search results from June to November 2024 across several US locations. The authors classify sources by type, then for journalistic sources measure political leaning by matching domains to the MediaBiasFactCheck (MBFC) database, recoding MBFC categories into binary left- and right-leaning indicators. They test whether Democratic-focused queries increase the presence of left-leaning media (H1a) and Republican-focused queries increase the presence of right-leaning media (H1b), and they examine moderation by location and time. The findings show that left-leaning sources are generally more prevalent overall, while H1b is consistently supported across engine, collection period, and interface element; H1a is only partially supported, with null results for several daily and Newsblock models. The paper concludes that search engines mirror partisan divides and may contribute to perceived polarization.

Significance. If the findings hold, this is a valuable large-scale algorithmic audit with several genuine strengths: a large and clearly described data collection, controlled virtual agents with randomized simultaneous queries, repeated daily and monthly collections, separate treatment of Google's Newsblock, and mixed-effects models with random intercepts by agent. The relative partisan-gap result for right-leaning media under Republican-focused queries is consistent and robust within the paper's coding, and it is a meaningful contribution to the search-engine-bias literature. The absolute claim that both engines 'prioritize left-leaning media sources' is more exposed to measurement assumptions, but the underlying data collection appears sound and the paper is transparent about many design limitations. These features make the manuscript a plausible candidate for publication after the robustness concerns below are addressed.

major comments (3)
  1. [Measures] The central absolute claim that both search engines 'prioritize left-leaning media sources' is directly determined by the decision to collapse MBFC's 'left-center' and 'left' ratings into a single left-leaning dummy. Because the most frequent sources in Tables B2.1 and B2.2 (cnn.com, nytimes.com, washingtonpost.com, nbcnews.com, etc.) are classified by MBFC as left-center, the descriptive statement that roughly half of media sources have at least some leaning toward the political left is largely a statement about mainstream outlets. The paper provides no sensitivity analysis under alternative codings, no report of the MBFC category distribution in the matched sample, and no release of data or code to audit the recoding. I request a robustness section that (i) reports the share of each MBFC category among matched journalistic domains, (ii) re-estimates the main models excluding left-center and right-center from the leaning dummies or using a three-level coding, and (iii) explicitly states whether the abstract's absolute 'prioritize left-leaning media' claim survives the stricter coding.
  2. [Results — User-side factors of political information curation] H1 is only partially supported. The monthly Bing (OR = 1.23, p < 0.001) and Google organic (OR = 1.36, p = 0.005) results support H1a, but the monthly Google Newsblock result is null (OR = 1.05, p = 0.420), and the daily Google organic (OR = 0.96, p = 0.264) and daily Google Newsblock (OR = 1.04, p = 0.063) results do not reach statistical significance. The consistent left-side evidence in the daily data is that Republican queries decrease the presence of left-leaning sources, not that Democratic queries increase it. The abstract's statement that both search engines tend to prioritize left-leaning media sources overstates the evidence and should be rephrased to separate the absolute prevalence of left-leaning sources from the relative partisan-query gap and to acknowledge the pronounced engine-, time-, and interface-dependence of the H1a results.
  3. [Measures] Source-type classification relies on 'existing lists of domains ... from earlier auditing studies produced by the authors' plus manual labeling by one author, with no reported inter-rater reliability and no publication of the lists. Because the journalistic-source subset is the population on which all MBFC matching and the main models are based, a systematic error in this step could affect which domains enter the analysis and, in principle, the estimated partisan gap. Please provide the classification lists or a transparency appendix with coding statistics and a reliability check (for example, a second coder on a sample of domains).
minor comments (4)
  1. [Data collection] The list of Republican-focused queries appears truncated in the text: after 'electionsdonaldtrump' the string 'jdvance' is not wrapped in quotation marks and the query is incomplete; please correct this typo so that the query set is unambiguous.
  2. [References] The spelling of Diakopoulos is inconsistent: 'Diakopoulous' appears in the introduction and elsewhere, while the reference list uses 'Diakopoulos'; please unify the spelling throughout.
  3. [Appendix Tables C1–C4] The random-effects grouping variable is labeled 'html_name' in the full regression tables; a more descriptive label such as 'agent_id' would clarify that the random intercepts are at the virtual-agent level.
  4. [Discussion] The limitations acknowledged in the Discussion (small query pool, limited geographic locations with almost no battleground states, and analysis of political leaning only for journalistic media) are relevant and should be echoed at the end of the Results section so that readers weigh the descriptive claims accordingly.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the partisan-curation claim is computed from external MBFC labels and observed search results, not from a fitted input or self-citation chain.

full rationale

The paper's core derivation is an empirical audit rather than a formal derivation. Political leaning of journalistic domains is externally determined by the MediaBiasFactCheck database, and query slant is an experimentally manipulated input; the generalized linear mixed-effects models estimate associations between these independent measurements. No parameter is fitted to a subset of data and then renamed as a prediction; the reported predicted probabilities are standard model outputs. The only self-referential element is the reuse of the authors' prior domain-classification lists for source type, which is a data-processing convenience and does not by itself determine the partisan outcome, because political leaning comes from MBFC, not from those lists. No uniqueness theorem or ansatz is imported from self-citations, and no known result is renamed. The treatment of MBFC 'left-center' outlets as left-leaning is a measurement-validity concern that could affect the absolute claim, but it is not a circularity within the paper's own inference chain.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is an empirical audit, so there are no free parameters in the sense of a derivation. The load-bearing assumptions are about measurement validity (MBFC labels and binary recoding), the representativeness of datacenter IPs, the exposure window (first page only), and the query set. No new entities are postulated.

assumptions (4)
  • domain assumption MBFC political leaning labels are valid, and the binary recoding (left-center or left counts as left; right-center or right counts as right) correctly measures media slant.
    Measures section. All GLMM outcome variables are built from this recoding; if MBFC is systematically biased, the 'left-leaning prevalence' result is an artifact of the label source.
  • domain assumption Search results returned to virtual agents on Google Compute Engine IP addresses are representative of what typical US users would see.
    Data collection section. Datacenter IPs may be treated differently by search engines (e.g., less personalization, different localization), so the measured partisan gap may not generalize to residential users.
  • domain assumption The first page of organic results (7-10 links) is the relevant information exposure window.
    Data collection section. The audit only saves the first page and defines 'prioritize' as appearing on it; deeper results and rank positions are ignored.
  • domain assumption The fixed set of 13-14 queries captures the range of partisan and neutral search behavior for the 2024 US election.
    Data collection section. The paper acknowledges the small query pool as a limitation; different phrasings or issues could change the observed slant.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Search engines in polarized media environment: Auditing political information curation on Google and Bing prior to 2024 US elections." pith.science (2026). https://pith.science/paper/FTTZVXMU

@misc{pith2026250104763,
  author       = {Pith},
  title        = {Pith review of: Search engines in polarized media environment: Auditing political information curation on Google and Bing prior to 2024 US elections},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FTTZVXMU}},
  note         = {Machine review of arXiv:2501.04763}
}
read the original abstract

Search engines play an important role in the context of modern elections. By curating information in response to user queries, search engines influence how individuals are informed about election-related developments and perceive the media environment in which elections take place. It has particular implications for (perceived) polarization, especially if search engines' curation results in a skewed treatment of information sources based on their political leaning. Until now, however, it is unclear whether such a partisan gap emerges through information curation on search engines and what user- and system-side factors affect it. To address this shortcoming, we audit the two largest Western search engines, Google and Bing, prior to the 2024 US presidential elections and examine how these search engines' organic search results and additional interface elements represent election-related information depending on the queries' slant, user location, and time when the search was conducted. Our findings indicate that both search engines tend to prioritize left-leaning media sources, with the exact scope of search results' ideological slant varying between Democrat- and Republican-focused queries. We also observe limited effects of location- and time-based factors on organic search results, whereas results for additional interface elements were more volatile over time and specific US states. Together, our observations highlight that search engines' information curation actively mirrors the partisan divides present in the US media environments and has the potential to contribute to (perceived) polarization within these environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023

    cs.CY 2025-07 conditional novelty 5.0 of 10

    In the 2023 Swiss federal election, Google showed men candidates more prominent news coverage than women and depicted women with more smiling and positive images, and these search patterns were associated with the can...

Reference graph

Works this paper leans on

21 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Annalsof theAmericanAssociationof Geographers,107(5),1194-1215

    Ballatore,A.,Graham,M.,&Sen,S.(2017).Digitalhegemonies:thelocalnessofsearchengineresults. Annalsof theAmericanAssociationof Geographers,107(5),1194-1215. Baly,R.etal.(2018).Predictingfactualityofreportingandbiasofnewsmediasources.InProceedings of the 2018 Conference onEmpirical MethodsinNatural LanguageProcessing(pp.3528-3539).ACM. Bandy,J.(2021).Problema...

  2. [2]

    Bandy, J., & Diakopoulos, N. (2020). Auditing news curation systems: A case studyexaminingalgorithmicandeditoriallogicinAppleNews.In Proceedings of the InternationalAAAI ConferenceonWebandSocial Media(pp.36-47).AAAI. Bartley,N.etal.(2021).AuditingalgorithmicbiasonTwitter.InProceedingsof the13thACMWebScienceConference2021(pp.65-73).ACM Blassnig,S.etal.(202...

  3. [3]

    et al.(2016).Shouldweworryaboutfilterbubbles? Internet Policy Review,5(1),1-16

    Borgesius, F. et al.(2016).Shouldweworryaboutfilterbubbles? Internet Policy Review,5(1),1-16. 22

  4. [4]

    Bragazzi, N. et al. (2017). Howoften people google for vaccination: Qualitative andquantitativeinsightsfromasystematicsearchof theweb-basedactivitiesusingGoogleTrends. HumanVaccines&Immunotherapeutics, 13(2),464-469. Bruns,A.(2019). AreFilter BubblesReal?JohnWiley&Sons. Castillo,C.(2019).Fairnessandtransparencyinranking. ACMSIGIRForum,52(2),64-71. Daucé,F...

  5. [5]

    Dobber, T. (2023). Microtargeting, privacy, andtheneedforregulatingalgorithms.InTheRoutledgeHandbookofPrivacyandSocialMedia(pp.237-245).Routledge. Dutton,W.etal.(2017). Searchandpolitics: Theusesandimpactsof searchinBritain,France, Germany, Italy, Poland, Spain, andtheUnitedStates.SSRN.https://dx.doi.org/10.2139/ssrn.2960697 Epstein,R.,&Robertson,R.(2015)...

  6. [6]

    Hoskins, A., &Tulloch, J. (2016). Risk and hyperconnectivity: Media and memories ofneoliberalism.OxfordUniversityPress

  7. [7]

    Jasper, S. (2023). How we’re approaching the 2024 U.S. elections. Google.https://blog.google/outreach-initiatives/civics/how-were-approaching-the-2024-us-elections/

  8. [8]

    Kang, J. (2024). How Biased Is the Media, Really? The New Yorker.https://www.newyorker.com/news/fault-lines/how-biased-is-the-media-really Kliman-Silver,C.etal.(2015).Location,location,location:Theimpactofgeolocationonwebsearchpersonalization.In Proceedings of the 2015 Internet Measurement Conference(pp.121-127).ACM. 23 Kubin,E.,&vonSikorski,C.(2023).Thec...

Show all 21 references
  1. [9]

    Kuznetsova, E., & Makhortykh, M. (2023). Blame it on the algorithm? Russiangovernment-sponsoredmediaandalgorithmiccurationofpoliticalinformationonFacebook.International Journal of Communication,17,971-992

  2. [10]

    Kuznetsova, E. et al. (2024). Algorithmically curated lies: How search engines handlemisinformationabout USbiolabsinUkraine.arXiv.https://doi.org/10.48550/arXiv.2401.13832

  3. [11]

    Lowell, H. (2024). Trump vows to seek criminal charges against Googleif re-electedpresident. The Guardian.https://www.theguardian.com/us-news/2024/sep/27/trump-google-threat-criminal-charges

  4. [12]

    Ludwig, K. et al. (2023). Dividedbythealgorithm?the(limited) effectsof content-andsentiment-basednewsrecommendationonaffective,ideological,andperceivedpolarization.SocialScienceComputerReview,41(6),2188-2210. Makhortykh,M.,Urman,A.,&Ulloa,R.(2020).Howsearchenginesdisseminatein...

  5. [13]

    (2022).Memory,counter-memoryanddenialism:How search engines circulate information about the Holodomor-related memory wars.MemoryStudies, 15(6),1330-1345

    Makhortykh, M., Urman, A., &Ulloa, R. (2022).Memory,counter-memoryanddenialism:How search engines circulate information about the Holodomor-related memory wars.MemoryStudies, 15(6),1330-1345

  6. [14]

    Makhortykh, M. et al. (2024). Popular Votes and Algorithms in Switzerland: IntransparentPriorisationof Political Information.Reatch

  7. [15]

    Makhortykh, M. et al. (2024). Doesit get betterwithtime?Websearchconsistencyandrelevanceinthevisual representationof theHolocaust. InE. Pfanzelter, D. Rupnow, É.KovácsandM. Windsperger(Eds.) Connected Histories: Memories and Narratives of theHolocaust inDigital Space(pp.13-33)...

  8. [16]

    Mittelstadt, B. (2016). Auditing for transparency in content personalization systems.International Journal of Communication,10,4991–5002. Nechushtai,E.,Zamith,R.,&Lewis,S.C.(2024).Moreofthesame?HomogenizationinnewsrecommendationswhenuserssearchonGoogle,YouTube,Facebook,andTwit...

  9. [17]

    Perreault, B. et al. (2024). Algorithmicmisjudgement inGooglesearchresults:EvidencefromauditingtheUSonlineelectoralinformationenvironment.In Proceedings of the 2024ACMConferenceonFairness, Accountability, andTransparency(pp.433-443).ACM. Pradel,F.(2021).Biasedrepresentationofp...

  10. [18]

    Puschmann, C. (2019). Beyond the bubble: Assessingthediversityof political searchresults. Digital Journalism, 7(6),824-843. Rader,E.,&Gray,R.(2015).UnderstandinguserbeliefsaboutalgorithmiccurationintheFacebooknewsfeed. In Proceedings of the 33rd Annual ACM Conference on HumanF...

  11. [19]

    Thorson, K. (2020). Attractingthenews: Algorithms, platforms, andreframingincidentalexposure. Journalism, 21(8),1067-1082

  12. [20]

    Foreignbeautieswanttomeetyou

    Trielli, D., &Diakopoulos, N. (2022). PartisansearchbehaviorandGoogleresultsinthe2018USmidtermelections. Information, Communication&Society,25(1),145-161. Ulloa,R.etal.(2024a).Noveltyinnewssearch:Alongitudinalstudyofthe2020USelections. Social ScienceComputer Review,42(3),700-7...

  13. [21]

    Wilkinson, M. (2023). Bingvs. Google: ComparingtheTwoSearchEngines. Semrush.https://www.semrush.com/blog/bing-vs-google/ Wilson,A.,Parker,V.,&Feinberg,M.(2020).Polarizationinthecontemporarypoliticalandmedialandscape. Current OpinioninBehavioral Sciences, 34,223-228. Yang,C.eta...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.