Pith. sign in

REVIEW 4 major objections 5 minor 3 references

Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that Google's text and image search results during the 2023 Swiss federal election campaign were gendered, with men receiving more news media links and women receiving more stereotypically pleasant images, and that these…

desk verdict The descriptive gender-bias audit is solid and worth publishing; the electoral-performance 'prediction' is an in-sample correlation and should be toned down. read the letter →

arxiv 2507.06018 v1 pith:MWPAQ3LX submitted 2025-07-08 cs.CY

classification cs.CY
keywords algorithmauditGooglesearchgenderbiasmediasourceprominenceimageelectoralperformanceSwissFederalElections
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that Google's text and image search output about candidates in the 2023 Swiss federal election campaign was gendered and that these search patterns carried electoral signal. Text searches were dominated by government and news media sources, but queries for men returned more and higher-ranked news media links than queries for women. Image searches showed women candidates slightly underrepresented relative to their share on the ballot lists and portrayed them with more smiling and positive affect and less negative affect, especially for women from conservative parties. The paper claims that adding text and image search measures to models of candidates' personal votes explains an additional 6 to 8 percent of the variance, making search-engine representation a measurable part of electoral performance. The expected electoral penalty for women displaying negative emotions was not found, so the gendered portrayal is not a simple one-way disadvantage.

What carries the argument

Two instruments carry the argument. The first is media source prominence, a rank-weighted share of news media domains defined as $\sum_{i\in\text{media}} 1/r_i$ over $\sum_{j=1}^N 1/r_j$, where $r_i$ is the rank of the $i$-th organic result; it converts Google's source mix and ordering into a single candidate-level visibility score. The second is a set of visual-stereotyping measures produced by a commercial computer-vision API from over a million returned images: share of women depicted, average smile probability, average positive affect (happiness plus calmness), and average negative affect (fear, anger, sadness, disgust). These measures are aggregated per candidate per wave and entered into mixed-effects regressions with by-canton random intercepts, so the audit infrastructure, which used virtual machines with Swiss IPs and cookie-cleared browser agents querying google.ch in two waves, makes the comparison controlled and the link to official election data possible.

What would settle it

Re-run the hierarchical regressions with the candidates' personal vote share from the 2019 Swiss federal election, or an equivalent prior-popularity measure, included as a control; if the 6 to 8 percent variance attributed to text and image search measures collapses toward zero, the predictive claim is an artifact of successful candidates generating more search results.

Watch

Extended reading notes

Core claim

The paper's central claim is that Google's selection and ranking of candidate information during the 2023 Swiss Federal Elections was systematically different for men and women, and that the resulting search-representation measures were associated with actual voting outcomes. In text results, media source prominence was higher for men candidates (b = -3.94 in wave 1; b = -2.72 in wave 2), meaning men received more news media links and better placement. In image results, women made up about 38.7% to 39.1% of depicted persons versus 42.5% of candidates on official lists, and queries for women returned images with more smiling (about 5.4 percentage points more), more positive affect (b = 1.82 and 2.22), and less negative affect (b = -1.05 and -1.08). Hierarchical regressions with by-canton random intercepts found that text and image search blocks together added 6 to 8 percent explained variance to models of log-transformed personal votes, with media source prominence positively predicting votes and, in wave 1, a significant interaction showing women benefited more from media prominence than men. The paper does not claim that negative affect is punished for women; that hypothesis was not supported.

Load-bearing premise

The load-bearing premise is that the observed correlation between Google search representation and personal votes is not driven by reverse causality, meaning that popular candidates are not simply generating more and more positive search results; the models do not control for prior vote share, offline media coverage, or campaign spending.

Editorial extensions

If this is right

  • Women candidates start from a lower base of rank-weighted news media visibility, and because media source prominence is positively associated with personal votes, the gender gap in Google text results implies a corresponding electoral disadvantage.
  • The consistent gender-by-party pattern in image output means that women from conservative parties are most exposed to stereotypically positive and smiling portrayals, making the interaction between gender and party part of the algorithmic-curation story.
  • Search-based measures explain 6 to 8 percent of variance in personal votes with the included controls, a magnitude the paper argues is politically consequential in elections decided by a few percentage points.
  • The null result for the negative-affect penalty indicates that not every gendered stereotype translates into an electoral punishment, so the double bind for women candidates is more conditional than the visual stereotype literature might suggest.
  • The two-wave design shows the main gendered patterns in text and image output were stable across the final month of the campaign, which points to a persistent, not transient, algorithmic environment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The most serious rival explanation, not settled by the paper, is reverse causality: candidates who are expected to win attract more media coverage and more image material, so the 6 to 8 percent predictive variance may partly be popularity rather than search-engine influence. A natural extension is to add prior vote share or campaign spending to the models and see whether the search coefficients su
  • Because the image results may mirror candidates' own campaign photography and party visual strategy, an editorial inference is that comparing Google image output with the candidates' official portraits or social-media feeds would separate algorithmic stereotyping from self-presentation.
  • A randomized or natural-experiment perturbation of search rankings, for example a temporary change in a candidate's media prominence that is unrelated to their popularity, could convert the predictive association into a causal test; the paper's design is observational and does not itself identify causal direction.
  • If the reverse-causality concern is borne out, regulators would need to focus on whether Google amplifies existing inequality rather than creates it; the paper leaves that distinction explicitly open in its discussion.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a large-scale algorithm audit of Google text and image search results for all 5,883 candidates in the 2023 Swiss federal elections, with data collected three weeks and one week before election day. It finds that text search results give men candidates higher media source prominence than women candidates, and that image search results for women show more smiling, more positive affect, and less negative affect, especially for right-leaning candidates. The paper further claims that these search representation measures are predictive of electoral performance, with search-derived variables adding 6–8% of explained variance to models of candidates' log-transformed personal votes.

Significance. If the descriptive findings hold, the study provides valuable large-scale evidence on algorithmic gender bias in political search output, extending prior work on media visibility and visual stereotyping to a full-candidate audit in a proportional electoral context. The audit design is a strength: it covers all candidates, uses two temporally separated waves, employs virtual agents to reduce personalization, and combines manual source coding with computer-vision affect measurement. The gender differences in media source prominence (b = -3.94 and -2.72) and in positive/negative affect (b = 1.82/2.22 and -1.05/-1.08) are consistent across waves and reported with confidence intervals. The weakest part is the electoral-performance analysis: the cross-sectional regressions cannot distinguish search-engine influence from reverse causality or omitted confounders, so the 'predictive' claim in the abstract and H2 is not currently supported.

major comments (4)
  1. [Search engines and electoral results; Discussion] The claim that search-output measures are 'predictive of electoral performance' (Abstract; Table 2) is not identified as a causal or predictive effect. The hierarchical regressions control gender, party, incumbency, list position, and canton random intercepts, but not prior vote share, offline media coverage volume, or campaign spending. Since the paper itself states that 'the presence in traditional media is a necessary condition for media source prominence' (Section 'Text searches: Algorithmic curation and media source prominence'), the news-media-prominence coefficients (Table 2; wave 1 b = 0.27) plausibly reflect pre-existing candidate popularity and news coverage rather than search-engine influence on voters. The Discussion acknowledges that distinguishing amplification from distortion is 'essential but challenging', but the Abstract's 'predictive' wording and the H2 formulation go beyond what the cross-sectional design can establish. Please either add controls or robustness checks (e.g., prior election results, offline coverage volume) or reframe the claim as an association.
  2. [Table 2 vs. Results text] There is a direct inconsistency between the text and Table 2 for the wave-2 media source prominence effect. The text reports b = 1.34, 95% CI [1.001, 1.58] for media prominence in wave 2, but Table 2 lists 'News media' as b = 0.37 (SE = 0.11) and 'Social media' as b = 1.34 (SE = 0.18). Please correct the reported coefficient/CI or the table, and ensure the interpretation (news media vs. social media) is consistent.
  3. [Methodology/Data analysis; Table 2] The sample size is inconsistent. The methods state that each wave contains n = 5,883 candidates (Table 1, k = 5,883), but Table 2 reports 5,952 and 5,948 observations for the two waves. Please clarify whether some candidates are represented by multiple agent-level records after aggregation, or whether the N in Table 2 should be 5,883; this affects degrees of freedom and the reported fit statistics.
  4. [Search engines and electoral results] The 'predictive' language is not supported by out-of-sample evaluation. The likelihood-ratio tests and Δ marginal R² values are computed on the same data used to fit the models; they quantify in-sample explanatory power, not predictive performance. If the authors intend 'predictive' in the forecasting sense, cross-validation or a temporal holdout is needed; otherwise, terms like 'associated with' or 'explained variance' should be used consistently.
minor comments (5)
  1. [Introduction] The introduction states that 'text search output included more and higher rank links to media sources for queries of women candidates', which contradicts the Results, Abstract, and Discussion, where women have lower media source prominence; please correct this sentence.
  2. [Image search results] The positive-affect finding is attributed to H4a, but H4a is about smiling; H4b is about more positive affect and H4c about less negative affect. Please fix the hypothesis labels in this section.
  3. [Methodology/Data analysis strategy] The models predicting log-transformed personal votes are described as 'generalised linear mixed-effects regression models'; unless a non-Gaussian family and link function are specified, these should be described as linear mixed-effects models.
  4. [Online Appendix, Table S4] Table S4 contains a duplicated column header ('affect positive'); please clean up the table formatting.
  5. [Data availability] No data or code availability statement is included; please add one or state that materials are available upon request, as this is standard for algorithm audits.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the gender/party differences are measured audit outputs, and the 'predictive' performance claim is an in-sample regression fit, not a derivation that reduces to its own inputs.

full rationale

This paper is an empirical algorithm audit rather than a first-principles derivation. The gender and party differences in Google text and image results are directly measured from the collected search outputs (e.g., media-source prominence differences of b = -3.94 and b = -2.72; image affect differences), so they are observations, not fitted targets. The central 'predictive of electoral performance' claim (Abstract; 'Search engines and electoral results', Table 2) is based on hierarchical mixed-effects regressions fitted to the same Swiss 2023 election outcomes; the reported 6-8% explained variance is an in-sample model-fit increment, not an out-of-sample forecast. This is a presentation limitation, but it is not circular by construction: the coefficients and likelihood-ratio tests could have come out null, and no parameter is defined in terms of the outcome it is said to predict. The Discussion explicitly acknowledges the harder identification issue: it is 'unclear whether these algorithms merely amplify existing social biases or create distortions in social reality' and that 'Establishing a robust baseline for search engine curation each politician... is essential but challenging.' The cited prior work by the same authors (Rohrbach et al., 2024; Makhortykh et al., 2025a/b) supplies background and earlier audit evidence, but it is not invoked as a load-bearing uniqueness theorem or an ansatz that forces the present results. A separate textual inconsistency exists (the Introduction says women received more media links while the Results say men did), but that is an editing error, not a circular step. Overall, the audit's measured outputs are self-contained, and the predictive language, while not a genuine forecast, does not make the derivation circular.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper's empirical estimates rest on several unverified measurement and sampling assumptions: the virtual agents mimic real voters, the emotion classifier is accurate for political images, the manual source coding is reliable, and the regression controls suffice to rule out reverse causality. None of these are independently evidenced in the manuscript, and no code or data is released to test them.

free parameters (3)
  • Inverse rank weighting in media source prominence = 1/rank
    The media source prominence metric arbitrarily weights results by inverse rank; other weighting schemes (log-rank, top-k) would change the measure and potentially the results.
  • Affect valence groupings = positive = happiness + calmness; negative = fear + anger + sadness + disgust
    Analytic choice: calmness is grouped with happiness and surprise/confusion are excluded without empirical justification.
  • Most prominent face smile score = probability score for the single most prominent face
    The aggregation from images to candidate-level averages is a constructed measure; no sensitivity analysis is reported for alternative face-selection rules.
assumptions (5)
  • domain assumption Google search results are a relevant source of political information for voters in Switzerland.
    The motivating literature is cited, but the audit does not measure actual exposure or effects; it only records search outputs. Entered in the Introduction.
  • domain assumption Virtual agents with cleared cookies and static IPs produce search results representative of ordinary Swiss voters.
    Personalization is minimized but not eliminated; the agents' results may differ from those of real users with different search histories. Stated in Methodology, Data collection.
  • domain assumption Amazon Rekognition's emotion labels accurately measure displayed emotions in political images.
    No validation or error analysis is provided; the model is treated as ground truth. Used in 'To enrich image search results' in Data processing.
  • domain assumption The single student assistant's manual coding of unknown domains is reliable.
    Only one coder after training; no inter-coder reliability statistic is reported. See Data processing and Table S1.
  • domain assumption The regression controls (incumbency, list position, party, canton, gender) are sufficient to prevent reverse causality from explaining the search-vote association.
    This is a strong assumption given that candidate quality and prior media attention are not directly measured. It underpins the 'predictive of electoral performance' claim in Results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023." pith.science (2026). https://pith.science/paper/MWPAQ3LX

@misc{pith2026250706018,
  author       = {Pith},
  title        = {Pith review of: Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MWPAQ3LX}},
  note         = {Machine review of arXiv:2507.06018}
}
read the original abstract

Search engines like Google have become major sources of information for voters during election campaigns. To assess potential biases across candidates' gender and partisan identities in the algorithmic curation of candidate information, we conducted a large-scale algorithm audit analyzing Google's selection and ranking of information about candidates for the 2023 Swiss Federal Elections, three and one week before the election day. Results indicate that text searches prioritize media sources in search output but less so for women politicians. Image searches revealed a tendency to reinforce stereotypes about women candidates, marked by a disproportionate focus on stereotypically pleasant emotions for women, particularly among right-leaning candidates. Crucially, we find that patterns of candidates' representation in Google text and image searches are predictive of their electoral performance.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 1 canonical work pages

  1. [1]

    Aaldering, L., Van der Meer, T., & Van der Brug, W. (2018). Mediated Leader Effects: The Impact of Newspapers’ Portrayal of Party Leadership on Electoral Support. The International Journal of Press/Politics , 23 (1), 70–94. https://doi.org/10.1177/1940161217740696 Bandy, J., & Diakopoulos, N. (2020). Auditing News Curation Systems: A Case Study Examining ...

  2. [12]

    Foreign beauties want to meet you

    Newman, N., Fletcher, R., Eddy, K., Robertson, C. T., & Nielsen, R. K. (2023). Reuters Institute digital news report 2023 . Reuters Institute for the study of Journalism. https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2023-06/Digital_News_Repo rt_2023.pdf Norocel, O. C., & Lewandowski, D. (2023). Google, data voids, and the dynamics of the...

  3. [33]

    https://doi.org/10.1007/s42001-025-00361-3 Makhortykh, M., Rohrbach, T., Sydorova, M., & Kuznetsova, E. (2025b). Search engines in polarized media environment: Auditing political information curation on Google and Bing prior to 2024 US elections. arXiv preprint arXiv:2501.04763. Mittelstadt, B. (2016). Automation, algorithms, and politics| auditing for tr...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.