{"id":"78e1daee-c2d2-43f2-8df4-ee621c0c4721","arxiv_id":"2507.06018","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In the 2023 Swiss federal election, Google showed men candidates more prominent news coverage than women and depicted women with more smiling and positive images, and these search patterns were associated with the candidates' final vote shares.","lead":"The authors audited Google text and image searches for all 5,883 candidates in Switzerland's 2023 federal election and linked what Google returned to actual vote results. They found women candidates received fewer and lower-ranked news links and more stereotypically positive images, with search representation correlated with electoral support.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reverse causality: search-output measures are downstream of offline media coverage and candidate strength, so the 6–8% explained variance in votes cannot support the 'predictive'/'impact' claim without controlling prior vote share or coverage volume.","rationale":"The reader's weakest assumption is exactly the one I would stress. The descriptive findings are credible: consistent gender differences across two waves in media source prominence and in emotion labels, with plausible effect sizes. However, the strongest claim—that search measures predict electoral performance and by implication shape outcomes—requires that the association survives adjustment for what a candidate brings to the search index. The available controls (incumbency, list position, party) are coarse proxies for candidate quality; incumbency alone does not capture the wide variation in name recognition among challengers. Since Google indexes offline media content, better-known candidates mechanically accumulate more media links; the regression then re-discovers this. The proposed test is feasible because Switzerland publishes prior election results and media coverage is trackable. I would not reject the paper; the descriptive audit has value and the two waves provide consistency. But the predictive claim should be downgraded or reworded as 'associated with' pending the robustness check. The reader's CONDITIONAL verdict therefore stands unchanged; the manuscript should add the omitted-variable control or soften the causal/predictive language.","tokens_in":29363,"tokens_out":5415,"duration_ms":61036,"concrete_test":"Re-run the Table 2 hierarchical regressions with an added control for candidates' 2019 personal votes (from the Swiss Federal Statistical Office; for first-time candidates use 0 plus an indicator) and, if available, a campaign-period offline media coverage count per candidate from Swissdox. Then compare the news-media coefficient and the block likelihood-ratio ΔR² for the text/image search measures. If the search block no longer adds significant variance or the coefficients shrink toward zero, the 'predictive' claim is explained by omitted candidate strength and media attention rather than by Google output.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that search representation is 'predictive of electoral performance' (Abstract; Section 'Search engines and electoral results', Table 2) treats Google output as an influence on votes. The regression controls gender, incumbency, party, list position, and canton random intercepts, but not prior vote share, offline media coverage volume, or campaign spending. The paper itself states that 'the presence in traditional media is a necessary condition for media source prominence' (Section 'Text searches: Algorithmic curation and media source prominence'). Because Google text results are built from the same media pages that also predict electoral success, the news-media-prominence coefficient (wave 1 b = 0.27; wave 2 b = 1.34) plausibly reflects pre-existing popularity and news coverage rather than search-engine influence. The image measures (positive/negative affect, smiles) similarly originate from campaign materials and media photos, which correlate with resources and expected success. The Discussion acknowledges the difficulty of separating 'amplify existing social biases' from 'create distortions' and calls a robust baseline 'essential but challenging'. Thus the cross-sectional regression cannot identify the direction of the association; the causal/predictive interpretation is the least secure part of the argument. A separate text inconsistency (Introduction says women receive more media links while Results say men do) should be corrected, but the identification problem is the load-bearing issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale algorithm audit of Google text and image search results for all 5,883 candidates in the 2023 Swiss federal elections, with data collected three weeks and one week before election day. It finds that text search results give men candidates higher media source prominence than women candidates, and that image search results for women show more smiling, more positive affect, and less negative affect, especially for right-leaning candidates. The paper further claims that these search representation measures are predictive of electoral performance, with search-derived variables adding 6–8% of explained variance to models of candidates' log-transformed personal votes.","tokens_in":29623,"tokens_out":5741,"duration_ms":61668,"significance":"If the descriptive findings hold, the study provides valuable large-scale evidence on algorithmic gender bias in political search output, extending prior work on media visibility and visual stereotyping to a full-candidate audit in a proportional electoral context. The audit design is a strength: it covers all candidates, uses two temporally separated waves, employs virtual agents to reduce personalization, and combines manual source coding with computer-vision affect measurement. The gender differences in media source prominence (b = -3.94 and -2.72) and in positive/negative affect (b = 1.82/2.22 and -1.05/-1.08) are consistent across waves and reported with confidence intervals. The weakest part is the electoral-performance analysis: the cross-sectional regressions cannot distinguish search-engine influence from reverse causality or omitted confounders, so the 'predictive' claim in the abstract and H2 is not currently supported.","major_comments":[{"comment":"The claim that search-output measures are 'predictive of electoral performance' (Abstract; Table 2) is not identified as a causal or predictive effect. The hierarchical regressions control gender, party, incumbency, list position, and canton random intercepts, but not prior vote share, offline media coverage volume, or campaign spending. Since the paper itself states that 'the presence in traditional media is a necessary condition for media source prominence' (Section 'Text searches: Algorithmic curation and media source prominence'), the news-media-prominence coefficients (Table 2; wave 1 b = 0.27) plausibly reflect pre-existing candidate popularity and news coverage rather than search-engine influence on voters. The Discussion acknowledges that distinguishing amplification from distortion is 'essential but challenging', but the Abstract's 'predictive' wording and the H2 formulation go beyond what the cross-sectional design can establish. Please either add controls or robustness checks (e.g., prior election results, offline coverage volume) or reframe the claim as an association.","section":"Search engines and electoral results; Discussion"},{"comment":"There is a direct inconsistency between the text and Table 2 for the wave-2 media source prominence effect. The text reports b = 1.34, 95% CI [1.001, 1.58] for media prominence in wave 2, but Table 2 lists 'News media' as b = 0.37 (SE = 0.11) and 'Social media' as b = 1.34 (SE = 0.18). Please correct the reported coefficient/CI or the table, and ensure the interpretation (news media vs. social media) is consistent.","section":"Table 2 vs. Results text"},{"comment":"The sample size is inconsistent. The methods state that each wave contains n = 5,883 candidates (Table 1, k = 5,883), but Table 2 reports 5,952 and 5,948 observations for the two waves. Please clarify whether some candidates are represented by multiple agent-level records after aggregation, or whether the N in Table 2 should be 5,883; this affects degrees of freedom and the reported fit statistics.","section":"Methodology/Data analysis; Table 2"},{"comment":"The 'predictive' language is not supported by out-of-sample evaluation. The likelihood-ratio tests and Δ marginal R² values are computed on the same data used to fit the models; they quantify in-sample explanatory power, not predictive performance. If the authors intend 'predictive' in the forecasting sense, cross-validation or a temporal holdout is needed; otherwise, terms like 'associated with' or 'explained variance' should be used consistently.","section":"Search engines and electoral results"}],"minor_comments":[{"comment":"The introduction states that 'text search output included more and higher rank links to media sources for queries of women candidates', which contradicts the Results, Abstract, and Discussion, where women have lower media source prominence; please correct this sentence.","section":"Introduction"},{"comment":"The positive-affect finding is attributed to H4a, but H4a is about smiling; H4b is about more positive affect and H4c about less negative affect. Please fix the hypothesis labels in this section.","section":"Image search results"},{"comment":"The models predicting log-transformed personal votes are described as 'generalised linear mixed-effects regression models'; unless a non-Gaussian family and link function are specified, these should be described as linear mixed-effects models.","section":"Methodology/Data analysis strategy"},{"comment":"Table S4 contains a duplicated column header ('affect positive'); please clean up the table formatting.","section":"Online Appendix, Table S4"},{"comment":"No data or code availability statement is included; please add one or state that materials are available upon request, as this is standard for algorithm audits.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The audit component is strong and publication-worthy if the descriptive gender-bias findings are the focus. The main risk is overinterpretation of the cross-sectional election model as showing search-engine influence on votes. The coefficient inconsistency between text and Table 2, and the sample-size mismatch, must be fixed before acceptance. If the authors add a robustness check with prior vote share or offline coverage, or clearly reframe the claim as associational, the paper could be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core descriptive findings are real and worth taking seriously. This is a full-population audit of all 5,883 Swiss federal candidates, run twice during the campaign, covering both text and image search output. The gender differences in media source prominence (men getting more and higher-ranked news links) and in image affect (women getting more smiles and positive affect, less negative affect) are consistent across waves and estimated with tight confidence intervals. Those results alone are a useful contribution to the algorithmic-fairness and political-communication literatures. The scale and the candidate-level linkage to electoral data are genuinely new for Switzerland.\n\nThe soft spot is exactly where the stress-test says it is. The claim that search representation is \"predictive of electoral performance\" rests on regressions of personal votes on search measures measured at the same time, with controls for gender, incumbency, party, list position, and canton, but not for prior vote share, offline media coverage volume, or campaign spending. Google text results are built largely from the same media pages that influence and reflect candidate strength, so the 6–8% explained variance could easily be capturing popularity rather than search-engine influence. The paper itself concedes that disentangling algorithmic amplification from social bias is \"essential but challenging\" and calls for a robust baseline. Given that, the abstract's \"predictive\" language is overstated. This is not fatal to the descriptive contribution, but it needs a rewrite.\n\nThere are also a few smaller issues. The Introduction says text search output included more media links for women candidates, while the Results say men received more. That contradiction must be fixed. No inter-coder reliability is reported for the manual source coding, and the Amazon Rekognition emotion labels are not validated against any ground truth. No data or code are released, which limits reproducibility for an audit study. These are all fixable or at least disclosable.\n\nThe paper is honest about its limitations in the Discussion, and the main descriptive findings hold up. I think this deserves a serious referee slot, not a desk rejection. The authors should be pushed to reframe the electoral-performance analysis as correlational, add whatever robustness checks are possible (even a lagged or placebo analysis), and release the aggregated data. With those changes it would be a solid publication.","headline":"The descriptive gender-bias audit is solid and worth publishing; the electoral-performance 'prediction' is an in-sample correlation and should be toned down.","tokens_in":136,"tokens_out":1248,"would_cite":true,"duration_ms":37633,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Google's text and image search results during the 2023 Swiss federal election campaign were gendered, with men receiving more news media links and women receiving more stereotypically pleasant images, and that these…","keywords":["algorithm audit","Google search","gender bias","media source prominence","image search","electoral performance","Swiss Federal Elections"],"falsifier":"Re-run the hierarchical regressions with the candidates' personal vote share from the 2019 Swiss federal election, or an equivalent prior-popularity measure, included as a control; if the 6 to 8 percent variance attributed to text and image search measures collapses toward zero, the predictive claim is an artifact of successful candidates generating more search results.","tokens_in":29165,"feed_emoji":"🔍","tokens_out":10033,"duration_ms":107345,"temperature":0.7,"pith_summary":"The paper tries to establish that Google's text and image search output about candidates in the 2023 Swiss federal election campaign was gendered and that these search patterns carried electoral signal. Text searches were dominated by government and news media sources, but queries for men returned more and higher-ranked news media links than queries for women. Image searches showed women candidates slightly underrepresented relative to their share on the ballot lists and portrayed them with more smiling and positive affect and less negative affect, especially for women from conservative parties. The paper claims that adding text and image search measures to models of candidates' personal votes explains an additional 6 to 8 percent of the variance, making search-engine representation a measurable part of electoral performance. The expected electoral penalty for women displaying negative emotions was not found, so the gendered portrayal is not a simple one-way disadvantage.","feed_headline":"Google search output favored men and predicted Swiss votes","feed_subtitle":"An audit of all 5,883 candidates links Google's gendered text and image results to actual election outcomes.","key_machinery":"Two instruments carry the argument. The first is media source prominence, a rank-weighted share of news media domains defined as $\\sum_{i\\in\\text{media}} 1/r_i$ over $\\sum_{j=1}^N 1/r_j$, where $r_i$ is the rank of the $i$-th organic result; it converts Google's source mix and ordering into a single candidate-level visibility score. The second is a set of visual-stereotyping measures produced by a commercial computer-vision API from over a million returned images: share of women depicted, average smile probability, average positive affect (happiness plus calmness), and average negative affect (fear, anger, sadness, disgust). These measures are aggregated per candidate per wave and entered into mixed-effects regressions with by-canton random intercepts, so the audit infrastructure, which used virtual machines with Swiss IPs and cookie-cleared browser agents querying google.ch in two waves, makes the comparison controlled and the link to official election data possible.","core_discovery":"The paper's central claim is that Google's selection and ranking of candidate information during the 2023 Swiss Federal Elections was systematically different for men and women, and that the resulting search-representation measures were associated with actual voting outcomes. In text results, media source prominence was higher for men candidates (b = -3.94 in wave 1; b = -2.72 in wave 2), meaning men received more news media links and better placement. In image results, women made up about 38.7% to 39.1% of depicted persons versus 42.5% of candidates on official lists, and queries for women returned images with more smiling (about 5.4 percentage points more), more positive affect (b = 1.82 and 2.22), and less negative affect (b = -1.05 and -1.08). Hierarchical regressions with by-canton random intercepts found that text and image search blocks together added 6 to 8 percent explained variance to models of log-transformed personal votes, with media source prominence positively predicting votes and, in wave 1, a significant interaction showing women benefited more from media prominence than men. The paper does not claim that negative affect is punished for women; that hypothesis was not supported.","pith_inferences":["The most serious rival explanation, not settled by the paper, is reverse causality: candidates who are expected to win attract more media coverage and more image material, so the 6 to 8 percent predictive variance may partly be popularity rather than search-engine influence. A natural extension is to add prior vote share or campaign spending to the models and see whether the search coefficients su","Because the image results may mirror candidates' own campaign photography and party visual strategy, an editorial inference is that comparing Google image output with the candidates' official portraits or social-media feeds would separate algorithmic stereotyping from self-presentation.","A randomized or natural-experiment perturbation of search rankings, for example a temporary change in a candidate's media prominence that is unrelated to their popularity, could convert the predictive association into a causal test; the paper's design is observational and does not itself identify causal direction.","If the reverse-causality concern is borne out, regulators would need to focus on whether Google amplifies existing inequality rather than creates it; the paper leaves that distinction explicitly open in its discussion."],"forward_implications":["Women candidates start from a lower base of rank-weighted news media visibility, and because media source prominence is positively associated with personal votes, the gender gap in Google text results implies a corresponding electoral disadvantage.","The consistent gender-by-party pattern in image output means that women from conservative parties are most exposed to stereotypically positive and smiling portrayals, making the interaction between gender and party part of the algorithmic-curation story.","Search-based measures explain 6 to 8 percent of variance in personal votes with the included controls, a magnitude the paper argues is politically consequential in elections decided by a few percentage points.","The null result for the negative-affect penalty indicates that not every gendered stereotype translates into an electoral punishment, so the double bind for women candidates is more conditional than the visual stereotype literature might suggest.","The two-wave design shows the main gendered patterns in text and image output were stable across the final month of the campaign, which points to a persistent, not transient, algorithmic environment."],"supporting_citations":[{"why":"Supplies the meta-analytic baseline that women are underrepresented in political media coverage, which the text-search hypotheses extend to search engines.","marker":"Van der Pas & Aaldering, 2022"},{"why":"Provides eye-tracking evidence that users click top-ranked links, justifying the inverse-rank weighting used to measure media source prominence.","marker":"Granka et al. (2004)"},{"why":"Demonstrates that search engines favor high-authority media domains, framing the expected dominance of news sources in Google text results.","marker":"Haim et al. (2018)"},{"why":"Shows that image search outputs gender-stereotyped representations of competent men and warm women, motivating the visual stereotyping measures.","marker":"Otterbacher et al. (2017)"},{"why":"Its analysis of Bing images of 2019 European Election candidates supplies the operationalization and comparative finding that women's images show more pleasant emotions and men's show more anger.","marker":"Jungblut & Haim (2023)"},{"why":"Provides large-scale evidence that online images amplify gender bias, the external benchmark against which the modest Swiss underrepresentation is interpreted.","marker":"Guilbeault et al. (2024)"},{"why":"Prior audit of political Google searches reporting a roughly 29 percent share of women in image output; this study extends that line to Switzerland and connects search measures to election results.","marker":"Rohrbach et al. (2024)"},{"why":"Experimental evidence that search-ranking manipulation can shift voter preferences, motivating the claim that search output can matter for electoral performance.","marker":"Epstein & Robertson (2015)"},{"why":"Describes the scaled-up virtual-agent infrastructure for search-engine audits, which the authors use to run many machines and control for personalization.","marker":"Ulloa et al. (2024c)"}],"fun_headline_variants":["Google search bias in Swiss election predicted voter outcomes","Swiss election audit: Google images favored men, stereotyped women","Google's gendered results in Swiss vote foreshadowed ballots","Audit of Swiss vote finds Google search favored men candidates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed correlation between Google search representation and personal votes is not driven by reverse causality, meaning that popular candidates are not simply generating more and more positive search results; the models do not control for prior vote share, offline media coverage, or campaign spending.","fun_headline_variants_meta":{"raw":{"variants":["Google search bias in Swiss election predicted voter outcomes","Swiss election audit: Google images favored men, stereotyped women","Google's gendered results in Swiss vote foreshadowed ballots","Audit of Swiss vote finds Google search favored men candidates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1727,"prompt_tokens":929,"completion_tokens":798,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":731}},"tokens_in":545,"tokens_out":798,"duration_ms":9509,"temperature":1.0,"reasoning_tokens":731,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:12:48.619184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the hierarchical regressions with the candidates' personal vote share from the 2019 Swiss federal election, or an equivalent prior-popularity measure, included as a control; if the 6 to 8 percent variance attributed to text and image search measures collapses toward zero, the predictive claim is an artifact of successful candidates generating more search results.","supporting_citations":[],"review_version":1}