REVIEW 4 major objections 5 minor 2 cited by
Generative Engine Optimization: How to Dominate AI Search
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This paper tries to establish that generative AI search systems—ChatGPT, Perplexity, Gemini, Claude—share a systematic sourcing bias: they draw their answers overwhelmingly from earned media (third-party editorial and review sites) rather t
desk verdict Useful empirical snapshot of AI search sourcing, but the universal "earned media bias" claim rests on ranking-style prompts and no statistics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The empirical engine is a controlled comparative pipeline: ranking-style prompts issued to each engine, citation URLs normalized to registrable domains, domains classified by GPT-4o-search-preview into Brand / Earned / Social, and outputs compared with overlap metrics (Coverage@k, Jaccard index). This machinery lets the authors separate stable structural tendencies (earned-media dominance) from engine-specific and language-specific variation.
What would settle it
Run the same citation-classification pipeline on conversational and task-oriented queries from the paper's own taxonomy (coding help, self-improvement, creative writing) and check whether the earned-media share falls below the measured 60–90%; or repeat the August 2025 measurements in a later window and check whether the brand/social shares have converged toward Google's mix.
Extended reading notes
Core claim
The central discovery is that under standardized consumer ranking queries, web-enabled AI engines allocate most of their citations to earned media domains (typically 60–90% depending on engine and vertical), nearly exclude social sources (often 0–5%), and relegate brand-owned sites to a secondary role, while Google returns a more balanced mix of brand, earned, and social domains. The paper also documents that engines differ in how they handle language: Claude reuses English authority domains across languages, GPT swaps to local-language ecosystems, and Perplexity/Gemini sit in between; paraphrase shifts move citations less than language shifts; and local-service queries produce very low over
Load-bearing premise
The results rest on the assumption that standardized ranking-style prompts (e.g., 'Top 10... brands') in a handful of consumer verticals capture how people actually use AI search; if real queries are predominantly conversational, long-tail, or non-ranking, the measured dominance of earned media may be an artifact of the prompt format.
Editorial extensions
If this is right
- Content strategists should treat earned media placements as the primary lever for AI search visibility; on-page SEO remains necessary but not sufficient.
- Websites should be structured for machine scannability (schema markup, comparison tables, explicit value propositions) so AI agents can extract justification.
- Multilingual visibility requires a dual path: local-language earned media for GPT/Perplexity, English-language authority for Claude.
- Niche brands face an inherent big-brand bias; they need niche authority and Perplexity-style social/YouTube presence to break in.
- Because engine behaviors are a moving target, conclusions need periodic re-measurement.
Reading between the lines
- The results imply a shift in marketing budgets: if AI search favors earned media, brand-owned content's ROI may decline relative to PR and third-party coverage; brands that currently invest heavily in owned content may need to rebalance.
- The near-zero social share in GPT answers suggests that community platforms (Reddit, forums) are losing a channel in AI-mediated discovery, which could alter traffic patterns for those platforms.
- The fact that engines overlap little with each other means the same query can yield materially different brands/answers across assistants; this raises a testable question about verifiability: a user who relies on only one assistant may get systematically different recommendations.
- The methodology (ranking-style prompts) suggests an untested boundary: the ground truth for 'earned bias' under conversational or long-tail queries may differ, so the next natural experiment is to run the pipeline on the taxonomy's non-ranking categories.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares web-enabled AI search systems (ChatGPT/GPT-4o-search-preview, Perplexity, Gemini, Claude) with traditional Google search across consumer verticals, languages, and paraphrases. It measures domain overlap, media-type mix (Brand/Earned/Social), freshness, and cross-engine diversity, concluding that AI search exhibits a systematic and overwhelming bias toward earned media relative to Google, and that engines differ in diversity, freshness, cross-language stability, and paraphrase sensitivity. Based on these results, the paper proposes a strategic Generative Engine Optimization (GEO) agenda for brands and practitioners.
Significance. If the empirical claims hold, the paper provides timely, multi-engine evidence about a major shift in information access, with direct implications for SEO/GEO practice and for the study of AI-mediated discovery. The strengths are the breadth of the comparison (four AI engines plus Google, multiple verticals and languages), the use of direct API outputs rather than simulated relevance judgments, and the explicit acknowledgement of temporal and black-box limitations in Section 7. However, the central empirical claim is currently supported by a restricted query protocol and by measurement instruments that share the same model family as the objects of study; these issues need to be addressed before the broader conclusions can be accepted.
major comments (4)
- [§4.1.1, §4.2.1, §4.3] The abstract's universal claim that 'AI Search exhibit a systematic and overwhelming bias towards Earned media' is established almost exclusively through ranking-style templates. §4.1.1 states that prompts are standardized to 'Top 10 ...' and §4.2.1 uses 1,000 such prompts; §4.2.3–4.2.4 and §5.2.x similarly use ranking or identification prompts. §4.3 is the only section with non-ranking intents, but it gives no query inventory, sample size, or engine coverage, and Figures 18–20 appear without methodology. Ranking prompts naturally elicit listicles and review roundups, so the measured earned-media dominance and near-zero social share may be a template artifact. The paper's own §3 taxonomy includes coding, creative writing, and self-improvement queries that are not 'Top 10' lists. Please either restrict the central claim to ranking-style prompts or add a comparable non-ranking experiment u
- [§4.1.4, §5.1] The Brand/Earned/Social classification is performed by GPT-4o-search-preview, and §5.1 further describes 'GPT-assisted classification prompts' for the brand experiments. The classifier is from the same model family as several systems under study (GPT itself, and to a lesser extent the other LLM-based engines), so the measured category shares could reflect the classifier's own sourcing biases rather than the engines' behavior. No validation against human labels, inter-annotator agreement, or an independent classifier is reported. Please provide a human-labeled validation set, report classification accuracy, and show that the main distributions are robust to using a rule-based or external classifier.
- [§4.2.1, §5.2.2] All reported percentages and overlap values are point estimates. For example, Section 4.2.1 gives figures such as 69.1% Earned and 0% Social for AI search in Canada, and Section 5.2.2 reports Claude and GPT at approximately 93.7% and 93.6% Earned, but no confidence intervals, hypothesis tests, or query-level variance are provided. The claims of 'systematic' and 'overwhelming' differences require statistical support, especially because several comparisons are based on 10 verticals and 100 queries, where sampling noise is non-negligible. Report per-query distributions, error bars, and formal tests (e.g., paired comparisons of the Earned share between AI and Google).
- [§5.2.6] The 'Big Brand Bias' experiment depends on curated sets of 20 major and 20 niche brands (Table 1). The selection of these specific brands is not justified, and the conclusion that Coca-Cola and Pepsi dominate may be a property of the chosen lists rather than a general property of the systems. Please describe how the brand sets were constructed, or show that the result is stable under alternative brand lists.
minor comments (5)
- [§4.2.4, §5.2.5] Cross-reference errors: §4.2.4's interpretation refers to 'language choice (§4.2.2)', but the language experiment is §4.2.3; §5.2.5's objective refers to 'prior experiments (§5.3)', which should be §5.2.4 or §5.2.2.
- [References] Reference [10] is labeled 'WIRED' but points to a Search Engine Land URL; the inline citation to 'reddit []' has an empty bracket. Please correct these.
- [Acknowledgments] The acknowledgment to ktau.ai, a commercial GEO company, should be accompanied by a conflict-of-interest statement, given the paper's prescriptive GEO agenda.
- [§4.3] The Informational/Consideration/Transactional taxonomy in §4.3 is not defined or linked to the Reddit-derived taxonomy in §3. It would be helpful to state how the two taxonomies relate and to provide example queries for each intent bucket.
- [Reproducibility] The paper states reproducibility 'in mind' (e.g., §5.2.4) but does not release query lists, raw outputs, or code. Given the API-based methodology, making these available as an artifact would substantially strengthen the paper.
Circularity Check
No significant circularity: the paper is an empirical measurement study; its central claims do not reduce to fitted parameters, self-citations, or definitional equivalences.
full rationale
The paper's derivation chain is a direct API-measurement pipeline: it issues queries to Google and AI engines, extracts cited domains, classifies them into Brand/Earned/Social, and aggregates shares. No parameter is fitted and then renamed a prediction; no uniqueness theorem is imported from the authors' prior work; and the only prior-work citation (Aggarwal et al. for the GEO term) is not load-bearing and is not a self-citation. The central 'earned media bias' claim is an empirical distribution, not an identity: using ranking-style prompts may affect external validity, but the observed dominance of earned domains is not true by construction. The use of GPT-4o-search-preview as the domain classifier is a measurement-validity concern, since the instrument overlaps with one studied engine; the paper itself discloses the subjectivity of its classification in §7 (Assumptions and Limitations). This is a methodological limitation, not a circular reduction. Similarly, the Reddit-derived taxonomy in §3 covers a broader query space than the standardized prompts, which is a generalizability limitation rather than circularity. The cola 'big brand bias' experiment also uses popularity-framed queries, but this is an experimental-design confound, not a derivation circle. Overall, no circular step meeting the evidence bar is present.
Assumptions & free parameters
free parameters (3)
- top-k cutoff =
10 (and 5 in some analyses)
- freshness scoring formula =
freshness = mean(1/(1+age)); coverage-adjusted = freshness * coverage
- curated major/niche brand sets =
20 major and 20 niche cola brands listed in Table 1
assumptions (5)
- domain assumption Citations returned by web-enabled AI engines are an accurate proxy for the sources used to generate the answer and for visibility in AI search.
- domain assumption GPT-4o-based classification of domains into Brand/Earned/Social is sufficiently accurate to support the reported percentages.
- domain assumption Ranking-style prompts are representative of real user interactions with AI search.
- domain assumption The outputs of these AI services are stable enough within the data-collection window that pooling queries and citing domains is meaningful.
- ad hoc to paper The three-category Brand/Earned/Social taxonomy is a valid projection of the information ecosystem.
Cite this review
Pith. "Pith review of Generative Engine Optimization: How to Dominate AI Search." pith.science (2026). https://pith.science/paper/5ZSFY6RB
@misc{pith2026250908919,
author = {Pith},
title = {Pith review of: Generative Engine Optimization: How to Dominate AI Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/5ZSFY6RB}},
note = {Machine review of arXiv:2509.08919}
}
read the original abstract
The rapid adoption of generative AI-powered search engines like ChatGPT, Perplexity, and Gemini is fundamentally reshaping information retrieval, moving from traditional ranked lists to synthesized, citation-backed answers. This shift challenges established Search Engine Optimization (SEO) practices and necessitates a new paradigm, which we term Generative Engine Optimization (GEO). This paper presents a comprehensive comparative analysis of AI Search and traditional web search (Google). Through a series of large-scale, controlled experiments across multiple verticals, languages, and query paraphrases, we quantify critical differences in how these systems source information. Our key findings reveal that AI Search exhibit a systematic and overwhelming bias towards Earned media (third-party, authoritative sources) over Brand-owned and Social content, a stark contrast to Google's more balanced mix. We further demonstrate that AI Search services differ significantly from each other in their domain diversity, freshness, cross-language stability, and sensitivity to phrasing. Based on these empirical results, we formulate a strategic GEO agenda. We provide actionable guidance for practitioners, emphasizing the critical need to: (1) engineer content for machine scannability and justification, (2) dominate earned media to build AI-perceived authority, (3) adopt engine-specific and language-aware strategies, and (4) overcome the inherent "big brand bias" for niche players. Our work provides the foundational empirical analysis and a strategic framework for achieving visibility in the new generative search landscape.
Figures
Figures from the paper (39 more)
Forward citations
Cited by 2 Pith papers
-
Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)
A critical review of GEO research concludes that already-retrieved content can improve citation and use, but no tested technique reliably raises organic discoverability or downstream traffic across engines.
-
AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework
A new audit framework, GEO-16, links 16 on-page quality signals to AI answer engine citations, identifying thresholds associated with higher citation odds.
Reference graph
Works this paper leans on
-
[1]
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande. 2024. GEO: Generative Engine Optimization. arXiv:2311.09735 [cs.LG] https://arxiv.org/abs/2311.09735
arXiv 2024
-
[2]
2025.CEO says Perplexity hit 780M queries in May
Perplexity AI. 2025.CEO says Perplexity hit 780M queries in May
2025
-
[3]
2025.About a third of U.S
Pew Research Center. 2025.About a third of U.S. adults have used ChatGPT; usage has increased over the past year. https://www.pewresearch.org/short-reads/2025/ 06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/ [Online]. Available
2025
-
[4]
2025.Google users are less likely to click on links when an AI summary appears in the results
Pew Research Center. 2025.Google users are less likely to click on links when an AI summary appears in the results. https://www.pewresearch.org/short- reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai- summary-appears-in-the-results/ [Online]. Available
2025
-
[5]
Aounon Kumar and Himabindu Lakkaraju. 2024. Manipulating Large Language Models to Increase Product Visibility. arXiv:2404.07981 [cs.IR] https://arxiv.org/ abs/2404.07981
arXiv 2024
-
[6]
2024.ChatGPT Surges to 3.1B Visits in September 2024
Similarweb. 2024.ChatGPT Surges to 3.1B Visits in September 2024. https://www.similarweb.com/blog/insights/ai-news/chatgpt-topped-3- billion-visits-in-september/ [Online]. Available
2024
-
[7]
2025.AI Chatbot Market Share Worldwide
StatCounter Global Stats. 2025.AI Chatbot Market Share Worldwide. https: //gs.statcounter.com/ai-chatbot-market-share [Online]. Available
2025
-
[8]
2025.Search Engine Market Share Worldwide
StatCounter Global Stats. 2025.Search Engine Market Share Worldwide. https: //gs.statcounter.com/search-engine-market-share [Online]. Available
2025
Show all 11 references
-
[9]
Alexander Wan, Eric Wallace, and Dan Klein. 2024. What Evidence Do Language Models Find Convincing? arXiv:2402.11782 [cs.CL] https://arxiv.org/abs/2402. 11782
2024 arXiv
-
[10]
2024.Google Cut Back AI Overviews in Search Even Before Its ’Pizza Glue’ Fiasco
WIRED. 2024.Google Cut Back AI Overviews in Search Even Before Its ’Pizza Glue’ Fiasco. https://searchengineland.com/google-ai-overviews-visibility-drops-15- percent-queries-442850?utm_source=chatgpt.com [Online]. Available
2024
-
[2025]
Available
https://www.perplexity.ai/page/ceo-says-perplexity-hit-780m-q- dENgiYOuTfaMEpxLQc2bIQ [Online]. Available
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.