REVIEW 4 major objections 7 minor 1 cited by
AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini
T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Across 236 countries, gpt-4o-mini tells the same story: a protagonist returns to a small town, restores tradition, and organises a community event.
desk verdict A useful, honest dataset and a genuinely interesting qualitative reading, but the paper's headline prevalence claim—'overwhelmingly conform'—is not actually measured, so the verdict stands or falls on whether the authors can either code a sample or soften the claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the dataset probe itself: a single minimal prompt repeated 11,850 times (50 stories for each of 236 countries plus 50 stories with no demonym), together with the concept of a synthetic imaginary—the model's statistically learned space of word associations that it draws on to imitate a “{demonym} story”. The minimal prompt is designed to push the model toward its default output, and the analysis then extracts the shared plot formula through word-frequency counts, word trees, sentiment summaries, and close readings of selected countries. The formula—return to a small town, minor conflict, community event, restored tradition—is the load-bearing object that carries the homogenisation claim.
What would settle it
Annotate a random sample of, say, 200 of the 11,800 stories for plot elements (protagonist returns home, small-town setting, community event as resolution, presence of romance, presence of direct conflict) and compare the distribution across countries with the distribution of the same elements in human-authored stories matched by country. If within-country variance in the generated stories is as large as between-country variance, the single-formula claim fails; if human-authored stories show the same concentration of return-and-community-event plots, the “human stories are more diverse” contrast fails.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that gpt-4o-mini's synthetic imaginary has a default story grammar. Whatever the nationality requested, the model produces a protagonist who returns home to a small town or village, faces a conflict that is downplayed or externalised, and resolves it by reconnecting with tradition and organising the community. Even stories for countries in active conflict follow this pattern: Palestinian stories emphasise standing firm and community organising rather than direct confrontation, Israeli stories individualise conflict behind vague opponents, and both lean on olive trees as symbols. The American stories are the clearest instance, with 23 of 50 titles beginning “The Last Train” and a structure close to the Hallmark-movie formula. The paper proposes calling this narrative standardisation: a formal bias at the level of plot, distinct from the word-and-image representational bias that dominates AI-bias research.
Load-bearing premise
The load-bearing premise is that one minimal English prompt and one model snapshot reveal the model's default story structure, rather than being an artifact of that exact wording, language, or model version.
Editorial extensions
If this is right
- If the default story grammar is as stable as claimed, AI-assisted writing tools that suggest or revise prose will tend to push storytellers toward nostalgia, reconciliation, and community-organising endings, crowding out plots built on conflict, change, or romance.
- AI-generated stories already posted to online story sites are becoming training data for future models, so the pattern can feed on itself unless deliberately countered.
- Measuring AI bias only at the level of words and images misses a systematic bias at the level of plot; cultural-alignment benchmarks would need to include narrative structure, not just values or stereotypes.
- For narratology, the finding supports the idea of “surface narration”: generated stories look story-like sentence by sentence but lack the coherence and causal drive of human-authored narratives.
Reading between the lines
- If the same experiment is repeated with prompts that specify genre, era, or emotional register, the uniformity may persist or break; the paper's own prompt-dependence caveat suggests this is the next test.
- A quantitative human baseline—annotating human-authored stories from the same countries with the same plot categories—would settle whether the claimed homogenisation is a property of the model or of the prompt.
- The narrow model choice (gpt-4o-mini, English prompts, early-2025 snapshot) means the finding is a lower bound on narrative standardisation; other models and multilingual prompts could show more or less flattening.
- The paper's closing anecdote about an AI-assisted journal edit hints that the same normalising pressure may extend to academic prose, not only fiction, though that connection is left undeveloped.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes a corpus of 11,850 stories generated by gpt-4o-mini (50 stories for each of 236 countries, plus 50 with no country specified) from the single prompt "Write a 1500 word potential {demonym} story." The authors report that, despite surface-level national symbols, the stories overwhelmingly conform to a single plot structure: a protagonist in a small town or village, often returning from a city, resolves a minor conflict by reconnecting with tradition and organizing a community event. They argue that real-world conflicts are sanitized, romance is nearly absent, and nostalgia and reconciliation replace narrative tension, constituting a distinct form of AI bias termed "narrative standardisation." The evidence combines word-frequency analyses, word trees, sentiment analysis on 50-word summaries, and close reading of stories from four countries plus sampling from others.
Significance. If the central empirical claim holds, this is a valuable and thought-provoking contribution. The dataset and analysis code are openly available, and the paper connects narratology, cultural AI bias, and computational literary studies in a way that few studies do. The word-tree comparisons for Palestinian and Israeli stories, the train symbolism analysis in American stories, and the plot diagram for Norwegian stories are illuminating and provide concrete, falsifiable observations. The proposed distinction between representational bias and structural or narrative bias is conceptually useful. However, the headline claim that the stories 'overwhelmingly conform to a single narrative plot structure across countries' is currently supported primarily by qualitative close reading rather than by a systematic, reproducible measure of plot structure. The paper also relies on a single prompt, a single model, and no human-story baseline, which limits the generality of the conclusion.
major comments (4)
- [Methodology ("Generating the dataset") and "Norwegian stories" (Figure 6)] The central claim that stories 'overwhelmingly conform to a single narrative plot structure across countries' (Abstract) is not supported by a direct quantitative measurement of plot structure. The authors state that 'The qualitative analysis was primarily done by Jill Walker Rettberg, based on a close reading of the Norwegian, American, Palestinian and Israeli stories and sampling stories from many other countries,' and Figure 6 is explicitly described as 'not generated computationally but by reading the stories and manually annotating them to identify shared plot points.' The word-frequency and sentiment analyses (Figures 1-2) measure surface vocabulary and affect, not plot events or their sequence. The paper needs a reproducible coding scheme for plot elements, applied to a representative sample (or, ideally, the full dataset), with inter-rater reliability and reported prevalence rates for the proposed standard plot. Without this, the claim of cross-country 'overwhelming' uniformity is an overgeneralization from a small hand-read subset.
- [Conclusion ("This requires further research") and "Generating the dataset"] The authors themselves raise the question, 'Is the similarity caused by our prompt, or is it a clue to how LLMs are inferring narrative structures from the training data...?' This is a load-bearing issue. With one prompt, one model (gpt-4o-mini), one generation window, and no comparison across prompt variants or models, the observed plot pattern could be an artifact of the particular instruction 'Write a 1500 word potential {demonym} story,' the length constraint, or the specific model version. The paper should include at least a small prompt-variation study (e.g., different phrasings, no length specification) and/or a second model to establish that the stability-and-tradition plot is a default narrative tendency rather than a response to this specific wording.
- [Introduction and Conclusion] The contrast with human-authored stories is asserted but not demonstrated: the paper states that 'Human-authored stories are far more diverse' and presents the Hallmark-movie comparison and the Norwegian Askeladden example as illustrations. Since the framing of the paper is about homogenisation relative to human narrative diversity, the authors should provide a matched baseline of human-authored stories from the same countries coded with the same plot scheme, or explicitly rescope the claim to state that the finding concerns homogeneity within the AI-generated corpus. Without a human baseline, the broader cultural-loss claim remains an interpretive leap rather than an empirical result.
- ["Generating the dataset" and "Word frequency"] The computational analyses do not actually cover all 236 countries on an equal footing. The authors note that 'Most, but not all, of the French, German and Indonesian stories are in French, German and Indonesian' and 'Many of the Danish stories are in Danish,' yet the word-frequency and noun-phrase analyses use an English-language spaCy pipeline and English lemmatization. Consequently, the reported counts for these countries substantially underrepresent the vocabulary of the generated stories. This limitation should be stated explicitly in the figure captions and in the limitations paragraph, and it further qualifies the cross-national comparisons in Figures 1 and 2.
minor comments (7)
- [Introduction] The word 'abandonned' should be 'abandoned' in the sentence about the train station.
- ["Norwegian stories"] The word 'reminiscient' should be 'reminiscent.'
- ["Palestinian and Israeli stories"] The word 'sanatised' should be 'sanitised.'
- ["American stories"] The phrase 'Mason-Dixie country' appears to be a typo for 'Mason-Dixon' (or the Mason-Dixon line).
- [Footnote 5] The aside about Trump tariffs is informal and likely to date the paper; consider removing it or moving it to a general note about the country list.
- [Conclusion] The anecdote about Paperpal Preflight is interesting but digresses from the narrative-standardisation argument; consider moving it to a separate section on AI in scientific publishing or shortening it substantially.
- ["Generating the dataset"] The sentiment analysis is performed on 50-word summaries generated by gpt-4o-mini rather than on the full stories, and the sentiment model was trained on Twitter data with no 'neutral' class. This is a limitation of the exploratory sentiment findings and should be mentioned wherever such findings are interpreted.
Circularity Check
No circularity: the paper's central claim is an empirical observation of LLM outputs, supported by human close reading and independent computational tools, with no fitted parameter, self-referential derivation, or equation-level reduction.
full rationale
The paper's load-bearing claim is that gpt-4o-mini generates stories that overwhelmingly conform to a single plot structure. This claim is an empirical description of generated outputs, not a derived quantity. The plot structure was identified by human close reading: the authors state that 'The qualitative analysis was primarily done by Jill Walker Rettberg, based on a close reading of the Norwegian, American, Palestinian and Israeli stories and sampling stories from many other countries', and that Figure 6's plot diagram 'was not generated computationally but by reading the stories and manually annotating them to identify shared plot points'. No parameter is fitted to a subset of the data and then presented as a prediction; no equation is defined in terms of the conclusion; no uniqueness theorem or prior result by the authors is invoked to force the interpretation. The ancillary use of gpt-4o-mini to produce 50-word plot summaries is explicitly limited to sentiment analysis with Distilbert, and the authors acknowledge its limitations; it does not generate the structural claim. The paper's own caveat that 'Is the similarity caused by our prompt, or is it a clue to how LLMs are inferring narrative structures from the training data...?' shows that the prompt-dependence concern is honestly flagged rather than hidden. Self-citations to the authors' own dataset and software are transparent data-availability statements, not load-bearing argumentation. The weakness that the 'overwhelmingly' prevalence claim rests on qualitative reading rather than a systematic coded sample is a methodological evidence concern, not circularity. Under the given criteria, there is no circular step to quote, and the correct finding is no significant circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption gpt-4o-mini is representative of 'AI-generated stories' generally.
- ad hoc to paper The prompt 'Write a 1500 word potential {demonym} story' elicits the model's default story structure rather than a prompt artifact.
- domain assumption Close reading of Norwegian, American, Palestinian and Israeli stories plus sampling from other countries is representative of all 236 countries.
- domain assumption Word frequency and sentiment analyses can indicate plot-structure uniformity.
- domain assumption Human-authored stories are more diverse than the generated corpus.
Cite this review
Pith. "Pith review of AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini." pith.science (2026). https://pith.science/paper/RUI4XQ4F
@misc{pith2026250722445,
author = {Pith},
title = {Pith review of: AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini},
year = {2026},
howpublished = {\url{https://pith.science/paper/RUI4XQ4F}},
note = {Machine review of arXiv:2507.22445}
}
read the original abstract
Can a language model trained largely on Anglo-American texts generate stories that are culturally relevant to other nationalities? To find out, we generated 11,800 stories - 50 for each of 236 countries - by sending the prompt "Write a 1500 word potential {demonym} story" to OpenAI's model gpt-4o-mini. Although the stories do include surface-level national symbols and themes, they overwhelmingly conform to a single narrative plot structure across countries: a protagonist lives in or returns home to a small town and resolves a minor conflict by reconnecting with tradition and organising community events. Real-world conflicts are sanitised, romance is almost absent, and narrative tension is downplayed in favour of nostalgia and reconciliation. The result is a narrative homogenisation: an AI-generated synthetic imaginary that prioritises stability above change and tradition above growth. We argue that the structural homogeneity of AI-generated narratives constitutes a distinct form of AI bias, a narrative standardisation that should be acknowledged alongside the more familiar representational bias. These findings are relevant to literary studies, narratology, critical AI studies, NLP research, and efforts to improve the cultural alignment of generative AI.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
The AI Fiction Paradox
AI's inability to generate compelling long-form fiction stems from narrative causation, informational revaluation, and multi-scale emotional architecture—three constraints current transformer architectures lack.
Reference graph
Works this paper leans on
-
[1]
Adams J, Brückner H, Naslund C: Who counts as a notable sociologist on Wikipedia? Gender, Race, and the ‘Professor Test.’ Socius. 2019; 5: 2378023118823946. Publisher Full Text Adilazuarda MF, Mukherjee S, Lavania P , et al.: Towards measuring and modeling ‘Culture’ in LLMs: a survey. arXiv
work page 2019
-
[8]
In: The Routledge Handbook of AI and Literature
Reference Source Ensslin A, Nelson J: Co-Creative multimodal authorship as procedural performance with DALL-E. In: The Routledge Handbook of AI and Literature. edited by Will Slocombe and Genevieve Liveley, 1st ed., New York: Routledge, 2024; 315–31. Publisher Full Text Ervik A: Generative AI and the collective imaginary. Image. 2023; 37(1): 42–57. Publis...
work page 2024
-
[11]
Reference Source Johnston DJ: Difference indifference. CounterText. 2024; 10(3): 230–31. Publisher Full Text Julien E: The Extroverted African Novel. The Novel, Volume 1, edited by Franco Moretti, Princeton University Press, 2006; 667–700. Publisher Full Text Lindsey J, Gurnee W, Ameisen E, et al.: On the biology of a large language model. Transformer Circuits
work page 2024
-
[12]
Reference Source Malevé N: On the data set’s ruins. AI & Society. 2021; 36(4): 1117–31. Publisher Full Text Malevé N: Lost in compression: models of authorship in generative AI. Media Theory. 2024; 8(1): 205–28. Publisher Full Text Marx L: The machine in the garden: technology and the pastoral ideal in America. [Thirty-Fitfth Anniversary Edition], Oxford:...
work page 2021
-
[15]
Reference Source Rettberg JW: Repeating ourselves with generative AI. CounterText. 2024; 10(3): 232–37. Publisher Full Text Rettberg S: Fin Du Monde: AI text-to-image writing and the digital unconscious. The Digital Review. 2024a
work page 2024
-
[16]
http://www.doi.org/10.18710/VM2K4O Rettberg S: Cyborg authorship: writing with AI – part 1: the trouble(s) with ChatGPT
-
[17]
Publisher Full Text Salvaggio E: How to read an AI image: toward a media studies methodology for the analysis of synthetic images. IMAGE. 2023; 37(1): 83–99. Publisher Full Text Saravia E, Liu HCT, Huang YH, et al.: CARER: Contextualized Affect Representations for Emotion Recognition. In: Proceedings of the 2018 Conference on Empirical Methods in Natural ...
work page 2023
-
[19]
Reference Source Sun Y: The presentation of modernity by trains in twentieth-century American literature. In: Proceedings of the 2022 4th International Conference on Literature, Art and Human Development (ICLAHD 2022). edited by Bootheina Majoul, Digvijay Pandya, and Lin Wang, Paris: Atlantis Press SARL, 2023; 1434–43. Publisher Full Text Tao Y, Viberg O,...
work page 2022
Show all 21 references
-
[20]
IEEE Trans Vis Comput Graph
Publisher Full Text Wattenberg M, Viegas FB: The word tree, an interactive visual concordance. IEEE Trans Vis Comput Graph. 2008; 14(6): 1221–28. PubMed Abstract | Publisher Full Text White J, Fu Q, Hays S, et al.: A prompt pattern catalog to enhance prompt engineering with ChatGPT
2008
-
[25]
2025; 1–21
New York, NY, USA: Association for Computing Machinery. 2025; 1–21. Publisher Full Text Baack S: A critical analysis of the largest source for generative AI training data: common crawl. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency. Rio de Janeiro B...
2025
-
[1991]
Feminist Stud
Reference Source Haraway D: Situated knowledges: the science question in feminism and the privilege of partial perspective. Feminist Stud. 1988; 14(3): 575–99. Publisher Full Text Hongisto T: Advertising with AI - On the presentation of authorship of ChatGPT-Generated Books. E...
1988
-
[2000]
In: Proceedings of the Fifth International Joint Conference on Artificial Intelligence
Reference Source Meehan JR: TALE-SPIN, an interactive program that writes stories. In: Proceedings of the Fifth International Joint Conference on Artificial Intelligence. MIT Cambridge, MA, USA, 1977; 91–98. Reference Source Munn L, Henrickson L: Tell me a story: a framework f...
1977
-
[2004]
Reference Source Page 17 of 17 Open Research Europe 2025, 5:202 Last updated: 29 JUL 2025
2025
-
[2006]
In: The Routledge Handbook of AI and Literature
Reference Source Ghosal T: Towards narrative AI studies. In: The Routledge Handbook of AI and Literature. by Will Slocombe and Genevieve Liveley, 1st ed., New York: Routledge, 2024; 269–87. Publisher Full Text Gillespie T: Generative AI and the politics of visibility. Big Data...
2024
-
[2016]
Television & New Media
Publisher Full Text Braithwaite A: From brand to genre: the hallmark movie. Television & New Media. 2023; 24(5): 488–98. Publisher Full Text Brown TB, Mann B, Ryder N, et al.: Language models are few-shot learners. arXiv
2023
-
[2020]
Information, Communication & Society
Publisher Full Text Bucher T: The algorithmic imaginary: exploring the ordinary affects of Facebook algorithms. Information, Communication & Society. 2016; 1–15. Publisher Full Text Carter R: Machine visions: mapping depictions of machine vision through AI image synthesis. Ope...
2016
-
[2021]
Emerging Media
Publisher Full Text Barroso Da Silveira J, Lima EA: Racial biases in AIs and Gemini’s inability to write narratives about black people. Emerging Media. 2024; 2(2): 277–87. Publisher Full Text Beduschi A: Synthetic data protection: towards a paradigm change in data regulation? ...
2024
-
[2022]
M/C Journal
Publisher Full Text Srdarov S, Leaver T: Generative AI glitches: the artificial everything. M/C Journal. 2024; 27(6). Publisher Full Text Sun J, Peng N: Men are elected, women are married: events gender bias on Wikipedia. arXiv,
2024
-
[2023]
In: Computational creativity research: towards creative machines
Reference Source Pérez y Pérez R: From MEXICA to MEXICA-Impro: the evolution of a computer model for plot generation. In: Computational creativity research: towards creative machines. edited by Tarek R. Besold, Marco Schorlemmer, and Alan Smaill. Atlantis Thinking Machines. Pa...
2015
-
[2024]
In: Proceedings of the 2025 CHI Conference on human factors in computing systems
Publisher Full Text Agarwal D, Naaman M, Vashistha A: AI suggestions homogenize writing toward western styles and diminish cultural nuances. In: Proceedings of the 2025 CHI Conference on human factors in computing systems. CHI ’
2025
-
[2025]
Examining the risks of A.I
Reference Source De Ninno F, Lacriola M: Mussolini and ChatGPT. Examining the risks of A.I. writing historical narratives on fascism. J Mod Ital Stud. 2025; 30(2): 187–209. Publisher Full Text Page 16 of 17 Open Research Europe 2025, 5:202 Last updated: 29 JUL 2025 De Seta G: ...
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.