REVIEW 6 cited by
We're Different, We're the Same: Creative Homogeneity Across LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Numerous powerful large language models (LLMs) are now available for use as writing support tools, idea generators, and beyond. Although these LLMs are marketed as helpful creative assistants, several works have shown that using an LLM as a creative partner results in a narrower set of creative outputs. However, these studies only consider the effects of interacting with a single LLM, begging the question of whether such narrowed creativity stems from using a particular LLM -- which arguably has a limited range of outputs -- or from using LLMs in general as creative assistants. To study this question, we elicit creative responses from humans and a broad set of LLMs using standardized creativity tests and compare the population-level diversity of responses. We find that LLM responses are much more similar to other LLM responses than human responses are to each other, even after controlling for response structure and other key variables. This finding of significant homogeneity in creative outputs across the LLMs we evaluate adds a new dimension to the ongoing conversation about creativity and LLMs. If today's LLMs behave similarly, using them as a creative partners -- regardless of the model used -- may drive all users towards a limited set of "creative" outputs.
Forward citations
Cited by 6 Pith papers
-
Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation
LLM outputs are usually plausible but cover far less of the human response space than matched human samples do, with the biggest gap at the periphery of the human distribution.
-
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.
-
Optimization Is Not All You Need
Optimization can measure how improbable generated text is but cannot tell whether that unlikelihood is error or invention, yet it now sets the protocols of legitimate language.
-
Correlated Errors in Large Language Models
Large language models from different providers and architectures often make the same errors, and more accurate models are especially likely to share mistakes.
-
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis
A meta-analysis of 28 studies finds no average creativity gap between GenAI and humans, a small boost when humans collaborate with GenAI, and a large drop in idea diversity in those collaborations.
-
Avoidance Decoding for Diverse Multi-Branch Story Generation
Avoidance Decoding penalizes token choices that resemble previously generated story branches, using a hybrid concept-level and narrative-level similarity penalty, and reports large diversity gains across several LLMs.
Discussion (0). Continue with ORCID to comment.