Pith. sign in

REVIEW 2 cited by

Probing the Creativity of Large Language Models: Can models produce divergent semantic association?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.11158 v1 pith:ZUDB6OZ6 submitted 2023-10-17 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelslanguagecreativitylargedivergentsemanticassociationcreative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models possess remarkable capacity for processing language, but it remains unclear whether these models can further generate creative content. The present study aims to investigate the creative thinking of large language models through a cognitive perspective. We utilize the divergent association task (DAT), an objective measurement of creativity that asks models to generate unrelated words and calculates the semantic distance between them. We compare the results across different models and decoding strategies. Our findings indicate that: (1) When using the greedy search strategy, GPT-4 outperforms 96% of humans, while GPT-3.5-turbo exceeds the average human level. (2) Stochastic sampling and temperature scaling are effective to obtain higher DAT scores for models except GPT-4, but face a trade-off between creativity and stability. These results imply that advanced large language models have divergent semantic associations, which is a fundamental process underlying creativity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. We're Different, We're the Same: Creative Homogeneity Across LLMs

    cs.CY 2025-01 conditional novelty 6.0 of 10

    Across three divergent-thinking tests, responses from seven LLM families were substantially more similar to one another than responses from 102 humans were to one another.

  2. Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLM-as-judge scoring of multi-round lateral thinking tasks can be fooled by answer leakage and question substitution, so response-based metrics may overstate reasoning ability.

Pith tools