Pith. sign in

REVIEW 4 cited by

Art or Artifice? Large Language Models and the False Promise of Creativity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.14556 v3 pith:DXDGZHIF submitted 2023-09-25 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords ttcwcreativityllmsstoriescreativewritingassessmentlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback

    cs.CL 2025-07 conditional novelty 6.0 of 10

    On a new controlled dataset of corrupted short stories, eight LLMs produce mostly correct, specific writing feedback but often fail to identify the biggest writing issue and are poor at deciding when to say a story is...

  2. Dynamic Reinforcement Learning for Actors

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A reinforcement learning update that adjusts each neuron's input-output sensitivity using TD error can replace external exploration noise and backpropagation through time in small actor-critic tasks.

  3. Polymind: Parallel Visual Diagramming with Large Language Models to Support Prewriting Through Microtasks

    cs.HC 2025-02 conditional novelty 6.0 of 10

    Polymind introduces parallel, configurable LLM microtasks on a diagramming canvas for prewriting, and a small user study indicates it affords users more control and customization than turn-taking chatbot interaction.

  4. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

Pith tools