Pith. sign in

REVIEW 3 cited by

Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11547 v2 pith:BSCGZIVI submitted 2024-09-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords humanhumanslanguagemodelscreativecreativitymodelwriting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we evaluate the creative fiction writing abilities of a fine-tuned small language model (SLM), BART-large, and compare its performance to human writers and two large language models (LLMs): GPT-3.5 and GPT-4o. Our evaluation consists of two experiments: (i) a human study in which 68 participants rated short stories from humans and the SLM on grammaticality, relevance, creativity, and attractiveness, and (ii) a qualitative linguistic analysis examining the textual characteristics of stories produced by each model. In the first experiment, BART-large outscored average human writers overall (2.11 vs. 1.85), a 14% relative improvement, though the slight human advantage in creativity was not statistically significant. In the second experiment, qualitative analysis showed that while GPT-4o demonstrated near-perfect coherence and used less cliche phrases, it tended to produce more predictable language, with only 3% of its synopses featuring surprising associations (compared to 15% for BART). These findings highlight how model size and fine-tuning influence the balance between creativity, fluency, and coherence in creative writing tasks, and demonstrate that smaller models can, in certain contexts, rival both humans and larger models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ArgCMV: An Argument Summarization Benchmark for the LLM-era

    cs.CL 2025-08 conditional novelty 6.0 of 10

    ArgCMV is a new LLM-curated benchmark of about 12,000 arguments from r/ChangeMyView, and current key point extraction methods transfer poorly to it.

  2. Pencils to Pixels: A Systematic Study of Creative Drawings across Children, Adults and AI

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A new dataset and computational framework quantify style and content in children's, adults', and DALL-E drawings, showing that expert and automated creativity scores disagree across groups.

  3. Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior?

    cs.SE 2025-11 reject novelty 3.0 of 10

    A five-model benchmark on a self-built malware dataset shows the smallest strong model, Phi-4-mini, outperforming 7-8B models, contradicting the paper's own 'larger models win' framing.

Pith tools