Pith. sign in

REVIEW 2 cited by

Are Large Language Models Capable of Generating Human-Level Narratives?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13248 v2 pith:WBQBGHNP submitted 2024-07-18 cs.CL

classification cs.CL
keywords narrativestoriesstorytellingarousaldiscoursellmsnarrativesabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper investigates the capability of LLMs in storytelling, focusing on narrative development and plot progression. We introduce a novel computational framework to analyze narratives through three discourse-level aspects: i) story arcs, ii) turning points, and iii) affective dimensions, including arousal and valence. By leveraging expert and automatic annotations, we uncover significant discrepancies between the LLM- and human- written stories. While human-written stories are suspenseful, arousing, and diverse in narrative structures, LLM stories are homogeneously positive and lack tension. Next, we measure narrative reasoning skills as a precursor to generative capacities, concluding that most LLMs fall short of human abilities in discourse understanding. Finally, we show that explicit integration of aforementioned discourse features can enhance storytelling, as is demonstrated by over 40% improvement in neural storytelling in terms of diversity, suspense, and arousal.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can LLMs Generate Good Stories? Insights and Challenges from a Narrative Planning Perspective

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new automatically verified benchmark shows GPT-4 tier LLMs can do small-scale causal story planning, but character intentionality and dramatic conflict remain hard except for reasoning models like o1.

  2. Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

    cs.CY 2026-06 conditional novelty 5.0 of 10

    Across 25,000 stories from five LLMs, an LLM judge rated stories mentioning intellectual disabilities as more infantile, paternalistic, dependent, and inspirational than stories without the label.

Pith tools