Pith. sign in

REVIEW 17 cited by

Is Temperature the Creativity Parameter of Large Language Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00492 v1 pith:V44WXGM5 submitted 2024-05-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords creativitytemperatureparameteroutputsclaimcohesioncorrelatedgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are applied to all sorts of creative tasks, and their outputs vary from beautiful, to peculiar, to pastiche, into plain plagiarism. The temperature parameter of an LLM regulates the amount of randomness, leading to more diverse outputs; therefore, it is often claimed to be the creativity parameter. Here, we investigate this claim using a narrative generation task with a predetermined fixed context, model and prompt. Specifically, we present an empirical analysis of the LLM output for different temperature values using four necessary conditions for creativity in narrative generation: novelty, typicality, cohesion, and coherence. We find that temperature is weakly correlated with novelty, and unsurprisingly, moderately correlated with incoherence, but there is no relationship with either cohesion or typicality. However, the influence of temperature on creativity is far more nuanced and weak than suggested by the "creativity parameter" claim; overall results suggest that the LLM generates slightly more novel outputs as temperatures get higher. Finally, we discuss ideas to allow more controlled LLM creativity, rather than relying on chance via changing the temperature parameter.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 43 citations worldwide. Full citation record

  1. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0 of 10

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  2. Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A sequential greedy-then-high-temperature decoding method with a learned refusal-prefix gate preserves 91-99% of greedy refusal responses at high temperatures.

  3. Assessing the Business Process Modeling Competences of Large Language Models

    cs.SE 2026-01 conditional novelty 6.0 of 10

    Open-source LLMs can produce BPMN process models that rival human experts on syntax and readability, but they lag on semantic accuracy and frequently generate invalid BPMN-XML.

  4. An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    EASI-RAG is a structured agile method for deploying RAG tools in industrial SMEs, validated by one case study where a no-experience team built a working assistant in three weeks.

  5. Intent Factored Generation: Unleashing the Diversity in Your Language Model

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Intent Factored Generation samples a high-temperature intent, such as keywords or a summary, and then samples the final response at lower temperature conditioned on that intent, increasing semantic diversity while kee...

  6. Evaluating the Evaluation of Diversity in Commonsense Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Content-based diversity metrics, such as Vendi Score and Chamfer distance, agree with LLM-based diversity ratings far better than form-based metrics like self-BLEU across three commonsense generation datasets.

  7. EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EarthSE provides a two-level QA benchmark and an open-ended dialogue benchmark for Earth science and shows current LLMs perform poorly on both.

  8. Dynamic Reinforcement Learning for Actors

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A reinforcement learning update that adjusts each neuron's input-output sensitivity using TD error can replace external exploration noise and backpropagation through time in small actor-critic tasks.

  9. Enhance-A-Video: Better Generated Video for Free

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.

  10. White Hat Search Engine Optimization using Large Language Models

    cs.IR 2025-02 conditional novelty 6.0 of 10

    LLM prompts that include past rankings produce document edits that improve retrieval ranking more than human students and a feature-based baseline, while keeping the text faithful.

  11. Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems

    cs.CY 2026-01 conditional novelty 5.0 of 10

    A RAG-LLM pipeline extracts transportation and energy policy items from U.S. climate equity plans and recommends cities with similar policy practices, but extraction is not validated against human coding.

  12. Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models

    cs.SE 2025-08 conditional novelty 5.0 of 10

    A keyframe-plus-GPT-4o pipeline retrieves the single most representative frame for a reported gameplay bug, with F1@1 of 0.79 and Accuracy@1 of 0.89 on industrial bug-report videos.

  13. Investigating the Performance of Small Language Models in Detecting Test Smells in Manual Test Cases

    cs.SE 2025-07 conditional novelty 5.0 of 10

    Small language models with a targeted prompting scheme detect seven test-smell types in natural-language Ubuntu manual tests, with pass@2 scores of 90-97% across three models.

  14. Cognitive Agents Powered by Large Language Models for Agile Software Project Management

    cs.SE 2025-08 reject novelty 4.0 of 10

    LLM agents acting as Agile roles produced plausible project artifacts in simulation, but the claimed improvements over human teams are unsupported because no comparison or validated metrics are provided.

  15. An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A broad benchmark of six open-weights LLMs shows prompt design and chunking affect summarization quality more than model size alone.

  16. The Paradox of Stochasticity: Limited Creativity and Computational Decoupling in Temperature-Varied LLM Outputs of Structured Fictional Data

    cs.LG 2025-02 conditional novelty 4.0 of 10

    In three LLMs generating fictional names and birthdates, model choice dominates processing time and default name archetypes persist across temperature, while rare names appear mainly at mid-range temperatures.

  17. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    cs.LG 2025-02

Pith tools