REVIEW 17 cited by
Is Temperature the Creativity Parameter of Large Language Models?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are applied to all sorts of creative tasks, and their outputs vary from beautiful, to peculiar, to pastiche, into plain plagiarism. The temperature parameter of an LLM regulates the amount of randomness, leading to more diverse outputs; therefore, it is often claimed to be the creativity parameter. Here, we investigate this claim using a narrative generation task with a predetermined fixed context, model and prompt. Specifically, we present an empirical analysis of the LLM output for different temperature values using four necessary conditions for creativity in narrative generation: novelty, typicality, cohesion, and coherence. We find that temperature is weakly correlated with novelty, and unsurprisingly, moderately correlated with incoherence, but there is no relationship with either cohesion or typicality. However, the influence of temperature on creativity is far more nuanced and weak than suggested by the "creativity parameter" claim; overall results suggest that the LLM generates slightly more novel outputs as temperatures get higher. Finally, we discuss ideas to allow more controlled LLM creativity, rather than relying on chance via changing the temperature parameter.
Forward citations
Cited by 17 Pith papers
-
More Is Not More: What Matters for Diversity in LLM Opinions?
Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.
-
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
A sequential greedy-then-high-temperature decoding method with a learned refusal-prefix gate preserves 91-99% of greedy refusal responses at high temperatures.
-
Assessing the Business Process Modeling Competences of Large Language Models
Open-source LLMs can produce BPMN process models that rival human experts on syntax and readability, but they lag on semantic accuracy and frequently generate invalid BPMN-XML.
-
An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs
EASI-RAG is a structured agile method for deploying RAG tools in industrial SMEs, validated by one case study where a no-experience team built a working assistant in three weeks.
-
Intent Factored Generation: Unleashing the Diversity in Your Language Model
Intent Factored Generation samples a high-temperature intent, such as keywords or a summary, and then samples the final response at lower temperature conditioned on that intent, increasing semantic diversity while kee...
-
Evaluating the Evaluation of Diversity in Commonsense Generation
Content-based diversity metrics, such as Vendi Score and Chamfer distance, agree with LLM-based diversity ratings far better than form-based metrics like self-BLEU across three commonsense generation datasets.
-
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
EarthSE provides a two-level QA benchmark and an open-ended dialogue benchmark for Earth science and shows current LLMs perform poorly on both.
-
Dynamic Reinforcement Learning for Actors
A reinforcement learning update that adjusts each neuron's input-output sensitivity using TD error can replace external exploration noise and backpropagation through time in small actor-critic tasks.
-
Enhance-A-Video: Better Generated Video for Free
Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.
-
White Hat Search Engine Optimization using Large Language Models
LLM prompts that include past rankings produce document edits that improve retrieval ranking more than human students and a feature-based baseline, while keeping the text faithful.
-
Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems
A RAG-LLM pipeline extracts transportation and energy policy items from U.S. climate equity plans and recommends cities with similar policy practices, but extraction is not validated against human coding.
-
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
A keyframe-plus-GPT-4o pipeline retrieves the single most representative frame for a reported gameplay bug, with F1@1 of 0.79 and Accuracy@1 of 0.89 on industrial bug-report videos.
-
Investigating the Performance of Small Language Models in Detecting Test Smells in Manual Test Cases
Small language models with a targeted prompting scheme detect seven test-smell types in natural-language Ubuntu manual tests, with pass@2 scores of 90-97% across three models.
-
Cognitive Agents Powered by Large Language Models for Agile Software Project Management
LLM agents acting as Agile roles produced plausible project artifacts in simulation, but the claimed improvements over human teams are unsupported because no comparison or validated metrics are provided.
-
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
A broad benchmark of six open-weights LLMs shows prompt design and chunking affect summarization quality more than model size alone.
-
The Paradox of Stochasticity: Limited Creativity and Computational Decoupling in Temperature-Varied LLM Outputs of Structured Fictional Data
In three LLMs generating fictional names and birthdates, model choice dominates processing time and default name archetypes persist across temperature, while rare names appear mainly at mid-range temperatures.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Discussion (0). Continue with ORCID to comment.