Pith. sign in

REVIEW 2 cited by

A Comparison of Large Language Model and Human Performance on Random Number Generation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.09656 v2 pith:U6FEYCSC submitted 2024-08-19 cs.AI cs.CLq-bio.NC

classification cs.AIcs.CLq-bio.NC
keywords numberrandomgenerationhumanchatgpt-3cognitivefrequencieshumans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Random Number Generation Tasks (RNGTs) are used in psychology for examining how humans generate sequences devoid of predictable patterns. By adapting an existing human RNGT for an LLM-compatible environment, this preliminary study tests whether ChatGPT-3.5, a large language model (LLM) trained on human-generated text, exhibits human-like cognitive biases when generating random number sequences. Initial findings indicate that ChatGPT-3.5 more effectively avoids repetitive and sequential patterns compared to humans, with notably lower repeat frequencies and adjacent number frequencies. Continued research into different models, parameters, and prompting methodologies will deepen our understanding of how LLMs can more closely mimic human random generation behaviors, while also broadening their applications in cognitive and behavioral science research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

    cs.CR 2026-07 conditional novelty 7.0 of 10

    Single-token answer distributions to everyday prompts fingerprint 165 served LLMs, recover family lineage, and verify claimed identity at 7.3% equal-error rate.

  2. In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models

    cs.AI 2026-04 conditional novelty 7.0 of 10

    VLMs can run Picbreeder but produce less refined, more mode-collapsed archives than humans; modest selection noise, short context, and many prompted personalities improve diversity metrics at quality cost.

Pith tools