Pith. sign in

REVIEW 4 cited by

How Random is Random? Evaluating the Randomness and Humaness of LLMs' Coin Flips

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00092 v1 pith:RC63OWBO submitted 2024-05-31 cs.AI cs.LG

classification cs.AIcs.LG
keywords randomhumanllmsrandomnessbehaviorhumanessapproachbias
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One uniquely human trait is our inability to be random. We see and produce patterns where there should not be any and we do so in a predictable way. LLMs are supplied with human data and prone to human biases. In this work, we explore how LLMs approach randomness and where and how they fail through the lens of the well studied phenomena of generating binary random sequences. We find that GPT 4 and Llama 3 exhibit and exacerbate nearly every human bias we test in this context, but GPT 3.5 exhibits more random behavior. This dichotomy of randomness or humaness is proposed as a fundamental question of LLMs and that either behavior may be useful in different circumstances.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

    cs.CR 2026-07 conditional novelty 7.0 of 10

    Single-token answer distributions to everyday prompts fingerprint 165 served LLMs, recover family lineage, and verify claimed identity at 7.3% equal-error rate.

  2. In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models

    cs.AI 2026-04 conditional novelty 7.0 of 10

    VLMs can run Picbreeder but produce less refined, more mode-collapsed archives than humans; modest selection noise, short context, and many prompted personalities improve diversity metrics at quality cost.

  3. CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting

    cs.LG 2025-05 accept novelty 7.0 of 10

    Capping achievable accuracy with randomized correct answers turns any model that exceeds the cap into a detectable contamination alarm.

  4. B-score: Detecting biases in large language models using response history

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs self-correct toward uniform answers in multi-turn repetition, and the gap between single-turn and multi-turn answer rates (B-score) flags biased answers better than verbalized confidence.

Pith tools