Pith. sign in

REVIEW 12 cited by

The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10457 v1 pith:QJIWELKM submitted 2024-07-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsnon-determinismsamplinggreedyperformancealignmentdecodingevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current evaluations of large language models (LLMs) often overlook non-determinism, typically focusing on a single output per example. This limits our understanding of LLM performance variability in real-world applications. Our study addresses this issue by exploring key questions about the performance differences between greedy decoding and sampling, identifying benchmarks' consistency regarding non-determinism, and examining unique model behaviors. Through extensive experiments, we observe that greedy decoding generally outperforms sampling methods for most evaluated tasks. We also observe consistent performance across different LLM sizes and alignment methods, noting that alignment can reduce sampling variance. Moreover, our best-of-N sampling approach demonstrates that smaller LLMs can match or surpass larger models such as GPT-4-Turbo, highlighting the untapped potential of smaller LLMs. This research shows the importance of considering non-determinism in LLM evaluations and provides insights for future LLM development and evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Verifier-free Test-Time Sampling for Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 7.0 of 10

    A verifier-free test-time sampling method for vision-language-action models that selects actions by KL divergence to a condition-masked reference distribution, improving task success rates.

  2. From Queries to Criteria: Understanding How Astronomers Evaluate LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A user study of an astronomy RAG bot identifies the question types and evaluation criteria astronomers actually use, and turns them into a 40-item benchmark.

  3. Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures

    cs.AI 2025-06 conditional novelty 7.0 of 10

    Xent Games formalize a large family of LLM evaluation tasks as games whose rewards and constraints are signed cross-entropy sums, and propose using them to build general capability measures.

  4. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  5. xpSHACL: Explainable SHACL Validation using Retrieval-Augmented Generation and Large Language Models

    cs.DB 2025-07 conditional novelty 6.0 of 10

    xpSHACL combines a rule-based trace of why a SHACL constraint failed with RAG and an LLM to generate human-readable, cached explanations for RDF validation violations.

  6. INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

    cs.CV 2025-07 conditional novelty 6.0 of 10

    INTER is a training-free logit-correction method that adds Harsanyi interaction scores to selected keyword tokens, lowering hallucination on six LVLM benchmarks.

  7. When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A systematic evaluation shows that masking the interlocutor's persona lowers target speaker identification accuracy, and that zero-shot models often copy biography details, making identification easier but dialogues m...

  8. The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.

  9. DecoRTL: A Run-time Decoding Framework for RTL Code Generation with LLMs

    cs.PL 2025-07 conditional novelty 5.0 of 10

    DecoRTL combines token-class-aware temperature adjustment with contrastive top-K reranking to improve synthesizability and functional correctness of LLM-generated Verilog.

  10. Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning

    cs.CV 2025-06 reject novelty 5.0 of 10

    A two-stage, value-guided decoding strategy with a margin-based reward adjustment is claimed to yield more faithful, detailed VLM captions at about a quarter of VisVM's inference cost.

  11. Position: Intelligent Coding Systems Should Write Programs with Justifications

    cs.SE 2025-08 conditional novelty 4.0 of 10

    A position paper advocating that intelligent coding systems should accompany code with justified explanations that are cognitively aligned and semantically faithful.

  12. Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A model is 'relatively biased' when its responses deviate from the consensus of a baseline LLM set, and this deviation can be scored by embedding distances or LLM judges plus equivalence tests.

Pith tools