Pith. sign in

REVIEW 8 cited by

Are Your LLMs Capable of Stable Reasoning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.13147 v5 pith:ZQAKAFPH submitted 2024-12-17 cs.AI cs.CL

classification cs.AIcs.CL
keywords reasoningevaluationllmscapabilitiescomplexconsistencyg-passlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The rapid advancement of large language models (LLMs) has shown remarkable progress in complex reasoning tasks. However, a significant disparity exists between benchmark performances and real-world applications. We attribute this gap primarily to current evaluation protocols and metrics, which inadequately capture the full spectrum of LLM capabilities, especially in complex reasoning tasks where both accuracy and consistency are essential. In this paper, we introduce G-Pass@$k$, a novel evaluation metric that continuously assesses model performance across multiple sampling attempts, quantifying both the model's performance potential and its stability. Through extensive experiments on various public and newly constructed benchmarks, we employ G-Pass@$k$ in conjunction with state-of-the-art large language models to provide comprehensive insights into their potential capabilities and operational consistency. Our findings reveal a significant opportunity to enhance the realistic reasoning abilities of LLMs, underscoring the necessity for more robust evaluation metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Spurious RLVR makes Qwen2.5-Math retrieve memorized answers via a layer 18-20 anchor and layer 21+ adapters, a shortcut that can be steered by scaling specific MLP keys.

  2. The Incomplete Bridge: How AI Research (Mis)Engages with Psychology

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A citation-based study of 1,006 LLM papers finds psychology is increasingly cited, concentrated in psychometrics and neural mechanisms, and identifies repeated misapplications of Theory of Mind.

  3. Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Relay-Bench, a 30-problem benchmark of chained multi-domain tasks with encoded prompts, resists saturation: the best tested model, GPT-5.5, scores 43.3% Pass@1.

  4. Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A 2.75B-active-parameter MoE model, trained with a token-budget-stabilized RL method (C3PO), matches or exceeds several 7-8B dense reasoning models on AIME, LiveCodeBench, and GPQA benchmarks.

  5. Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Reasoning trajectories are formalized as pseudo-gradient descent on LLM parameters, making LLM reasoning training a MAML-style meta-learning problem.

  6. The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants

    cs.CL 2025-05 reject novelty 5.0 of 10

    Clustering-based routing plus self-consistency voting among ten 7B open models reportedly outranks GPT-4.1 and GPT-4.5 on average over 15 diverse benchmarks.

  7. Swarm Intelligence Enhanced Reasoning: A Density-Driven Framework for LLM-Based Multi-Agent Optimization

    cs.MA 2025-05 conditional novelty 5.0 of 10

    SIER uses kernel density estimation and Pareto-style selection to preserve diversity in PRM-guided multi-agent reasoning, improving pass@8 and prm@8 on math benchmarks at higher token cost.

  8. Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

    cs.CL 2025-02 reject novelty 5.0 of 10

    OREAL shows that outcome-reward RL with best-of-N positive behavior cloning, negative reward shaping, and token-level reweighting reaches state-of-the-art MATH-500 accuracy at 7B and 32B scale.

Pith tools