Pith. sign in

REVIEW 4 cited by

Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.17306 v1 pith:6GMXH2LG submitted 2023-05-26 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelsreasoninglanguagelargellmscapabilitieschain-of-thoughtchallenging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As large language models (LLMs) are continuously being developed, their evaluation becomes increasingly important yet challenging. This work proposes Chain-of-Thought Hub, an open-source evaluation suite on the multi-step reasoning capabilities of large language models. We are interested in this setting for two reasons: (1) from the behavior of GPT and PaLM model family, we observe that complex reasoning is likely to be a key differentiator between weaker and stronger LLMs; (2) we envisage large language models to become the next-generation computational platform and foster an ecosystem of LLM-based new applications, this naturally requires the foundation models to perform complex tasks that often involve the composition of linguistic and logical operations. Our approach is to compile a suite of challenging reasoning benchmarks to track the progress of LLMs. Our current results show that: (1) model scale clearly correlates with reasoning capabilities; (2) As of May 2023, Claude-v1.3 and PaLM-2 are the only two models that are comparable with GPT-4, while open-sourced models still lag behind; (3) LLaMA-65B performs closely to code-davinci-002, indicating that with successful further development such as reinforcement learning from human feedback (RLHF), it has great potential to be close to GPT-3.5-Turbo. Our results also suggest that for the open-source efforts to catch up, the community may focus more on building better base models and exploring RLHF.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing

    cs.LG 2025-05 conditional novelty 7.0 of 10

    SharePrefill accelerates long-context LLM prefilling by clustering similar attention heads offline and sharing exact block-sparse attention patterns among them during inference.

  2. Unveiling Confirmation Bias in Chain-of-Thought Reasoning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    LLMs exhibit confirmation bias in chain-of-thought: strong internal beliefs, approximated by direct answer probabilities, skew both reasoning generation and how the final answer is chosen, which helps explain why CoT ...

  3. AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Post-training an open Qwen3-30B agent on 52,361 verifiable tasks in 5,018 synthesized stateful environments lifts its average across four agent benchmarks from 22.9% to 41.7%.

  4. Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought

    cs.CL 2026-05 reject novelty 5.0 of 10

    MoE routers allocate expert diversity in proportion to operation rarity (the "Frequency-Diversity Law"), and subset-difference pruning can expose this pattern when load-balancing creates redundant experts.

Pith tools