Pith. sign in

REVIEW 9 cited by

ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.07919 v2 pith:IXC5F7PD submitted 2022-12-15 cs.CL cs.LG

classification cs.CLcs.LG
keywords reasoningmetricsroscoeevaluationfinalstep-by-stepautomaticbaseline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models show improved downstream task performance when prompted to generate step-by-step reasoning to justify their final answers. These reasoning steps greatly improve model interpretability and verification, but objectively studying their correctness (independent of the final answer) is difficult without reliable methods for automatic evaluation. We simply do not know how often the stated reasoning steps actually support the final end task predictions. In this work, we present ROSCOE, a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics. To evaluate ROSCOE against baseline metrics, we design a typology of reasoning errors and collect synthetic and human evaluation scores on commonly used reasoning datasets. In contrast with existing metrics, ROSCOE can measure semantic consistency, logicality, informativeness, fluency, and factuality - among other traits - by leveraging properties of step-by-step rationales. We empirically verify the strength of our metrics on five human annotated and six programmatically perturbed diagnostics datasets - covering a diverse set of tasks that require reasoning skills and show that ROSCOE can consistently outperform baseline metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 28 citations worldwide. Full citation record

  1. More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A three-agent homogeneous debate panel does not reliably improve groundedness verification: gains and losses are task-dependent, and its main mechanism is threshold recalibration rather than evidence acquisition.

  2. Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

    cs.CL 2025-12 conditional novelty 6.0 of 10

    Monitoring the semantic type of reasoning steps with a lightweight BERT classifier can drive interpretable early stopping that cuts generation tokens by 20-50% at modest accuracy cost.

  3. Rethinking Human Preference Evaluation of LLM Rationales

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A fine-grained attribute-based evaluation of LLM rationales can explain human preferences and reveal model trade-offs that binary comparisons obscure.

  4. VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 7-billion-parameter multimodal model fine-tuned on 2,500 expert critiques of data visualizations matches or beats much larger models at identifying visualization defects.

  5. Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLM reasoning can be scored separately for knowledge and step-by-step information gain, and doing so shows SFT and RL affect these two capacities differently across medicine and math.

  6. ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ARB provides 1,356 Arabic multimodal questions with 5,119 human-reviewed reasoning steps and shows leading models score much higher on reasoning fluency than on correct answers.

  7. Diagnosing Pathological Chain-of-Thought in Reasoning Models

    cs.AI 2026-02 conditional novelty 5.0 of 10

    Three log-probability-difference metrics — Necessity, Paraphrasability, Substantivity — are proposed and tested on deliberately fine-tuned 'model organisms' to detect post-hoc, encoded, and internalized chain-of-thoug...

  8. Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation

    cs.CL 2025-06 reject novelty 4.0 of 10

    Prompt rewrites of LeetCode problems cause large accuracy swings in nine LLMs, but invalid negation test cases and inconsistent tables make the headline numbers unreliable.

  9. Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.

Pith tools