Pith. sign in

REVIEW 3 cited by

Rationale-Augmented Ensembles in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.00747 v1 pith:OCVMF4C6 submitted 2022-07-02 cs.CL

classification cs.CL
keywords rationale-augmentedpromptingrationalesensemblesoutputperformancedemonstrateexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent research has shown that rationales, or step-by-step chains of thought, can be used to improve performance in multi-step reasoning tasks. We reconsider rationale-augmented prompting for few-shot in-context learning, where (input -> output) prompts are expanded to (input, rationale -> output) prompts. For rationale-augmented prompting we demonstrate how existing approaches, which rely on manual prompt engineering, are subject to sub-optimal rationales that may harm performance. To mitigate this brittleness, we propose a unified framework of rationale-augmented ensembles, where we identify rationale sampling in the output space as the key component to robustly improve performance. This framework is general and can easily be extended to common natural language processing tasks, even those that do not traditionally leverage intermediate steps, such as question answering, word sense disambiguation, and sentiment analysis. We demonstrate that rationale-augmented ensembles achieve more accurate and interpretable results than existing prompting approaches--including standard prompting without rationales and rationale-based chain-of-thought prompting--while simultaneously improving interpretability of model predictions through the associated rationales.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 29 citations worldwide. Full citation record

  1. What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale

    cs.LG 2025-09 conditional novelty 6.0 of 10

    An instruction-tuned LLaMA 3.1 8B model, trained on LLM-generated pseudo-labels, is claimed to identify IoT device vendors from passive network metadata with 98.25% top-1 accuracy across 2,015 vendors.

  2. VulRTex: A Reasoning-Guided Approach to Identify Vulnerabilities from Rich-Text Issue Report

    cs.SE 2025-09 conditional novelty 6.0 of 10

    A retrieval-augmented LLM approach that identifies vulnerability-related issue reports and CWE types from screenshots and code snippets, improving F1 by 11 points and AUPRC by 20 points over baselines.

  3. Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning

    cs.CL 2025-02 conditional novelty 6.0 of 10

    TICL improves style personalization by iteratively adding model-generated negative examples and explanations to an in-context prompt, beating fine-tuned baselines in LLM-judged comparisons without any parameter updates.

Pith tools