Pith. sign in

REVIEW 10 cited by

Can language models learn from explanations in context?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.02329 v4 pith:O2DSCINK submitted 2022-04-05 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords explanationsexamplesmodelsperformancetasksbenefitschallengingcontrols
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language Models (LMs) can perform new tasks by adapting to a few in-context examples. For humans, explanations that connect examples to task principles can improve learning. We therefore investigate whether explanations of few-shot examples can help LMs. We annotate questions from 40 challenging tasks with answer explanations, and various matched control explanations. We evaluate how different types of explanations, instructions, and controls affect zero- and few-shot performance. We analyze these results using statistical multilevel modeling techniques that account for the nested dependencies among conditions, tasks, prompts, and models. We find that explanations can improve performance -- even without tuning. Furthermore, explanations hand-tuned for performance on a small validation set offer substantially larger benefits, and building a prompt by selecting examples and explanations together substantially improves performance over selecting examples alone. Finally, even untuned explanations outperform carefully matched controls, suggesting that the benefits are due to the link between an example and its explanation, rather than lower-level features. However, only large models benefit. In summary, explanations can support the in-context learning of large LMs on challenging tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search

    cs.LG 2025-05 conditional novelty 7.0 of 10

    ReGUIDE reaches state-of-the-art GUI grounding accuracy using 0.2% of the usual training data by adding self-generated reasoning, spatial-consistency training, and test-time KDE coordinate search.

  2. Test-Time Scaling via Error Localization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    TTEL uses feedback-induced token probability drops to localize the first error in a failed reasoning trace and branch a new generation from that prefix, improving pass@k per token on coding and math benchmarks.

  3. Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

  4. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  5. Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new patent novelty benchmark from real examiner rejections shows large language models can classify novelty at about 62% accuracy, while smaller classification models perform at chance.

  6. Data-Efficient Adaptation of LLMs via Attention Head Reweighting

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Learning a single scalar per attention head lets LLMs adapt to few-shot text classification better than LoRA, with 200–1000x fewer trainable parameters.

  7. Reasoning or Overthinking: Evaluating Large Language Models on Financial Sentiment Analysis

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On financial sentiment classification, zero-shot LLMs match human labels better without chain-of-thought reasoning than with it.

  8. TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Keyword distillation of clinical notes improved BERT predictions and explanation quality in small tests, but effects are modest and some important details are unverified.

  9. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

    cs.RO 2025-07 reject novelty 4.0 of 10

    An LLM plus LTL-based verification module that reorders, inserts, and removes steps in household robot plans, reporting reduced ordering errors but with weak experimental support.

  10. Towards Transparent AI: A Survey on Explainable Large Language Models

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A review that groups LLM explainability methods by transformer architecture and discusses their evaluation and applications.

Pith tools