Pith. sign in

REVIEW 3 cited by

When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08993 v1 pith:EWI4AYZH submitted 2023-11-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords tasksllmsspecification-heavyalignmentcomplicatedfailurehandlinghumans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In-context learning (ICL) has become the default method for using large language models (LLMs), making the exploration of its limitations and understanding the underlying causes crucial. In this paper, we find that ICL falls short of handling specification-heavy tasks, which are tasks with complicated and extensive task specifications, requiring several hours for ordinary humans to master, such as traditional information extraction tasks. The performance of ICL on these tasks mostly cannot reach half of the state-of-the-art results. To explore the reasons behind this failure, we conduct comprehensive experiments on 18 specification-heavy tasks with various LLMs and identify three primary reasons: inability to specifically understand context, misalignment in task schema comprehension with humans, and inadequate long-text understanding ability. Furthermore, we demonstrate that through fine-tuning, LLMs can achieve decent performance on these tasks, indicating that the failure of ICL is not an inherent flaw of LLMs, but rather a drawback of existing alignment methods that renders LLMs incapable of handling complicated specification-heavy tasks via ICL. To substantiate this, we perform dedicated instruction tuning on LLMs for these tasks and observe a notable improvement. We hope the analyses in this paper could facilitate advancements in alignment methods enabling LLMs to meet more sophisticated human demands.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

    cs.AI 2025-05 conditional novelty 7.0 of 10

    AgentIF introduces a realistic, long-form instruction-following benchmark for agentic scenarios and shows that current LLMs follow fewer than 30% of such instructions perfectly.

  2. Detecting Conversational Mental Manipulation with Intent-Aware Prompting

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Adding per-speaker intent summaries to an LLM prompt reduces false negatives in mental manipulation detection by 30.5% versus zero-shot prompting on the MentalManip dataset.

  3. Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A template-plus-LLM pipeline generates 23,000 tabular math word problems with illustrative solutions, and fine-tuning on them raises the accuracy of 7B-8B LLMs on TabMWP by roughly four percentage points.

Pith tools