REVIEW 3 cited by
When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In-context learning (ICL) has become the default method for using large language models (LLMs), making the exploration of its limitations and understanding the underlying causes crucial. In this paper, we find that ICL falls short of handling specification-heavy tasks, which are tasks with complicated and extensive task specifications, requiring several hours for ordinary humans to master, such as traditional information extraction tasks. The performance of ICL on these tasks mostly cannot reach half of the state-of-the-art results. To explore the reasons behind this failure, we conduct comprehensive experiments on 18 specification-heavy tasks with various LLMs and identify three primary reasons: inability to specifically understand context, misalignment in task schema comprehension with humans, and inadequate long-text understanding ability. Furthermore, we demonstrate that through fine-tuning, LLMs can achieve decent performance on these tasks, indicating that the failure of ICL is not an inherent flaw of LLMs, but rather a drawback of existing alignment methods that renders LLMs incapable of handling complicated specification-heavy tasks via ICL. To substantiate this, we perform dedicated instruction tuning on LLMs for these tasks and observe a notable improvement. We hope the analyses in this paper could facilitate advancements in alignment methods enabling LLMs to meet more sophisticated human demands.
Forward citations
Cited by 3 Pith papers
-
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
AgentIF introduces a realistic, long-form instruction-following benchmark for agentic scenarios and shows that current LLMs follow fewer than 30% of such instructions perfectly.
-
Detecting Conversational Mental Manipulation with Intent-Aware Prompting
Adding per-speaker intent summaries to an LLM prompt reduces false negatives in mental manipulation detection by 30.5% versus zero-shot prompting on the MentalManip dataset.
-
Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation
A template-plus-LLM pipeline generates 23,000 tabular math word problems with illustrative solutions, and fine-tuning on them raises the accuracy of 7B-8B LLMs on TabMWP by roughly four percentage points.
Discussion (0). Continue with ORCID to comment.