REVIEW 4 cited by
Cutting Down on Prompts and Parameters: Simple Few-Shot Learning with Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Prompting language models (LMs) with training examples and task descriptions has been seen as critical to recent successes in few-shot learning. In this work, we show that finetuning LMs in the few-shot setting can considerably reduce the need for prompt engineering. In fact, one can use null prompts, prompts that contain neither task-specific templates nor training examples, and achieve competitive accuracy to manually-tuned prompts across a wide range of tasks. While finetuning LMs does introduce new parameters for each downstream task, we show that this memory overhead can be substantially reduced: finetuning only the bias terms can achieve comparable or better accuracy than standard finetuning while only updating 0.1% of the parameters. All in all, we recommend finetuning LMs for few-shot learning as it is more accurate, robust to different prompts, and can be made nearly as efficient as using frozen LMs.
Forward citations
Cited by 4 Pith papers
-
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
SelfPrompt makes an LLM generate adversarial prompts from domain-specific knowledge graph triples and then use them to compute its own robustness score.
-
QuaLLM-Health: An Adaptation of an LLM-Based Framework for Quantitative Data Extraction from Online Health Discussions
QuaLLM-Health claims GPT-4o-mini can extract clinical variables from GLP-1 Reddit discussions with macro F1 above 0.90, but the evaluation uses the same gold standard for prompt tuning and testing.
-
Trusting CHATGPT: how minor tweaks in the prompts lead to major differences in sentiment classification
Minor prompt rewording produces statistically significant shifts in GPT-4o mini's Spanish sentiment labels, yet overall agreement between prompts stays between 92% and 98%.
-
When IoT Meet LLMs: Applications and Challenges
A survey of LLM-IoT integration plus an unvalidated conceptual system model for Tree of Thought based predictive maintenance in industrial IoT.
Discussion (0). Continue with ORCID to comment.