REVIEW 5 cited by
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Prevailing methods for mapping large generative language models to supervised tasks may fail to sufficiently probe models' novel capabilities. Using GPT-3 as a case study, we show that 0-shot prompts can significantly outperform few-shot prompts. We suggest that the function of few-shot examples in these cases is better described as locating an already learned task rather than meta-learning. This analysis motivates rethinking the role of prompts in controlling and evaluating powerful language models. In this work, we discuss methods of prompt programming, emphasizing the usefulness of considering prompts through the lens of natural language. We explore techniques for exploiting the capacity of narratives and cultural anchors to encode nuanced intentions and techniques for encouraging deconstruction of a problem into components before producing a verdict. Informed by this more encompassing theory of prompt programming, we also introduce the idea of a metaprompt that seeds the model to generate its own natural language prompts for a range of tasks. Finally, we discuss how these more general methods of interacting with language models can be incorporated into existing and future benchmarks and practical applications.
Forward citations
Cited by 5 Pith papers
-
Agentic Synthesis against Counterexample-Supplemented Sketches
This paper proposes counterexample-supplemented sketches: a repository workflow where human operators approve policy changes and a dual gate (deterministic replay plus sketch review) controls agent-driven code synthesis.
-
Empirical Prompt Engineering for Construct Identification with Large Language Models
For LLM classification of psychological constructs, selecting the best prompt from many variants improves human-model agreement more than personas, chain-of-thought, or explanations.
-
Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
Synthetic customer agents built from real bank data can mimic customer semantics and personality well enough to serve as scalable chatbot validation proxies.
-
Large Language Models as Autonomous Spacecraft Operators in Kerbal Space Program
LLM agents using prompt engineering and fine-tuning ranked second in the Kerbal Space Program Differential Games pursuit-evasion challenge.
-
A Mathematical Theory of Discursive Networks
A two-state Markov model of error propagation suggests that small amounts of cross-agent peer review can flip a network of fallible language models from a falsehood-dominant to a truth-dominant state.
Discussion (0). Continue with ORCID to comment.