REVIEW 6 cited by
Prompting GPT-3 To Be Reliable
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) show impressive abilities via few-shot prompting. Commercialized APIs such as OpenAI GPT-3 further increase their use in real-world language applications. However, the crucial problem of how to improve the reliability of GPT-3 is still under-explored. While reliability is a broad and vaguely defined term, we decompose reliability into four main facets that correspond to the existing framework of ML safety and are well-recognized to be important: generalizability, social biases, calibration, and factuality. Our core contribution is to establish simple and effective prompts that improve GPT-3's reliability as it: 1) generalizes out-of-distribution, 2) balances demographic distribution and uses natural language instructions to reduce social biases, 3) calibrates output probabilities, and 4) updates the LLM's factual knowledge and reasoning chains. With appropriate prompts, GPT-3 is more reliable than smaller-scale supervised models on all these facets. We release all processed datasets, evaluation scripts, and model predictions. Our systematic empirical study not only sheds new insights on the reliability of prompting LLMs, but more importantly, our prompting strategies can help practitioners more reliably use LLMs like GPT-3.
Forward citations
Cited by 6 Pith papers
-
Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
EAACD reduces hallucination in MoE LLMs by contrasting predictions of high-reliability expert groups against hallucination-amplified low-reliability expert groups.
-
Novobo: Supporting Teachers' Peer Learning of Instructional Gestures by Teaching a Mentee AI-Agent Together
Novobo, a teachable AI agent that acts as an apprentice teacher, helped 30 teachers in 10 sessions externalize and co-construct knowledge about instructional gestures through group discussion and embodied demonstration.
-
S2LPP: Small-to-Large Prompt Prediction across LLMs
Small and large LLMs often share the same optimal prompt, so a small model can select effective prompts for a larger model at lower cost.
-
Advertising in AI systems: Society must be vigilant
Generative AI outputs will likely carry embedded commercial content, and the paper proposes design principles, provenance tracking, and two debiasing strategies to preserve transparency.
-
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
Entity popularity and entity co-occurrence in Wikipedia correlate with LLM QA accuracy, confidence, and calibration, and combining them with confidence improves answer-correctness prediction by 5.24% on average.
-
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
A model is 'relatively biased' when its responses deviate from the consensus of a baseline LLM set, and this deviation can be scored by embedding distances or LLM judges plus equivalence tests.
Discussion (0). Continue with ORCID to comment.