Pith. sign in

REVIEW 4 cited by

Fairness-guided Few-shot Prompting for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13217 v3 pith:ZTSQACTX submitted 2023-03-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords in-contextlearningpromptmodelsbiasperformancepredictivedownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models have demonstrated surprising ability to perform in-context learning, i.e., these models can be directly applied to solve numerous downstream tasks by conditioning on a prompt constructed by a few input-output examples. However, prior research has shown that in-context learning can suffer from high instability due to variations in training examples, example order, and prompt formats. Therefore, the construction of an appropriate prompt is essential for improving the performance of in-context learning. In this paper, we revisit this problem from the view of predictive bias. Specifically, we introduce a metric to evaluate the predictive bias of a fixed prompt against labels or a given attributes. Then we empirically show that prompts with higher bias always lead to unsatisfactory predictive quality. Based on this observation, we propose a novel search strategy based on the greedy search to identify the near-optimal prompt for improving the performance of in-context learning. We perform comprehensive experiments with state-of-the-art mainstream models such as GPT-3 on various downstream tasks. Our results indicate that our method can enhance the model's in-context learning performance in an effective and interpretable manner.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. Addressing speaker gender bias in large scale speech translation systems

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Fine-tuning a large speech translation model on GPT-4-reformulated gender-balanced training data raises MuST-SHE feminine-form accuracy from about 10% to over 84% without BLEU loss.

  2. Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Output token count, observable through response timing, can reveal a user's target language or classification result with 70-87% accuracy in the authors' experiments.

  3. Towards a Unified Paradigm: Integrating Recommendation Systems as a New Language in Large Models

    cs.IR 2024-12 conditional novelty 5.0 of 10

    RSLLM mixes item ID embeddings from classical recommenders with text titles inside an LLM prompt and uses two-stage contrastive fine-tuning to improve sequential recommendation.

  4. Bias Mitigation Agent: Optimizing Source Selection for Fair and Balanced Knowledge Retrieval

    cs.AI 2025-08 reject novelty 4.0 of 10

    A multi-agent retrieval system that filters sources by a bias classifier reports an 81.82% relative drop in bias rate, but the evaluation uses the same classifier as the filter.

Pith tools