REVIEW 1 cited by
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation. We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark for QA (BBQ), designed to evaluate social bias in long-form generation by having LLMs generate continuations of story prompts. Building our benchmark in English and Korean, we measure the probability of neutral and biased generations across ten LLMs. We also compare our long-form story generation evaluation results with multiple-choice BBQ evaluation, showing that the two approaches produce inconsistent results.
Forward citations
Cited by 1 Pith paper
-
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
EsBBQ and CaBBQ are new Spanish and Catalan bias benchmarks for multiple-choice QA, built with survey-validated stereotypes from Spain and evaluated on 17 language models.
Discussion (0). Continue with ORCID to comment.