Pith. sign in

REVIEW 2 cited by

Enhancing Distractor Generation for Multiple-Choice Questions with Retrieval Augmented Pretraining and Knowledge Graph Integration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13578 v1 pith:WVZTQLGP submitted 2024-06-19 cs.CL

classification cs.CL
keywords pretrainingaugmenteddatasetdistractorgenerationintegrationknowledgemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we tackle the task of distractor generation (DG) for multiple-choice questions. Our study introduces two key designs. First, we propose \textit{retrieval augmented pretraining}, which involves refining the language model pretraining to align it more closely with the downstream task of DG. Second, we explore the integration of knowledge graphs to enhance the performance of DG. Through experiments with benchmarking datasets, we show that our models significantly outperform the state-of-the-art results. Our best-performing model advances the F1@3 score from 14.80 to 16.47 in MCQ dataset and from 15.92 to 16.50 in Sciq dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    AutoConverter converts open-ended VQA questions into multiple-choice format via multi-agent GPT-4o, and VMCBench applies it to 20 datasets to evaluate 33 vision-language models.

  2. LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Using Llama-3.1-8B to generate and score synthetic MCQA data, then distilling those soft labels into DeBERTa-v3-base, improves few-shot MMLU accuracy from 28.9% to 39.3%.

Pith tools