A new evaluation framework finds that even the strongest LLMs still hallucinate a large share of references when asked to write literature reviews, with accuracy varying across disciplines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
A new evaluation framework finds that even the strongest LLMs still hallucinate a large share of references when asked to write literature reviews, with accuracy varying across disciplines.