On a 375-sample HaluBench test set, MIPROv2 and Bootstrap Few Shot with Random Search achieve the highest weighted and macro F1 scores for LLM-based hallucination detection, though without statistical significance tests.
How to choose a t hreshold for an evaluation metric for large language models, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
other 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
On a 375-sample HaluBench test set, MIPROv2 and Bootstrap Few Shot with Random Search achieve the highest weighted and macro F1 scores for LLM-based hallucination detection, though without statistical significance tests.