Pith. sign in

REVIEW 3 cited by

Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15721 v2 pith:MHW33XBF submitted 2024-02-24 cs.AI cs.CL

classification cs.AIcs.CL
keywords hallucinationsevaluationhallucinationlvlmsdataeventframeworklanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision Language Models exhibit remarkable capabilities but struggle with hallucinations inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs efficacy in handling hallucinations. We will release our code and data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free bottom-up reasoning framework that verifies scene graphs and commonsense with external tools reduces hallucinations in multimodal LLMs.

  2. Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Decomposing MLLM responses into atomic verification tasks and checking them with an ensemble of open-source expert models yields preference data that reduces hallucination in LLaVA and Qwen-VL-Chat.

  3. RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance

    cs.LG 2025-01 reject novelty 5.0 of 10

    A framework using fine-tuned LLaVA and VILA models to score relevance and correctness in multimodal retrieval-augmented generation, with reported 88 percent test accuracy and 91 percent human agreement.

Pith tools