Pith. sign in

REVIEW 3 cited by

LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09962 v2 pith:ACU4XPXZ submitted 2024-10-13 cs.CV

classification cs.CV
keywords hallucinationdatalonghalqamllmscomplexdiscriminativeevaluatorslong
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hallucination, a phenomenon where multimodal large language models~(MLLMs) tend to generate textual responses that are plausible but unaligned with the image, has become one major hurdle in various MLLM-related applications. Several benchmarks have been created to gauge the hallucination levels of MLLMs, by either raising discriminative questions about the existence of objects or introducing LLM evaluators to score the generated text from MLLMs. However, the discriminative data largely involve simple questions that are not aligned with real-world text, while the generative data involve LLM evaluators that are computationally intensive and unstable due to their inherent randomness. We propose LongHalQA, an LLM-free hallucination benchmark that comprises 6K long and complex hallucination text. LongHalQA is featured by GPT4V-generated hallucinatory data that are well aligned with real-world scenarios, including object/image descriptions and multi-round conversations with 14/130 words and 189 words, respectively, on average. It introduces two new tasks, hallucination discrimination and hallucination completion, unifying both discriminative and generative evaluations in a single multiple-choice-question form and leading to more reliable and efficient evaluations without the need for LLM evaluators. Further, we propose an advanced pipeline that greatly facilitates the construction of future hallucination benchmarks with long and complex questions and descriptions. Extensive experiments over multiple recent MLLMs reveal various new challenges when they are handling hallucinations with long and complex textual data. Dataset and evaluation code are available at https://github.com/hanqiu-hq/LongHalQA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring Epistemic Humility in Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new 22,831-question visual benchmark shows that major multimodal LLMs struggle to reject false answer options, often scoring near random when abstaining is the only correct response.

  2. Machine Mirages: Defining the Undefined

    cs.AI 2025-06 reject novelty 4.0 of 10

    A conceptual paper defines 34 AI failure modes with formula-like conditions, proposes expectile-value-at-risk quantification, and proves mostly tautological or trivial existence and impossibility results without experiments.

  3. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools