Pith. sign in

REVIEW 8 cited by

Visual Hallucinations of Multi-modal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.14683 v2 pith:UBEL7N3Z submitted 2024-02-22 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords instancesexistingvhtestbenchmarkfindimagevisualbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual hallucination (VH) means that a multi-modal LLM (MLLM) imagines incorrect details about an image in visual question answering. Existing studies find VH instances only in existing image datasets, which results in biased understanding of MLLMs' performance under VH due to limited diversity of such VH instances. In this work, we propose a tool called VHTest to generate a diverse set of VH instances. Specifically, VHTest finds some initial VH instances in existing image datasets (e.g., COCO), generates a text description for each VH mode, and uses a text-to-image generative model (e.g., DALL-E-3) to generate VH images based on the text descriptions. We collect a benchmark dataset with 1,200 VH instances in 8 VH modes using VHTest. We find that existing MLLMs such as GPT-4V, LLaVA-1.5, and MiniGPT-v2 hallucinate for a large fraction of the instances in our benchmark. Moreover, we find that fine-tuning an MLLM using our benchmark dataset reduces its likelihood to hallucinate without sacrificing its performance on other benchmarks. Our benchmarks are publicly available: https://github.com/wenhuang2000/VHTest.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Fine-tuning LVLMs on original/altered image pairs with targeted visual instructions reduces object and attribute hallucinations on POPE, LLaVA-Bench, and MMHal-Bench.

  2. AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AutoV selects instance- and query-specific visual prompts via loss-based pairwise ranking, consistently improving LVLMs across many benchmarks with no backbone fine-tuning.

  3. ChartLens: Fine-grained Visual Attribution in Charts

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ChartLens uses segmentation and set-of-marks prompting to attribute chart-based answers to specific visual elements, and the authors release a new benchmark for evaluating such attribution.

  4. Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Uncertainty-o estimates uncertainty in large multimodal models by perturbing prompts and computing entropy over semantically clustered answers, improving hallucination detection across five modalities.

  5. Unified Multimodal Understanding via Byte-Pair Visual Encoding

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Priority-guided byte-pair encoding of quantized image patches plus curriculum training yields an 8B discrete-token MLLM competitive with continuous-embedding models on VQA and multimodal benchmarks.

  6. ReFrame: Rectification Framework for Image Explaining Architectures

    cs.CV 2025-06 reject novelty 4.0 of 10

    ReFrame wraps image captioning, VQA, and GPT-4 with a Mask R-CNN rectifier, reporting big gains on metrics defined against that same detector.

  7. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

  8. Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model

    cs.CV 2025-05 reject novelty 2.0 of 10

    The paper reports lower hallucination scores when selecting the best of three filtered image variants, but the selection uses the ground truth, so the improvement is an artifact of choosing the minimum.

Pith tools