REVIEW 6 cited by
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The advancement of Large Vision-Language Models (LVLMs) has increasingly highlighted the critical issue of their tendency to hallucinate non-existing objects in the images. To address this issue, previous works focused on using specially curated datasets or powerful LLMs to rectify the outputs of LVLMs. However, these approaches require either costly training or fine-tuning, or API access to proprietary LLMs for post-generation correction. In response to these limitations, we propose Mitigating hallucinAtion via image-gRounded guIdaNcE (MARINE), a framework that is both training-free and API-free. MARINE effectively and efficiently reduces object hallucinations during inference by introducing image-grounded guidance to LVLMs. This is achieved by leveraging open-source vision models to extract object-level information, thereby enhancing the precision of LVLM-generated content. Our framework's flexibility further allows for the integration of multiple vision models, enabling more reliable and robust object-level guidance. Through comprehensive evaluations across 5 popular LVLMs with diverse evaluation metrics and benchmarks, we demonstrate the effectiveness of MARINE, which even outperforms existing fine-tuning-based methods. Remarkably, it reduces hallucinations consistently in GPT-4V-assisted evaluation while maintaining the detailedness of LVLMs' generations. We release our code at https://github.com/Linxi-ZHAO/MARINE.
Forward citations
Cited by 6 Pith papers
-
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
EraseLoRA removes masked objects by having an MLLM separate target, non-target foreground, and background, then test-time LoRA optimization aggregates background subtypes to reconstruct the occluded region.
-
Controlling Multimodal LLMs via Reward-guided Decoding
MRGD guides MLLM decoding with a learned hallucination reward and a detector-based recall reward, allowing users to trade off object precision, recall, and test-time compute while reducing object hallucinations on CHA...
-
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
CAI reduces object hallucination in LVLMs by injecting caption-query attention patterns into selected attention heads at inference time.
-
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
MMGrounded-PostAlign trains MLLMs to produce a grounded object token or a rejection token plus selective rationales, improving hallucination and VQA benchmarks.
-
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
A study of InstructBLIP and mPLUG-Owl2 finds that scene words like grass and tree co-occur with hallucinated objects, and a two-step foreground/background prompt lowers hallucination scores.
-
Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild
A self-questioning training and inference framework for lightweight multimodal LLMs is claimed to reduce hallucinations and improve zero-shot visual reasoning.
Discussion (0). Continue with ORCID to comment.