Pith. sign in

REVIEW 2 cited by

TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04107 v1 pith:ZE6GDHAE submitted 2024-10-05 cs.CV cs.CL

classification cs.CVcs.CL
keywords questionslvlmstubenchunanswerablereasoningvisualevaluateimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision-Language Models (LVLMs) have achieved remarkable progress on visual perception and linguistic interpretation. Despite their impressive capabilities across various tasks, LVLMs still suffer from the issue of hallucination, which involves generating content that is incorrect or unfaithful to the visual or textual inputs. Traditional benchmarks, such as MME and POPE, evaluate hallucination in LVLMs within the scope of Visual Question Answering (VQA) using answerable questions. However, some questions are unanswerable due to insufficient information in the images, and the performance of LVLMs on such unanswerable questions remains underexplored. To bridge this research gap, we propose TUBench, a benchmark specifically designed to evaluate the reliability of LVLMs using unanswerable questions. TUBench comprises an extensive collection of high-quality, unanswerable questions that are meticulously crafted using ten distinct strategies. To thoroughly evaluate LVLMs, the unanswerable questions in TUBench are based on images from four diverse domains as visual contexts: screenshots of code snippets, natural images, geometry diagrams, and screenshots of statistical tables. These unanswerable questions are tailored to test LVLMs' trustworthiness in code reasoning, commonsense reasoning, geometric reasoning, and mathematical reasoning related to tables, respectively. We conducted a comprehensive quantitative evaluation of 28 leading foundational models on TUBench, with Gemini-1.5-Pro, the top-performing model, achieving an average accuracy of 69.2%, and GPT-4o, the third-ranked model, reaching 66.7% average accuracy, in determining whether questions are answerable. TUBench is available at https://github.com/NLPCode/TUBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Look Again Before You Abstain:Budgeted Conformal Evidence Acquisition for Reliable Vision-Language Model

    cs.CV 2026-06 conditional novelty 6.0 of 10

    Folding visual evidence acquisition into the conformal score and re-calibrating on post-acquisition scores preserves the hallucination-rate guarantee while recovering coverage in LVLM selective prediction.

  2. Explainability for Vision Foundation Models: A Survey

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).

Pith tools