Pith. sign in

REVIEW 2 cited by

FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.14715 v2 pith:YXTLSHPZ submitted 2024-04-23 cs.CV cs.CL

classification cs.CVcs.CL
keywords imagetextmodelsfine-grainedfinematchvlmsaspect-basedcorrection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimodal content. Despite the impressive ability to perform complex reasoning for VLMs, current models often struggle to effectively and precisely capture the compositional information on both the image and text sides. To address this, we propose FineMatch, a new aspect-based fine-grained text and image matching benchmark, focusing on text and image mismatch detection and correction. This benchmark introduces a novel task for boosting and evaluating the VLMs' compositionality for aspect-based fine-grained text and image matching. In this task, models are required to identify mismatched aspect phrases within a caption, determine the aspect's class, and propose corrections for an image-text pair that may contain between 0 and 3 mismatches. To evaluate the models' performance on this new task, we propose a new evaluation metric named ITM-IoU for which our experiments show a high correlation to human evaluation. In addition, we also provide a comprehensive experimental analysis of existing mainstream VLMs, including fully supervised learning and in-context learning settings. We have found that models trained on FineMatch demonstrate enhanced proficiency in detecting fine-grained text and image mismatches. Moreover, models (e.g., GPT-4V, Gemini Pro Vision) with strong abilities to perform multimodal in-context learning are not as skilled at fine-grained compositional image and text matching analysis. With FineMatch, we are able to build a system for text-to-image generation hallucination detection and correction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Refine-by-Align uses diffusion cross-attention maps to locate the reference region matching a masked artifact, then re-inpaints the artifact with that reference detail.

  2. FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FINECAPTION captions any masked image region across 18 compositional attributes by fusing a mask-aware encoder with high-resolution encoders, and the paper introduces the COMPOSITIONCAP dataset to train and test it.

Pith tools