Pith. sign in

REVIEW 4 cited by

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04594 v3 pith:YTEP53SV submitted 2024-08-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords datadatasetdifferenceimagemllmsmodelsobjectsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

High-performance Multimodal Large Language Models (MLLMs) are heavily dependent on data quality. To advance fine-grained image recognition within MLLMs, we introduce a novel data synthesis method inspired by contrastive learning and image difference captioning. Our key idea involves challenging the model to discern both matching and distinct elements by scrutinizing object differences in detailed regions across similar images. We begin by generating pairs of similar images that emphasize object variations. Following this, we employ a Difference Area Generator to pinpoint object differences, and subsequently, a Difference Captions Generator to articulate these differences. This process results in a high-quality dataset of "object replacement" samples, termed Img-Diff, which can be scaled as needed due to its automated nature. We leverage this generated dataset to fine-tune state-of-the-art (SOTA) MLLMs, such as InternVL2, achieving substantial improvements across various image difference and Visual Question Answering tasks. Notably, the trained models significantly outperform existing SOTA models like GPT-4V and Gemini on the MMVP benchmark. Additionally, we conduct comprehensive evaluations to validate the dataset's diversity, quality, and robustness, offering several insights into the synthesis of such contrastive datasets. We release our codes and dataset to encourage further research on multimodal data synthesis and MLLMs' fundamental capabilities for image understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    DeGF reduces hallucinations in vision-language models by generating an image from the model's own response and using the divergence between predictions on original and generated images to switch between complementary ...

  2. Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Migician is an instruction-tuned MLLM that performs free-form grounding across multiple images, with a new 630k dataset and a 10-task benchmark, but the evaluation is weakened by source overlap between training and be...

  3. Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A contrastive matching loss plus homography-based alignment and Hungarian matching improves zero-shot change detection and predicts correspondences between detected changes.

  4. TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

    cs.CV 2024-12

Pith tools