Pith. sign in

REVIEW 1 cited by

CLIP4IDC: CLIP for Image Difference Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.00629 v2 pith:ZCGUROLF submitted 2022-06-01 cs.CV

classification cs.CV
keywords clipvisualclip4idcimageimagescaptioningdatasetsdifference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image Difference Captioning (IDC) aims at generating sentences to describe differences between two similar-looking images. Conventional approaches learn an IDC model with a pre-trained and usually frozen visual feature extractor. Accordingly, two major issues may arise: (1) a large domain gap usually exists between the pre-training datasets used for training such a visual encoder and that of the downstream IDC task, and (2) the visual feature extractor, when separately encoding two images, often does not effectively encode the visual changes between two images. Due to the excellent zero-shot performance of the recently proposed CLIP, we thus propose CLIP4IDC to transfer a CLIP model for the IDC task to address those issues. Different from directly fine-tuning CLIP to generate sentences, we introduce an adaptation training process to adapt CLIP's visual encoder to capture and align differences in image pairs based on the textual descriptions. Experiments on three IDC benchmark datasets, CLEVR-Change, Spot-the-Diff, and Image-Editing-Request, demonstrate the effectiveness of CLIP4IDC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A contrastive matching loss plus homography-based alignment and Hungarian matching improves zero-shot change detection and predicts correspondences between detected changes.

Pith tools