Pith. sign in

REVIEW 2 cited by

UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.14019 v1 pith:HAQ6JFBU submitted 2021-06-26 cs.CL cs.CV

classification cs.CLcs.CV
keywords captionsumicimagemetriccaptioningdatasetannotationsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the success of various text generation metrics such as BERTScore, it is still difficult to evaluate the image captions without enough reference captions due to the diversity of the descriptions. In this paper, we introduce a new metric UMIC, an Unreferenced Metric for Image Captioning which does not require reference captions to evaluate image captions. Based on Vision-and-Language BERT, we train UMIC to discriminate negative captions via contrastive learning. Also, we observe critical problems of the previous benchmark dataset (i.e., human annotations) on image captioning metric, and introduce a new collection of human annotations on the generated captions. We validate UMIC on four datasets, including our new dataset, and show that UMIC has a higher correlation than all previous metrics that require multiple references. We release the benchmark dataset and pre-trained models to compute the UMIC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A distilled 99M-parameter CLIP gives L-CLIPScore, a lightweight caption metric that matches CLIPScore on human correlation and best improves captioning models when mixed with CIDEr.

  2. Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new benchmark and fine-tuned models show that multi-task training improves average perceptual-similarity accuracy on known tasks but does not generalize to out-of-distribution perceptual tasks.

Pith tools