REVIEW 4 cited by
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Most existing image captioning evaluation metrics focus on assigning a single numerical score to a caption by comparing it with reference captions. However, these methods do not provide an explanation for the assigned score. Moreover, reference captions are expensive to acquire. In this paper, we propose FLEUR, an explainable reference-free metric to introduce explainability into image captioning evaluation metrics. By leveraging a large multimodal model, FLEUR can evaluate the caption against the image without the need for reference captions, and provide the explanation for the assigned score. We introduce score smoothing to align as closely as possible with human judgment and to be robust to user-defined grading criteria. FLEUR achieves high correlations with human judgment across various image captioning evaluation benchmarks and reaches state-of-the-art results on Flickr8k-CF, COMPOSITE, and Pascal-50S within the domain of reference-free evaluation metrics. Our source code and results are publicly available at: https://github.com/Yebin46/FLEUR.
Forward citations
Cited by 4 Pith papers
-
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
A new reference-free metric, SPECS, fine-tunes LongCLIP with a specificity objective and reaches LLM-level human correlation on long captions at a fraction of the computational cost.
-
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
A fine-tuned multilingual CLIP model rates image captions in ten languages with human-judgment correlation as high as English-only models on English data, using machine-translated benchmarks and native multicultural tests.
-
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
Reference-free caption quality is scored by the downstream vision-language accuracy of a caption-conditioned reconstructed image, via a new CTTD benchmark.
-
RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
RIVAL iteratively re-trains a reward model adversarially against the current translator and adds a BLEU-predicting head, improving in-domain WMT and subtitle translation over SFT baselines.
Discussion (0). Continue with ORCID to comment.