Pith. sign in

REVIEW 4 cited by

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.12329 v1 pith:HGO6IF7R submitted 2025-03-16 cs.CV cs.CL

classification cs.CVcs.CL
keywords captioningcaparenadetailedhumanimagecaptionmetricsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive image descriptions. However, benchmarking the quality of such captions remains unresolved. This paper addresses two key questions: (1) How well do current VLMs actually perform on image captioning, particularly compared to humans? We built CapArena, a platform with over 6000 pairwise caption battles and high-quality human preference votes. Our arena-style evaluation marks a milestone, showing that leading models like GPT-4o achieve or even surpass human performance, while most open-source models lag behind. (2) Can automated metrics reliably assess detailed caption quality? Using human annotations from CapArena, we evaluate traditional and recent captioning metrics, as well as VLM-as-a-Judge. Our analysis reveals that while some metrics (e.g., METEOR) show decent caption-level agreement with humans, their systematic biases lead to inconsistencies in model ranking. In contrast, VLM-as-a-Judge demonstrates robust discernment at both the caption and model levels. Building on these insights, we release CapArena-Auto, an accurate and efficient automated benchmark for detailed captioning, achieving 94.3% correlation with human rankings at just $4 per test. Data and resources will be open-sourced at https://caparena.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    EviSelect uses the target multimodal model's internal attention as a prior to dynamically select frames, sampling rates, and resolutions, achieving about 50% token reduction and a 3.9x speedup with better benchmark accuracy.

  2. Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A weakly supervised, retrieval-augmented vision-language framework generates structured SOAP notes from lesion images and sparse clinical text, with evaluation against GPT-4o, Claude, and Janus Pro on a small set of cases.

  3. A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Reference-free caption quality is scored by the downstream vision-language accuracy of a caption-conditioned reconstructed image, via a new CTTD benchmark.

  4. Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Skin-SOAP is a weakly supervised multimodal system that turns a skin lesion image and sparse clinical text into structured SOAP notes, evaluated with two new metrics.

Pith tools