Pith. sign in

REVIEW 13 cited by

Benchmarking and Improving Detail Image Caption

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19092 v4 pith:PVKPIFFD submitted 2024-05-29 cs.CV

classification cs.CV
keywords captionevaluationdataimagedetaillvlmcaptioningcapture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image captioning has long been regarded as a fundamental task in visual understanding. Recently, however, few large vision-language model (LVLM) research discusses model's image captioning performance because of the outdated short-caption benchmarks and unreliable evaluation metrics. In this work, we propose to benchmark detail image caption task by curating high-quality evaluation datasets annotated by human experts, GPT-4V and Gemini-1.5-Pro. We also design a more reliable caption evaluation metric called CAPTURE (CAPtion evaluation by exTracting and coUpling coRE information). CAPTURE extracts visual elements, e.g., objects, attributes and relations from captions, and then matches these elements through three stages, achieving the highest consistency with expert judgements over other rule-based or model-based caption metrics. The proposed benchmark and metric provide reliable evaluation for LVLM's detailed image captioning ability. Guided by this evaluation, we further explore to unleash LVLM's detail caption capabilities by synthesizing high-quality data through a five-stage data construction pipeline. Our pipeline only uses a given LVLM itself and other open-source tools, without any human or GPT-4V annotation in the loop. Experiments show that the proposed data construction strategy significantly improves model-generated detail caption data quality for LVLMs with leading performance, and the data quality can be further improved in a self-looping paradigm. All code and dataset will be publicly available at https://github.com/foundation-multimodal-models/CAPTURE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. Holo-Captioning: Toward the Text Equivalent of 3D Scenes

    cs.CV 2026-07 conditional novelty 6.5 of 10

    HoloScribe generates pure-text structured captions of all entities, boxes, attributes and relations in 3D indoor scenes and outperforms dense captioners and 3D LLMs on a new 15K-scene benchmark.

  3. Reliability-Prioritized Fine-Grained Generation in Multimodal Large

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Proposes GranFact benchmark with coarse-to-fine annotations and a DPO variant that penalizes unreliable fine-grained claims to improve reliable specificity in MLLM outputs.

  4. SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new reference-free metric, SPECS, fine-tunes LongCLIP with a specificity objective and reaches LLM-level human correlation on long captions at a fraction of the computational cost.

  5. SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An RL framework that trains vision-language models to self-correct captions via a scene-graph-based reward outperforms SFT and DPO on caption quality.

  6. Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark and model for sentence-level factuality checking and error explanation of long image captions, with a correction pipeline that improves caption accuracy.

  7. RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RICO refines image captions by reconstructing them into images with a text-to-image model and asking GPT-4o to fix discrepancies against the original, iteratively, with a DPO-distilled fast variant.

  8. GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A multi-agent vision-language framework that extracts event, time, and location from public event images, evaluated with a new soft metric on VLM-augmented datasets.

  9. Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProxyV introduces proxy vision tokens that take over expensive attention and feed-forward computation in later layers of decoder-only multimodal models, cutting FLOPs by 25-46% while retaining or improving accuracy on...

  10. CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Caption quality can be split into Coverage and Precision; in controlled training runs, Coverage predicts VLM understanding while Precision predicts T2I generation.

  11. Exploring The Visual Feature Space for Multimodal Neural Decoding

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VINDEX decodes fMRI into nested 9-token CLIP features that feed a frozen MLLM, improving detailed brain captioning and QA over a regression baseline, with a new benchmark.

  12. Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A detailness score combining object coverage and per-object description depth selects 20% of captions that train a text-to-image model better than the full dataset.

  13. SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.

Pith tools