Pith. sign in

REVIEW 2 cited by

Improving Image Captioning with Better Use of Captions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.11807 v1 pith:LQO4VHO4 submitted 2020-06-21 cs.CV cs.CL

classification cs.CVcs.CL
keywords imagecaptioningvisualbettercaptionsextensivegenerationlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics available in captions and leverage that to enhance both image representation and caption generation. Our models first construct caption-guided visual relationship graphs that introduce beneficial inductive bias using weakly supervised multi-instance learning. The representation is then enhanced with neighbouring and contextual nodes with their textual and visual features. During generation, the model further incorporates visual relationships using multi-task learning for jointly predicting word and object/predicate tag sequences. We perform extensive experiments on the MSCOCO dataset, showing that the proposed framework significantly outperforms the baselines, resulting in the state-of-the-art performance under a wide range of evaluation metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trade-offs in Image Generation: How Do Different Dimensions Interact?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.

  2. Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A continuous-time RL algorithm that treats diffusion scores as actions fine-tunes text-to-image models with a Girsanov-based KL regularizer, showing stability across different denoising step counts.

Pith tools