Pith. sign in

REVIEW 1 cited by

Boosting Entity-aware Image Captioning with Multi-modal Knowledge Graph

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.11970 v1 pith:B7ZJZPTR submitted 2021-07-26 cs.CV

classification cs.CV
keywords entitiesimageknowledgenamedgrapharticlecaptioningmulti-modal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Entity-aware image captioning aims to describe named entities and events related to the image by utilizing the background knowledge in the associated article. This task remains challenging as it is difficult to learn the association between named entities and visual cues due to the long-tail distribution of named entities. Furthermore, the complexity of the article brings difficulty in extracting fine-grained relationships between entities to generate informative event descriptions about the image. To tackle these challenges, we propose a novel approach that constructs a multi-modal knowledge graph to associate the visual objects with named entities and capture the relationship between entities simultaneously with the help of external knowledge collected from the web. Specifically, we build a text sub-graph by extracting named entities and their relationships from the article, and build an image sub-graph by detecting the objects in the image. To connect these two sub-graphs, we propose a cross-modal entity matching module trained using a knowledge base that contains Wikipedia entries and the corresponding images. Finally, the multi-modal knowledge graph is integrated into the captioning model via a graph attention mechanism. Extensive experiments on both GoodNews and NYTimes800k datasets demonstrate the effectiveness of our method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model

    cs.CV 2025-05 conditional novelty 3.0 of 10

    Applying beam search, patch self-attention, and cosine scheduling to the K-Replay captioning framework improves knowledge-keyword recognition on KnowCap, though the full combined model is not reported.

Pith tools