Pith. sign in

REVIEW 6 cited by

Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1412.6632 v5 pith:HOBVZACR submitted 2014-12-20 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords modelm-rnndeepimagemultimodalnetworkneuralrecurrent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel image captions. It directly models the probability distribution of generating a word given previous words and an image. Image captions are generated by sampling from this distribution. The model consists of two sub-networks: a deep recurrent neural network for sentences and a deep convolutional network for images. These two sub-networks interact with each other in a multimodal layer to form the whole m-RNN model. The effectiveness of our model is validated on four benchmark datasets (IAPR TC-12, Flickr 8K, Flickr 30K and MS COCO). Our model outperforms the state-of-the-art methods. In addition, we apply the m-RNN model to retrieval tasks for retrieving images or sentences, and achieves significant performance improvement over the state-of-the-art methods which directly optimize the ranking objective function for retrieval. The project page of this work is: www.stat.ucla.edu/~junhua.mao/m-RNN.html .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aesthetic Image Captioning From Weakly-Labelled Photographs

    cs.CV 2019-08 conditional novelty 6.0 of 10

    By filtering noisy web comments, the authors built AVA-Captions, a 230,000-image aesthetic captioning dataset, and showed a weakly supervised CNN can match ImageNet-pretrained features for this task.

  2. Dynamic Stale Synchronous Parallel Distributed Training for Deep Learning

    cs.DC 2019-08 conditional novelty 6.0 of 10

    DSSP dynamically sets the staleness threshold in stale synchronous parallel training, reducing waiting time and reaching target accuracy faster than BSP and SSP in the tested GPU clusters.

  3. Image Captioning using Facial Expression and Attention

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Facial expression features, especially with attention, yield small captioning improvements on face-containing Flickr images, driven mostly by more diverse verbs.

  4. UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer

    cs.CV 2024-12 reject novelty 4.0 of 10

    The authors combine factual and stylized image captioning with a transformer summarizer to output a single caption containing factual, romantic, and humorous elements.

  5. Semi Supervised Phrase Localization in a Bidirectional Caption-Image Retrieval Framework

    cs.CV 2019-08 conditional novelty 4.0 of 10

    A retrieval-trained neural network produces word and phrase localization maps, reaching 51.06 pointing-game accuracy on Flickr30K Entities, the best reported score among weakly supervised methods.

  6. Scene-based Factored Attention for Image Captioning

    cs.CV 2019-08 conditional novelty 4.0 of 10

    A scene-conditioned factored attention module, which multiplies attention weights by the image's predicted scene vector, improves MS COCO captioning scores over an Up-Down re-implementation.

Pith tools