REVIEW 6 cited by
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel image captions. It directly models the probability distribution of generating a word given previous words and an image. Image captions are generated by sampling from this distribution. The model consists of two sub-networks: a deep recurrent neural network for sentences and a deep convolutional network for images. These two sub-networks interact with each other in a multimodal layer to form the whole m-RNN model. The effectiveness of our model is validated on four benchmark datasets (IAPR TC-12, Flickr 8K, Flickr 30K and MS COCO). Our model outperforms the state-of-the-art methods. In addition, we apply the m-RNN model to retrieval tasks for retrieving images or sentences, and achieves significant performance improvement over the state-of-the-art methods which directly optimize the ranking objective function for retrieval. The project page of this work is: www.stat.ucla.edu/~junhua.mao/m-RNN.html .
Forward citations
Cited by 6 Pith papers
-
Aesthetic Image Captioning From Weakly-Labelled Photographs
By filtering noisy web comments, the authors built AVA-Captions, a 230,000-image aesthetic captioning dataset, and showed a weakly supervised CNN can match ImageNet-pretrained features for this task.
-
Dynamic Stale Synchronous Parallel Distributed Training for Deep Learning
DSSP dynamically sets the staleness threshold in stale synchronous parallel training, reducing waiting time and reaching target accuracy faster than BSP and SSP in the tested GPU clusters.
-
Image Captioning using Facial Expression and Attention
Facial expression features, especially with attention, yield small captioning improvements on face-containing Flickr images, driven mostly by more diverse verbs.
-
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer
The authors combine factual and stylized image captioning with a transformer summarizer to output a single caption containing factual, romantic, and humorous elements.
-
Semi Supervised Phrase Localization in a Bidirectional Caption-Image Retrieval Framework
A retrieval-trained neural network produces word and phrase localization maps, reaching 51.06 pointing-game accuracy on Flickr30K Entities, the best reported score among weakly supervised methods.
-
Scene-based Factored Attention for Image Captioning
A scene-conditioned factored attention module, which multiplies attention weights by the image's predicted scene vector, improves MS COCO captioning scores over an Up-Down re-implementation.
Discussion (0). Continue with ORCID to comment.