Pith. sign in

REVIEW 1 cited by

Learning Convolutional Text Representations for Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.06824 v2 pith:IAE2CFUL submitted 2017-05-18 cs.LG cs.CLcs.NEstat.ML

classification cs.LGcs.CLcs.NEstat.ML
keywords questionansweringvisualconvolutionallearningnetworksneuraltasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual question answering is a recently proposed artificial intelligence task that requires a deep understanding of both images and texts. In deep learning, images are typically modeled through convolutional neural networks, and texts are typically modeled through recurrent neural networks. While the requirement for modeling images is similar to traditional computer vision tasks, such as object recognition and image classification, visual question answering raises a different need for textual representation as compared to other natural language processing tasks. In this work, we perform a detailed analysis on natural language questions in visual question answering. Based on the analysis, we propose to rely on convolutional neural networks for learning textual representations. By exploring the various properties of convolutional neural networks specialized for text data, such as width and depth, we present our "CNN Inception + Gate" model. We show that our model improves question representations and thus the overall accuracy of visual question answering models. We also show that the text representation requirement in visual question answering is more complicated and comprehensive than that in conventional natural language processing tasks, making it a better task to evaluate textual representation methods. Shallow models like fastText, which can obtain comparable results with deep learning models in tasks like text classification, are not suitable in visual question answering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Performance Analysis of Traditional VQA Models Under Limited Computational Resources

    cs.CV 2025-02 reject novelty 2.0 of 10

    An empirical comparison claims BidGRU with embedding size 300 and vocabulary 3000 is the best resource-constrained VQA configuration, but the paper lacks dataset and statistical details.

Pith tools