Pith. sign in

REVIEW

Learning to Predict: A Fast Re-constructive Method to Generate Multimodal Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.08737 v1 pith:U5GD3BJX submitted 2017-03-25 stat.ML

classification stat.ML
keywords multimodalmethodbuildembeddingsinformationlearningre-constructiverepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrating visual and linguistic information into a single multimodal representation is an unsolved problem with wide-reaching applications to both natural language processing and computer vision. In this paper, we present a simple method to build multimodal representations by learning a language-to-vision mapping and using its output to build multimodal embeddings. In this sense, our method provides a cognitively plausible way of building representations, consistent with the inherently re-constructive and associative nature of human memory. Using seven benchmark concept similarity tests we show that the mapped vectors not only implicitly encode multimodal information, but also outperform strong unimodal baselines and state-of-the-art multimodal methods, thus exhibiting more "human-like" judgments---particularly in zero-shot settings.

Discussion (0). Continue with ORCID to comment.

Pith tools