Pith. sign in

REVIEW 2 cited by

Probing the Need for Visual Context in Multimodal Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.08678 v2 pith:7WSAS4LD submitted 2019-03-20 cs.CL

classification cs.CL
keywords visualcontextmodelsmodalitytextualcurrenteithermachine
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Current work on multimodal machine translation (MMT) has suggested that the visual modality is either unnecessary or only marginally beneficial. We posit that this is a consequence of the very simple, short and repetitive sentences used in the only available dataset for the task (Multi30K), rendering the source text sufficient as context. In the general case, however, we believe that it is possible to combine visual and textual information in order to ground translations. In this paper we probe the contribution of the visual modality to state-of-the-art MMT models by conducting a systematic analysis where we partially deprive the models from source-side textual context. Our results show that under limited textual context, models are capable of leveraging the visual input to generate better translations. This contradicts the current belief that MMT models disregard the visual modality because of either the quality of the image features or the way they are integrated into the model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries

    cs.CL 2025-05 conditional novelty 6.0 of 10

    This paper builds TopicVD, a topic-based documentary video-subtitle translation dataset, and shows with a cross-modal attention model that visual and contextual information improve BLEU scores.

  2. Predicting Actions to Help Predict Translations

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Action-aware visual features improve English-to-Portuguese translation on How2 by up to 0.4 BLEU, with the largest gains when verbs are masked in the source text.

Pith tools