Pith. sign in

REVIEW 1 cited by

Improving End-to-End Text Image Translation From the Auxiliary Text Translation Task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03887 v1 pith:6PLMZY5T submitted 2022-10-08 cs.CL

classification cs.CL
keywords texttranslationend-to-endimageauxiliarymodelmulti-tasktasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end text image translation (TIT), which aims at translating the source language embedded in images to the target language, has attracted intensive attention in recent research. However, data sparsity limits the performance of end-to-end text image translation. Multi-task learning is a non-trivial way to alleviate this problem via exploring knowledge from complementary related tasks. In this paper, we propose a novel text translation enhanced text image translation, which trains the end-to-end model with text translation as an auxiliary task. By sharing model parameters and multi-task training, our model is able to take full advantage of easily-available large-scale text parallel corpus. Extensive experimental results show our proposed method outperforms existing end-to-end methods, and the joint multi-task learning with both text translation and recognition tasks achieves better results, proving translation and recognition auxiliary tasks are complementary.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring In-Image Machine Translation with Real-World Background

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DebackX translates text inside images by separating text from the background, translating the text-image directly, and fusing it back, outperforming prior IIMT models on a new real-background dataset.

Pith tools