Pith. sign in

Marian: Cost-effective High-Quality Neural Machine Translation in C++

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper describes the submissions of the "Marian" team to the WNMT 2018 shared task. We investigate combinations of teacher-student training, low-precision matrix products, auto-tuning and other methods to optimize the Transformer model on GPU and CPU. By further integrating these methods with the new averaging attention networks, a recently introduced faster Transformer variant, we create a number of high-quality, high-performance models on the GPU and CPU, dominating the Pareto frontier for this shared task.

fields

cs.CL 1

years

2019 1

verdicts

CONDITIONAL 1

representative citing papers

Adaptively Sparse Transformers

cs.CL · 2019-08-30 · conditional · novelty 7.0

An adaptively sparse Transformer with per-head learned α-entmax attention yields sparser, more confident attention heads and slight BLEU gains over softmax Transformers on four machine translation datasets.

citing papers explorer

Showing 1 of 1 citing paper.

  • Adaptively Sparse Transformers cs.CL · 2019-08-30 · conditional · none · ref 17 · internal anchor

    An adaptively sparse Transformer with per-head learned α-entmax attention yields sparser, more confident attention heads and slight BLEU gains over softmax Transformers on four machine translation datasets.