An adaptively sparse Transformer with per-head learned α-entmax attention yields sparser, more confident attention heads and slight BLEU gains over softmax Transformers on four machine translation datasets.
Marian: Cost-effective High-Quality Neural Machine Translation in C++
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper describes the submissions of the "Marian" team to the WNMT 2018 shared task. We investigate combinations of teacher-student training, low-precision matrix products, auto-tuning and other methods to optimize the Transformer model on GPU and CPU. By further integrating these methods with the new averaging attention networks, a recently introduced faster Transformer variant, we create a number of high-quality, high-performance models on the GPU and CPU, dominating the Pareto frontier for this shared task.
fields
cs.CL 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adaptively Sparse Transformers
An adaptively sparse Transformer with per-head learned α-entmax attention yields sparser, more confident attention heads and slight BLEU gains over softmax Transformers on four machine translation datasets.