Pith. sign in

REVIEW 1 cited by

ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00850 v2 pith:VHHKA3XC submitted 2020-05-02 cs.CL cs.LG

classification cs.CLcs.LG
keywords non-autoregressivemodelautoregressiveinferencetranslationapproachenergyenergy-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose to train a non-autoregressive machine translation model to minimize the energy defined by a pretrained autoregressive model. In particular, we view our non-autoregressive translation system as an inference network (Tu and Gimpel, 2018) trained to minimize the autoregressive teacher energy. This contrasts with the popular approach of training a non-autoregressive model on a distilled corpus consisting of the beam-searched outputs of such a teacher model. Our approach, which we call ENGINE (ENerGy-based Inference NEtworks), achieves state-of-the-art non-autoregressive results on the IWSLT 2014 DE-EN and WMT 2016 RO-EN datasets, approaching the performance of autoregressive models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A multi-level optimal transport loss combining sequence-level ranking, top-k truncation, and Sinkhorn sequence distance outperforms earlier cross-tokenizer distillation losses on QA and summarization.

Pith tools