Pith. sign in

REVIEW 3 cited by

Glancing Transformer for Non-Autoregressive Neural Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.07905 v3 pith:2AUMUSOP submitted 2020-08-18 cs.CL

classification cs.CL
keywords transformertranslationdecodingglancingglatmachinenon-autoregressiveparallel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work on non-autoregressive neural machine translation (NAT) aims at improving the efficiency by parallel decoding without sacrificing the quality. However, existing NAT methods are either inferior to Transformer or require multiple decoding passes, leading to reduced speedup. We propose the Glancing Language Model (GLM), a method to learn word interdependency for single-pass parallel generation models. With GLM, we develop Glancing Transformer (GLAT) for machine translation. With only single-pass parallel decoding, GLAT is able to generate high-quality translation with 8-15 times speedup. Experiments on multiple WMT language directions show that GLAT outperforms all previous single pass non-autoregressive methods, and is nearly comparable to Transformer, reducing the gap to 0.25-0.9 BLEU points.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenHMR: Generative Human Mesh Recovery

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GenHMR applies masked generative token prediction and 2D-pose-guided latent refinement to monocular human mesh recovery, reporting state-of-the-art MPJPE on Human3.6M, 3DPW, and EMDB.

  2. Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge

    cs.SD 2025-05 conditional novelty 5.0 of 10

    Training a neural speech enhancer on pseudo labels derived from aligned close-talk recordings improves far-field meeting ASR, reaching second place in the MISP-Meeting Challenge.

  3. Self-Evolution Knowledge Distillation for LLM-based Machine Translation

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A token-adaptive distillation method that mixes teacher and ground-truth targets only for hard tokens yields consistent BLEU gains in LLM translation.

Pith tools