Pith. sign in

REVIEW 1 cited by

Candidate Soups: Fusing Candidate Results Improves Translation Quality for Non-Autoregressive Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.11503 v1 pith:FGE2FXNA submitted 2023-01-27 cs.CL

classification cs.CL
keywords translationcandidatemodelinferencequalitysoupsfullymethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Non-autoregressive translation (NAT) model achieves a much faster inference speed than the autoregressive translation (AT) model because it can simultaneously predict all tokens during inference. However, its translation quality suffers from degradation compared to AT. And existing NAT methods only focus on improving the NAT model's performance but do not fully utilize it. In this paper, we propose a simple but effective method called "Candidate Soups," which can obtain high-quality translations while maintaining the inference speed of NAT models. Unlike previous approaches that pick the individual result and discard the remainders, Candidate Soups (CDS) can fully use the valuable information in the different candidate translations through model uncertainty. Extensive experiments on two benchmarks (WMT'14 EN-DE and WMT'16 EN-RO) demonstrate the effectiveness and generality of our proposed method, which can significantly improve the translation quality of various base models. More notably, our best variant outperforms the AT model on three translation tasks with 7.6 times speedup.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PARA: Parameter-Efficient Fine-tuning with Prompt Aware Representation Adjustment

    cs.CL 2025-02 conditional novelty 6.0 of 10

    PARA generates prompt-conditioned scaling vectors for Q, V, and FFN activations, outperforming (IA)^3 and LoRA-style baselines on several benchmarks with similar parameter counts and lower multi-tenant inference latency.

Pith tools