Pith. sign in

REVIEW 6 cited by

Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08771 v2 pith:TQC4MN5T submitted 2021-04-18 cs.CL

classification cs.CL
keywords translationcross-attentionfine-tuningmachineexperimentsextendlanguagemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the power of cross-attention in the Transformer architecture within the context of transfer learning for machine translation, and extend the findings of studies into cross-attention when training from scratch. We conduct a series of experiments through fine-tuning a translation model on data where either the source or target language has changed. These experiments reveal that fine-tuning only the cross-attention parameters is nearly as effective as fine-tuning all parameters (i.e., the entire translation model). We provide insights into why this is the case and observe that limiting fine-tuning in this manner yields cross-lingually aligned embeddings. The implications of this finding for researchers and practitioners include a mitigation of catastrophic forgetting, the potential for zero-shot translation, and the ability to extend machine translation models to several new language pairs with reduced parameter storage overhead.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization

    cs.SD 2026-01 unverdicted novelty 6.0 of 10

    LLMs using in-context learning and fine-tuning on listener experiment data generate equalization settings that align better with population preferences than random sampling or static presets.

  2. PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    PLoP selects LoRA adapter placement by ranking normalized feature norms and placing adapters on the lowest-scoring module types, using only forward passes.

  3. Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Krul dynamically selects per-conversation cross-layer KV cache compression from attention similarity patterns, cutting TTFT by 1.28x to 2.68x and KV storage by 1.33x to 2.35x with less than 1% average accuracy loss on...

  4. TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages

    cs.CL 2025-06 reject novelty 5.0 of 10

    A trainable pipeline of translation models and a frozen LLM improves Hebrew last-word prediction accuracy to 5.59%, about twice the best baseline.

  5. BlastOFormer: Attention and Neural Operator Deep Learning Methods for Explosive Blast Prediction

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BlastOFormer, a transformer with signed-distance inputs and an auxiliary unscaling CNN, predicts full-field blast peak pressures with R2 up to 0.9516 in the unscaled domain and about 100,000x faster than blastFoam CFD.

  6. Mixture of Low Rank Adaptation with Partial Parameter Sharing for Time Series Forecasting

    cs.LG 2025-05 conditional novelty 4.0 of 10

    MoLA adapts a pre-trained short-horizon forecaster to multiple forecast steps via segment-specific mixtures of shared low-rank adapters, reporting modest mean-squared-error gains over the base models on most of eight ...

Pith tools