Pith. sign in

REVIEW 2 cited by

Contributions of Transformer Attention Heads in Multi- and Cross-lingual Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.08375 v1 pith:LVI2C3FO submitted 2021-08-18 cs.CL cs.LG

classification cs.CLcs.LG
keywords headsmulti-lingualtasksattentioncross-lingualexperimentspruninglanguages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies the relative importance of attention heads in Transformer-based models to aid their interpretability in cross-lingual and multi-lingual tasks. Prior research has found that only a few attention heads are important in each mono-lingual Natural Language Processing (NLP) task and pruning the remaining heads leads to comparable or improved performance of the model. However, the impact of pruning attention heads is not yet clear in cross-lingual and multi-lingual tasks. Through extensive experiments, we show that (1) pruning a number of attention heads in a multi-lingual Transformer-based model has, in general, positive effects on its performance in cross-lingual and multi-lingual tasks and (2) the attention heads to be pruned can be ranked using gradients and identified with a few trial experiments. Our experiments focus on sequence labeling tasks, with potential applicability on other cross-lingual and multi-lingual tasks. For comprehensiveness, we examine two pre-trained multi-lingual models, namely multi-lingual BERT (mBERT) and XLM-R, on three tasks across 9 languages each. We also discuss the validity of our findings and their extensibility to truly resource-scarce languages and other task settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Task-KV identifies 'heterogeneous' attention heads by distance from a per-task semantic center and allocates differentiated KV cache budgets, achieving modest average gains on long-context benchmarks at reduced memory.

  2. FewTopNER: Integrating Few-Shot Learning with Topic Modeling and Named Entity Recognition in a Multilingual Framework

    cs.CL 2025-02 conditional novelty 5.0 of 10

    FewTopNER reports that adding a topic-modeling branch to a prototype-based few-shot NER model improves multilingual F1 by 2.5 to 4.0 points and increases topic coherence scores.

Pith tools