REVIEW 2 cited by
Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their Applications
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recently, large language models (LLMs) have achieved tremendous breakthroughs in the field of NLP, but still lack understanding of their internal neuron activities when processing different languages. We designed a method to convert dense LLMs into fine-grained MoE architectures, and then visually studied the multilingual activation patterns of LLMs through expert activation frequency heatmaps. Through comprehensive experiments on different model families, different model sizes, and different variants, we analyzed the similarities and differences in the internal neuron activation patterns of LLMs when processing different languages. Specifically, we investigated the distribution of high-frequency activated experts, multilingual shared experts, whether multilingual activation patterns are related to language families, and the impact of instruction tuning on activation patterns. We further explored leveraging the discovered differences in expert activation frequencies to guide sparse activation and pruning. Experimental results demonstrated that our method significantly outperformed random expert pruning and even exceeded the performance of unpruned models in some languages. Additionally, we found that configuring different pruning rates for different layers based on activation level differences could achieve better results. Our findings reveal the multilingual processing mechanisms within LLMs and utilize these insights to offer new perspectives for applications such as sparse activation and model pruning.
Forward citations
Cited by 2 Pith papers
-
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
Sparse-autoencoder analysis of Gemma-2-2B shows that medium-to-low resource languages get up to 26% lower activations than English, and LoRA fine-tuning that explicitly minimizes the activation gap raises activations ...
-
Multilingual Large Language Models: A Systematic Survey
This is a systematic review that categorizes research on multilingual LLMs into architecture, corpora, tuning, evaluation, interpretability, and applications, with a public curated paper list.
Discussion (0). Continue with ORCID to comment.