Pith. sign in

REVIEW 2 cited by

Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.16367 v3 pith:DBL2REE3 submitted 2024-02-26 cs.CL

classification cs.CL
keywords activationdifferentllmsmultilingualpatternspruningdifferencesexpert
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, large language models (LLMs) have achieved tremendous breakthroughs in the field of NLP, but still lack understanding of their internal neuron activities when processing different languages. We designed a method to convert dense LLMs into fine-grained MoE architectures, and then visually studied the multilingual activation patterns of LLMs through expert activation frequency heatmaps. Through comprehensive experiments on different model families, different model sizes, and different variants, we analyzed the similarities and differences in the internal neuron activation patterns of LLMs when processing different languages. Specifically, we investigated the distribution of high-frequency activated experts, multilingual shared experts, whether multilingual activation patterns are related to language families, and the impact of instruction tuning on activation patterns. We further explored leveraging the discovered differences in expert activation frequencies to guide sparse activation and pruning. Experimental results demonstrated that our method significantly outperformed random expert pruning and even exceeded the performance of unpruned models in some languages. Additionally, we found that configuring different pruning rates for different layers based on activation level differences could achieve better results. Our findings reveal the multilingual processing mechanisms within LLMs and utilize these insights to offer new perspectives for applications such as sparse activation and model pruning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders

    cs.CL 2025-07 reject novelty 4.0 of 10

    Sparse-autoencoder analysis of Gemma-2-2B shows that medium-to-low resource languages get up to 26% lower activations than English, and LoRA fine-tuning that explicitly minimizes the activation gap raises activations ...

  2. Multilingual Large Language Models: A Systematic Survey

    cs.CL 2024-11 conditional novelty 3.0 of 10

    This is a systematic review that categorizes research on multilingual LLMs into architecture, corpora, tuning, evaluation, interpretability, and applications, with a public curated paper list.

Pith tools