EuroBERT: Scaling Multilingual Encoders for European Languages

Andr\'e Martins; Ayoub Hammal; Caio Corro; C\'eline Hudelot; Duarte M. Alves; Emmanuel Malherbe; Etienne Malaboeuf; Fanny Jourdan; Gabriel Hautreux; Hippolyte Gisserot-Boukhlef

arxiv: 2503.05500 · v3 · pith:PK3E3VS4new · submitted 2025-03-07 · 💻 cs.CL · cs.AI

EuroBERT: Scaling Multilingual Encoders for European Languages

Nicolas Boizard , Hippolyte Gisserot-Boukhlef , Duarte M. Alves , Andr\'e Martins , Ayoub Hammal , Caio Corro , C\'eline Hudelot , Emmanuel Malherbe

show 11 more authors

Etienne Malaboeuf Fanny Jourdan Gabriel Hautreux Jo\~ao Alves Kevin El Haddad Manuel Faysse Maxime Peyrard Nuno M. Guerreiro Patrick Fernandes Ricardo Rei Pierre Colombo

This is my paper

classification 💻 cs.CL cs.AI

keywords multilingualencoderseurobertmodelstrainingadvanceseuropeanlanguages

0 comments

read the original abstract

General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in generative decoder-only models. However, many innovations driving this progress are not inherently tied to decoders. In this paper, we revisit the development of multilingual encoders through the lens of these advances, and introduce EuroBERT, a family of multilingual encoders covering European and widely spoken global languages. Our models outperform existing alternatives across a diverse range of tasks, spanning multilingual capabilities, mathematics, and coding, and natively supporting sequences of up to 8,192 tokens. We also examine the design decisions behind EuroBERT, offering insights into our dataset composition and training pipeline. We publicly release the EuroBERT models, including intermediate training checkpoints, together with our training framework.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

LEDGER: A Long-Context Benchmark of Corporate Annual Reports for Grounded Financial Retrieval and Extraction
cs.CL 2026-06 unverdicted novelty 7.0

LEDGER provides a corpus of 4,999 annual reports with 31 labeled KPIs and three benchmarks for page-level retrieval, needle-in-haystack lookup, and full KPI extraction from long documents.
Is She Even Relevant? When BERT Ignores Explicit Gender Cues
cs.CL 2026-05 conditional novelty 7.0

A Dutch BERT model encodes gender linearly by epoch 20 but does not dynamically update its representations when explicit female cues contradict learned stereotypical associations in short sentence templates.
Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining
cs.CL 2026-06 unverdicted novelty 6.0

A web curation recipe using medical-term density filtering and LLM signal-amplifying rephrasing produces FineMed, a French medical pretraining corpus, and DoctoBERT encoders that outperform prior approaches on medical...
Should We Still Pretrain Encoders with Masked Language Modeling?
cs.CL 2025-07 accept novelty 6.0

Controlled ablations of 38 models find MLM superior to CLM on representation benchmarks while CLM offers better data efficiency and stability; a biphasic CLM-then-MLM schedule is optimal under fixed compute and improv...
jina-embeddings-v5-text: Task-Targeted Embedding Distillation
cs.CL 2026-02 unverdicted novelty 5.0

A distillation-plus-task-contrastive training regimen yields compact embedding models that match or exceed state-of-the-art performance for their size while supporting 32k-token contexts and quantization.
DunbaaBERT: From Sacrifice to Semantics
cs.CL 2026-05 unverdicted novelty 4.0

DunbaaBERT releases competitive Urdu encoder models trained from scratch, with the 32k-vocab variant showing the best efficiency profile across acceptability, classification, and sentiment tasks.
Ideology Prediction of German Political Texts
cs.CL 2026-05 unverdicted novelty 4.0

Transformer models predict German political ideology on a continuous left-right scale, reaching F1 0.844 in-domain and MAE 0.172 on newspaper out-of-domain tests.