Pith. sign in

REVIEW 5 cited by

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2509.06888 v1 pith:C6W5OZB6 submitted 2025-09-08 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords languagesmodelsmmbertclassificationincludinglow-resourcemultilingualdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Encoder-only languages models are frequently used for a variety of standard machine learning tasks, including classification and retrieval. However, there has been a lack of recent research for encoder models, especially with respect to multilingual models. We introduce mmBERT, an encoder-only language model pretrained on 3T tokens of multilingual text in over 1800 languages. To build mmBERT we introduce several novel elements, including an inverse mask ratio schedule and an inverse temperature sampling ratio. We add over 1700 low-resource languages to the data mix only during the decay phase, showing that it boosts performance dramatically and maximizes the gains from the relatively small amount of training data. Despite only including these low-resource languages in the short decay phase we achieve similar classification performance to models like OpenAI's o3 and Google's Gemini 2.5 Pro. Overall, we show that mmBERT significantly outperforms the previous generation of models on classification and retrieval tasks -- on both high and low-resource languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    cs.CL 2026-07 conditional novelty 7.0 of 10

    With matched open data and backbones, ColBERT-style late interaction turns English translate-train into multilingual generalization, while dense retrieval stays mostly inside the translated languages.

  2. Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders

    cs.IR 2026-07 conditional novelty 7.0 of 10

    Bekko a8m, with 7.7M active parameters, scores 56.2 on MMTEB Multilingual v2 Retrieval, beating mE5 models and BGE-M3, while a25m reaches 57.5, on par with gte-multilingual-base.

  3. Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

    cs.IR 2026-01 conditional novelty 6.0 of 10

    A new benchmark dataset and evaluation framework for detecting intertextual references in Latin literature, with baseline results showing moderate performance of dense retrieval and classification models.

  4. HalleluBERT: Let Every Token That Has Meaning Bear Its Weight

    cs.CL 2025-10 conditional novelty 5.0 of 10

    HalleluBERT, a Hebrew-only RoBERTa encoder family trained from scratch at scale, reports the highest unweighted mean scores on BMC, NEMO, and SMCD benchmarks, but without statistical significance testing.

  5. SindBERT, the Sailor: Charting the Seas of Turkish NLP

    cs.CL 2025-10 conditional novelty 5.0 of 10

    SindBERT releases Turkish RoBERTa base/large models trained on 312GB of text; they match existing models, with the large variant best on two of four tasks and little scaling gain.

Pith tools