REVIEW 26 cited by
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce the Aya Expanse model family, a new generation of 8B and 32B parameter multilingual language models, aiming to address the critical challenge of developing highly performant multilingual models that match or surpass the capabilities of monolingual models. By leveraging several years of research at Cohere For AI and Cohere, including advancements in data arbitrage, multilingual preference training, and model merging, Aya Expanse sets a new state-of-the-art in multilingual performance. Our evaluations on the Arena-Hard-Auto dataset, translated into 23 languages, demonstrate that Aya Expanse 8B and 32B outperform leading open-weight models in their respective parameter classes, including Gemma 2, Qwen 2.5, and Llama 3.1, achieving up to a 76.6% win-rate. Notably, Aya Expanse 32B outperforms Llama 3.1 70B, a model with twice as many parameters, achieving a 54.0% win-rate. In this short technical report, we present extended evaluation results for the Aya Expanse model family and release their open-weights, together with a new multilingual evaluation dataset m-ArenaHard.
Forward citations
Cited by 26 Pith papers
-
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
Multilingual LLM evaluators systematically inflate scores for lower-resource languages, and the standard pairwise-accuracy metric cannot detect the resulting safety-threshold disparities.
-
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.
-
Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon
Cross-lingual transfer of cultural knowledge is bidirectional for high-resource languages and asymmetric for low-resource ones, with corpus frequency correlating with transfer success.
-
LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World
LEMONADE is a new 20-language, expert-annotated conflict event dataset for abstractive event extraction, and ZEST, a zero-shot retrieval entity linker, beats prior zero-shot baselines but trails supervised models.
-
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
A multilingual extension of the CUTE benchmark shows that LLM token-manipulation performance varies by language and script, with surprisingly strong results on low-resource languages and weak results on sub-character ...
-
Disentangling Language Modeling and Boundaries
The paper hypothesizes that next-byte and boundary distributions in byte-level LMs can be disentangled, proposes two experiments to test it, but provides no experimental results.
-
LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal
Large language models encode query-document relevance as a linearly decodable internal signal that strengthens in middle-to-late layers and, in several models, outperforms their own generated judgments.
-
Massively Multilingual Joint Segmentation and Glossing
A jointly trained multilingual seq2seq model (PolyGloss) predicts morphological segmentation and glosses together, achieving lower morpheme error rates than GlossLM and open LLMs while giving perfectly aligned outputs...
-
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...
-
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.
-
When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
Projecting away the estimated Modern Standard Arabic subspace during fine-tuning improves generation across 25 Arabic dialects by up to +4.9 chrF++, evidence that subspace dominance by a high-resource variety restrict...
-
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
MELAC introduces 19 Persian and Iranian-culture evaluation datasets and benchmarks 41 LLMs, showing weak performance on Iranian-specific content.
-
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.
-
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
Hedged sampling, checklist-based one-pass selection (CHOPS), and cross-lingual MBR (X-MBR) improve multilingual LLM output quality when scaling from one to five samples.
-
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
A multilingual LLM training method that groups similar languages, converts high-deviation layers into mixture-of-experts layers, and assigns one expert per language group improves perplexity across 18 to 128 languages.
-
On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures
LLMs are less accurate when asked to report facts in non-default measurement systems, and chain-of-thought restores accuracy only at a 180-300 percent increase in test-time compute.
-
M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
M-Wanda improves one-shot pruning for multilingual LLMs by combining cross-lingual activation variance, activation probability, and correlation-weighted layerwise sparsity, reducing perplexity and boosting downstream ...
-
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
The new Fann or Flop benchmark measures LLM comprehension of Arabic poetry through expert-written verse explanations and shows current LLMs perform poorly on interpretive depth.
-
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
KoBALT, an expert-crafted 700-question Korean linguistic benchmark, finds that even the best LLM answers only 61% correctly, with human preference ratings correlating moderately with benchmark accuracy.
-
Gaokerena: A Small Persian Medical Language Model Family
Fine-tuned Persian medical language models reach 49-53% on translated medical MMLU, with datasets released, but the reasoning variant's gain depends on extra test-time compute and a verifier.
-
Zoom In Disparities in Healthcare LLM Q&A
Health Q&A chatbots align their answers with English Wikipedia even for non-English prompts, and injecting a non-English Wikipedia excerpt at query time shifts answers toward local references.
-
Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
A 7B open-weight translation model matches or outperforms far larger commercial systems across 28 languages in automatic and human evaluations.
-
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
A large LLM-to-LLM math tutoring simulation across 11 languages shows English-language hints often yield the largest accuracy gains for student models, but the low-resource-language results lack statistical support.
-
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
Repeated sampling with a verifier improves open-ended multilingual generation, and reward-based verifiers are needed for math and code tasks.
-
CSIRO-LT at SemEval-2025 Task 11: Adapting LLMs for Emotion Recognition for Multiple Languages
On the SemEval-2025 multilingual emotion recognition task, directly fine-tuning a multilingual LLM with LoRA separately for each language beat zero-shot, few-shot, and English-bridged adaptation in the authors' experiments.
-
From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation
On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.
Discussion (0). Sign in to comment.