Pith. sign in

REVIEW 14 cited by

Aya 23: Open Weight Releases to Further Multilingual Progress

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15032 v2 pith:DMRFVOC2 submitted 2024-05-23 cs.CL

classification cs.CL
keywords multilinguallanguagesmodelmodelslanguageexpandinglikeopen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This technical report introduces Aya 23, a family of multilingual language models. Aya 23 builds on the recent release of the Aya model (\"Ust\"un et al., 2024), focusing on pairing a highly performant pre-trained model with the recently released Aya collection (Singh et al., 2024). The result is a powerful multilingual large language model serving 23 languages, expanding state-of-art language modeling capabilities to approximately half of the world's population. The Aya model covered 101 languages whereas Aya 23 is an experiment in depth vs breadth, exploring the impact of allocating more capacity to fewer languages that are included during pre-training. Aya 23 outperforms both previous massively multilingual models like Aya 101 for the languages it covers, as well as widely used models like Gemma, Mistral and Mixtral on an extensive range of discriminative and generative tasks. We release the open weights for both the 8B and 35B models as part of our continued commitment for expanding access to multilingual progress.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Hallucination signals from LLM internals transfer across English–Arabic and Arabic domains for most models, depending on class separability and feature-space language alignment.

  2. ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

    cs.CL 2025-10 conditional novelty 7.0 of 10

    ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.

  3. Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A 1,834-question Indonesian dilemma benchmark shows leading LLMs match Indonesian human choices only about half the time and struggle most on religion and unity.

  4. Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Incidental multilingualism from uneven web training makes LLMs unequal, brittle, and opaque across languages.

  5. Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG

    cs.CL 2025-09 conditional novelty 6.0 of 10

    In multilingual retrieval-augmented generation, models cite English evidence more accurately than translated evidence, and this language preference can outweigh document relevance.

  6. When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Projecting away the estimated Modern Standard Arabic subspace during fine-tuning improves generation across 25 Arabic dialects by up to +4.9 chrF++, evidence that subspace dominance by a high-resource variety restrict...

  7. VN-MTEB: Vietnamese Massive Text Embedding Benchmark

    cs.CL 2025-07 conditional novelty 6.0 of 10

    VN-MTEB is a new 41-dataset Vietnamese benchmark for text embeddings, built by machine-translating MTEB datasets with embedding-based and LLM-based quality filters.

  8. MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 11-task Maltese benchmark shows that 55 large language models lag behind small fine-tuned models, with prior Maltese exposure the strongest predictor.

  9. Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Multilingual Statement-Tuning gives encoder-only models zero-shot cross-lingual classification, matching or beating multilingual LLMs up to 72B on three of four benchmarks.

  10. Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuned mT5-large outperforms prompt-based LLMs on Czech anaphora resolution (88% vs 74.5% accuracy) on a new dataset derived from the Prague Dependency Treebank.

  11. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

    cs.CL 2025-05 reject novelty 5.0 of 10

    A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.

  12. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

  13. Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A comprehensive survey of cross-lingual aspect-based sentiment analysis that catalogs tasks, datasets, modeling paradigms, and cross-lingual transfer techniques, and identifies research gaps.

  14. FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation

    cs.CL 2025-05 reject novelty 3.0 of 10

    FuxiMT combines a frozen BLOOMz model with sparse mixture-of-experts layers, Chinese-first pretraining, and curriculum learning to translate into Chinese from 65 languages, with claimed low-resource gains that the pap...

Pith tools