REVIEW 14 cited by
Aya 23: Open Weight Releases to Further Multilingual Progress
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This technical report introduces Aya 23, a family of multilingual language models. Aya 23 builds on the recent release of the Aya model (\"Ust\"un et al., 2024), focusing on pairing a highly performant pre-trained model with the recently released Aya collection (Singh et al., 2024). The result is a powerful multilingual large language model serving 23 languages, expanding state-of-art language modeling capabilities to approximately half of the world's population. The Aya model covered 101 languages whereas Aya 23 is an experiment in depth vs breadth, exploring the impact of allocating more capacity to fewer languages that are included during pre-training. Aya 23 outperforms both previous massively multilingual models like Aya 101 for the languages it covers, as well as widely used models like Gemma, Mistral and Mixtral on an extensive range of discriminative and generative tasks. We release the open weights for both the 8B and 35B models as part of our continued commitment for expanding access to multilingual progress.
Forward citations
Cited by 14 Pith papers
-
CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?
Hallucination signals from LLM internals transfer across English–Arabic and Arabic domains for most models, depending on class separability and feature-space language alignment.
-
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.
-
Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila
A 1,834-question Indonesian dilemma benchmark shows leading LLMs match Indonesian human choices only about half the time and struggle most on religion and unity.
-
Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs
Incidental multilingualism from uneven web training makes LLMs unequal, brittle, and opaque across languages.
-
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
In multilingual retrieval-augmented generation, models cite English evidence more accurately than translated evidence, and this language preference can outweigh document relevance.
-
When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
Projecting away the estimated Modern Standard Arabic subspace during fine-tuning improves generation across 25 Arabic dialects by up to +4.9 chrF++, evidence that subspace dominance by a high-resource variety restrict...
-
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
VN-MTEB is a new 41-dataset Vietnamese benchmark for text embeddings, built by machine-translating MTEB datasets with embedding-based and LLM-based quality filters.
-
MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP
A new 11-task Maltese benchmark shows that 55 large language models lag behind small fine-tuned models, with prior Maltese exposure the strongest predictor.
-
Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models
Multilingual Statement-Tuning gives encoder-only models zero-shot cross-lingual classification, matching or beating multilingual LLMs up to 72B on three of four benchmarks.
-
Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution
Fine-tuned mT5-large outperforms prompt-based LLMs on Czech anaphora resolution (88% vs 74.5% accuracy) on a new dataset derived from the Prague Dependency Treebank.
-
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model
A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.
-
Salamandra Technical Report
Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.
-
Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
A comprehensive survey of cross-lingual aspect-based sentiment analysis that catalogs tasks, datasets, modeling paradigms, and cross-lingual transfer techniques, and identifies research gaps.
-
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation
FuxiMT combines a frozen BLOOMz model with sparse mixture-of-experts layers, Chinese-first pretraining, and curriculum learning to translate into Chinese from 65 languages, with claimed low-resource gains that the pap...
Discussion (0). Continue with ORCID to comment.