Pith. sign in

REVIEW 18 cited by

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07827 v1 pith:NWGG2EIG submitted 2024-02-12 cs.CL

classification cs.CL
keywords languagesmodellanguagemultilingualtasksbreakthroughsbroadenevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively multilingual generative language model that follows instructions in 101 languages of which over 50% are considered as lower-resourced. Aya outperforms mT0 and BLOOMZ on the majority of tasks while covering double the number of languages. We introduce extensive new evaluation suites that broaden the state-of-art for multilingual eval across 99 languages -- including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance. Furthermore, we conduct detailed investigations on the optimal finetuning mixture composition, data pruning, as well as the toxicity, bias, and safety of our models. We open-source our instruction datasets and our model at https://hf.co/CohereForAI/aya-101

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

    cs.CL 2025-10 conditional novelty 7.0 of 10

    ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.

  2. In-Place Tokenizer Expansion for Pre-trained LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Continuing a model's own BPE merges and training only new embedding rows preserves quality while cutting token counts 2.4–4× for previously under-tokenized languages.

  3. Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations

    math.ST 2025-09 conditional novelty 6.0 of 10

    A new hypothesis test for binary LLM responses treats semantically equivalent query perturbations as an unknown null set and gives asymptotic validity and consistency guarantees under a uniformity assumption.

  4. Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms

    math.NT 2025-08 unverdicted novelty 6.0 of 10

    The paper proves finiteness and structural results for preperiodic points of algebraic morphisms, including Burnside-type and Northcott-type theorems.

  5. Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Marco-Bench-MIF localizes IFEval into 30 languages with cultural adaptation and finds large resource gaps and scale effects in multilingual instruction following.

  6. Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An empirical study showing that smaller monolingual language models are more vulnerable to adversarial attacks than larger multilingual models across 70 languages, though multilinguality alone does not guarantee security.

  7. When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Hedged sampling, checklist-based one-pass selection (CHOPS), and cross-lingual MBR (X-MBR) improve multilingual LLM output quality when scaling from one to five samples.

  8. The Emergence of Abstract Thought in Large Language Models Beyond Any Language

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Across 20 open LLMs, shared multilingual neurons grow in number and per-neuron importance over release generations, which the authors interpret as evidence of language-agnostic abstract thought and use to guide neuron...

  9. CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    CC-Tuning fuses English feed-forward activations into non-English inputs during multilingual supervised fine-tuning, using a trainable Decision Maker and a least-squares Transform Matrix to simulate the connection at ...

  10. Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A Thai 70B model trained with an SFT-plus-DARE-merge recipe matches DeepSeek R1 on reasoning benchmarks while retaining most Thai language quality.

  11. Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis

    cs.CL 2025-07 reject novelty 5.0 of 10

    The authors release eight Bengali benchmarks translated from English and report that models with more fragmented Bengali tokenization tend to score lower.

  12. Text2Cypher Across Languages: Evaluating and Finetuning LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new multilingual Text2Cypher benchmark shows LLMs rank English highest, Spanish next, and Turkish lowest, and multilingual finetuning narrows the language gap more than English-only finetuning.

  13. Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components

    cs.IR 2025-06 conditional novelty 5.0 of 10

    For Arabic retrieval-augmented generation, sentence-aware chunking, BGE-M3 and Multilingual-E5-large embeddings, bge-reranker-v2-m3, and Aya-8B yield the highest RAGAS scores across six Arabic datasets.

  14. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

    cs.CL 2025-05 reject novelty 5.0 of 10

    A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.

  15. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

  16. CSIRO-LT at SemEval-2025 Task 11: Adapting LLMs for Emotion Recognition for Multiple Languages

    cs.CL 2025-08 conditional novelty 4.0 of 10

    On the SemEval-2025 multilingual emotion recognition task, directly fine-tuning a multilingual LLM with LoRA separately for each language beat zero-shot, few-shot, and English-bridged adaptation in the authors' experiments.

  17. Scaling Decentralized Learning with FLock

    cs.LG 2025-07 reject novelty 4.0 of 10

    FLock claims the first secure decentralized fine-tuning of a 70B-class LLM, but the experiments omit the validator mechanism and compare against weak baselines.

  18. In-Domain African Languages Translation Using LLMs and Multi-armed Bandits

    cs.CL 2025-05 reject novelty 4.0 of 10

    Bandit-based model selection matches or slightly improves on the best single NMT system for in-domain English-to-African translation, but the claimed high-confidence statistical support is absent.

Pith tools