REVIEW 11 cited by
Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Jais and Jais-chat, new state-of-the-art Arabic-centric foundation and instruction-tuned open generative large language models (LLMs). The models are based on the GPT-3 decoder-only architecture and are pretrained on a mixture of Arabic and English texts, including source code in various programming languages. With 13 billion parameters, they demonstrate better knowledge and reasoning capabilities in Arabic than any existing open Arabic and multilingual models by a sizable margin, based on extensive evaluation. Moreover, the models are competitive in English compared to English-centric open models of similar size, despite being trained on much less English data. We provide a detailed description of the training, the tuning, the safety alignment, and the evaluation of the models. We release two open versions of the model -- the foundation Jais model, and an instruction-tuned Jais-chat variant -- with the aim of promoting research on Arabic LLMs. Available at https://huggingface.co/inception-mbzuai/jais-13b-chat
Forward citations
Cited by 11 Pith papers
-
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
ALiBi's linearly growing positional bias underflows floating-point attention in long contexts, zeroing out distant attention weights, with measurable but task-dependent effects on retrieval.
-
Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation
Arabizi spelling varies systematically across five Arabic dialects, and speakers can often recognize their own dialect's Arabizi, but the recognition result is partly confounded by authors judging their own transcriptions.
-
ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification
ArabicDialectSafety, a 25,071-prompt, six-dialect Arabic safety benchmark, shows fine-tuned MARBERTv2 reaches 0.95 binary and 0.90 granular Macro-F1, outperforming prompted frontier LLMs.
-
Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning
OG-MAR improves LLM prediction of survey responses by retrieving demographically matched World Values Survey profiles and ontology-derived value relations, then aggregating persona-agent answers with a judge agent.
-
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
Llama-GENBA-10B is a 10B-parameter trilingual model that reports top Bavarian scores among sub-10B models on a machine-translated benchmark the authors built.
-
BALSAM: A Platform for Benchmarking Arabic Large Language Models
BALSAM is a new Arabic LLM benchmark with blind test sets, and the paper argues that LLM-based judging should replace n-gram and embedding metrics for scoring it.
-
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
AraTable is the first Arabic tabular QA benchmark; its experiments show LLMs are much weaker at reasoning over Arabic tables than at direct lookup.
-
SpeLLM: Character-Level Multi-Head Decoding
SpeLLM converts a standard token-based LLM into a character-spelling model with multiple parallel output heads, achieving competitive downstream performance with a 5.1% average decoding speedup.
-
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
A fully open 518M Arabic-specialized LLM, built by vocabulary injection and standard post-training on Qwen2.5-0.5B, beats same-class multilingual baselines and ships at 398 MB quantized.
-
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
AraHalluEval introduces a 12-indicator Arabic hallucination taxonomy and finds factual errors dominate, with Allam competitive against reasoning models.
-
Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks
Across Arabic, English, and Kannada benchmarks, 4-bit and 8-bit quantization preserves most accuracy while aggressive pruning degrades larger multilingual models more than smaller ones.
Discussion (0). Sign in to comment.