Pith. sign in

REVIEW 3 cited by

Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.04685 v1 pith:PPL2RADB submitted 2024-05-07 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords turkishlanguagedatalanguagesllmsmodelfine-tuninglow-resource
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) are becoming crucial across various fields, emphasizing the urgency for high-quality models in underrepresented languages. This study explores the unique challenges faced by low-resource languages, such as data scarcity, model selection, evaluation, and computational limitations, with a special focus on Turkish. We conduct an in-depth analysis to evaluate the impact of training strategies, model choices, and data availability on the performance of LLMs designed for underrepresented languages. Our approach includes two methodologies: (i) adapting existing LLMs originally pretrained in English to understand Turkish, and (ii) developing a model from the ground up using Turkish pretraining data, both supplemented with supervised fine-tuning on a novel Turkish instruction-tuning dataset aimed at enhancing reasoning capabilities. The relative performance of these methods is evaluated through the creation of a new leaderboard for Turkish LLMs, featuring benchmarks that assess different reasoning and knowledge skills. Furthermore, we conducted experiments on data and model scaling, both during pretraining and fine-tuning, simultaneously emphasizing the capacity for knowledge transfer across languages and addressing the challenges of catastrophic forgetting encountered during fine-tuning on a different language. Our goal is to offer a detailed guide for advancing the LLM framework in low-resource linguistic contexts, thereby making natural language processing (NLP) benefits more globally accessible.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Make Satire Boring Again: Reducing Stylistic Bias of Satirical Corpus by Utilizing Generative LLMs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A generative-LLM debiasing pipeline that rewrites satirical Turkish news into plainer language improves cross-lingual and cross-domain satire and irony detection for masked language models, while having limited effect...

  2. Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis

    cs.CL 2025-07 reject novelty 5.0 of 10

    The authors release eight Bengali benchmarks translated from English and report that models with more fragmented Bengali tokenization tend to score lower.

  3. Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A systematic review of 54 studies finds that generative language modelling for low-resource languages relies mostly on transformer models, covers only a small set of languages, and lacks consistent evaluation.

Pith tools