REVIEW 3 cited by
DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Language technologies should be judged on their usefulness in real-world use cases. An often overlooked aspect in natural language processing (NLP) research and evaluation is language variation in the form of non-standard dialects or language varieties (hereafter, varieties). Most NLP benchmarks are limited to standard language varieties. To fill this gap, we propose DIALECTBENCH, the first-ever large-scale benchmark for NLP on varieties, which aggregates an extensive set of task-varied variety datasets (10 text-level tasks covering 281 varieties). This allows for a comprehensive evaluation of NLP system performance on different language varieties. We provide substantial evidence of performance disparities between standard and non-standard language varieties, and we also identify language clusters with large performance divergence across tasks. We believe DIALECTBENCH provides a comprehensive view of the current state of NLP for language varieties and one step towards advancing it further. Code/data: https://github.com/ffaisal93/DialectBench
Forward citations
Cited by 3 Pith papers
-
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
LLMs predict sentiment worse on Taiwan Mandarin than Mainland Mandarin reviews, using a new contextually paired dataset from Booking.com.
-
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic
LLMs understand dialectal Arabic better than they generate it, and current post-training appears to bias them toward Modern Standard Arabic.
-
Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek
Benchmarking Llama-70b against GPT-4o mini on Greek shows task-specific strengths, but the contamination-probe and legal-clustering claims need stronger baselines.
Discussion (0). Continue with ORCID to comment.